ci: add advisory macOS Apple Container integration runner

Adds an `integration-macos` job that runs the integration suite against
BOT_BOTTLE_BACKEND=macos-container on a self-hosted host-mode macOS runner
(label `macos`), plus the PRD and provisioning docs.

The job is advisory (push-to-main + workflow_dispatch only, never PRs, not
in coverage.needs) since it targets a single non-redundant laptop. It
preflights `container`/`backend status` so a misprovisioned runner fails
loudly, serializes on a concurrency group, and tears down the
`bot-bottle-mac-infra` singleton on exit (#425).

Also relaxes the TestSandboxEscape CI skip guard: it skipped every backend
but firecracker under GITEA_ACTIONS, which would also skip on a host-mode
macOS runner. The guard's real target is the containerized act_runner, so
it now allows both host-mode backends (firecracker, macos-container)
through — otherwise the macOS job would go green while skipping the one
end-to-end test that proves the backend launches.

Closes #426

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-25 18:47:02 -04:00
parent 0e70d26af4
commit ce1d16dd5b
5 changed files with 231 additions and 4 deletions
+63
View File
@@ -212,6 +212,69 @@ jobs:
name: firecracker-inputs
path: /var/cache/bot-bottle-fc/dropbear
# Integration tests against the macOS Apple Container backend. Runs on a
# self-hosted macOS runner (label `macos`) registered in HOST mode — Apple
# Container needs the host `container` CLI + virtualization framework and
# cannot run inside a Linux container, so this cannot reuse the KVM runner.
#
# Advisory only: push-to-main and workflow_dispatch, never pull_request. A
# single non-redundant laptop that sleeps/roams must not be able to block a
# PR merge, so this job is deliberately NOT in the `coverage` job's `needs`
# and its coverage never feeds the diff-coverage gate. Not gating on PRs also
# means no fork PR ever executes on the host-mode runner.
#
# The infra container is a singleton (`bot-bottle-mac-infra`); the
# `concurrency` group serializes runs so two never collide on it (#425), and
# the always-run teardown removes it so a crashed run can't wedge the next.
#
# Runner prerequisites (provision once; see README "macOS Apple Container"):
# the `container` CLI on PATH with `container system status` running, and a
# Python >=3.11 with `coverage` importable on the launchd service PATH.
integration-macos:
runs-on: [self-hosted, macos]
if: >-
github.event_name == 'push' ||
github.event_name == 'workflow_dispatch'
concurrency:
group: integration-macos-infra
cancel-in-progress: false
steps:
- name: Checkout
uses: actions/checkout@v4
# Fail loudly if the backend this job promises isn't actually usable,
# rather than letting every test silently `unittest.skip` and the job go
# green on zero coverage. `backend status` exits non-zero (and prints the
# per-check summary) when the `container` CLI or its system service is
# missing — the same readiness check the skip guards gate on.
- name: Preflight — Apple Container backend is ready
run: |
command -v container >/dev/null || {
echo "container CLI not on PATH — provision the runner (README: macOS Apple Container)"; exit 1; }
container system status || {
echo "container system service not running — run 'container system start'"; exit 1; }
python3 cli.py backend status --backend=macos-container
# `coverage` comes from the runner's provisioned Python (no pip install
# into the host interpreter). Advisory job: report coverage in-line for
# visibility but don't upload — it never feeds the combined gate.
- name: Run integration tests (macos-container) with coverage
env:
BOT_BOTTLE_BACKEND: macos-container
COVERAGE_FILE: ${{ github.workspace }}/.coverage.macos
run: python3 -m coverage run -m unittest discover -t . -s tests/integration -v
- name: Report macos coverage
env:
COVERAGE_FILE: ${{ github.workspace }}/.coverage.macos
run: python3 -m coverage report -m
# Remove the singleton infra container so a crashed or cancelled run
# cannot leave `bot-bottle-mac-infra` wedged for the next job.
- name: Teardown infra singleton
if: always()
run: python3 -c 'from bot_bottle.backend.macos_container.infra import MacosInfraService; MacosInfraService().stop()'
# Combined coverage gate: aggregates .coverage.* artifacts uploaded by each
# test job, then runs the diff-coverage gate (new/changed lines >= 90%).
#
+2
View File
@@ -75,6 +75,8 @@ On compatible macOS hosts, the default backend requires Apple's `container` CLI
Use `BOT_BOTTLE_BACKEND=docker ./cli.py start <agent>` on hosts where neither Apple Container nor KVM is available and Docker is the desired backend.
> **CI (macOS Apple Container):** the `integration-macos` job (`.gitea/workflows/test.yml`) runs the integration suite against `BOT_BOTTLE_BACKEND=macos-container` on a self-hosted macOS runner labelled `macos`, because Apple Container needs the host virtualization framework and cannot run in a Linux container (so it can't reuse the `kvm` runner). Provision an Apple Silicon host with the `container` CLI on `PATH` and `container system status` running, then register the runner in **host mode** (not docker mode) with the `macos` label — `brew install gitea-runner` (the `act_runner` rename). Give it a Python ≥ 3.11 with `coverage` importable on the launchd service's `PATH` (a launchd service doesn't inherit your shell profile, so pin `node` and the Python env explicitly). The job is **advisory** — push-to-main and `workflow_dispatch` only, never a required PR check — since a single laptop that sleeps/roams must not block merges; its coverage doesn't feed the gate. The infra container is a singleton (`bot-bottle-mac-infra`), so keep runner concurrency at 1.
### Containers inside a bottle
A bottle may set `nested_containers: true`. On the macOS backend this starts a
+12 -1
View File
@@ -2,11 +2,22 @@
The test workflow lives at [`.gitea/workflows/test.yml`](../.gitea/workflows/test.yml).
It runs the unit suite plus one integration job per backend
(`integration-docker`, `integration-firecracker`) on:
(`integration-docker`, `integration-firecracker`, `integration-macos`) on:
- every push to a branch with an open pull request, and
- every push to `main`.
`integration-macos` is the exception: it is **advisory**, running only on
push-to-`main` and `workflow_dispatch`, never on pull requests. It targets the
Apple Container backend on a self-hosted macOS runner (label `macos`,
registered in host mode — Apple Container can't run in a Linux container, so it
can't reuse the `kvm` runner). A single non-redundant laptop must not be able
to block a PR merge, so the job stays out of the `coverage` job's `needs` and
its coverage never feeds the diff-coverage gate. Because the infra container is
a singleton (`bot-bottle-mac-infra`), the job declares a `concurrency` group
and tears the container down on exit; keep runner concurrency at 1. See the
README "macOS Apple Container" CI note for runner provisioning.
Each integration job selects its backend via `BOT_BOTTLE_BACKEND` and
runs a **preflight** (`./cli.py backend status --backend=<name>`) that
prints a clear per-check readiness summary and fails the job when the
@@ -0,0 +1,140 @@
# PRD prd-new: macOS (Apple Container) CI runner
- **Status:** Draft
- **Author:** Claude
- **Created:** 2026-07-25
- **Issue:** #426
## Summary
CI has no runner for the `macos-container` (Apple Container) backend.
`.gitea/workflows/test.yml` exercises Docker (`ubuntu-latest`) and
Firecracker (self-hosted `kvm`) but never the macOS backend. This PRD adds a
self-hosted macOS runner (label `macos`) and an advisory `integration-macos`
job that runs the integration suite against `BOT_BOTTLE_BACKEND=macos-container`,
so the backend that is the default on macOS stops shipping unexercised.
## Problem
The gap is not theoretical. `5ad3449` moved `bot_bottle` from flat files under
`/app` into a pip-installed package but left init scripts spawning the
supervisor as `python3 /app/gateway_init.py`, which no longer exists. Both the
Firecracker and macOS backends carried the identical bug:
- **firecracker** — caught and fixed in `127ba49` because the KVM runner
(added in `c193b04`, PR #349) runs that backend's integration suite.
- **macos-container** — survived on `main` and only surfaced when a human ran
`bot-bottle start` by hand.
The failure mode is expensive to debug: the supervisor never starts, so
mitmdump never generates its CA, and launch dies downstream with
`GatewayError: gateway CA not available`, which points at TLS rather than at
the supervisor. Unit tests did not help — `test_macos_infra` asserted the
substring `"gateway_init.py"`, which the *broken* path satisfies. (That
specific assertion has since been tightened to the module form
`bot_bottle.gateway_init`, matching its Firecracker twin, so the exact
regression is now covered on `ubuntu-latest`. What remains missing is the
end-to-end runner that would catch the *next* macOS-only launch regression.)
PR #470 (#414) already made the integration suite backend-agnostic:
`skip_unless_selected_backend_available()` gates on the *selected* backend's
own `is_backend_ready()` rather than `docker_available()`, and each
integration job runs `./cli.py backend status --backend=<name>` as a preflight
that fails loudly when the backend is missing. That is the machinery this job
plugs into; this PRD supplies the runner and the job.
## Goals / Success criteria
- A macOS runner is registered and picks up jobs by the `macos` label.
- An `integration-macos` job runs the integration suite against
`BOT_BOTTLE_BACKEND=macos-container`.
- The job **fails, not skips**, when the backend is unavailable on the runner
(via the `backend status` preflight).
- Reverting the `macos_container/infra.py` supervisor fix makes the job fail:
the broken supervisor path throws `GatewayError` at bottle launch, which is
`TestSandboxEscape.setUpClass`, failing the whole class before any individual
attack runs.
## Non-goals
- Making `integration-macos` a **required** PR check. It runs on push-to-main
and `workflow_dispatch` only. A single non-redundant laptop that sleeps and
roams must never be able to block a PR merge, and it is deliberately kept out
of the `coverage` job's `needs` so the diff-coverage gate never depends on it.
- Multi-machine or hosted macOS runners. Apple Container needs the host
virtualization framework, so the runner must be a physical/VM macOS host on
Apple Silicon — it cannot reuse the KVM runner or run in a Linux container.
- Coverage aggregation from the macOS job into the combined gate (would couple
the gate to the laptop).
## Design
### Runner (operational, provisioned once)
- Apple Silicon macOS host with Apple's `container` CLI installed and
`container system status` reporting `running`.
- Install the runner: `brew install gitea-runner` (the `act_runner` rename),
registered in **host mode** with label `macos` — not docker mode, because
Apple Container needs the host `container` CLI and virtualization framework,
not a nested container.
- A Python ≥ 3.11 with `coverage` importable on the runner's `PATH`. Because a
launchd service does not inherit an interactive shell's `PATH`, pin `node`
(for the JS `actions/*`) and the Python env explicitly in the service
environment rather than relying on `nvm`/shell profile.
- Concurrency 1. The infra container is a singleton (`bot-bottle-mac-infra`),
so two simultaneous runs on one host collide (#425). The job also declares a
`concurrency` group as belt-and-suspenders and tears the singleton down after
each run.
### `integration-macos` job
Modeled on `integration-firecracker`:
- `runs-on: [self-hosted, macos]`.
- `if:` push-to-main OR `workflow_dispatch` only (advisory; never PRs, so no
fork-PR exposure and no merge-blocking).
- `concurrency: { group: integration-macos-infra, cancel-in-progress: false }`
to serialize runs against the singleton.
- **Preflight** — `command -v container`, `container system status`, then
`./cli.py backend status --backend=macos-container`; any failure exits
non-zero so a misprovisioned runner fails loudly instead of silently
skipping.
- Run the integration suite under coverage with
`BOT_BOTTLE_BACKEND=macos-container` and print a `coverage report -m` for
visibility (no upload, not in the gate).
- **Teardown** (`if: always()`) — `MacosInfraService().stop()` removes the
singleton so a crashed run cannot wedge the next one.
### The `test_sandbox_escape` CI guard (the trap #470 left)
`TestSandboxEscape` is the only backend-agnostic integration test that boots a
real bottle, so it is the one that would catch a macOS launch regression. It
still carries a second guard that skips under `GITEA_ACTIONS` for every backend
except `firecracker`:
```python
@unittest.skipIf(
os.environ.get("GITEA_ACTIONS") == "true"
and os.environ.get("BOT_BOTTLE_BACKEND") != "firecracker",
...,
)
```
The skip exists because the *containerized* `act_runner` (docker on
`ubuntu-latest`) can't see a host bind mount and hides sibling-gateway network
topology. Those constraints do not apply to a **host-mode** runner — neither
the KVM host runner nor a macOS host runner is containerized. This PRD relaxes
the guard to allow both host-mode backends (`firecracker`, `macos-container`)
through while still skipping on the containerized Docker job. Without this
change the macOS job would run green while skipping the exact test that proves
the backend launches — the very false-green this issue is about.
## Open questions
- Some sandbox-escape attacks depend on `bottle.git` / git-gate, which is
intentionally deferred on the macOS backend until a safe Apple Container
key-delivery path exists. First runner bring-up may reveal individual attacks
that need a `skipUnless` guard for `macos-container`. This does not affect the
infra-regression signal (that fails at `setUpClass`/launch, before any
attack), but it is expected first-run shakeout and is left to the bring-up
PR that actually has the runner in hand.
+14 -3
View File
@@ -67,14 +67,25 @@ _DUMMY_HOST_KEY = (
)
# Backends whose CI runner is HOST-mode (self-hosted), so the test process
# and the backend share a host. The containerized act_runner (docker on
# ubuntu-latest) is the one that can't see the host bind mount egress_tls_init
# uses and hides sibling-gateway network topology; host-mode runners
# (firecracker/KVM, macos-container) don't have those constraints, so the test
# runs there. Keep this in sync with the `runs-on` labels in
# .gitea/workflows/test.yml.
_HOST_MODE_CI_BACKENDS = frozenset({"firecracker", "macos-container"})
@skip_unless_selected_backend_available()
@unittest.skipIf(
os.environ.get("GITEA_ACTIONS") == "true"
and os.environ.get("BOT_BOTTLE_BACKEND") != "firecracker",
"skipped under act_runner unless BOT_BOTTLE_BACKEND=firecracker: "
and os.environ.get("BOT_BOTTLE_BACKEND") not in _HOST_MODE_CI_BACKENDS,
"skipped under the containerized act_runner (docker on ubuntu-latest): "
"egress_tls_init uses a host bind mount the runner container can't "
"see, and the network topology hides sibling-gateway visibility — "
"these constraints don't apply on the self-hosted KVM runner",
"these constraints don't apply on the self-hosted host-mode runners "
"(firecracker/KVM, macos-container)",
)
class TestSandboxEscape(unittest.TestCase):
"""End-to-end attacks against a real bottle. The bottle stays