ci: add advisory macOS Apple Container integration runner

Adds an `integration-macos` job that runs the integration suite against
BOT_BOTTLE_BACKEND=macos-container on a self-hosted host-mode macOS runner
(label `macos`), plus the PRD and provisioning docs.

The job is advisory (push-to-main + workflow_dispatch only, never PRs, not
in coverage.needs) since it targets a single non-redundant laptop. It
preflights `container`/`backend status` so a misprovisioned runner fails
loudly, serializes on a concurrency group, and tears down the
`bot-bottle-mac-infra` singleton on exit (#425).

Also relaxes the TestSandboxEscape CI skip guard: it skipped every backend
but firecracker under GITEA_ACTIONS, which would also skip on a host-mode
macOS runner. The guard's real target is the containerized act_runner, so
it now allows both host-mode backends (firecracker, macos-container)
through — otherwise the macOS job would go green while skipping the one
end-to-end test that proves the backend launches.

Closes #426

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-25 18:47:02 -04:00
parent 2644759b0d
commit 238f5f7614
5 changed files with 231 additions and 4 deletions
@@ -0,0 +1,140 @@
# PRD prd-new: macOS (Apple Container) CI runner
- **Status:** Draft
- **Author:** Claude
- **Created:** 2026-07-25
- **Issue:** #426
## Summary
CI has no runner for the `macos-container` (Apple Container) backend.
`.gitea/workflows/test.yml` exercises Docker (`ubuntu-latest`) and
Firecracker (self-hosted `kvm`) but never the macOS backend. This PRD adds a
self-hosted macOS runner (label `macos`) and an advisory `integration-macos`
job that runs the integration suite against `BOT_BOTTLE_BACKEND=macos-container`,
so the backend that is the default on macOS stops shipping unexercised.
## Problem
The gap is not theoretical. `5ad3449` moved `bot_bottle` from flat files under
`/app` into a pip-installed package but left init scripts spawning the
supervisor as `python3 /app/gateway_init.py`, which no longer exists. Both the
Firecracker and macOS backends carried the identical bug:
- **firecracker** — caught and fixed in `127ba49` because the KVM runner
(added in `c193b04`, PR #349) runs that backend's integration suite.
- **macos-container** — survived on `main` and only surfaced when a human ran
`bot-bottle start` by hand.
The failure mode is expensive to debug: the supervisor never starts, so
mitmdump never generates its CA, and launch dies downstream with
`GatewayError: gateway CA not available`, which points at TLS rather than at
the supervisor. Unit tests did not help — `test_macos_infra` asserted the
substring `"gateway_init.py"`, which the *broken* path satisfies. (That
specific assertion has since been tightened to the module form
`bot_bottle.gateway_init`, matching its Firecracker twin, so the exact
regression is now covered on `ubuntu-latest`. What remains missing is the
end-to-end runner that would catch the *next* macOS-only launch regression.)
PR #470 (#414) already made the integration suite backend-agnostic:
`skip_unless_selected_backend_available()` gates on the *selected* backend's
own `is_backend_ready()` rather than `docker_available()`, and each
integration job runs `./cli.py backend status --backend=<name>` as a preflight
that fails loudly when the backend is missing. That is the machinery this job
plugs into; this PRD supplies the runner and the job.
## Goals / Success criteria
- A macOS runner is registered and picks up jobs by the `macos` label.
- An `integration-macos` job runs the integration suite against
`BOT_BOTTLE_BACKEND=macos-container`.
- The job **fails, not skips**, when the backend is unavailable on the runner
(via the `backend status` preflight).
- Reverting the `macos_container/infra.py` supervisor fix makes the job fail:
the broken supervisor path throws `GatewayError` at bottle launch, which is
`TestSandboxEscape.setUpClass`, failing the whole class before any individual
attack runs.
## Non-goals
- Making `integration-macos` a **required** PR check. It runs on push-to-main
and `workflow_dispatch` only. A single non-redundant laptop that sleeps and
roams must never be able to block a PR merge, and it is deliberately kept out
of the `coverage` job's `needs` so the diff-coverage gate never depends on it.
- Multi-machine or hosted macOS runners. Apple Container needs the host
virtualization framework, so the runner must be a physical/VM macOS host on
Apple Silicon — it cannot reuse the KVM runner or run in a Linux container.
- Coverage aggregation from the macOS job into the combined gate (would couple
the gate to the laptop).
## Design
### Runner (operational, provisioned once)
- Apple Silicon macOS host with Apple's `container` CLI installed and
`container system status` reporting `running`.
- Install the runner: `brew install gitea-runner` (the `act_runner` rename),
registered in **host mode** with label `macos` — not docker mode, because
Apple Container needs the host `container` CLI and virtualization framework,
not a nested container.
- A Python ≥ 3.11 with `coverage` importable on the runner's `PATH`. Because a
launchd service does not inherit an interactive shell's `PATH`, pin `node`
(for the JS `actions/*`) and the Python env explicitly in the service
environment rather than relying on `nvm`/shell profile.
- Concurrency 1. The infra container is a singleton (`bot-bottle-mac-infra`),
so two simultaneous runs on one host collide (#425). The job also declares a
`concurrency` group as belt-and-suspenders and tears the singleton down after
each run.
### `integration-macos` job
Modeled on `integration-firecracker`:
- `runs-on: [self-hosted, macos]`.
- `if:` push-to-main OR `workflow_dispatch` only (advisory; never PRs, so no
fork-PR exposure and no merge-blocking).
- `concurrency: { group: integration-macos-infra, cancel-in-progress: false }`
to serialize runs against the singleton.
- **Preflight** — `command -v container`, `container system status`, then
`./cli.py backend status --backend=macos-container`; any failure exits
non-zero so a misprovisioned runner fails loudly instead of silently
skipping.
- Run the integration suite under coverage with
`BOT_BOTTLE_BACKEND=macos-container` and print a `coverage report -m` for
visibility (no upload, not in the gate).
- **Teardown** (`if: always()`) — `MacosInfraService().stop()` removes the
singleton so a crashed run cannot wedge the next one.
### The `test_sandbox_escape` CI guard (the trap #470 left)
`TestSandboxEscape` is the only backend-agnostic integration test that boots a
real bottle, so it is the one that would catch a macOS launch regression. It
still carries a second guard that skips under `GITEA_ACTIONS` for every backend
except `firecracker`:
```python
@unittest.skipIf(
os.environ.get("GITEA_ACTIONS") == "true"
and os.environ.get("BOT_BOTTLE_BACKEND") != "firecracker",
...,
)
```
The skip exists because the *containerized* `act_runner` (docker on
`ubuntu-latest`) can't see a host bind mount and hides sibling-gateway network
topology. Those constraints do not apply to a **host-mode** runner — neither
the KVM host runner nor a macOS host runner is containerized. This PRD relaxes
the guard to allow both host-mode backends (`firecracker`, `macos-container`)
through while still skipping on the containerized Docker job. Without this
change the macOS job would run green while skipping the exact test that proves
the backend launches — the very false-green this issue is about.
## Open questions
- Some sandbox-escape attacks depend on `bottle.git` / git-gate, which is
intentionally deferred on the macOS backend until a safe Apple Container
key-delivery path exists. First runner bring-up may reveal individual attacks
that need a `skipUnless` guard for `macos-container`. This does not affect the
infra-regression signal (that fails at `setUpClass`/launch, before any
attack), but it is expected first-run shakeout and is left to the bring-up
PR that actually has the runner in hand.