Files
bot-bottle/docs/prds/0077-macos-container-ci-runner.md
T
didericis cd9f023f3d
prd-number-check / require-numbered-prds (pull_request) Successful in 11s
tracker-policy-pr / check-pr (pull_request) Successful in 14s
test / unit (pull_request) Successful in 59s
test / integration-docker (pull_request) Successful in 1m6s
test / coverage (push) Successful in 21s
test / image-input-builds (push) Successful in 44s
test / image-input-builds (pull_request) Successful in 1m15s
test / unit (push) Successful in 57s
test / coverage (pull_request) Successful in 20s
lint / lint (push) Successful in 1m3s
Update Quality Badges / update-badges (push) Successful in 1m12s
test / integration-docker (push) Successful in 1m0s
docs: name bot-bottle, not ./cli.py, in every instruction
`doctor` told users to run `./cli.py backend setup`. cli.py is a four-line
wrapper at the repo root that calls bot_bottle.cli:main, and it does not ship —
[tool.setuptools.packages.find] includes only bot_bottle*, so anyone who
installed rather than cloned has no such file. Confirmed against a real
install: the user's tree contains exactly one executable, `bot-bottle`, and no
cli.py anywhere, while doctor recommended `./cli.py` three times. Invisible in
development, where ./cli.py works fine from a checkout, which is why only a
clean-install test surfaced it.

The two entry points are the same code, so the fix is to name the one that
always exists. 141 replacements across 45 files: the runtime messages that
caused this, plus README, docs, PRDs, research notes and test prose, so nothing
teaches the invocation a user cannot run.

scripts/demo.sh is deliberately untouched — it *executes* ./cli.py from a
checkout, where that is the correct and available path.

Verified end to end: a sandbox install now reports
"Run: bot-bottle backend setup --backend=firecracker", and bot-bottle is on
that user's PATH. Unit suite unchanged against the pre-existing baseline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEfZZhakx13bxTfXcZCoS5
2026-07-27 10:20:17 -04:00

6.9 KiB

PRD 0077: macOS (Apple Container) CI runner

  • Status: Active
  • Author: Claude
  • Created: 2026-07-25
  • Issue: #426

Summary

CI has no runner for the macos-container (Apple Container) backend. .gitea/workflows/test.yml exercises Docker (ubuntu-latest) and Firecracker (self-hosted kvm) but never the macOS backend. This PRD adds a self-hosted macOS runner (label macos) and an advisory integration-macos job that runs the integration suite against BOT_BOTTLE_BACKEND=macos-container, so the backend that is the default on macOS stops shipping unexercised.

Problem

The gap is not theoretical. 5ad3449 moved bot_bottle from flat files under /app into a pip-installed package but left init scripts spawning the supervisor as python3 /app/gateway_init.py, which no longer exists. Both the Firecracker and macOS backends carried the identical bug:

  • firecracker — caught and fixed in 127ba49 because the KVM runner (added in c193b04, PR #349) runs that backend's integration suite.
  • macos-container — survived on main and only surfaced when a human ran bot-bottle start by hand.

The failure mode is expensive to debug: the supervisor never starts, so mitmdump never generates its CA, and launch dies downstream with GatewayError: gateway CA not available, which points at TLS rather than at the supervisor. Unit tests did not help — test_macos_infra asserted the substring "gateway_init.py", which the broken path satisfies. (That specific assertion has since been tightened to the module form bot_bottle.gateway_init, matching its Firecracker twin, so the exact regression is now covered on ubuntu-latest. What remains missing is the end-to-end runner that would catch the next macOS-only launch regression.)

PR #470 (#414) already made the integration suite backend-agnostic: skip_unless_selected_backend_available() gates on the selected backend's own is_backend_ready() rather than docker_available(), and each integration job runs bot-bottle backend status --backend=<name> as a preflight that fails loudly when the backend is missing. That is the machinery this job plugs into; this PRD supplies the runner and the job.

Goals / Success criteria

  • A macOS runner is registered and picks up jobs by the macos label.
  • An integration-macos job runs the integration suite against BOT_BOTTLE_BACKEND=macos-container.
  • The job fails, not skips, when the backend is unavailable on the runner (via the backend status preflight).
  • Reverting the macos_container/infra.py supervisor fix makes the job fail: the broken supervisor path throws GatewayError at bottle launch, which is TestSandboxEscape.setUpClass, failing the whole class before any individual attack runs.

Non-goals

  • Making integration-macos a required PR check. It runs on workflow_dispatch (manual dispatch) only — never on push or PRs. A single non-redundant laptop that sleeps and roams must never be able to block a PR merge or churn unattended on every push to main, and it is deliberately kept out of the coverage job's needs so the diff-coverage gate never depends on it.
  • Multi-machine or hosted macOS runners. Apple Container needs the host virtualization framework, so the runner must be a physical/VM macOS host on Apple Silicon — it cannot reuse the KVM runner or run in a Linux container.
  • Coverage aggregation from the macOS job into the combined gate (would couple the gate to the laptop).

Design

Runner (operational, provisioned once)

  • Apple Silicon macOS host with Apple's container CLI installed and container system status reporting running.
  • Install the runner: brew install gitea-runner (the act_runner rename), registered in host mode with label macos — not docker mode, because Apple Container needs the host container CLI and virtualization framework, not a nested container.
  • A Python ≥ 3.11 with coverage importable on the runner's PATH. Because a launchd service does not inherit an interactive shell's PATH, pin node (for the JS actions/*) and the Python env explicitly in the service environment rather than relying on nvm/shell profile.
  • Concurrency 1. The infra container is a singleton (bot-bottle-mac-infra), so two simultaneous runs on one host collide (#425). The job also declares a concurrency group as belt-and-suspenders and tears the singleton down after each run.

integration-macos job

Modeled on integration-firecracker:

  • runs-on: [self-hosted, macos].
  • if: workflow_dispatch only (advisory, manual dispatch; never push or PRs, so no fork-PR exposure, no merge-blocking, and no unattended runs on push).
  • concurrency: { group: integration-macos-infra, cancel-in-progress: false } to serialize runs against the singleton.
  • Preflightcommand -v container, container system status, then bot-bottle backend status --backend=macos-container; any failure exits non-zero so a misprovisioned runner fails loudly instead of silently skipping.
  • Run the integration suite under coverage with BOT_BOTTLE_BACKEND=macos-container and print a coverage report -m for visibility (no upload, not in the gate).
  • Teardown (if: always()) — MacosInfraService().stop() removes the singleton so a crashed run cannot wedge the next one.

The test_sandbox_escape CI guard (the trap #470 left)

TestSandboxEscape is the only backend-agnostic integration test that boots a real bottle, so it is the one that would catch a macOS launch regression. It still carries a second guard that skips under GITEA_ACTIONS for every backend except firecracker:

@unittest.skipIf(
    os.environ.get("GITEA_ACTIONS") == "true"
    and os.environ.get("BOT_BOTTLE_BACKEND") != "firecracker",
    ...,
)

The skip exists because the containerized act_runner (docker on ubuntu-latest) can't see a host bind mount and hides sibling-gateway network topology. Those constraints do not apply to a host-mode runner — neither the KVM host runner nor a macOS host runner is containerized. This PRD relaxes the guard to allow both host-mode backends (firecracker, macos-container) through while still skipping on the containerized Docker job. Without this change the macOS job would run green while skipping the exact test that proves the backend launches — the very false-green this issue is about.

Open questions

  • None known that block the job. git-gate is fully implemented on the macOS backend (the gateway's consolidated git-http daemon plus dynamic key provisioning/revocation), so TestSandboxEscape attack 5 — secret exfil pushed through git-gate, rejected by the gitleaks hook before the upstream push — runs the same as on the other backends. Any genuinely macOS-specific test adjustment would surface at first runner bring-up, but none is anticipated from the current backend implementation.