The per-bottle egress-secret encryption (secret_store.py) was
unauthenticated CTR/XOR: decrypting with the WRONG ENV_VAR_SECRET
produced garbage that decrypt_value only rejected when it wasn't valid
UTF-8. For short token values that garbage is coincidentally valid UTF-8
~5% of the time, so `reprovision_from_secret` would occasionally "succeed"
with a wrong key and inject a garbage egress credential — and
test_reprovision_rejects_missing_rows_and_wrong_key failed ~5% of runs
(flaky CI, surfaced by this stack's unit job).
Switch to authenticated encrypt-then-MAC: append an HMAC-SHA256 tag over
`nonce || ciphertext`, keyed by a domain-separated MAC subkey derived from
the ENV_VAR_SECRET. decrypt_value verifies the tag (constant-time) before
returning any plaintext, so a wrong key or tampered ciphertext is rejected
deterministically. Blob format is now `nonce || ciphertext || tag`
(the stored rows are transient — re-written every launch — so no
migration is needed).
Pre-existing bug on main, unrelated to the transport work, but it blocks
this stack's CI. Tests: wrong key rejected 200/200; tampered ciphertext
rejected; round-trips unchanged. Deterministic now (was ~5% flaky).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Codex review on #496:
- **High — ambiguous delivery no longer orphans a launched bottle.** A
timeout / dropped response from the host controller is now the ambiguous
BrokerUnavailableError (distinct from the definite BrokerAuthError /
BrokerClientError). OrchestratorCore.launch_bottle keeps the registry
row on the ambiguous case instead of deregistering — deregistering would
orphan a running container with no record (reconcile reaps rows, never
containers). The row is left for reconcile to reap iff the bottle is not
actually live. Definite failures still roll back, so a real failure
leaves no orphan row.
- **Medium — the privileged endpoint bounds request bodies.** The host
server rejects an oversized Content-Length with 413 before reading it,
and sets a per-request socket timeout, so a caller that can merely reach
the socket (no signed token) can't exhaust memory or a handler thread.
Tests: ambiguous-keep vs definite-rollback in the launch path; the
BrokerUnavailableError/BrokerClientError split in BrokerClient; the 413
body cap + handler error paths (driven in-thread, since daemon request
threads lose coverage) plus a deterministic real-socket check that
declares an oversized Content-Length but sends a sliver (rejection on the
header, no unread-body reset race); and the __main__ entrypoint broker
selection. Diff-coverage 98%; pyright clean; pylint 9.8.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Chunk 1 of the host-control-server stack: close the PRD's **transport**
gap. Today LaunchBroker.submit(token) is an in-process method call from
OrchestratorCore; this makes it a real out-of-process service reached
over HTTP.
- host_server.py: the host control server. A pure dispatch() (POST
/broker verifies a signed token via the existing verify_request +
_launch/_teardown path, GET /health) wrapped by a thin http.server
adapter, mirroring orchestrator/server.py. Only the signed token
crosses the wire; provenance/schema failures are fail-closed 401s that
never touch the backend, a backend launch failure is a 502.
- broker_client.py: BrokerClient — a drop-in submit(token) that POSTs the
signed token to the host controller. A 401 re-raises as BrokerAuthError
so the launch path's rollback is identical local or remote.
- broker.py: SubmitBroker Protocol — the one method OrchestratorCore
depends on, satisfied by both LaunchBroker and BrokerClient, so the
core is unchanged (service.py annotation only).
- __main__.py: wire `--broker http` behind the shared-secret env var
(BOT_BOTTLE_BROKER_SECRET, hex) — a chunk-1 stopgap the durable
TrustDomain key (chunk 2, #476) replaces.
Tested: pure-dispatch cases, BrokerClient with HTTP mocked, and a
real-socket sign -> POST -> verify -> act round-trip (incl. fail-closed
forged token). pyright clean; pylint 9.86.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Addresses the third review round on PR #481.
- `bot-bottle doctor` now checks `is_backend_ready()` (a full backend
status() probe: daemon reachable, network pool present, KVM usable)
instead of the cheap PATH-only `is_backend_available()`. A host with a
stopped Docker daemon or half-configured Firecracker no longer reports
`ok: backend` / exit 0 when `start` can't actually work; each not-ready
backend prints its own diagnostics, and doctor passes only if at least
one backend is ready.
- `install.sh` resolves the pip `--user` scripts directory from the
interpreter (`sysconfig.get_path("scripts", get_preferred_scheme("user"))`)
instead of hardcoding `~/.local/bin`, which is wrong on a python.org
macOS interpreter (`~/Library/Python/<X.Y>/bin`). The PATH guidance now
prints the actual directory.
Tests: doctor tests mock `is_backend_ready` (the readiness contract) and
cover the not-ready → fail path; a new install-script test drives the
macOS `osx_framework_user` scheme and asserts it resolves a
non-~/.local/bin directory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Addresses the review on PR #481.
Self-contained wheel (review point 1): the gateway/infra/orchestrator
images build from a context that must hold bot_bottle/, pyproject.toml,
and the root-level Dockerfiles. Modules previously located these by
walking __file__ to the repo root, so an installed wheel (package in
site-packages, no repo root) passed `doctor` but failed `start`.
- Add bot_bottle/resources.py: build_root() returns the repo root in a
checkout (unchanged) or a staged copy from the wheel's bundled
_resources/ otherwise; dockerfile()/nix_netpool_module()/
netpool_script() derive from it.
- setup.py bundles the root Dockerfiles, nix module, netpool script, and
pyproject.toml into bot_bottle/_resources/ at build; MANIFEST.in ships
them in the sdist.
- Route every _REPO_ROOT/_REPO_DIR call site (docker/macos launch, macos
infra, firecracker infra_vm/infra_artifact/setup, orchestrator
lifecycle/gateway) through resources. Checkout behavior is unchanged.
install.sh prerequisites (review point 2): check for git when installing
a git+ spec, and — before the pip fallback — that pip is usable and the
interpreter isn't externally managed (PEP 668), pointing at pipx.
Tests: test_resources covers checkout + staged-wheel layouts;
test_wheel_install builds the wheel, installs it into an isolated venv,
and asserts `doctor` runs and build_root() yields a valid context.
Running `start` end-to-end still needs a Docker/KVM host (CI).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Give bot-bottle a real distribution path so new users can install
without cloning the repo:
- pyproject.toml: full project metadata, a `bot-bottle` console-script
entry point (bot_bottle.cli:main), and package-data for the runtime
assets (Dockerfiles, egress entrypoint, netpool defaults, macos init).
Still zero runtime pip dependencies.
- install.sh: POSIX, sudo-free, idempotent bootstrapper — checks Python
>= 3.11, creates ~/.bot-bottle/{agents,bottles,contrib}, installs via
pipx (pip --user fallback), then runs `bot-bottle doctor`.
- `bot-bottle doctor`: new store-free subcommand reporting Python
version, backend availability (reuses is_backend_available rather than
hardcoding Docker), and config-dir presence. Exits non-zero when a hard
prerequisite is unmet.
- PRD prd-new-install-script and unit tests for doctor, the packaging
contract, and the install script.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Remove the empty-key opt-out codex flagged: ControlPlaneProvisioning
no longer lets the orchestrator start without a signing key when a
backend declares an "isolated" topology. Host separation is not the
safety condition — the gateway (and any other caller) must reach the
control-plane listener for /resolve, so an open orchestrator would
still grant them full `cli`. The signing key is now mandatory for
every backend.
Drops the now-purposeless Topology/COLOCATED machinery (no backend
declared a non-default topology) and the contractual
test_orchestrator_key_allows_empty_when_isolated. Updates the PRD's
invariant 4 and lifecycle docstring accordingly. Docs/behaviour only
otherwise.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Address review on #482: drop the generic "credential boundary" framing in the
PRD and docstrings and talk about the specific services — the orchestrator holds
the control-plane key; the host controller (#468) gets a separate key the
orchestrator never holds, so the orchestrator can't mint the credentials it uses
to talk to the host controller that owns its lifecycle. Lead with that concrete
win. No code behaviour change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Hoist control-plane auth provisioning out of the per-backend launchers into
one shared contract, parameterized per trust domain (#476). Every blocking
finding in PR #471 was the same integration-bug class: each launcher
re-derived, by hand, how to generate the signing key, scope it to the
orchestrator, mint the gateway JWT, and keep the host key canonical.
Introduces `trust_domain.py`:
* `TrustDomain` — one credential boundary (host-canonical key file + role
set + env vars). `mint`/`verify` are scoped to the domain's roles, so a
future host-controller domain (#468) uses its own key/verifier/roles
rather than a `host` role on the control plane's frozenset (which the
orchestrator key could then forge).
* `ControlPlaneProvisioning` — the single seam answering the four
invariants: host-canonical key, split key-vs-token credential, CLI token
valid across co-running backends, and fail-closed (no open mode) for any
co-located topology.
* `Topology` — the backend declares what it is; the default is co-located +
fail-closed, so a backend need not redeclare it.
The `Orchestrator` ABC gets `control_plane_key()` (fail-closed) and routes
`mint_gateway_token()` through the contract; docker/macOS/firecracker
orchestrators, the server (verify), and the host CLI client (mint cli) all go
through the domain instead of reading the host key directly. `orchestrator_auth`
gains an optional `roles=` arg (default unchanged) so a domain scopes its own
role set; `paths.host_signing_key(filename)` generalizes host_orchestrator_token.
Adds unit coverage for the domain boundary + provisioning invariants and a PRD
capturing the durable rationale. No change to the auth primitive's HMAC, the
plane split, or the server's documented open-mode fallback.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Closes#476
`_supervise_token_block` and `_await_token_response` called the synchronous
`PolicyResolver` directly from mitmproxy's async request handling. Those RPCs
funnel through `PolicyResolver._post_json`, which blocks on
`urllib.request.urlopen`. Because the poll loop runs for the entire
operator-approval window, a slow or unreachable orchestrator would repeatedly
freeze the proxy event loop — stalling every other bottle's traffic — despite
`_await_token_response`'s contract of not blocking the loop (#471 review,
review #443).
Dispatch both the propose and poll RPCs via `asyncio.to_thread` so the
blocking urllib call runs in a worker thread and the event loop keeps serving
other flows. Fail-closed semantics are unchanged: `PolicyResolveError` still
propagates through the await and is caught exactly as before.
Regression test records the thread each RPC executes on and asserts it is not
the loop thread; it fails against the pre-fix inline calls and passes with the
dispatch. Full egress-addon + policy-resolver + supervise-server suites green
(155 tests).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The previous fix interpolated the real per-bottle identity token into the
`warn()` message shown when `claude mcp add supervise` fails, so a
registration failure would leak the credential into host terminal output
and any collected launch logs (#476 review, PR #471 comment).
Render a `<bottle-identity-token>` placeholder instead. The recovery hint
still shows the required `--header x-bot-bottle-identity: …` shape, and the
operator substitutes the value from inside the bottle (it rides in the
agent's HTTPS_PROXY credentials) — so the token never reaches the host log.
The token truthiness still gates whether the header is shown at all.
Test updated to assert the header name and placeholder are present and the
raw token is absent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The manual registration command shown in the fallback warning when
`claude mcp add supervise` fails was missing the `--header
x-bot-bottle-identity: <token>` option. Since (source_ip,
identity_token) attribution is now mandatory on the supervise server,
the missing header caused every subsequent supervise request to fail
closed as unattributed — defeating supervision for the entire bottle.
Reuse the already-constructed `token` variable to build the same
`manual_header` string the automatic path uses; include it in the
printed command. Tests verify both the happy path and the fallback
include (or correctly omit) the header.
Addresses P2 finding in #471 (comment 7, didericis-codex 2026-07-25).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Remove symbols with no remaining consumer:
- `gateway_name` property on Docker/MacosInfraService — nothing uses it; the
launch flow reads `gateway().name` off the Gateway service. (+ its docker test.)
- firecracker `infra_vm` EGRESS_PORT / SUPERVISE_PORT / GIT_HTTP_PORT — dead
duplicates; launch.py uses supervisor.types / docker.egress / a local const.
- macOS `INFRA_DB_VOLUME` back-compat alias — unused since the DB-volume test
moved to `ORCHESTRATOR_DB_VOLUME`.
- docker `infra.py`: the unused `ORCHESTRATOR_SOURCE_HASH_LABEL` import and the
`__all__` re-exports of the orchestrator constants (consumers import them from
`docker.orchestrator`, not `docker.infra`).
- macOS `infra.py`: the unused `GatewayError` import + re-export.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a backend-neutral `InfraService` ABC (backend/infra_service.py) so all three
backends present the same composer contract: `orchestrator()` / `gateway()`
accessors, `ensure_running(*, startup_timeout) -> str` (the host control-plane
URL), `stop()`, plus concrete `url()` / `is_healthy()` that delegate to the
orchestrator. `DockerInfraService` and `MacosInfraService` now subclass it.
Give firecracker the missing class: `FirecrackerInfraService`
(backend/firecracker/infra.py) wraps the `infra_vm` substrate as the pair
coordinator (adopt-or-boot-both under the singleton flock). `infra_vm` keeps only
the plane-agnostic substrate; its `ensure_running` / `_adopt` / `InfraEndpoint`
move to the class, and the coordinator helpers it now calls cross-module go
public (`singleton_lock` / `expected_version` / `adoptable` /
`record_booted_version`).
Unify the return type on the way: `ensure_running` returns the URL string
everywhere (was `str` for docker, two different `InfraEndpoint`s for
macOS/firecracker), and those `InfraEndpoint`s are deleted — the services are the
source of truth (callers read `gateway().address()` / `orchestrator().url()`).
docker's `url` property becomes the inherited `url()`. backend.py /
consolidated_launch / image_builder + tests updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sweep for vestiges of the old combined-plane model and the pre-split shared
rootfs. Two are load-bearing, the rest are stale docs/comments:
- Bug: macOS `enumerate_active` only excluded the gateway container from the
agent list, so after the split the orchestrator container
(`bot-bottle-mac-orchestrator`, also `bot-bottle-`-prefixed) was enumerated as
a phantom agent. Exclude both infra containers; test covers it.
- Dead code: the gateway `bootstrap.py` still carried an `orchestrator` daemon
spec + `_OPT_IN_DAEMONS` + a signing-key/JWT env branch, all for the old
combined container where the gateway process could also run the control plane.
No backend ever requests it now — removed; the key-stripping stays as
defense-in-depth.
Stale-comment reframes: "the/single infra container" -> the orchestrator +
gateway pair (or the specific plane); "shared rootfs / bb_role init / one
published rootfs" -> the per-plane rootfs + `role_init`; the deleted
Dockerfile.infra references in Dockerfile.orchestrator/.gateway; and the macOS
"one infra container ... same address" docstring + its now-false
share-one-address test (the planes are distinct containers with distinct
addresses).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Each infra VM now boots its own rootfs instead of a shared combined image, so
the exposed gateway VM no longer carries buildah + the control-plane code it
never runs (a fatter, less-isolated exposed surface — the opposite of the plane
split's intent). The combined image bought only "one artifact"; each VM already
kept a full copy of the shared rootfs, so nothing was saved at boot.
* orchestrator rootfs — control plane + buildah (Dockerfile.orchestrator.fc,
FROM orchestrator); the slim gateway rootfs boots bot-bottle-gateway:latest
directly. Dockerfile.infra.fc (the combined image) is deleted.
* `_infra_init` splits into `_orchestrator_init` / `_gateway_init` (shared
preamble via `_init_head`); each per-plane rootfs bakes only its role init,
so the `bb_role` cmdline branch is gone.
* infra_artifact is role-parametrized: a per-role package
(bot-bottle-firecracker-<role>), version hash, URL, cache dir, and candidate
subdir. `ensure_built` pulls both; `_expected_version` combines both markers.
* publish_infra builds the images once, then builds + publishes an
orchestrator and a gateway artifact under DIR/<role>/.
coverage.sh's candidate-dir plumbing is layout-agnostic (publish_infra --output
now fills DIR/<role>/, which ensure_artifact_gz reads). Full suite green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The orchestrator/gateway container split retired the docker backend's combined
`bot-bottle-infra` container, leaving `Dockerfile.infra` used only as the base
layer for the firecracker rootfs image — and its header still (falsely) claimed
"used directly by the Docker backend."
Fold its one `COPY --from orchestrator` into `Dockerfile.infra.fc` (now
`FROM bot-bottle-gateway` directly + the copy + buildah), delete Dockerfile.infra,
and drop the now-dead build step + artifact-version entry. `build_infra_images_with_docker`
builds three images instead of four; the firecracker rootfs is unchanged in
content (still gateway + orchestrator package + buildah, one artifact).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pull the control plane's host-side logic out of `infra_vm` into
`FirecrackerOrchestrator` (backend/firecracker/orchestrator.py): booting the
orchestrator microVM on its link with the persistent registry volume (/dev/vdb),
seeding the host-canonical signing key over SSH, waiting for /health, and the
buildah SSH target. `url()` == `gateway_url()` (the gateway resolves the
orchestrator by guest IP via bb_orch). This completes the trio — docker, macOS,
firecracker — behind both the Gateway and Orchestrator ABCs.
`infra_vm` stays the pair coordinator + shared substrate: `ensure_running` now
mints nothing itself — it constructs both services (lazy import to break the
cycle), calls `orchestrator.ensure_running()` first, then
`gateway.connect_to_orchestrator(orch.gateway_url(), orch.mint_gateway_token())`.
The plane-agnostic boot primitives graduate to public API (`_boot_vm`/
`_push_secret` -> `boot_vm`/`push_secret`) now that they're consumed only across
modules; `InfraVm` loses its `orchestrator_url` (moved to the service) and
`InfraEndpoint.orchestrator` is now the `FirecrackerOrchestrator` service.
Container-lifecycle tests split into test_firecracker_orchestrator; the infra_vm
tests keep the substrate + pair-coordinator coverage.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Extract the Apple-Container control-plane lifecycle out of `MacosInfraService`
into `MacosOrchestrator` (backend/macos_container/orchestrator.py): the
container run on the host-only control network, the container-only DB volume,
the source-hash recreate gate, health polling against the resolved
control-network address, and `probe_orchestrator_url`. `url()` == `gateway_url()`
here (Apple has no container DNS, so one resolved address serves the CLI and the
gateway).
`MacosInfraService` now composes `orchestrator()` + `gateway()`: ensure the
networks, build both images (each service self-builds via `ensure_built` — added
to `MacosGateway` too), bring the orchestrator up first, then connect the gateway
with the orchestrator-minted token. `ORCHESTRATOR_IMAGE` + the DB-volume constant
move to the orchestrator module (with back-compat aliases where imported).
Container-lifecycle tests split into test_macos_orchestrator; test_macos_infra
now covers the composition.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The two-VM stop() only knew the new orchestrator/gateway config paths
(<infra>/{orchestrator,gateway}/config.json), so a host still running the
OLD single combined-VM (booted from <infra>/config.json on the orchestrator
link) kept it alive across the migration — and it held the orchestrator TAP,
so the first split boot died with "Open tap device failed: Resource busy"
(seen on the CI runner's bborchci0 pool).
stop() now also reaps that legacy singleton (both its drifted pidfile and any
firecracker whose --config-file is the legacy path), so the cutover is
self-healing with no manual per-host teardown. Scoped to the legacy infra
config path, so pool agent VMs are untouched.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Mirror the Gateway work for the control plane. Add the backend-neutral
`Orchestrator` service ABC in orchestrator/lifecycle.py — build the image/rootfs,
`ensure_running` (start + health-gate), `is_running`/`is_healthy`/`stop`,
`url()` (host-facing) vs `gateway_url()` (what the gateway resolves against),
and a concrete `mint_gateway_token()` (the orchestrator holds the signing key,
so it issues the gateway's token — #469).
`DockerOrchestrator` (backend/docker/orchestrator.py) is the first impl,
extracted out of `DockerInfraService`: the control-plane container run, the
`--internal` control network, the source-hash recreate gate, health polling,
and the orchestrator constants (name/label/network/image/dockerfile/hash-label).
`DockerInfraService` now composes `orchestrator()` + `gateway()` — each builds
its own image via `ensure_built`, the orchestrator comes up first, then the
gateway is connected with the orchestrator-minted token. Its `url`/`is_healthy`
delegate to the orchestrator service.
Split the container-lifecycle tests into test_docker_orchestrator; test_docker_infra
now covers the composition. macOS + firecracker follow.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Free the `Orchestrator` name for the incoming host-side control-plane service
ABC (parallel to `Gateway`). The in-guest control-plane core (registry + broker,
in orchestrator/service.py) becomes `OrchestratorCore`; server.py, __main__.py,
the lazy facade, and the two service/server tests updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rename the module files to match their functions' verb-first names:
`bottle_provision.py` -> `provision_bottle.py`,
`gateway_provision.py` -> `provision_gateway.py` (and the test file to match).
Imports, docstrings, and the git_gate/service.py pointer updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The backends' public teardown entry point now pairs with `launch_consolidated`
under the provision/deprovision verb used everywhere else
(`deprovision_bottle`, `deprovision_git_gate`). Renamed across all three
backends' consolidated_launch + launch modules and the tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rename `consolidated_util` -> `bottle_provision` and align it with
`gateway_provision` as the two provisioning domains: `bottle_provision` owns
the bottle lifecycle (`provision_bottle` + `deprovision_bottle`, the latter
renamed from `teardown_consolidated`), `gateway_provision` owns the git-gate
state pushed through the transport (`provision_git_gate`/`deprovision_git_gate`).
The three backends' consolidated_launch modules import from the new module;
their own public `teardown_consolidated` wrappers (called by launch.py) keep
their names.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
`backend/docker/gateway_provision.py` had no docker-specific code left once
`DockerGatewayTransport` moved out — `provision_git_gate` / `deprovision_git_gate`
drive any `GatewayTransport`, and the guest-side paths they write
(`/git-gate/creds/<id>`, `/git/<id>`, `/etc/git-gate/...`) are identical inside
every backend's gateway. Yet the neutral `backend/consolidated_util.py` reached
into the docker package to import them.
Move it to `backend/gateway_provision.py`, a sibling of its only importer. The
gateway *package* can't host it (git_gate already imports gateway.git_gate,
so gateway importing git_gate would cycle), but the backend layer uses git_gate
freely. Docstrings + the git_gate/service.py pointer updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Give every backend a `gateway_transport.py` holding just its `GatewayTransport`
impl, named consistently with the backend: `DockerGatewayTransport`,
`MacosGatewayTransport` (was `AppleGatewayTransport`), and
`FirecrackerGatewayTransport` (was `SshGatewayTransport`).
- docker: split `DockerGatewayTransport` out of `gateway_provision.py`, which
keeps only the backend-neutral provisioning logic (provision_git_gate /
deprovision_git_gate) the other backends share.
- macOS: `gateway_provision.py` held only the transport, so it's replaced by
`gateway_transport.py`.
- firecracker: move the transport out of `gateway.py`.
Behaviour-preserving; importers + tests updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The `backend/docker/provision` package held no modules — a placeholder kept
"in case backend-specific provisioners are added later" after PRD 0050 moved
provisioning to the AgentProvider plugin. Nothing imports it, so drop it; a
new package is cheap to add if that need ever materializes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pull the gateway's host-side logic out of `infra_vm` and into
`FirecrackerGateway`'s own methods, so the gateway is self-contained and
`infra_vm` shrinks toward being just the pair coordinator (a step toward
removing it). Now living in the gateway service: the gateway VM boot +
`bb_orch` cmdline, the pre-minted JWT push, the mitmproxy CA fetch over SSH,
the `SshGatewayTransport` exec/cp, and the lifecycle predicates
(is_running/stop/address). `ensure_running` / `_adopt` construct the gateway
and call `connect_to_orchestrator` (lazy import to break the cycle); the
InfraEndpoint's gateway is now the `FirecrackerGateway` service.
What deliberately stays shared in `infra_vm` (documented at the top of
gateway.py): the plane-agnostic VM substrate the orchestrator VM also needs
(`_boot_vm`, the stable SSH keypair, the secret-push retry, PID lifecycle), the
pair coordinator (`ensure_running`: orchestrator-first health gate + singleton
lock + adoption/version marker), and the single shared `bb_role`-branched init
baked into the one published rootfs both VMs boot — the guest-side gateway
startup can't live in a host method. These are the seam that moves to a neutral
module when the Orchestrator service lands and `infra_vm` dissolves.
CA fetch now raises GatewayError (was die/SystemExit) to match the ABC.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add `FirecrackerGateway`, the Firecracker implementation of the shared
`Gateway` service — completing the trio (docker, macOS, firecracker) behind one
uniform contract. The gateway here is a whole microVM, so the adapter composes
`infra_vm`'s VM primitives (boot, SSH secret push, netpool links) into the ABC
surface: `connect_to_orchestrator(url, token)` parses the orchestrator guest IP
off the URL and boots the gateway VM seeding the pre-minted token; `address()`
is the gateway link's guest IP; `ca_cert_pem()` fetches over SSH (adapting the
running VM when this process didn't boot it); `provisioning_transport()` is the
SSH exec/cp transport.
Trust-model alignment with the other backends: minting moves out of
`_push_gateway_jwt` to the host composition point — `ensure_running` mints the
role-scoped `gateway` JWT and threads it through `boot_gateway(orch_ip, token)`,
so the push step only writes the token the gateway presents (#469).
`infra_vm` keeps driving the orchestrator+gateway pair as a singleton and gains
gateway-only helpers (`gateway_running`/`stop_gateway`/`gateway_guest_ip`/
`adopted_gateway_vm`); the consolidated launch reads the gateway's transport +
CA off the service.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Extract the Apple-Container gateway data plane out of `MacosInfraService` into
`MacosGateway`, the macOS implementation of the shared `Gateway` service.
`connect_to_orchestrator(url, gateway_token)` runs the triple-homed gateway
container (egress + agent + control networks) carrying the mitmproxy CA and the
pre-minted `gateway` token; `address()` / `ca_cert_pem()` /
`provisioning_transport()` round out the contract.
The infra service now composes the gateway: `ensure_running` mints the
role-scoped `gateway` JWT (it holds the signing key; the gateway never does —
#469) and hands it to `gateway().connect_to_orchestrator(...)`, and
`ca_cert_pem` delegates to the service. The inlined `_ensure_gateway_container`
+ CA-read are gone. `AppleGatewayTransport` + the gateway container name/label
now live with the gateway module, breaking the old infra→provision coupling.
Splits the gateway-run + CA tests out of test_macos_infra into a dedicated
test_macos_gateway; the infra tests mock `svc.gateway`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replace the docker backend's ad-hoc gateway wiring with the shared `Gateway`
service ABC (in the gateway package), the first of the two per-host service
classes that supersede the per-backend infra glue.
The contract is backend-neutral: `connect_to_orchestrator(url, gateway_token)`
binds the gateway to its control plane and brings it up, `address()` reports
the agent-facing proxy target, `ca_cert_pem()` vends the mitmproxy CA, and
`provisioning_transport()` hands back the exec/cp seam git-gate provisioning
stages per-bottle repos + deploy keys through. `GatewayProvisionError` +
`GatewayTransport` move to the ABC so every backend shares them.
Key trust-model change: the gateway no longer mints its own token. It never
holds the signing key (#469), so the orchestrator mints the role-scoped
`gateway` JWT and hands it in via `connect_to_orchestrator`; the gateway only
injects it. `DockerGateway.ensure_running` becomes `connect_to_orchestrator`
(stashing the URL + token as instance state), and the docker launch flow reads
`address()` / `provisioning_transport()` off the service rather than
re-deriving the container IP + transport itself.
macOS + firecracker gateways move under the ABC in following commits.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Now that #469 got the DB off the data plane, the Firecracker infra runs as
two microVMs instead of one — mirroring the docker/macos plane split:
* orchestrator VM (ORCH_IFACE) — control plane + buildah image builds; sole
DB opener; host-seeded signing key. No gateway daemons.
* gateway VM (new GW_IFACE) — egress / git-http / supervise data plane;
mitmproxy CA + a host-minted `gateway` JWT (never the key). Reaches the
orchestrator only over the one nft forward rule its link allows.
Both boot the SAME shared infra rootfs; a `bb_role=` kernel-cmdline arg
selects which plane a VM's PID-1 init starts, so there is still one published
artifact. The gateway learns the orchestrator's address via `bb_orch=` on the
cmdline (no IP baked into the artifact).
Isolation is nearly free: agents were already nft-dropped except the DNAT'd
gateway ports, so re-pointing that single DNAT rule at the gateway VM
(`dnat to gw_guest`) severs every agent's L3 route to the control plane. The
only added nft is the second infra link's mirror block (masquerade egress +
forward accept, which subsumes gateway->orchestrator) in the shared shell
script and the NixOS module.
netpool gains GW_IFACE + gw_slot() (the /31 above the orch link);
firecracker_vm.boot gains extra_boot_args for the role cmdline; infra_vm
ensure_running() boots + adopts the pair (orchestrator first, then the gateway
that resolves policy against it) and returns an InfraEndpoint mirroring the
docker/macos shape. Builds stay in the orchestrator (PRD 0070 v1); the gateway
is the slim unit.
Unit-tested (test_firecracker_infra_vm rewritten for two VMs; gw_slot helper
test added); the KVM boot / L3-isolation checks are validated on a Firecracker
host.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The single Apple container existed only because two guests writing one
bot-bottle.db over virtiofs would race incoherent fcntl locks; #469 removed
that (the data plane no longer opens the DB), so the macOS backend now runs
the planes as two containers like docker:
* bot-bottle-mac-orchestrator — lean control plane on the host-only
`bot-bottle-mac-control` network only (image Dockerfile.orchestrator,
`-m bot_bottle.orchestrator`). Sole mounter of the container-only DB
volume; holds the signing key. The CLI + gateway reach it at its
control-network address.
* bot-bottle-mac-infra — the gateway, triple-homed on the NAT egress net,
the host-only agent net, and the control net. Resolves the orchestrator by
IP (Apple has no container DNS) via BOT_BOTTLE_ORCHESTRATOR_URL; holds the
CA + gateway JWT.
Agents sit on the agent network only, so they have no route to the control
plane. ensure_networks gains the control network; MacosInfraService brings up
the orchestrator then the gateway; probe_orchestrator_url + ca_cert_pem target
the right containers. pyright 0 errors; unit suite green (2274). Needs a real
Apple-container host to validate the networking.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Now that #469 got the DB off the data plane, the docker backend runs the
control plane and data plane as two containers instead of the combined
`bot-bottle-infra` container:
* bot-bottle-orchestrator — the lean control-plane container (image
Dockerfile.orchestrator, `-m bot_bottle.orchestrator`). Joins the new
`bot-bottle-orchestrator` control network (`--internal`) only, plus a
host-loopback publish for the CLI. Sole opener of bot-bottle.db; holds the
signing key; no CA, no gateway daemons.
* bot-bottle-orch-gateway — the data-plane container (DockerGateway),
dual-homed on the agent `bot-bottle-gateway` network AND the control
network, so it resolves the orchestrator by name
(http://bot-bottle-orchestrator:8099) while agents — never on the control
network — have no route to the control plane (the L3 block, not just the
JWT). Holds the gateway JWT + the mitmproxy CA.
DockerInfraService now orchestrates the pair (builds gateway + orchestrator
images, brings up the orchestrator then the gateway) and exposes gateway_name;
the launch flow attributes agents against the gateway container. DockerGateway
gains a control_network it joins after run. rotate_ca / the CA read target the
gateway container. The combined Dockerfile.infra is no longer built by docker
(firecracker/macOS still use it until their splits).
pyright 0 errors; unit suite green (2273). Integration tests updated for the
two-container shape but need a real-docker run to validate the networking.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Brings in main's 8 commits — #470 (backend-agnostic CI guards: BackendStatus
enum, quiet status(), is_backend_available/is_backend_ready, tests/_backend.py
skip_unless_backend) plus firecracker status() readiness checks (kvm/kernel/
dropbear/mke2fs).
Reconciled #470's additions into the reorg'd backend structure:
* BackendStatus enum + the status(*, quiet=) ABC signature -> backend/base.py
* is_backend_available / is_backend_ready -> backend/selection.py
* exposed all three through the lazy backend facade (__init__).
Integration-test conflicts were our control_plane->orchestrator renames vs
main's skip_unless_docker -> skip_unless_backend swap: kept our refactored
import paths + main's guard. Repointed main's new is_backend_* tests to patch
selection._backends (our moved home) instead of the facade.
pyright: 0 errors. Full unit suite green (2272).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two reviewer findings addressed:
1. _kvm_accessible() now opens /dev/kvm with O_RDWR|O_CLOEXEC instead
of read-only. VM creation requires write access; a read-only
descriptor can satisfy KVM_GET_API_VERSION but fails at boot time.
2. status() now checks every hard prerequisite that require_firecracker()
checks at launch: guest kernel image, static dropbear binary, and
mke2fs. Previously a host with a configured TAP pool but missing
artifacts could pass status() and then fail during launch.
Regression coverage: TestFirecrackerKvmCheck gains test_kvm_open_rdwr_fails;
TestFirecrackerArtifactCheck covers each missing artifact; existing
TestFirecrackerStatus fixtures updated to stub the new checks.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds `_firecracker_binary_ok()` (runs `firecracker --version`) and
`_kvm_accessible()` (opens /dev/kvm, issues KVM_GET_API_VERSION ioctl)
and gates `status()` on both, so the launch preflight reports missing
binary or inaccessible KVM before the caller ever attempts a boot.
Regression tests in TestFirecrackerBinaryCheck, TestFirecrackerKvmCheck,
and TestFirecrackerStatusRuntime cover all branches. Updated the
pre-existing TestFirecrackerStatus fixture to also stub the two new
helpers so it still exercises only the tap-pool/nft logic it was
designed for.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>