feat(firecracker): split orchestrator and gateway into separate VMs (PRD 0070)
tracker-policy-pr / check-pr (pull_request) Successful in 12s
test / integration-docker (pull_request) Successful in 19s
lint / lint (push) Successful in 53s
test / unit (pull_request) Failing after 1m46s
test / integration-firecracker (pull_request) Failing after 2m42s
test / coverage (pull_request) Has been skipped
test / publish-infra (pull_request) Has been skipped
tracker-policy-pr / check-pr (pull_request) Successful in 12s
test / integration-docker (pull_request) Successful in 19s
lint / lint (push) Successful in 53s
test / unit (pull_request) Failing after 1m46s
test / integration-firecracker (pull_request) Failing after 2m42s
test / coverage (pull_request) Has been skipped
test / publish-infra (pull_request) Has been skipped
Now that #469 got the DB off the data plane, the Firecracker infra runs as two microVMs instead of one — mirroring the docker/macos plane split: * orchestrator VM (ORCH_IFACE) — control plane + buildah image builds; sole DB opener; host-seeded signing key. No gateway daemons. * gateway VM (new GW_IFACE) — egress / git-http / supervise data plane; mitmproxy CA + a host-minted `gateway` JWT (never the key). Reaches the orchestrator only over the one nft forward rule its link allows. Both boot the SAME shared infra rootfs; a `bb_role=` kernel-cmdline arg selects which plane a VM's PID-1 init starts, so there is still one published artifact. The gateway learns the orchestrator's address via `bb_orch=` on the cmdline (no IP baked into the artifact). Isolation is nearly free: agents were already nft-dropped except the DNAT'd gateway ports, so re-pointing that single DNAT rule at the gateway VM (`dnat to gw_guest`) severs every agent's L3 route to the control plane. The only added nft is the second infra link's mirror block (masquerade egress + forward accept, which subsumes gateway->orchestrator) in the shared shell script and the NixOS module. netpool gains GW_IFACE + gw_slot() (the /31 above the orch link); firecracker_vm.boot gains extra_boot_args for the role cmdline; infra_vm ensure_running() boots + adopts the pair (orchestrator first, then the gateway that resolves policy against it) and returns an InfraEndpoint mirroring the docker/macos shape. Builds stay in the orchestrator (PRD 0070 v1); the gateway is the slim unit. Unit-tested (test_firecracker_infra_vm rewritten for two VMs; gw_slot helper test added); the KVM boot / L3-isolation checks are validated on a Firecracker host. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
@@ -1,17 +1,17 @@
|
||||
"""Docker-free agent-image builds for the Firecracker backend (PRD 0069 Stage 3).
|
||||
|
||||
Agent Dockerfiles build **inside the persistent per-host infra VM**
|
||||
Agent Dockerfiles build **inside the persistent per-host orchestrator VM**
|
||||
(`infra_vm.py`), which carries buildah (rootless, daemonless): no host Docker
|
||||
daemon, no root-equivalent `docker` group. The build runs over SSH against the
|
||||
infra VM and its rootfs streams back to the host, where the existing
|
||||
orchestrator VM and its rootfs streams back to the host, where the existing
|
||||
`mke2fs -d` path (`util.build_rootfs_ext4`) turns it into a bootable ext4.
|
||||
|
||||
Building in the infra VM — rather than a throwaway builder VM — means there is
|
||||
one buildah image (`bot-bottle-infra`) and no contention for the orchestrator
|
||||
TAP. Tradeoff: an untrusted Dockerfile's `RUN` steps share the VM with the
|
||||
control plane + gateway (buildah `--isolation chroot` isn't a hard boundary) —
|
||||
the accepted single-VM blast-radius tradeoff, re-splittable into a disposable
|
||||
builder (booted from this same image on its own TAP) later.
|
||||
Builds run on the control-plane (orchestrator) side — not the data-plane
|
||||
gateway — per PRD 0070 v1 ("builds stay in the orchestrator"): it owns
|
||||
launches and already carries buildah. Tradeoff: an untrusted Dockerfile's
|
||||
`RUN` steps share the VM with the control plane (buildah `--isolation chroot`
|
||||
isn't a hard boundary) — the accepted blast-radius tradeoff, re-splittable
|
||||
into a disposable builder (booted from this same image on its own TAP) later.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -121,11 +121,12 @@ def _build_lock() -> Generator[None, None, None]:
|
||||
def _build_in_infra(
|
||||
dockerfile: Path, base: Path, smoke_test: tuple[str, ...], digest: str,
|
||||
) -> None:
|
||||
"""Ensure the infra VM is up, `buildah build` the Dockerfile in it, smoke
|
||||
test the image, and stream its rootfs into `base`. The infra VM persists;
|
||||
only the per-build container/image/context are cleaned up."""
|
||||
infra = infra_vm.ensure_running()
|
||||
key, ip = infra.private_key, infra.guest_ip
|
||||
"""Ensure the infra VMs are up, `buildah build` the Dockerfile in the
|
||||
orchestrator VM (the build-capable control plane), smoke test the image,
|
||||
and stream its rootfs into `base`. The infra VMs persist; only the
|
||||
per-build container/image/context are cleaned up."""
|
||||
orchestrator = infra_vm.ensure_running().orchestrator
|
||||
key, ip = orchestrator.private_key, orchestrator.guest_ip
|
||||
tag = f"bot-bottle-agent-build-{digest}"
|
||||
ctx = f"/tmp/agent-build-{digest}"
|
||||
smoke_ctr, export_ctr = f"{tag}-smoke", f"{tag}-export"
|
||||
|
||||
Reference in New Issue
Block a user