fix(supervise): get bot-bottle.db off the data plane (supervise + egress proposals over RPC) #471
@@ -353,6 +353,116 @@ in memory). Rotation does not auto-re-provision the new CA into running bottles
|
||||
— those re-attach to install the new anchor — so it is an operator action with
|
||||
a brief egress interruption, never an implicit one.
|
||||
|
||||
### Separating the planes: distinct orchestrator and gateway runtimes
|
||||
|
||||
The combined per-host infra unit above (one VM on Firecracker/macOS, one
|
||||
container on Docker running both planes) was forced by **DB coherence**, not
|
||||
preference: two guest kernels sharing a virtiofs-mounted `bot-bottle.db` don't
|
||||
get coherent `fcntl` locks, so both planes had to live in the one guest that
|
||||
owned the file. **Issue #469 removed that constraint** — the data plane no
|
||||
longer opens the DB at all; it reaches the queue/registry over the
|
||||
control-plane RPC (see §State). With the coupling gone, the orchestrator
|
||||
(control plane) and the gateway (data plane) can run as **separate runtimes**:
|
||||
separate VMs on Firecracker, separate containers on macOS and Docker. (This
|
||||
only re-splits the *orchestrator↔gateway* boundary; agents stay one
|
||||
unit-each, unchanged.)
|
||||
|
||||
**Why do it:** less surface to compromise (a breached gateway shares no
|
||||
runtime with the control plane), independent lifecycle/placement (scale or
|
||||
relocate the gateway; put the orchestrator elsewhere), and separation of
|
||||
concerns the topology *enforces* rather than the process boundary merely
|
||||
implying it. **The cost is lifecycle** — two units to build/start/adopt/
|
||||
health-check/stop per backend — not access (below) and not, right-sized,
|
||||
memory (below).
|
||||
|
||||
#### Access: keep the agent's isolation, gain it where it's missing
|
||||
|
||||
The only relationship that changes is **gateway → orchestrator**: loopback
|
||||
(`127.0.0.1:8099`) inside one unit today, a network hop when split. Everything
|
||||
else is unchanged — CLI/console reach the orchestrator as before, and the
|
||||
agent reaches only the gateway's data ports and **never** needs the
|
||||
orchestrator (it flows agent → gateway supervise MCP → orchestrator, never
|
||||
agent → orchestrator directly).
|
||||
|
||||
The invariant to preserve is **agent cannot reach the orchestrator** — and
|
||||
today it holds *unevenly*:
|
||||
|
||||
| Backend | agent → orchestrator today | enforced by |
|
||||
|---|---|---|
|
||||
| Docker | reachable at `gateway_ip:8099` | **token only** — shared plain bridge, orchestrator binds `0.0.0.0` |
|
||||
| macOS | reachable at `infra_ip:8099` | **token only** — no `pfctl` rule |
|
||||
| Firecracker | **unreachable** | nft: /31 TAP, forward chain drops all but the DNAT'd gateway ports |
|
||||
|
||||
So on Docker/macOS the agent is one `curl` from the control plane, gated
|
||||
*only* by the JWT (one open-mode bug — exactly the class of the #469
|
||||
Firecracker regression — or one token leak from a bypass). Firecracker already
|
||||
blocks it at L3. **The split is the moment to fix that, not a risk to it:**
|
||||
|
||||
- **Docker / macOS — a dedicated `bot-bottle-orchestrator` control network.**
|
||||
A second `--internal` network that **only** the orchestrator and gateway
|
||||
join. The gateway is **dual-homed** (data network `bot-bottle-gateway` +
|
||||
control network `bot-bottle-orchestrator`); the orchestrator joins the
|
||||
control network only (plus the host-loopback publish the CLI already uses).
|
||||
The agent is never attached to the control network, so it has **no route**
|
||||
to the orchestrator — the same L3 block Firecracker already has. The one
|
||||
wiring change: gateway daemons point `BOT_BOTTLE_ORCHESTRATOR_URL` at
|
||||
`http://bot-bottle-orchestrator:8099` (control-net address) instead of
|
||||
loopback. `DockerGateway` already takes an `orchestrator_url` and owns its
|
||||
network — it needs the second `--network` and the URL retargeted; the
|
||||
orchestrator side is `Dockerfile.orchestrator` run as its own container on
|
||||
the control net. macOS mirrors this with a second container network and the
|
||||
orchestrator's pinned control-net IP (no Apple DNS — reuse the
|
||||
`gateway_hosts` `/etc/hosts` mechanism).
|
||||
- **Firecracker — nearly free.** The nft forward chain already drops
|
||||
everything a VM sends except the DNAT'd gateway ports, so a *separate*
|
||||
orchestrator VM is unreachable by agent VMs with **no new agent rules**. The
|
||||
new piece is a gateway-VM ↔ orchestrator-VM link — the same shape as today's
|
||||
orch `/31` link, but between the two infra VMs instead of loopback within
|
||||
one — with a forward-accept rule for that link alone. Agent egress DNATs to
|
||||
the **gateway** VM's data ports; nothing DNATs to the orchestrator VM.
|
||||
|
||||
Role-scoped JWT auth (#469) stays as defense-in-depth on all three,
|
||||
unchanged. A *compromised gateway* can already reach the orchestrator today
|
||||
(loopback, with its `gateway` JWT) and role scoping is what contains that — so
|
||||
the split adds **no new surface** for the gateway-compromise case; it only
|
||||
closes the *agent* path on Docker/macOS.
|
||||
|
||||
#### Memory: right-size, don't double
|
||||
|
||||
Firecracker guest memory is a **fixed ceiling** set at boot (`mem_size_mib`);
|
||||
it's demand-paged by the host (RSS grows with what the guest touches) but there
|
||||
is **no reclaim without a balloon device**, and none is configured — so plan
|
||||
for the ceiling. The combined infra VM is **4096 MiB**, but that headroom is
|
||||
for **buildah** (in-guest agent-image builds), not the daemons, which are
|
||||
light.
|
||||
|
||||
Split naïvely that reads as 2×, but right-sized it is **less** than today:
|
||||
|
||||
- **Orchestrator** — stdlib-only Python HTTP + SQLite, no builds → **256–512
|
||||
MiB**.
|
||||
- **Gateway** — mitmproxy (TLS bump + DLP body buffering) + git-http +
|
||||
supervise + gitleaks → **512 MiB–1 GB**, *if it does not build images*.
|
||||
|
||||
The real lever is **where agent-image builds run** (buildah's 2–4 GB working
|
||||
set). Move builds off the gateway — to a host-controller / dedicated build unit
|
||||
(a #468-adjacent decision) — and a slim orchestrator (~512 MiB) + slim gateway
|
||||
(~1 GB) ≈ **1.5 GB total vs. today's 4 GB**. A per-unit **balloon device** to
|
||||
reclaim idle memory is a later option, not needed for v1.
|
||||
|
||||
#### iroh belongs in the orchestrator
|
||||
|
||||
The remote-access transport (iroh — `web console -(iroh)-> orchestrator`, #468)
|
||||
terminates at the **orchestrator**: everything the console/phone drives —
|
||||
launch, teardown, list, supervise approvals, policy — is a control-plane
|
||||
operation, so the trusted control plane is the single remote entry point. Keep
|
||||
it **out of the gateway**: the data plane is the more-exposed, agent-facing
|
||||
unit, and a goal of the split is to shrink its surface and keep its image slim
|
||||
(no iroh/Rust dependency). Agent-PTY-to-mobile streaming is a *data-plane*
|
||||
stream and a separate concern; if it must reach the phone over the same
|
||||
endpoint, **relay it through the orchestrator's iroh door** rather than opening
|
||||
a second door in the gateway — one authenticated remote entry point, plane
|
||||
split intact.
|
||||
|
||||
## Sequencing
|
||||
|
||||
Jump straight to the **virtualized** end state (not a host-daemon stepping
|
||||
@@ -440,3 +550,12 @@ Keep the gateway **service one shared thing** throughout.
|
||||
plane (add/remove routes/keys/proposals without a restart).
|
||||
- **Identity-token delivery:** exactly how the per-bottle token is placed
|
||||
where the agent can present it but not swap in another bottle's.
|
||||
- **Where agent-image builds run once the planes split:** buildah's 2–4 GB
|
||||
working set is what keeps the combined infra VM at 4 GB. Keeping the gateway
|
||||
slim (~1 GB) means builds move off it — to a host-controller / dedicated
|
||||
build unit (#468-adjacent) — or the gateway keeps the build headroom and the
|
||||
memory win is smaller. Decides the gateway VM/container size.
|
||||
- **Agent-PTY-to-mobile streaming path:** relay it through the orchestrator's
|
||||
iroh endpoint (default — one authenticated remote door, plane split intact),
|
||||
or give the data-plane stream its own transport. Only matters once remote
|
||||
terminal mirroring lands.
|
||||
|
||||
Reference in New Issue
Block a user