Files
bot-bottle/docs/prds/0083-reprovision-gateway-state-on-bringup.md
T
didericis-claude cf6bbc43da
test / integration-docker (pull_request) Has been cancelled
prd-number-check / require-numbered-prds (pull_request) Successful in 14s
lint / lint (push) Successful in 1m4s
test / image-input-builds (pull_request) Successful in 59s
tracker-policy-pr / check-pr (pull_request) Successful in 11s
test / unit (pull_request) Failing after 13m55s
test / coverage (pull_request) Has been skipped
docs(prd): renumber 0081 -> 0083 (0081 already taken)
PRD number 0081 was claimed by another PRD before this one merged (main is on
0082); renumber this design to the next free slot, 0083. File + title only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-27 15:14:54 +00:00

6.0 KiB

PRD 0083: Reprovision gateway-dependent state on gateway bring-up

  • Status: Active
  • Author: claude
  • Created: 2026-07-26
  • Issue: #510, #512

Summary

When the gateway is (re)built, reconcile every already-running bottle against the fresh gateway instead of persisting the gateway's state. On a gateway cold boot, the gateway reconciles all live bottles in one flow: replace each agent's trusted CA with the freshly-minted gateway CA, re-provision each bottle's git-gate repos + creds onto the gateway, and restore each bottle's egress tokens. One mechanism across all three services; the CA rotates for free on every bring-up.

Problem

A gateway rebuild/restart silently breaks every already-running bottle:

  • CA (#510). The gateway's mitmproxy mints a new CA on a fresh rootfs; agents still trust the old one, so egress fails TLS verification (SSL certificate verification failed).
  • git-gate (#512). Per-bottle bare repos (/git/<id>) + deploy creds (/git-gate/creds/<id>) live in the gateway's ephemeral rootfs; a rebuild wipes them and the agent 404s on fetch/push.

Both are the same root cause: per-bottle gateway-dependent state is provisioned once, at bottle launch, and nothing restores it for already-running bottles when the gateway comes back. The existing launch-time reprovision (reprovision_bottles) restores only egress tokens, and only as a side effect of the next launch.

Persisting the state (a host bind-mount / a per-VM data volume per service) was prototyped and rejected: it differs per service (a volume for the CA, another for git-gate, reprovision for tokens), it pins the CA static forever (no rotation), and it adds volume surface that docker volume prune / a wiped cache can silently destroy (the original #450 failure mode).

Goals / Success Criteria

  • A gateway (re)boot restores all running bottles' gateway-dependent state with no manual step and no relaunch — the agent's next egress / fetch / push just works.
  • One reconcile flow covering CA, git-gate, and egress tokens, rather than a different mechanism per service.
  • The CA rotates on every gateway bring-up (no long-lived CA), distributed to running agents by the same reconcile.
  • Per-bottle failures are tolerated: one unreachable or malformed bottle does not block the others or the gateway coming up.
  • All backends (Firecracker, docker, macOS) reconcile through the same contract — an abstract method on the backend base class, so a new backend cannot forget to implement it and none drifts onto a bespoke mechanism.

Non-goals

  • Deliberate mid-session CA rotation without a gateway restart — rotate_ca stays for that operator action.
  • Changing source-IP attribution, /resolve, or the plane split (#469).

Design

The contract — attach_bottled_agents_to_gateway() on the backend ABC. Reconciling running bottles against the current gateway is a backend responsibility (only the backend can enumerate its agents and reach them — firecracker over SSH, docker/macOS over exec/cp), so it is an @abc.abstractmethod on BottleBackend (backend/base.py). Every backend implements it; the host calls it whenever the gateway is (re)brought up. This is what makes the fix cross-backend by construction rather than a per-backend follow-up. It reprovisions all registered bottles' gateway-dependent state: CA, git-gate, and egress tokens.

Trigger — the gateway bring-up path. The host calls attach_bottled_agents_to_gateway() only when the gateway was actually (re)brought up — the cold-boot branch of the infra bring-up (e.g. FirecrackerInfraService.ensure_running after it boots a fresh pair), never on an adopt of a healthy, current gateway (state intact). So it fires exactly when the gateway was (re)booted — including orchestrator restarts, since the pair boots together — and there is no bare-restart path that bypasses bring-up.

Per-backend implementation. Each backend's attach_bottled_agents_to_gateway enumerates its live bottles and, for each, reconciles the three services against the current gateway. The firecracker implementation, once the gateway VM is up and its CA is available:

  1. Map each live bottle's guest IP → bottle_id from the orchestrator registry (list_bottles).
  2. Install the current shared git-gate hooks (git_gate_render_hook / git_gate_render_access_hook) into the fresh gateway once — rendered from code, never from a bottle's possibly-stale state dir.
  3. For each live agent VM (enumerated from its run dir):
    • CA: SSH the current gateway CA into the agent's trust store and run update-ca-certificates (unconditional replace — there is one gateway, so no fingerprint match is needed).
    • git-gate: rebuild the bottle's upstreams from its persisted git-gate state dir (deploy key, known_hosts, upstream URL) and re-init its bare repos
      • per-repo creds under /git/<bottle_id>.
    • egress token: read the agent's ENV_VAR_SECRET and feed reprovision_bottles, restoring the orchestrator's in-memory tokens.
  4. Every per-bottle step is wrapped so one failure is logged and skipped.

Retire the persistence prototype. No CA data volume, no git-gate data volume (the abandoned PRs #511 / #513). The launch-time _reprovision_running_bottles call folds into this bring-up reconcile, so egress tokens are restored on the same cold-boot trigger (an adopt needs no restore — the orchestrator never restarted).

CA rotation. Because the gateway rootfs is ephemeral, every cold boot mints a fresh CA; reconcile is what distributes it, so a routine gateway rebuild doubles as a CA rotation with zero extra machinery.

Open questions

  • Docker's existing host-bind-mounted CA (host_gateway_ca_dir): once docker's attach_bottled_agents_to_gateway pushes the CA to running agents on bring-up, the bind-mount is redundant — drop it (so docker rotates like firecracker) or keep it as belt-and-suspenders? Leaning drop, for one behaviour across backends.