Files
bot-bottle/docs/prds/0081-reprovision-gateway-state-on-bringup.md
T
didericis cecc1b4d60 docs(prd): reprovision gateway-dependent state on gateway bring-up (0081)
Reconcile every running bottle against a freshly-booted gateway (replace
each agent's CA, re-provision git-gate, restore egress tokens) instead of
persisting gateway state on volumes. Supersedes the abandoned CA-volume
(#511) and git-gate-volume (#513) PRs.

Refs #510, #512

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-27 15:14:33 +00:00

5.2 KiB

PRD 0081: Reprovision gateway-dependent state on gateway bring-up

  • Status: Active
  • Author: claude
  • Created: 2026-07-26
  • Issue: #510, #512

Summary

When the gateway is (re)built, reconcile every already-running bottle against the fresh gateway instead of persisting the gateway's state. On a gateway cold boot, the gateway reconciles all live bottles in one flow: replace each agent's trusted CA with the freshly-minted gateway CA, re-provision each bottle's git-gate repos + creds onto the gateway, and restore each bottle's egress tokens. One mechanism across all three services; the CA rotates for free on every bring-up.

Problem

A gateway rebuild/restart silently breaks every already-running bottle:

  • CA (#510). The gateway's mitmproxy mints a new CA on a fresh rootfs; agents still trust the old one, so egress fails TLS verification (SSL certificate verification failed).
  • git-gate (#512). Per-bottle bare repos (/git/<id>) + deploy creds (/git-gate/creds/<id>) live in the gateway's ephemeral rootfs; a rebuild wipes them and the agent 404s on fetch/push.

Both are the same root cause: per-bottle gateway-dependent state is provisioned once, at bottle launch, and nothing restores it for already-running bottles when the gateway comes back. The existing launch-time reprovision (reprovision_bottles) restores only egress tokens, and only as a side effect of the next launch.

Persisting the state (a host bind-mount / a per-VM data volume per service) was prototyped and rejected: it differs per service (a volume for the CA, another for git-gate, reprovision for tokens), it pins the CA static forever (no rotation), and it adds volume surface that docker volume prune / a wiped cache can silently destroy (the original #450 failure mode).

Goals / Success Criteria

  • A gateway (re)boot restores all running bottles' gateway-dependent state with no manual step and no relaunch — the agent's next egress / fetch / push just works.
  • One reconcile flow covering CA, git-gate, and egress tokens, rather than a different mechanism per service.
  • The CA rotates on every gateway bring-up (no long-lived CA), distributed to running agents by the same reconcile.
  • Per-bottle failures are tolerated: one unreachable or malformed bottle does not block the others or the gateway coming up.

Non-goals

  • Deliberate mid-session CA rotation without a gateway restart — rotate_ca stays for that operator action.
  • The docker / macOS backends. This lands the reconcile for Firecracker (the backend that boots a fresh gateway rootfs each time); docker keeps its existing host-bind-mounted CA until the same seam is extended there (follow-up). The unification target is per-service, not per-backend.
  • Changing source-IP attribution, /resolve, or the plane split (#469).

Design

Trigger — the gateway startup flow. FirecrackerGateway.connect_to_orchestrator is the cold-boot path: FirecrackerInfraService.ensure_running calls it only when it boots a fresh pair, and skips it when it adopts a healthy, current gateway (state intact). So reconcile fires exactly when the gateway was (re)booted — including orchestrator restarts, since the pair boots together — and never on an adopt. There is no bare-restart path that bypasses ensure_running.

reconcile_running_bottles(gateway, orchestrator_url) (new firecracker/reconcile.py), called at the end of connect_to_orchestrator once the gateway VM is up and its CA is available:

  1. Map each live bottle's guest IP → bottle_id from the orchestrator registry (list_bottles).
  2. Install the current shared git-gate hooks (git_gate_render_hook / git_gate_render_access_hook) into the fresh gateway once — rendered from code, never from a bottle's possibly-stale state dir.
  3. For each live agent VM (enumerated from its run dir):
    • CA: SSH the current gateway CA into the agent's trust store and run update-ca-certificates (unconditional replace — there is one gateway, so no fingerprint match is needed).
    • git-gate: rebuild the bottle's upstreams from its persisted git-gate state dir (deploy key, known_hosts, upstream URL) and re-init its bare repos
      • per-repo creds under /git/<bottle_id>.
    • egress token: read the agent's ENV_VAR_SECRET and feed reprovision_bottles, restoring the orchestrator's in-memory tokens.
  4. Every per-bottle step is wrapped so one failure is logged and skipped.

Retire the persistence prototype. No CA data volume, no git-gate data volume (the abandoned PRs #511 / #513). The launch-time _reprovision_running_bottles call folds into this bring-up reconcile, so egress tokens are restored on the same cold-boot trigger (an adopt needs no restore — the orchestrator never restarted).

CA rotation. Because the gateway rootfs is ephemeral, every cold boot mints a fresh CA; reconcile is what distributes it, so a routine gateway rebuild doubles as a CA rotation with zero extra machinery.

Open questions

  • Docker / macOS parity: extend the same reconcile seam (docker exec instead of SSH) and drop docker's CA bind-mount, or leave docker on persistence? Deferred to a follow-up; not blocking the Firecracker fix.