diff --git a/docs/prds/0081-reprovision-gateway-state-on-bringup.md b/docs/prds/0081-reprovision-gateway-state-on-bringup.md new file mode 100644 index 00000000..58e88a06 --- /dev/null +++ b/docs/prds/0081-reprovision-gateway-state-on-bringup.md @@ -0,0 +1,105 @@ +# PRD 0081: Reprovision gateway-dependent state on gateway bring-up + +- **Status:** Active +- **Author:** claude +- **Created:** 2026-07-26 +- **Issue:** #510, #512 + +## Summary + +When the gateway is (re)built, reconcile every already-running bottle against the +fresh gateway instead of persisting the gateway's state. On a gateway cold boot, +the gateway reconciles all live bottles in one flow: **replace** each agent's +trusted CA with the freshly-minted gateway CA, **re-provision** each bottle's +git-gate repos + creds onto the gateway, and **restore** each bottle's egress +tokens. One mechanism across all three services; the CA rotates for free on every +bring-up. + +## Problem + +A gateway rebuild/restart silently breaks every already-running bottle: + +- **CA (#510).** The gateway's mitmproxy mints a new CA on a fresh rootfs; agents + still trust the old one, so egress fails TLS verification (`SSL certificate + verification failed`). +- **git-gate (#512).** Per-bottle bare repos (`/git/`) + deploy creds + (`/git-gate/creds/`) live in the gateway's ephemeral rootfs; a rebuild wipes + them and the agent 404s on fetch/push. + +Both are the same root cause: per-bottle gateway-dependent state is provisioned +**once, at bottle launch**, and nothing restores it for already-running bottles +when the gateway comes back. The existing launch-time reprovision +(`reprovision_bottles`) restores **only** egress tokens, and only as a side effect +of the *next* launch. + +Persisting the state (a host bind-mount / a per-VM data volume per service) was +prototyped and rejected: it differs per service (a volume for the CA, another for +git-gate, reprovision for tokens), it pins the CA static forever (no rotation), +and it adds volume surface that `docker volume prune` / a wiped cache can silently +destroy (the original #450 failure mode). + +## Goals / Success Criteria + +- A gateway (re)boot restores **all** running bottles' gateway-dependent state + with no manual step and no relaunch — the agent's next egress / fetch / push + just works. +- **One** reconcile flow covering CA, git-gate, and egress tokens, rather than a + different mechanism per service. +- The CA **rotates** on every gateway bring-up (no long-lived CA), distributed to + running agents by the same reconcile. +- Per-bottle failures are tolerated: one unreachable or malformed bottle does not + block the others or the gateway coming up. + +## Non-goals + +- Deliberate mid-session CA rotation *without* a gateway restart — `rotate_ca` + stays for that operator action. +- The docker / macOS backends. This lands the reconcile for Firecracker (the + backend that boots a fresh gateway rootfs each time); docker keeps its existing + host-bind-mounted CA until the same seam is extended there (follow-up). The + unification target is *per-service*, not *per-backend*. +- Changing source-IP attribution, `/resolve`, or the plane split (#469). + +## Design + +**Trigger — the gateway startup flow.** `FirecrackerGateway.connect_to_orchestrator` +is the cold-boot path: `FirecrackerInfraService.ensure_running` calls it only when +it boots a fresh pair, and skips it when it adopts a healthy, current gateway +(state intact). So reconcile fires exactly when the gateway was (re)booted — +including orchestrator restarts, since the pair boots together — and never on an +adopt. There is no bare-restart path that bypasses `ensure_running`. + +**`reconcile_running_bottles(gateway, orchestrator_url)`** (new +`firecracker/reconcile.py`), called at the end of `connect_to_orchestrator` once +the gateway VM is up and its CA is available: + +1. Map each live bottle's guest IP → `bottle_id` from the orchestrator registry + (`list_bottles`). +2. Install the **current** shared git-gate hooks (`git_gate_render_hook` / + `git_gate_render_access_hook`) into the fresh gateway once — rendered from + code, never from a bottle's possibly-stale state dir. +3. For each live agent VM (enumerated from its run dir): + - **CA:** SSH the current gateway CA into the agent's trust store and run + `update-ca-certificates` (unconditional replace — there is one gateway, so no + fingerprint match is needed). + - **git-gate:** rebuild the bottle's upstreams from its persisted git-gate + state dir (deploy key, known_hosts, upstream URL) and re-init its bare repos + + per-repo creds under `/git/`. + - **egress token:** read the agent's `ENV_VAR_SECRET` and feed + `reprovision_bottles`, restoring the orchestrator's in-memory tokens. +4. Every per-bottle step is wrapped so one failure is logged and skipped. + +**Retire the persistence prototype.** No CA data volume, no git-gate data volume +(the abandoned PRs #511 / #513). The launch-time `_reprovision_running_bottles` +call folds into this bring-up reconcile, so egress tokens are restored on the same +cold-boot trigger (an adopt needs no restore — the orchestrator never restarted). + +**CA rotation.** Because the gateway rootfs is ephemeral, every cold boot mints a +fresh CA; reconcile is what distributes it, so a routine gateway rebuild doubles +as a CA rotation with zero extra machinery. + +## Open questions + +- Docker / macOS parity: extend the same reconcile seam (docker `exec` instead of + SSH) and drop docker's CA bind-mount, or leave docker on persistence? Deferred + to a follow-up; not blocking the Firecracker fix.