diff --git a/docs/prds/prd-new-egress-control-plane.md b/docs/prds/prd-new-egress-metering-budgets-cutoff.md similarity index 53% rename from docs/prds/prd-new-egress-control-plane.md rename to docs/prds/prd-new-egress-metering-budgets-cutoff.md index f7bb414f..69a3be76 100644 --- a/docs/prds/prd-new-egress-control-plane.md +++ b/docs/prds/prd-new-egress-metering-budgets-cutoff.md @@ -1,4 +1,4 @@ -# PRD prd-new: Egress control plane — metering, budgets, and forced cutoff +# PRD prd-new: Egress metering, budgets, and forced cutoff - **Status:** Draft - **Author:** didericis @@ -7,25 +7,32 @@ ## Summary -Add an **out-of-band egress enforcement & observability plane**: meter every -agent's token usage at the egress proxy, decrement budgets without the agent's -cooperation, and forcibly cut a bottle's egress when a budget is exhausted — -either automatically or on command from a host-level dashboard. The trigger -(usage threshold) and the action (route-drop / freeze / kill) both live in the -egress plane and run with no agent in the loop. This is distinct from the -supervise sidecar (PRD 0013), which is agent-initiated and therefore cannot -enforce a cost cutoff on a runaway agent. State (usage ledger, budgets, audit) -moves into a host-level SQLite database behind a thin repository API, the first -SQL store in an otherwise flat-file repo. +Add **out-of-band cost enforcement** on top of the egress plane: meter every +agent's authoritative token usage at the egress proxy, decrement per-scope +budgets without the agent's cooperation, and forcibly cut a bottle's egress when +a budget is exhausted. The trigger (usage threshold) and the action (route-drop +/ freeze / kill) both run with **no agent in the loop** — this is enforcement, +distinct from the supervise sidecar (PRD 0013), which is agent-initiated and so +cannot stop a runaway agent. + +This builds on infrastructure that already exists and is **not** re-designed +here: + +- **The egress chokepoint** — the MITM proxy every agent's traffic flows through + (PRD 0001 / 0006 / 0017) — is where metering reads and cutoff acts. +- **The host SQLite store** (PRD 0067) is where usage, budget, and enforcement + state persist, behind its existing repository API. +- **Observability** — the host dashboard and egress traffic logging (PRD 0055) — + already surfaces what agents are doing. + +This PRD adds only the three missing pieces: **metering** (authoritative +accounting), **budgets**, and **forced cutoff**. ## Problem -bot-bottle can't currently do two things the cost-overrun case demands: - -1. **Forced egress shutdown on limit.** When an agent crosses a token - threshold, kill its egress automatically — no human in the loop. -2. **Remote (host-level) management.** Drive agents from a single surface: - see usage, cut egress, stop bottles, to prevent cost overruns. +bot-bottle can't meter agent token usage or enforce a cost limit. When an agent +crosses a token threshold, there is no way to kill its egress automatically — no +human in the loop. The existing supervise sidecar (PRD 0013) is **entirely agent-initiated**: every action begins with the agent voluntarily calling an MCP tool and an operator @@ -36,20 +43,23 @@ it mandatory (#249) would not deliver forced cost-cutoff. The requirement forces a distinction the current design blurs: -- **Plane A — enforcement / observability (this PRD).** System → infrastructure. - Meter usage, cut egress on threshold or command, account for cost. - Out-of-band; independent of the agent. **Unconditional** — an enforcement - plane you can opt out of isn't enforcement. +- **Plane A — enforcement (this PRD).** System → infrastructure. Meter usage, + cut egress on threshold, account for cost. Out-of-band; independent of the + agent. **Unconditional** — an enforcement plane you can opt out of isn't + enforcement. - **Plane B — agent-facing recovery (the existing supervise sidecar).** Agent → operator, approval-gated. Useful interactively; meaningless for a headless agent with no operator watching its queue. Remains optional. -This PRD builds Plane A. It reframes the "always-on control" invariant of #249 -as "the egress control plane is always present" — a more defensible property -than "every agent runs the agent-facing supervisor." Unsupervised -(headless/CI/ephemeral) agents stay first-class: still subject to the mandatory -meter + kill switch, they simply lack the agent-facing proposal tools they -couldn't use anyway. +This PRD builds the **enforcement** slice of Plane A — metering, budgets, and +forced cutoff. The chokepoint that makes interception possible and the +observability surfaces that display usage already exist; what's missing is +turning measured usage into an enforced budget. It reframes the "always-on +control" invariant of #249 as "the egress enforcement plane is always present" — +a more defensible property than "every agent runs the agent-facing supervisor." +Unsupervised (headless/CI/ephemeral) agents stay first-class: still subject to +the mandatory meter + kill switch, they simply lack the agent-facing proposal +tools they couldn't use anyway. ## Goals / Success Criteria @@ -61,28 +71,37 @@ couldn't use anyway. most-specific applicable budget governs. - When usage crosses a budget, the bottle's configured **cutoff policy** (`cutoff` | `freeze` | `kill`) fires automatically, executed host-side on the - egress plane — never via the supervise queue. -- An operator can, from a single **host-level TUI dashboard**, see live per-bottle - usage against budget and command a cutoff/stop on demand. + egress plane — never via the supervise queue. An operator can also trigger the + same cutoff on demand through the **existing** host controller/dashboard + surface (that surface is not built here). - Host budgets, default cutoff policy, and per-provider limits are declared in a new host-level `~/.bot-bottle/settings.yml`, parseable by `yaml_subset.py`. -- All usage, budget state, and enforcement actions persist in a host-level - SQLite DB behind a thin repository API, so the store can later be swapped for - a cross-host cloud service. +- Usage, budget state, and enforcement actions persist in the **existing** host + SQLite store (PRD 0067) behind its repository API. ## Non-goals +- **The egress chokepoint — already implemented.** The MITM egress proxy + (PRD 0001 / 0006 / 0017) is the interception point; this PRD hooks metering and + cutoff into it, it does not build or modify the proxy. +- **Observability / dashboard — already implemented.** Live usage display, the + host dashboard, and egress traffic logging (PRD 0055) exist; this PRD writes + the usage/enforcement data they surface but adds no new dashboard and no new + visibility surface. +- **The SQLite store foundation — already implemented (PRD 0067).** This PRD adds + metering/budget/enforcement tables behind the existing repository API; it does + not introduce the store. - **Remote control / cross-host control plane.** Web + mobile remote control, cross-host budgets, and the authn/transport they require are explicitly - deferred. v1 is a **host-only TUI** with no remote surface. + deferred. This is host-only. - **Dollar-denominated budgets.** Budgets are token counts keyed by agent provider, not currency. Price tables are out of scope. - **Migrating existing flat-file state into SQLite.** Resume `metadata.json`, transcripts, Dockerfile overrides, the supervise queue, and audit logs stay on the filesystem. Only the *new* metering/budget/enforcement ledger is SQL. - **Making the supervise sidecar (Plane B) mandatory.** Out of scope here; this - PRD is the answer to "what should be unconditional" (Plane A), leaving #249's - Plane-B question open. + PRD is the answer to "what should be unconditional" (Plane A enforcement), + leaving #249's Plane-B question open. - **Per-request hard pre-send blocking as the primary mechanism.** The gate is budget-crossing detected at/after metering; a pre-flight estimator (below) is a refinement, not the core enforcement path. @@ -129,11 +148,10 @@ bottles already use). Four scopes, most-specific wins: agent → bottle → parent bottle → global (host) ``` -The global host budget is the highest-priority feature to ship (the cross-host -control plane will eventually consume it); per-agent and per-bottle budgets -override it for finer control. A budget can also be supplied **at bottle -launch** (`--budget` or equivalent), overriding the settings.yml defaults for -that run. Enforcement evaluates the effective budget as the +The global host budget is the highest-priority feature to ship; per-agent and +per-bottle budgets override it for finer control. A budget can also be supplied +**at bottle launch** (`--budget` or equivalent), overriding the settings.yml +defaults for that run. Enforcement evaluates the effective budget as the nearest-defined scope at decrement time. ### `~/.bot-bottle/settings.yml` @@ -152,8 +170,9 @@ shutdown: cutoff # default cutoff policy: cutoff | freeze | kill ### Forced cutoff and cutoff policy -On budget exhaustion (or an operator command), the configured per-bottle cutoff -policy fires. The three policies map onto primitives that already exist: +On budget exhaustion (or an operator command via the existing surface), the +configured per-bottle cutoff policy fires. The three policies map onto +primitives that already exist: - **`cutoff`** (default) — drop the bottle's `routes.yaml` to empty and reload (or isolate the bottle from the egress network); the agent/bottle keeps @@ -167,63 +186,40 @@ The trigger lives in the metering path and the action in the egress/backend plane; **neither touches the supervise proposal queue** (design constraint from #251). -### Host-level SQLite store +### Persistence -**Decision: introduce SQLite now, narrowly.** +Usage, budget state, and enforcement-audit rows live in the **existing** host +SQLite store (PRD 0067) at `~/.bot-bottle/bot-bottle.db`, behind its repository +API — this PRD adds tables, not a store. SQLite is the right fit for the +concurrency this introduces: a *global* token budget decremented by N egress +sidecars in parallel is a read-modify-write race that atomic transactions + WAL +handle for free, and the per-scope precedence rollup plus "sum across all +bottles" is a `GROUP BY` rather than an N-directory rescan. The new tables carry +a `schema_version`-managed migration like the rest of the store. -- **The dependency objection doesn't apply.** `sqlite3` is in the Python stdlib, - so it does not break the AGENTS.md stdlib-first / no-runtime-pip stance — same - category as the hand-rolled `yaml_subset.py`, except the stdlib already ships - the whole engine. -- **It fits the problem.** A *global* token budget decremented concurrently by N - egress sidecars (today `~/.bot-bottle/` already has `state/`, `audit/`, - `queue/` written by parallel bottles) is a read-modify-write race. Over JSON - that means hand-rolled file locking; SQLite gives atomic transactions + WAL for - free. The per-agent/per-bottle precedence rollup plus "sum across all bottles" - is a `GROUP BY`, not an N-directory rescan. -- **It rehearses the cloud swap.** "Wrap operations in an API so we can swap to a - cloud service" maps directly onto a thin repository/DAO over SQLite → Postgres - later. A JSON-file store is a worse rehearsal than SQL. +### Where enforcement runs -**Costs (real but bounded):** a new paradigm in a flat-file repo needs a -`schema_version` table + idempotent startup migrations; SQLite serializes -writers, so WAL mode + `busy_timeout` are required (a non-issue at a handful of -bottles); test fixtures need temp DBs. - -**Scope of the store:** one DB at `~/.bot-bottle/bot-bottle.db` behind a thin -repository API. Only the **new** metering/budget/enforcement-audit ledger lives -there. Existing per-bottle blobs (resume `metadata.json`, transcripts, -Dockerfile overrides, supervise queue) stay on the filesystem — migrating them -now is churn for no benefit and they lack the concurrency/aggregation problem. - -### Host-level controller + dashboard - -A single **host-level controller** owns the meter, budget evaluation, and the -cutoff actions across all bottles (cf. `bot_bottle/cli/supervise.py`'s -cross-bottle view), rather than a per-bottle daemon. v1 ships one host-level -**TUI dashboard** that reads live usage-vs-budget from the SQLite store and -offers on-demand cutoff/stop. The existing supervisor UI should eventually fold -into this same dashboard; this PRD lays the host-level surface it will move to. +Metering, budget evaluation, and the cutoff actions run **out-of-band in the +host-level controller** — cross-bottle, independent of any agent, not a +per-bottle daemon. The controller writes usage/enforcement state that the +already-implemented observability surfaces (dashboard, traffic log) read; this +PRD does not build or modify those surfaces. ## Implementation chunks Ordered, individually mergeable: -1. **SQLite repository foundation.** `~/.bot-bottle/bot-bottle.db`, schema + - `schema_version` migrations, WAL + `busy_timeout`, thin repository API, - temp-DB test fixtures. No behavior wired yet. -2. **Metering at the egress proxy.** Parse authoritative response `usage` +1. **Metering at the egress proxy.** Parse authoritative response `usage` (including SSE final-usage tailing) in the egress addon `response` hook; - write per-bottle / per-provider usage rows to the ledger. -3. **`settings.yml` + budget model.** Host-level `~/.bot-bottle/settings.yml` + write per-bottle / per-provider usage rows to the ledger (new tables in the + existing PRD 0067 store). +2. **`settings.yml` + budget model.** Host-level `~/.bot-bottle/settings.yml` parsed by `yaml_subset.py`; budget precedence (agent → bottle → parent → global) and the `--budget` launch flag. -4. **Forced cutoff + cutoff policy.** Wire the threshold trigger to the +3. **Forced cutoff + cutoff policy.** Wire the threshold trigger to the `cutoff` / `freeze` / `kill` primitives on the egress/backend plane; record enforcement actions to the audit ledger. -5. **Host-level TUI dashboard.** Live usage-vs-budget view + on-demand - cutoff/stop, reading the store. -6. **`count_tokens` pre-flight gate (optional refinement).** Abstract method + +4. **`count_tokens` pre-flight gate (optional refinement).** Abstract method + stdlib estimator default; Anthropic/OpenAI endpoints for built-in claude/codex; optional pre-send block. @@ -235,13 +231,10 @@ Ordered, individually mergeable: stream is interrupted mid-flight? - **Crossing mid-request.** A single response can push usage past budget only *after* it's already been delivered. Is post-hoc cutoff (next request blocked) - sufficient, or is a pre-flight estimator gate (chunk 6) required for v1? + sufficient, or is a pre-flight estimator gate (chunk 4) required for v1? - **Provider name ↔ metered host mapping.** How does the proxy attribute a flow to an agent-provider budget key — by destination host, by bottle identity, or both? - **Parent-bottle budget semantics.** For `bottle extends` (PRD 0025 / 0065) chains, does "parent bottle" mean the manifest parent, the launching bottle, or the full ancestry summed? -- **Dashboard ↔ controller transport (even host-only).** In-process, a local - socket, or polling the SQLite store directly? Picks the seam the future remote - control plane will extend.