docs: narrow egress PRD to metering, budgets, and forced cutoff

Per owner request on PR #285: scope this PRD to the three still-missing
enforcement capabilities and drop what already ships.

- Retitle "Egress control plane" -> "Egress metering, budgets, and
  forced cutoff"; rename file to match.
- Move to already-implemented (out of scope, referenced not rebuilt):
  the egress chokepoint / MITM proxy (PRD 0001/0006/0017), observability
  (host dashboard + egress traffic logging PRD 0055), and the host
  SQLite store (PRD 0067). The "introduce SQLite now" justification
  collapses to "add tables to the existing store."
- Drop the SQLite-foundation and dashboard implementation chunks (both
  done); renumber to metering / settings+budgets / forced cutoff /
  optional count_tokens gate.
- Trim the "& observability plane" framing and the dashboard-transport
  open question; keep metering, budget precedence, settings.yml, and the
  cutoff/freeze/kill design intact.

Issue: #251

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-26 05:57:46 +00:00
parent 2cdedbb7ca
commit b563443605
@@ -1,4 +1,4 @@
# PRD prd-new: Egress control plane — metering, budgets, and forced cutoff # PRD prd-new: Egress metering, budgets, and forced cutoff
- **Status:** Draft - **Status:** Draft
- **Author:** didericis - **Author:** didericis
@@ -7,25 +7,32 @@
## Summary ## Summary
Add an **out-of-band egress enforcement & observability plane**: meter every Add **out-of-band cost enforcement** on top of the egress plane: meter every
agent's token usage at the egress proxy, decrement budgets without the agent's agent's authoritative token usage at the egress proxy, decrement per-scope
cooperation, and forcibly cut a bottle's egress when a budget is exhausted — budgets without the agent's cooperation, and forcibly cut a bottle's egress when
either automatically or on command from a host-level dashboard. The trigger a budget is exhausted. The trigger (usage threshold) and the action (route-drop
(usage threshold) and the action (route-drop / freeze / kill) both live in the / freeze / kill) both run with **no agent in the loop** — this is enforcement,
egress plane and run with no agent in the loop. This is distinct from the distinct from the supervise sidecar (PRD 0013), which is agent-initiated and so
supervise sidecar (PRD 0013), which is agent-initiated and therefore cannot cannot stop a runaway agent.
enforce a cost cutoff on a runaway agent. State (usage ledger, budgets, audit)
moves into a host-level SQLite database behind a thin repository API, the first This builds on infrastructure that already exists and is **not** re-designed
SQL store in an otherwise flat-file repo. here:
- **The egress chokepoint** — the MITM proxy every agent's traffic flows through
(PRD 0001 / 0006 / 0017) — is where metering reads and cutoff acts.
- **The host SQLite store** (PRD 0067) is where usage, budget, and enforcement
state persist, behind its existing repository API.
- **Observability** — the host dashboard and egress traffic logging (PRD 0055) —
already surfaces what agents are doing.
This PRD adds only the three missing pieces: **metering** (authoritative
accounting), **budgets**, and **forced cutoff**.
## Problem ## Problem
bot-bottle can't currently do two things the cost-overrun case demands: bot-bottle can't meter agent token usage or enforce a cost limit. When an agent
crosses a token threshold, there is no way to kill its egress automatically — no
1. **Forced egress shutdown on limit.** When an agent crosses a token human in the loop.
threshold, kill its egress automatically — no human in the loop.
2. **Remote (host-level) management.** Drive agents from a single surface:
see usage, cut egress, stop bottles, to prevent cost overruns.
The existing supervise sidecar (PRD 0013) is **entirely agent-initiated**: every The existing supervise sidecar (PRD 0013) is **entirely agent-initiated**: every
action begins with the agent voluntarily calling an MCP tool and an operator action begins with the agent voluntarily calling an MCP tool and an operator
@@ -36,20 +43,23 @@ it mandatory (#249) would not deliver forced cost-cutoff.
The requirement forces a distinction the current design blurs: The requirement forces a distinction the current design blurs:
- **Plane A — enforcement / observability (this PRD).** System → infrastructure. - **Plane A — enforcement (this PRD).** System → infrastructure. Meter usage,
Meter usage, cut egress on threshold or command, account for cost. cut egress on threshold, account for cost. Out-of-band; independent of the
Out-of-band; independent of the agent. **Unconditional** — an enforcement agent. **Unconditional** — an enforcement plane you can opt out of isn't
plane you can opt out of isn't enforcement. enforcement.
- **Plane B — agent-facing recovery (the existing supervise sidecar).** - **Plane B — agent-facing recovery (the existing supervise sidecar).**
Agent → operator, approval-gated. Useful interactively; meaningless for a Agent → operator, approval-gated. Useful interactively; meaningless for a
headless agent with no operator watching its queue. Remains optional. headless agent with no operator watching its queue. Remains optional.
This PRD builds Plane A. It reframes the "always-on control" invariant of #249 This PRD builds the **enforcement** slice of Plane A — metering, budgets, and
as "the egress control plane is always present" — a more defensible property forced cutoff. The chokepoint that makes interception possible and the
than "every agent runs the agent-facing supervisor." Unsupervised observability surfaces that display usage already exist; what's missing is
(headless/CI/ephemeral) agents stay first-class: still subject to the mandatory turning measured usage into an enforced budget. It reframes the "always-on
meter + kill switch, they simply lack the agent-facing proposal tools they control" invariant of #249 as "the egress enforcement plane is always present" —
couldn't use anyway. a more defensible property than "every agent runs the agent-facing supervisor."
Unsupervised (headless/CI/ephemeral) agents stay first-class: still subject to
the mandatory meter + kill switch, they simply lack the agent-facing proposal
tools they couldn't use anyway.
## Goals / Success Criteria ## Goals / Success Criteria
@@ -61,28 +71,37 @@ couldn't use anyway.
most-specific applicable budget governs. most-specific applicable budget governs.
- When usage crosses a budget, the bottle's configured **cutoff policy** - When usage crosses a budget, the bottle's configured **cutoff policy**
(`cutoff` | `freeze` | `kill`) fires automatically, executed host-side on the (`cutoff` | `freeze` | `kill`) fires automatically, executed host-side on the
egress plane — never via the supervise queue. egress plane — never via the supervise queue. An operator can also trigger the
- An operator can, from a single **host-level TUI dashboard**, see live per-bottle same cutoff on demand through the **existing** host controller/dashboard
usage against budget and command a cutoff/stop on demand. surface (that surface is not built here).
- Host budgets, default cutoff policy, and per-provider limits are declared in a - Host budgets, default cutoff policy, and per-provider limits are declared in a
new host-level `~/.bot-bottle/settings.yml`, parseable by `yaml_subset.py`. new host-level `~/.bot-bottle/settings.yml`, parseable by `yaml_subset.py`.
- All usage, budget state, and enforcement actions persist in a host-level - Usage, budget state, and enforcement actions persist in the **existing** host
SQLite DB behind a thin repository API, so the store can later be swapped for SQLite store (PRD 0067) behind its repository API.
a cross-host cloud service.
## Non-goals ## Non-goals
- **The egress chokepoint — already implemented.** The MITM egress proxy
(PRD 0001 / 0006 / 0017) is the interception point; this PRD hooks metering and
cutoff into it, it does not build or modify the proxy.
- **Observability / dashboard — already implemented.** Live usage display, the
host dashboard, and egress traffic logging (PRD 0055) exist; this PRD writes
the usage/enforcement data they surface but adds no new dashboard and no new
visibility surface.
- **The SQLite store foundation — already implemented (PRD 0067).** This PRD adds
metering/budget/enforcement tables behind the existing repository API; it does
not introduce the store.
- **Remote control / cross-host control plane.** Web + mobile remote control, - **Remote control / cross-host control plane.** Web + mobile remote control,
cross-host budgets, and the authn/transport they require are explicitly cross-host budgets, and the authn/transport they require are explicitly
deferred. v1 is a **host-only TUI** with no remote surface. deferred. This is host-only.
- **Dollar-denominated budgets.** Budgets are token counts keyed by agent - **Dollar-denominated budgets.** Budgets are token counts keyed by agent
provider, not currency. Price tables are out of scope. provider, not currency. Price tables are out of scope.
- **Migrating existing flat-file state into SQLite.** Resume `metadata.json`, - **Migrating existing flat-file state into SQLite.** Resume `metadata.json`,
transcripts, Dockerfile overrides, the supervise queue, and audit logs stay on transcripts, Dockerfile overrides, the supervise queue, and audit logs stay on
the filesystem. Only the *new* metering/budget/enforcement ledger is SQL. the filesystem. Only the *new* metering/budget/enforcement ledger is SQL.
- **Making the supervise sidecar (Plane B) mandatory.** Out of scope here; this - **Making the supervise sidecar (Plane B) mandatory.** Out of scope here; this
PRD is the answer to "what should be unconditional" (Plane A), leaving #249's PRD is the answer to "what should be unconditional" (Plane A enforcement),
Plane-B question open. leaving #249's Plane-B question open.
- **Per-request hard pre-send blocking as the primary mechanism.** The gate is - **Per-request hard pre-send blocking as the primary mechanism.** The gate is
budget-crossing detected at/after metering; a pre-flight estimator (below) is a budget-crossing detected at/after metering; a pre-flight estimator (below) is a
refinement, not the core enforcement path. refinement, not the core enforcement path.
@@ -129,11 +148,10 @@ bottles already use). Four scopes, most-specific wins:
agent → bottle → parent bottle → global (host) agent → bottle → parent bottle → global (host)
``` ```
The global host budget is the highest-priority feature to ship (the cross-host The global host budget is the highest-priority feature to ship; per-agent and
control plane will eventually consume it); per-agent and per-bottle budgets per-bottle budgets override it for finer control. A budget can also be supplied
override it for finer control. A budget can also be supplied **at bottle **at bottle launch** (`--budget` or equivalent), overriding the settings.yml
launch** (`--budget` or equivalent), overriding the settings.yml defaults for defaults for that run. Enforcement evaluates the effective budget as the
that run. Enforcement evaluates the effective budget as the
nearest-defined scope at decrement time. nearest-defined scope at decrement time.
### `~/.bot-bottle/settings.yml` ### `~/.bot-bottle/settings.yml`
@@ -152,8 +170,9 @@ shutdown: cutoff # default cutoff policy: cutoff | freeze | kill
### Forced cutoff and cutoff policy ### Forced cutoff and cutoff policy
On budget exhaustion (or an operator command), the configured per-bottle cutoff On budget exhaustion (or an operator command via the existing surface), the
policy fires. The three policies map onto primitives that already exist: configured per-bottle cutoff policy fires. The three policies map onto
primitives that already exist:
- **`cutoff`** (default) — drop the bottle's `routes.yaml` to empty and reload - **`cutoff`** (default) — drop the bottle's `routes.yaml` to empty and reload
(or isolate the bottle from the egress network); the agent/bottle keeps (or isolate the bottle from the egress network); the agent/bottle keeps
@@ -167,63 +186,40 @@ The trigger lives in the metering path and the action in the egress/backend
plane; **neither touches the supervise proposal queue** (design constraint from plane; **neither touches the supervise proposal queue** (design constraint from
#251). #251).
### Host-level SQLite store ### Persistence
**Decision: introduce SQLite now, narrowly.** Usage, budget state, and enforcement-audit rows live in the **existing** host
SQLite store (PRD 0067) at `~/.bot-bottle/bot-bottle.db`, behind its repository
API — this PRD adds tables, not a store. SQLite is the right fit for the
concurrency this introduces: a *global* token budget decremented by N egress
sidecars in parallel is a read-modify-write race that atomic transactions + WAL
handle for free, and the per-scope precedence rollup plus "sum across all
bottles" is a `GROUP BY` rather than an N-directory rescan. The new tables carry
a `schema_version`-managed migration like the rest of the store.
- **The dependency objection doesn't apply.** `sqlite3` is in the Python stdlib, ### Where enforcement runs
so it does not break the AGENTS.md stdlib-first / no-runtime-pip stance — same
category as the hand-rolled `yaml_subset.py`, except the stdlib already ships
the whole engine.
- **It fits the problem.** A *global* token budget decremented concurrently by N
egress sidecars (today `~/.bot-bottle/` already has `state/`, `audit/`,
`queue/` written by parallel bottles) is a read-modify-write race. Over JSON
that means hand-rolled file locking; SQLite gives atomic transactions + WAL for
free. The per-agent/per-bottle precedence rollup plus "sum across all bottles"
is a `GROUP BY`, not an N-directory rescan.
- **It rehearses the cloud swap.** "Wrap operations in an API so we can swap to a
cloud service" maps directly onto a thin repository/DAO over SQLite → Postgres
later. A JSON-file store is a worse rehearsal than SQL.
**Costs (real but bounded):** a new paradigm in a flat-file repo needs a Metering, budget evaluation, and the cutoff actions run **out-of-band in the
`schema_version` table + idempotent startup migrations; SQLite serializes host-level controller** — cross-bottle, independent of any agent, not a
writers, so WAL mode + `busy_timeout` are required (a non-issue at a handful of per-bottle daemon. The controller writes usage/enforcement state that the
bottles); test fixtures need temp DBs. already-implemented observability surfaces (dashboard, traffic log) read; this
PRD does not build or modify those surfaces.
**Scope of the store:** one DB at `~/.bot-bottle/bot-bottle.db` behind a thin
repository API. Only the **new** metering/budget/enforcement-audit ledger lives
there. Existing per-bottle blobs (resume `metadata.json`, transcripts,
Dockerfile overrides, supervise queue) stay on the filesystem — migrating them
now is churn for no benefit and they lack the concurrency/aggregation problem.
### Host-level controller + dashboard
A single **host-level controller** owns the meter, budget evaluation, and the
cutoff actions across all bottles (cf. `bot_bottle/cli/supervise.py`'s
cross-bottle view), rather than a per-bottle daemon. v1 ships one host-level
**TUI dashboard** that reads live usage-vs-budget from the SQLite store and
offers on-demand cutoff/stop. The existing supervisor UI should eventually fold
into this same dashboard; this PRD lays the host-level surface it will move to.
## Implementation chunks ## Implementation chunks
Ordered, individually mergeable: Ordered, individually mergeable:
1. **SQLite repository foundation.** `~/.bot-bottle/bot-bottle.db`, schema + 1. **Metering at the egress proxy.** Parse authoritative response `usage`
`schema_version` migrations, WAL + `busy_timeout`, thin repository API,
temp-DB test fixtures. No behavior wired yet.
2. **Metering at the egress proxy.** Parse authoritative response `usage`
(including SSE final-usage tailing) in the egress addon `response` hook; (including SSE final-usage tailing) in the egress addon `response` hook;
write per-bottle / per-provider usage rows to the ledger. write per-bottle / per-provider usage rows to the ledger (new tables in the
3. **`settings.yml` + budget model.** Host-level `~/.bot-bottle/settings.yml` existing PRD 0067 store).
2. **`settings.yml` + budget model.** Host-level `~/.bot-bottle/settings.yml`
parsed by `yaml_subset.py`; budget precedence (agent → bottle → parent → parsed by `yaml_subset.py`; budget precedence (agent → bottle → parent →
global) and the `--budget` launch flag. global) and the `--budget` launch flag.
4. **Forced cutoff + cutoff policy.** Wire the threshold trigger to the 3. **Forced cutoff + cutoff policy.** Wire the threshold trigger to the
`cutoff` / `freeze` / `kill` primitives on the egress/backend plane; record `cutoff` / `freeze` / `kill` primitives on the egress/backend plane; record
enforcement actions to the audit ledger. enforcement actions to the audit ledger.
5. **Host-level TUI dashboard.** Live usage-vs-budget view + on-demand 4. **`count_tokens` pre-flight gate (optional refinement).** Abstract method +
cutoff/stop, reading the store.
6. **`count_tokens` pre-flight gate (optional refinement).** Abstract method +
stdlib estimator default; Anthropic/OpenAI endpoints for built-in stdlib estimator default; Anthropic/OpenAI endpoints for built-in
claude/codex; optional pre-send block. claude/codex; optional pre-send block.
@@ -235,13 +231,10 @@ Ordered, individually mergeable:
stream is interrupted mid-flight? stream is interrupted mid-flight?
- **Crossing mid-request.** A single response can push usage past budget only - **Crossing mid-request.** A single response can push usage past budget only
*after* it's already been delivered. Is post-hoc cutoff (next request blocked) *after* it's already been delivered. Is post-hoc cutoff (next request blocked)
sufficient, or is a pre-flight estimator gate (chunk 6) required for v1? sufficient, or is a pre-flight estimator gate (chunk 4) required for v1?
- **Provider name ↔ metered host mapping.** How does the proxy attribute a - **Provider name ↔ metered host mapping.** How does the proxy attribute a
flow to an agent-provider budget key — by destination host, by bottle flow to an agent-provider budget key — by destination host, by bottle
identity, or both? identity, or both?
- **Parent-bottle budget semantics.** For `bottle extends` (PRD 0025 / 0065) - **Parent-bottle budget semantics.** For `bottle extends` (PRD 0025 / 0065)
chains, does "parent bottle" mean the manifest parent, the launching bottle, chains, does "parent bottle" mean the manifest parent, the launching bottle,
or the full ancestry summed? or the full ancestry summed?
- **Dashboard ↔ controller transport (even host-only).** In-process, a local
socket, or polling the SQLite store directly? Picks the seam the future remote
control plane will extend.