PRD: canonical tamper-evident audit-event schema (#487) #495

Closed
didericis wants to merge 5 commits from didericis/prd-audit-event-schema AGit into main
+706
View File
@@ -0,0 +1,706 @@
# PRD prd-new: Canonical tamper-evident audit-event schema and local query contract
- **Status:** Draft
- **Author:** didericis-claude
- **Created:** 2026-07-26
- **Issue:** #487
## Summary
bot-bottle already emits security- and provenance-relevant events from
several producers — supervise operator decisions (PRD 0013's
`AuditStore`), egress allow/block enforcement, git-gate push decisions,
control-plane token minting, and (next) host-controller lifecycle
transitions — but each writes its own shape to its own sink. There is no
shared envelope, no tamper-evidence, and no single place to search. Local
incident reconstruction means grepping several stores that don't agree on
field names, timestamps, or how a bottled agent is identified.
This PRD defines **one canonical audit-event contract** every producer
emits into:
1. A **versioned envelope** — schema version, event id, event/observed
timestamps, host-attributed identity (`bottle`/`bottled_agent`/
`activation`), provenance (`manifest_digest` — the manifest is the
policy — and `engine` = bot-bottle version/SHA),
host-observed `actor`/`action`/`resource`/`outcome`, correlation/
causation ids, a sensitivity class, a typed payload, and an explicit
**trust boundary** between host-supplied and agent-claimed fields.
2. **Canonical JSON serialization + a per-writer hash chain** (with
normative test vectors), so any deletion, edit, or reorder of a past
record breaks the chain and is detectable offline.
3. An **append-only JSONL journal as the source of truth**, with a
**rebuildable SQLite index** and a local `audit query`/`verify` surface
— no paid platform, no network dependency.
4. An **initial event registry** covering lifecycle, host-controller,
supervise decision, egress (request/decision/cutoff/anomaly), git-gate
and signed-commit (#480), auth/authz, and audit self-events — each with
its trusted-vs-claimed fields and redaction rules.
5. A **stable export projection** (CloudEvents / OpenTelemetry Logs) and
the **#324 delivery contract** (payload, `(epoch, seq)` cursor, dedup,
backpressure, retention ordering).
It is explicitly scheduled to land **immediately after the host
controller (#468)** so the host controller's lifecycle transitions are the
first producer wired onto the new contract (per the directive on #487).
## Problem
Audit infrastructure is fragmented across #468, #324, and #480 with no
shared schema. Concretely:
- **No shared envelope.** `supervise_audit_entries` (PRD 0013) has
`timestamp, bottle_slug, component, operator_action, ...`. The egress
proxy and git-gate log their own ad-hoc lines. There is no common
`event_id`, `event_type`, or version, so cross-producer correlation
("what did bottled agent X do between its start and this rejected push?") is
manual and lossy.
- **No tamper-evidence.** The audit store is a plain SQLite table. Anyone
who can write the DB can delete or rewrite a row and leave no trace.
Audit that an attacker (or a buggy agent) can silently rewrite is not
audit.
- **Trusted and untrusted data are mixed.** A bottled agent is attributed by
**source IP → slug** at the gateway (host-supplied, trustworthy). An
agent can also *claim* things about itself in a tool call
(agent-claimed, adversarial). Today nothing in the record marks which is
which, so a reader can be misled by an agent-supplied field that looks
authoritative.
- **No local search.** Reconstructing an incident means reading multiple
sinks with different schemas. There is no query contract and no promise
that the index can be rebuilt from the journal if it drifts or is lost.
- **No redaction rule.** Nothing prohibits a producer from writing a raw
token or secret into an audit record, which would turn the audit log
itself into a credential store.
## Goals / Success Criteria
- A single `AuditEvent` envelope type, versioned, that every producer
emits. The trust boundary is **one `untrusted` region**: everything
outside it is host-established and trusted (source-IP → bottled-agent
attribution, host wall-clock, producer identity, chain metadata);
`untrusted` is the sole place anything an agent or a remote claimed may
go. The boundary is structural, not a convention.
- **Canonical serialization** (`sort_keys`, `(",", ":")` separators,
UTF-8, `ensure_ascii=False`) is defined once and reused, so the same
logical event always hashes identically across producers and hosts.
- Each writer maintains a **hash chain**: `hash = sha256(prev_hash ||
canonical(event))`. Deleting or editing any past record breaks every
subsequent link; a standalone verifier detects the break offline with no
secret material.
- The **JSONL journal is the source of truth**; the **SQLite index is
fully rebuildable** from it (`audit rebuild` reconstructs the DB and
re-verifies the chain).
- **Local query** works with no paid platform and no egress: filter by
bottled agent, event type, time range, and producer, and follow a bottled
agent's events in order.
- A **redaction rule** is enforced at the envelope boundary: known
credential-shaped fields are rejected/redacted before a record is
written; the writer refuses raw secrets rather than storing them.
- The envelope **projects onto the OpenTelemetry Logs data model and a
CloudEvents JSON envelope** by field re-mapping alone (no reformat),
preserving the trust boundary and carrying the integrity fields — per
#487's export/interop requirement. (The export adapters are follow-up;
the *schema* must make them a re-map.)
- The **host controller (#468)** emits `lifecycle.*` events through this
contract as the first consumer; existing supervise/egress producers are
migrated behind the same envelope without changing operator-facing
behavior.
## Non-goals
- **Cross-host aggregation / shipping.** This PRD makes each host's journal
canonical and correlatable *by construction* (stable ids, hash chain),
but the transport that merges multiple hosts into one timeline is a
follow-up (#324). The schema is designed so that merge is a later append,
not a reformat.
- **Cryptographic signing / external anchoring.** Hash-chaining gives
tamper-**evidence** (you can detect edits), not tamper-**resistance**
against an attacker who can rewrite the whole chain. Per-writer signing
keys and periodic external anchoring are a follow-up; the chain-head hash
is the seam they attach to.
- **Real-time alerting / SIEM rules.** Query is local and pull-based here.
- **Retention / rotation policy.** Journal rotation and TTL are operator
policy, tracked separately; the format must survive rotation (chain head
carried across segments) but this PRD does not set the schedule.
- **Replacing PRD 0013's operator queue.** The supervise proposal/response
queue is unchanged; only its terminal *audit* record is re-emitted onto
the new envelope.
Review

why not put this in trusted?

why not put this in trusted?
Review

Agreed — it's host-supplied, so it doesn't belong outside trusted. Moved ts_wall/ts_mono and producer into the trusted block; only chain metadata (v/id/type/epoch/seq/prev/hash) stays at the top level.

Agreed — it's host-supplied, so it doesn't belong outside `trusted`. Moved `ts_wall`/`ts_mono` **and** `producer` into the `trusted` block; only chain metadata (`v/id/type/epoch/seq/prev/hash`) stays at the top level.
Review

Actually let's change this a bit: everything outside of "untrusted" is trusted, (including epoc, seq) so we should just remove the "trusted" subsection.

Actually let's change this a bit: everything outside of "untrusted" is trusted, (including epoc, seq) so we should just remove the "trusted" subsection.
Review

Done — removed the trusted sub-block. The envelope is now flat: v/id/type/epoch/seq/prev/hash/producer/host/bottled_agent/ts_wall/ts_mono all sit at the top level and are trusted by construction; untrusted is the one region for agent/remote-claimed data. The rule is now "a field is trusted unless it's deliberately placed under untrusted," and the construction API takes trusted fields positionally + a single untrusted mapping, so a producer can't accidentally emit a top-level field that looks authoritative but isn't. (Head 5a9428c.)

Done — removed the `trusted` sub-block. The envelope is now flat: `v/id/type/epoch/seq/prev/hash/producer/host/bottled_agent/ts_wall/ts_mono` all sit at the top level and are trusted **by construction**; `untrusted` is the one region for agent/remote-claimed data. The rule is now "a field is trusted unless it's deliberately placed under `untrusted`," and the construction API takes trusted fields positionally + a single `untrusted` mapping, so a producer can't accidentally emit a top-level field that looks authoritative but isn't. (Head `5a9428c`.)
## Design
Review

How does this account for host controller restarts?

How does this account for host controller restarts?
Review

Handled explicitly now. Ordering is (epoch, seq), not the clock: on restart the writer reads the last journal line, adopts its hash as the next prev (chain stays continuous), bumps epoch, and resets seq. ts_mono orders only within an epoch. See the new Single writer; ordering across restarts section.

Handled explicitly now. Ordering is `(epoch, seq)`, not the clock: on restart the writer reads the last journal line, adopts its `hash` as the next `prev` (chain stays continuous), bumps `epoch`, and resets `seq`. `ts_mono` orders only *within* an epoch. See the new **Single writer; ordering across restarts** section.
### The envelope
One dataclass, `AuditEvent`, serialized to a JSON object with a small,
Review

should be precise/label this a "bottled-agent"

should be precise/label this a "bottled-agent"
Review

Renamed to bottled_agent — both the trusted.bottled_agent field and the lifecycle.bottled_agent_* leaves, so one term for the subject everywhere.

Renamed to `bottled_agent` — both the `trusted.bottled_agent` field and the `lifecycle.bottled_agent_*` leaves, so one term for the subject everywhere.
stable top level:
```
{
// ---- schema + integrity (host-owned) ----
"v": 1, // schema version — bumped only on a breaking change
"id": "<uuid4>", // globally unique event id; stable across export/replay (dedup key)
"type": "egress.decision", // dotted event type from the registry
"epoch": 7, // writer-boot counter, bumped once per host-controller (writer) start
"seq": 1287, // monotonic sequence within this epoch (gap-detectable)
"segment": "20260726T000000Z", // journal segment id (rotation boundary); chain continues across segments
"prev": "<hex>", // hash of the previous record in the chain ("" for a segment genesis)
"hash": "<hex>", // sha256(prev + canonical(this event with hash=""))
// ---- timestamps (host-owned) ----
"ts_event": "2026-07-26T18:22:04.061Z", // when the underlying event occurred at the boundary
"ts_recorded": "2026-07-26T18:22:04.113Z", // when the single writer appended it (authoritative)
"ts_mono": 90142.55, // monotonic secs since this epoch's boot (intra-epoch ordering only)
// ---- attribution + provenance (host-established) ----
"producer": "egress", // host component that emitted the event
"host": "mac-studio-1",
"engine": "bot-bottle/0.1.0+abc1234", // bot-bottle version + git SHA of the enforcing host code
"bottle": "amber-fox", // bottle (container/VM) identity
"bottled_agent": "amber-fox-12", // bottled-agent slug from source-IP attribution (null for host-level events)
"activation": "01J8Z...", // activation id: one run/session of the bottled agent (null if n/a)
"manifest_digest": "sha256:9f2…", // digest of the manifest — which IS the policy (egress routes etc.); null if n/a
// ---- semantics: host-observed facts of what happened ----
"actor": "bottled-agent:amber-fox-12", // who acted, as a host-attributed identity
"action": "egress.connect", // what was attempted / done
"resource": "registry.npmjs.org:443", // what it acted on, as observed at the boundary
"outcome": "blocked", // host-decided result: allowed|blocked|deferred|success|failure
"sensitivity": "security", // classification: normal|security|restricted (drives redaction + export)
// ---- correlation (host-assigned) ----
"correlation_id": "flow-9c2a…", // groups a related flow (request → decision → cutoff)
"causation_id": "<event id>", // the event that directly caused this one ("" if root)
// ---- typed, trusted, event-specific payload (shape fixed per type in the registry) ----
"payload": {
"route_id": 4,
"detector": "token_patterns"
},
// ---- the ONLY untrusted region: agent- or remote-claimed data ----
"untrusted": {
"reason": "npm install needs registry.npmjs.org" // the agent's stated justification
}
}
```
The **trust boundary is a single region, not a split.** Everything outside
`untrusted` is trusted by construction — the host established it: schema and
chain metadata, both timestamps, the attribution/provenance fields
(source-IP → `bottled_agent`/`bottle`/`activation`, `manifest_digest`,
`engine`), the host-observed semantics
(`actor`/`action`/`resource`/`outcome`), the host-assigned correlation ids,
and the typed `payload`. `untrusted` is the **one** place anything an agent
or a remote claimed may go (e.g. the agent's free-text `reason`). A reader
(or a future policy engine) trusts every field outside `untrusted` for
attribution and treats `untrusted.*` — and only `untrusted.*` — as
adversarial claims.
Framing it as "one untrusted region, everything else trusted" removes the
mistake where a producer forgets to mark a claimed field: a field is
trusted unless it is deliberately placed inside `untrusted`. The
Review

annoying, but should be consistent: bottled_agent_start, bottled_agent_stop, bottled_agent_crash

annoying, but should be consistent: bottled_agent_start, bottled_agent_stop, bottled_agent_crash
Review

Done: lifecycle.bottled_agent_start / _stop / _crash, matching the bottled_agent field name.

Done: `lifecycle.bottled_agent_start` / `_stop` / `_crash`, matching the `bottled_agent` field name.
construction API enforces this — producers pass trusted fields explicitly
and hand all agent/remote-claimed data as the single `untrusted` mapping,
so there is no way to emit a top-level field that *looks* authoritative but
isn't.
**Two timestamps** because they answer different questions and can diverge
under backpressure: `ts_event` is when the thing happened at the boundary
(the proxy saw the connect, the gate saw the push); `ts_recorded` is when
the single writer durably appended it. Ordering and the chain use
`(epoch, seq)`, never either wall clock. Both are host-set — a bottled
agent never supplies a timestamp.
**The manifest *is* the policy.** bot-bottle has no separate policy
artifact — a bottled agent's egress routes and other constraints are
declared in its manifest (`bot_bottle/manifest/egress.py`), so
`manifest_digest` already pins the ruleset in force; there is no distinct
`policy_version`. Given a fixed manifest, the only other thing that can
change a decision's outcome is the enforcing code — captured by `engine`
didericis marked this conversation as resolved
Review

Would be nice if we use the same credential shape detectors we use on egress, but not sure whether it's worth it or if that would be brittle/concerns would diverge. Evaluate how reasonable that is.

Would be nice if we use the same credential shape detectors we use on egress, but not sure whether it's worth it or if that would be brittle/concerns would diverge. Evaluate how reasonable that is.
Review

Evaluated: reuse the deterministic detectors in gateway/egress/dlp_detectors.py (scan_token_patterns / redact_tokens, plus scan_known_secrets) — pure-Python and already the project's source of truth for credential shapes, so wire-egress and journal-redaction can't drift. Excluded scan_entropy: it's tuned for large streamed bodies and would false-positive on the short high-entropy values audit records legitimately carry (hashes, uuids, base64 ids), redacting the very fingerprints we want to keep. So: share the deterministic detectors, keep entropy egress-only.

Evaluated: reuse the deterministic detectors in `gateway/egress/dlp_detectors.py` (`scan_token_patterns` / `redact_tokens`, plus `scan_known_secrets`) — pure-Python and already the project's source of truth for credential shapes, so wire-egress and journal-redaction can't drift. **Excluded `scan_entropy`**: it's tuned for large streamed bodies and would false-positive on the short high-entropy values audit records legitimately carry (hashes, uuids, base64 ids), redacting the very fingerprints we want to keep. So: share the deterministic detectors, keep entropy egress-only.
(bot-bottle version + git SHA). So two `egress.decision` records with the
same `resource` but different `outcome` are explained by exactly one of:
different `manifest_digest` (the rules changed) or different `engine` (the
enforcer changed). Runtime operator overrides (a supervise `egress-allow`)
are themselves audit events, so the effective ruleset at any instant is
`manifest_digest` plus the logged, approved deltas — reconstructable from
the chain, not from a version stamp.
**Optionality.** `bottle`/`bottled_agent`/`activation`, `manifest_digest`,
and `payload`/`untrusted` are `null`/absent for events that have no such
subject (a host-level `hostctl.*` or `audit.*` event has no bottled agent).
Absent ≠ empty: a reader distinguishes "no subject" from "unknown". `id`,
`type`, the chain fields, both timestamps, `producer`, `host`, `engine`,
`actor`, `action`, `outcome`, and `sensitivity` are always present.
#### Trust provenance of every common field
| Field | Trust | Set by |
|---|---|---|
| `v` `id` `type` `epoch` `seq` `segment` `prev` `hash` | trusted | the single writer |
| `ts_event` | trusted | emitting host component (boundary) |
| `ts_recorded` `ts_mono` | trusted | the single writer |
| `producer` `host` `engine` | trusted | the single writer |
| `bottle` `bottled_agent` `activation` | trusted | gateway source-IP → slug attribution |
| `manifest_digest` | trusted | control plane (the manifest = the policy in force) |
| `actor` `action` `resource` `outcome` | trusted | host component that observed/decided it |
| `sensitivity` | trusted | registry default for `type`, overridable up (never down) by the producer |
| `correlation_id` `causation_id` | trusted | the single writer (assigned as it threads the flow) |
| `payload.*` | trusted | emitting host component (shape fixed per `type`) |
| `untrusted.*` | **claimed** | copied verbatim from a bottle / gateway / forge / remote |
Every registry entry (below) restates, per event type, which `payload`
Review

Single writer, definitely

Single writer, definitely
Review

Locked in as decided — single writer, host controller owns it. Removed the equivocation; per-producer chains are noted only as a future scaling path.

Locked in as decided — single writer, host controller owns it. Removed the equivocation; per-producer chains are noted only as a future scaling path.
keys are required and names any `untrusted` keys it carries — so "trusted
vs claimed" is explicit for every event-specific attribute, not just the
common ones.
### Canonical serialization + hash chain
Serialization is defined once (extends the existing `sha256_hex` /
`util.py` helpers):
```
def canonical(event: dict) -> str:
return json.dumps(event, sort_keys=True, separators=(",", ":"),
Review

Yes

Yes
Review

Adopted — the rotated-out head becomes the new segment's genesis prev, so the verifier still trusts the current head across a rotation. Folded into the retention follow-up.

Adopted — the rotated-out head becomes the new segment's genesis `prev`, so the verifier still trusts the current head across a rotation. Folded into the retention follow-up.
ensure_ascii=False)
```
The `hash` field is computed over the canonical form of the event **with
`hash` set to `""`**, prefixed by the previous record's hash:
```
digest = sha256_hex(prev_hash + canonical({**event, "hash": ""}))
```
`prev` is the prior record's `hash`; a segment genesis uses `prev = ""`.
Two exact rules pin the bytes so the chain is reproducible anywhere:
1. **Serialize the record with its own `hash` field set to `""`** (present,
empty), never omitted — the key set is identical before and after
hashing.
2. **Digest = `sha256_hex(prev + canonical(record_with_empty_hash))`**,
where `prev` is the previous record's `hash` string (`""` at genesis),
`+` is string concatenation, and `canonical` is the function above.
`ts_mono`, being a float, is serialized by Python's shortest-round-trip
`repr` via `json.dumps`; producers therefore emit it as a JSON number
they do not post-process. (All other fields are strings/ints/objects,
which serialize unambiguously.)
Editing or deleting record *n* changes its hash, so record *n+1*'s `prev`
no longer matches — the break is local and names the tampered record.
Verification needs only the journal itself (no keys), so it runs offline
and in CI.
#### Test vectors (normative)
Two records, reduced to the chain-relevant fields, demonstrate the exact
serialization and linkage. An implementation is conformant iff it
reproduces these bytes and hashes.
```
# Record 0 — segment genesis (prev = "")
canonical(record0, hash=""):
{"hash":"","id":"11111111-1111-4111-8111-111111111111","prev":"","seq":0,"type":"audit.segment_open"}
hash0 = sha256("" + canonical) =
942ea5729bcac6efdbdea942396bfa574ab0d6ebf5615402595359422f2aeb83
# Record 1 — chains onto record 0 (prev = hash0)
canonical(record1, hash=""):
{"hash":"","id":"22222222-2222-4222-8222-222222222222","prev":"942ea5729bcac6efdbdea942396bfa574ab0d6ebf5615402595359422f2aeb83","seq":1,"type":"lifecycle.bottled_agent_start"}
hash1 = sha256(hash0 + canonical) =
bc082347680405fee50b60a9c304611aa026950b15d869b7e3ae56e1c451b856
# Tamper check: flip record0.type → recompute →
# 4553eda647f33f0c608cfea44be28efbbaca45ed30b873fcbd4405fa5ce737ed
# which no longer equals record1.prev (942ea5…) — the break is detected at record1.
```
The implementation PR ships these plus full-envelope vectors (every field
populated, and a redaction case) as committed fixtures, so a schema-version
bump that changes the bytes fails a golden test loudly.
### Ordering, idempotency, and duplicate handling
- **Ordering.** `(epoch, seq)` is a strict total order per host and, because
the writer is single, a strict order per bottle/activation within that
host — satisfying "at least strict causal order per activation/bottle".
`causation_id` records the explicit cause edges (a DAG) on top of the
total order, so a consumer can reconstruct request → decision → cutoff
even if unrelated events interleave between them.
- **Idempotency.** `id` is the idempotency key. A producer that retries an
emit (e.g. after a writer restart mid-handoff) **reuses the same `id`**;
the writer drops a second append bearing an `id` already present in the
current segment's in-memory set, and the index `UPSERT`s by `id`, so a
duplicate never double-counts or forks the chain.
- **Deduplication downstream.** Because `id` is stable across export and
replay, #324's cursor replay and any cross-host merge dedup on `id` — no
consumer needs to invent a second identity.
### Behavior across rotation, restart, import, truncation
- **Rotation.** At a segment boundary the writer opens a new segment file,
sets its `segment` id, and carries the rotated-out segment's head as the
new genesis `prev` — so the chain is continuous *across* segments while
each file stays independently openable. `verify` walks segments in order
and checks the head-to-genesis link at each seam.
- **Restart.** Covered above: read last line → adopt its `hash` as `prev`,
bump `epoch`, reset `seq`. The chain never restarts even though the
counters do.
- **Import.** `audit import <segment>` appends an externally supplied
segment (e.g. recovered from another host or a backup). Import verifies
the incoming chain in isolation first, then links it only if its genesis
`prev` matches a known head or is explicitly grafted; imported records
keep their original `id` (dedup) and are marked with their origin host so
attribution is not laundered.
- **Truncation.** A crash can leave a partial final line; `verify` reports
it as `truncated-tail` (recoverable — replay resumes from the last intact
record). A chain that ends before a persisted head, or a missing interior
`seq`, is reported as `gap`/`missing-suffix` (evidence of deletion, not a
clean crash). The two are distinguished so an operator can tell "power
loss" from "someone trimmed the log".
### Single writer; ordering across restarts
**Decided: one writer per host** (reviewed — the host controller owns it).
Producers hand events to the host controller, which is the sole appender,
so the chain has one well-defined total order and one `seq`/`epoch`
counter. This ties audit availability to the host controller being up,
which is acceptable because the host controller already gates every
lifecycle transition; per-producer chains are noted only as a future
scaling path, not built now.
**Restarts** are handled by the chain, not the clock. `ts_mono` resets to
~0 on every writer start, so it orders events only *within* one boot. On
start the writer:
1. reads the last line of the journal, adopts its `hash` as the next
record's `prev` (the chain is continuous across the restart), and
2. bumps `epoch` (persisted alongside the chain head) and resets `seq` to
0 for the new boot.
Total order is therefore `(epoch, seq)` — monotonic across restarts by
construction — with `ts_event`/`ts_recorded` for human reading and `ts_mono` for
sub-second ordering inside an epoch. A crash mid-append truncates at most
the last (partial) line; the verifier flags it and replay resumes from the
last intact record.
### Journal (source of truth) + SQLite index (rebuildable)
- **Journal:** one append-only JSONL file per host (path from `paths.py`,
alongside `host_db_path()`), one canonical event per line, opened
`O_APPEND`. This is authoritative.
- **Index:** a new `audit_events` table via the existing `DbStore` /
`TableMigrations` machinery. It is a **derived cache, not a second source
of truth**: `audit rebuild` truncates and replays the journal,
re-verifying the chain as it goes, so a deleted or drifted DB is
regenerated from the journal with no data loss. On a `verify` failure
during rebuild it stops and reports rather than indexing past a break.
(This supersedes the free-standing `supervise_audit_entries` table, which
becomes a producer onto the new index.)
**Indexable fields** (columns + indices): `ts_event`, `ts_recorded`,
`type`, `host`, `bottle`, `bottled_agent`, `activation`, `actor`,
`outcome`, `sensitivity`, `correlation_id`, `causation_id`, plus two
event-specific projections promoted out of `payload` for query —
`repository` and `commit_sha` (populated for `forge.*`/`commit.*`, null
otherwise). The full canonical record is stored verbatim in a `raw` column
so the index never loses fidelity to the journal.
**Local query surface** — `audit query`, no egress, no paid platform:
```
audit query \
[--since T] [--until T] [--type egress.*] [--host H] [--bottle B] \
[--activation A] [--agent SLUG] [--actor ID] [--outcome blocked] \
[--repository R] [--correlation-id C] [--commit SHA] \
[--follow BOTTLE] # a bottled agent's events in (epoch, seq) order
[--json | --table]
audit verify [--segment S] # offline chain check; exit non-zero on any break
audit rebuild # drop + replay journal → index
audit import <segment> # graft an external segment (see above)
```
Type filters accept a `group.*` glob. A read-only local HTTP endpoint
mirrors the same filters for the future review console; both are pure reads
over the index and can never mutate the journal.
### Event registry (initial)
Dotted `type` names, grouped. The registry is a table mapping each type to
its required `payload` keys, its `untrusted` keys (if any), a default
`sensitivity`, and its correlation behavior — so producers and the verifier
agree on shape and "trusted vs claimed" is pinned per type. Initial
coverage (the issue's mandated set):
| Group / type | Producer | Required `payload` (trusted) | `untrusted` | Default sensitivity |
|---|---|---|---|---|
| **lifecycle.*** — `bottled_agent_start` / `_stop` / `_crash` | host-controller (#468) | `manifest_digest`, `exit` (for stop/crash) | — | normal |
| **hostctl.*** — `broker_launch`, `broker_teardown`, `broker_reject` | host-controller (#468) | `op`, `request_digest` | — | security |
| **decision.*** — `proposed`, `resolved` | supervise | `tool`, `operator_action`, `justification`, `diff_digest` | `agent_rationale` | security |
| **egress.*** — `request`, `decision`, `cutoff`, `anomaly` | egress proxy | `route_id`, `detector` (on match), `bytes` (cutoff) | `reason`, `target_claimed` | security |
| **forge.*** — `push_accepted`, `push_rejected`, `pr_opened` | git-gate | `repository`, `ref`, `gitleaks_result` | `title`, `description` | security |
| **commit.signed** (#480) | git-gate | `repository`, `commit_sha`, `activation_key_id`, `signature_ref` | `commit_message` | security |
| **auth.*** — `token_minted`, `token_rejected`, `authz_denied` | control plane | `role`, `token_id`, `reason_code` | — | security |
| **audit.*** — `segment_open`, `verify_failed`, `truncation_detected`, `export_failed` | audit writer/verifier | `segment`, `detail` | — | security |
Notes:
- **`egress.request` vs `egress.decision`** share a `correlation_id`; the
`decision`'s `causation_id` points at the `request`, and a later `cutoff`
chains onto the `decision` — so a flow is reconstructable.
- **`audit.*` self-events** make the audit subsystem audit itself: a failed
verification, a detected truncation, or a dropped export is itself a
chained, tamper-evident record — you cannot silence the alarm without
breaking the chain that carries it.
- **Free-text and remote-echoed fields are always `untrusted`** (`reason`,
`agent_rationale`, PR `title`/`description`, `target_claimed`), because
they originate in the bottle or a remote response; the host-observed
counterpart (`resource`, `outcome`, `gitleaks_result`) is the trusted
fact.
**Sensitivity + redaction per type.** Every type's default `sensitivity`
is listed above; a producer may raise it (never lower it). `restricted`
events keep their `payload` in the journal but the export projection ships
only the envelope + a payload digest unless the consumer is authorized —
so a `security`/`restricted` record is still counted and correlated
downstream without leaking its body. The credential-shape redaction rules
(next) apply to **every** type regardless of sensitivity.
**Schema evolution & backward-compatible readers.** The registry is
append-only: **adding** a type, an optional `payload` key, or an
`untrusted` key does **not** bump `v`; readers ignore unknown fields
(forward-compatible) and treat absent optional fields as `null`.
**Removing** or **re-typing** a field, or making an optional field
required, bumps `v`. A reader declares the max `v` it understands and
refuses to *interpret* a higher-`v` record, but the **verifier is
version-agnostic** — the hash covers whatever fields exist, so chain
integrity is checkable across versions without understanding semantics.
Every `v` bump ships a migration note and updated golden vectors.
### Redaction rule
Redaction runs at the envelope boundary, before a record is written, in two
layers:
1. **Key deny-list (structural).** A field whose *key* matches a known
credential shape (`token`, `secret`, `password`, `authorization`,
`*_key`) is refused — the producer must pass a reference (a token *id*
or `sha256` fingerprint), never the raw value. `auth.token_minted`
therefore records the token id and role, not the JWT. This is the
primary guard: it is cheap, deterministic, and catches the intended
mistake (a producer stuffing a credential into a named field).
2. **Value scan — reuse the egress DLP detectors.** Per review, the value
layer reuses the *same* deterministic credential-shape detectors the
egress proxy already ships:
`bot_bottle/gateway/egress/dlp_detectors.py` —
`scan_token_patterns` / `redact_tokens` (and `scan_known_secrets` for
host-known secret material). They are pure-Python, mitmproxy-free, and
already the project's source of truth for "what a leaked credential
looks like," so a single detector set governs both what may leave over
the wire and what may land in the journal — they can't drift apart.
**Scoped deliberately:** only the pattern/known-secret detectors are
reused, **not** `scan_entropy`. Entropy scoring is tuned for large
streamed request bodies; on the short, high-entropy structured values an
audit event legitimately carries (hashes, uuids, base64 ids) it would
false-positive and start redacting the very fingerprints the log needs.
So the shared layer is the deterministic detectors; entropy stays an
egress-only concern. (This is the "evaluate how reasonable that is" from
review: reuse the deterministic detectors — yes; share the entropy
heuristic — no.)
On a value-layer match the default is **redact** (scrub to a placeholder
and keep the event) rather than drop, so a producer bug can never make an
audit event vanish; the key deny-list stays a hard refusal because a
credential in a named field is always a producer bug worth surfacing.
**No raw-payload capture by default.** The envelope carries *decisions and
metadata*, not traffic. Prompts, model responses, request/response bodies,
and file contents are **not** recorded unless a producer opts a specific,
reviewed field in — and such a field is `untrusted` and subject to both
redaction layers. This keeps the audit log from becoming a covert copy of
the very data the sandbox exists to contain (the issue's "unsafe payload
capture" non-goal).
### Export / interoperability (CloudEvents, OpenTelemetry Logs)
#487 requires the envelope to map onto the **OpenTelemetry Logs data
model** and/or a **CloudEvents JSON** envelope *without losing integrity or
attribution semantics*. The flattened shape (one `untrusted` region,
everything else trusted at top level) does **not** conflict with either — it
maps *more* cleanly than a nested `trusted`/`untrusted` pair would, because
both target models expect a flat set of top-level fields plus one payload
subtree.
**CloudEvents.** Context attributes MUST be scalar simple types — a map
cannot be a context attribute — so a nested `trusted` block would have had
to be flattened for CloudEvents anyway. Our flat top level maps directly:
`id`→`id`, `type`→`type`, `producer`+`host`→`source`,
`bottled_agent`→`subject`, `ts_event`→`time` (`ts_recorded` as an extension); the integrity/chain fields
(`epoch`, `seq`, `prev`, `hash`, `v`) ride as **extension attributes**
(scalars — legal). The `untrusted` map goes in `data`. Only mechanical
transform needed: extension attribute names must be lowercase-alphanumeric,
so `bottled_agent`/`ts_mono`/etc. are renamed at export (e.g. a
`botbottle`-prefixed form) — a naming rule, not a schema conflict.
**OpenTelemetry Logs.** `ts_event`→`Timestamp`, `ts_recorded`→`ObservedTimestamp`; `type`→the `event.name`
attribute; the flat trusted fields → `Attributes` under a `botbottle.*`
namespace (`botbottle.bottled_agent`, `botbottle.producer`,
`botbottle.chain.hash`, …); `untrusted.*` → `Attributes` under
`botbottle.untrusted.*` (or `Body`). OTel attributes are a dotted map that
happily carries the nested subtree.
**Attribution is preserved** precisely because the boundary is now
structural: on export, top-level fields become trusted context/attributes
and the `untrusted` subtree stays a single, clearly-named region — so a
downstream consumer still sees exactly which fields an agent claimed.
Nothing agent-claimed is promoted to a trusted-looking position.
**Integrity has one deliberate caveat.** CloudEvents/OTel are
representation envelopes with their own (or no) canonicalization; `hash`
and `prev` are computed over **our** canonical JSON, not over the exported
form. So the chain fields travel *as data* for reference, but
tamper-evidence is always verified against the **native journal** (the
source of truth) — never re-derived from an exported CloudEvents/OTel
record, whose key ordering / number formatting the exporter may change.
Export is thus a lossless-for-attribution **projection** that carries the
integrity fields along; verification stays on the canonical journal. This
satisfies "without losing integrity or attribution semantics": both are
carried, neither is *relied upon* in the foreign format.
The export adapters themselves are follow-up implementation — this PRD
fixes the *schema* so that projection is a field re-map, never a reformat.
#### The #324 delivery contract (payload, cursor, backpressure)
#324 transports events off-box; it must not invent a second envelope. This
PRD fixes the contract it depends on:
- **Payload.** #324 ships the **native canonical record verbatim** (the
exact bytes the hash covers), optionally wrapped in the CloudEvents
projection whose `data` *is* that record. Either way the integrity fields
travel intact and the receiver can verify against the same bytes.
- **Cursor.** The export cursor is `(epoch, seq)` (equivalently the last
exported `hash`). It advances **only on acknowledgement**, so delivery is
at-least-once and gap-free; a crash re-sends from the last acked cursor.
- **Idempotency / replay.** Dedup is on `id` (stable across replay), so
at-least-once delivery is safe — the receiver collapses re-sends.
- **Backpressure.** The outbox is the journal itself plus a cursor; when
the endpoint is slow the cursor simply lags — the writer never blocks on
export, and audit never applies backpressure to the data plane it
records.
- **Retention interaction.** Retention/rotation **must not** prune a
segment whose records are still behind the export cursor; the reaper
honors `min(cursor)` across all configured consumers. (The schedule
itself stays the retention follow-up; this is the *ordering* constraint
that follow-up must respect.)
#### #480 signed-commit attribution maps in without weakening it
#480 binds a commit's bytes to a per-activation signing key. It maps to the
`commit.signed` event: `payload` carries `repository`, `commit_sha`,
`activation_key_id`, and a `signature_ref` (the detached-signature
location or its digest) — **not** the private key and not a re-derived
signature. The audit event therefore *references and timestamps* #480's
existing byte-to-activation-key proof inside the tamper-evident chain; it
does not re-implement or replace it, so #480's guarantee is unweakened —
the signature still verifies against the commit bytes independently, and
the audit record adds only "this binding was observed at this point in the
chain". The trusted `actor`/`activation` fields and the `commit_sha`
payload are host-observed at the gate, so attribution cannot be forged by
the committing agent.
## Implementation chunks
1. **(this PR — PRD only.)** The contract above. No code; scheduled to land
right after #468.
2. **Envelope + canonical + chain core.** `AuditEvent` dataclass,
`canonical()`, chain hashing, and the single-writer journal appender in
`bot_bottle/store/` (reusing `sha256_hex`); redaction wired to the
existing `gateway/egress/dlp_detectors` (`scan_token_patterns` /
`redact_tokens`); embed the git SHA at build so `engine` is populated
(only `version = "0.1.0"` exists in `pyproject.toml` today — the build
must stamp the SHA); unit tests for determinism, chain-break detection,
`epoch`/`seq` continuity across a simulated restart, and redaction of
both a deny-listed key and a token-shaped value.
3. **SQLite index + `audit` CLI.** New `audit_events` migration (indexable
fields above); replay-from-journal; offline chain verifier
(`truncated-tail` vs `gap`); `query` / `verify` / `rebuild` / `import`;
idempotent `UPSERT` by `id`.
4. **Host controller as first producer (#468).** Wire
`lifecycle.bottled_agent_*` and `hostctl.*` emission into the host
controller; establish the `epoch` bump + chain-head carry + segment
rotation on writer restart here (it owns the single writer).
5. **Migrate existing producers.** Re-emit supervise `decision.*` (retiring
the standalone `supervise_audit_entries` shape behind the index), egress
`egress.*`, git-gate `forge.*` + `commit.signed` (#480), control-plane
`auth.*`; add the `audit.*` self-events (verify/truncation/export
failure).
6. **CloudEvents / OTel export adapters + #324 delivery.** Projection layer
(field re-map per *Export / interoperability*) plus the outbox cursor,
ack-driven advance, and retention-ordering guard the #324 contract
specifies.
7. **(follow-up.)** Cross-host merge transport; per-writer signing +
external anchoring on the chain head; retention/rotation *schedule*.
## Acceptance-criteria coverage (#487)
The issue defines the contract; implementation is explicitly split into
follow-up PRs. This PRD is the durable decision record; each acceptance box
maps to a section:
| #487 acceptance criterion | Where |
|---|---|
| Durable PRD defines versioned envelope + initial registry | *The envelope*, *Event registry* |
| Canonical JSON + hash-chain rules, unambiguous, with test vectors | *Canonical serialization + hash chain* → *Test vectors* |
| Trust provenance explicit for every common + event-specific field | *Trust provenance of every common field*; per-type `untrusted` in *Event registry* |
| Redaction prohibits credentials / raw secrets / unsafe capture by default | *Redaction rule*; `untrusted`-only claims; sensitivity classes |
| JSONL journal canonical; SQLite index fully rebuildable | *Journal + SQLite index* (`audit rebuild`) |
| Minimum local search/query contract | *Journal + SQLite index* → *Local query surface* |
| #324 can transport/replay without a second envelope | *The #324 delivery contract* |
| #480 maps in without weakening its byte-to-activation-key guarantee | *#480 signed-commit attribution maps in…* |
| Schema evolution + backward-compatible readers | *Schema evolution & backward-compatible readers* |
| Integrity detects modification / deletion / reorder / bad continuation | *Test vectors* (tamper), *Behavior across rotation…truncation*, `audit verify` |
Two acceptance items are **specified here, implemented later** by design
(the issue permits this): the concrete test-vector *fixtures* and the
`audit` CLI land in impl chunks 23; the #324 outbox lands in chunk 6.
Nothing in the contract is left undefined — only its code is deferred.
## Resolved in review (#495)
- **Single writer per host — decided.** The host controller owns the sole
appender; per-producer chains are a future scaling path only. (Design →
*Single writer; ordering across restarts*.)
- **Restarts — decided.** An `epoch` counter (bumped per writer boot) plus
carrying the last chain head as the next `prev` gives a total order of
`(epoch, seq)` that survives restarts; `ts_mono` orders only within an
epoch. (Design → *ordering across restarts*.)
- **Flatten to one `untrusted` region — decided.** Everything outside
`untrusted` (chain metadata, `producer`/`host`, `bottled_agent`, `ts_*`)
is trusted by construction, so the separate `trusted` sub-block is
removed; a field is trusted unless deliberately placed under `untrusted`.
(Design → *The envelope*.)
- **Subject term is `bottled_agent` everywhere** — the top-level field and
the `lifecycle.bottled_agent_*` leaf names. (Design → *The envelope* /
*Event registry*.)
- **Retention head-carry — yes.** When a journal segment is rotated out,
the new segment's genesis `prev` is the rotated-out head, so the verifier
still trusts the current head across a rotation. (Folds into the
retention follow-up.)
- **Redaction reuses the egress detectors — yes, scoped.** Reuse the
deterministic `dlp_detectors` (`scan_token_patterns` / `redact_tokens` /
`scan_known_secrets`); exclude `scan_entropy` as brittle on the short,
high-entropy structured values audit records carry. (Design → *Redaction
rule*.)
## Open questions
- **Value-scan cost on the hot path.** The single writer runs the reused
detectors on every event's `untrusted` block inline. Is that cheap enough
at lifecycle-event volume, or should the value scan move to index-build
time (journal stays raw, index stores the redacted view)? Leaning inline
so the raw journal never contains a leaked value in the first place.
- **`epoch` persistence location.** Store the per-writer `epoch` + chain
head in the SQLite index (rebuildable, but then the writer needs the DB
at boot) or in a tiny sidecar file next to the journal (independent of
the index)? Leaning sidecar, so the writer can start and append without
the index present.