Codex review on #496:
- **High — ambiguous delivery no longer orphans a launched bottle.** A
timeout / dropped response from the host controller is now the ambiguous
BrokerUnavailableError (distinct from the definite BrokerAuthError /
BrokerClientError). OrchestratorCore.launch_bottle keeps the registry
row on the ambiguous case instead of deregistering — deregistering would
orphan a running container with no record (reconcile reaps rows, never
containers). The row is left for reconcile to reap iff the bottle is not
actually live. Definite failures still roll back, so a real failure
leaves no orphan row.
- **Medium — the privileged endpoint bounds request bodies.** The host
server rejects an oversized Content-Length with 413 before reading it,
and sets a per-request socket timeout, so a caller that can merely reach
the socket (no signed token) can't exhaust memory or a handler thread.
Tests: ambiguous-keep vs definite-rollback in the launch path; the
BrokerUnavailableError/BrokerClientError split in BrokerClient; the 413
body cap + handler error paths (driven in-thread, since daemon request
threads lose coverage) plus a deterministic real-socket check that
declares an oversized Content-Length but sends a sliver (rejection on the
header, no unread-body reset race); and the __main__ entrypoint broker
selection. Diff-coverage 98%; pyright clean; pylint 9.8.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Chunk 1 of the host-control-server stack: close the PRD's **transport**
gap. Today LaunchBroker.submit(token) is an in-process method call from
OrchestratorCore; this makes it a real out-of-process service reached
over HTTP.
- host_server.py: the host control server. A pure dispatch() (POST
/broker verifies a signed token via the existing verify_request +
_launch/_teardown path, GET /health) wrapped by a thin http.server
adapter, mirroring orchestrator/server.py. Only the signed token
crosses the wire; provenance/schema failures are fail-closed 401s that
never touch the backend, a backend launch failure is a 502.
- broker_client.py: BrokerClient — a drop-in submit(token) that POSTs the
signed token to the host controller. A 401 re-raises as BrokerAuthError
so the launch path's rollback is identical local or remote.
- broker.py: SubmitBroker Protocol — the one method OrchestratorCore
depends on, satisfied by both LaunchBroker and BrokerClient, so the
core is unchanged (service.py annotation only).
- __main__.py: wire `--broker http` behind the shared-secret env var
(BOT_BOTTLE_BROKER_SECRET, hex) — a chunk-1 stopgap the durable
TrustDomain key (chunk 2, #476) replaces.
Tested: pure-dispatch cases, BrokerClient with HTTP mocked, and a
real-socket sign -> POST -> verify -> act round-trip (incl. fail-closed
forged token). pyright clean; pylint 9.86.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Per PR review, replay protection is too heavy for the host control
server MVP. Drop it from the four-gap framing (now three gaps), remove
the enforcement design section and implementation chunk, and track the
iat-window + jti-cache work in #494 as an independent in-process change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Promote the in-process launch broker into a standalone host control
server: the single privileged host component that brokers launches, owns
orchestrator lifecycle, and is the sole writer of host-durable state.
Closes the four broker gaps (transport, durable provisioned secret,
replay protection, disciplined op vocabulary) and splits host state by
owner and lifetime (orchestrator SQLite / host JSONL audit / gateway
none). The payoff is dropping the Docker socket from the CLI.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
On a docker daemon that default-enables IPv6 (default-address-pools),
creating the gateway network with only `--subnet` lets the daemon also
attach an fdd0::/64 IPv6 subnet. Its gateway is stored as `::1/64`,
which trips docker's own netip.ParseAddr in `network inspect`/`ls`:
ParseAddr("fdd0:0:0:6::1/64"): unexpected character, want colon
That poisons every `_network_cidr`/`network ls` read and fails the
docker integration suite intermittently (whichever run the runner's
IPv6 pool index lands on a broken network). bot-bottle attribution
pins IPv4 source IPs and has no IPv6 support, so pass `--ipv6=false`
explicitly at network create to keep the gateway network IPv4-only
regardless of the daemon default.
Note: an already-poisoned runner still needs a one-time
`docker network rm bot-bottle-gateway` (and possibly a daemon
restart) to clear the malformed network.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Gitea AGit accepts pushes to refs/for/*, refs/draft/*, and
refs/for-review/* and opens pull requests backed by server-managed
refs/pull/<n>/head refs rather than ordinary refs/heads/* branches.
That breaks the git-gate branch workflow: follow-up commits can't be
pushed back through the branch, and Gitea rejects later direct updates
to the generated review ref, so recovery means recreating the PR.
Add a Phase 0 guard to the shared pre-receive hook that rejects
creation or update of those AGit review refs before any gitleaks scan
or upstream forward, with a message pointing callers at the
branch-backed PR workflow. Deletions (new == zero) stay allowed so
legacy AGit refs can still be cleaned up; normal branches and tags are
untouched.
Closes#506
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The require-numbered-prds gate previously ran on every pull request.
Scope its trigger to PRs whose base branch is main, so numbering is
only enforced at the point of merging into main.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>