Compare commits

..

3 Commits

Author SHA1 Message Date
didericis-claude 985e79bd8e fix(backend): fix pyright errors in lazy-load implementation
lint / lint (push) Failing after 46s
test / unit (pull_request) Failing after 46s
test / integration (pull_request) Successful in 47s
test / coverage (pull_request) Successful in 1m24s
- Rename _BACKENDS → _backends: pyright treats uppercase module-level
  names as constants and flags the reassignment in _get_backends() as
  reportConstantRedefinition; lowercase avoids this.
- Add TYPE_CHECKING guard importing CommitCancelled/Freezer/get_freezer
  from .freeze: pyright cannot see module-level __getattr__ bindings, so
  reportUnsupportedDunderAll fired for those three __all__ entries; the
  guard makes them visible to the type checker without running at import
  time.
- Update test_backend_selection.py to patch _backends (lowercase).
2026-07-18 04:58:15 -04:00
didericis-claude e574b7b99f fix(backend): silence pylint false positives from lazy-load pattern
`undefined-all-variable` fires on CommitCancelled / Freezer / get_freezer
in __all__ because pylint can't see module-level __getattr__ bindings;
`global-statement` fires on the _BACKENDS singleton setter. Both are
intentional patterns — add inline disables rather than suppress globally.
2026-07-18 04:58:15 -04:00
didericis-claude 196a76dfe8 perf: lazy-load backend modules and consolidate docker subprocess helpers
Importing backend.docker.util previously triggered eager loading of all
three backend packages (~76 modules) because backend/__init__.py imported
DockerBottleBackend, FirecrackerBottleBackend, and MacosContainerBottleBackend
at module scope. This made the module prohibitively expensive to import
from the orchestrator layer and elsewhere.

The three backend imports are now deferred into _get_backends(), which
loads all three on first call and caches the result in the module-level
_BACKENDS variable (initially None). Module-level __getattr__ exposes
backend classes and freeze symbols lazily for existing import/patch sites.

backend/docker/util.py raw subprocess.run(["docker", ...]) calls are
replaced with the shared run_docker primitive from docker_cmd, eliminating
the duplication between the backend and orchestrator implementations.
_silent_run() is removed; image_exists() is inlined directly onto
run_docker. The commit_container test is updated to patch run_docker
instead of subprocess.run.
2026-07-18 04:58:15 -04:00
178 changed files with 1417 additions and 10672 deletions
-4
View File
@@ -1,10 +1,6 @@
[run]
branch = True
source = .
# Store paths relative to the project root so .coverage.* files produced on
# different runners (ubuntu-latest vs self-hosted KVM) can be combined by the
# coverage job without a [paths] remapping section.
relative_files = True
[report]
# Coverage policy: see docs/decisions/0004-coverage-policy.md.
+5 -2
View File
@@ -22,7 +22,10 @@ jobs:
- name: Checkout
uses: actions/checkout@v4
# No actions/setup-python: canaries are stdlib unittest on the image's
# system Python 3.12 (older act_runner mishandles setup-python's PATH).
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Run canaries
run: python3 -m unittest discover -t . -s tests/canaries -v
+10 -20
View File
@@ -13,30 +13,20 @@ jobs:
steps:
- uses: actions/checkout@v3
# No actions/setup-python: the runner image already ships Python 3.12,
# and older act_runner engines mishandle setup-python's PATH. Install
# into the ephemeral job container's system Python — the pylint/pyright
# console scripts land on /usr/local/bin (on PATH) so the steps below
# still resolve. --break-system-packages is safe: the container is
# disposable.
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: "3.12"
- name: Install dev dependencies
run: python3 -m pip install --break-system-packages -r requirements-dev.txt
run: |
python -m pip install --upgrade pip
pip install -r requirements-dev.txt
- name: Run pylint
run: |
# Pylint's normal exit code is nonzero for any emitted finding,
# regardless of --fail-under. Preserve the full report but enforce
# the aggregate score this workflow promises.
set +e
find . -name '*.py' -not -path './.venv/*' -not -path './.git/*' \
| xargs pylint --fail-under=8.0 \
| tee /tmp/pylint-output.txt
set -e
SCORE=$(sed -n \
's/^Your code has been rated at \([-0-9.]*\)\/10.*/\1/p' \
/tmp/pylint-output.txt | tail -1)
test -n "$SCORE"
awk -v score="$SCORE" 'BEGIN { exit !(score >= 8.0) }'
# Run pylint on all Python files in the repo
find . -name '*.py' -not -path './.venv/*' -not -path './.git/*' | xargs pylint --fail-under=8.0
- name: Run pyright
run: |
+5 -2
View File
@@ -37,8 +37,11 @@ jobs:
fetch-depth: 0
token: ${{ secrets.GITHUB_TOKEN }}
# No actions/setup-python: the inline script is stdlib-only on the
# image's system Python 3.12 (older act_runner mishandles its PATH).
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Configure git
run: |
git config user.name "github-actions[bot]"
+40 -244
View File
@@ -1,21 +1,19 @@
# Run the project's test suite when package or runtime inputs change on a PR
# or on push to main.
# Run the project's test suite on every PR push and on push to main.
#
# The suite uses stdlib `unittest` discovery — no external Python
# dependencies are required to execute it. Tests are split by directory:
#
# tests/unit/ — pure unit tests; always run
# tests/integration/ — need a reachable backend; skip cleanly when
# the backend isn't available on the runner
# tests/integration/ — need a reachable Docker daemon; skip cleanly
# (via tests/_docker.py:skip_unless_docker) when
# Docker isn't available on the runner
# tests/canaries/ — upstream regression canaries; run on a separate
# schedule (see canaries.yml), not here
#
# Each test job runs once under coverage and uploads a small .coverage.*
# artifact. The `coverage` job combines them — no test reruns, no KVM
# dependency on that job. For main-branch pushes only, the tested rootfs
# and matching dropbear are uploaded so `publish-infra` can publish the
# byte-identical artifact that was tested. PRs avoid the ~194 MB rootfs
# transfer entirely.
# This workflow assumes the Gitea Actions runner exposes the host Docker
# socket to the job container so `docker` commands inside the job can
# reach the daemon. If that's not yet configured on the runner the
# integration tests will skip rather than fail.
name: test
@@ -24,35 +22,10 @@ on:
branches:
- main
paths:
- 'bot_bottle/**'
- 'tests/**/*.py'
- 'cli.py'
- 'scripts/coverage.sh'
- 'scripts/critical-modules.txt'
- 'scripts/diff_coverage.py'
- 'scripts/tracker_policy.py'
- 'scripts/firecracker-netpool.sh'
- 'Dockerfile*'
- 'pyproject.toml'
- 'requirements-dev.txt'
- '.coveragerc'
- '.dockerignore'
- '**.py'
pull_request:
paths:
- 'bot_bottle/**'
- 'tests/**/*.py'
- 'cli.py'
- 'scripts/coverage.sh'
- 'scripts/critical-modules.txt'
- 'scripts/diff_coverage.py'
- 'scripts/tracker_policy.py'
- 'scripts/firecracker-netpool.sh'
- 'Dockerfile*'
- 'pyproject.toml'
- 'requirements-dev.txt'
- '.coveragerc'
- '.dockerignore'
workflow_dispatch:
- '**.py'
jobs:
unit:
@@ -61,47 +34,30 @@ jobs:
- name: Checkout
uses: actions/checkout@v4
# No actions/setup-python: the runner image already ships Python 3.12,
# and older act_runner engines mishandle setup-python's PATH (coverage
# lands in one interpreter, `python3` resolves to another). Install
# straight into the ephemeral job container's system Python —
# --break-system-packages is safe because the container is disposable.
- name: Install dev requirements
run: python3 -m pip install --break-system-packages -r requirements-dev.txt
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Run unit tests with coverage
env:
COVERAGE_FILE: ${{ github.workspace }}/.coverage.unit
- name: Install dev requirements
run: python3 -m pip install -r requirements-dev.txt
- name: Run unit tests
run: python3 -m coverage run -m unittest discover -t . -s tests/unit -v
- name: Report unit coverage
env:
COVERAGE_FILE: ${{ github.workspace }}/.coverage.unit
run: python3 -m coverage report -m
# upload-artifact@v3's glob skips dotfiles, so a bare `.coverage.unit`
# silently uploads nothing ("No files were found"). Stage it under a
# non-dot name; the coverage job renames it back before `coverage
# combine`. `cp` also fails loudly if coverage never wrote the file.
- name: Stage unit coverage for upload
run: cp .coverage.unit coverage-unit.dat
- name: Upload unit coverage artifact
uses: actions/upload-artifact@v3
with:
name: coverage-unit
path: coverage-unit.dat
integration-docker:
integration:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
# No actions/setup-python (see the note in the `unit` job); the
# container's system Python 3.12 runs the stdlib test suite directly.
- name: Install coverage
run: python3 -m pip install --break-system-packages coverage
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Show environment
run: |
@@ -112,193 +68,33 @@ jobs:
echo "docker not on PATH — integration tests will skip"
fi
- name: Run integration tests (docker) with coverage
env:
BOT_BOTTLE_BACKEND: docker
COVERAGE_FILE: ${{ github.workspace }}/.coverage.docker
run: python3 -m coverage run -m unittest discover -t . -s tests/integration -v
- name: Run integration tests
run: python3 -m unittest discover -t . -s tests/integration -v
# Non-dot name so upload-artifact's dotfile-skipping glob picks it up.
- name: Stage docker coverage for upload
run: cp .coverage.docker coverage-docker.dat
- name: Upload docker coverage artifact
uses: actions/upload-artifact@v3
with:
name: coverage-docker
path: coverage-docker.dat
# Integration tests against the Firecracker backend. Runs on a self-hosted
# KVM runner (label `kvm`) where /dev/kvm and the TAP/nft pool are available.
# Combined unit+integration coverage report (informational). See
# docs/decisions/0004-coverage-policy.md.
#
# Restricted to same-repo PRs, push to main, and workflow_dispatch — fork
# PRs don't execute untrusted code on the privileged runner.
#
# Runner prerequisites (provision once; see README "Firecracker on Linux"):
# `firecracker` on PATH, `/dev/kvm` accessible, cached kernel +
# static dropbear at /var/cache/bot-bottle-fc/dropbear, and the pool as a
# persistent systemd unit.
#
# The infra candidate is built here directly (no artifact download) to
# eliminate the ~70 s ubuntu-latest upload + ~83 s combined download that
# the old build-infra → integration-firecracker + coverage chain incurred.
# For main-branch pushes the tested rootfs and matching dropbear are
# uploaded so publish-infra can publish the byte-identical artifact; PRs
# skip those uploads entirely.
integration-firecracker:
runs-on: [self-hosted, kvm]
if: >-
github.event_name == 'push' ||
github.event_name == 'workflow_dispatch' ||
(github.event_name == 'pull_request' &&
github.event.pull_request.head.repo.full_name == github.repository)
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Preflight — Firecracker host is ready
run: |
command -v firecracker >/dev/null || {
echo "firecracker not on PATH — provision the runner (README: Firecracker on Linux)"; exit 1; }
test -e /dev/kvm || { echo "/dev/kvm missing — KVM not available on this runner"; exit 1; }
# `backend status` exits non-zero unless the TAP pool is up + no
# range overlap; it prints the exact `backend setup` fix.
python3 cli.py backend status --backend=firecracker
- name: Build infra candidate from this checkout
env:
BOT_BOTTLE_FC_DROPBEAR: /var/cache/bot-bottle-fc/dropbear
run: python3 -m bot_bottle.backend.firecracker.publish_infra --output infra-candidate --reuse-published
- name: Replace the persistent infra VM with the candidate
run: python3 -c 'from bot_bottle.backend.firecracker import infra_vm; infra_vm.stop()'
# No dev-requirements install: `coverage` is already provided by the
# self-hosted runner's Nix python env, and that env has no `pip`
# module to install into anyway.
- name: Run integration tests (firecracker) with coverage
env:
BOT_BOTTLE_BACKEND: firecracker
BOT_BOTTLE_INFRA_ARTIFACT_DIR: ${{ github.workspace }}/infra-candidate
COVERAGE_FILE: ${{ github.workspace }}/.coverage.firecracker
run: python3 -m coverage run -m unittest discover -t . -s tests/integration -v
# Non-dot name so upload-artifact's dotfile-skipping glob picks it up.
- name: Stage firecracker coverage for upload
run: cp .coverage.firecracker coverage-firecracker.dat
- name: Upload firecracker coverage artifact
uses: actions/upload-artifact@v3
with:
name: coverage-firecracker
path: coverage-firecracker.dat
# Only upload the large rootfs artifact on main-branch pushes;
# PRs avoid the ~194 MB transfer. publish-infra only runs on main
# and downloads these to publish the byte-identical tested rootfs.
- name: Upload tested rootfs (main branch only)
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
uses: actions/upload-artifact@v3
with:
name: infra-candidate
path: infra-candidate/
- name: Upload dropbear for publish verification (main branch only)
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
uses: actions/upload-artifact@v3
with:
name: firecracker-inputs
path: /var/cache/bot-bottle-fc/dropbear
# Combined coverage gate: aggregates .coverage.* artifacts uploaded by each
# test job, then runs the diff-coverage gate (new/changed lines >= 90%).
#
# Runs on ubuntu-latest — no KVM needed, no test reruns. Coverage files use
# relative_files = True (.coveragerc) so they combine cleanly across runners.
# Each test job sets COVERAGE_FILE to an absolute path so coverage.py writes
# to a known location that upload-artifact can find regardless of runner env.
#
# Restricted to the same events as integration-firecracker: it depends on
# that job's coverage artifact and skips for fork PRs alongside it.
# The hard diff-coverage gate (changed lines >= 90%) is DEFERRED: the
# Firecracker backend's VM/SSH orchestration is covered by the integration
# suite, which needs /dev/kvm + the provisioned TAP/nft pool — a
# container-based runner skips it and those lines read uncovered, so the
# gate can't pass here. Re-enabling it on a self-hosted KVM runner is
# tracked separately (see PRD 0069 / #348 and the ci-runner branch).
coverage:
needs: [unit, integration-docker, integration-firecracker]
timeout-minutes: 15
runs-on: ubuntu-latest
if: >-
github.event_name == 'push' ||
github.event_name == 'workflow_dispatch' ||
(github.event_name == 'pull_request' &&
github.event.pull_request.head.repo.full_name == github.repository)
steps:
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Install coverage
run: python3 -m pip install --break-system-packages coverage
- name: Download unit coverage artifact
uses: actions/download-artifact@v3
- name: Set up Python
uses: actions/setup-python@v5
with:
name: coverage-unit
path: ${{ github.workspace }}
python-version: "3.12"
- name: Download docker coverage artifact
uses: actions/download-artifact@v3
with:
name: coverage-docker
path: ${{ github.workspace }}
- name: Install dev requirements
run: python3 -m pip install -r requirements-dev.txt
- name: Download firecracker coverage artifact
uses: actions/download-artifact@v3
with:
name: coverage-firecracker
path: ${{ github.workspace }}
# Rename the non-dot upload names back to the .coverage.* files that
# `coverage combine` discovers (see the staging steps in each test job).
- name: Reassemble coverage data files
run: |
mv coverage-unit.dat .coverage.unit
mv coverage-docker.dat .coverage.docker
mv coverage-firecracker.dat .coverage.firecracker
- name: Combined coverage (unit + integration, incl. firecracker)
run: PYTHON=python3 bash scripts/coverage.sh aggregate critical
- name: Diff-coverage gate (changed lines >= 90%)
run: |
git fetch --no-tags origin main:refs/remotes/origin/main
python3 scripts/diff_coverage.py --base origin/main --min 90
publish-infra:
needs: [unit, integration-docker, integration-firecracker, coverage]
runs-on: ubuntu-latest
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
steps:
- name: Checkout the tested revision
uses: actions/checkout@v4
- name: Download the tested rootfs
uses: actions/download-artifact@v3
with:
name: infra-candidate
path: infra-candidate
# publish_infra re-derives the version from the checkout to confirm the
# bundle matches before uploading, and the version hashes the dropbear
# bytes. Download the SAME dropbear integration-firecracker used, or
# the recheck computes a "<missing>"-dropbear version and rejects the
# candidate.
- name: Download the staged dropbear (matches build's version)
uses: actions/download-artifact@v3
with:
name: firecracker-inputs
path: firecracker-inputs
- name: Publish the tested candidate
env:
BOT_BOTTLE_INFRA_ARTIFACT_TOKEN: ${{ secrets.BOT_BOTTLE_INFRA_ARTIFACT_TOKEN }}
BOT_BOTTLE_FC_DROPBEAR: ${{ github.workspace }}/firecracker-inputs/dropbear
run: python3 -m bot_bottle.backend.firecracker.publish_infra --publish-dir infra-candidate
- name: Combined coverage report (unit + integration)
run: PYTHON=python3 bash scripts/coverage.sh critical
@@ -1,17 +0,0 @@
name: tracker-policy-issues
on:
issues:
types: [opened, unlabeled]
jobs:
label-issue:
runs-on: ubuntu-latest
permissions:
issues: write
steps:
- uses: actions/checkout@v4
- name: Ensure the issue has a label
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: python3 scripts/tracker_policy.py label-issue
-18
View File
@@ -1,18 +0,0 @@
name: tracker-policy-pr
on:
pull_request:
types: [opened, edited, reopened, synchronize, labeled, unlabeled]
jobs:
check-pr:
runs-on: ubuntu-latest
permissions:
issues: read
pull-requests: read
steps:
- uses: actions/checkout@v4
- name: Require an unlabeled PR linked to an issue
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: python3 scripts/tracker_policy.py check-pr
+12 -12
View File
@@ -14,27 +14,27 @@ on:
jobs:
update-badges:
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- uses: actions/checkout@v3
with:
fetch-depth: 0
token: ${{ secrets.BADGE_PUSH_TOKEN }}
token: ${{ secrets.GITHUB_TOKEN }}
- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: '3.12'
# No actions/setup-python: the runner image ships Python 3.12 and older
# act_runner engines mishandle setup-python's PATH. Install into the
# ephemeral job container's system Python (--break-system-packages is
# safe because the container is disposable).
- name: Install dev dependencies
run: python3 -m pip install --break-system-packages -r requirements-dev.txt
run: |
python -m pip install --upgrade pip
pip install -r requirements-dev.txt
- name: Run coverage and extract percentage
id: coverage
run: |
python3 -m coverage run -m unittest discover -t . -s tests/unit > /dev/null 2>&1 || true
PERCENT=$(python3 -m coverage report 2>/dev/null | grep '^TOTAL' | grep -oP '\d+(?=%)' | tail -1)
python -m coverage run -m unittest discover -t . -s tests/unit > /dev/null 2>&1 || true
PERCENT=$(python -m coverage report 2>/dev/null | grep '^TOTAL' | grep -oP '\d+(?=%)' | tail -1)
echo "percent=$PERCENT" >> $GITHUB_OUTPUT
echo "Coverage: $PERCENT%"
@@ -45,7 +45,7 @@ jobs:
# the single source of truth in scripts/critical-modules.txt; every
# core module is unit-tested, so the unit-only run is accurate for it.
INCLUDE=$(grep -vE '^[[:space:]]*(#|$)' scripts/critical-modules.txt | paste -sd, -)
PERCENT=$(python3 -m coverage report --include="$INCLUDE" 2>/dev/null | grep '^TOTAL' | grep -oP '\d+(?=%)' | tail -1)
PERCENT=$(python -m coverage report --include="$INCLUDE" 2>/dev/null | grep '^TOTAL' | grep -oP '\d+(?=%)' | tail -1)
echo "percent=$PERCENT" >> $GITHUB_OUTPUT
echo "Core coverage: $PERCENT%"
+29 -19
View File
@@ -16,12 +16,10 @@
# Layout:
#
# /usr/bin/gitleaks gitleaks binary
# /app/egress_addon.py mitmproxy addon entry point
# /app/egress_addon.py + siblings mitmproxy addon (egress)
# /app/egress-entrypoint.sh mitmdump launcher
# /usr/local/lib/python*/bot_bottle/ installed package (all daemons + shared modules)
# /app/egress_addon.py one-line shim: re-exports addons from package
# (mitmdump -s requires a file path, not a module)
# /etc/egress/routes.yaml bind-mounted at run time
# /app/supervise_server.py + .py supervise MCP server
# /app/gateway_init.py PID 1 supervisor
# /etc/git-gate/pre-receive docker-cp'd at start time
# /git-gate-entrypoint.sh docker-cp'd at start time
# /git-gate/creds/* docker-cp'd at start time
@@ -89,19 +87,27 @@ RUN arch="${TARGETARCH:-$(dpkg --print-architecture)}" \
&& tar -xzf /tmp/gitleaks.tar.gz -C /usr/bin gitleaks \
&& rm /tmp/gitleaks.tar.gz
# Install bot_bottle as a proper package so entry-point scripts can use
# `from bot_bottle.X import Y` absolute imports. A rename or a missing
# module is caught at pip-install time — not at container runtime.
COPY pyproject.toml /src/
COPY bot_bottle/ /src/bot_bottle/
RUN pip install --no-cache-dir /src/
# mitmdump -s requires a file path, not a module. Write a one-line shim that
# re-exports `addons` from the installed package; mitmdump finds it there.
# WORKDIR here also creates /app so the shim + COPYs below can write into it
# (nothing created /app before this point).
WORKDIR /app
RUN printf 'from bot_bottle.egress_addon import addons\n' > /app/egress_addon.py
# Project Python: addon + server modules + the init supervisor.
# Kept flat under /app/ so mitmdump's loader resolves them as
# top-level siblings (absolute imports), matching the prior
# Dockerfile.egress / Dockerfile.supervise layout.
COPY bot_bottle/egress_addon_core.py /app/egress_addon_core.py
COPY bot_bottle/egress_dlp_config.py /app/egress_dlp_config.py
COPY bot_bottle/egress_addon.py /app/egress_addon.py
COPY bot_bottle/policy_resolver.py /app/policy_resolver.py
COPY bot_bottle/dlp_detectors.py /app/dlp_detectors.py
COPY bot_bottle/yaml_subset.py /app/yaml_subset.py
COPY bot_bottle/paths.py /app/paths.py
COPY bot_bottle/migrations.py /app/migrations.py
COPY bot_bottle/db_store.py /app/db_store.py
COPY bot_bottle/supervise_types.py /app/supervise_types.py
COPY bot_bottle/queue_store.py /app/queue_store.py
COPY bot_bottle/audit_store.py /app/audit_store.py
COPY bot_bottle/store_manager.py /app/store_manager.py
COPY bot_bottle/supervise.py /app/supervise.py
COPY bot_bottle/supervise_server.py /app/supervise_server.py
COPY bot_bottle/gateway_init.py /app/gateway_init.py
COPY bot_bottle/git_http_backend.py /app/git_http_backend.py
COPY bot_bottle/egress_entrypoint.sh /app/egress-entrypoint.sh
RUN chmod +x /app/egress-entrypoint.sh
@@ -120,6 +126,10 @@ RUN mkdir -p \
# subset the bottle uses.
EXPOSE 8888 9099 9418 9420 9100
# WORKDIR matches Dockerfile.supervise's prior layout so the
# in-app same-dir import in supervise_server.py stays deterministic.
WORKDIR /app
# PID 1 is the supervisor. It owns signal handling and exit-code
# propagation; no `exec` chain in the entrypoint itself.
ENTRYPOINT ["python3", "-m", "bot_bottle.gateway_init"]
ENTRYPOINT ["python3", "/app/gateway_init.py"]
+36 -13
View File
@@ -1,22 +1,45 @@
# Shared infra image: gateway data plane + orchestrator control plane.
# Firecracker single infra-VM image (PRD 0070 Stage B).
#
# Used directly by the Docker backend (run as one `bot-bottle-infra`
# container, replacing the prior two-container split). The Firecracker
# backend extends this via Dockerfile.infra.fc, adding buildah/crun/
# netavark for in-VM agent-image building.
#
# Dockerfile.orchestrator is the single definition of the orchestrator
# content (the lean `bot_bottle` package on python:3.12-slim). Both this
# image and Dockerfile.infra.fc pull it in via `COPY --from`.
# The per-host infra VM runs the orchestrator control plane, the gateway
# data plane, AND builds agent images (buildah) — all in one microVM (see
# backend/firecracker/infra_vm.py). It composes:
# * FROM the gateway image (mitmproxy / git / gitleaks / supervise + the
# flat daemon modules) — now trixie-based, so buildah 1.39 is available;
# * `COPY --from` the orchestrator image's content (the single definition
# of the control-plane payload — see Dockerfile.orchestrator), so this
# VM and the docker backend share one orchestrator definition; and
# * buildah, installed HERE only (the docker orchestrator/gateway images
# never carry it).
#
# multi-`FROM` can't union two bases (that's multi-stage, not multiple
# inheritance), so the orchestrator content is pulled in via `COPY --from`
# rather than a second base. Both images share the trixie `python:3.12-slim`
# base, so the copy is clean (same python; future installed deps copy too).
#
# The docker backend keeps orchestrator + gateway as separate images; this
# combined image exists only for the Firecracker single-VM cut. Splitting a
# service back into its own VM later is a routing change, not a repackaging
# (PRD 0070's "secret concentration"; a disposable builder can boot from
# this same image on its own TAP).
FROM bot-bottle-gateway:latest
# The orchestrator content, from its single definition. The gateway image
# already has the flat daemon modules under /app; this adds the full
# `bot_bottle` package so `python3 -m bot_bottle.orchestrator` resolves —
# used by gateway_init when BOT_BOTTLE_GATEWAY_DAEMONS includes `orchestrator`.
# --- in-VM agent-image builder (PRD 0069 Stage 3) -------------------
# The Firecracker backend builds users' agent Dockerfiles *inside this VM*
# with buildah (rootless, daemonless) instead of on the host — no host
# Docker daemon, no root-equivalent `docker` group. `crun` is the OCI
# runtime; `netavark` + `aardvark-dns` are the network backend for `FROM`
# pulls + `RUN` egress. Requires the trixie base (buildah 1.39: bookworm's
# 1.28 can't parse Dockerfile heredocs that agent images use).
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
buildah crun netavark aardvark-dns \
&& rm -rf /var/lib/apt/lists/*
# vfs + chroot: buildah works as root in the bare microVM (no
# fuse-overlayfs / overlay module / subuid maps). Matches image_builder.
ENV STORAGE_DRIVER=vfs \
BUILDAH_ISOLATION=chroot
# The orchestrator content, pulled from its single definition. The gateway
# image already has the flat daemon modules under /app; this adds the full
# `bot_bottle` package so `python3 -m bot_bottle.orchestrator` resolves.
COPY --from=bot-bottle-orchestrator:latest /app/bot_bottle /app/bot_bottle
-23
View File
@@ -1,23 +0,0 @@
# Firecracker infra VM image (PRD 0070 Stage B).
#
# Extends the shared infra base (Dockerfile.infra: gateway + orchestrator
# control plane) with the in-VM agent-image builder. The Firecracker backend
# builds users' agent Dockerfiles *inside this VM* with buildah (rootless,
# daemonless) instead of on the host — no host Docker daemon, no
# root-equivalent `docker` group.
#
# Requires the trixie base from bot-bottle-gateway (buildah 1.39: bookworm's
# 1.28 can't parse Dockerfile heredocs that agent images use).
#
# `crun` is the OCI runtime; `netavark` + `aardvark-dns` are the network
# backend for `FROM` pulls + `RUN` egress. `vfs` + `chroot`: buildah works
# as root in the bare microVM (no fuse-overlayfs / overlay module / subuid
# maps). Matches image_builder.
FROM bot-bottle-infra:latest
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
buildah crun netavark aardvark-dns \
&& rm -rf /var/lib/apt/lists/*
ENV STORAGE_DRIVER=vfs \
BUILDAH_ISOLATION=chroot
+2 -95
View File
@@ -5,8 +5,8 @@
# bot-bottle
[![test](https://gitea.dideric.is/didericis/bot-bottle/actions/workflows/test.yml/badge.svg?branch=main)](https://gitea.dideric.is/didericis/bot-bottle/actions?workflow=test.yml)
[![coverage](https://img.shields.io/badge/coverage-83%25-brightgreen)](https://coverage.readthedocs.io/)
[![core coverage](https://img.shields.io/badge/core%20coverage-94%25-brightgreen)](https://gitea.dideric.is/didericis/bot-bottle/src/branch/main/docs/decisions/0004-coverage-policy.md)
[![coverage](https://img.shields.io/badge/coverage-82%25-brightgreen)](https://coverage.readthedocs.io/)
[![core coverage](https://img.shields.io/badge/core%20coverage-95%25-brightgreen)](https://gitea.dideric.is/didericis/bot-bottle/src/branch/main/docs/decisions/0004-coverage-policy.md)
**Problem:** Developer wants to run a coding agent without supervision, but they don't want a prompt injected or misbehaving agent wrecking their environment or exfiltrating sensitive data.
@@ -75,88 +75,6 @@ On compatible macOS hosts, the default backend requires Apple's `container` CLI
Use `BOT_BOTTLE_BACKEND=docker ./cli.py start <agent>` on hosts where neither Apple Container nor KVM is available and Docker is the desired backend.
### Containers inside a bottle
A bottle may set `nested_containers: true`. On the macOS backend this starts a
guest-local, rootless **podman** service after the bottle is registered and
exposes its Docker-compatible API socket, so the agent still runs `docker` and
`docker compose`. Nothing is mounted from the host: Docker Desktop's socket
stays out of the bottle and the guest gains no outer VM capabilities. Backends
that cannot do this (`docker`, `firecracker`) reject the flag rather than
silently ignore it.
Rootless Docker was tried first and does not work here at all: Apple
Container's capability bounding set omits `CAP_SYS_ADMIN`, which the kernel
requires to write a multi-range `uid_map`. See
[`docs/research/rootless-docker-in-apple-container-spike.md`](docs/research/rootless-docker-in-apple-container-spike.md).
The tradeoff to understand before enabling it: podman avoids that requirement
by falling back to a single-UID mapping, so nested containers provide **no
isolation from the agent itself** — `root` inside a nested container is the
agent user outside it. Nested containers are a build/test convenience, not a
security boundary. The bottle remains the boundary.
Pulling images goes through the bottle's egress proxy like every other
request, so each registry needs a route — **and so does the CDN it redirects
layer blobs to**, which is a different host. Without the CDN route the pull
authenticates, fetches the manifest, then 403s partway through.
Docker Hub and GHCR additionally need `preserve_auth: true`: their token dance
uses a client-fetched per-scope bearer token that the proxy would otherwise
strip. Turn DLP off on every registry and CDN route — the bodies are
compressed layer blobs that no detector can read, and buffering them is what
triggers the shared-proxy OOM in #455.
```yaml
nested_containers: true
egress:
routes:
# Docker Hub: registry, token endpoint, blob CDN.
- host: registry-1.docker.io
preserve_auth: true
dlp: { outbound_detectors: false, inbound_detectors: false }
- host: auth.docker.io
preserve_auth: true
dlp: { outbound_detectors: false, inbound_detectors: false }
- host: production.cloudfront.docker.com
dlp: { outbound_detectors: false, inbound_detectors: false }
# GHCR: registry + blob CDN.
- host: ghcr.io
preserve_auth: true
dlp: { outbound_detectors: false, inbound_detectors: false }
- host: pkg-containers.githubusercontent.com
dlp: { outbound_detectors: false, inbound_detectors: false }
# quay.io: registry + blob CDNs. No preserve_auth needed for public pulls.
- host: quay.io
dlp: { outbound_detectors: false, inbound_detectors: false }
- host: cdn01.quay.io
dlp: { outbound_detectors: false, inbound_detectors: false }
- host: cdn02.quay.io
dlp: { outbound_detectors: false, inbound_detectors: false }
- host: cdn03.quay.io
dlp: { outbound_detectors: false, inbound_detectors: false }
```
`mcr.microsoft.com` and `registry.k8s.io` follow the same shape and also
redirect blobs elsewhere (`*.data.mcr.microsoft.com` and
`us-*-docker.pkg.dev` respectively); route whichever host the 403 names.
Inside a nested container the same allowlist applies: an allowlisted host
returns 200 and anything else gets a 403 straight from the proxy. The
gateway's CA bundle and proxy settings are wired in automatically, so
`docker run … curl https://…` works with no extra flags — no `--add-host`,
`-e`, or `-v`.
Two things worth knowing when testing that:
- Public DNS inside a nested container fails **by design**. Everything
egresses through the proxy, so `nslookup` failing is expected and is not
evidence of a problem.
- Alpine's BusyBox `wget` drops the connection after the proxy's TLS
interception and reports `error getting response`, even though the proxy
logs the decrypted request and returns a response. Use `curl` to test
egress; BusyBox `wget` will lie to you.
### Firecracker on Linux
On Linux, a KVM-capable host defaults to the Firecracker backend. It needs:
@@ -172,8 +90,6 @@ BOT_BOTTLE_BACKEND=firecracker ./cli.py start <agent>
> **NixOS:** enable `virtualisation.docker`, ensure the KVM module is loaded (`boot.kernelModules = [ "kvm-intel" ];` or `kvm-amd`), and add your user to the `kvm` and `docker` groups. For the network pool, consume the flake module — `imports = [ inputs.bot-bottle.nixosModules.firecracker-netpool ]; services.bot-bottle-firecracker = { enable = true; owner = "you"; };` — then `nixos-rebuild switch` (imperative nft/TAP rules don't survive a rebuild; channel users can `imports = [ <bot-bottle>/nix/firecracker-netpool.nix ]`). `firecracker` isn't in nixpkgs by default as a user binary — install the release binary (pin the version) and put it on `PATH`.
> **CI:** the coverage gate (`.gitea/workflows/test.yml` → `coverage` job) runs on a self-hosted runner labelled `kvm`, because the Firecracker backend's VM/SSH orchestration is exercised only by the integration suite, which needs `/dev/kvm` + the provisioned pool (a container runner would skip it and read as uncovered). Provision that runner exactly like a normal Firecracker host — `firecracker` on `PATH`, `/dev/kvm`, the cached guest kernel + static dropbear, and the pool installed as the persistent systemd unit — then register it with the `kvm` label. A Docker-capable hosted job builds the candidate once; KVM tests boot those exact bytes, and a successful main run publishes them. The unit/lint jobs still run on `ubuntu-latest`.
```sh
./cli.py start <agent> # builds the image on first run, drops you into claude
```
@@ -255,15 +171,6 @@ When an outbound DLP detector matches a token, the route's `dlp.outbound_on_matc
More examples in `examples/`. Full design lives under `docs/prds/`; the trust-boundary rationale is in `docs/prds/0011-per-file-md-manifest.md`.
## Tracker policy
Issues are the canonical work items and own all tracker labels; every issue
must have at least one. Pull requests stay unlabeled and deliberately reference
an issue with `Closes #…`, `Part of #…`, or another form defined in
[`ADR 0005`](docs/decisions/0005-issues-own-tracker-metadata.md). Gitea Actions
enforces the convention for new work from 2026-07-18 onward. Earlier closed
PRs are grandfathered rather than given artificial retrospective issues.
## Trademarks
bot-bottle is an independent project and is not affiliated with, endorsed by, or sponsored by Anthropic, PBC. "Claude" and "Claude Code" are trademarks of Anthropic, PBC; the project name uses "claude" descriptively to indicate that the tool runs Claude Code inside a sandbox.
+2 -43
View File
@@ -45,10 +45,6 @@ PROVIDER_TEMPLATES = frozenset({PROVIDER_CLAUDE, PROVIDER_CODEX, PROVIDER_PI})
# forward_host_credentials is enabled. Pipelock must pass these through
# (no TLS MITM) or its header DLP blocks the injected JWT.
CODEX_HOST_CREDENTIAL_HOSTS = ("api.openai.com", "chatgpt.com")
# Host that egress injects the host Claude bearer on when Claude
# forward_host_credentials is enabled.
CLAUDE_HOST_CREDENTIAL_HOSTS = ("api.anthropic.com",)
PromptMode = Literal[
"append_file",
"read_prompt_file",
@@ -261,28 +257,7 @@ class AgentProvider(ABC):
Default: Debian/node — writes the git-gate insteadOf gitconfig
and sets user.name/email as node. Workspace copy runs through
BottleBackend.provision_workspace against the running bottle."""
from .log import die, info
# Firecracker exports image rootfs files through an unprivileged host
# tar extraction, so image-time ownership of XDG directories is not
# preserved. Git consults ~/.config/git even when the actual config
# is ~/.gitconfig; an unreadable directory there can prevent the
# git-gate insteadOf rules below from taking effect. Repair this at
# runtime, after every backend's copy/export path has completed.
git_xdg_dir = f"{plan.guest_home}/.config/git"
repair = bottle.exec(
f"chown node:node {shlex.quote(plan.guest_home)} && "
f"chmod 755 {shlex.quote(plan.guest_home)} && "
f"mkdir -p {shlex.quote(git_xdg_dir)} && "
f"chown -R node:node {shlex.quote(f'{plan.guest_home}/.config')} && "
f"chmod -R u+rwX,go+rX {shlex.quote(f'{plan.guest_home}/.config')}",
user="root",
)
if repair.returncode != 0:
die(
"git provisioning: could not make the runtime Git config "
f"directory readable: {(repair.stderr or repair.stdout).strip()}"
)
from .log import info
manifest_bottle = plan.manifest.bottle
if manifest_bottle.git:
@@ -305,27 +280,11 @@ class AgentProvider(ABC):
f"{len(manifest_bottle.git)} insteadOf rule(s)"
)
bottle.cp_in(str(config_file), guest_gitconfig)
permissions = bottle.exec(
bottle.exec(
f"chown node:node {shlex.quote(guest_gitconfig)} && "
f"chmod 644 {shlex.quote(guest_gitconfig)}",
user="root",
)
if permissions.returncode != 0:
die(
"git provisioning: could not set ownership on "
f"{guest_gitconfig}: "
f"{(permissions.stderr or permissions.stdout).strip()}"
)
configured = bottle.exec(
"git config --global --get-regexp '^url\\..*\\.insteadof$'",
user="node",
)
if configured.returncode != 0:
die(
"git provisioning: the runtime user cannot read the "
f"git-gate insteadOf rules from {guest_gitconfig}: "
f"{(configured.stderr or configured.stdout).strip()}"
)
gu = manifest_bottle.git_user
if not gu.is_empty():
+2 -2
View File
@@ -42,7 +42,7 @@ class AuditStore(DbStore):
super().__init__(db_path or host_db_path(), migrations)
def write_audit_entry(self, entry: AuditEntry) -> Path:
with self._connection() as conn:
with self._connect() as conn:
conn.execute(
"""
INSERT INTO supervise_audit_entries (
@@ -66,7 +66,7 @@ class AuditStore(DbStore):
def read_audit_entries(self, component: str, slug: str) -> list[AuditEntry]:
if not self.db_path.is_file():
return []
with self._connection() as conn:
with self._connect() as conn:
rows = conn.execute(
"""
SELECT * FROM supervise_audit_entries
+12 -120
View File
@@ -37,16 +37,15 @@ import os
import shlex
import sys
from abc import ABC, abstractmethod
from contextlib import AbstractContextManager, contextmanager
from contextlib import AbstractContextManager
from dataclasses import dataclass
from pathlib import Path
from typing import TYPE_CHECKING, Any, Generator, Generic, Sequence, TypeVar
from typing import TYPE_CHECKING, Any, Generic, Sequence, TypeVar
from ..agent_provider import AgentProvisionPlan, get_provider, build_agent_provision_plan
from ..egress import EgressPlan
from ..git_gate import GitGatePlan
from ..log import die, info, warn
from ..util import read_tty_line
from ..log import die, info
from ..manifest import Manifest, ManifestIndex
from ..supervise import SupervisePlan
from ..util import expand_tilde
@@ -83,9 +82,6 @@ class BottleSpec:
# True when launched via --headless (no TTY, no interactive prompts).
# The git-gate host-key preflight uses this to error rather than prompt.
headless: bool = False
# Image startup policy. "fresh" preserves the normal build path;
# "cached" reuses the current local image/artifact without rebuilding.
image_policy: str = "fresh"
@dataclass(frozen=True)
@@ -281,18 +277,6 @@ PlanT = TypeVar("PlanT", bound=BottlePlan)
CleanupT = TypeVar("CleanupT", bound=BottleCleanupPlan)
@dataclass(frozen=True)
class BottleImages:
"""Resolved image references (or artifact paths) for a bottle launch.
For Docker/macOS-container backends, `agent` and `sidecar` are string
image refs. For the smolmachines backend they are Path objects pointing
to pre-built `.smolmachine` artifacts."""
agent: str | Path
sidecar: str | Path = ""
class BottleBackend(ABC, Generic[PlanT, CleanupT]):
"""Abstract base for selectable bottle backends. Concrete subclasses
(e.g. DockerBottleBackend) own their own prepare/launch impls.
@@ -302,11 +286,6 @@ class BottleBackend(ABC, Generic[PlanT, CleanupT]):
name: str
# Whether this backend can run a container engine *inside* the bottle.
# Backends that cannot must reject `nested_containers: true` rather than
# reach for a host daemon socket (issue #392).
supports_nested_containers: bool = False
def prepare(self, spec: BottleSpec, stage_dir: Path) -> PlanT:
"""Template method: run cross-backend host-side validation, then
delegate to the subclass's `_resolve_plan` for the
@@ -320,16 +299,12 @@ class BottleBackend(ABC, Generic[PlanT, CleanupT]):
prepare_egress,
prepare_git_gate,
prepare_supervise,
reject_nested_containers,
resolve_manifest_dockerfile,
write_launch_metadata,
)
manifest = self._validate(spec)
if not self.supports_nested_containers:
reject_nested_containers(self.name, manifest)
self._preflight()
from ..git_gate_host_key import preflight_host_keys
@@ -461,27 +436,9 @@ class BottleBackend(ABC, Generic[PlanT, CleanupT]):
prompt file, Dockerfile path, and guest home all live on
`agent_provision_plan` — the source of truth."""
def prelaunch_checks(self, plan: PlanT) -> None:
"""Raise StaleImageError if any cached image used by this plan is stale.
No-op default; backends override to call the shared check_stale*
helpers on their image/artifact timestamps. Called by the CLI before
launch so the operator can be prompted outside the launch context."""
@contextmanager
def launch(self, plan: PlanT) -> Generator[Bottle, None, None]:
"""Template: build or load images, then delegate to _launch_impl."""
images = self._build_or_load_images(plan)
with self._launch_impl(plan, images) as bottle:
yield bottle
@abstractmethod
def _build_or_load_images(self, plan: PlanT) -> BottleImages:
"""Return the agent and sidecar image references (or artifact paths)
for this plan, building fresh images when the policy requires it."""
@abstractmethod
def _launch_impl(self, plan: PlanT, images: BottleImages) -> AbstractContextManager[Bottle]:
"""Bring up the bottle using pre-resolved images; yield a handle; tear down on exit."""
def launch(self, plan: PlanT) -> AbstractContextManager[Bottle]:
"""Build/run the bottle and yield a handle; tear down on exit."""
def provision(self, plan: PlanT, bottle: "Bottle") -> str | None:
"""Copy host-side files (CA cert, prompt, skills, .git) into
@@ -691,26 +648,19 @@ def __getattr__(name: str) -> Any:
def get_bottle_backend(
name: str | None = None,
*,
prompt: bool = True,
) -> BottleBackend[Any, Any]:
"""Resolve the bottle backend.
`name` precedence:
1. explicit arg (e.g. resume passes the recorded backend name)
1. explicit arg (CLI `--backend=<name>` passes through here)
2. BOT_BOTTLE_BACKEND env var
3. auto-selection: VM backend first, docker fallback with prompt
`prompt` controls whether auto-selection may block on an interactive
[i/d/q] prompt when falling back to docker. Pass `prompt=False` in
non-interactive contexts (headless launches, CI) so the call dies
with an actionable message instead of hanging.
3. `macos-container` on compatible macOS hosts
4. `firecracker` on KVM-capable Linux hosts
5. default `docker`
Dies with a pointer at the known backends if the chosen name
isn't implemented."""
resolved = name or os.environ.get("BOT_BOTTLE_BACKEND")
if resolved is None:
resolved = _auto_select_backend(prompt=prompt)
resolved = name or os.environ.get("BOT_BOTTLE_BACKEND") or _default_backend_name()
backends = _get_backends()
if resolved not in backends:
known = ", ".join(sorted(backends))
@@ -718,35 +668,7 @@ def get_bottle_backend(
return backends[resolved]
def _platform_vm_suggestion() -> str:
"""Platform-appropriate VM backend name for install suggestions."""
return "macos-container" if sys.platform == "darwin" else "firecracker"
def _print_vm_install_instructions() -> None:
"""Print platform-appropriate VM backend install instructions to stderr."""
vm = _platform_vm_suggestion()
if vm == "macos-container":
info("Install Apple Container: https://github.com/apple/container/releases")
info("Then start the service: container system start")
else:
info("Install Firecracker: https://github.com/firecracker-microvm/firecracker/releases")
info("Configure the host: ./cli.py backend setup")
def _auto_select_backend(prompt: bool = True) -> str:
"""Tier-1 / tier-2 backend auto-selection.
Tier 1: VM backend — macos-container on macOS when Apple Container is
installed; firecracker on KVM-capable Linux even before the binary is
present (its preflight prints an install pointer).
Tier 2: docker, with a security warning and an interactive prompt.
When `prompt=False` (headless / CI), dies with an actionable message
instead of blocking on a TTY read. When docker is also absent, prints
VM install instructions and exits.
"""
# --- Tier 1: VM backend -----------------------------------------
def _default_backend_name() -> str:
if has_backend("macos-container"):
return "macos-container"
# A KVM-capable Linux host defaults to firecracker even when the
@@ -756,37 +678,7 @@ def _auto_select_backend(prompt: bool = True) -> str:
from .firecracker import FirecrackerBottleBackend
if FirecrackerBottleBackend.is_host_capable():
return "firecracker"
# --- Tier 2: docker fallback ------------------------------------
if not has_backend("docker"):
info("No backend available on this host.")
_print_vm_install_instructions()
die("no backend available; install a VM backend and re-run")
vm = _platform_vm_suggestion()
warn(
"docker is less secure than VM backends — "
"containers share the host kernel."
)
if not prompt:
die(
f"no VM backend available; set BOT_BOTTLE_BACKEND=docker to proceed "
f"with docker, or install the {vm!r} backend."
)
sys.stderr.write(
f"bot-bottle: For better isolation, install the {vm!r} backend.\n"
f" [i] show {vm} install instructions and exit\n"
" [d] use docker anyway\n"
" [q] quit\n"
"bot-bottle: choice [i/d/q]: "
)
sys.stderr.flush()
reply = read_tty_line().strip().lower()
if reply == "d":
return "docker"
if reply == "i":
_print_vm_install_instructions()
die("not proceeding with docker; install a VM backend or set BOT_BOTTLE_BACKEND=docker")
return "docker"
def known_backend_names() -> tuple[str, ...]:
-60
View File
@@ -1,60 +0,0 @@
"""Shared helpers for the consolidated launch sequence (PRD 0070).
Logic that was duplicated across the docker, macos_container, and
firecracker consolidated_launch modules — extracted so each backend
imports it rather than re-implementing it.
"""
from __future__ import annotations
from ..egress import EgressPlan
from ..git_gate import GitGatePlan
from ..orchestrator.client import OrchestratorClient
from ..orchestrator.registration import registration_inputs
from .docker.gateway_provision import GatewayTransport, deprovision_git_gate, provision_git_gate
def provision_bottle(
client: OrchestratorClient,
source_ip: str,
egress_plan: EgressPlan,
git_gate_plan: GitGatePlan,
transport: GatewayTransport,
*,
image_ref: str = "",
tokens: dict[str, str] | None = None,
):
"""Register the bottle and provision its git-gate state. Rolls back the
registration if provisioning fails so no orphan is left. Returns the
`RegisteredBottle` from the orchestrator."""
inputs = registration_inputs(egress_plan)
reg = client.register_bottle(
source_ip, image_ref=image_ref, policy=inputs.policy,
metadata=inputs.metadata, tokens=tokens,
)
try:
provision_git_gate(transport, reg.bottle_id, git_gate_plan)
except Exception:
client.teardown_bottle(reg.bottle_id)
raise
return reg
def teardown_consolidated(
bottle_id: str,
transport: GatewayTransport,
*,
orchestrator_url: str,
timeout: float | None = None,
) -> None:
"""Deregister the bottle and remove its git-gate state. Both steps are
idempotent so this is safe from a cleanup trap."""
from ..orchestrator.config_store import DEFAULT_TEARDOWN_TIMEOUT_SECONDS
OrchestratorClient(
orchestrator_url,
timeout=timeout if timeout is not None else DEFAULT_TEARDOWN_TIMEOUT_SECONDS,
).teardown_bottle(bottle_id)
deprovision_git_gate(transport, bottle_id)
__all__ = ["provision_bottle", "teardown_consolidated"]
+3 -9
View File
@@ -31,7 +31,7 @@ from ...env import ResolvedEnv
from ...git_gate import GitGatePlan
from ...supervise import SupervisePlan
from ...manifest import Manifest
from .. import ActiveAgent, BottleBackend, BottleImages, BottleSpec
from .. import ActiveAgent, BottleBackend, BottleSpec
from . import cleanup as _cleanup
from . import enumerate as _enumerate
from . import launch as _launch
@@ -100,15 +100,9 @@ class DockerBottleBackend(BottleBackend["DockerBottlePlan", "DockerBottleCleanup
stage_dir=stage_dir,
)
def prelaunch_checks(self, plan: DockerBottlePlan) -> None:
_launch.stale_checks(plan)
def _build_or_load_images(self, plan: DockerBottlePlan) -> BottleImages:
return _launch.build_or_load_images(plan)
@contextmanager
def _launch_impl(self, plan: DockerBottlePlan, images: BottleImages) -> Generator[DockerBottle, None, None]:
with _launch.launch(plan, images, provision=self.provision) as bottle:
def launch(self, plan: DockerBottlePlan) -> Generator[DockerBottle, None, None]:
with _launch.launch(plan, provision=self.provision) as bottle:
yield bottle
def ensure_orchestrator(self) -> str:
@@ -1,13 +1,19 @@
"""Consolidated bottle launch sequence for the docker backend (PRD 0070).
Composes the orchestrator primitives into the register/teardown sequence:
Composes the orchestrator primitives into the register/teardown sequence that
replaces the per-bottle gateway:
1. ensure the single infra container (control plane + gateway) is up;
2. allocate the bottle a pinned source IP on the gateway network;
3. register it and provision its git-gate repos/creds into the gateway.
1. ensure the orchestrator control plane + shared gateway are up;
2. allocate the bottle a pinned source IP on the gateway network (the
attribution key), skipping the gateway's own address + live bottles;
3. register it (egress policy blob + slug metadata) bottle id + identity
token;
4. provision its git-gate repos/creds into the running gateway.
Returns a `LaunchContext` with everything the agent container needs to
attach. The agent `docker run` itself is the backend's job; this owns the
It returns a `LaunchContext` with everything the agent container needs to
attach network, pinned IP, the gateway's address (its proxy target), the
orchestrator URL, and the identity token. The agent `docker run` itself is
the backend's job (it owns provider provisioning); this owns the
orchestrator-facing wiring so that sequence stays testable in isolation.
"""
@@ -19,12 +25,15 @@ from ...docker_cmd import run_docker
from ...egress import EgressPlan
from ...git_gate import GitGatePlan
from ...orchestrator.client import OrchestratorClient
from ...orchestrator.gateway import GATEWAY_NETWORK
from ...orchestrator.lifecycle import INFRA_NAME, OrchestratorService
from ..consolidated_util import provision_bottle
from ..consolidated_util import teardown_consolidated as _teardown_util
from .gateway_provision import DockerGatewayTransport
from ...orchestrator.gateway import GATEWAY_NAME, GATEWAY_NETWORK
from ...orchestrator.lifecycle import OrchestratorService
from ...orchestrator.registration import registration_inputs
from .gateway_net import next_free_ip
from .gateway_provision import (
DockerGatewayTransport,
deprovision_git_gate,
provision_git_gate,
)
class ConsolidatedLaunchError(RuntimeError):
@@ -66,21 +75,24 @@ def _container_ip(name: str, network: str) -> str:
ip = proc.stdout.strip()
if proc.returncode != 0 or not ip:
raise ConsolidatedLaunchError(
f"container {name} has no address on {network}: {proc.stderr.strip()}"
f"gateway {name} has no address on {network}: {proc.stderr.strip()}"
)
return ip
def _network_container_ips(network: str) -> list[str]:
"""Every address currently assigned on the gateway network — the ground
truth for "in use": the infra container and every live agent. Read from
the network so a new bottle can't collide with anything actually attached."""
truth for "in use": the gateway + orchestrator infrastructure containers
and every live agent. Read from the network so a new bottle can't collide
with anything actually attached (a registry-only view would miss the
orchestrator/gateway containers)."""
proc = run_docker([
"docker", "network", "inspect", "--format",
"{{range .Containers}}{{.IPv4Address}} {{end}}", network,
])
ips: list[str] = []
for entry in proc.stdout.split():
# entries look like "172.20.0.2/16" — keep the address.
ips.append(entry.split("/", 1)[0])
return ips
@@ -92,24 +104,33 @@ def launch_consolidated(
image_ref: str = "",
tokens: dict[str, str] | None = None,
service: OrchestratorService | None = None,
infra_name: str = INFRA_NAME,
gateway_name: str = GATEWAY_NAME,
network: str = GATEWAY_NETWORK,
) -> LaunchContext:
"""Ensure the infra container is up, allocate + register the bottle, and
provision its git-gate state. Returns the agent's attach context."""
"""Ensure the orchestrator + gateway are up, allocate + register the
bottle, and provision its git-gate state. Returns the agent's attach
context. Raises `ConsolidatedLaunchError` (or the primitives' own errors)
if any step fails the caller tears down on failure."""
service = service or OrchestratorService()
url = service.ensure_running()
client = OrchestratorClient(url)
cidr = _network_cidr(network)
gateway_ip = _container_ip(infra_name, network)
gateway_ip = _container_ip(gateway_name, network)
source_ip = next_free_ip(cidr, _network_container_ips(network))
transport = DockerGatewayTransport(infra_name)
reg = provision_bottle(
client, source_ip, egress_plan, git_gate_plan, transport,
image_ref=image_ref, tokens=tokens,
inputs = registration_inputs(egress_plan)
reg = client.register_bottle(
source_ip, image_ref=image_ref, policy=inputs.policy,
metadata=inputs.metadata, tokens=tokens,
)
try:
provision_git_gate(
DockerGatewayTransport(gateway_name), reg.bottle_id, git_gate_plan)
except Exception:
# Roll the registration back so a provisioning failure leaves no orphan.
client.teardown_bottle(reg.bottle_id)
raise
return LaunchContext(
bottle_id=reg.bottle_id,
identity_token=reg.identity_token,
@@ -121,12 +142,12 @@ def launch_consolidated(
def teardown_consolidated(
bottle_id: str, *, orchestrator_url: str, infra_name: str = INFRA_NAME,
timeout: float | None = None,
bottle_id: str, *, orchestrator_url: str, gateway_name: str = GATEWAY_NAME,
) -> None:
"""Deregister the bottle and remove its git-gate state. Idempotent."""
_teardown_util(bottle_id, DockerGatewayTransport(infra_name),
orchestrator_url=orchestrator_url, timeout=timeout)
"""Deregister the bottle and remove its git-gate state from the gateway.
Both steps are idempotent so this is safe from a cleanup trap."""
OrchestratorClient(orchestrator_url).teardown_bottle(bottle_id)
deprovision_git_gate(DockerGatewayTransport(gateway_name), bottle_id)
__all__ = [
+23 -59
View File
@@ -42,9 +42,7 @@ from ...git_gate import (
provision_git_gate_dynamic_keys,
revoke_git_gate_provisioned_keys,
)
from ...image_cache import check_stale
from ...log import die, info, warn
from .. import BottleImages
from ...log import info, warn
from . import util as docker_mod
from .bottle import DockerBottle
from .bottle_plan import DockerBottlePlan
@@ -64,7 +62,6 @@ from .compose import (
write_compose_file,
)
from .consolidated_compose import consolidated_agent_compose
from ...orchestrator.config_store import resolve_teardown_timeout
from .consolidated_launch import launch_consolidated, teardown_consolidated
from ...orchestrator.gateway import DockerGateway
@@ -73,47 +70,16 @@ from ...orchestrator.gateway import DockerGateway
_REPO_DIR = str(Path(__file__).resolve().parent.parent.parent.parent)
def build_or_load_images(plan: DockerBottlePlan) -> BottleImages:
"""Resolve the agent image ref for this plan.
Returns the committed snapshot if one exists, the cached image when the
policy is 'cached', or builds a fresh image and returns that."""
committed = read_committed_image(plan.slug)
if committed and docker_mod.image_exists(committed):
info(f"using committed image {committed!r}")
return BottleImages(agent=committed)
if plan.spec.image_policy == "cached":
if not docker_mod.image_exists(plan.image):
die(
f"cached agent image {plan.image!r} not found; "
"run without --cached-images to build it"
)
info(f"using cached agent image {plan.image!r}")
return BottleImages(agent=plan.image)
docker_mod.build_image(plan.image, _REPO_DIR, dockerfile=plan.dockerfile_path)
docker_mod.verify_agent_image(
plan.image, runtime_for(plan.agent_provider_template).smoke_test,
)
return BottleImages(agent=plan.image)
@contextmanager
def launch(
plan: DockerBottlePlan,
images: BottleImages,
*,
provision: Callable[[DockerBottlePlan, "DockerBottle"], str | None],
) -> Generator[DockerBottle, None, None]:
"""Launch and provision a Docker bottle via compose. Teardown on exit."""
"""Build, launch, and provision a Docker bottle via compose.
Teardown on exit."""
stack = ExitStack()
# Stamp the resolved agent image ref into the plan so compose rendering
# picks up the right image (may be a committed snapshot or cached ref).
plan = dataclasses.replace(
plan,
agent_provision=dataclasses.replace(plan.agent_provision, image=str(images.agent)),
)
_bottle_for_revoke = plan.manifest.bottle
_git_gate_dir_for_revoke = git_gate_state_dir(plan.slug)
@@ -130,6 +96,25 @@ def launch(
)
try:
# Step 1: agent image. Use a committed snapshot when one exists
# and is present in the local daemon; otherwise build from the
# Dockerfile. (The gateway image is built by the orchestrator.)
committed = read_committed_image(plan.slug)
if committed and docker_mod.image_exists(committed):
info(f"using committed image {committed!r}")
plan = dataclasses.replace(
plan,
agent_provision=dataclasses.replace(plan.agent_provision, image=committed),
)
else:
docker_mod.build_image(
plan.image, _REPO_DIR,
dockerfile=plan.dockerfile_path,
)
docker_mod.verify_agent_image(
plan.image, runtime_for(plan.agent_provider_template).smoke_test,
)
# Step 2: mint the git-gate dynamic (gitea) deploy keys, if any, before
# provisioning the bottle's repos into the shared gateway.
git_gate_plan = plan.git_gate_plan
@@ -148,14 +133,11 @@ def launch(
token_values = egress_resolve_token_values(
plan.egress_plan.token_env_map, effective_env,
)
teardown_timeout = resolve_teardown_timeout()
ctx = launch_consolidated(
plan.egress_plan, git_gate_plan, image_ref=plan.image, tokens=token_values,
)
stack.callback(
teardown_consolidated, ctx.bottle_id,
orchestrator_url=ctx.orchestrator_url,
timeout=teardown_timeout,
teardown_consolidated, ctx.bottle_id, orchestrator_url=ctx.orchestrator_url,
)
# Step 4: install the SHARED gateway CA into the agent (replaces the
@@ -225,21 +207,3 @@ def launch(
yield bottle
finally:
teardown()
def stale_checks(plan: DockerBottlePlan) -> None:
"""Raise StaleImageError if a cached image is older than the configured
threshold. Only runs when image_policy is 'cached'. Called by the backend
class's _image_stale_checks before _launch_impl starts any resources."""
if plan.spec.image_policy != "cached":
return
committed = read_committed_image(plan.slug)
if committed and docker_mod.image_exists(committed):
ts = docker_mod.image_created_at(committed)
if ts is not None:
check_stale(f"agent image {committed!r}", ts)
return
if docker_mod.image_exists(plan.image):
ts = docker_mod.image_created_at(plan.image)
if ts is not None:
check_stale(f"agent image {plan.image!r}", ts)
+4 -8
View File
@@ -27,14 +27,10 @@ def _docker_on_path() -> bool:
def _daemon_reachable() -> bool:
if not _docker_on_path():
return False
try:
return subprocess.run(
["docker", "info"],
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
check=False, timeout=5,
).returncode == 0
except subprocess.TimeoutExpired:
return False
return subprocess.run(
["docker", "info"],
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, check=False,
).returncode == 0
def _print_install_pointer() -> None:
-44
View File
@@ -5,7 +5,6 @@ existence, and building images."""
from __future__ import annotations
import os
from datetime import datetime, timezone
import re
import shutil
import subprocess
@@ -198,46 +197,3 @@ def commit_container(container_name: str, image_tag: str) -> None:
f"{(result.stderr or '').strip() or '<no stderr>'}"
)
info(f"committed {container_name!r}{image_tag!r}")
def image_created_at(ref: str) -> datetime | None:
"""Return Docker's image Created timestamp as an aware UTC datetime, or
None when the field is absent or unparseable. Callers should skip the
stale check when None is returned."""
r = subprocess.run(
["docker", "image", "inspect", "--format", "{{.Created}}", ref],
capture_output=True,
text=True,
check=False,
)
if r.returncode != 0:
die(
f"docker image inspect for {ref!r} failed: "
f"{(r.stderr or '').strip() or '<no stderr>'}"
)
raw = r.stdout.strip()
if not raw:
return None
try:
return _parse_docker_timestamp(raw)
except ValueError:
return None
def _parse_docker_timestamp(raw: str) -> datetime:
text = raw.strip()
if text.endswith("Z"):
text = text[:-1] + "+00:00"
dot = text.find(".")
if dot != -1:
tz_plus = text.find("+", dot)
tz_minus = text.find("-", dot)
tz_candidates = [pos for pos in (tz_plus, tz_minus) if pos != -1]
if tz_candidates:
tz_pos = min(tz_candidates)
frac = text[dot + 1:tz_pos]
text = text[:dot + 1] + frac[:6].ljust(6, "0") + text[tz_pos:]
dt = datetime.fromisoformat(text)
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt.astimezone(timezone.utc)
+4 -11
View File
@@ -18,7 +18,7 @@ from ...env import ResolvedEnv
from ...git_gate import GitGatePlan
from ...manifest import Manifest
from ...supervise import SupervisePlan
from .. import ActiveAgent, BottleBackend, BottleImages, BottleSpec
from .. import ActiveAgent, BottleBackend, BottleSpec
from . import cleanup as _cleanup
from . import enumerate as _enumerate
from . import launch as _launch
@@ -92,18 +92,11 @@ class FirecrackerBottleBackend(
stage_dir=stage_dir,
)
def _build_or_load_images(self, plan: FirecrackerBottlePlan) -> BottleImages:
return BottleImages(agent=_launch.build_or_load_agent_base(plan))
def prelaunch_checks(self, plan: FirecrackerBottlePlan) -> None:
_launch.stale_checks(plan)
@contextmanager
def _launch_impl(
self, plan: FirecrackerBottlePlan, images: BottleImages,
def launch(
self, plan: FirecrackerBottlePlan
) -> Generator[FirecrackerBottle, None, None]:
assert isinstance(images.agent, Path)
with _launch.launch(plan, images.agent, provision=self.provision) as bottle:
with _launch.launch(plan, provision=self.provision) as bottle:
yield bottle
def prepare_cleanup(self) -> FirecrackerBottleCleanupPlan:
+16 -66
View File
@@ -1,23 +1,8 @@
"""Cleanup for the Firecracker backend.
Reaps *orphans* only resources with no live VM behind them:
* orphan run dirs: a per-bottle run dir (holding the ~1G rootfs.ext4)
whose firecracker process has exited. These leak when a launch is
hard-killed before its teardown runs (host OOM/crash, a cancelled CI
job, `kill -9`); the clean-exit path already removes its own dir in
launch.py.
* orphan VM pids: a firecracker process whose run dir is already gone
a VMM left lingering after its dir was removed.
A run dir with a *live* firecracker process is a running bottle and is
left strictly alone: it is neither killed nor removed. (The backend's
`enumerate_active` registry is still a stub #354 — so a live process
is the only reliable "this bottle is in use" signal we have. Once the
registry lands, registry-orphaned-but-running VMs can be reaped too.)
TAP slots free themselves (the flock drops when the launcher exits), so
there is nothing to reclaim there.
Orphans are: firecracker VMM processes whose config lives under our run
dir, and the per-bottle run dirs. TAP slots free themselves (the flock
drops when the launcher exits), so there is nothing to reclaim there.
"""
from __future__ import annotations
@@ -37,73 +22,38 @@ def _run_root() -> Path:
return util.cache_dir() / "run"
def _run_dir_of(cmd: str, run_root: Path) -> Path | None:
"""The bottle run dir a firecracker cmdline belongs to, or None.
A bottle VM is launched with `--config-file <run_root>/<slug>/config.json`,
so the run dir is the config file's parent when it sits directly under
the run root. Anything else (a builder VM, the infra VM elsewhere) is
not ours to reap here.
"""
toks = cmd.split()
for i, tok in enumerate(toks):
if tok == "--config-file" and i + 1 < len(toks):
parent = Path(toks[i + 1]).parent
if parent.parent == run_root:
return parent
return None
def _scan_processes(run_root: Path) -> tuple[set[str], list[int]]:
"""Inspect running firecracker VMs under ``run_root``.
Returns ``(live_run_dirs, orphan_pids)``:
* ``live_run_dirs`` run dirs backed by a running VM (never reaped);
* ``orphan_pids`` firecracker pids whose run dir no longer exists
(a lingering VMM to kill).
"""
def _orphan_vm_pids() -> list[int]:
"""firecracker processes whose --config-file is under our run dir."""
run_root = str(_run_root())
result = subprocess.run(
["pgrep", "-a", "firecracker"],
capture_output=True, text=True, check=False,
)
if result.returncode != 0:
return set(), []
live: set[str] = set()
orphan_pids: list[int] = []
return []
pids: list[int] = []
for line in result.stdout.splitlines():
parts = line.split(None, 1)
if len(parts) != 2:
if len(parts) != 2 or run_root not in parts[1]:
continue
try:
pid = int(parts[0])
pids.append(int(parts[0]))
except ValueError:
continue
run_dir = _run_dir_of(parts[1], run_root)
if run_dir is None:
continue
if run_dir.is_dir():
live.add(str(run_dir))
else:
orphan_pids.append(pid)
return live, orphan_pids
return pids
def _orphan_run_dirs(run_root: Path, live: set[str]) -> list[str]:
"""Run dirs with no live VM behind them — the leaked ones to remove."""
def _run_dirs() -> list[str]:
run_root = _run_root()
if not run_root.is_dir():
return []
return sorted(
str(p) for p in run_root.iterdir()
if p.is_dir() and str(p) not in live
)
return sorted(str(p) for p in run_root.iterdir() if p.is_dir())
def prepare_cleanup() -> FirecrackerBottleCleanupPlan:
run_root = _run_root()
live, orphan_pids = _scan_processes(run_root)
return FirecrackerBottleCleanupPlan(
vm_pids=tuple(orphan_pids),
run_dirs=tuple(_orphan_run_dirs(run_root, live)),
vm_pids=tuple(_orphan_vm_pids()),
run_dirs=tuple(_run_dirs()),
)
@@ -33,7 +33,8 @@ from ...orchestrator.client import OrchestratorClient
from ...orchestrator.lifecycle import (
OrchestratorStartError, # re-exported so callers can catch it
)
from ..consolidated_util import provision_bottle, teardown_consolidated as _teardown_util
from ...orchestrator.registration import registration_inputs
from ..docker.gateway_provision import deprovision_git_gate, provision_git_gate
from . import infra_vm
@@ -67,11 +68,18 @@ def launch_consolidated(
url = infra.control_plane_url
client = OrchestratorClient(url)
transport = infra_vm.gateway_transport()
reg = provision_bottle(
client, guest_ip, egress_plan, git_gate_plan, transport,
image_ref=image_ref, tokens=tokens,
inputs = registration_inputs(egress_plan)
reg = client.register_bottle(
guest_ip, image_ref=image_ref, policy=inputs.policy,
metadata=inputs.metadata, tokens=tokens,
)
try:
provision_git_gate(
infra_vm.gateway_transport(), reg.bottle_id, git_gate_plan)
except Exception:
client.teardown_bottle(reg.bottle_id)
raise
# The shared gateway CA every agent on this host trusts for TLS
# interception — fetched from the infra VM over SSH.
return LaunchContext(
@@ -83,15 +91,13 @@ def launch_consolidated(
)
def teardown_consolidated(
bottle_id: str, *, orchestrator_url: str, timeout: float | None = None,
) -> None:
def teardown_consolidated(bottle_id: str, *, orchestrator_url: str) -> None:
"""Deregister the bottle and remove its git-gate state from the gateway
VM. Both steps are idempotent so this is safe from a cleanup trap. Does
NOT stop the infra VM it's a persistent per-host singleton shared by
every bottle."""
_teardown_util(bottle_id, infra_vm.gateway_transport(),
orchestrator_url=orchestrator_url, timeout=timeout)
OrchestratorClient(orchestrator_url).teardown_bottle(bottle_id)
deprovision_git_gate(infra_vm.gateway_transport(), bottle_id)
__all__ = [
+1 -1
View File
@@ -5,7 +5,7 @@ backend — we stream the guest root filesystem out over the control
channel (SSH here). Unlike the other backends this needs no Docker: the
tar *is* the resumable artifact. `resume` extracts it and rebuilds a
fresh per-bottle ext4 with `mke2fs -d` (see `util.build_committed_rootfs_dir`
and `launch.build_or_load_agent_base`). The bottle keeps running after the
and `launch._build_agent_base`). The bottle keeps running after the
snapshot.
"""
@@ -38,47 +38,27 @@ _BUILD_TIMEOUT_SECONDS = 900.0
def _dockerfile_hash(dockerfile: Path) -> str:
"""The Dockerfile's content hash. The shipped agent Dockerfiles COPY
nothing from the build context (see .dockerignore), so their content fully
determines the built image; a Dockerfile that adds COPY will want the
"""Cache key: the Dockerfile's content. The shipped agent Dockerfiles
COPY nothing from the build context (see .dockerignore), so their content
fully determines the image; a Dockerfile that adds COPY will want the
context folded in here too."""
return hashlib.sha256(dockerfile.read_bytes()).hexdigest()[:16]
def _rootfs_digest(dockerfile: Path) -> str:
"""Cache key for the built AND boot-injected agent rootfs. Two inputs
determine the on-disk rootfs: the Dockerfile (the image) and the guest init
injected into it (`util._GUEST_INIT`). Folding the init in means a fix to
it e.g. making /tmp world-writable busts the cache instead of silently
reusing a stale rootfs built with the old init."""
h = hashlib.sha256()
h.update(_dockerfile_hash(dockerfile).encode())
h.update(b"\0")
h.update(util._GUEST_INIT.encode())
return h.hexdigest()[:16]
def cached_agent_rootfs_dir(dockerfile: Path) -> Path | None:
"""Return the ready cached rootfs for ``dockerfile``, if one exists."""
base = util.cache_dir() / "rootfs" / f"agent-{_rootfs_digest(dockerfile)}"
return base if (base / ".bb-ready").is_file() else None
def build_agent_rootfs_dir(
dockerfile: Path, *, image_tag: str, smoke_test: tuple[str, ...] = (),
) -> Path:
"""Build `dockerfile` in the infra VM (buildah, no host docker), export its
rootfs, inject the guest boot bits, and return the cached base dir the
same shape `util.build_rootfs_ext4` consumes. Cached by Dockerfile content
+ injected guest init, so a repeat launch skips the rebuild but an init or
Dockerfile change rebuilds.
same shape `util.build_rootfs_ext4` consumes. Cached by Dockerfile content,
so a repeat launch skips the rebuild.
`smoke_test` (the provider's declared argv, e.g. `("claude","--version")`)
is run in the freshly built image before export, catching an npm
silent-failure image at build time rather than at first agent use."""
digest = _rootfs_digest(dockerfile)
digest = _dockerfile_hash(dockerfile)
base = util.cache_dir() / "rootfs" / f"agent-{digest}"
if cached_agent_rootfs_dir(dockerfile) is not None:
if (base / ".bb-ready").is_file():
info(f"using cached agent rootfs {base.name}")
return base
@@ -41,7 +41,7 @@ from . import util
_ARTIFACT_FORMAT = "1"
_REPO_ROOT = Path(__file__).resolve().parents[3]
_DOCKERFILES = ("Dockerfile.orchestrator", "Dockerfile.gateway", "Dockerfile.infra", "Dockerfile.infra.fc")
_DOCKERFILES = ("Dockerfile.orchestrator", "Dockerfile.gateway", "Dockerfile.infra")
_DEFAULT_BASE = "https://gitea.dideric.is"
_DEFAULT_OWNER = "didericis"
@@ -85,11 +85,6 @@ def infra_artifact_version(init_script: str, *, repo_root: Path = _REPO_ROOT) ->
h.update(name.encode())
h.update(b"\0")
h.update((repo_root / name).read_bytes())
h.update(b"pyproject.toml\0")
h.update((repo_root / "pyproject.toml").read_bytes())
h.update(b"dropbear\0")
dropbear = util.dropbear_path()
h.update(dropbear.read_bytes() if dropbear.is_file() else b"<missing>")
h.update(b"init\0")
h.update(init_script.encode())
return h.hexdigest()[:16]
@@ -116,7 +111,6 @@ def artifact_url(version: str, filename: str) -> str:
_GZ_NAME = "rootfs.ext4.gz"
_SHA_NAME = "rootfs.ext4.gz.sha256"
_CANDIDATE_DIR_ENV = "BOT_BOTTLE_INFRA_ARTIFACT_DIR"
def _cache_root(version: str) -> Path:
@@ -166,33 +160,6 @@ def ensure_artifact_gz(version: str) -> Path:
"""The verified, cached `rootfs.ext4.gz` for `version` — downloading it (and
its `.sha256`) once, then reusing it. Fail-closed on a checksum mismatch:
the partial is removed and we die rather than boot an unverified rootfs."""
candidate_dir = os.environ.get(_CANDIDATE_DIR_ENV, "").strip()
if candidate_dir:
root = Path(candidate_dir)
version_file = root / "version.txt"
# Guard the read so a missing version.txt is a clean error, not a raw
# FileNotFoundError.
if not version_file.is_file():
die(f"infra candidate bundle is incomplete: {root}")
declared = version_file.read_text(encoding="utf-8").strip()
if declared != version:
die(
f"infra candidate version mismatch: expected {version}, "
f"bundle contains {declared or '<empty>'}"
)
gz = root / _GZ_NAME
sha = root / _SHA_NAME
if not gz.is_file() or not sha.is_file():
die(f"infra candidate bundle is incomplete: {root}")
expected = sha.read_text().split()[0].strip().lower()
actual = _sha256_file(gz)
if actual != expected:
die(
f"infra candidate checksum mismatch for {version}:\n"
f" expected {expected}\n actual {actual}"
)
return gz
root = _cache_root(version)
root.mkdir(parents=True, exist_ok=True)
gz = root / _GZ_NAME
+18 -82
View File
@@ -33,7 +33,6 @@ from pathlib import Path
from typing import Generator
from ...log import die, info
from .. import util as backend_util
from ..docker import util as docker_mod
from ..docker.gateway_provision import GatewayProvisionError
from . import firecracker_vm, infra_artifact, netpool, util
@@ -94,18 +93,19 @@ class InfraVm:
"""The gateway's mitmproxy CA (PEM) that agents install to trust its
TLS interception. Generated a moment after boot, so this polls over
SSH until it appears (mirrors DockerGateway.ca_cert_pem)."""
def _fetch() -> str | None:
deadline = time.monotonic() + timeout
while True:
proc = subprocess.run(
util.ssh_base_argv(self.private_key, self.guest_ip)
+ [f"cat {_GATEWAY_CA_PATH}"],
capture_output=True, text=True, timeout=15, check=False,
)
ok = proc.returncode == 0 and "BEGIN CERTIFICATE" in proc.stdout
return proc.stdout if ok else None
try:
return backend_util.poll_ca_cert(_fetch, timeout=timeout)
except TimeoutError as exc:
die(str(exc))
if proc.returncode == 0 and "BEGIN CERTIFICATE" in proc.stdout:
return proc.stdout
if time.monotonic() >= deadline:
die(f"gateway CA not available after {timeout:g}s: "
f"{proc.stderr.strip() or 'empty'}")
time.sleep(_HEALTH_POLL_SECONDS)
def ensure_built() -> None:
@@ -125,19 +125,16 @@ def ensure_built() -> None:
def build_infra_images_with_docker() -> None:
"""Build the four fixed images from source with host Docker: orchestrator,
gateway, the shared infra base (Dockerfile.infra), then the Firecracker
infra image (Dockerfile.infra.fc: FROM infra + buildah). The launch host
uses this only in `BOT_BOTTLE_INFRA_BUILD=local` mode; `publish_infra`
uses it off-host to produce the published artifact."""
"""Build the three fixed images from source with host Docker: orchestrator,
gateway, then the combined infra image (`COPY --from` orchestrator, `FROM`
gateway). The launch host uses this only in `BOT_BOTTLE_INFRA_BUILD=local`
mode; `publish_infra` uses it off-host to produce the published artifact."""
docker_mod.build_image(
_ORCHESTRATOR_IMAGE, str(_REPO_ROOT), dockerfile="Dockerfile.orchestrator")
docker_mod.build_image(
_GATEWAY_IMAGE, str(_REPO_ROOT), dockerfile="Dockerfile.gateway")
docker_mod.build_image(
"bot-bottle-infra:latest", str(_REPO_ROOT), dockerfile="Dockerfile.infra")
docker_mod.build_image(
_INFRA_IMAGE, str(_REPO_ROOT), dockerfile="Dockerfile.infra.fc")
_INFRA_IMAGE, str(_REPO_ROOT), dockerfile="Dockerfile.infra")
def build_infra_rootfs_dir() -> Path:
@@ -164,23 +161,20 @@ def ensure_running() -> InfraVm:
slot = netpool.orch_slot()
url = f"http://{slot.guest_ip}:{CONTROL_PLANE_PORT}"
key = _infra_dir() / "id_ed25519"
want = _expected_version()
if _adoptable(key, url, want):
if key.exists() and _health_ok(url):
info(f"adopting running infra VM at {url}")
return InfraVm(guest_ip=slot.guest_ip, private_key=key)
with _singleton_lock():
# Re-check under the lock: another launcher may have booted it while
# we waited for the lock (double-checked, so we adopt not re-boot).
if _adoptable(key, url, want):
if key.exists() and _health_ok(url):
info(f"adopting running infra VM at {url}")
return InfraVm(guest_ip=slot.guest_ip, private_key=key)
# Clear a stale/hung/OUTDATED VM holding the link before booting fresh.
stop()
stop() # clear a stale/hung VM holding the link before booting fresh
ensure_built()
infra = boot()
wait_for_health(infra)
_record_booted_version(want)
return infra
@@ -199,15 +193,9 @@ def _singleton_lock() -> Generator[None, None, None]:
def stop() -> None:
"""Stop the infra VM singleton (idempotent — absent is success). Reaps the
recorded VMM AND any orphaned firecracker still bound to the infra config
the PID file drifts after crashes / out-of-band kills, and a survivor would
hold the orchestrator TAP so the next boot dies with "tap … Resource busy".
Drops the version marker so a stopped VM is never treated as adoptable."""
"""Stop the infra VM singleton (idempotent — absent is success)."""
_kill_pidfile()
_kill_infra_firecrackers()
_pid_file().unlink(missing_ok=True)
_version_file().unlink(missing_ok=True)
def boot() -> InfraVm:
@@ -250,36 +238,6 @@ def _pid_file() -> Path:
return _infra_dir() / "vm.pid"
def _version_file() -> Path:
"""Records the infra-artifact version the *running* VM booted from, so a
later launcher can tell whether the singleton it found is the current code.
Without it, a healthy VM built from an older image gets adopted forever and
the new code never boots every infra change would need an out-of-band
kill to dislodge the stale VM (and races whatever launched next)."""
return _infra_dir() / "booted-version"
def _expected_version() -> str:
return infra_artifact.infra_artifact_version(_infra_init())
def _adoptable(key: Path, url: str, want: str) -> bool:
"""Adopt a running infra VM only if it booted from the CURRENT version and
its control plane is healthy. A missing/mismatched marker means a prior
launcher booted an older infra image reboot rather than reuse stale code."""
if not key.exists():
return False
try:
booted = _version_file().read_text(encoding="utf-8").strip()
except OSError:
return False
return booted == want and _health_ok(url)
def _record_booted_version(version: str) -> None:
_version_file().write_text(version + "\n", encoding="utf-8")
# The registry "volume": a host-side ext4 file attached to the infra VM as a
# second virtio-block device (guest /dev/vdb), mounted at the control plane's
# DB dir. It outlives the ephemeral rootfs, so the bottle registry survives an
@@ -351,28 +309,6 @@ def _kill_pidfile() -> None:
pass
def _kill_infra_firecrackers(proc_root: Path = Path("/proc")) -> None:
"""SIGKILL any firecracker VMM whose `--config-file` is this host's infra
config, independent of the PID file reaps orphans it lost track of so the
orchestrator TAP is free to rebind. Scoped to the infra config path, so the
interactive pool's agent/infra VMs (other config paths) are untouched."""
cfg = str(_infra_dir() / "config.json")
for entry in proc_root.iterdir():
if not entry.name.isdigit():
continue
try:
if (entry / "comm").read_text().strip() != "firecracker":
continue
args = (entry / "cmdline").read_bytes().split(b"\0")
except OSError:
continue # process vanished / not ours
if any(a.decode("utf-8", "replace") == cfg for a in args):
try:
os.kill(int(entry.name), signal.SIGKILL)
except (OSError, ValueError):
pass
def _health_ok(url: str) -> bool:
try:
with urllib.request.urlopen(f"{url}/health", timeout=1.0) as resp:
@@ -497,7 +433,7 @@ BOT_BOTTLE_ROOT=/var/lib/bot-bottle python3 -m bot_bottle.orchestrator \\
BOT_BOTTLE_GATEWAY_DAEMONS=egress,git-http,supervise \\
BOT_BOTTLE_ORCHESTRATOR_URL=http://127.0.0.1:{CONTROL_PLANE_PORT} \\
SUPERVISE_DB_PATH=/var/lib/bot-bottle/db/bot-bottle.db \\
python3 -m bot_bottle.gateway_init &
python3 /app/gateway_init.py &
# Reap as PID 1; children are backgrounded, so `wait` blocks.
while : ; do wait ; done
+13 -47
View File
@@ -26,7 +26,6 @@ from __future__ import annotations
import dataclasses
import os
import shutil
from contextlib import ExitStack, contextmanager
from pathlib import Path
from typing import Callable, Generator
@@ -46,15 +45,13 @@ from ...git_gate import (
provision_git_gate_dynamic_keys,
revoke_git_gate_provisioned_keys,
)
from ...image_cache import check_stale_path
from ...log import die, info, warn
from ...log import info, warn
from ...supervise import SUPERVISE_PORT
from ..docker.egress import EGRESS_PORT
from ..util import AGENT_CA_BUNDLE, AGENT_CA_PATH
from . import firecracker_vm, image_builder, isolation_probe, netpool, util
from .bottle import FirecrackerBottle
from .bottle_plan import FirecrackerBottlePlan
from ...orchestrator.config_store import resolve_teardown_timeout
from .consolidated_launch import (
launch_consolidated,
teardown_consolidated,
@@ -67,7 +64,6 @@ _GIT_HTTP_PORT = 9420
@contextmanager
def launch(
plan: FirecrackerBottlePlan,
agent_base: Path,
*,
provision: Callable[[FirecrackerBottlePlan, "FirecrackerBottle"], str | None],
) -> Generator[FirecrackerBottle, None, None]:
@@ -89,9 +85,11 @@ def launch(
raise teardown_exc
try:
# Step 1 (rootfs resolution/build) runs in BottleBackend.launch before
# this context starts resources. ``agent_base`` is the selected cache,
# fresh build, or committed snapshot.
# Step 1: agent rootfs. Built from the Dockerfile inside a Firecracker
# builder VM (buildah, no host docker); a committed snapshot is reused
# when present. Returns the base dir the per-bottle ext4 is made from.
plan, agent_base = _build_agent_base(plan)
# Step 2: mint the git-gate dynamic (gitea) deploy keys, if any.
git_gate_plan = plan.git_gate_plan
if git_gate_plan.upstreams:
@@ -114,7 +112,6 @@ def launch(
token_values = egress_resolve_token_values(
plan.egress_plan.token_env_map, effective_env,
)
teardown_timeout = resolve_teardown_timeout()
ctx = launch_consolidated(
plan.egress_plan, git_gate_plan,
guest_ip=slot.guest_ip,
@@ -124,7 +121,6 @@ def launch(
stack.callback(
teardown_consolidated, ctx.bottle_id,
orchestrator_url=ctx.orchestrator_url,
timeout=teardown_timeout,
)
# Step 5: install the SHARED gateway CA (replaces the per-bottle CA).
@@ -168,10 +164,6 @@ def launch(
# Step 6: build the per-bottle rootfs + SSH key, then boot.
run_dir = util.cache_dir() / "run" / plan.slug
run_dir.mkdir(parents=True, exist_ok=True)
# Remove the run dir on teardown so the per-bottle rootfs.ext4 (~1G)
# doesn't leak. Registered before vm.terminate below so it runs *after*
# it (ExitStack is LIFO): the VM is gone before we rm its rootfs.
stack.callback(lambda: shutil.rmtree(run_dir, ignore_errors=True))
rootfs = run_dir / "rootfs.ext4"
util.build_rootfs_ext4(agent_base, rootfs)
private_key, pubkey = util.generate_keypair(run_dir)
@@ -214,7 +206,9 @@ def launch(
teardown()
def build_or_load_agent_base(plan: FirecrackerBottlePlan) -> Path:
def _build_agent_base(
plan: FirecrackerBottlePlan,
) -> tuple[FirecrackerBottlePlan, Path]:
"""Produce the agent's base rootfs dir. Primary path: build the Dockerfile
inside a Firecracker builder VM (buildah, no host docker), smoke-testing
the image before export. A committed snapshot (freeze/migrate) is resumed
@@ -223,36 +217,13 @@ def build_or_load_agent_base(plan: FirecrackerBottlePlan) -> Path:
committed_tar = committed_rootfs_path(plan.slug)
if committed and committed_tar.is_file():
info(f"resuming from committed rootfs {committed_tar}")
return util.build_committed_rootfs_dir(committed_tar)
dockerfile = Path(plan.dockerfile_path)
if plan.spec.image_policy == "cached":
cached = image_builder.cached_agent_rootfs_dir(dockerfile)
if cached is None:
die(
f"cached agent rootfs for {plan.image!r} not found; "
"run without --cached-images to build it"
)
info(f"using cached agent rootfs {cached.name}")
return cached
return image_builder.build_agent_rootfs_dir(
dockerfile,
return plan, util.build_committed_rootfs_dir(committed_tar)
base = image_builder.build_agent_rootfs_dir(
Path(plan.dockerfile_path),
image_tag=plan.image,
smoke_test=runtime_for(plan.agent_provider_template).smoke_test,
)
def stale_checks(plan: FirecrackerBottlePlan) -> None:
"""Raise when the cached rootfs selected by this plan is stale."""
if plan.spec.image_policy != "cached":
return
committed = read_committed_image(plan.slug)
committed_tar = committed_rootfs_path(plan.slug)
if committed and committed_tar.is_file():
check_stale_path(f"agent rootfs {committed_tar}", committed_tar)
return
cached = image_builder.cached_agent_rootfs_dir(Path(plan.dockerfile_path))
if cached is not None:
check_stale_path(f"agent rootfs {cached}", cached / ".bb-ready")
return plan, base
# --- agent guest env -------------------------------------------------
@@ -268,11 +239,6 @@ def _agent_guest_env(plan: FirecrackerBottlePlan, host_ip: str) -> dict[str, str
"HTTPS_PROXY": proxy_url, "HTTP_PROXY": proxy_url,
"https_proxy": proxy_url, "http_proxy": proxy_url,
"NO_PROXY": no_proxy, "no_proxy": no_proxy,
# Rootfs export can leave Git's implicit XDG paths unreadable even
# after the runtime repair. Bypass that discovery and name the
# provisioned global config explicitly so insteadOf can never fall
# through to the credential-bearing upstream URL.
"GIT_CONFIG_GLOBAL": f"{plan.guest_home}/.gitconfig",
"NODE_EXTRA_CA_CERTS": AGENT_CA_PATH,
"SSL_CERT_FILE": AGENT_CA_BUNDLE,
"REQUESTS_CA_BUNDLE": AGENT_CA_BUNDLE,
+21 -97
View File
@@ -11,8 +11,7 @@ The `<version>` is `infra_artifact.infra_artifact_version(...)`, the content
hash of the rootfs inputs, so a launch host at the same code checkout resolves
the exact artifact this produced.
python3 -m bot_bottle.backend.firecracker.publish_infra --output DIR
python3 -m bot_bottle.backend.firecracker.publish_infra --publish-dir DIR
python3 -m bot_bottle.backend.firecracker.publish_infra [--dry-run] [--force]
Auth: a token with `write:package` on the target owner, from
`BOT_BOTTLE_INFRA_ARTIFACT_TOKEN`.
@@ -25,6 +24,7 @@ import gzip
import hashlib
import shutil
import sys
import tempfile
import urllib.error
import urllib.request
from pathlib import Path
@@ -131,112 +131,36 @@ def build_artifact(out_dir: Path) -> tuple[str, Path, Path]:
return version, gz, sha
def _try_download_published(out_dir: Path) -> tuple[str, Path, Path] | None:
"""If this version's artifact is already in the registry, download the gz
and sha to out_dir and return (version, gz_path, sha_path). Returns None
when not yet published."""
version = infra_artifact.infra_artifact_version(infra_vm._infra_init())
sha_url = infra_artifact.artifact_url(version, "rootfs.ext4.gz.sha256")
try:
with urllib.request.urlopen(infra_artifact._open(sha_url)):
pass
except urllib.error.HTTPError as e:
if e.code == 404:
return None
raise SystemExit(f"registry check failed (HTTP {e.code}): {sha_url}")
except urllib.error.URLError as e:
raise SystemExit(f"registry unreachable: {sha_url} ({e.reason})")
print(f"infra rootfs {version} already published — downloading instead of building")
gz = out_dir / "rootfs.ext4.gz"
sha = out_dir / "rootfs.ext4.gz.sha256"
infra_artifact._download(infra_artifact.artifact_url(version, "rootfs.ext4.gz"), gz)
infra_artifact._download(infra_artifact.artifact_url(version, "rootfs.ext4.gz.sha256"), sha)
return version, gz, sha
def _publish_bundle(root: Path, token: str) -> str:
version_file = root / "version.txt"
# Guard the read so a missing version.txt is a clean error, not a raw
# FileNotFoundError.
if not version_file.is_file():
raise SystemExit(f"incomplete artifact bundle: {root}")
version = version_file.read_text(encoding="utf-8").strip()
expected = infra_artifact.infra_artifact_version(infra_vm._infra_init())
if version != expected:
raise SystemExit(
f"artifact bundle version {version!r} does not match checkout {expected!r}"
)
gz = root / "rootfs.ext4.gz"
sha = root / "rootfs.ext4.gz.sha256"
if not gz.is_file() or not sha.is_file():
raise SystemExit(f"incomplete artifact bundle: {root}")
expected_sha = sha.read_text().split()[0].strip().lower()
if _sha256(gz) != expected_sha:
raise SystemExit("artifact bundle checksum mismatch")
gz_url = infra_artifact.artifact_url(version, gz.name)
sha_url = infra_artifact.artifact_url(version, sha.name)
about_url = infra_artifact.artifact_url(version, _ABOUT_NAME)
# Publishing is idempotent. If this exact complete artifact is already
# present, a test-only main commit is a no-op. Otherwise clear any partial
# upload left by an interrupted prior attempt and upload the complete set.
try:
with urllib.request.urlopen(infra_artifact._open(sha_url)) as resp:
remote_sha = resp.read().decode("utf-8").split()[0].strip().lower()
except urllib.error.HTTPError as e:
if e.code != 404:
raise SystemExit(f"checking existing artifact failed (HTTP {e.code})")
remote_sha = ""
except urllib.error.URLError as e:
raise SystemExit(f"registry unreachable: {sha_url} ({e.reason})")
if remote_sha == expected_sha:
print(f"infra rootfs {version} already published")
return version
for url in (gz_url, sha_url, about_url):
_delete(url, token)
_put(gz_url, gz, token)
_put(sha_url, sha.read_bytes(), token)
_put(about_url, _ABOUT_TEXT.encode(), token)
return version
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
prog="publish_infra", description="Build + publish the infra rootfs artifact.")
mode = parser.add_mutually_exclusive_group(required=True)
mode.add_argument("--output", type=Path,
help="build a candidate bundle in DIR without publishing")
mode.add_argument("--publish-dir", type=Path,
help="publish an already-built and tested candidate bundle")
parser.add_argument("--reuse-published", action="store_true",
help="with --output: download from registry if already published instead of building")
parser.add_argument("--dry-run", action="store_true",
help="build the artifact but do not upload")
parser.add_argument("--force", action="store_true",
help="overwrite an already-published artifact of this version")
args = parser.parse_args(argv)
_, _, token = infra_artifact._config()
if args.publish_dir is not None and not token:
if not args.dry_run and not token:
raise SystemExit(
"no publish token: set BOT_BOTTLE_INFRA_ARTIFACT_TOKEN to a token "
"with write:package")
if args.output is not None:
args.output.mkdir(parents=True, exist_ok=True)
reused = None
if args.reuse_published:
reused = _try_download_published(args.output)
if reused is not None:
version, _, _ = reused
(args.output / "version.txt").write_text(version + "\n", encoding="utf-8")
print(f"reused published infra rootfs candidate {version}")
with tempfile.TemporaryDirectory(prefix="bb-publish-infra.") as tmp:
version, gz, sha = build_artifact(Path(tmp))
gz_url = infra_artifact.artifact_url(version, gz.name)
sha_url = infra_artifact.artifact_url(version, sha.name)
about_url = infra_artifact.artifact_url(version, _ABOUT_NAME)
if args.dry_run:
print(f"dry-run: would upload -> {gz_url}")
return 0
version, _gz, _sha = build_artifact(args.output)
(args.output / "version.txt").write_text(version + "\n", encoding="utf-8")
print(f"built infra rootfs candidate {version}")
return 0
assert args.publish_dir is not None
version = _publish_bundle(args.publish_dir, token)
if args.force:
_delete(gz_url, token)
_delete(sha_url, token)
_delete(about_url, token)
_put(gz_url, gz, token) # streamed from disk (hundreds of MB)
_put(sha_url, sha.read_bytes(), token) # tiny, in-memory is fine
_put(about_url, _ABOUT_TEXT.encode(), token) # package description
print(f"published infra rootfs {version}")
return 0
+1 -12
View File
@@ -88,7 +88,7 @@ def require_firecracker() -> None:
booting a VM without it."""
if not is_linux():
die("firecracker backend is only supported on Linux (KVM). "
"On macOS use the macos-container backend.")
"On macOS use --backend=macos-container.")
if shutil.which("firecracker") is None:
info("Firecracker is required but was not found on PATH.")
info("Install: https://github.com/firecracker-microvm/firecracker/releases")
@@ -368,17 +368,6 @@ mount -t devtmpfs dev /dev 2>/dev/null
mkdir -p /dev/pts && mount -t devpts devpts /dev/pts 2>/dev/null
mount -o remount,rw / 2>/dev/null
# /tmp must be world-writable + sticky. The rootless rootfs build can land
# it 0755/root-owned, leaving the agent (uid 1000 node) unable to create
# scratch dirs there — git worktrees, build temp, `git init /tmp/...`, etc.
mkdir -p /tmp && chmod 1777 /tmp
# Rootfs export also maps the image's original owners to the unprivileged
# host build uid. That uid is not guaranteed to be node's uid in the guest;
# restore the home-directory boundary before any SSH provisioning runs.
chown node:node /home/node 2>/dev/null || true
chmod 755 /home/node 2>/dev/null || true
# Install the per-bottle SSH pubkey from the kernel cmdline.
KEY=$(sed -n 's/.*bb_pubkey=\([^ ]*\).*/\1/p' /proc/cmdline | base64 -d 2>/dev/null)
if [ -n "$KEY" ]; then
+4 -11
View File
@@ -12,7 +12,7 @@ from ...env import ResolvedEnv
from ...git_gate import GitGatePlan
from ...supervise import SupervisePlan
from ...manifest import Manifest
from .. import ActiveAgent, BottleBackend, BottleImages, BottleSpec
from .. import ActiveAgent, BottleBackend, BottleSpec
from . import cleanup as _cleanup
from . import enumerate as _enumerate
from . import launch as _launch
@@ -31,7 +31,6 @@ class MacosContainerBottleBackend(
`--backend=macos-container`."""
name = "macos-container"
supports_nested_containers = True
@classmethod
def is_available(cls) -> bool:
@@ -83,17 +82,11 @@ class MacosContainerBottleBackend(
stage_dir=stage_dir,
)
def prelaunch_checks(self, plan: MacosContainerBottlePlan) -> None:
_launch.stale_checks(plan)
def _build_or_load_images(self, plan: MacosContainerBottlePlan) -> BottleImages:
return _launch.build_or_load_images(plan)
@contextmanager
def _launch_impl(
self, plan: MacosContainerBottlePlan, images: BottleImages
def launch(
self, plan: MacosContainerBottlePlan
) -> Generator[MacosContainerBottle, None, None]:
with _launch.launch(plan, images, provision=self.provision) as bottle:
with _launch.launch(plan, provision=self.provision) as bottle:
yield bottle
def ensure_orchestrator(self) -> str:
+3 -8
View File
@@ -68,14 +68,9 @@ class MacosContainerBottle(Bottle):
# reaches the agent (PRD 0070): registration mints it *after* the
# container exists — its source IP is the registration key and Apple
# Container assigns that by DHCP — so it cannot be in the run-time env
# the way docker's compose spec does it.
#
# `container exec --env` does NOT override a run-time value — it
# appends, leaving duplicate entries in the agent's `environ` whose
# resolution is runtime-specific (Node last-wins, Rust first-wins). So
# nothing here may rely on superseding: the proxy vars are supplied
# *only* at exec time and are deliberately absent from the run-time
# env. See `launch._agent_env_entries`.
# the way docker's compose spec does it. `container exec --env` wins
# over the run-time value, so the token-bearing proxy URL set here
# supersedes the token-less one baked in at launch.
self._exec_env = dict(exec_env or {})
self._closed = False
@@ -20,9 +20,6 @@ class MacosContainerBottlePlan(BottlePlan):
# bottle is registered. See launch.py's stamp for why it lives here and not
# only in the exec-time proxy env.
identity_token: str = ""
# Guest-local container engine (issue #392). Gates the derived image, the
# device-mode relaxation, and the resident podman service.
nested_containers: bool = False
@property
def container_name(self) -> str:
@@ -36,11 +36,9 @@ from dataclasses import dataclass
from ...egress import EgressPlan
from ...git_gate import GitGatePlan
from ...log import info
from ...orchestrator.client import OrchestratorClient, OrchestratorClientError
from ..consolidated_util import provision_bottle, teardown_consolidated as _teardown_util
from . import util as container_mod
from .enumerate import CONTAINER_NAME_PREFIX, EnumerationError, enumerate_active
from ...orchestrator.client import OrchestratorClient
from ...orchestrator.registration import registration_inputs
from ..docker.gateway_provision import deprovision_git_gate, provision_git_gate
from .gateway import GATEWAY_NETWORK
from .gateway_provision import AppleGatewayTransport
from .infra import MacosInfraService, OrchestratorStartError
@@ -91,32 +89,6 @@ def ensure_gateway(
)
def live_source_ips(network: str) -> list[str]:
"""Every running agent container's address on `network`.
The reconciliation input: the orchestrator lives inside the infra
container and cannot enumerate the host's containers, so the host has to
tell it which bottles are actually up. Containers that have not been
assigned an address yet contribute nothing the reap's grace window, not
this list, is what protects an in-flight launch.
Raises `EnumerationError` when the live set cannot be determined
authoritatively: either the container listing fails or any individual
inspect fails. Callers must skip reconciliation in that case to avoid
unregistering healthy bottles."""
ips: list[str] = []
for agent in enumerate_active():
name = f"{CONTAINER_NAME_PREFIX}{agent.slug}"
ip = container_mod.inspect_container_network_ip(name, network)
if ip is None:
raise EnumerationError(
f"container inspect {name!r} failed; live set is not authoritative"
)
if ip:
ips.append(ip)
return ips
def register_agent(
egress_plan: EgressPlan,
git_gate_plan: GitGatePlan,
@@ -131,20 +103,17 @@ def register_agent(
container it is the attribution key the gateway resolves policy by.
Raises on failure; the caller tears down."""
client = OrchestratorClient(endpoint.orchestrator_url)
# Self-heal before registering: a launcher that died hard (SIGKILL, closed
# terminal, host sleep) never ran its teardown callback, leaving an active
# row with no container. vmnet recycles addresses, so such a row can
# collide with this bottle's — and `by_source_ip` fail-closes on ambiguity,
# which would resolve no policy at all and deny every host. Best-effort: a
# reconciliation failure must not block an otherwise-fine launch.
try:
client.reconcile(live_source_ips(endpoint.network))
except (OrchestratorClientError, EnumerationError) as e:
info(f"registry reconciliation skipped: {e}")
reg = provision_bottle(
client, source_ip, egress_plan, git_gate_plan, AppleGatewayTransport(),
image_ref=image_ref, tokens=tokens,
inputs = registration_inputs(egress_plan)
reg = client.register_bottle(
source_ip, image_ref=image_ref, policy=inputs.policy,
metadata=inputs.metadata, tokens=tokens,
)
try:
provision_git_gate(AppleGatewayTransport(), reg.bottle_id, git_gate_plan)
except Exception:
# Roll the registration back so a provisioning failure leaves no orphan.
client.teardown_bottle(reg.bottle_id)
raise
return LaunchContext(
bottle_id=reg.bottle_id,
identity_token=reg.identity_token,
@@ -155,21 +124,18 @@ def register_agent(
)
def teardown_consolidated(
bottle_id: str, *, orchestrator_url: str, timeout: float | None = None,
) -> None:
def teardown_consolidated(bottle_id: str, *, orchestrator_url: str) -> None:
"""Deregister the bottle and remove its git-gate state from the gateway.
Both steps are idempotent so this is safe from a cleanup trap. Does NOT
stop the gateway it's a persistent per-host singleton."""
_teardown_util(bottle_id, AppleGatewayTransport(),
orchestrator_url=orchestrator_url, timeout=timeout)
OrchestratorClient(orchestrator_url).teardown_bottle(bottle_id)
deprovision_git_gate(AppleGatewayTransport(), bottle_id)
__all__ = [
"GatewayEndpoint",
"LaunchContext",
"ensure_gateway",
"live_source_ips",
"register_agent",
"teardown_consolidated",
"ConsolidatedLaunchError",
@@ -8,21 +8,13 @@ from ...bottle_state import read_metadata
from .. import ActiveAgent
from .infra import INFRA_NAME
# The name every agent container carries: `bot-bottle-<slug>`. Exported
# because callers that act on a running bottle (gateway-host rewrites,
# registry reconciliation) have to map an enumerated slug back to a
# container name.
CONTAINER_NAME_PREFIX = "bot-bottle-"
_PREFIX = "bot-bottle-"
# The shared per-host infra container carries the same prefix as agent
# containers but is infrastructure, not a bottle — one control plane + gateway
# serves every agent, so listing it as an agent would invent one per host.
_INFRA_NAMES = frozenset({INFRA_NAME})
class EnumerationError(RuntimeError):
"""container list failed; the resulting live set is not authoritative."""
def enumerate_active() -> list[ActiveAgent]:
result = subprocess.run(
["container", "list", "--quiet"],
@@ -31,15 +23,12 @@ def enumerate_active() -> list[ActiveAgent]:
check=False,
)
if result.returncode != 0:
raise EnumerationError(
f"container list failed: "
f"{(result.stderr or '').strip() or '<no stderr>'}"
)
return []
out: list[ActiveAgent] = []
for name in sorted(line.strip() for line in result.stdout.splitlines()):
if not name.startswith(CONTAINER_NAME_PREFIX) or name in _INFRA_NAMES:
if not name.startswith(_PREFIX) or name in _INFRA_NAMES:
continue
slug = name[len(CONTAINER_NAME_PREFIX):]
slug = name[len(_PREFIX):]
metadata = read_metadata(slug)
out.append(ActiveAgent(
backend_name="macos-container",
@@ -1,96 +0,0 @@
"""Stable gateway name for macOS agents, via each bottle's `/etc/hosts`.
The shared gateway's address is assigned by vmnet's DHCP and changes whenever
the infra container is recreated a source-hash bump, an image upgrade, a
crash. Every agent-facing URL (egress proxy, git-http, supervise) embeds that
address, and the proxy URL reaches the agent as **process environment** at
`container exec` time. A running process's `environ` cannot be rewritten from
outside, so a moved gateway used to strand every running bottle permanently:
not degraded, unreachable, until the bottle was relaunched and its session
thrown away.
So the agent never learns the address. It is given a stable *name*
(`GATEWAY_HOSTNAME`) in every URL, resolved through its own `/etc/hosts`.
Unlike `environ`, that is a file it can be rewritten inside a container that
is already running, so a gateway that comes back at a new address is picked up
by live bottles instead of orphaning them.
Apple Container 1.0 offers no container-name DNS on a user network (the only
nameserver an agent sees is vmnet's, which does not know container names) and
`container run` has no `--add-host`, so the entry is written by exec after the
container starts.
Writing it needs root, and the agent runs as `node`: the agent therefore
cannot repoint its own gateway name, while the host (which drives `container
exec --user root`) can. That asymmetry is deliberate keep it.
"""
from __future__ import annotations
from ...log import warn
from . import util as container_mod
from .enumerate import CONTAINER_NAME_PREFIX, enumerate_active
# The name every agent-facing gateway URL uses. Must not collide with a real
# DNS name the agent might resolve; it is bottle-local by construction.
GATEWAY_HOSTNAME = "bot-bottle-gateway"
# Marker so the rewrite is idempotent and only ever touches our own line —
# the rest of /etc/hosts (localhost, the container's own name) is preserved.
_MARKER = "# bot-bottle gateway"
def _rewrite_script(gateway_ip: str) -> str:
"""A shell one-liner that replaces our managed line in `/etc/hosts`.
Rewrites in place via a temp file + `cat` rather than `mv`, so the file
keeps its original inode, ownership, and mode a bind-mounted or
pre-created `/etc/hosts` must not be replaced by a root-owned 0644 copy
that the runtime then refuses to update.
"""
return (
"set -e; "
f"grep -v '{_MARKER}' /etc/hosts > /tmp/.bb-hosts || true; "
f"printf '%s %s %s\\n' '{gateway_ip}' '{GATEWAY_HOSTNAME}' "
f"'{_MARKER}' >> /tmp/.bb-hosts; "
"cat /tmp/.bb-hosts > /etc/hosts; "
"rm -f /tmp/.bb-hosts"
)
def set_gateway_host(container_name: str, gateway_ip: str) -> None:
"""Point `GATEWAY_HOSTNAME` at `gateway_ip` inside one running container.
Must run before the agent is exec'd: the agent's proxy URL names the
gateway, so the entry has to exist for its first connection. Idempotent
re-running with the same address is a no-op in effect.
"""
container_mod.exec_container_as_root(
container_name, ["sh", "-c", _rewrite_script(gateway_ip)],
)
def refresh_gateway_host(gateway_ip: str) -> list[str]:
"""Re-point every running bottle at the current gateway address.
Called once the shared gateway is known to be up, so a bottle stranded by
an earlier gateway restart re-attaches instead of needing a relaunch.
Returns the containers updated.
Best-effort per bottle: one container that refuses the write (already
exiting, say) must not stop the others from being repaired, and must not
fail the launch that triggered the sweep.
"""
updated: list[str] = []
for agent in enumerate_active():
name = f"{CONTAINER_NAME_PREFIX}{agent.slug}"
try:
set_gateway_host(name, gateway_ip)
updated.append(name)
# One bad bottle must not stop the sweep, so this is deliberately broad.
except Exception as e: # noqa: BLE001 # pylint: disable=broad-exception-caught
warn(f"could not re-point {name} at the gateway: {e}")
return updated
__all__ = ["GATEWAY_HOSTNAME", "set_gateway_host", "refresh_gateway_host"]
+15 -23
View File
@@ -41,7 +41,7 @@ from dataclasses import dataclass
from pathlib import Path
from ... import log
from ...orchestrator.gateway import GATEWAY_CA_CERT, MITMPROXY_HOME
from ...orchestrator.gateway import GATEWAY_CA_CERT
from ...orchestrator.lifecycle import (
DEFAULT_PORT,
DEFAULT_STARTUP_TIMEOUT_SECONDS,
@@ -52,9 +52,7 @@ from ...paths import (
CONTROL_PLANE_TOKEN_ENV,
HOST_DB_FILENAME,
host_control_plane_token,
host_gateway_ca_dir,
)
from .. import util as backend_util
from . import util as container_mod
from .gateway import (
DEFAULT_CA_TIMEOUT_SECONDS,
@@ -103,7 +101,7 @@ def _init_script(port: int) -> str:
# Gateway data plane, multi-tenant against the local control plane.
f"( cd /app && BOT_BOTTLE_GATEWAY_DAEMONS={_GATEWAY_DAEMONS} "
f"BOT_BOTTLE_ORCHESTRATOR_URL=http://127.0.0.1:{port} "
f"SUPERVISE_DB_PATH={_DB_PATH_IN_CONTAINER} python3 -m bot_bottle.gateway_init ) &\n"
f"SUPERVISE_DB_PATH={_DB_PATH_IN_CONTAINER} python3 /app/gateway_init.py ) &\n"
"while : ; do wait ; done\n"
)
@@ -219,14 +217,6 @@ class MacosInfraService:
# Container-only DB volume: one kernel writes bot-bottle.db, never
# shared with the host or another guest.
"--volume", f"{self._db_volume}:{_DB_ROOT_IN_CONTAINER}",
# The DB needs a container-only ext4 volume for coherent SQLite
# locking, but the CA has no such constraint. Keep it in the host
# app-data root so infra-container recreation and Apple Container
# volume pruning cannot silently rotate every bottle's trust
# anchor (issue #450).
"--mount",
container_mod.bind_mount_spec(
str(host_gateway_ca_dir()), MITMPROXY_HOME),
# Bind-mount the control-plane source (read-only); a code change
# takes effect on relaunch with no image rebuild.
"--mount",
@@ -270,19 +260,21 @@ class MacosInfraService:
def ca_cert_pem(self, *, timeout: float = DEFAULT_CA_TIMEOUT_SECONDS) -> str:
"""The gateway's mitmproxy CA (PEM) agents install to trust its TLS
interception. Read through the container path backed by the persistent
host CA directory; polls because mitmproxy writes it a beat after
start."""
def _fetch() -> str | None:
interception. Read out of the container (the CA lives on a
container-internal path, not a host mount); polls because mitmproxy
writes it a beat after start."""
deadline = time.monotonic() + timeout
while True:
result = container_mod.run_container_argv(
["container", "exec", self._name, "cat", GATEWAY_CA_CERT])
return result.stdout if result.returncode == 0 and result.stdout.strip() else None
try:
return backend_util.poll_ca_cert(_fetch, timeout=timeout)
except TimeoutError as exc:
raise GatewayError(
f"gateway CA not available in {self._name} after {timeout:g}s"
) from exc
if result.returncode == 0 and result.stdout.strip():
return result.stdout
if time.monotonic() >= deadline:
raise GatewayError(
f"gateway CA not available in {self._name} after {timeout:g}s: "
f"{(result.stderr or '').strip() or 'empty'}"
)
time.sleep(_CA_POLL_SECONDS)
def stop(self) -> None:
"""Remove the infra container (idempotent). The DB volume persists."""
+41 -147
View File
@@ -53,22 +53,13 @@ from ...git_gate import (
revoke_git_gate_provisioned_keys,
)
from ...git_http_backend import DEFAULT_PORT as _GIT_HTTP_PORT
from ...image_cache import check_stale
from ...log import die, info, warn
from .. import BottleImages
from ...supervise import SUPERVISE_PORT
from ..docker.egress import EGRESS_PORT
from ..util import AGENT_CA_BUNDLE, AGENT_CA_PATH
from . import util as container_mod
from .bottle import MacosContainerBottle
from .gateway_hosts import (
GATEWAY_HOSTNAME,
refresh_gateway_host,
set_gateway_host,
)
from . import nested_containers as nested_containers_mod
from .bottle_plan import MacosContainerBottlePlan
from ...orchestrator.config_store import resolve_teardown_timeout
from .consolidated_launch import (
GatewayEndpoint,
ensure_gateway,
@@ -80,69 +71,18 @@ _REPO_DIR = str(Path(__file__).resolve().parent.parent.parent.parent)
_AGENT_SLEEP_SECONDS = "2147483647"
def build_or_load_images(plan: MacosContainerBottlePlan) -> BottleImages:
"""Resolve the agent image ref for this plan. The gateway's own image is
built by `ensure_gateway` it belongs to the shared singleton."""
return BottleImages(agent=_layer_nested_containers(plan, _agent_image(plan)))
def _agent_image(plan: MacosContainerBottlePlan) -> str:
committed = read_committed_image(plan.slug)
if committed and container_mod.image_exists(committed):
info(f"using committed image {committed!r}")
return committed
if plan.spec.image_policy == "cached":
if not container_mod.image_exists(plan.image):
die(
f"cached agent image {plan.image!r} not found; "
"run without --cached-images to build it"
)
info(f"using cached agent image {plan.image!r}")
return plan.image
container_mod.build_image(plan.image, _REPO_DIR, dockerfile=plan.dockerfile_path)
return plan.image
def _layer_nested_containers(
plan: MacosContainerBottlePlan, agent_image: str,
) -> str:
"""Add the guest-local container tooling on top of the agent image.
A separate derived tag, not the provider Dockerfile, so bottles that never
ask for nested containers carry none of its weight.
"""
if not plan.nested_containers:
return agent_image
derived = f"{agent_image}{nested_containers_mod.IMAGE_SUFFIX}"
if plan.spec.image_policy == "cached":
if not container_mod.image_exists(derived):
die(
f"cached nested-container image {derived!r} not found; "
"run without --cached-images to build it"
)
info(f"using cached nested-container image {derived!r}")
return derived
return nested_containers_mod.build_image(agent_image, container_mod.build_image)
@contextmanager
def launch(
plan: MacosContainerBottlePlan,
images: BottleImages,
*,
provision: Callable[[MacosContainerBottlePlan, "MacosContainerBottle"], str | None],
) -> Generator[MacosContainerBottle, None, None]:
"""Run, register, provision, and yield an Apple Container bottle on the
shared per-host gateway."""
"""Build, run, register, provision, and yield an Apple Container bottle on
the shared per-host gateway."""
stack = ExitStack()
bottle_for_revoke = plan.manifest.bottle
git_gate_dir_for_revoke = git_gate_state_dir(plan.slug)
plan = dataclasses.replace(
plan,
agent_provision=dataclasses.replace(plan.agent_provision, image=str(images.agent)),
)
def teardown() -> None:
teardown_exc: BaseException | None = None
try:
@@ -155,14 +95,11 @@ def launch(
raise teardown_exc
try:
plan = _build_images(plan)
# Step 1: the per-host singletons. Must precede the agent run — its
# proxy env needs the gateway's address at `container run` time.
endpoint = ensure_gateway()
# The gateway's address may have changed since these bottles launched
# (any infra recreate re-runs DHCP). They name the gateway rather than
# address it, so re-pointing /etc/hosts re-attaches them in place
# instead of leaving them stranded until relaunch.
refresh_gateway_host(endpoint.gateway_ip)
# Step 2: mint this bottle's deploy keys, then point it at the SHARED
# gateway's CA + git-http/supervise ports.
@@ -180,9 +117,6 @@ def launch(
# attribution key; `--cap-drop CAP_NET_RAW` at run is what makes it
# unforgeable. Poll: `container run --detach` can return before vmnet's
# DHCP has assigned the address.
# Resolve the gateway name before anything execs: every agent-facing
# URL uses it, so the entry must exist for the first connection.
set_gateway_host(plan.container_name, endpoint.gateway_ip)
source_ip = container_mod.wait_container_ipv4_on_network(
plan.container_name, endpoint.network,
)
@@ -195,7 +129,6 @@ def launch(
token_values = egress_resolve_token_values(
plan.egress_plan.token_env_map, effective_env,
)
teardown_timeout = resolve_teardown_timeout()
ctx = register_agent(
plan.egress_plan,
plan.git_gate_plan,
@@ -207,7 +140,6 @@ def launch(
stack.callback(
teardown_consolidated, ctx.bottle_id,
orchestrator_url=ctx.orchestrator_url,
timeout=teardown_timeout,
)
info(
f"agent {plan.container_name} registered "
@@ -223,10 +155,6 @@ def launch(
# token above, so — unlike the run-time env — the plan CAN carry it.
plan = dataclasses.replace(plan, identity_token=ctx.identity_token)
exec_env = {
**_identity_proxy_env(endpoint, ctx.identity_token),
**nested_containers_mod.guest_env(plan.nested_containers),
}
bottle = MacosContainerBottle(
plan.container_name,
teardown,
@@ -240,38 +168,31 @@ def launch(
),
terminal_color=plan.spec.color,
agent_workdir=plan.workspace_plan.workdir,
exec_env=exec_env,
exec_env=_identity_proxy_env(endpoint, ctx.identity_token),
)
bottle.prompt_path = provision(plan, bottle)
if plan.nested_containers:
nested_containers_mod.prepare_guest_devices(
plan.container_name, container_mod.exec_container_as_root,
)
nested_containers_mod.start(bottle)
yield bottle
finally:
teardown()
def stale_checks(plan: MacosContainerBottlePlan) -> None:
"""Raise StaleImageError if a cached image is older than the configured
threshold. Only runs when image_policy is 'cached'. Called by the backend
class's _image_stale_checks before _launch_impl starts any resources."""
if plan.spec.image_policy != "cached":
return
def _build_images(plan: MacosContainerBottlePlan) -> MacosContainerBottlePlan:
"""Build the agent image. The gateway's own image is built by
`ensure_gateway` it belongs to the shared singleton, not to a bottle."""
committed = read_committed_image(plan.slug)
if committed and container_mod.image_exists(committed):
ts = container_mod.image_created_at(committed)
if ts is not None:
check_stale(f"agent image {committed!r}", ts)
return
if container_mod.image_exists(plan.image):
ts = container_mod.image_created_at(plan.image)
if ts is not None:
check_stale(f"agent image {plan.image!r}", ts)
info(f"using committed image {committed!r}")
return dataclasses.replace(
plan,
agent_provision=dataclasses.replace(
plan.agent_provision, image=committed,
),
)
container_mod.build_image(
plan.image, _REPO_DIR, dockerfile=plan.dockerfile_path,
)
return plan
def _provision_git_gate_keys(
@@ -310,20 +231,13 @@ def _stamp_agent_urls(
) -> MacosContainerBottlePlan:
"""Point the agent's git-gate insteadOf rewrites + supervise MCP at the
shared gateway's ports. Both bypass the egress proxy (NO_PROXY covers the
gateway name).
Addressed by `GATEWAY_HOSTNAME`, never by IP: these URLs are baked into
the agent's gitconfig and MCP config at provision time, so an address here
would strand the bottle the moment the gateway moved. The name is resolved
per connection through `/etc/hosts`, which stays rewritable while the
bottle runs."""
del endpoint # addressed by name; the address reaches the bottle via /etc/hosts
gateway address)."""
git_gate_url = (
f"http://{GATEWAY_HOSTNAME}:{_GIT_HTTP_PORT}"
f"http://{endpoint.gateway_ip}:{_GIT_HTTP_PORT}"
if plan.git_gate_plan.upstreams else ""
)
supervise_url = (
f"http://{GATEWAY_HOSTNAME}:{SUPERVISE_PORT}/"
f"http://{endpoint.gateway_ip}:{SUPERVISE_PORT}/"
if plan.supervise_plan is not None else ""
)
return dataclasses.replace(
@@ -333,43 +247,30 @@ def _stamp_agent_urls(
)
def _proxy_url(identity_token: str = "") -> str:
def _proxy_url(gateway_ip: str, identity_token: str = "") -> str:
"""The agent's egress proxy URL. The identity token rides as proxy
credentials the gateway reads Proxy-Authorization, resolves the
(source_ip, token) pair against the control plane, and strips it before
upstream. Without a valid pair `/resolve` denies the request (#366).
Names the gateway rather than addressing it: this URL reaches the agent as
process environment, which cannot be rewritten once the agent is running,
so an address baked here is unfixable if the gateway moves."""
upstream. Without a valid pair `/resolve` denies the request (#366)."""
cred = f"bottle:{identity_token}@" if identity_token else ""
return f"http://{cred}{GATEWAY_HOSTNAME}:{EGRESS_PORT}"
return f"http://{cred}{gateway_ip}:{EGRESS_PORT}"
def _no_proxy() -> str:
def _no_proxy(gateway_ip: str) -> str:
# git-http + supervise live on the gateway and must NOT go through the
# egress proxy — the agent reaches them directly by name. Deliberately
# address-free: NO_PROXY is baked into the run-time env and is therefore
# just as unfixable as the proxy URL if the gateway moves.
return f"localhost,127.0.0.1,{GATEWAY_HOSTNAME}"
# egress proxy — the agent reaches them directly by its address.
return f"localhost,127.0.0.1,{gateway_ip}"
def _identity_proxy_env(
endpoint: GatewayEndpoint, identity_token: str,
) -> dict[str, str]:
"""The token-bearing proxy env applied at `container exec` — the only way
to get the token in, since it does not exist until after the container
runs (registration keys on the DHCP-assigned address).
This is the *sole* source of `*_PROXY` for the agent. It deliberately does
not rely on overriding a run-time value: `container exec --env` appends
rather than replaces, so a run-time `HTTPS_PROXY` would survive alongside
this one and first-wins runtimes would read the wrong entry. See
`_agent_env_entries`."""
"""The token-bearing proxy env applied at `container exec`. It supersedes
the token-less run-time value (exec `--env` wins), which is the only way to
get the token in: it does not exist until after the container runs."""
if not identity_token:
return {}
del endpoint # the gateway is named, not addressed
url = _proxy_url(identity_token)
url = _proxy_url(endpoint.gateway_ip, identity_token)
return {
"HTTPS_PROXY": url, "HTTP_PROXY": url,
"https_proxy": url, "http_proxy": url,
@@ -416,23 +317,16 @@ def _agent_run_argv(
def _agent_env_entries(
plan: MacosContainerBottlePlan, endpoint: GatewayEndpoint,
) -> tuple[str, ...]:
# No `*_PROXY` here on purpose. The token-bearing URL is applied at
# `container exec` (`_identity_proxy_env`), and Apple's `container exec
# --env` **appends** to the run-time environment rather than replacing it:
# setting a token-less value here leaves two `HTTPS_PROXY` entries in the
# agent's `environ`, token-less first. Which one a runtime reads is then
# pure luck — Node takes the last (and worked), Rust's `std::env::var`
# takes the first, so Codex proxied without its identity token and
# `/resolve` fail-closed on every request.
#
# A token-less proxy URL has no legitimate consumer anyway: the init
# process is `sleep` and everything that egresses arrives via exec. Its
# only value was a tidy 403 for unattributed callers, which is not worth
# silently dropping attribution for. Without it a process that egresses
# before the exec-time env still fails closed — the agent network is
# host-only, so there is no route off it except the gateway.
no_proxy = _no_proxy()
# Token-less at run time — the token does not exist yet (see
# `_identity_proxy_env`). Anything egressing before the exec-time override
# is denied by `/resolve`, which is the safe direction.
proxy_url = _proxy_url(endpoint.gateway_ip)
no_proxy = _no_proxy(endpoint.gateway_ip)
env = [
f"HTTPS_PROXY={proxy_url}",
f"HTTP_PROXY={proxy_url}",
f"https_proxy={proxy_url}",
f"http_proxy={proxy_url}",
f"NO_PROXY={no_proxy}",
f"no_proxy={no_proxy}",
f"NODE_EXTRA_CA_CERTS={AGENT_CA_PATH}",
@@ -1,210 +0,0 @@
#!/bin/sh
set -eu
uid="$(id -u)"
if [ "$uid" -eq 0 ]; then
echo "refusing to run the guest container engine as root" >&2
exit 1
fi
# Every piece of podman 5's networking stack is checked here, because each
# one fails at a different and misleading layer if it is absent: no pasta and
# nothing starts at all; no nft and netavark cannot build the bridge every
# compose file expects; no aardvark-dns and DNS inside nested containers fails
# while everything else looks healthy.
for command in podman docker fuse-overlayfs pasta nft slirp4netns; do
command -v "$command" >/dev/null 2>&1 || {
echo "missing nested-container prerequisite: $command" >&2
exit 1
}
done
# The inverse of the rootless-Docker check, and the whole point of the podman
# variant: a subordinate range would push podman onto newuidmap, which cannot
# write a multi-range uid_map without CAP_SYS_ADMIN in this guest. An empty
# range keeps it on the single-UID self-mapping an unprivileged process may
# write itself.
if grep -q "^$(id -un):" /etc/subuid 2>/dev/null; then
echo "unexpected subordinate UID range for $(id -un): podman would" >&2
echo "require CAP_SYS_ADMIN via newuidmap in this guest" >&2
exit 1
fi
for device in /dev/fuse /dev/net/tun; do
[ -r "$device" ] && [ -w "$device" ] || {
echo "device $device is not readable/writable by $(id -un)" >&2
exit 1
}
done
# Short by necessity, not by accident: conmon's attach socket lives under
# this directory and must fit in a 108-byte sun_path. See nested_containers.py.
# Must stay in step with AGENT_CA_BUNDLE in bot_bottle/backend/util.py; a unit
# test pins the two together.
CA_BUNDLE="/etc/ssl/certs/ca-certificates.crt"
[ -r "$CA_BUNDLE" ] || {
echo "gateway CA bundle $CA_BUNDLE is missing or unreadable" >&2
exit 1
}
# The proxy URL the agent inherits names `bot-bottle-gateway`, which resolves
# only through this bottle's /etc/hosts. A nested container gets its own hosts
# file, so it cannot resolve the name and dies at "Could not resolve proxy".
#
# podman's containers.conf `hosts_file` would fix that, except the
# Docker-compatible API ignores it — it only takes effect for native
# `podman run`, and the agent types `docker`. So the name is resolved *here*
# and the address, not the name, goes into the proxy URL the nested container
# receives. Verified on macOS 26 / podman 5.4.2: with the address in place,
# https://quay.io returns 200 and a non-allowlisted host still gets 403, so
# the egress boundary applies inside nested containers too.
GATEWAY_NAME="bot-bottle-gateway"
gateway_ip="$(
awk -v name="$GATEWAY_NAME" '$2 == name { print $1; exit }' /etc/hosts
)"
[ -n "$gateway_ip" ] || {
echo "no /etc/hosts entry for $GATEWAY_NAME; the gateway address is" >&2
echo "needed so nested containers can reach the egress proxy" >&2
exit 1
}
export XDG_RUNTIME_DIR="${XDG_RUNTIME_DIR:-/tmp/bbp}"
config="$HOME/.config/containers"
mkdir -p "$XDG_RUNTIME_DIR" "$config"
chmod 700 "$XDG_RUNTIME_DIR"
# ignore_chown_errors is required, not incidental: with a single-UID mapping
# there is no second UID for image layers to be chowned to, so layers that
# record other owners would otherwise fail to extract.
cat > "$config/storage.conf" <<'CONF'
[storage]
driver="overlay"
[storage.options.overlay]
mount_program="/usr/bin/fuse-overlayfs"
ignore_chown_errors="true"
CONF
# No cgroup delegation reaches this guest, so asking podman to manage cgroups
# fails; events_logger=file avoids the journald socket that is equally absent.
#
# The rest of this config is what lets a nested container reach the network:
#
# hosts_file only takes effect for native `podman run` — the
# Docker-compatible API ignores it, and the agent types
# `docker`. Kept anyway because it costs nothing and makes
# podman-native use behave; the compat path is covered by the
# address-bearing proxy URL below.
# volumes/env the gateway TLS-intercepts, so a container that does not
# trust the bottle's CA bundle gets "unable to get local issuer
# certificate". Mounting the bundle read-only and pointing the
# usual env vars at it covers curl, wget, python, and node
# without distro-specific trust commands.
#
# The proxy URL carries the bottle's identity token. podman already forwards
# that same URL into every nested container from the agent's own environment,
# so writing it to a 0600 file inside this disposable VM hands it to nobody
# new. It is never echoed.
CA_BUNDLE="$CA_BUNDLE" GATEWAY_NAME="$GATEWAY_NAME" GATEWAY_IP="$gateway_ip" \
CONTAINERS_CONF="$config/containers.conf" python3 - <<'PY'
import os
from pathlib import Path
ca = os.environ["CA_BUNDLE"]
name = os.environ["GATEWAY_NAME"]
ip = os.environ["GATEWAY_IP"]
entries = [
f"SSL_CERT_FILE={ca}",
f"CURL_CA_BUNDLE={ca}",
f"REQUESTS_CA_BUNDLE={ca}",
f"NODE_EXTRA_CA_CERTS={ca}",
]
# The gateway name resolves only through the bottle's /etc/hosts, which a
# nested container does not inherit, so hand it the address instead.
for var in ("HTTP_PROXY", "HTTPS_PROXY", "http_proxy", "https_proxy"):
value = os.environ.get(var)
if value:
entries.append(f"{var}={value.replace(name, ip)}")
# NO_PROXY keeps the name: it is matched against what a client asks for, and
# code inside a nested container still says "bot-bottle-gateway".
for var in ("NO_PROXY", "no_proxy"):
value = os.environ.get(var)
if value:
entries.append(f"{var}={value}")
path = Path(os.environ["CONTAINERS_CONF"])
path.write_text("\n".join([
"[containers]",
'cgroups="disabled"',
# podman copies the host's proxy vars into every container by default,
# and that copy *wins* over the env below — putting the unresolvable
# gateway name back. Turn it off so the address-bearing URLs stand.
"http_proxy=false",
'hosts_file="/etc/hosts"',
f'volumes=["{ca}:{ca}:ro"]',
"env=[",
*[f' "{entry}",' for entry in entries],
"]",
"[engine]",
'cgroup_manager="cgroupfs"',
'events_logger="file"',
"",
]), encoding="utf-8")
path.chmod(0o600)
PY
# Registry pulls egress through the bottle's proxy like everything else. The
# token-bearing proxy URL is already in the agent's environment; persisting it
# inside this disposable VM does not broaden its authority.
#
# This file is also what the Docker CLI copies into every container it starts,
# and being client-side it beats anything the podman service does — it is why
# containers.conf `env`, `http_proxy=false`, and the service's own environment
# all failed to change what a nested container saw. The address goes in here
# for the same reason it goes everywhere else: `bot-bottle-gateway` resolves
# in the bottle, never inside a nested container.
GATEWAY_NAME="$GATEWAY_NAME" GATEWAY_IP="$gateway_ip" python3 - <<'PY'
import json
import os
from pathlib import Path
name = os.environ["GATEWAY_NAME"]
ip = os.environ["GATEWAY_IP"]
proxy = os.environ.get("HTTPS_PROXY") or os.environ.get("https_proxy", "")
# NO_PROXY keeps the name: it is matched against what a client asks for, and
# code inside a nested container still says "bot-bottle-gateway".
no_proxy = os.environ.get("NO_PROXY") or os.environ.get("no_proxy", "")
config = {"proxies": {"default": {
"httpProxy": proxy.replace(name, ip),
"httpsProxy": proxy.replace(name, ip),
"noProxy": no_proxy,
}}}
path = Path.home() / ".docker" / "config.json"
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(json.dumps(config), encoding="utf-8")
path.chmod(0o600)
PY
if docker info >/dev/null 2>&1; then
exit 0
fi
# Belt to the ~/.docker/config.json braces above, which is what actually
# decides this for `docker run`. The service environment is what podman falls
# back to for anything the CLI does not stamp — its own registry pulls, and
# containers created through the API by something other than the Docker CLI.
# Cheap, and it keeps the address consistent across both paths.
#
# Assigned via parameter expansion, never echoed: these carry the bottle's
# identity token.
for var in HTTP_PROXY HTTPS_PROXY http_proxy https_proxy; do
eval "value=\${$var:-}"
[ -n "$value" ] || continue
eval "export $var=\"\${value%%$GATEWAY_NAME*}$gateway_ip\${value#*$GATEWAY_NAME}\""
done
log=/tmp/bot-bottle-nested-containers.log
nohup podman system service --time=0 \
"unix://$XDG_RUNTIME_DIR/podman.sock" \
>"$log" 2>&1 </dev/null &
@@ -1,162 +0,0 @@
"""Guest-local container engine for Apple-container bottles (issue #392).
The service and every nested container remain inside the existing per-bottle
VM. This module refuses to compensate for missing prerequisites with outer
capabilities, a privileged container, or a host Docker socket.
Podman is used rather than rootless Docker for one specific reason: Apple
Container's capability bounding set omits `CAP_SYS_ADMIN`, which the kernel
requires to write a multi-range `uid_map` via `newuidmap`. Rootless Docker
has no path that avoids that write. Podman does with no subordinate UID
range configured it falls back to a single-UID self-mapping, which an
unprivileged process may write itself. See
`docs/research/rootless-docker-in-apple-container-spike.md`.
That fallback is why `build_image` *removes* the agent user's `/etc/subuid`
and `/etc/subgid` entries instead of adding them: their presence is precisely
what would send podman down the `newuidmap` path that cannot work here.
The agent still talks to `docker` and `docker compose`; those speak to
podman's Docker-compatible API socket, so nothing in the agent's habits
changes.
Nested containers run *within* the bottle boundary, not inside a new one: the
single-UID mapping means `root` in a nested container is the agent user
outside it. This is for build and test workloads, not for sandboxing
untrusted code.
"""
from __future__ import annotations
import shlex
import shutil
import tempfile
import time
from pathlib import Path
from typing import Callable
from ...log import die, info
_INIT = "/usr/local/libexec/bot-bottle/nested-containers-init"
# Deliberately cryptic and short. podman derives conmon's attach socket as
# `$XDG_RUNTIME_DIR/libpod/tmp/socket/<64-hex-id>/attach`, and a Unix socket
# path may not exceed 108 bytes (`sun_path`). The descriptive
# `/tmp/bot-bottle-podman-run` produced a 116-byte path — over the limit, so
# attach would have broken as soon as anything got far enough to attach. Do
# not lengthen this for readability; it buys 8 bytes of headroom.
_RUNTIME_DIR = "/tmp/bbp"
_SOCKET = f"{_RUNTIME_DIR}/podman.sock"
_LOG = "/tmp/bot-bottle-nested-containers.log"
IMAGE_SUFFIX = "-nested-containers"
READY_RETRIES = 30
# Apple Container creates both device nodes 0600 root:root, so the agent user
# cannot open them: /dev/fuse blocks the fuse-overlayfs storage driver and
# /dev/net/tun blocks slirp4netns, which rootless podman uses for the default
# bridge network that stock compose files expect. Relaxing the modes needs no
# capability the bottle does not already hold — unlike CAP_SYS_ADMIN, which is
# what killed the rootless-Docker approach.
_GUEST_DEVICES = ("/dev/fuse", "/dev/net/tun")
def build_image(
base_image: str,
build: Callable[..., None],
) -> str:
"""Layer the nested-container tooling onto an already-built agent image.
Only what the flag is meant to gate lands here. Podman itself is already
in every built-in agent image (issue #451); the storage/network helpers,
the Docker CLI, and the compose plugin are the ~100MB this flag buys.
"""
image = f"{base_image}{IMAGE_SUFFIX}"
init_script = Path(__file__).with_name("nested-containers-init.sh")
with tempfile.TemporaryDirectory(prefix="bot-bottle-nested-containers.") as tmp:
context = Path(tmp)
shutil.copy2(init_script, context / "nested-containers-init.sh")
(context / "Dockerfile").write_text(
"FROM docker:28-cli AS docker_cli\n"
f"FROM {base_image}\n"
"USER root\n"
"COPY --from=docker_cli /usr/local/bin/docker /usr/local/bin/docker\n"
"COPY --from=docker_cli /usr/local/libexec/docker/cli-plugins/"
"docker-compose /usr/local/libexec/docker/cli-plugins/docker-compose\n"
"RUN apt-get update \\\n"
# podman 5's networking stack, installed explicitly because the
# base image's `--no-install-recommends` podman does not pull it
# in and each piece fails at a different, misleading layer:
# passt -> `pasta`, the default rootless netns helper
# (podman 4 used slirp4netns); without it
# nothing starts: "could not find pasta"
# nftables -> `nft`, which netavark shells out to for the
# bridge network every compose file expects
# aardvark-dns -> name resolution *inside* nested containers;
# without it DNS fails while everything else
# looks healthy
# slirp4netns stays as the documented fallback for pasta.
" && apt-get install -y --no-install-recommends "
"aardvark-dns fuse-overlayfs netavark nftables passt "
"slirp4netns uidmap \\\n"
" && rm -rf /var/lib/apt/lists/* \\\n"
# Deliberate: an empty subordinate range keeps podman on the
# single-UID mapping that needs no CAP_SYS_ADMIN. Adding ranges
# here would reintroduce the newuidmap failure this design exists
# to route around.
" && sed -i '/^node:/d' /etc/subuid /etc/subgid\n"
"COPY nested-containers-init.sh "
"/usr/local/libexec/bot-bottle/nested-containers-init\n"
"RUN chmod 0755 /usr/local/libexec/bot-bottle/nested-containers-init\n"
"USER node\n",
encoding="utf-8",
)
build(image, str(context), dockerfile=str(context / "Dockerfile"))
return image
def guest_env(enabled: bool) -> dict[str, str]:
"""Environment consumed by the Docker CLI inside an enabled bottle."""
if not enabled:
return {}
return {
"DOCKER_HOST": f"unix://{_SOCKET}",
"XDG_RUNTIME_DIR": _RUNTIME_DIR,
}
def prepare_guest_devices(container_name: str, exec_as_root: Callable[..., None]) -> None:
"""Make /dev/fuse and /dev/net/tun openable by the agent user.
Runs as root inside the bottle because the agent must not be able to
re-mode device nodes itself. No outer capability is involved.
"""
exec_as_root(
container_name,
["sh", "-c", f"chmod 0666 {' '.join(_GUEST_DEVICES)}"],
)
def start(bottle: object) -> None:
"""Start and verify the unprivileged service through the bottle exec API."""
info("starting guest-local container engine")
result = bottle.exec(shlex.quote(_INIT)) # type: ignore[attr-defined]
if result.returncode != 0:
detail = (result.stderr or result.stdout or "").strip()
die(f"nested-container bootstrap failed: {detail or '<no output>'}")
for _ in range(READY_RETRIES):
result = bottle.exec("docker info >/dev/null 2>&1") # type: ignore[attr-defined]
if result.returncode == 0:
info("guest-local container engine is ready")
return
time.sleep(0.2)
logs = bottle.exec( # type: ignore[attr-defined]
f"tail -n 80 {_LOG} 2>/dev/null || true"
)
die(
"guest-local container engine did not become ready without additional "
f"outer privileges:\n{(logs.stdout or logs.stderr or '<no log>').strip()}"
)
__all__ = ["build_image", "guest_env", "prepare_guest_devices", "start"]
@@ -44,5 +44,4 @@ def resolve_plan(
egress_plan=egress_plan,
supervise_plan=supervise_plan,
agent_provision=agent_provision_plan,
nested_containers=manifest.bottle.nested_containers,
)
@@ -10,7 +10,6 @@ import shutil
import subprocess
import tempfile
import time
from datetime import datetime, timezone
from typing import Iterable
from ...log import die, info
@@ -361,21 +360,6 @@ def exec_container(name: str, argv: list[str]) -> None:
)
def exec_container_as_root(name: str, argv: list[str]) -> None:
"""`exec_container`, but as uid 0 inside the container.
For host-driven maintenance the agent itself must not be able to perform
rewriting `/etc/hosts` to point the gateway name at an address. The agent
runs as `node`, so it cannot repoint its own gateway; the host can.
"""
result = _run_container_op([_CONTAINER, "exec", "--user", "root", name, *argv])
if result.returncode != 0:
die(
f"container exec (root) in {name} failed: "
f"{(result.stderr or '').strip() or '<no stderr>'}"
)
def _run_container_op(cmd: list[str]) -> subprocess.CompletedProcess[str]:
result = subprocess.run(
cmd,
@@ -573,41 +557,6 @@ def try_container_ipv4_on_network(name: str, network: str) -> str:
return ""
def inspect_container_network_ip(name: str, network: str) -> str | None:
"""IP of `name` on `network`, distinguishing inspect failure from "not yet".
Returns:
- the IP string when the container has one on `network`
- "" when inspect succeeds but no address is assigned yet (in-flight DHCP)
- None when the inspect command itself fails (authoritative list impossible)
"""
result = subprocess.run(
[_CONTAINER, "inspect", name],
capture_output=True, text=True, check=False,
)
if result.returncode != 0:
return None
try:
data = json.loads(result.stdout or "[]")
except json.JSONDecodeError:
return None
if isinstance(data, list):
data = data[0] if data else {}
if not isinstance(data, dict):
return None
status = data.get("status")
networks = status.get("networks") if isinstance(status, dict) else None
if not isinstance(networks, list):
return ""
for entry in networks:
if not isinstance(entry, dict) or entry.get("network") != network:
continue
raw = entry.get("ipv4Address")
if isinstance(raw, str) and raw:
return raw.split("/", 1)[0]
return ""
def wait_container_ipv4_on_network(
name: str, network: str, *, timeout: float = 15.0, poll: float = 0.25,
) -> str:
@@ -662,39 +611,6 @@ def image_id(ref: str) -> str:
raise AssertionError("unreachable")
def image_created_at(ref: str) -> datetime | None:
"""Return the image creation timestamp as an aware UTC datetime, or None
when the field is absent or unparseable (e.g. FROM-scratch images, images
pulled from registries that omit the field). Callers should skip the stale
check when None is returned rather than treating it as an error."""
result = subprocess.run(
[_CONTAINER, "image", "inspect", ref],
capture_output=True,
text=True,
check=False,
)
if result.returncode != 0:
die(
f"container image inspect for {ref!r} failed: "
f"{(result.stderr or '').strip() or '<no stderr>'}"
)
try:
data = json.loads(result.stdout or "{}")
except json.JSONDecodeError as exc:
die(f"container image inspect for {ref!r} returned malformed JSON: {exc}")
if isinstance(data, list) and data:
data = data[0]
if isinstance(data, dict):
value = data.get("created") or data.get("Created")
if isinstance(value, str) and value:
try:
ts = value.rstrip("Z")
return datetime.fromisoformat(ts).replace(tzinfo=timezone.utc)
except ValueError:
pass
return None
def save(ref: str, output: str) -> None:
subprocess.run([_CONTAINER, "image", "save", ref, "-o", output], check=True)
-18
View File
@@ -26,7 +26,6 @@ from ..bottle_state import (
)
from ..egress import Egress, EgressPlan
from ..git_gate import GitGate, GitGatePlan
from ..log import die
from ..manifest import Manifest, ManifestBottle
from ..supervise import Supervise, SupervisePlan
from . import BottleSpec
@@ -113,22 +112,6 @@ def merge_provision_env_vars(provision: AgentProvisionPlan) -> AgentProvisionPla
return replace(provision, guest_env=merged)
def reject_nested_containers(backend: str, manifest: Manifest) -> None:
"""Fail loudly when a backend cannot honor `nested_containers: true`.
Silently ignoring it would hand the agent a bottle where `docker` is not
there and the only sound alternatives on these backends (a host daemon
socket, a privileged container) are exactly what issue #392 rules out.
"""
if not manifest.bottle.nested_containers:
return
die(
f"nested_containers is not supported on the {backend} backend. "
"Only macos-container runs a guest-local container engine today; "
"mounting the host Docker socket is not an option bot-bottle offers."
)
def resolve_manifest_dockerfile(path_value: str, spec: BottleSpec) -> str:
"""Resolve a manifest-supplied dockerfile path relative to user_cwd."""
path = Path(os.path.expanduser(path_value))
@@ -139,7 +122,6 @@ def resolve_manifest_dockerfile(path_value: str, spec: BottleSpec) -> str:
__all__ = [
"merge_provision_env_vars",
"reject_nested_containers",
"mint_slug",
"prepare_agent_state_dir",
"prepare_egress",
-20
View File
@@ -7,8 +7,6 @@ from __future__ import annotations
import hashlib
import os
import ssl
import time
from collections.abc import Callable
from pathlib import Path
from typing import TYPE_CHECKING
@@ -17,24 +15,6 @@ from ..log import die, info
if TYPE_CHECKING:
from ..egress import EgressPlan
_CA_POLL_INTERVAL = 0.5
def poll_ca_cert(fetch: Callable[[], str | None], *, timeout: float) -> str:
"""Poll `fetch` until it returns a non-empty PEM string or `timeout` expires.
`fetch` should return the PEM on success and `None` (or empty string) when
the cert is not yet available. Raises `TimeoutError` if the cert never
appears within `timeout` seconds."""
deadline = time.monotonic() + timeout
while True:
result = fetch()
if result:
return result
if time.monotonic() >= deadline:
raise TimeoutError(f"CA cert not available after {timeout:g}s")
time.sleep(_CA_POLL_INTERVAL)
# Debian-family CA layout, shared by every backend (all guest images
# are Debian-family). AGENT_CA_PATH is the source path that
+5 -11
View File
@@ -19,7 +19,6 @@ from .commit import cmd_commit
from .edit import cmd_edit
from .info import cmd_info
from .init import cmd_init
from .login import cmd_login
from .resume import cmd_resume
from .start import cmd_start
from .supervise import cmd_supervise
@@ -34,19 +33,11 @@ COMMANDS = {
"info": cmd_info,
"init": cmd_init,
"list": cmd_list,
"login": cmd_login,
"resume": cmd_resume,
"start": cmd_start,
"supervise": cmd_supervise,
}
# Commands that manage host prerequisites (or are otherwise store-free) and
# must run before — or without — a migrated DB. `backend` provisions/probes
# the host (TAP pool, /dev/kvm, firecracker) and never opens the store, so
# gating it on the schema breaks preflight on a fresh CI runner where stdin
# isn't a TTY and the migration prompt can't be answered.
NO_MIGRATION_COMMANDS = frozenset({"backend", "login"})
def usage() -> None:
sys.stderr.write(f"usage: {PROG} <command> [args...]\n\n")
@@ -58,7 +49,6 @@ def usage() -> None:
sys.stderr.write(" info print env, skills, and prompt details for a named agent\n")
sys.stderr.write(" init interactively create a new agent and add it to bot-bottle.json\n")
sys.stderr.write(" list list available agents or active containers\n")
sys.stderr.write(" login register this host with a bot-bottle console\n")
sys.stderr.write(
" resume re-launch a bottle by its identity "
"(continues state from PRD 0016)\n"
@@ -90,7 +80,7 @@ def main(argv: list[str] | None = None) -> int:
usage()
die(f"unknown command: {command}")
mgr = StoreManager.instance()
if command not in NO_MIGRATION_COMMANDS and not mgr.is_migrated():
if not mgr.is_migrated():
sys.stderr.write("bot-bottle: database schema is out of date\n")
sys.stderr.write("Migrate now? [y/N] ")
sys.stderr.flush()
@@ -114,3 +104,7 @@ def main(argv: list[str] | None = None) -> int:
return e.code if isinstance(e.code, int) else 1
except KeyboardInterrupt:
return 130
if __name__ == "__main__":
sys.exit(main())
-15
View File
@@ -1,15 +0,0 @@
"""Entry point for `python -m bot_bottle.cli`.
`cli.py` at the repo root is the usual way in; this makes the package
runnable too, so the CLI works from an installed copy where there is no
`cli.py` on disk to point at.
"""
from __future__ import annotations
import sys
from . import main
if __name__ == "__main__":
sys.exit(main())
+10 -2
View File
@@ -3,10 +3,18 @@
from __future__ import annotations
import os
import sys
from pathlib import Path
from ..util import read_tty_line as read_tty_line
PROG = "cli.py"
USER_CWD = os.getcwd()
REPO_DIR = str(Path(__file__).resolve().parent.parent.parent)
def read_tty_line() -> str:
"""Mirror `IFS= read -r REPLY </dev/tty`. Falls back to stdin."""
try:
with open("/dev/tty", "r", encoding="utf-8") as tty:
return tty.readline().rstrip("\n")
except OSError:
return sys.stdin.readline().rstrip("\n")
+3 -7
View File
@@ -21,20 +21,16 @@ from __future__ import annotations
import sys
from ..backend import get_bottle_backend, has_backend, known_backend_names
from ..backend import get_bottle_backend, known_backend_names
from ..log import info
from ._common import read_tty_line
def cmd_cleanup(_argv: list[str]) -> int:
# Order: stable backend iteration so the y/N output is
# deterministic across runs. Skip backends whose runtime
# isn't available on this host so e.g. macos-container
# doesn't error on Linux.
# deterministic across runs.
plans = [
(name, get_bottle_backend(name))
for name in known_backend_names()
if has_backend(name)
(name, get_bottle_backend(name)) for name in known_backend_names()
]
prepared = [(name, b, b.prepare_cleanup()) for name, b in plans]
-168
View File
@@ -1,168 +0,0 @@
"""bb login — register this host with a bot-bottle console.
Opens a device-authorization flow against the target console, waits for the
operator to approve, then writes access and refresh tokens to
~/.bot-bottle/console.json (or $BOT_BOTTLE_ROOT/console.json).
Usage:
bb login [--console-url URL] [--label LABEL]
Flags:
--console-url URL Target console URL (overrides BB_CONSOLE_URL env var)
--label LABEL Host label shown in the console (default: hostname)
"""
from __future__ import annotations
import json
import os
import socket
import sys
import tempfile
import time
import urllib.error
import urllib.request
from pathlib import Path
from typing import Any
from ..paths import bot_bottle_root
_CONSOLE_URL_ENV = "BB_CONSOLE_URL"
_POLL_SLEEP = 2 # seconds between polls; matches console's poll_interval default
def _usage() -> None:
sys.stderr.write(
"usage: bb login [--console-url URL] [--label LABEL]\n"
"\n"
"Options:\n"
" --console-url URL Console base URL (or BB_CONSOLE_URL env var)\n"
" --label LABEL Host label shown in the console (default: hostname)\n"
)
def _flag(argv: list[str], name: str) -> str | None:
for i, arg in enumerate(argv):
if arg == name and i + 1 < len(argv):
return argv[i + 1]
if arg.startswith(f"{name}="):
return arg[len(name) + 1:]
return None
def _post(url: str, payload: dict[str, Any]) -> dict[str, Any]:
data = json.dumps(payload).encode()
req = urllib.request.Request(
url, data=data, headers={"Content-Type": "application/json"}
)
with urllib.request.urlopen(req, timeout=10) as resp:
return json.loads(resp.read())
def _get(url: str) -> tuple[int, dict[str, Any]]:
req = urllib.request.Request(url)
try:
with urllib.request.urlopen(req, timeout=10) as resp:
return resp.status, json.loads(resp.read())
except urllib.error.HTTPError as e:
return e.code, {}
def _save_credentials(
console_url: str, host_id: str, access_token: str, refresh_token: str
) -> Path:
path = bot_bottle_root() / "console.json"
path.parent.mkdir(parents=True, exist_ok=True)
content = (
json.dumps(
{
"url": console_url,
"host_id": host_id,
"access_token": access_token,
"refresh_token": refresh_token,
},
indent=2,
)
+ "\n"
)
fd, tmp_path_str = tempfile.mkstemp(dir=path.parent, prefix=".console-")
tmp = Path(tmp_path_str)
try:
tmp.chmod(0o600)
with os.fdopen(fd, "w") as f:
f.write(content)
os.replace(tmp, path)
except OSError:
try:
tmp.unlink()
except OSError:
pass
raise
return path
def cmd_login(argv: list[str]) -> int:
if "--help" in argv or "-h" in argv:
_usage()
return 0
console_url = _flag(argv, "--console-url") or os.environ.get(_CONSOLE_URL_ENV)
if not console_url:
sys.stderr.write(
"bb login: --console-url or BB_CONSOLE_URL is required\n"
)
return 1
console_url = console_url.rstrip("/")
label = _flag(argv, "--label") or socket.gethostname()
try:
resp = _post(f"{console_url}/api/v1/hosts/authorize", {"label": label})
except (OSError, ValueError) as exc:
sys.stderr.write(f"bb login: failed to start authorization: {exc}\n")
return 1
device_code = resp["device_code"]
user_code = resp["user_code"]
expires_in = resp.get("expires_in", 300)
poll_sleep = max(1, min(int(resp.get("poll_interval", _POLL_SLEEP)), 60))
sys.stderr.write(
f"\nOpen this URL in your browser to authorize this host:\n\n"
f" {console_url}/hosts/authorize?code={user_code}\n\n"
f"Waiting for approval"
)
deadline = time.monotonic() + expires_in
while time.monotonic() < deadline:
sys.stderr.write(".")
sys.stderr.flush()
time.sleep(poll_sleep)
try:
code, result = _get(
f"{console_url}/api/v1/hosts/authorize/{device_code}"
)
except (OSError, ValueError):
continue
if code == 410:
break
st = result.get("status")
if st == "approved":
sys.stderr.write("\n\nApproved.\n")
path = _save_credentials(
console_url,
result["host_id"],
result["access_token"],
result["refresh_token"],
)
sys.stderr.write(f"Credentials saved to {path}\n")
return 0
if st == "denied":
sys.stderr.write("\n\nDenied by operator.\n")
return 1
sys.stderr.write("\n\nAuthorization timed out.\n")
return 1
+22 -42
View File
@@ -27,6 +27,7 @@ from ..backend import (
BottleSpec,
enumerate_active_agents,
get_bottle_backend,
known_backend_names,
)
from ..backend.docker import util as docker_mod
from ..backend.docker.bottle_plan import DockerBottlePlan
@@ -35,7 +36,6 @@ from ..bottle_state import (
is_preserved,
mark_preserved,
)
from ..image_cache import StaleImageError
from ..log import info, die
from ..manifest import Manifest, ManifestIndex
from ._common import PROG, USER_CWD, read_tty_line
@@ -57,6 +57,15 @@ def cmd_start(argv: list[str]) -> int:
"into a cached layer."
),
)
parser.add_argument(
"--backend",
choices=known_backend_names(),
default=None,
help=(
"backend to launch the bottle on (default: $BOT_BOTTLE_BACKEND "
"or host auto-selection). Overrides the env var when set."
),
)
parser.add_argument(
"--headless",
action="store_true",
@@ -65,14 +74,6 @@ def cmd_start(argv: list[str]) -> int:
"skip all prompts. For orchestrators, CI, and webhooks."
),
)
parser.add_argument(
"--cached-images",
action="store_true",
help=(
"quickstart with existing local agent and sidecar images; "
"only valid with --headless"
),
)
parser.add_argument(
"--bottle",
action="append",
@@ -105,8 +106,6 @@ def cmd_start(argv: list[str]) -> int:
help="agent name defined in bot-bottle.json (omit to pick interactively)",
)
args = parser.parse_args(argv)
if args.cached_images and not args.headless:
die("--cached-images is only supported with --headless")
dry_run = args.dry_run or os.environ.get("BOT_BOTTLE_DRY_RUN") == "1"
if args.no_cache or os.environ.get("BOT_BOTTLE_NO_CACHE") == "1":
@@ -116,10 +115,11 @@ def cmd_start(argv: list[str]) -> int:
os.environ["BOT_BOTTLE_NO_CACHE"] = "1"
manifest = ManifestIndex.resolve(USER_CWD)
backend_name: str | None = args.backend
if args.headless:
return _start_headless(
manifest, args, dry_run=dry_run
manifest, args, dry_run=dry_run, backend_name=backend_name
)
agent_name: str | None = args.name
@@ -158,10 +158,6 @@ def cmd_start(argv: list[str]) -> int:
label, color = tui.name_color_modal(default_label=agent_name)
label, color = _resolve_unique_label(label, color)
image_policy = _select_image_policy()
if image_policy is None:
return 0
spec = BottleSpec(
manifest=manifest,
agent_name=agent_name,
@@ -170,11 +166,11 @@ def cmd_start(argv: list[str]) -> int:
label=label,
color=color,
bottle_names=bottle_names,
image_policy=image_policy,
)
return _launch_bottle(
spec,
dry_run=dry_run,
backend_name=backend_name,
)
@@ -186,6 +182,7 @@ def _start_headless(
args: argparse.Namespace,
*,
dry_run: bool,
backend_name: str | None,
) -> int:
"""Non-interactive launch path for orchestrators / CI / webhooks.
@@ -229,11 +226,11 @@ def _start_headless(
color=args.color or "",
bottle_names=bottle_names,
headless=True,
image_policy="cached" if args.cached_images else "fresh",
)
return _launch_bottle(
spec,
dry_run=dry_run,
backend_name=backend_name,
assume_yes=True,
headless_prompt_text=prompt,
)
@@ -271,18 +268,15 @@ def prepare_with_preflight(
injected callable, prompt y/N via the injected callable.
`backend_name` selects which backend prepares the plan
(`None` `$BOT_BOTTLE_BACKEND` host auto-selection).
When `spec.headless` is True the docker-fallback prompt is suppressed:
auto-selection dies with an actionable message rather than blocking
on a TTY read (which would hang CI, webhook dispatch, and orchestrators).
(`None` `$BOT_BOTTLE_BACKEND` host auto-selection). The CLI
passes whatever `--backend` resolved to.
Returns `(plan, identity)`. `plan` is None on dry-run or
operator-N, but `identity` is set as soon as `backend.prepare`
returns so callers can reap the prepare-time state dir via
`settle_state(identity)` in their finally exactly the existing
semantics."""
backend = get_bottle_backend(backend_name, prompt=not spec.headless)
backend = get_bottle_backend(backend_name)
plan = backend.prepare(spec, stage_dir=stage_dir)
identity = _identity_from_plan(plan)
@@ -412,13 +406,6 @@ def _text_prompt_yes() -> bool:
return reply in ("y", "Y", "yes", "YES")
def _select_image_policy() -> str | None:
return tui.filter_select(
["fresh", "cached"],
title="Select image startup mode",
)
def _text_render_preflight():
def _render(plan: DockerBottlePlan, backend_name: str) -> None:
print(file=sys.stderr)
@@ -524,8 +511,6 @@ def _manifest_to_yaml(manifest: Manifest) -> str:
lines.append(f" scheme: {r.AuthScheme}")
lines.append(f" supervise: {'true' if bottle.supervise else 'false'}")
if bottle.nested_containers:
lines.append(" nested_containers: true")
return "\n".join(lines)
@@ -563,15 +548,6 @@ def _launch_bottle(
return 0
backend = get_bottle_backend(backend_name)
try:
backend.prelaunch_checks(plan)
except StaleImageError as exc:
if assume_yes:
die(str(exc))
sys.stderr.write(f"bot-bottle: {exc}\nLaunch anyway? [y/N] ")
sys.stderr.flush()
if read_tty_line() not in ("y", "Y", "yes", "YES"):
return 0
with backend.launch(plan) as bottle:
agent_provider_template = getattr(plan, "agent_provider_template", "claude")
extra_args: tuple[str, ...] = ()
@@ -590,6 +566,10 @@ def _launch_bottle(
f"session ended (exit {exit_code}); "
f"container {bottle.name} will be removed"
)
# While the container is still alive: always snapshot the
# transcript and — if the agent exited non-zero — mark
# the state for preservation. This picks up crashes /
# Ctrl-Cs / OOM kills before cleanup removes the state dir.
if agent_provider_template == "claude":
capture_claude_session_state(identity, exit_code)
return 0
-71
View File
@@ -1,71 +0,0 @@
"""SQLite-backed bot-bottle configuration store."""
from __future__ import annotations
from pathlib import Path
try:
from .db_store import DbStore
from .migrations import TableMigrations
from .paths import host_db_path
except ImportError:
from db_store import DbStore # type: ignore[import-not-found] # pylint: disable=import-error,no-name-in-module
from migrations import TableMigrations # type: ignore[import-not-found] # pylint: disable=import-error,no-name-in-module
from paths import host_db_path # type: ignore[import-not-found] # pylint: disable=import-error,no-name-in-module
DEFAULT_CACHED_IMAGE_STALE_WARNING_DAYS = 1
class ConfigStore(DbStore):
"""SQLite configuration for host-side bot-bottle settings."""
def __init__(self, db_path: Path | None = None) -> None:
migrations = TableMigrations("config_store", [
# v1 — host-side bot-bottle settings
"""
CREATE TABLE IF NOT EXISTS bot_bottle_config (
id INTEGER PRIMARY KEY CHECK (id = 1),
cached_image_stale_warning_days INTEGER NOT NULL DEFAULT 1
)
""",
])
super().__init__(db_path or host_db_path(), migrations)
def cached_image_stale_warning_days(self) -> int:
if not self.db_path.is_file():
return DEFAULT_CACHED_IMAGE_STALE_WARNING_DAYS
with self._connect() as conn:
row = conn.execute(
"""
SELECT cached_image_stale_warning_days
FROM bot_bottle_config
WHERE id = 1
""",
).fetchone()
if row is None:
return DEFAULT_CACHED_IMAGE_STALE_WARNING_DAYS
try:
return int(row["cached_image_stale_warning_days"])
except (TypeError, ValueError):
return DEFAULT_CACHED_IMAGE_STALE_WARNING_DAYS
def set_cached_image_stale_warning_days(self, days: int) -> Path:
with self._connect() as conn:
conn.execute(
"""
INSERT INTO bot_bottle_config (id, cached_image_stale_warning_days)
VALUES (1, ?)
ON CONFLICT(id) DO UPDATE SET
cached_image_stale_warning_days = excluded.cached_image_stale_warning_days
""",
(days,),
)
self._chmod()
return self.db_path
__all__ = [
"DEFAULT_CACHED_IMAGE_STALE_WARNING_DAYS",
"ConfigStore",
]
-17
View File
@@ -1,17 +0,0 @@
"""Shared wire-protocol constants for gateway-bundled modules.
Single source of truth for values that appear across the egress addon,
git-http backend, supervise server, and git-gate renderer. Importing
from this module instead of duplicating the literals means a rename is
a one-line change and is caught by the type checker at the import site."""
# App-layer identity token header. Delivered as proxy credentials
# (HTTPS_PROXY=http://<bottle_id>:<token>@gw) by launch; the egress
# addon reads and strips it, the supervise server and git-http backend
# read it for attribution, and none of them forward it upstream.
IDENTITY_HEADER = "x-bot-bottle-identity"
# Shared timeout (seconds) for all git-gate subprocess and CGI calls:
# git daemon (--timeout/--init-timeout), the access-hook subprocess in
# git_http_backend, and the git http-backend CGI subprocess.
GIT_GATE_TIMEOUT_SECS = 15
+2 -15
View File
@@ -10,7 +10,7 @@
# Current Node LTS; slim variant keeps the image small while still
# providing apt-get for any future additions.
FROM node:22-trixie-slim
FROM node:22-slim
# Install runtime system deps. claude-code shells out to git for several
# features (status checks, commits, PR creation) — without git in the
@@ -21,15 +21,7 @@ FROM node:22-trixie-slim
# to it) works against egress's bumped TLS without the agent needing
# local DNS.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
git \
ca-certificates \
curl \
openssh-client \
podman \
ripgrep \
iproute2 \
dnsutils \
&& apt-get install -y --no-install-recommends git ca-certificates curl ripgrep iproute2 dnsutils \
&& rm -rf /var/lib/apt/lists/*
# App-specific deps. Python isn't required by claude-code itself
@@ -47,11 +39,6 @@ RUN apt-get update \
RUN npm install -g --no-fund --no-audit @anthropic-ai/claude-code@2.1.172 \
&& npm cache clean --force
# Git reads both ~/.gitconfig and ~/.config/git/config. Keep its XDG config
# path traversable by the non-root runtime user so permission errors do not
# suppress bot-bottle's git-gate insteadOf rules.
RUN install -d -o node -g node -m 755 /home/node/.config /home/node/.config/git
# Run as a non-root user. The node image already provides a `node` user
# (uid 1000) with a home directory, which is where claude-code will write
# its session state.
+5 -17
View File
@@ -23,9 +23,8 @@ from ...agent_provider import (
provider_startup_args,
)
from ...backend.docker import util as docker_mod
from ...egress import CLAUDE_HOST_CREDENTIAL_TOKEN_REF, EgressRoute
from ...egress import EgressRoute
from ...log import die, info, warn
from .claude_auth import claude_host_access_token
if TYPE_CHECKING:
@@ -119,6 +118,7 @@ class ClaudeAgentProvider(AgentProvider):
color: str = "",
provider_settings: dict[str, object] | None = None,
) -> AgentProvisionPlan:
del forward_host_credentials, host_env
resolved_guest_env = dict(guest_env or {})
startup_args = provider_startup_args(provider_settings)
guest_home = self.guest_home
@@ -180,24 +180,13 @@ class ClaudeAgentProvider(AgentProvider):
claude_settings,
f"{guest_home}/.claude/settings.json",
))
provisioned_env: dict[str, str] = {}
if forward_host_credentials:
_host_env = host_env or dict(os.environ)
provisioned_env[CLAUDE_HOST_CREDENTIAL_TOKEN_REF] = (
claude_host_access_token(_host_env)
)
cred_token_ref = (
CLAUDE_HOST_CREDENTIAL_TOKEN_REF if forward_host_credentials
else auth_token
)
egress_routes = (EgressRoute(
host="api.anthropic.com",
auth_scheme="Bearer" if (auth_token or forward_host_credentials) else "",
token_ref=cred_token_ref,
auth_scheme="Bearer" if auth_token else "",
token_ref=auth_token,
),)
hidden_env_names: frozenset[str] = frozenset()
if auth_token or forward_host_credentials:
if auth_token:
env_vars["CLAUDE_CODE_OAUTH_TOKEN"] = "egress-placeholder"
hidden_env_names = frozenset({"CLAUDE_CODE_OAUTH_TOKEN"})
@@ -219,7 +208,6 @@ class ClaudeAgentProvider(AgentProvider):
files=tuple(files),
egress_routes=egress_routes,
hidden_env_names=hidden_env_names,
provisioned_env=provisioned_env,
)
def provision_skills(self, plan: "BottlePlan", bottle: "Bottle") -> None:
-114
View File
@@ -1,114 +0,0 @@
"""Host Claude auth helpers.
Reads the host's Claude Code credentials and returns only the access
token needed by egress. Does not expose refresh tokens or raw payloads.
Credential storage by platform:
Linux ~/.claude/.credentials.json
macOS macOS Keychain, service "Claude Code-credentials"
(file path is tried first; Keychain is the fallback)
"""
from __future__ import annotations
import json
import os
import subprocess
import sys
from datetime import datetime, timezone
from pathlib import Path
from ...log import die
_KEYCHAIN_SERVICE = "Claude Code-credentials"
def claude_auth_path(host_env: dict[str, str] | None = None) -> Path:
env = os.environ if host_env is None else host_env
home = env.get("HOME")
if home:
return Path(home) / ".claude" / ".credentials.json"
return Path.home() / ".claude" / ".credentials.json"
def _read_keychain() -> dict[str, object] | None:
"""Try the macOS Keychain. Returns parsed JSON dict or None."""
if sys.platform != "darwin":
return None
try:
result = subprocess.run(
["security", "find-generic-password", "-s", _KEYCHAIN_SERVICE, "-w"],
capture_output=True,
text=True,
timeout=10,
)
except (FileNotFoundError, subprocess.TimeoutExpired):
return None
if result.returncode != 0 or not result.stdout.strip():
return None
try:
raw = json.loads(result.stdout.strip())
except json.JSONDecodeError:
return None
return raw if isinstance(raw, dict) else None
def claude_host_access_token(
host_env: dict[str, str] | None = None,
*,
now: datetime | None = None,
) -> str:
path = claude_auth_path(host_env)
raw: dict[str, object] | None = None
if path.is_file():
try:
raw = json.loads(path.read_text())
except (OSError, json.JSONDecodeError) as e:
die(f"claude host credentials: could not read valid JSON at {path}: {e}")
if not isinstance(raw, dict):
die(f"claude host credentials: {path} must contain a JSON object")
else:
raw = _read_keychain()
if raw is None:
die(
f"claude host credentials: auth file missing at {path} and "
f"macOS Keychain lookup for '{_KEYCHAIN_SERVICE}' failed. "
"Run `claude login` on the host or disable "
"agent_provider.forward_host_credentials."
)
oauth = raw.get("claudeAiOauth")
if not isinstance(oauth, dict):
die(
"claude host credentials: claudeAiOauth is missing from credentials. "
"Run `claude login` on the host or disable "
"agent_provider.forward_host_credentials."
)
access_token = oauth.get("accessToken")
if not isinstance(access_token, str) or not access_token:
die(
"claude host credentials: claudeAiOauth.accessToken is missing or empty. "
"Run `claude login` on the host and restart the bottle."
)
# expiresAt is in milliseconds
expires_at = oauth.get("expiresAt")
if isinstance(expires_at, (int, float)):
check_now = now or datetime.now(timezone.utc)
exp_dt = datetime.fromtimestamp(float(expires_at) / 1000.0, timezone.utc)
if exp_dt <= check_now:
die(
"claude host credentials: host Claude access token is expired. "
"Run `claude login` on the host and restart the bottle."
)
return access_token
__all__ = [
"claude_auth_path",
"claude_host_access_token",
]
+2 -11
View File
@@ -3,17 +3,10 @@
# Mirrors the default Claude image shape: Node LTS, git/network tooling,
# non-root node user, and the provider CLI installed for that user.
FROM node:22-trixie-slim
FROM node:22-slim
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
git \
ca-certificates \
curl \
openssh-client \
podman \
procps \
ripgrep \
&& apt-get install -y --no-install-recommends git ca-certificates curl procps ripgrep \
&& rm -rf /var/lib/apt/lists/*
# App-specific deps. Python isn't required by codex itself
@@ -24,8 +17,6 @@ RUN apt-get update \
&& apt-get install -y --no-install-recommends python3 python3-pip python3-venv \
&& rm -rf /var/lib/apt/lists/*
RUN install -d -o node -g node -m 755 /home/node/.config /home/node/.config/git
USER node
WORKDIR /home/node
+2 -5
View File
@@ -2,7 +2,7 @@
#
# Node LTS, git/network tooling, and the Pi coding-agent CLI installed globally.
FROM node:22-trixie-slim
FROM node:22-slim
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
@@ -10,8 +10,6 @@ RUN apt-get update \
ca-certificates \
curl \
fd-find \
openssh-client \
podman \
ripgrep \
&& ln -s /usr/bin/fdfind /usr/local/bin/fd \
&& rm -rf /var/lib/apt/lists/*
@@ -23,8 +21,7 @@ RUN apt-get update \
RUN npm install -g --ignore-scripts --no-fund --no-audit @earendil-works/pi-coding-agent \
&& npm cache clean --force
RUN install -d -o node -g node -m 755 /home/node/.config /home/node/.config/git \
&& mkdir -p /home/node/.pi/agent \
RUN mkdir -p /home/node/.pi/agent \
/home/node/.pi/context-mode/sessions \
/tmp/pi-subagents-uid-1000 \
&& chown -R node:node /home/node/.pi /tmp \
+2 -12
View File
@@ -3,7 +3,6 @@
from __future__ import annotations
import sqlite3
from contextlib import contextmanager
from pathlib import Path
try:
@@ -29,21 +28,12 @@ class DbStore:
conn.row_factory = sqlite3.Row
return conn
@contextmanager
def _connection(self):
conn = self._connect()
try:
with conn:
yield conn
finally:
conn.close()
def is_migrated(self) -> bool:
"""Return True if the DB is fully up-to-date, False if migration is needed."""
if not self.db_path.exists():
return False
try:
with self._connection() as conn:
with self._connect() as conn:
row = conn.execute(
"SELECT version FROM schema_versions WHERE module = ?",
(self._migrations.schema_key,),
@@ -55,7 +45,7 @@ class DbStore:
def migrate(self) -> None:
"""Apply any pending migrations and set permissions on the DB file."""
with self._connection() as conn:
with self._connect() as conn:
self._migrations.apply(conn)
self._chmod()
+7 -3
View File
@@ -3,8 +3,9 @@
Pure Python, no mitmproxy dependency. Each detector is a module-level
function returning `ScanResult | None`.
Available in the gateway via the installed `bot_bottle` package
(see `Dockerfile.gateway`).
Ships flat into the gateway image alongside
`egress_addon_core.py` both this file and the package source use
the same try/except import shim pattern.
"""
from __future__ import annotations
@@ -19,7 +20,10 @@ from math import log2
from collections import Counter
from urllib.parse import quote as url_quote
from .egress_addon_core import ScanResult
try:
from egress_addon_core import ScanResult # type: ignore[import-not-found]
except ImportError: # pragma: no cover - host-side path
from .egress_addon_core import ScanResult
# ---------------------------------------------------------------------------
-7
View File
@@ -30,7 +30,6 @@ if TYPE_CHECKING:
from .manifest import ManifestBottle
CODEX_HOST_CREDENTIAL_TOKEN_REF = "BOT_BOTTLE_CODEX_HOST_ACCESS_TOKEN"
CLAUDE_HOST_CREDENTIAL_TOKEN_REF = "BOT_BOTTLE_CLAUDE_HOST_ACCESS_TOKEN"
EGRESS_HOSTNAME = "egress"
@@ -146,7 +145,6 @@ def egress_manifest_routes(
outbound_detectors=r.OutboundDetectors,
inbound_detectors=r.InboundDetectors,
outbound_on_match=r.OutboundOnMatch,
preserve_auth=r.PreserveAuth,
))
return tuple(out)
@@ -255,8 +253,6 @@ def _route_to_yaml_fields(r: Route) -> dict[str, object]:
fields["matches"] = matches_data
if r.git_fetch:
fields["git"] = {"fetch": True}
if r.preserve_auth:
fields["preserve_auth"] = True
if (
r.outbound_detectors is not None
or r.inbound_detectors is not None
@@ -337,8 +333,6 @@ def egress_render_routes(
lines.append(" git:")
if git_dict.get("fetch") is True:
lines.append(" fetch: true")
if f.get("preserve_auth") is True:
lines.append(" preserve_auth: true")
if "dlp" in f:
dlp_dict: dict[str, object] = f["dlp"] # type: ignore
lines.append(" dlp:")
@@ -406,7 +400,6 @@ class Egress(ABC):
)
__all__ = [
"CLAUDE_HOST_CREDENTIAL_TOKEN_REF",
"CODEX_HOST_CREDENTIAL_TOKEN_REF",
"EGRESS_HOSTNAME",
"EGRESS_ROUTES_FILENAME",
+27 -10
View File
@@ -15,9 +15,7 @@ import typing
from mitmproxy import http # type: ignore[import-not-found] # pylint: disable=import-error
from bot_bottle.constants import IDENTITY_HEADER
from bot_bottle.dlp_detectors import redact_tokens, strip_crlf
from bot_bottle.egress_addon_core import (
from egress_addon_core import ( # type: ignore[import-not-found] # pylint: disable=import-error
LOG_BLOCKS,
LOG_FULL,
DEFAULT_OUTBOUND_ON_MATCH,
@@ -40,8 +38,24 @@ from bot_bottle.egress_addon_core import (
scan_inbound,
scan_outbound,
)
from bot_bottle import supervise as _sv
from bot_bottle.policy_resolver import PolicyResolver
try:
from dlp_detectors import redact_tokens, strip_crlf # type: ignore[import-not-found]
except ImportError: # pragma: no cover - host-side path
from bot_bottle.dlp_detectors import ( # type: ignore[import-not-found]
redact_tokens,
strip_crlf,
)
try:
import supervise as _sv # type: ignore[import-not-found]
except ImportError: # pragma: no cover - host-side path
from bot_bottle import supervise as _sv # type: ignore[import-not-found]
try:
from policy_resolver import PolicyResolver # type: ignore[import-not-found]
except ImportError: # pragma: no cover - host-side path
from bot_bottle.policy_resolver import PolicyResolver
INTROSPECT_HOST = "_egress.local"
@@ -52,6 +66,13 @@ INTROSPECT_HOST = "_egress.local"
# back to — so an unset value is a fatal misconfiguration (see __init__).
ORCHESTRATOR_URL_ENV = "BOT_BOTTLE_ORCHESTRATOR_URL"
# App-layer identity token. Delivered as proxy credentials
# (`HTTPS_PROXY=http://<bottle_id>:<token>@gw`): clients honor it as part of
# the proxy protocol without app changes, and the addon reads + strips it so
# it never leaks upstream. The legacy `x-bot-bottle-identity` request header
# is still stripped defensively (git-http uses that header on its own port).
IDENTITY_HEADER = "x-bot-bottle-identity"
# Per-flow key under which `request()` stashes the resolved (Config, supervise
# slug, env) so the later `response()` and `websocket_message()` hooks scan
# against the *calling bottle's* policy — the same one the request was decided
@@ -367,10 +388,7 @@ class EgressAddon:
# Strip agent-set Authorization after DLP scan so smuggled tokens
# are caught above; the route may inject gateway-owned auth below.
# Routes with preserve_auth=True pass the header through as-is so the
# agent's own credentials (e.g. registry bearer tokens) reach the upstream.
if route is None or not route.preserve_auth:
flow.request.headers.pop("authorization", None)
flow.request.headers.pop("authorization", None)
# Build headers mapping for match evaluation
req_headers = {k.lower(): v for k, v in flow.request.headers.items()}
@@ -382,7 +400,6 @@ class EgressAddon:
env,
request_method=flow.request.method,
request_headers=req_headers,
deny_reason=config.deny_reason,
)
if decision.action == "block":
+40 -74
View File
@@ -6,9 +6,9 @@ exercise the parse + decision functions without depending on the
`mitmproxy.http.HTTPFlow` API and is loaded inside the gateway
container.
Imports: stdlib + sibling package modules (`yaml_subset`,
`egress_dlp_config`). Available in the gateway via the installed
`bot_bottle` package (see `Dockerfile.gateway`)."""
Imports: stdlib + `yaml_subset` (which is itself stdlib-only and
ships flat into the gateway image alongside this file
see `Dockerfile.gateway`)."""
from __future__ import annotations
@@ -16,20 +16,36 @@ import re
import typing
from dataclasses import dataclass
from .yaml_subset import YamlSubsetError, parse_yaml_subset
try:
from yaml_subset import YamlSubsetError, parse_yaml_subset # type: ignore[import-not-found]
except ImportError: # pragma: no cover - host-side path
from .yaml_subset import YamlSubsetError, parse_yaml_subset
# DLP detector-config parsing lives in a sibling module. Re-exported below
# so existing `from egress_addon_core import ON_MATCH_*` callers keep working.
from .egress_dlp_config import (
DEFAULT_OUTBOUND_ON_MATCH,
INBOUND_DETECTOR_NAMES,
ON_MATCH_BLOCK,
ON_MATCH_REDACT,
ON_MATCH_SUPERVISE,
OUTBOUND_DETECTOR_NAMES,
OUTBOUND_ON_MATCH_VALUES,
parse_dlp_block,
)
# DLP detector-config parsing lives in a sibling module (also flat-bundled
# into the gateway — see Dockerfile.gateway). Re-exported below so existing
# `from egress_addon_core import ON_MATCH_*` callers keep working.
try:
from egress_dlp_config import ( # type: ignore[import-not-found]
DEFAULT_OUTBOUND_ON_MATCH,
INBOUND_DETECTOR_NAMES,
ON_MATCH_BLOCK,
ON_MATCH_REDACT,
ON_MATCH_SUPERVISE,
OUTBOUND_DETECTOR_NAMES,
OUTBOUND_ON_MATCH_VALUES,
parse_dlp_block,
)
except ImportError: # pragma: no cover - host-side path
from .egress_dlp_config import (
DEFAULT_OUTBOUND_ON_MATCH,
INBOUND_DETECTOR_NAMES,
ON_MATCH_BLOCK,
ON_MATCH_REDACT,
ON_MATCH_SUPERVISE,
OUTBOUND_DETECTOR_NAMES,
OUTBOUND_ON_MATCH_VALUES,
parse_dlp_block,
)
# ---------------------------------------------------------------------------
@@ -78,7 +94,6 @@ class Route:
inbound_detectors: tuple[str, ...] | None = None
# "" means unset → DEFAULT_OUTBOUND_ON_MATCH. See OUTBOUND_ON_MATCH_VALUES.
outbound_on_match: str = ""
preserve_auth: bool = False
LOG_OFF = 0 # no logging
@@ -90,14 +105,6 @@ LOG_FULL = 2 # log block/warn events + full request and response bodies
class Config:
routes: tuple[Route, ...]
log: int = LOG_OFF
# Why this Config is a deny-all, when it is one for a reason *other* than
# the bottle's own policy genuinely not listing the host. A deny-all is
# indistinguishable from "policy loaded, host not allowed" at the decision
# point — both are simply "no matching route" — so without this the
# operator sees `host X is not in the allowlist` and goes hunting for a
# missing route that was never the problem. Empty for a normally-parsed
# policy; `decide` prefers it over the allowlist wording when set.
deny_reason: str = ""
@dataclass(frozen=True)
@@ -309,18 +316,11 @@ def _parse_one(idx: int, raw: object) -> Route:
idx, host, raw_dict,
)
preserve_auth_raw = raw_dict.get("preserve_auth", False)
if preserve_auth_raw is not True and preserve_auth_raw is not False:
raise ValueError(
f"{label} ({host}): 'preserve_auth' must be a boolean"
)
preserve_auth: bool = preserve_auth_raw
for k in raw_dict:
if k not in ("host", "matches", "auth_scheme", "token_env", "dlp", "git", "preserve_auth"):
if k not in ("host", "matches", "auth_scheme", "token_env", "dlp", "git"):
raise ValueError(
f"{label} ({host}): unknown key {k!r}; accepted keys "
f"are 'host', 'matches', 'auth_scheme', 'token_env', 'dlp', 'git', 'preserve_auth'"
f"are 'host', 'matches', 'auth_scheme', 'token_env', 'dlp', 'git'"
)
return Route(
@@ -332,7 +332,6 @@ def _parse_one(idx: int, raw: object) -> Route:
outbound_detectors=outbound_detectors,
inbound_detectors=inbound_detectors,
outbound_on_match=outbound_on_match,
preserve_auth=preserve_auth,
)
@@ -385,8 +384,6 @@ def route_to_yaml_dict(r: Route) -> dict[str, object]:
dlp["outbound_on_match"] = r.outbound_on_match
if dlp:
d["dlp"] = dlp
if r.preserve_auth:
d["preserve_auth"] = True
return d
@@ -424,40 +421,16 @@ class PolicyResolverLike(typing.Protocol):
...
# Deny-all explanations. Each names the *actual* failure so an operator isn't
# sent looking for a missing egress route when the bottle never had a policy
# to begin with — the failure mode that made a bricked registration read like
# a misconfigured allowlist.
DENY_UNATTRIBUTED = (
"egress: this request was not attributed to any bottle, so no egress "
"policy applies and every host is denied. Either the bottle's registry "
"row is missing/ambiguous (torn down, or another bottle claimed its "
"source IP), or the request carried no matching identity token — check "
"that the caller's proxy URL includes it. This is not an allowlist problem."
)
DENY_UNPARSEABLE = (
"egress: this bottle's egress policy could not be parsed, so it is being "
"treated as deny-all. Fix the bottle's egress.routes; every host is denied "
"until it loads."
)
DENY_RESOLVER_ERROR = (
"egress: the orchestrator could not be reached to resolve this bottle's "
"egress policy, so every host is denied (fail-closed). Check that the "
"control plane is up; this is not an allowlist problem."
)
def _config_from_policy(policy: "str | None") -> "Config":
"""Parse a resolved policy blob into a Config, fail-closed: None / empty /
unparseable all become a deny-all Config (no routes every request
blocked). Each deny-all carries the reason it is one, so the block message
names the real fault instead of blaming the allowlist."""
blocked)."""
if not policy:
return Config(routes=(), deny_reason=DENY_UNATTRIBUTED)
return Config(routes=()) # unattributed or empty → deny-all
try:
return load_config(policy)
except ValueError:
return Config(routes=(), deny_reason=DENY_UNPARSEABLE)
return Config(routes=()) # unparseable policy → deny
def resolve_client_config(
@@ -471,7 +444,7 @@ def resolve_client_config(
try:
policy = resolver.resolve(client_ip, identity_token)
except Exception: # noqa: BLE001 # pylint: disable=broad-exception-caught
return Config(routes=(), deny_reason=DENY_RESOLVER_ERROR)
return Config(routes=()) # orchestrator unreachable/errored → deny
return _config_from_policy(policy)
@@ -500,7 +473,7 @@ def resolve_client_context(
client_ip, identity_token,
)
except Exception: # noqa: BLE001 # pylint: disable=broad-exception-caught
return Config(routes=(), deny_reason=DENY_RESOLVER_ERROR), "", {}
return Config(routes=()), "", {} # orchestrator unreachable/errored → deny
return _config_from_policy(policy), (bottle_id or ""), tokens
@@ -615,16 +588,12 @@ def decide(
*,
request_method: str = "GET",
request_headers: typing.Mapping[str, str] | None = None,
deny_reason: str = "",
) -> Decision:
"""`deny_reason` is `Config.deny_reason`: when the deny-all came from a
missing/unparseable policy rather than the bottle's own allowlist, report
that instead of implying a route is merely absent."""
route = match_route(routes, request_host)
if route is None:
return Decision(
action="block",
reason=deny_reason or (
reason=(
f"egress: host {request_host!r} is not in the "
f"bottle's egress.routes allowlist. Declare a "
f"route for it or remove the request."
@@ -899,9 +868,6 @@ __all__ = [
"is_git_push_request",
"is_git_fetch_request",
"load_config",
"DENY_UNATTRIBUTED",
"DENY_UNPARSEABLE",
"DENY_RESOLVER_ERROR",
"resolve_client_config",
"resolve_client_context",
"PolicyResolverLike",
+17 -37
View File
@@ -61,11 +61,6 @@ class _DaemonSpec:
_EGRESS_ONLY_ENV_PREFIXES: tuple[str, ...] = ("EGRESS_TOKEN_",)
_READY_GATED_DAEMONS: tuple[str, ...] = ("git-gate", "git-http")
# Daemons that must be requested explicitly via BOT_BOTTLE_GATEWAY_DAEMONS
# and are NOT started in the default (env-var-unset) case. The orchestrator
# only runs in the combined infra container, never in a standalone gateway.
_OPT_IN_DAEMONS: frozenset[str] = frozenset({"orchestrator"})
def _env_for_daemon(name: str, base_env: dict[str, str]) -> dict[str, str]:
"""Egress sees the full bundle env. Everyone else gets a copy
@@ -80,18 +75,11 @@ def _env_for_daemon(name: str, base_env: dict[str, str]) -> dict[str, str]:
}
# The orchestrator is listed first so it starts before the gateway daemons,
# giving the control plane a head start to accept /resolve calls. The gateway
# daemons tolerate early /resolve failures and retry per-request.
_DAEMONS: tuple[_DaemonSpec, ...] = (
_DaemonSpec("orchestrator", (
"python3", "-m", "bot_bottle.orchestrator",
"--host", "0.0.0.0", "--port", "8099", "--broker", "stub",
)),
_DaemonSpec("egress", ("/bin/sh", "/app/egress-entrypoint.sh")),
_DaemonSpec("git-gate", ("/bin/sh", "/git-gate-entrypoint.sh")),
_DaemonSpec("git-http", ("python3", "-m", "bot_bottle.git_http_backend")),
_DaemonSpec("supervise", ("python3", "-m", "bot_bottle.supervise_server")),
_DaemonSpec("git-http", ("python3", "/app/git_http_backend.py")),
_DaemonSpec("supervise", ("python3", "/app/supervise_server.py")),
)
@@ -115,20 +103,18 @@ def _selected_daemons(
env: dict[str, str],
all_daemons: Sequence[_DaemonSpec] | None = None,
) -> tuple[_DaemonSpec, ...]:
"""Filter the daemon set by the BOT_BOTTLE_GATEWAY_DAEMONS env var.
"""Filter the daemon set by the BOT_BOTTLE_GATEWAY_DAEMONS env
var. Unknown names in the list are ignored the caller is the
source of truth for which daemons are wired.
When the var is unset/empty, return all non-opt-in daemons (the
standard gateway subset). Opt-in daemons (e.g. `orchestrator`) only
run when explicitly named they never start in a plain gateway
container that doesn't set the env var. Unknown names are ignored.
`all_daemons` defaults to `_DAEMONS` resolved at call time (not at
definition time), so tests can pass a custom list."""
`all_daemons` defaults to `_DAEMONS` resolved at call time (not
at definition time), so tests can monkey-patch the module-level
`_DAEMONS` and have the new value take effect."""
if all_daemons is None:
all_daemons = _DAEMONS
raw = env.get("BOT_BOTTLE_GATEWAY_DAEMONS", "").strip()
if not raw:
return tuple(d for d in all_daemons if d.name not in _OPT_IN_DAEMONS)
return tuple(all_daemons)
wanted = {n.strip() for n in raw.split(",") if n.strip()}
return tuple(d for d in all_daemons if d.name in wanted)
@@ -150,7 +136,7 @@ def _pump(name: str, stream: IO[bytes]) -> None:
def _spawn(spec: _DaemonSpec) -> subprocess.Popen[bytes]:
env = _env_for_daemon(spec.name, dict(os.environ))
proc = subprocess.Popen( # pylint: disable=consider-using-with
proc = subprocess.Popen(
_argv_for_daemon(spec.name, spec.argv, env),
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
@@ -197,14 +183,6 @@ class _Supervisor:
except ProcessLookupError:
pass
def _sigkill_all(self) -> None:
for _, p in self.procs:
if p.poll() is None:
try:
p.kill()
except ProcessLookupError:
pass
def request_restart(self, daemon_name: str) -> bool:
"""Queue a daemon restart for the main loop to process.
@@ -257,7 +235,12 @@ class _Supervisor:
f"grace ({_GRACE_SECONDS:.0f}s) elapsed; SIGKILL on "
f"{', '.join(still_running)}"
)
self._sigkill_all()
for _, p in self.procs:
if p.poll() is None:
try:
p.kill()
except ProcessLookupError:
pass
done = all(p.poll() is not None for _, p in self.procs)
if done:
@@ -378,10 +361,7 @@ def main(argv: Sequence[str] | None = None) -> int:
# --signal HUP <bundle>` after writing routes.yaml. The kernel
# delivers SIGHUP to PID 1 (this supervisor); forward it to
# mitmdump so it reloads its addon.
signal.signal(
signal.SIGHUP,
lambda *_: sup.forward_signal(signal.SIGHUP, "egress"), # type: ignore[misc]
)
signal.signal(signal.SIGHUP, lambda *_: sup.forward_signal(signal.SIGHUP, "egress")) # type: ignore
while not sup.tick():
time.sleep(_POLL_INTERVAL)
+19 -19
View File
@@ -14,12 +14,18 @@ import shlex
from dataclasses import dataclass
from pathlib import Path
from .constants import GIT_GATE_TIMEOUT_SECS, IDENTITY_HEADER
from .manifest import ManifestBottle, ManifestGitEntry
# Short network alias for git-gate inside the gateway. The
# agent's `.gitconfig` insteadOf rewrites resolve through this name.
GIT_GATE_HOSTNAME = "git-gate"
# App-layer identity token header the agent's git sends to git-http and the
# gateway validates (mirrors egress_addon / git_http_backend IDENTITY_HEADER).
IDENTITY_HEADER = "x-bot-bottle-identity"
# Shared timeout (seconds) for all git-gate subprocess and CGI calls:
# git daemon (--timeout/--init-timeout), the access-hook subprocess in
# git_http_backend, and the git http-backend CGI subprocess.
GIT_GATE_TIMEOUT_SECS = 15
@dataclass(frozen=True)
@@ -419,24 +425,18 @@ PY
while IFS=' ' read -r old new ref; do
[ -z "$ref" ] && continue
[ "$new" = "$zero" ] && continue
# Scan only the commits this push introduces — those reachable from
# $new but not from any ref the gate already has. Everything already
# on the gate arrived via upstream mirror-fetch or a previously
# gitleaks-scanned push, so it's already-upstream or already-scanned;
# re-scanning it only resurfaces historical fixture findings.
#
# Applies to both new refs and updates. The old existing-branch range
# `$old..$new` walks commits reachable from the new tip but not the
# *old branch tip*: on a rebase/force-push onto a freshly-advanced
# main that pulls in all of main's new history (incl. the deliberate
# sandbox-escape gitleaks fixtures), blocking the push. `--not --all`
# excludes anything already on the gate regardless of ancestry, so it
# is also correct for non-fast-forward pushes (a rebase can skip
# commits off the direct path). Security-equivalent per PRD 0028's
# analysis: the bare repo's refs come only from trusted upstream
# mirror-fetch or gitleaks-gated pushes.
# See PRD 0028 (open question) / issues #106, #346.
log_opts="$new --not --all"
if [ "$old" = "$zero" ]; then
# New ref: scan only the commits this push introduces — those
# reachable from $new but not from any ref the gate already has.
# Everything already on the gate arrived via upstream mirror-fetch
# or a previously gitleaks-scanned push, so it's already-upstream
# or already-scanned; re-scanning it (the old `$new` full-ancestry
# range) only resurfaces historical findings and blocks every new
# branch. See PRD 0028 / issue #106.
log_opts="$new --not --all"
else
log_opts="$old..$new"
fi
echo "git-gate: gitleaks scanning $ref ($log_opts)" >&2
if ! gitleaks git --log-opts="$log_opts" --no-banner --redact 1>&2; then
echo "git-gate: gitleaks rejected push to $ref" >&2
+23 -2
View File
@@ -26,8 +26,16 @@ from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from urllib.parse import urlsplit
from bot_bottle.constants import GIT_GATE_TIMEOUT_SECS, IDENTITY_HEADER
from bot_bottle.policy_resolver import PolicyResolveError, PolicyResolver
# policy_resolver ships flat alongside this file in the gateway
# image (see Dockerfile.gateway); the bot_bottle.* fallback is the
# host-side / test path. Mirrors egress_addon's import shape.
try:
from policy_resolver import ( # type: ignore[import-not-found]
PolicyResolveError,
PolicyResolver,
)
except ImportError: # pragma: no cover - host-side path
from bot_bottle.policy_resolver import PolicyResolveError, PolicyResolver
DEFAULT_PORT = 9420
@@ -38,6 +46,12 @@ DEFAULT_PORT = 9420
# repo-root fallback.
ORCHESTRATOR_URL_ENV = "BOT_BOTTLE_ORCHESTRATOR_URL"
# App-layer identity token (defense-in-depth over the source-IP invariant);
# the agent injects it, the backend reads it for attribution and never
# forwards it to `git http-backend`. Mirrors egress_addon.IDENTITY_HEADER
# (duplicated, not imported: egress_addon pulls in mitmproxy).
IDENTITY_HEADER = "x-bot-bottle-identity"
# The base under which each bottle's `<bottle_id>` repo namespace is nested.
DEFAULT_REPO_ROOT = "/git"
@@ -75,6 +89,13 @@ def resolve_sandbox_root(
return None # bottle_id tried to escape the root → deny
return namespace
# Mirrors git_gate_render.GIT_GATE_TIMEOUT_SECS. Duplicated rather than
# imported: this module ships as a flat top-level sibling in the gateway
# bundle image (see Dockerfile.gateway), not as part of the bot_bottle
# package, so `bot_bottle.git_gate` and its dependency chain aren't
# available at runtime.
GIT_GATE_TIMEOUT_SECS = 15
# Bound memory use while still allowing ordinary git push packfiles.
MAX_BODY_BYTES = 100 * 1024 * 1024
-42
View File
@@ -1,42 +0,0 @@
"""Shared helpers for cached-image quickstart stale checks."""
from __future__ import annotations
from datetime import datetime, timezone
from pathlib import Path
try:
from .config_store import ConfigStore
except ImportError:
from config_store import ConfigStore # type: ignore[import-not-found] # pylint: disable=import-error,no-name-in-module
class StaleImageError(Exception):
"""Raised when a cached image or artifact exceeds the configured staleness
threshold. Callers can catch this to prompt interactively; headless paths
let it propagate as a fatal error."""
def check_stale(label: str, created_at: datetime) -> None:
"""Raise StaleImageError if `created_at` is older than the configured
stale-warning threshold. Negative threshold disables the check."""
threshold_days = ConfigStore().cached_image_stale_warning_days()
if threshold_days < 0:
return
now = datetime.now(timezone.utc)
created = created_at.astimezone(timezone.utc)
age = now - created
if age.total_seconds() <= threshold_days * 86400:
return
raise StaleImageError(
f"cached {label} is {age.days} day(s) old; "
"quickstart does not verify it matches the current Dockerfile/context"
)
def check_stale_path(label: str, path: Path) -> None:
"""Raise StaleImageError if `path`'s mtime exceeds the staleness threshold."""
check_stale(label, datetime.fromtimestamp(path.stat().st_mtime, tz=timezone.utc))
__all__ = ["StaleImageError", "check_stale", "check_stale_path"]
-1
View File
@@ -20,7 +20,6 @@ Bottle schema (frontmatter):
egress: { routes: [ <egress-route>, ... ] }
# route keys: host, matches, auth, role, dlp
supervise: <bool> # optional (default true)
nested_containers: <bool> # optional (default false)
Agent schema (frontmatter):
bottle: <bottle-name> # required
+4 -10
View File
@@ -25,9 +25,8 @@ class ManifestAgentProvider:
header, and sets a placeholder CLAUDE_CODE_OAUTH_TOKEN in the agent
so the Claude Code CLI starts.
`forward_host_credentials` forwards the host provider auth token into
the egress sidecar (Codex and Claude). For Codex this reads
`~/.codex/auth.json`; for Claude it reads `~/.claude/.credentials.json`.
`forward_host_credentials` forwards the host Codex auth token into
the egress daemon (Codex only).
"""
template: str = "claude"
@@ -93,15 +92,10 @@ class ManifestAgentProvider:
f"is only supported for built-in templates "
f"({', '.join(sorted(PROVIDER_TEMPLATES))})"
)
if forward_host_credentials and template not in {"codex", "claude"}:
if forward_host_credentials and template != "codex":
raise ManifestError(
f"bottle '{bottle_name}' agent_provider.forward_host_credentials "
"is only supported for templates 'codex' and 'claude'"
)
if forward_host_credentials and auth_token:
raise ManifestError(
f"bottle '{bottle_name}' agent_provider.forward_host_credentials "
"and auth_token both set; use one or the other"
"is currently only supported for template 'codex'"
)
settings = _parse_provider_settings(bottle_name, template, d.get("settings"))
return cls(
-13
View File
@@ -44,11 +44,6 @@ class ManifestBottle:
# daemon that exposes egress MCP tools to the agent. Set
# `supervise: false` to skip the gateway.
supervise: bool = True
# Guest-local container engine (issue #392). Not a host-daemon grant:
# backends implement it inside the bottle or reject it. Gated because it
# costs image weight, a resident service, and relaxed guest device modes
# that the majority of bottles never need.
nested_containers: bool = False
@classmethod
def from_dict(cls, name: str, raw: object) -> "ManifestBottle":
@@ -128,15 +123,7 @@ class ManifestBottle:
f"(was {type(supervise_raw).__name__})"
)
nested_raw = d.get("nested_containers", False)
if not isinstance(nested_raw, bool):
raise ManifestError(
f"bottle '{name}' nested_containers must be a boolean "
f"(was {type(nested_raw).__name__})"
)
return cls(
env=env, agent_provider=agent_provider, git=git,
git_user=git_user, egress=egress, supervise=supervise_raw,
nested_containers=nested_raw,
)
+2 -15
View File
@@ -71,7 +71,6 @@ class ManifestEgressRoute:
OutboundDetectors: tuple[str, ...] | None = None
InboundDetectors: tuple[str, ...] | None = None
OutboundOnMatch: str = ""
PreserveAuth: bool = False
@classmethod
def from_dict(cls, bottle_name: str, idx: int, raw: object) -> "ManifestEgressRoute":
@@ -191,22 +190,11 @@ class ManifestEgressRoute:
f"only 'fetch' is accepted"
)
# --- preserve_auth ---
preserve_auth = False
if "preserve_auth" in d:
raw_preserve_auth = d.get("preserve_auth")
if not isinstance(raw_preserve_auth, bool):
raise ManifestError(
f"{label} preserve_auth must be a boolean "
f"(was {type(raw_preserve_auth).__name__})"
)
preserve_auth = raw_preserve_auth
for k in d:
if k not in ("host", "matches", "auth", "role", "dlp", "git", "preserve_auth"):
if k not in ("host", "matches", "auth", "role", "dlp", "git"):
raise ManifestError(
f"{label} has unknown key {k!r}; accepted keys are "
f"'host', 'matches', 'auth', 'role', 'dlp', 'git', 'preserve_auth'"
f"'host', 'matches', 'auth', 'role', 'dlp', 'git'"
)
return cls(
@@ -219,7 +207,6 @@ class ManifestEgressRoute:
OutboundDetectors=outbound_detectors,
InboundDetectors=inbound_detectors,
OutboundOnMatch=outbound_on_match,
PreserveAuth=preserve_auth,
)
+1 -18
View File
@@ -15,16 +15,7 @@ def merge_bottles_runtime(bottles: "list[ManifestBottle]") -> "ManifestBottle":
the same field-merge rules as the file-based extends machinery:
env: dict merge, later wins; git_user: per-field overlay, later
wins on non-empty; git (repos): union by name, later wins; egress
routes: concatenate; agent_provider, supervise: later
replaces; nested_containers: OR (see below).
nested_containers is OR'd rather than replaced because these objects
are already resolved: a bottle that never mentions the key is
indistinguishable from one that sets it false, so "later replaces"
would let any bottle composed after a container-enabled one silently
drop the capability. The file-based `extends:` path still sees the
raw keys, so there an explicit `nested_containers: false` in a child
turns it back off.
routes: concatenate; agent_provider, supervise: later replaces.
"""
if not bottles:
raise ValueError("merge_bottles_runtime requires at least one bottle")
@@ -63,7 +54,6 @@ def _merge_two_bottles_runtime(base: "ManifestBottle", override: "ManifestBottle
git_user=merged_git_user,
egress=merged_egress,
supervise=override.supervise,
nested_containers=base.nested_containers or override.nested_containers,
)
@@ -216,7 +206,6 @@ def _fold_two_bottles(
git_user=merged_git_user,
egress=merged_egress,
supervise=later.supervise,
nested_containers=earlier.nested_containers or later.nested_containers,
), merged_repos_raw
@@ -277,11 +266,6 @@ def _merge_bottles(
merged_supervise = (
child.supervise if "supervise" in child_raw else parent.supervise
)
merged_nested_containers = (
child.nested_containers
if "nested_containers" in child_raw
else parent.nested_containers
)
validate_egress_routes(name, merged_egress.routes)
return ManifestBottle(
@@ -291,7 +275,6 @@ def _merge_bottles(
git_user=merged_git_user,
egress=merged_egress,
supervise=merged_supervise,
nested_containers=merged_nested_containers,
)
+1 -4
View File
@@ -16,10 +16,7 @@ _FILENAME_RX = re.compile(r"^[a-z][a-z0-9-]*$")
# sets dies with a "did you mean" pointer: typos should not silently
# ghost into an empty config.
BOTTLE_KEYS = frozenset(
{
"env", "extends", "agent_provider", "git-gate", "egress",
"supervise", "nested_containers",
}
{"env", "extends", "agent_provider", "git-gate", "egress", "supervise"}
)
AGENT_KEYS_REQUIRED: frozenset[str] = frozenset()
AGENT_KEYS_OPTIONAL = frozenset({"bottle", "skills", "git-gate"})
-15
View File
@@ -15,7 +15,6 @@ from __future__ import annotations
import json
import urllib.error
import urllib.request
from collections.abc import Iterable
from dataclasses import dataclass
from ..paths import host_control_plane_token
@@ -148,20 +147,6 @@ class OrchestratorClient:
raise OrchestratorClientError(f"teardown {bottle_id}: HTTP {status}")
return True
def reconcile(
self, live_source_ips: Iterable[str], *, grace_seconds: float | None = None,
) -> list[str]:
"""Drop registry rows for bottles that are no longer running
(`POST /reconcile`), returning the reaped bottle ids. `live_source_ips`
is the caller's enumeration of its live bottles — the orchestrator
can't see the backend from inside the infra container."""
body: dict[str, object] = {"live_source_ips": list(live_source_ips)}
if grace_seconds is not None:
body["grace_seconds"] = grace_seconds
payload = self._ok("POST", "/reconcile", body)
reaped = payload.get("reaped")
return [r for r in reaped if isinstance(r, str)] if isinstance(reaped, list) else []
def set_policy(self, bottle_id: str, policy: str) -> bool:
"""Live-reload a bottle's policy (`PUT /bottles/<id>/policy`). False on
404 (unknown bottle)."""
-107
View File
@@ -1,107 +0,0 @@
"""Per-host orchestrator configuration store (settings in bot-bottle.db).
Co-tenants the shared `bot-bottle.db` via the `DbStore` framework. Settings
are readable by the host launch path directly (no HTTP round-trip to the
orchestrator), so they take effect even before the orchestrator is reachable.
"""
from __future__ import annotations
import os
import sqlite3
from pathlib import Path
from ..db_store import DbStore
from ..migrations import TableMigrations
from ..paths import host_db_path
TEARDOWN_TIMEOUT_ENV = "BOT_BOTTLE_ORCHESTRATOR_TEARDOWN_TIMEOUT_SECONDS"
DEFAULT_TEARDOWN_TIMEOUT_SECONDS = 30.0
_MIGRATIONS = TableMigrations(
"orchestrator_config",
[
"""
CREATE TABLE IF NOT EXISTS orchestrator_config (
id INTEGER PRIMARY KEY CHECK (id = 1),
teardown_timeout_seconds REAL
)
""",
],
)
class OrchestratorConfigStore(DbStore):
"""Orchestrator settings in the shared host DB."""
def __init__(self, db_path: Path | None = None) -> None:
super().__init__(db_path or host_db_path(), _MIGRATIONS)
def _connect(self) -> sqlite3.Connection:
conn = super()._connect()
conn.execute("PRAGMA busy_timeout=5000")
return conn
def get_teardown_timeout_seconds(self) -> float | None:
"""Return the configured teardown timeout, or None if not set."""
try:
with self._connection() as conn:
row = conn.execute(
"SELECT teardown_timeout_seconds FROM orchestrator_config WHERE id = 1"
).fetchone()
except sqlite3.OperationalError:
return None
return row["teardown_timeout_seconds"] if row else None
def set_teardown_timeout_seconds(self, value: float) -> None:
"""Persist the teardown timeout."""
with self._connection() as conn:
conn.execute(
"INSERT OR REPLACE INTO orchestrator_config"
" (id, teardown_timeout_seconds) VALUES (1, ?)",
(value,),
)
self._chmod()
def delete_teardown_timeout_seconds(self) -> bool:
"""Clear the stored teardown timeout. Returns True if a value existed."""
with self._connection() as conn:
cur = conn.execute(
"UPDATE orchestrator_config SET teardown_timeout_seconds = NULL"
" WHERE id = 1 AND teardown_timeout_seconds IS NOT NULL"
)
return cur.rowcount > 0
def resolve_teardown_timeout(db_path: Path | None = None) -> float:
"""Return the teardown timeout to use, in priority order:
1. ``BOT_BOTTLE_ORCHESTRATOR_TEARDOWN_TIMEOUT_SECONDS`` env var
2. ``teardown_timeout_seconds`` in the orchestrator config DB
3. ``DEFAULT_TEARDOWN_TIMEOUT_SECONDS`` (30 s)
"""
raw = os.environ.get(TEARDOWN_TIMEOUT_ENV, "").strip()
if raw:
try:
value = float(raw)
if value > 0:
return value
except ValueError:
pass
store = OrchestratorConfigStore(db_path)
if not store.is_migrated():
store.migrate()
db_value = store.get_teardown_timeout_seconds()
if db_value is not None and db_value > 0:
return db_value
return DEFAULT_TEARDOWN_TIMEOUT_SECONDS
__all__ = [
"OrchestratorConfigStore",
"resolve_teardown_timeout",
"TEARDOWN_TIMEOUT_ENV",
"DEFAULT_TEARDOWN_TIMEOUT_SECONDS",
]
-24
View File
@@ -13,9 +13,6 @@ vsock / unix-socket portability caveats):
PUT /bottles/<bottle_id>/policy -> 200 {"updated": true} | 404 (live reload)
body: {"policy"}
DELETE /bottles/<bottle_id> -> 200 {"torn_down": true} | 404 (teardown)
POST /reconcile -> 200 {"reaped": [bottle_id, ...]}
body: {"live_source_ips": [...],
["grace_seconds"]}
POST /attribute -> 200 {"bottle_id"} | 403
POST /resolve -> 200 {"bottle_id","policy"} | 403
body: {"source_ip","identity_token"}
@@ -144,27 +141,6 @@ def dispatch( # pylint: disable=too-many-return-statements,too-many-branches
return 200, {"torn_down": True}
return 404, {"error": "no such bottle"}
if method == "POST" and route == "/reconcile":
# Host-driven self-heal: the caller enumerates its live bottles (only
# the host can see the backend) and the orchestrator drops rows for
# every other active bottle. Trusted-caller only — an agent that could
# reach this would be able to unregister its neighbours.
try:
data = _parse_json_object(body)
except ValueError as e:
return 400, {"error": f"invalid JSON: {e}"}
raw_ips = data.get("live_source_ips")
if not isinstance(raw_ips, list):
return 400, {"error": "live_source_ips (list of strings) is required"}
live = [ip for ip in raw_ips if isinstance(ip, str) and ip]
grace = data.get("grace_seconds")
kwargs = (
{"grace_seconds": float(grace)}
if isinstance(grace, (int, float)) and not isinstance(grace, bool)
else {}
)
return 200, {"reaped": orch.reconcile(live, **kwargs)}
if method == "POST" and route == "/attribute":
try:
data = _parse_json_object(body)
+10 -41
View File
@@ -27,7 +27,6 @@ from ..paths import (
CONTROL_PLANE_TOKEN_ENV,
host_control_plane_token,
host_db_path,
host_gateway_ca_dir,
)
from ..supervise import DB_PATH_IN_CONTAINER
@@ -49,23 +48,14 @@ GATEWAY_LABEL = "bot-bottle-orch-gateway=1"
# the source IP the gateway attributes by is the address on this network.
GATEWAY_NETWORK = "bot-bottle-gateway"
# mitmproxy's CA dir in the bundle. The host's gateway-CA dir (see
# `host_gateway_ca_dir`) is bind-mounted here so the gateway's self-generated
# CA stays STABLE across container recreation — every agent installs this one
# CA to trust the shared gateway's TLS interception, so it must not rotate when
# the gateway restarts. A host bind-mount rather than a named volume: a named
# volume is silently wiped by `docker volume prune`, minting a fresh CA that
# breaks every running bottle (issue #450).
# mitmproxy's CA dir in the bundle. A persistent named volume here keeps the
# gateway's self-generated CA STABLE across container recreation — every agent
# installs this one CA to trust the shared gateway's TLS interception, so it
# must not rotate when the gateway restarts.
MITMPROXY_HOME = "/home/mitmproxy/.mitmproxy"
GATEWAY_CA_VOLUME = "bot-bottle-gateway-mitmproxy"
GATEWAY_CA_CERT = f"{MITMPROXY_HOME}/mitmproxy-ca-cert.pem"
# The CA material mitmproxy writes into its confdir. mitmproxy reuses these on
# startup when present and generates them only on first run, so persisting them
# is what makes the CA stable; deleting them (see `rotate_gateway_ca`) forces a
# fresh CA on the next start. `mitmproxy-ca.pem` (cert + private key) is the
# signing identity; the rest are derived encodings agents/clients consume.
GATEWAY_CA_GLOB = "mitmproxy-ca*"
# The gateway data-plane image + its Dockerfile. Kept as a local constant
# rather than imported from the backend layer, which would drag
# the whole backend layer into the lean orchestrator (see #359); unify when
@@ -83,26 +73,6 @@ def _host_db_dir() -> str:
return str(db_dir)
def rotate_gateway_ca(ca_dir: Path | None = None) -> list[Path]:
"""Delete the persisted mitmproxy CA so the next gateway start mints a
fresh one the explicit, deliberate CA-rollover path (issue #450).
Persistence keeps the CA stable across restarts precisely because mitmproxy
reuses the on-disk CA; rotation is therefore just removing that material.
Returns the files removed (empty when there was no CA yet); idempotent.
This only clears the on-disk CA. It does NOT stop the running gateway (whose
mitmproxy still holds the old CA in memory) or re-provision agents the
caller recreates the gateway to mint the new CA and re-attaches bottles.
`rotate-ca` on the orchestrator CLI wires those steps together."""
ca_dir = ca_dir if ca_dir is not None else host_gateway_ca_dir()
removed: list[Path] = []
for path in sorted(ca_dir.glob(GATEWAY_CA_GLOB)):
path.unlink()
removed.append(path)
return removed
class GatewayError(Exception):
"""The shared gateway failed to build/start/stop (non-zero `docker` exit)."""
@@ -251,10 +221,9 @@ class DockerGateway(Gateway):
"--name", self.name,
"--label", GATEWAY_LABEL,
"--network", self.network,
# Persist the self-generated CA on the host so it survives both
# container recreation AND docker volume pruning (agents trust it)
# — see host_gateway_ca_dir / issue #450.
"--volume", f"{host_gateway_ca_dir()}:{MITMPROXY_HOME}",
# Persist the self-generated CA so it survives restarts (agents
# trust it) — see GATEWAY_CA_VOLUME.
"--volume", f"{GATEWAY_CA_VOLUME}:{MITMPROXY_HOME}",
# Share the one host DB: the supervise daemon queues proposals
# into the same file the orchestrator (and the operator, over
# HTTP) reads — no second, disconnected DB in the container.
@@ -303,7 +272,7 @@ class DockerGateway(Gateway):
__all__ = [
"Gateway", "DockerGateway", "GatewayError", "rotate_gateway_ca",
"Gateway", "DockerGateway", "GatewayError",
"GATEWAY_NAME", "GATEWAY_LABEL", "GATEWAY_IMAGE", "GATEWAY_NETWORK",
"GATEWAY_CA_CERT", "GATEWAY_CA_GLOB",
"GATEWAY_CA_VOLUME", "GATEWAY_CA_CERT",
]
+164 -159
View File
@@ -1,15 +1,17 @@
"""Orchestrator + gateway lifecycle (PRD 0070, docker slice).
Runs both the orchestrator control plane and the gateway data plane inside
a single `bot-bottle-infra` container on the shared gateway network
matching the structure already used by the macOS and Firecracker backends.
`gateway_init` is PID 1 and supervises both; the infra container is an
idempotent per-host singleton.
Runs the orchestrator control plane **as a container** on the shared gateway
network, alongside the gateway container. This is the PRD's "virtualize the
orchestrator": container↔container between the gateway and the orchestrator
avoids the host firewall (which drops containerhost traffic), and the gateway
reaches the control plane by container name over docker DNS. The host CLI
reaches it via a published loopback port.
The combined container replaces the prior two-container split
(bot-bottle-orchestrator + bot-bottle-orch-gateway). The host CLI reaches
the control plane via a published loopback port; gateway daemons reach it
over 127.0.0.1 (same container).
The orchestrator runs with the **register-only broker** the *backend*
launches agent containers (compose), so the orchestrator needs no docker
socket. That keeps this control-plane container unprivileged; the host manages
both containers. `ensure_running` is an idempotent singleton (fixed container
names + the published port).
"""
from __future__ import annotations
@@ -23,75 +25,50 @@ from pathlib import Path
from .. import log
from ..docker_cmd import run_docker
from ..paths import (
CONTROL_PLANE_TOKEN_ENV,
bot_bottle_root,
host_control_plane_token,
host_gateway_ca_dir,
)
from ..supervise import DB_PATH_IN_CONTAINER
from .gateway import (
GATEWAY_DOCKERFILE,
GATEWAY_IMAGE,
GATEWAY_NETWORK,
GatewayError,
MITMPROXY_HOME,
_host_db_dir,
)
from ..paths import CONTROL_PLANE_TOKEN_ENV, bot_bottle_root, host_control_plane_token
from .gateway import GATEWAY_IMAGE, GATEWAY_NAME, GATEWAY_NETWORK, DockerGateway, GatewayError
DEFAULT_PORT = 8099
DEFAULT_STARTUP_TIMEOUT_SECONDS = 45.0
INFRA_NAME = "bot-bottle-infra"
INFRA_LABEL = "bot-bottle-infra=1"
# The combined infra image: gateway data plane + orchestrator content.
# Built from Dockerfile.infra (FROM gateway + COPY --from orchestrator).
INFRA_IMAGE = os.environ.get("BOT_BOTTLE_INFRA_IMAGE", "bot-bottle-infra:latest")
INFRA_DOCKERFILE = "Dockerfile.infra"
# Baked as a container label so `ensure_running` can detect whether the
# running container is executing the current bind-mounted source.
INFRA_SOURCE_HASH_LABEL = "bot-bottle-infra-source-hash"
# Orchestrator image: the single canonical definition of the control-plane
# content (lean: python:3.12-slim + bot_bottle package, no mitmproxy/git).
# Used as a build intermediate: `Dockerfile.infra` COPY --from this image.
ORCHESTRATOR_NAME = "bot-bottle-orchestrator"
ORCHESTRATOR_LABEL = "bot-bottle-orchestrator=1"
# The control-plane's own runtime image — lean (python + the stdlib-only
# `bot_bottle` package, bind-mounted at run time), distinct from the heavy
# gateway data-plane image it used to borrow (#384). Env override for
# operators pinning a published build.
ORCHESTRATOR_IMAGE = os.environ.get(
"BOT_BOTTLE_ORCHESTRATOR_IMAGE", "bot-bottle-orchestrator:latest"
)
ORCHESTRATOR_DOCKERFILE = "Dockerfile.orchestrator"
# Baked onto the container as a label so `ensure_running` can tell whether the
# running process is executing the *current* bind-mounted source — see
# `source_hash`.
ORCHESTRATOR_SOURCE_HASH_LABEL = "bot-bottle-orchestrator-source-hash"
# The gateway daemons + orchestrator the infra container runs.
# BOT_BOTTLE_GATEWAY_DAEMONS listing `orchestrator` opts it in to
# gateway_init's supervise tree (see gateway_init._OPT_IN_DAEMONS).
_INFRA_DAEMONS = "egress,git-http,supervise,orchestrator"
# The bind-mount path for the live control-plane source inside the
# container. Separate from /app so the gateway's baked scripts
# (egress_addon.py, egress-entrypoint.sh) are not overlaid.
_SRC_IN_CONTAINER = "/bot-bottle-src"
# Bot-bottle host-root bind-mount inside the container (DB + state).
# The repo root is bind-mounted into the control-plane container so
# `python -m bot_bottle.orchestrator` resolves the package (the orchestrator
# is stdlib-only, so the lean orchestrator image's python is enough).
_REPO_ROOT = Path(__file__).resolve().parents[2]
_APP_DIR = "/app"
_ROOT_IN_CONTAINER = "/bot-bottle-root"
# The supervise daemon writes proposals into the host DB directory.
_SUPERVISE_DB_DIR_IN_CONTAINER = os.path.dirname(DB_PATH_IN_CONTAINER)
_HEALTH_POLL_SECONDS = 0.25
DEFAULT_STARTUP_TIMEOUT_SECONDS = 45.0
_HEALTH_REQUEST_TIMEOUT_SECONDS = 1.0
_REPO_ROOT = Path(__file__).resolve().parents[2]
class OrchestratorStartError(RuntimeError):
"""The infra container did not become healthy within the timeout."""
"""The orchestrator container did not become healthy within the timeout."""
def source_hash(repo_root: Path) -> str:
"""Content hash of the orchestrator's bind-mounted Python source (the
`bot_bottle` package the control-plane process imports). Changes only
when the code that would actually run changes `ensure_running`
recreates the container on a mismatch so a code change takes effect,
but leaves a healthy up-to-date container alone to preserve in-memory
egress tokens."""
`bot_bottle` package the control-plane process imports). This only
changes when the code that would actually run inside the container
changes `ensure_running` recreates the container on a mismatch and
otherwise leaves a healthy one alone, so a bottle launch that isn't
accompanied by a code change doesn't restart the process and drop every
*other* active bottle's in-memory egress tokens (`Orchestrator._tokens`
in `service.py`, never persisted to disk by design)."""
h = hashlib.sha256()
for path in sorted((repo_root / "bot_bottle").rglob("*.py")):
h.update(str(path.relative_to(repo_root)).encode())
@@ -100,37 +77,57 @@ def source_hash(repo_root: Path) -> str:
class OrchestratorService:
"""Manages the single per-host infra container (control plane + gateway).
"""Manages the orchestrator control-plane container + the shared gateway.
Callers only need `ensure_running()` + `url`.
`infra_name` / `infra_label` let backends run independent infra containers
on the same host without name collisions (e.g. isolated integration tests
that can't share the production INFRA_NAME singleton)."""
`orchestrator_name` / `orchestrator_label` let backends run independent
orchestrators on the same host without name collisions (e.g. the
Firecracker backend uses `bot-bottle-fc-orchestrator` alongside the Docker
backend's `bot-bottle-orchestrator`); `gateway_name` gives the paired
gateway container the same treatment (e.g. isolated integration tests
that can't share the production `GATEWAY_NAME` singleton). Subclass and
override `_gateway()` for anything `_gateway_image`/`gateway_name` can't
express (a genuinely backend-specific gateway variant)."""
def __init__(
self,
*,
port: int = DEFAULT_PORT,
network: str = GATEWAY_NETWORK,
image: str = INFRA_IMAGE,
image: str = ORCHESTRATOR_IMAGE,
gateway_image: str = GATEWAY_IMAGE,
gateway_name: str = GATEWAY_NAME,
repo_root: Path = _REPO_ROOT,
host_root: Path | None = None,
infra_name: str = INFRA_NAME,
infra_label: str = INFRA_LABEL,
orchestrator_name: str = ORCHESTRATOR_NAME,
orchestrator_label: str = ORCHESTRATOR_LABEL,
) -> None:
self.port = port
self.network = network
# Two distinct images (#384): `image` is the lean control-plane
# runtime this container runs; `_gateway_image` is the heavy egress /
# git-gate / supervise data plane the gateway container runs. They
# were one conflated image before the split.
self.image = image
self._gateway_image = gateway_image
self._gateway_name = gateway_name
self._repo_root = repo_root
self._host_root = host_root or bot_bottle_root()
self._infra_name = infra_name
self._infra_label = infra_label
self._orchestrator_name = orchestrator_name
self._orchestrator_label = orchestrator_label
@property
def url(self) -> str:
"""Host-side control-plane URL (published loopback port)."""
return f"http://127.0.0.1:{self.port}"
@property
def internal_url(self) -> str:
"""Control-plane URL as the gateway container reaches it — by name over
docker DNS on the shared network. This is the gateway's
BOT_BOTTLE_ORCHESTRATOR_URL."""
return f"http://{self._orchestrator_name}:{self.port}"
def is_healthy(self, *, timeout: float = _HEALTH_REQUEST_TIMEOUT_SECONDS) -> bool:
try:
with urllib.request.urlopen(f"{self.url}/health", timeout=timeout) as resp:
@@ -142,131 +139,139 @@ class OrchestratorService:
proc = run_docker(["docker", "ps", "--filter", f"name=^/{name}$", "--format", "{{.Names}}"])
return name in proc.stdout.split()
def _infra_source_current(self, current_hash: str) -> bool:
"""True iff the running infra container was started from the current
bind-mounted source. Mirrors the macOS backend's `_source_current`."""
if not self._container_running(self._infra_name):
return False
proc = run_docker([
"docker", "inspect", "--format",
"{{ index .Config.Labels \"" + INFRA_SOURCE_HASH_LABEL + "\" }}",
self._infra_name,
])
if proc.returncode != 0:
return True # can't compare → don't churn a working container
return proc.stdout.strip() == current_hash
def _ensure_network(self) -> None:
if run_docker(["docker", "network", "inspect", self.network]).returncode == 0:
return
proc = run_docker(["docker", "network", "create", self.network])
if proc.returncode != 0 and "already exists" not in proc.stderr:
raise GatewayError(
f"gateway network {self.network} failed to create: {proc.stderr.strip()}"
)
def _build_images(self) -> None:
"""Build the gateway base, the orchestrator intermediate, then the
infra image. All are cache-aware: a no-op when nothing changed."""
for tag, dockerfile in (
(GATEWAY_IMAGE, GATEWAY_DOCKERFILE),
(ORCHESTRATOR_IMAGE, ORCHESTRATOR_DOCKERFILE),
(self.image, INFRA_DOCKERFILE),
):
argv = ["docker", "build", "-t", tag,
"-f", str(self._repo_root / dockerfile),
str(self._repo_root)]
if os.environ.get("BOT_BOTTLE_NO_CACHE"):
argv.insert(2, "--no-cache")
proc = run_docker(argv)
if proc.returncode != 0:
raise GatewayError(f"{dockerfile} build failed: {proc.stderr.strip()}")
def _run_infra_container(self, current_hash: str) -> None:
"""Start the combined infra container (idempotent: clears a stale
fixed-name container first). Labels the container with `current_hash`
so a later `ensure_running` can detect a real code change."""
self._ensure_network()
run_docker(["docker", "rm", "--force", self._infra_name])
def _run_orchestrator_container(self, current_hash: str) -> None:
"""Start the control-plane container (idempotent: clears a stale
fixed-name container first). Register-only broker no docker socket.
Labels the container with `current_hash` so a later `ensure_running`
can detect a real code change (see `source_hash`)."""
run_docker(["docker", "rm", "--force", self._orchestrator_name])
proc = run_docker([
"docker", "run", "--detach",
"--name", self._infra_name,
"--label", self._infra_label,
"--label", f"{INFRA_SOURCE_HASH_LABEL}={current_hash}",
"--name", self._orchestrator_name,
"--label", self._orchestrator_label,
"--label", f"{ORCHESTRATOR_SOURCE_HASH_LABEL}={current_hash}",
"--network", self.network,
# Host CLI reaches the control plane here (loopback only).
# gateway_init always starts the orchestrator on DEFAULT_PORT (8099)
# inside the container; self.port is the host-side published port.
"--publish", f"127.0.0.1:{self.port}:{DEFAULT_PORT}",
# Persist the mitmproxy CA on the host so it survives container
# recreation AND docker volume pruning (issue #450): every agent
# trusts this one CA, so a fresh one would break all running bottles.
"--volume", f"{host_gateway_ca_dir()}:{MITMPROXY_HOME}",
# Shared supervise DB (same file the operator reads over HTTP).
"--volume", f"{_host_db_dir()}:{_SUPERVISE_DB_DIR_IN_CONTAINER}",
"--env", f"SUPERVISE_DB_PATH={DB_PATH_IN_CONTAINER}",
# Live control-plane source, mounted to a path that does not
# overlay the gateway's baked /app scripts.
"--volume", f"{self._repo_root}:{_SRC_IN_CONTAINER}:ro",
# PYTHONPATH lets the orchestrator (and other Python daemons)
# import the live source ahead of the installed package.
"--env", f"PYTHONPATH={_SRC_IN_CONTAINER}",
# Orchestrator registry DB on the host (sole writer: control plane).
# Host CLI reaches the control plane here; bound to loopback so it
# is not exposed on the host's external interfaces. NOTE: the
# container is still on `self.network` (the shared gateway network),
# so agents can reach it by container IP — which is exactly why the
# control plane requires the secret below rather than trusting the
# network boundary.
"--publish", f"127.0.0.1:{self.port}:{self.port}",
"--volume", f"{self._repo_root}:{_APP_DIR}:ro",
"--workdir", _APP_DIR,
# Persist the registry DB on the host (sole-owner: only the
# orchestrator opens bot-bottle.db).
"--volume", f"{self._host_root}:{_ROOT_IN_CONTAINER}",
"--env", f"BOT_BOTTLE_ROOT={_ROOT_IN_CONTAINER}",
# Control-plane secret: required by the orchestrator (to enforce)
# and by the gateway daemons (to present on /resolve calls).
# The control-plane secret it requires on every route but /health.
# Bare `--env NAME` → docker inherits the value from the run env
# below, so the secret never lands on argv / `docker inspect`.
"--env", CONTROL_PLANE_TOKEN_ENV,
# Gateway daemons reach the orchestrator over loopback at its
# fixed internal port (DEFAULT_PORT), independent of self.port.
"--env", f"BOT_BOTTLE_ORCHESTRATOR_URL=http://127.0.0.1:{DEFAULT_PORT}",
# Opt the orchestrator into gateway_init's supervise tree.
"--env", f"BOT_BOTTLE_GATEWAY_DAEMONS={_INFRA_DAEMONS}",
"--entrypoint", "python3",
self.image,
"-m", "bot_bottle.orchestrator",
"--host", "0.0.0.0", "--port", str(self.port), "--broker", "stub",
], env={**os.environ, CONTROL_PLANE_TOKEN_ENV: host_control_plane_token()})
if proc.returncode != 0:
raise OrchestratorStartError(
f"infra container failed to start: {proc.stderr.strip()}"
f"orchestrator container failed to start: {proc.stderr.strip()}"
)
def _gateway(self) -> DockerGateway:
return DockerGateway(
self._gateway_image,
name=self._gateway_name,
network=self.network,
orchestrator_url=self.internal_url,
)
def _ensure_orchestrator_image(self) -> None:
"""Build the lean control-plane image from `Dockerfile.orchestrator`
when it's missing (#384). Cheap — a `FROM python:*-slim` base with no
deps to install, so the layer cache makes rebuilds a no-op. Unlike the
gateway image this is build-if-missing, not build-every-time: the
control plane bind-mounts its source, so a code change is caught by the
source-hash recreate (below), not by an image rebuild."""
if run_docker(["docker", "image", "inspect", self.image]).returncode == 0:
return
argv = ["docker", "build", "-t", self.image,
"-f", str(self._repo_root / ORCHESTRATOR_DOCKERFILE),
str(self._repo_root)]
if os.environ.get("BOT_BOTTLE_NO_CACHE"):
argv.insert(2, "--no-cache")
proc = run_docker(argv)
if proc.returncode != 0:
raise GatewayError(
f"orchestrator image build failed: {proc.stderr.strip()}"
)
def _orchestrator_source_current(self, current_hash: str) -> bool:
"""True iff the running orchestrator container was created from the
*current* bind-mounted source. Mirrors `DockerGateway`'s
image-staleness check, but by content hash rather than image id since
the orchestrator runs bind-mounted source, not a built image."""
if not self._container_running(self._orchestrator_name):
return False
proc = run_docker([
"docker", "inspect", "--format",
"{{ index .Config.Labels \"" + ORCHESTRATOR_SOURCE_HASH_LABEL + "\" }}",
self._orchestrator_name,
])
if proc.returncode != 0:
return True # can't compare -> don't churn a working container
return proc.stdout.strip() == current_hash
def ensure_running(
self, *, startup_timeout: float = DEFAULT_STARTUP_TIMEOUT_SECONDS,
) -> str:
"""Ensure the infra container (control plane + gateway) is up; return
the host control-plane URL. Idempotent a healthy container on current
source is left untouched. Raises `OrchestratorStartError` on timeout."""
self._build_images()
"""Ensure the control plane + shared gateway are up; return the host
control-plane URL. Idempotent a healthy control plane running
current code and a running gateway are left untouched. Raises
`OrchestratorStartError` on timeout."""
gateway = self._gateway()
gateway.ensure_built() # rebuild the bundle image on a source change
gateway.ensure_running() # creates the shared network + (re)starts gateway
# Recreate the orchestrator container only when its bind-mounted
# source has actually changed since it started — its Python process
# loaded that code at startup and won't reload, so a stale container
# would keep running OLD control-plane code. Recreating on *every*
# launch (the prior behaviour) would drop every other active
# bottle's in-memory egress tokens each time a new bottle starts,
# since the orchestrator process holds them only in memory (#381).
current_hash = source_hash(self._repo_root)
if self.is_healthy() and self._infra_source_current(current_hash):
if self.is_healthy() and self._orchestrator_source_current(current_hash):
return self.url
log.info("starting infra container", context={"name": self._infra_name})
self._run_infra_container(current_hash)
self._ensure_orchestrator_image()
log.info(
"starting orchestrator container",
context={"name": self._orchestrator_name},
)
self._run_orchestrator_container(current_hash)
deadline = time.monotonic() + startup_timeout
while time.monotonic() < deadline:
if self.is_healthy():
log.info("infra container healthy", context={"url": self.url})
log.info("orchestrator healthy", context={"url": self.url})
return self.url
time.sleep(_HEALTH_POLL_SECONDS)
raise OrchestratorStartError(
f"infra container at {self.url} did not become healthy within {startup_timeout:g}s"
f"orchestrator at {self.url} did not become healthy within {startup_timeout:g}s"
)
def stop(self) -> None:
"""Remove the infra container (idempotent)."""
run_docker(["docker", "rm", "--force", self._infra_name])
"""Remove the orchestrator + gateway containers (idempotent)."""
run_docker(["docker", "rm", "--force", self._orchestrator_name])
self._gateway().stop()
__all__ = [
"OrchestratorService",
"OrchestratorStartError",
"INFRA_NAME",
"INFRA_IMAGE",
"INFRA_SOURCE_HASH_LABEL",
"ORCHESTRATOR_NAME",
"ORCHESTRATOR_IMAGE",
"DEFAULT_PORT",
"DEFAULT_STARTUP_TIMEOUT_SECONDS",
"source_hash",
]
+6 -78
View File
@@ -32,7 +32,6 @@ import hmac
import secrets
import sqlite3
import time
from collections.abc import Iterable
from dataclasses import dataclass
from pathlib import Path
@@ -43,12 +42,6 @@ from ..paths import host_db_path
# 256 bits of urandom, URL-safe — unguessable per-bottle identity token.
IDENTITY_TOKEN_BYTES = 32
# How recently a row must have been registered to be exempt from
# `reap_absent`. Covers the window between `container run` and the address
# becoming visible to another launch's enumeration, so reconciliation never
# reaps a bottle that is still coming up.
DEFAULT_REAP_GRACE_SECONDS = 120.0
def new_identity_token() -> str:
"""A fresh per-bottle identity token (PRD 0070 attribution defence)."""
@@ -174,7 +167,7 @@ class RegistryStore(DbStore):
metadata=metadata,
policy=policy,
)
with self._connection() as conn:
with self._connect() as conn:
conn.execute(
"DELETE FROM orchestrator_bottles "
"WHERE source_ip = ? AND state = 'active' AND bottle_id != ?",
@@ -200,7 +193,7 @@ class RegistryStore(DbStore):
def set_policy(self, bottle_id: str, policy: str) -> bool:
"""Update a bottle's policy in place (live reload). Returns True if
the bottle exists."""
with self._connection() as conn:
with self._connect() as conn:
cur = conn.execute(
"UPDATE orchestrator_bottles SET policy = ? WHERE bottle_id = ?",
(policy, bottle_id),
@@ -210,7 +203,7 @@ class RegistryStore(DbStore):
def deregister(self, bottle_id: str) -> bool:
"""Remove a bottle. Returns True if a row was deleted."""
with self._connection() as conn:
with self._connect() as conn:
cur = conn.execute(
"DELETE FROM orchestrator_bottles WHERE bottle_id = ?", (bottle_id,)
)
@@ -218,7 +211,7 @@ class RegistryStore(DbStore):
def get(self, bottle_id: str) -> BottleRecord | None:
"""Return the bottle by id, or None if absent."""
with self._connection() as conn:
with self._connect() as conn:
row = conn.execute(
"SELECT * FROM orchestrator_bottles WHERE bottle_id = ?", (bottle_id,)
).fetchone()
@@ -226,76 +219,12 @@ class RegistryStore(DbStore):
def all(self) -> list[BottleRecord]:
"""Every registered bottle, oldest first."""
with self._connection() as conn:
with self._connect() as conn:
rows = conn.execute(
"SELECT * FROM orchestrator_bottles ORDER BY created_at"
).fetchall()
return [_row_to_record(r) for r in rows]
def reap_absent(
self,
live_source_ips: Iterable[str],
*,
grace_seconds: float = DEFAULT_REAP_GRACE_SECONDS,
now: float | None = None,
) -> list[BottleRecord]:
"""Delete active rows whose source IP is not held by a live bottle.
A row only ever leaves the registry two ways: an explicit
`teardown_bottle` (the launcher's cleanup callback) or the supersede
sweep in `register`. Neither runs when the launching CLI dies hard
SIGKILL, a closed terminal, a host sleep/crash so the row outlives
its container. That orphan is not inert: source IPs are recycled by
the backend's DHCP, and `by_source_ip` fail-closes on ambiguity, so a
leftover row at a reused address can brick the *next* bottle that
lands on it (no policy resolved -> every host denied, reported to the
agent as "not in the allowlist"). Reconciling against the live set at
launch keeps the registry from accumulating those landmines.
Restores the invariant the data plane needs: **at most one active row
per live address, and none at all for a dead one.** Two cases, because
a dead bottle's address may already have been handed to a live one:
* no live bottle holds the address every row there is an orphan;
* a live bottle holds it but several rows claim it the newest
registration is authoritative and the rest are orphans, the same
rule `register`'s same-IP supersede sweep applies. Without this
second case a recycled address stays ambiguous, which is exactly
the state that resolves no policy.
`grace_seconds` protects an in-flight launch: registration happens
moments after `container run`, and a concurrent launch's address may
not be visible to the caller's enumeration yet. Rows younger than the
grace window are never reaped, so reconciliation can't race a bottle
that is still coming up. Returns the deleted records."""
live = {ip for ip in live_source_ips if ip}
cutoff = (time.time() if now is None else now) - grace_seconds
with self._connection() as conn:
rows = conn.execute(
"SELECT * FROM orchestrator_bottles WHERE state = 'active'",
).fetchall()
by_ip: dict[str, list[BottleRecord]] = {}
for row in rows:
rec = _row_to_record(row)
by_ip.setdefault(rec.source_ip, []).append(rec)
candidates: list[BottleRecord] = []
for ip, recs in by_ip.items():
if ip not in live:
candidates.extend(recs)
continue
# Keep the newest claim on a live address; supersede the rest.
recs.sort(key=lambda r: r.created_at)
candidates.extend(recs[:-1])
doomed = [r for r in candidates if r.created_at <= cutoff]
for rec in doomed:
conn.execute(
"DELETE FROM orchestrator_bottles WHERE bottle_id = ?",
(rec.bottle_id,),
)
if doomed:
self._chmod()
return doomed
def by_source_ip(self, source_ip: str) -> BottleRecord | None:
"""Network-layer attribution: the single active bottle at this source
IP, or None if unknown or ambiguous (more than one a
@@ -303,7 +232,7 @@ class RegistryStore(DbStore):
source IP is unspoofable (Firecracker `/31` + nft) and the control
plane is reachable only by the trusted gateway; pair with the
identity token (`attribute`) elsewhere."""
with self._connection() as conn:
with self._connect() as conn:
rows = conn.execute(
"SELECT * FROM orchestrator_bottles "
"WHERE source_ip = ? AND state = 'active'",
@@ -333,5 +262,4 @@ __all__ = [
"new_identity_token",
"default_db_path",
"IDENTITY_TOKEN_BYTES",
"DEFAULT_REAP_GRACE_SECONDS",
]
-62
View File
@@ -1,62 +0,0 @@
"""Rotate the shared gateway's mitmproxy CA (issue #450).
python -m bot_bottle.orchestrator.rotate_ca
A deliberate CA rollover has two halves: drop the *persisted* CA so a fresh one
is minted, and drop the *running* gateway so its mitmproxy (which holds the old
CA in memory) is replaced. This one-shot command does both:
1. Delete the persisted CA under the host gateway-CA dir the next gateway
start generates a new one (mitmproxy reuses an existing CA, generates only
when absent).
2. Force-remove the infra / standalone-gateway containers so the stale
in-memory CA is gone; the next bottle launch's idempotent `ensure_running`
brings the gateway back up and mints the fresh CA.
It does NOT re-provision the new CA into already-running bottles those must be
re-attached so they install the new trust anchor. Rotation is thus an explicit,
operator-driven action with a brief egress interruption, not an automatic one.
"""
from __future__ import annotations
import sys
from pathlib import Path
from ..docker_cmd import run_docker
from ..paths import host_gateway_ca_dir
from .gateway import GATEWAY_NAME, rotate_gateway_ca
from .lifecycle import INFRA_NAME
# The containers whose mitmproxy would still be serving the old CA from memory:
# the consolidated infra container and the standalone per-host gateway.
_GATEWAY_CONTAINERS = (INFRA_NAME, GATEWAY_NAME)
def _out(msg: str) -> None:
sys.stdout.write(f"rotate-ca: {msg}\n")
def main(argv: list[str] | None = None) -> int:
del argv # no flags — a single deliberate action
ca_dir: Path = host_gateway_ca_dir()
removed = rotate_gateway_ca(ca_dir)
if removed:
_out(f"removed {len(removed)} CA file(s) from {ca_dir}")
else:
_out(f"no persisted CA under {ca_dir}; a fresh one is minted on next start")
# Drop any running gateway so its in-memory (now-stale) CA is replaced on
# the next launch. `rm --force` on an absent name is a tolerated no-op.
for name in _GATEWAY_CONTAINERS:
proc = run_docker(["docker", "rm", "--force", name])
if proc.returncode == 0 and proc.stdout.strip():
_out(f"removed running container {name}")
_out("done — the next bottle launch remints the CA; re-attach bottles to "
"install the new trust anchor")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1 -30
View File
@@ -13,20 +13,15 @@ Launch lifecycle:
and returns the record. If the broker rejects/fails, the registry entry
is rolled back so a failed launch leaves no orphan.
* `teardown_bottle` sends a signed teardown request, then deregisters.
* `reconcile` sweeps rows whose bottle is no longer running the
self-heal for the teardown paths that never got to run (a hard-killed
launcher), since an orphan row at a recycled source IP bricks the next
bottle that lands on it.
"""
from __future__ import annotations
import json
from collections.abc import Iterable
from datetime import datetime, timezone
from .broker import LaunchBroker, LaunchRequest, sign_request
from .registry import DEFAULT_REAP_GRACE_SECONDS, BottleRecord, RegistryStore
from .registry import BottleRecord, RegistryStore
from .gateway import Gateway
from ..supervise import (
AuditEntry,
@@ -122,30 +117,6 @@ class Orchestrator:
self._tokens.pop(bottle_id, None)
return True
def reconcile(
self,
live_source_ips: Iterable[str],
*,
grace_seconds: float = DEFAULT_REAP_GRACE_SECONDS,
) -> list[str]:
"""Drop registry rows for bottles that are no longer running, and
forget their in-memory egress tokens. Returns the reaped bottle ids.
The caller supplies the live set because only the host can enumerate
its own containers the orchestrator runs *inside* the infra
container and has no view of the backend. Deliberately does not
broker a teardown: the container is already gone, so there is nothing
to stop, and a broker error must not stop the sweep from clearing
the row that would otherwise brick the next bottle at that address.
See `RegistryStore.reap_absent` for why orphans accumulate and why
they are harmful rather than merely untidy."""
reaped = self.registry.reap_absent(
live_source_ips, grace_seconds=grace_seconds)
for rec in reaped:
self._tokens.pop(rec.bottle_id, None)
return [rec.bottle_id for rec in reaped]
def tokens_for(self, bottle_id: str) -> dict[str, str]:
"""The bottle's in-memory egress auth tokens (env_name -> value), or
empty. The gateway injects these per request; they are never
-26
View File
@@ -33,13 +33,6 @@ HOST_DB_FILENAME = "bot-bottle.db"
CONTROL_PLANE_TOKEN_FILENAME = "control-plane-token"
CONTROL_PLANE_TOKEN_ENV = "BOT_BOTTLE_CONTROL_PLANE_TOKEN"
# The host directory holding the gateway's persistent mitmproxy CA. Bind-mounted
# into the infra/gateway container at mitmproxy's confdir so the self-generated
# CA survives container recreation — every agent installs this one CA to trust
# the shared gateway's TLS interception, so it must not rotate on restart. See
# host_gateway_ca_dir() for why this is a host bind-mount, not a named volume.
GATEWAY_CA_DIRNAME = "gateway-ca"
def bot_bottle_root() -> Path:
"""The app data root — `$BOT_BOTTLE_ROOT` if set, else `~/.bot-bottle`."""
@@ -66,23 +59,6 @@ def host_db_dir() -> Path:
return db_dir
def host_gateway_ca_dir() -> Path:
"""The directory holding the gateway's persistent mitmproxy CA, created if
missing. Backends bind-mount this into the infra/gateway container at
mitmproxy's confdir so the CA persists across container recreation.
A host bind-mount under the app-data root deliberately NOT a Docker
named volume. A named volume survives `docker rm` but is silently wiped by
`docker volume prune` / `docker system prune --volumes` during routine host
maintenance; the gateway then mints a fresh CA that every already-running
bottle distrusts, failing the TLS handshake even after it reconnects to the
moved gateway (issue #450). A path under the root docker never prunes it,
and it stays directly inspectable + rotatable from the host."""
ca_dir = bot_bottle_root() / GATEWAY_CA_DIRNAME
ca_dir.mkdir(parents=True, exist_ok=True)
return ca_dir
def host_control_plane_token() -> str:
"""The per-host control-plane secret, minted (256-bit, url-safe) and
persisted 0600 on first use, then reused.
@@ -118,10 +94,8 @@ __all__ = [
"HOST_DB_FILENAME",
"CONTROL_PLANE_TOKEN_FILENAME",
"CONTROL_PLANE_TOKEN_ENV",
"GATEWAY_CA_DIRNAME",
"bot_bottle_root",
"host_db_path",
"host_db_dir",
"host_gateway_ca_dir",
"host_control_plane_token",
]
+2 -1
View File
@@ -22,7 +22,8 @@ closed too rather than silently serving stale or empty policy.
The resolved value is the policy blob the orchestrator stores verbatim; the
consumer parses it (e.g. the egress addon's `load_config`). This module is
stdlib-only and free of bot-bottle imports.
stdlib-only and free of bot-bottle imports so it can be COPYed flat into
the gateway.
"""
from __future__ import annotations
+7 -7
View File
@@ -66,7 +66,7 @@ class QueueStore(DbStore):
super().__init__(resolved, migrations)
def write_proposal(self, proposal: Proposal) -> Path:
with self._connection() as conn:
with self._connect() as conn:
conn.execute(
"""
INSERT OR REPLACE INTO supervise_proposals (
@@ -89,7 +89,7 @@ class QueueStore(DbStore):
return self.db_path
def read_proposal(self, proposal_id: str) -> Proposal:
with self._connection() as conn:
with self._connect() as conn:
row = conn.execute(
"""
SELECT * FROM supervise_proposals
@@ -104,7 +104,7 @@ class QueueStore(DbStore):
def list_pending_proposals(self) -> list[Proposal]:
if not self.db_path.is_file():
return []
with self._connection() as conn:
with self._connect() as conn:
rows = conn.execute(
"""
SELECT p.* FROM supervise_proposals p
@@ -125,7 +125,7 @@ class QueueStore(DbStore):
def list_all_pending_proposals(self) -> list[Proposal]:
if not self.db_path.is_file():
return []
with self._connection() as conn:
with self._connect() as conn:
rows = conn.execute(
"""
SELECT p.* FROM supervise_proposals p
@@ -142,7 +142,7 @@ class QueueStore(DbStore):
return [self._row_to_proposal(row) for row in rows]
def write_response(self, response: Response) -> Path:
with self._connection() as conn:
with self._connect() as conn:
conn.execute(
"""
INSERT OR REPLACE INTO supervise_responses (
@@ -161,7 +161,7 @@ class QueueStore(DbStore):
return self.db_path
def read_response(self, proposal_id: str) -> Response:
with self._connection() as conn:
with self._connect() as conn:
row = conn.execute(
"""
SELECT * FROM supervise_responses
@@ -176,7 +176,7 @@ class QueueStore(DbStore):
def archive_proposal(self, proposal_id: str) -> None:
if not self.db_path.is_file():
return
with self._connection() as conn:
with self._connect() as conn:
conn.execute(
"""
UPDATE supervise_proposals SET archived = 1
-4
View File
@@ -6,11 +6,9 @@ from pathlib import Path
try:
from .audit_store import AuditStore
from .config_store import ConfigStore
from .queue_store import QueueStore
except ImportError:
from audit_store import AuditStore # type: ignore[import-not-found] # pylint: disable=import-error,no-name-in-module
from config_store import ConfigStore # type: ignore[import-not-found] # pylint: disable=import-error,no-name-in-module
from queue_store import QueueStore # type: ignore[import-not-found] # pylint: disable=import-error,no-name-in-module
_instance: StoreManager | None = None
@@ -49,13 +47,11 @@ class StoreManager:
return (
QueueStore("", self.db_path).is_migrated()
and AuditStore(self.db_path).is_migrated()
and ConfigStore(self.db_path).is_migrated()
)
def migrate(self) -> None:
QueueStore("", self.db_path).migrate()
AuditStore(self.db_path).migrate()
ConfigStore(self.db_path).migrate()
__all__ = ["StoreManager"]
+34 -18
View File
@@ -37,23 +37,40 @@ from abc import ABC
from dataclasses import dataclass
from pathlib import Path
from .supervise_types import (
ACTION_OPERATOR_EDIT,
AuditEntry,
Proposal,
Response,
STATUSES,
STATUS_APPROVED,
STATUS_MODIFIED,
STATUS_REJECTED,
TOOLS,
TOOL_CHECK_PROPOSAL,
TOOL_EGRESS_ALLOW,
TOOL_EGRESS_BLOCK,
TOOL_EGRESS_TOKEN_ALLOW,
TOOL_GITLEAKS_ALLOW,
TOOL_LIST_EGRESS_ROUTES,
)
try:
from .supervise_types import (
ACTION_OPERATOR_EDIT,
AuditEntry,
Proposal,
Response,
STATUSES,
STATUS_APPROVED,
STATUS_MODIFIED,
STATUS_REJECTED,
TOOLS,
TOOL_EGRESS_ALLOW,
TOOL_EGRESS_BLOCK,
TOOL_EGRESS_TOKEN_ALLOW,
TOOL_GITLEAKS_ALLOW,
TOOL_LIST_EGRESS_ROUTES,
)
except ImportError:
from supervise_types import ( # type: ignore[import-not-found,no-redef] # pylint: disable=import-error,no-name-in-module
ACTION_OPERATOR_EDIT,
AuditEntry,
Proposal,
Response,
STATUSES,
STATUS_APPROVED,
STATUS_MODIFIED,
STATUS_REJECTED,
TOOLS,
TOOL_EGRESS_ALLOW,
TOOL_EGRESS_BLOCK,
TOOL_EGRESS_TOKEN_ALLOW,
TOOL_GITLEAKS_ALLOW,
TOOL_LIST_EGRESS_ROUTES,
)
try:
@@ -264,7 +281,6 @@ __all__ = [
"TOOLS",
"EGRESS_FORWARD_PROXY",
"EGRESS_INTROSPECT_URL",
"TOOL_CHECK_PROPOSAL",
"TOOL_EGRESS_ALLOW",
"TOOL_EGRESS_BLOCK",
"TOOL_GITLEAKS_ALLOW",
+29 -130
View File
@@ -2,24 +2,14 @@
Per-bottle MCP server exposing tools the agent calls to propose egress
config changes when stuck. The tools are `egress-allow`,
`egress-block`, `list-egress-routes`, and `check-proposal`.
`egress-block`, and `list-egress-routes`.
Each queued proposal tool call:
Each queued tool call:
1. Validates the proposed file syntactically.
2. Writes a Proposal to the host SQLite database.
3. Blocks polling for a matching Response row, up to a short grace
window (`SUPERVISE_RESPONSE_TIMEOUT_SECONDS`, default 30s).
4. On a decision within the window, returns the operator's
`{status, notes}`. On timeout, returns `status: pending` **with the
proposal id** and leaves the proposal queued the flow is
non-blocking past the grace window (PRD prd-new / issue #412).
`check-proposal` is the non-blocking companion: given a `proposal_id`
returned by a `pending` response, it reports the current decision
(`pending` | `approved` | `modified` | `rejected`) without re-proposing,
so an approval made out-of-band (e.g. a web review console) can be resumed
without holding an HTTP request open.
3. Blocks polling for a matching Response row.
4. Returns the operator's `{status, notes}` to the agent.
One shared server fronts every bottle (PRD 0070) and attributes each
proposal to the calling bottle by source IP, resolved from the orchestrator
@@ -32,14 +22,13 @@ Speaks MCP over HTTP+JSON-RPC. Methods handled:
* `initialize` handshake; returns server info + caps.
* `notifications/initialized` ack-only.
* `tools/list` returns the tool definitions.
* `tools/call` validates, queues, waits out the grace
window, returns (pending past it); or, for
`check-proposal`, a non-blocking status poll.
* `tools/call` validates, queues, blocks, returns.
Everything else returns JSON-RPC error -32601 (method not found).
The Dockerfile copies this script to /app/supervise_server.py and installs
the bot_bottle package so its `from bot_bottle.*` imports resolve.
Stdlib-only. The Dockerfile copies this file + bot_bottle/supervise.py
into the image; the server imports `supervise` for the queue / Proposal
plumbing.
"""
from __future__ import annotations
@@ -53,18 +42,29 @@ import time
import typing
from dataclasses import dataclass, replace
from bot_bottle.constants import IDENTITY_HEADER
from bot_bottle.egress_addon_core import (
LOG_OFF, load_config, resolve_client_context, route_to_yaml_dict,
)
from bot_bottle.policy_resolver import PolicyResolveError, PolicyResolver
from bot_bottle import supervise as _sv
try:
# Same-directory imports inside the bundle container; these files are
# COPYed flat under /app by Dockerfile.gateway.
from egress_addon_core import (
LOG_OFF, load_config, resolve_client_context, route_to_yaml_dict,
)
from policy_resolver import PolicyResolveError, PolicyResolver
import supervise as _sv
except ModuleNotFoundError:
# Package imports for host-side tests and tooling.
from .egress_addon_core import (
LOG_OFF, load_config, resolve_client_context, route_to_yaml_dict,
)
from .policy_resolver import PolicyResolveError, PolicyResolver
from . import supervise as _sv
# --- JSON-RPC / MCP plumbing ----------------------------------------------
MCP_PROTOCOL_VERSION = "2024-11-05"
# App-layer identity token header (mirrors egress_addon / git_http_backend).
IDENTITY_HEADER = "x-bot-bottle-identity"
SERVER_NAME = "bot-bottle-supervise"
SERVER_VERSION = "0.1.0"
@@ -244,31 +244,6 @@ TOOL_DEFINITIONS: list[dict[str, object]] = [
),
"inputSchema": _proposal_input_schema(),
},
{
"name": _sv.TOOL_CHECK_PROPOSAL,
"description": (
"Poll a previously queued proposal for the operator's decision "
"WITHOUT blocking or re-proposing. Pass the `proposal_id` you "
"got back when an `egress-allow`/`egress-block` call returned "
"`status: pending`. Returns the current status: `pending` (no "
"decision yet — poll again later), `approved`, `modified`, "
"`rejected`, or `unknown` (no such queued proposal — wrong id, "
"or it was already resolved and read)."
),
"inputSchema": {
"type": "object",
"properties": {
"proposal_id": {
"type": "string",
"description": (
"The proposal id from a `pending` response."
),
},
},
"required": ["proposal_id"],
"additionalProperties": False,
},
},
]
@@ -390,7 +365,7 @@ def handle_tools_call(
deadline=deadline,
)
except TimeoutError:
text = format_pending_response_text(proposal.id, config.response_timeout_seconds)
text = format_pending_response_text(config.response_timeout_seconds)
return {
"content": [{"type": "text", "text": text}],
"isError": False,
@@ -407,54 +382,6 @@ def handle_tools_call(
}
def handle_check_proposal(
params: dict[str, object],
config: ServerConfig,
) -> dict[str, object]:
"""Non-blocking poll of a queued proposal's decision, by id.
Never creates a Proposal (so `check-proposal` isn't in `TOOLS`); it only
reads the queue. Resolution order mirrors the synchronous path's terminal
step a decided proposal is archived here exactly as `handle_tools_call`
archives it after `wait_for_response`, so `pending` proposals stay visible
to the operator until they're both decided *and* polled."""
args_raw = params.get("arguments", {})
if not isinstance(args_raw, dict):
raise _RpcClientError(ERR_INVALID_PARAMS, "tools/call 'arguments' must be an object")
proposal_id = args_raw.get("proposal_id")
if not isinstance(proposal_id, str) or not proposal_id.strip():
raise _RpcClientError(
ERR_INVALID_PARAMS,
"check-proposal: 'proposal_id' is required and must be a non-empty string",
)
proposal_id = proposal_id.strip()
try:
response = _sv.read_response(config.bottle_slug, proposal_id)
except FileNotFoundError:
# No decision yet — distinguish "still queued" from "unknown id".
try:
_sv.read_proposal(config.bottle_slug, proposal_id)
except FileNotFoundError:
return {
"content": [{"type": "text", "text": format_unknown_proposal_text(proposal_id)}],
"isError": True,
}
return {
"content": [{"type": "text", "text": format_still_pending_text(proposal_id)}],
"isError": False,
}
try:
_sv.archive_proposal(config.bottle_slug, proposal_id)
except OSError as e:
raise _RpcInternalError(f"failed to archive proposal: {e}") from e
return {
"content": [{"type": "text", "text": format_response_text(response)}],
"isError": response.status == _sv.STATUS_REJECTED,
}
def format_response_text(response: "_sv.Response") -> str:
"""Pretty-print a Response for the tool's text content. The agent
reads the text and decides whether to retry / give up / surface."""
@@ -467,35 +394,12 @@ def format_response_text(response: "_sv.Response") -> str:
return "\n".join(lines)
def format_pending_response_text(proposal_id: str, timeout_seconds: float) -> str:
"""Grace-window timeout: the proposal stays queued, and the agent is
told the id so it can `check-proposal` instead of re-proposing."""
def format_pending_response_text(timeout_seconds: float) -> str:
return "\n".join([
"status: pending",
f"proposal_id: {proposal_id}",
(
f"notes: no operator decision within {timeout_seconds:g}s; the "
"proposal remains queued. Poll it (do not re-propose) by calling "
f"`check-proposal` with proposal_id={proposal_id!r}."
),
])
def format_still_pending_text(proposal_id: str) -> str:
return "\n".join([
"status: pending",
f"proposal_id: {proposal_id}",
"notes: still queued; no operator decision yet. Call `check-proposal` again later.",
])
def format_unknown_proposal_text(proposal_id: str) -> str:
return "\n".join([
"status: unknown",
f"proposal_id: {proposal_id}",
(
"notes: no queued proposal with this id for this bottle — the id "
"may be wrong, or the proposal was already resolved and read."
"notes: operator response timed out after "
f"{timeout_seconds:g}s; proposal remains queued"
),
])
@@ -590,11 +494,6 @@ class MCPHandler(http.server.BaseHTTPRequestHandler):
# — silently dropping base routes like api.anthropic.com on approval.
if req.params.get("name") == _sv.TOOL_LIST_EGRESS_ROUTES:
return self._resolved_routes_payload()
# `check-proposal` is a non-blocking read of the calling bottle's
# own queue — attributed by source IP like a proposal, but it
# never queues or blocks.
if req.params.get("name") == _sv.TOOL_CHECK_PROPOSAL:
return handle_check_proposal(req.params, self._attributed_config(config))
# Attribute the proposal to the source-IP-resolved bottle, so the one
# shared server queues each bottle's proposal under its own slug.
return handle_tools_call(req.params, self._attributed_config(config))
-5
View File
@@ -20,10 +20,6 @@ TOOL_EGRESS_ALLOW = "egress-allow"
TOOL_GITLEAKS_ALLOW = "gitleaks-allow"
TOOL_EGRESS_TOKEN_ALLOW = "egress-token-allow"
TOOL_LIST_EGRESS_ROUTES = "list-egress-routes"
# Read-only agent tool: poll a queued proposal for the operator's decision
# without blocking or re-proposing. It never becomes a `Proposal.tool` (no
# queue record is created for it), so it is intentionally NOT in `TOOLS`.
TOOL_CHECK_PROPOSAL = "check-proposal"
TOOLS: tuple[str, ...] = (
TOOL_EGRESS_ALLOW,
TOOL_EGRESS_BLOCK,
@@ -160,7 +156,6 @@ __all__ = [
"TOOLS",
"TOOL_EGRESS_ALLOW",
"TOOL_EGRESS_BLOCK",
"TOOL_CHECK_PROPOSAL",
"TOOL_EGRESS_TOKEN_ALLOW",
"TOOL_GITLEAKS_ALLOW",
"TOOL_LIST_EGRESS_ROUTES",
-10
View File
@@ -7,7 +7,6 @@ from __future__ import annotations
import ipaddress
import os
import sys
def is_ip_literal(value: str) -> bool:
@@ -18,15 +17,6 @@ def is_ip_literal(value: str) -> bool:
return True
def read_tty_line() -> str:
"""Mirror `IFS= read -r REPLY </dev/tty`. Falls back to stdin."""
try:
with open("/dev/tty", "r", encoding="utf-8") as tty:
return tty.readline().rstrip("\n")
except OSError:
return sys.stdin.readline().rstrip("\n")
def expand_tilde(path: str) -> str:
"""Expand a leading '~' to $HOME. Leaves paths without a leading
tilde unchanged. Falls back to the empty string if $HOME is unset
@@ -1,57 +0,0 @@
# ADR 0005: Keep tracker metadata on issues
- **Status:** Accepted
- **Date:** 2026-07-18
- **Deciders:** didericis
## Context
Gitea exposes labels on both issues and pull requests. Applying the same labels
to both copies planning metadata, creates a synchronization obligation, and
makes disagreements between the two records possible. At the same time,
unlabelled objects look accidental unless the repository states which object
owns the metadata.
The repository already uses issues as work items and PRs as implementations of
those work items. At this decision's cutoff, all open PRs reference issues, but
121 of 219 historically merged PRs do not. Manufacturing retrospective issues
for that history would create records that never participated in planning and
would make the issue history less truthful.
## Decision
Issues are the canonical tracker records and own labels. Every issue has at
least one label. An issue opened or left without labels receives
`Status/Needs Triage` automatically until it is classified.
Pull requests carry no labels. Every new PR deliberately references at least
one existing issue in its title or description with one of these forms:
- `Closes #123`, `Fixes #123`, or `Resolves #123` when merging completes it.
- `Part of #123`, `Related to #123`, `Refs #123`, or `References #123` when it
contributes without completing it.
Gitea Actions enforces both PR rules as a status check and repairs the empty
issue-label state. Branch protection makes the PR policy check required.
The policy applies from 2026-07-18 onward. Existing issues may be labelled as
they are encountered, but closed PRs are grandfathered: no retrospective
issues or PR labels are created solely to make history conform.
## Consequences
- Classification, priority, and workflow metadata have one source of truth.
- A PR's issue link is the navigation path to its planning metadata.
- Multi-PR issues do not require copied or synchronized labels.
- `Status/Needs Triage` is an intentional fallback, not a final
classification.
- Direct issue creation remains convenient; automation repairs a missing label
immediately after creation because Gitea has no native required-label rule.
- The required check must be configured in branch protection after this
workflow lands.
## Links
- Issue #405.
- `.gitea/workflows/tracker-policy.yml`.
- `scripts/tracker_policy.py`.
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0023: smolmachines bottle backend
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis
@@ -1,6 +1,6 @@
# PRD 0032: Decompose smolmachines launch and harden bringup sequencing
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-claude
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0038: smolmachines Env Contract and Secret-Safe Injection
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-codex
@@ -1,6 +1,6 @@
# PRD 0039: smolmachines Capability-Block Remediation
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-codex
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0042: smolmachines Cross-Backend Parity Tests
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-codex
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0057: Promote smolmachines to default backend; convert Docker to example-only
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0068: smolmachines backend on Linux
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** Claude
-38
View File
@@ -312,44 +312,6 @@ reaches over the RPC rather than a shared mount into the VM. WAL on the
shared DB is therefore a deliberate, tested future change — not enabled ad
hoc. `sqlite3` itself is stdlib, so "the host needs SQLite" is a non-cost.
### Gateway CA: host-resident, like the DB
The shared gateway bumps TLS with a self-generated mitmproxy CA, and **every
bottle installs that CA** into its trust store to accept the bumped leaves. So
the CA is durable per-host state with the same rule as the DB: it must outlive
any single gateway container, or a restart mints a fresh CA that every
already-running bottle distrusts — the TLS handshake then fails even after the
bottle re-resolves and reconnects to the moved gateway (issue #450, a
re-attachment blocker distinct from #443/#445).
The CA lives on the **host filesystem** at `bot_bottle_root()/gateway-ca`
(`host_gateway_ca_dir()`), bind-mounted into the container at mitmproxy's
confdir. This is deliberately a host bind-mount, **not a container-runtime
named volume**: a named volume survives ordinary container removal but can be
silently wiped by Docker's or Apple Container's volume-prune commands during
routine host maintenance, which is exactly how the ephemeral-CA symptom shows
up in practice. A path under the app-data root is not managed or pruned by the
container runtime, and stays directly inspectable and rotatable from the host.
mitmproxy reuses an existing CA and generates one only on first run, so the
bind-mount alone gives
"adopt-existing, generate-on-first-run" for free.
The macOS backend uses the same host-resident CA directory and bind-mounts it
into the consolidated Apple infra container. Its `bot-bottle-mac-db` named
volume remains container-only because that prevents incoherent cross-kernel
SQLite locking, but the CA is deliberately not stored there: Apple Container
also has a `container volume prune` operation, and the named volume is
temporarily unreferenced while the infra container is recreated. Keeping the
CA on the host makes both ordinary recreation and volume pruning safe.
**Deliberate rollover** is the explicit inverse: `rotate_gateway_ca()` removes
the persisted CA material so the next start remints it, and the
`python -m bot_bottle.orchestrator.rotate_ca` one-shot wires that together with
dropping the running gateway container (whose mitmproxy still holds the old CA
in memory). Rotation does not auto-re-provision the new CA into running bottles
— those re-attach to install the new anchor — so it is an operator action with
a brief egress interruption, never an implicit one.
## Sequencing
Jump straight to the **virtualized** end state (not a host-daemon stepping
-110
View File
@@ -1,110 +0,0 @@
# PRD prd-new: CI artifact-based coverage and local Firecracker candidate flow
- **Status:** Active
- **Author:** Claude
- **Created:** 2026-07-21
- **Issue:** #446
## Summary
Restructure the CI test pipeline to run each test suite exactly once, upload
small `.coverage.*` artifacts, and combine them in a lightweight aggregation
job. Move the infra build onto the KVM runner so the ~194 MB rootfs never
crosses the network for PRs. On main-branch pushes, publish the byte-identical
rootfs that was tested.
## Motivation
The prior pipeline had two redundant costs:
1. **Duplicate artifact transfers.** `build-infra` (ubuntu-latest) built and
uploaded the ~194 MB rootfs; `integration-firecracker` downloaded it; the
`coverage` job downloaded it a second time. Combined download overhead: ~83
seconds per run, plus the ~70-second upload.
2. **Duplicate test execution.** `integration-firecracker` ran the Firecracker
integration suite; `coverage` ran the entire unit + integration suite again
on the same KVM runner to collect coverage data. Every line of Firecracker
code was tested twice per CI run.
## Goals
- Each test suite (unit, integration-docker, integration-firecracker) executes
exactly once per workflow run.
- PRs incur no large artifact transfers — the rootfs stays on the KVM runner.
- Main-branch pushes publish a byte-for-byte identical rootfs to the one that
passed the integration tests.
- Concurrent workflow runs cannot cross-publish candidates (naturally enforced
by Gitea Actions' per-run artifact scoping).
- Failed or cancelled runs block publication (enforced by the `needs:` chain on
`publish-infra`).
## Non-goals
- Changing test semantics or the coverage policy (ADR 0004).
- Removing the KVM runner guard on `integration-firecracker` and `coverage`.
- Changing how `publish_infra.py` builds or uploads the rootfs.
## Design
### Job graph
```
unit ──────────────────────────────────┐
integration-docker ────────────────────┤──► coverage ──► publish-infra (main only)
integration-firecracker (KVM) ─────────┘
```
### `unit`
Unchanged except: `coverage run` writes `--data-file=.coverage.unit`; the file
is uploaded as the `coverage-unit` artifact.
### `integration-docker`
Adds a `coverage` install step. `coverage run` writes `--data-file=.coverage.docker`;
the file is uploaded as `coverage-docker`.
### `integration-firecracker` (KVM runner)
Replaces the old `stage-firecracker-inputs``build-infra` → download chain:
1. Builds the infra candidate locally with
`BOT_BOTTLE_FC_DROPBEAR=/var/cache/bot-bottle-fc/dropbear`.
2. Boots the candidate and runs integration tests with coverage, writing
`.coverage.firecracker`.
3. Uploads the small `coverage-firecracker` artifact unconditionally.
4. On main-branch pushes only, uploads the rootfs as `infra-candidate` and the
dropbear as `firecracker-inputs` so `publish-infra` can verify and publish
the byte-identical artifact.
### `coverage`
Moves from a KVM runner to `ubuntu-latest`. No tests are re-executed:
1. Downloads `coverage-unit`, `coverage-docker`, and `coverage-firecracker`.
2. Runs `scripts/coverage.sh aggregate critical`, which calls
`coverage combine` then `coverage report`.
3. Runs the diff-coverage gate (`scripts/diff_coverage.py`).
Coverage files use `relative_files = True` (`.coveragerc`) so they combine
cleanly across runners with different absolute workspace paths.
### `publish-infra`
Depends on all four predecessor jobs (unchanged gate). Downloads `infra-candidate`
and `firecracker-inputs` that were uploaded by `integration-firecracker` on
main — the same byte sequence that passed the integration tests.
### Eliminated jobs
- `stage-firecracker-inputs`: existed only to copy the dropbear to ubuntu-latest
for `build-infra`. No longer needed.
- `build-infra`: the infra candidate is now built on the KVM runner in
`integration-firecracker`.
### Script changes
`scripts/coverage.sh` gains an `aggregate` mode (`coverage.sh aggregate [critical]`)
that combines pre-existing `.coverage.*` files instead of re-running tests.
The existing run mode (`coverage.sh [critical]`) is preserved for local dev.
@@ -1,146 +0,0 @@
# PRD prd-new: Claude forward_host_credentials
- **Status:** Draft
- **Author:** claude
- **Created:** 2026-07-01
- **Issue:** #325
## Summary
Add `agent_provider.forward_host_credentials: true` support for the
`claude` template, mirroring the existing Codex flow. When enabled,
bot-bottle reads the host's Claude OAuth session key from
`~/.claude/.credentials.json` at launch, forwards it only to the egress sidecar,
and injects a placeholder `CLAUDE_CODE_OAUTH_TOKEN` into the agent so
Claude Code starts without ever seeing the real credential.
## Problem
Running a Claude agent in a container today requires the operator to
manually extract a long-lived OAuth token (`claude setup-token`), export
it as `BOT_BOTTLE_CLAUDE_OAUTH_TOKEN`, and reference it explicitly in
the manifest with `agent_provider.auth_token:
"BOT_BOTTLE_CLAUDE_OAUTH_TOKEN"`. This is a two-step manual ceremony
that is easy to skip or do incorrectly.
The host already stores a valid Claude session in `~/.claude/.credentials.json`
after `claude login`. Codex already automates an
equivalent extraction from `~/.codex/auth.json`. There is no reason
Claude bottles cannot do the same.
## Goals / Success Criteria
- A Claude bottle with `forward_host_credentials: true` in the manifest
uses the host's `~/.claude/.credentials.json` session key at launch with no
additional operator steps.
- The agent container receives only `CLAUDE_CODE_OAUTH_TOKEN=egress-placeholder`
— never the real token.
- The real session key lives only in the egress sidecar's environment.
- Missing, malformed, or expired host Claude auth fails launch with a
clear operator-facing message.
- Existing `auth_token` behavior is unchanged.
- `forward_host_credentials: true` is rejected in the manifest when both
`auth_token` and `forward_host_credentials` are set, since they serve
the same purpose.
## Non-goals
- Refreshing Claude OAuth tokens in the sidecar.
- Writing a dummy `~/.claude.json` auth state to the agent (unlike the
Codex flow, Claude Code reads its credential from `CLAUDE_CODE_OAUTH_TOKEN`
in env, not from an auth file — no guest-side auth marker is needed).
- Supporting `forward_host_credentials` for providers other than `codex`
and `claude`.
## Design
### Manifest schema
```yaml
agent_provider:
template: claude
forward_host_credentials: true
```
Rejects in manifest validation when:
- Template is not `codex` or `claude`.
- Both `auth_token` and `forward_host_credentials` are set.
### Host auth extraction (`contrib/claude/claude_auth.py`)
Claude Code credential storage varies by platform:
- **Linux**: `~/.claude/.credentials.json`
- **macOS**: macOS Keychain, service `"Claude Code-credentials"`
(the file path is tried first; Keychain is the fallback when the file
is absent)
`~/.claude.json` contains only UI state and profile metadata — no token.
The credentials JSON schema (same whether from file or Keychain):
```json
{
"claudeAiOauth": {
"accessToken": "<access-token>",
"refreshToken": "<refresh-token>",
"expiresAt": 1748276587173,
"scopes": ["user:inference", "user:profile"]
}
}
```
`expiresAt` is in **milliseconds** (not seconds).
At prepare/launch time, when `forward_host_credentials: true`:
1. Try `~/.claude/.credentials.json`; on macOS, if absent, run
`security find-generic-password -s "Claude Code-credentials" -w`
and parse its stdout as JSON.
2. Require a `claudeAiOauth` dict.
3. Require a non-empty `claudeAiOauth.accessToken` string.
4. If `claudeAiOauth.expiresAt` is present, divide by 1000 and require
the result to be in the future.
5. Return only the access token to the launch path.
Errors name the missing or invalid condition and point the operator at
`claude login`, without printing token values.
### Egress route
When `forward_host_credentials: true`:
- Provision the session key in `provisioned_env` under
`BOT_BOTTLE_CLAUDE_HOST_ACCESS_TOKEN` (new constant in `egress.py`).
- Set up the `api.anthropic.com` egress route with `auth_scheme: Bearer`
and `token_ref: BOT_BOTTLE_CLAUDE_HOST_ACCESS_TOKEN`.
- Set `CLAUDE_CODE_OAUTH_TOKEN=egress-placeholder` in the agent env and
add it to `hidden_env_names`.
No dummy auth file and no `verify` step are needed — Claude Code reads
the credential from the env var, not from a file.
### Constants
- `CLAUDE_HOST_CREDENTIAL_TOKEN_REF = "BOT_BOTTLE_CLAUDE_HOST_ACCESS_TOKEN"`
in `egress.py` (alongside the existing `CODEX_HOST_CREDENTIAL_TOKEN_REF`).
- `CLAUDE_HOST_CREDENTIAL_HOSTS = ("api.anthropic.com",)` in
`agent_provider.py` (alongside the existing `CODEX_HOST_CREDENTIAL_HOSTS`).
### Data flow
```
Host ~/.claude/.credentials.json → bot-bottle launch
├──► egress sidecar env (real token only)
└──► agent env: CLAUDE_CODE_OAUTH_TOKEN=egress-placeholder
Agent → HTTPS to api.anthropic.com (via egress)
Egress → injects Authorization: Bearer <real token>
Egress → forwards to api.anthropic.com
```
## Open questions
None — the Codex precedent makes the design clear.

Some files were not shown because too many files have changed in this diff Show More