Compare commits
7 Commits
e1f10fb9da
...
2e8fdc7234
| Author | SHA1 | Date | |
|---|---|---|---|
| 2e8fdc7234 | |||
| 02552c5196 | |||
| 10b8c49076 | |||
| 4f16a9876d | |||
| b25911c575 | |||
| 727eafe0f9 | |||
| 1ec114b6d7 |
+94
-26
@@ -4,16 +4,15 @@
|
|||||||
# dependencies are required to execute it. Tests are split by directory:
|
# dependencies are required to execute it. Tests are split by directory:
|
||||||
#
|
#
|
||||||
# tests/unit/ — pure unit tests; always run
|
# tests/unit/ — pure unit tests; always run
|
||||||
# tests/integration/ — need a reachable Docker daemon; skip cleanly
|
# tests/integration/ — need a reachable backend; skip cleanly when
|
||||||
# (via tests/_docker.py:skip_unless_docker) when
|
# the backend isn't available on the runner
|
||||||
# Docker isn't available on the runner
|
|
||||||
# tests/canaries/ — upstream regression canaries; run on a separate
|
# tests/canaries/ — upstream regression canaries; run on a separate
|
||||||
# schedule (see canaries.yml), not here
|
# schedule (see canaries.yml), not here
|
||||||
#
|
#
|
||||||
# This workflow assumes the Gitea Actions runner exposes the host Docker
|
# Integration tests run once per backend in separate jobs. Each job sets
|
||||||
# socket to the job container so `docker` commands inside the job can
|
# BOT_BOTTLE_BACKEND explicitly so the test suite uses the right backend.
|
||||||
# reach the daemon. If that's not yet configured on the runner the
|
# Backends that aren't available on the runner fail the preflight step
|
||||||
# integration tests will skip rather than fail.
|
# rather than silently skipping inside the test output.
|
||||||
|
|
||||||
name: test
|
name: test
|
||||||
|
|
||||||
@@ -23,9 +22,16 @@ on:
|
|||||||
- main
|
- main
|
||||||
paths:
|
paths:
|
||||||
- '**.py'
|
- '**.py'
|
||||||
|
- '.gitea/workflows/**.yml'
|
||||||
|
- 'scripts/**'
|
||||||
|
- 'README.md'
|
||||||
pull_request:
|
pull_request:
|
||||||
paths:
|
paths:
|
||||||
- '**.py'
|
- '**.py'
|
||||||
|
- '.gitea/workflows/**.yml'
|
||||||
|
- 'scripts/**'
|
||||||
|
- 'README.md'
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
jobs:
|
jobs:
|
||||||
unit:
|
unit:
|
||||||
@@ -48,7 +54,7 @@ jobs:
|
|||||||
- name: Report unit coverage
|
- name: Report unit coverage
|
||||||
run: python3 -m coverage report -m
|
run: python3 -m coverage report -m
|
||||||
|
|
||||||
integration:
|
integration-docker:
|
||||||
runs-on: ubuntu-latest
|
runs-on: ubuntu-latest
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout
|
- name: Checkout
|
||||||
@@ -65,33 +71,95 @@ jobs:
|
|||||||
echo "docker not on PATH — integration tests will skip"
|
echo "docker not on PATH — integration tests will skip"
|
||||||
fi
|
fi
|
||||||
|
|
||||||
- name: Run integration tests
|
- name: Run integration tests (docker)
|
||||||
|
env:
|
||||||
|
BOT_BOTTLE_BACKEND: docker
|
||||||
run: python3 -m unittest discover -t . -s tests/integration -v
|
run: python3 -m unittest discover -t . -s tests/integration -v
|
||||||
|
|
||||||
# Combined unit+integration coverage report (informational). See
|
# Integration tests against the Firecracker backend. Runs on a self-hosted
|
||||||
# docs/decisions/0004-coverage-policy.md.
|
# KVM runner (label `kvm`) where /dev/kvm and the TAP/nft pool are available.
|
||||||
#
|
#
|
||||||
# The hard diff-coverage gate (changed lines >= 90%) is DEFERRED: the
|
# Restricted to same-repo PRs, push to main, and workflow_dispatch — fork
|
||||||
# Firecracker backend's VM/SSH orchestration is covered by the integration
|
# PRs don't execute untrusted code on the privileged runner.
|
||||||
# suite, which needs /dev/kvm + the provisioned TAP/nft pool — a
|
#
|
||||||
# container-based runner skips it and those lines read uncovered, so the
|
# Runner prerequisites (provision once; see README "Firecracker on Linux"):
|
||||||
# gate can't pass here. Re-enabling it on a self-hosted KVM runner is
|
# `firecracker` on PATH, `/dev/kvm` accessible, Docker, cached kernel +
|
||||||
# tracked separately (see PRD 0069 / #348 and the ci-runner branch).
|
# static dropbear, and the pool as a persistent systemd unit.
|
||||||
|
integration-firecracker:
|
||||||
|
runs-on: [self-hosted, kvm]
|
||||||
|
if: >-
|
||||||
|
github.event_name == 'push' ||
|
||||||
|
github.event_name == 'workflow_dispatch' ||
|
||||||
|
(github.event_name == 'pull_request' &&
|
||||||
|
github.event.pull_request.head.repo.full_name == github.repository)
|
||||||
|
steps:
|
||||||
|
- name: Checkout
|
||||||
|
uses: actions/checkout@v4
|
||||||
|
|
||||||
|
- name: Preflight — Firecracker host is ready
|
||||||
|
run: |
|
||||||
|
command -v firecracker >/dev/null || {
|
||||||
|
echo "firecracker not on PATH — provision the runner (README: Firecracker on Linux)"; exit 1; }
|
||||||
|
test -e /dev/kvm || { echo "/dev/kvm missing — KVM not available on this runner"; exit 1; }
|
||||||
|
# `backend status` exits non-zero unless the TAP pool is up + no
|
||||||
|
# range overlap; it prints the exact `backend setup` fix.
|
||||||
|
python3 cli.py backend status --backend=firecracker
|
||||||
|
|
||||||
|
- name: Install dev requirements
|
||||||
|
run: python3 -m pip install --user -r requirements-dev.txt
|
||||||
|
|
||||||
|
- name: Run integration tests (firecracker)
|
||||||
|
env:
|
||||||
|
BOT_BOTTLE_BACKEND: firecracker
|
||||||
|
run: python3 -m unittest discover -t . -s tests/integration -v
|
||||||
|
|
||||||
|
# Combined unit+integration coverage + the diff-coverage gate (the hard
|
||||||
|
# gate: new/changed lines >= 90%). See docs/decisions/0004-coverage-policy.md.
|
||||||
|
#
|
||||||
|
# This runs on a self-hosted KVM runner (label `kvm`), NOT ubuntu-latest,
|
||||||
|
# because the Firecracker backend's subprocess/VM orchestration
|
||||||
|
# (launch/boot/SSH/isolation-probe) is covered by the integration suite,
|
||||||
|
# and that suite needs `/dev/kvm` + the provisioned TAP/nft pool — which a
|
||||||
|
# container-based runner doesn't have. On such a runner the firecracker
|
||||||
|
# integration test skips and its ~230 orchestration lines read as
|
||||||
|
# uncovered, so the gate can't pass there.
|
||||||
|
#
|
||||||
|
# Restricted to the same events as integration-firecracker (same-repo PRs,
|
||||||
|
# push, workflow_dispatch) for the same security reason.
|
||||||
|
#
|
||||||
|
# See #414 for the planned follow-up: artifact-based coverage combination
|
||||||
|
# (run tests once in their respective jobs, combine .coverage files here).
|
||||||
coverage:
|
coverage:
|
||||||
runs-on: ubuntu-latest
|
runs-on: [self-hosted, kvm]
|
||||||
|
if: >-
|
||||||
|
github.event_name == 'push' ||
|
||||||
|
github.event_name == 'workflow_dispatch' ||
|
||||||
|
(github.event_name == 'pull_request' &&
|
||||||
|
github.event.pull_request.head.repo.full_name == github.repository)
|
||||||
steps:
|
steps:
|
||||||
- name: Checkout
|
- name: Checkout
|
||||||
uses: actions/checkout@v4
|
uses: actions/checkout@v4
|
||||||
with:
|
with:
|
||||||
fetch-depth: 0
|
fetch-depth: 0
|
||||||
|
|
||||||
# No actions/setup-python: the runner image already ships Python 3.12,
|
- name: Preflight — Firecracker host is ready
|
||||||
# and older act_runner engines mishandle setup-python's PATH (coverage
|
run: |
|
||||||
# lands in one interpreter, `python3` resolves to another). Install
|
command -v firecracker >/dev/null || {
|
||||||
# straight into the ephemeral job container's system Python —
|
echo "firecracker not on PATH — provision the runner (README: Firecracker on Linux)"; exit 1; }
|
||||||
# --break-system-packages is safe because the container is disposable.
|
test -e /dev/kvm || { echo "/dev/kvm missing — KVM not available on this runner"; exit 1; }
|
||||||
- name: Install dev requirements
|
# `backend status` exits non-zero unless the TAP pool is up + no
|
||||||
run: python3 -m pip install --break-system-packages -r requirements-dev.txt
|
# range overlap; it prints the exact `backend setup` fix.
|
||||||
|
python3 cli.py backend status --backend=firecracker
|
||||||
|
|
||||||
- name: Combined coverage report (unit + integration)
|
- name: Install dev requirements
|
||||||
|
run: python3 -m pip install --user -r requirements-dev.txt
|
||||||
|
|
||||||
|
- name: Combined coverage (unit + integration, incl. firecracker)
|
||||||
|
env:
|
||||||
|
BOT_BOTTLE_BACKEND: firecracker
|
||||||
run: PYTHON=python3 bash scripts/coverage.sh critical
|
run: PYTHON=python3 bash scripts/coverage.sh critical
|
||||||
|
|
||||||
|
- name: Diff-coverage gate (changed lines >= 90%)
|
||||||
|
run: |
|
||||||
|
git fetch --no-tags origin main:refs/remotes/origin/main
|
||||||
|
python3 scripts/diff_coverage.py --base origin/main --min 90
|
||||||
|
|||||||
@@ -90,6 +90,8 @@ BOT_BOTTLE_BACKEND=firecracker ./cli.py start <agent>
|
|||||||
|
|
||||||
> **NixOS:** enable `virtualisation.docker`, ensure the KVM module is loaded (`boot.kernelModules = [ "kvm-intel" ];` or `kvm-amd`), and add your user to the `kvm` and `docker` groups. For the network pool, consume the flake module — `imports = [ inputs.bot-bottle.nixosModules.firecracker-netpool ]; services.bot-bottle-firecracker = { enable = true; owner = "you"; };` — then `nixos-rebuild switch` (imperative nft/TAP rules don't survive a rebuild; channel users can `imports = [ <bot-bottle>/nix/firecracker-netpool.nix ]`). `firecracker` isn't in nixpkgs by default as a user binary — install the release binary (pin the version) and put it on `PATH`.
|
> **NixOS:** enable `virtualisation.docker`, ensure the KVM module is loaded (`boot.kernelModules = [ "kvm-intel" ];` or `kvm-amd`), and add your user to the `kvm` and `docker` groups. For the network pool, consume the flake module — `imports = [ inputs.bot-bottle.nixosModules.firecracker-netpool ]; services.bot-bottle-firecracker = { enable = true; owner = "you"; };` — then `nixos-rebuild switch` (imperative nft/TAP rules don't survive a rebuild; channel users can `imports = [ <bot-bottle>/nix/firecracker-netpool.nix ]`). `firecracker` isn't in nixpkgs by default as a user binary — install the release binary (pin the version) and put it on `PATH`.
|
||||||
|
|
||||||
|
> **CI:** the coverage gate (`.gitea/workflows/test.yml` → `coverage` job) runs on a self-hosted runner labelled `kvm`, because the Firecracker backend's VM/SSH orchestration is exercised only by the integration suite, which needs `/dev/kvm` + the provisioned pool (a container runner would skip it and read as uncovered). Provision that runner exactly like a normal Firecracker host — `firecracker` on `PATH`, `/dev/kvm`, Docker, the cached guest kernel + static dropbear, and the pool installed as the persistent systemd unit — then register it with the `kvm` label. The unit/lint jobs still run on `ubuntu-latest`.
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
./cli.py start <agent> # builds the image on first run, drops you into claude
|
./cli.py start <agent> # builds the image on first run, drops you into claude
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -0,0 +1,308 @@
|
|||||||
|
# HN discourse on agent sandbox safety — June/July 2026
|
||||||
|
|
||||||
|
A survey of community opinion and notable security disclosures on Hacker
|
||||||
|
News and adjacent sources over June–July 2026. The question: what does
|
||||||
|
the current discourse say about whether sandboxes are sufficient for
|
||||||
|
agentic AI safety, and where does bot-bottle land against the issues
|
||||||
|
being raised?
|
||||||
|
|
||||||
|
Research conducted 2026-07-18.
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
The past month marks a turning point in community opinion. Earlier in
|
||||||
|
2026, the debate was mostly "which sandbox tool is best?" By June–July,
|
||||||
|
a cascade of critical CVEs and novel attack classes has shifted the
|
||||||
|
framing to "sandboxes are not enough — what else do you need?" The
|
||||||
|
attacks that drove this shift are structurally distinct: most route
|
||||||
|
through legitimate, trusted channels (Sentry issues, MCP descriptions,
|
||||||
|
README files) rather than exploiting the isolation boundary directly.
|
||||||
|
|
||||||
|
bot-bottle's architecture holds up well against the direct-escape class
|
||||||
|
(Firecracker/Apple Container default backends, credentials never in the
|
||||||
|
agent's env, harness entirely on the host). The remaining gap is prompt
|
||||||
|
injection — attacker-controlled data interpreted as model instructions.
|
||||||
|
Egress controls and prompt injection defenses are orthogonal: egress
|
||||||
|
limits what the agent can *send out*; injection is about what it is
|
||||||
|
*told to do*. The two don't substitute for each other. Inside a tightly-
|
||||||
|
egressed sandbox a successful injection can't exfiltrate to unknown
|
||||||
|
hosts, but it can still corrupt the work product, push malicious commits
|
||||||
|
past a secret scanner, or use allowlisted channels for exfiltration.
|
||||||
|
Those residual risks are addressed below.
|
||||||
|
|
||||||
|
## The sandboxing boom sets the stage
|
||||||
|
|
||||||
|
The preceding months generated a wave of sandbox tooling. A March 28
|
||||||
|
Ask HN thread
|
||||||
|
([#47444917](https://news.ycombinator.com/item?id=47444917)) catalogued
|
||||||
|
the explosion: E2B, AIO Sandbox, AgentSphere, Yolobox, Exe.dev,
|
||||||
|
AgentFence, DenoSandbox, Capsule (WASM), ERA, Vibekit, Daytona, Modal,
|
||||||
|
Nono, and more — all launched within roughly 12 months. A parallel March
|
||||||
|
9 thread ([#47185250](https://news.ycombinator.com/item?id=47185250))
|
||||||
|
surveyed what developers were actually deploying: "containers or YOLO"
|
||||||
|
dominated. The honest community mood was that most teams hadn't solved
|
||||||
|
this and were shipping anyway.
|
||||||
|
|
||||||
|
## The June–July attack cascade
|
||||||
|
|
||||||
|
Six attack patterns broke in quick succession. Together they form the
|
||||||
|
argument that the community's framing was wrong: the threat model for
|
||||||
|
agents isn't just "code that escapes its container" — it's also prompt
|
||||||
|
injection, where attacker-controlled data is interpreted as model
|
||||||
|
instructions regardless of whether any isolation boundary was crossed.
|
||||||
|
Sections 2–4 below are all the same attack class; the "trusted channel"
|
||||||
|
label describes the delivery vector, not a different threat.
|
||||||
|
|
||||||
|
### 1. Sandbox escape CVEs (DuneSlide, CVE-2026-39861)
|
||||||
|
|
||||||
|
Cato AI Labs disclosed **DuneSlide** (CVE-2026-50548/50549, CVSS 9.8),
|
||||||
|
a pair of flaws in Cursor 2.x. CVE-2026-50548 abuses the sandbox's
|
||||||
|
`working_directory` parameter to point writes at system files; CVE-26-50549
|
||||||
|
exploits a symlink-resolution fallback that fails open. Both start with
|
||||||
|
a prompt injection and end in sandbox escape — and Cato's framing was
|
||||||
|
blunt: "each CVE defeats a different guardrail; the problem is
|
||||||
|
structural, not a string of one-offs."
|
||||||
|
|
||||||
|
Claude Code's own sandbox had a similar escape this year:
|
||||||
|
**CVE-2026-39861** (symlink flaw). The CurXecute/MCPoison/CVE-2026-26268
|
||||||
|
chain from Cursor added a poisoned Slack message, a swap-after-approval
|
||||||
|
MCP config, and a Git hook as three more entry points in the same
|
||||||
|
attack class.
|
||||||
|
|
||||||
|
All patched, but the pattern holds: any application-level sandbox that
|
||||||
|
takes attacker-influenced values as path parameters is reachable from a
|
||||||
|
prompt injection.
|
||||||
|
|
||||||
|
### 2. Prompt injection via MCP data (Agentjacking)
|
||||||
|
|
||||||
|
Tenet's "Agentjacking" technique planted a fake bug report in Sentry's
|
||||||
|
MCP output. When an agent queries Sentry to fix open issues, the
|
||||||
|
malicious event is rendered as structured content visually
|
||||||
|
indistinguishable from a real Sentry event, and the agent executes the
|
||||||
|
embedded instructions with the developer's full privileges. Hit rate
|
||||||
|
across Claude Code and Cursor: **85%**. The route is entirely through a
|
||||||
|
legitimately-authorized MCP channel — no isolation boundary is crossed;
|
||||||
|
the injection arrives inbound through a channel the sandbox explicitly
|
||||||
|
trusts.
|
||||||
|
|
||||||
|
The Cloud Security Alliance's summary: treat observability, bug-report,
|
||||||
|
and integration data as **untrusted agent input**, not neutral
|
||||||
|
development metadata.
|
||||||
|
|
||||||
|
### 3. README-embedded prompt injection
|
||||||
|
|
||||||
|
A July disclosure showed malicious instructions hidden in `README.md`
|
||||||
|
— a file that receives no trust prompt and requires no elevated access.
|
||||||
|
When asked point-blank whether the repo held hidden instructions, both
|
||||||
|
Claude Sonnet 4.6 and GPT-5.5 said no. A payload written for Sonnet
|
||||||
|
4.6 transferred unchanged to Sonnet 5, Opus 4.8, and GPT-5.5. The
|
||||||
|
attack surface is every repo an agent is asked to work in.
|
||||||
|
|
||||||
|
### 4. Prompt injection via MCP tool descriptions
|
||||||
|
|
||||||
|
Microsoft research (June 30) showed that attacker-controlled MCP tool
|
||||||
|
description fields can silently redirect agent behavior. The injection
|
||||||
|
is embedded in metadata the model reads during tool selection — before
|
||||||
|
any sandbox enforcement or egress check runs, and entirely on the
|
||||||
|
inbound path that egress controls cannot touch.
|
||||||
|
|
||||||
|
### 5. MCP STDIO command injection (10 CVEs)
|
||||||
|
|
||||||
|
OX Security disclosed a systemic command injection class in Anthropic's
|
||||||
|
MCP protocol, covering 10 CVEs across multiple coding agents. The
|
||||||
|
Windsurf case (CVE-2026-30615): processing attacker-controlled HTML
|
||||||
|
causes the agent to auto-register a malicious MCP STDIO server and
|
||||||
|
execute arbitrary commands with no further user interaction.
|
||||||
|
|
||||||
|
### 6. LiteLLM gateway compromise (CVE-2026-40217, CVE-2026-42271)
|
||||||
|
|
||||||
|
CVE-2026-40217 exposes LiteLLM's guardrail sandbox via `exec()` with no
|
||||||
|
source filtering. CVE-2026-42271 (exploited in the wild, added to CISA's
|
||||||
|
KEV catalog) lets callers spawn subprocesses through MCP preview
|
||||||
|
endpoints. The threat extends to any agent routed through a compromised
|
||||||
|
LiteLLM proxy: the proxy can swap model responses for forged tool calls
|
||||||
|
in transit, giving the attacker a reverse shell from the developer's
|
||||||
|
machine.
|
||||||
|
|
||||||
|
## HN community opinion clusters
|
||||||
|
|
||||||
|
**"Move enforcement to the kernel, not the app"** — the Nono Show HN
|
||||||
|
([#46849615](https://news.ycombinator.com/item?id=46849615)) and a
|
||||||
|
kernel-sandbox thread
|
||||||
|
([#47066574](https://news.ycombinator.com/item?id=47066574)) both argued
|
||||||
|
that application-layer sandboxes are inherently bypassable by the code
|
||||||
|
they're sandboxing. The academic framing, from *Red-Teaming the Agentic
|
||||||
|
Red-Team* ([arXiv 2606.24496](https://arxiv.org/pdf/2606.24496)):
|
||||||
|
"enforcement should occur at the OS level via the kernel refusing system
|
||||||
|
calls that violate policy at runtime — not pre-execution argument
|
||||||
|
validation in tool calls."
|
||||||
|
|
||||||
|
**"The harness belongs outside the sandbox"** — a May thread
|
||||||
|
([#47990675](https://news.ycombinator.com/item?id=47990675)) converged
|
||||||
|
on clean architectural separation: harness in one VM, tool execution in
|
||||||
|
another. Top comment: "having the harness in one VM, and tool use applied
|
||||||
|
to user data in another, is about as safe as you can be at present."
|
||||||
|
Several replies described a hypervisor-like policy layer — sitting outside
|
||||||
|
both VMs — as the right long-term model.
|
||||||
|
|
||||||
|
**"Sandboxes are too coarse-grained"** — a Feb thread
|
||||||
|
([#47006445](https://news.ycombinator.com/item?id=47006445)) argued
|
||||||
|
that VMs don't answer the real question: knowing whether an agent
|
||||||
|
*should* be sending an email or making a transaction. "Everything's just
|
||||||
|
in the same big box." This framing picked up traction through June–July
|
||||||
|
as the trusted-channel attacks dominated.
|
||||||
|
|
||||||
|
**"MCP's trust model is the real problem"** — the month's recurring
|
||||||
|
theme. MCP by design gives agents access to authorized external services.
|
||||||
|
Once a trusted channel delivers a malicious payload, filesystem sandboxing
|
||||||
|
is irrelevant. The community call: treat all MCP tool metadata and return
|
||||||
|
values as untrusted input subject to policy validation before ingestion,
|
||||||
|
and disable automatic MCP server loading from untrusted repositories.
|
||||||
|
|
||||||
|
## How bot-bottle addresses these issues
|
||||||
|
|
||||||
|
### What it covers well
|
||||||
|
|
||||||
|
**Direct sandbox escape (CVEs, container breakout)**
|
||||||
|
|
||||||
|
bot-bottle's default backends are Firecracker microVM (KVM Linux) and
|
||||||
|
Apple Container (macOS). Both run the agent in a separate VM with a
|
||||||
|
dedicated kernel — the container-escape CVE class (Dirty Pipe, runc
|
||||||
|
escapes, DuneSlide's path-parameter abuse) requires escaping a real
|
||||||
|
hypervisor boundary, not just a namespace. On the legacy Docker backend,
|
||||||
|
gVisor auto-detection provides a userspace syscall barrier for hosts where
|
||||||
|
neither KVM nor Apple Container is available.
|
||||||
|
|
||||||
|
The bot-bottle process itself runs entirely on the host, outside the VM.
|
||||||
|
This is the "harness outside the sandbox" architecture the HN thread
|
||||||
|
converged on as best practice. The bottle manifest, egress rules, and
|
||||||
|
secrets never enter the agent VM.
|
||||||
|
|
||||||
|
**Credential theft on sandbox escape**
|
||||||
|
|
||||||
|
Even on a successful VM/container escape, the agent has nothing useful
|
||||||
|
to steal. Credentials are injected in-flight by the gateway proxy
|
||||||
|
(`auth.scheme` / `auth.token_ref` in the egress route config) — `printenv`
|
||||||
|
inside the agent shows proxy URLs only. The git-gate similarly holds the
|
||||||
|
upstream SSH credential on the host; the agent pushes through a
|
||||||
|
gitleaks-scanned daemon that forwards clean refs upstream. An escaped
|
||||||
|
agent gets the host filesystem, not the keys.
|
||||||
|
|
||||||
|
**Orphaned-agent credential risk**
|
||||||
|
|
||||||
|
bot-bottle is explicitly ephemeral: when the agent exits, `cli.py` tears
|
||||||
|
down every gateway and both networks — nothing persists between runs. The
|
||||||
|
agent never holds credentials, so there is nothing to orphan.
|
||||||
|
|
||||||
|
**MCP config redirection / STDIO auto-registration**
|
||||||
|
|
||||||
|
The trust boundary at `$HOME` means bottles live only under
|
||||||
|
`~/.bot-bottle/bottles/` — a cloned repo cannot add egress routes or
|
||||||
|
redirect env vars to attacker hosts (the design rationale is in
|
||||||
|
`docs/prds/0011-per-file-md-manifest.md`). Auto-registering a malicious
|
||||||
|
MCP STDIO server from within the agent is still sandboxed by the VM, and
|
||||||
|
any outbound calls from that server must pass the egress allowlist and
|
||||||
|
outbound DLP scanner.
|
||||||
|
|
||||||
|
**Outbound exfiltration (any injection class)**
|
||||||
|
|
||||||
|
Whatever triggers the agent — README injection, Agentjacking, MCP
|
||||||
|
description poisoning — the final step in most attacks is exfiltration.
|
||||||
|
bot-bottle's egress allowlist is default-deny with a per-bottle host
|
||||||
|
allowlist; unknown hosts get a hard 403. Outbound DLP scanning
|
||||||
|
(`outbound_detectors: [token_patterns, known_secrets]`) catches tokens
|
||||||
|
and secrets in outbound bodies; the `supervise` policy (default for
|
||||||
|
manifest routes) holds the request for operator approval rather than
|
||||||
|
silently blocking it. Together these limit what a successful injection
|
||||||
|
can *do* even if it succeeds at the model layer.
|
||||||
|
|
||||||
|
**LiteLLM / compromised-proxy attacks**
|
||||||
|
|
||||||
|
bot-bottle does not use LiteLLM. The model API route (e.g.
|
||||||
|
`api.anthropic.com`) is an auto-injected provider route on the egress
|
||||||
|
allowlist; the agent dials the gateway, not the model API directly.
|
||||||
|
A compromised third-party proxy is not in the architecture.
|
||||||
|
|
||||||
|
### Where it is weaker
|
||||||
|
|
||||||
|
**Prompt injection**
|
||||||
|
|
||||||
|
Egress controls and prompt injection defenses are orthogonal. Egress
|
||||||
|
limits what the agent can *send out* (outbound leg); prompt injection
|
||||||
|
is about what attacker-controlled data *tells the agent to do* (inbound
|
||||||
|
leg). The two don't substitute for each other and must be treated
|
||||||
|
separately.
|
||||||
|
|
||||||
|
The inbound DLP scanner (`inbound_detectors: [naive_injection_detection]`)
|
||||||
|
is the only runtime defense against injection arriving through allowlisted
|
||||||
|
channels — Sentry MCP responses, MCP tool descriptions, README content.
|
||||||
|
It is explicitly pattern-matching and will not catch a sufficiently
|
||||||
|
crafted payload. There is no semantic / intent-level gate between what
|
||||||
|
the model decides and what the agent executes.
|
||||||
|
|
||||||
|
**Blast radius within the permitted scope**
|
||||||
|
|
||||||
|
Inside a tightly-egressed sandbox a successful injection can't
|
||||||
|
exfiltrate to unknown hosts, but it still has real options:
|
||||||
|
|
||||||
|
- *Work product corruption.* The agent can modify, delete, or backdoor
|
||||||
|
files in the working directory. This is within its permitted scope;
|
||||||
|
egress controls have nothing to say about it.
|
||||||
|
|
||||||
|
- *Malicious commits past the git-gate.* The git-gate scans outbound
|
||||||
|
refs for secrets (gitleaks), not for semantic code intent. A prompt-
|
||||||
|
injected agent can commit subtly malicious code — logic bombs,
|
||||||
|
backdoored auth paths, code that exfiltrates data through the
|
||||||
|
application's own HTTP clients at runtime — that looks clean to a
|
||||||
|
secret scanner.
|
||||||
|
|
||||||
|
- *Exfiltration through allowlisted channels.* If an attacker knows or
|
||||||
|
can predict what hosts are in the egress allowlist, those channels are
|
||||||
|
available for exfiltration. A GitHub remote being allowlisted means
|
||||||
|
"push to an attacker-controlled fork" is viable. A logging endpoint
|
||||||
|
being allowlisted means structured data can leave through it. The
|
||||||
|
outbound DLP scanner catches credential tokens and known secrets but
|
||||||
|
not arbitrary business data.
|
||||||
|
|
||||||
|
- *Dependency installation within the sandbox.* An agent that runs
|
||||||
|
`npm install` or `pip install` on attacker-specified packages executes
|
||||||
|
code inside the sandbox with the same capabilities the agent has:
|
||||||
|
filesystem access, tool calls, calls to allowlisted hosts. Supply chain
|
||||||
|
injection via package names is in the same injection family, triggered
|
||||||
|
by the same prompt-injection path.
|
||||||
|
|
||||||
|
### What would close the remaining gaps
|
||||||
|
|
||||||
|
The blast-radius risks above point at two distinct mitigations that
|
||||||
|
don't yet exist in bot-bottle:
|
||||||
|
|
||||||
|
- *Outbound intent classification.* The egress addon today scans
|
||||||
|
outbound request content for token patterns. What it lacks is
|
||||||
|
awareness of context — it can't distinguish "agent is pushing a
|
||||||
|
legitimate commit" from "agent was injected and is pushing a backdoor."
|
||||||
|
The `supervise` policy is already the right shape for human-in-the-loop
|
||||||
|
review on sensitive outbound actions; extending it with context from
|
||||||
|
the agent's recent tool calls (what files were touched, what was the
|
||||||
|
triggering task) would narrow the gap.
|
||||||
|
|
||||||
|
- *Semantic code review on git push.* gitleaks is the wrong tool for
|
||||||
|
catching injected logic. A review step on outbound commits — even a
|
||||||
|
simple diff summary surfaced in `cli.py supervise` before the push is
|
||||||
|
forwarded — would close the malicious-commit path without requiring
|
||||||
|
the agent to be fully trusted.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- [Ask HN: The new wave of AI agent sandboxes? (Mar 2026)](https://news.ycombinator.com/item?id=47444917)
|
||||||
|
- [OK, let's survey how everybody is sandboxing AI coding agents (Mar 2026)](https://news.ycombinator.com/item?id=47185250)
|
||||||
|
- [The agent harness belongs outside the sandbox (May 2026)](https://news.ycombinator.com/item?id=47990675)
|
||||||
|
- [Show HN: Nono – Kernel-enforced sandboxing for AI agents (Feb 2026)](https://news.ycombinator.com/item?id=46849615)
|
||||||
|
- [Kernel-enforced sandbox for AI agents, MCP and LLM workloads (Feb 2026)](https://news.ycombinator.com/item?id=47066574)
|
||||||
|
- [Sandboxes will be left in 2026 (Feb 2026)](https://news.ycombinator.com/item?id=47006445)
|
||||||
|
- [Critical Cursor Flaws / DuneSlide – The Hacker News](https://thehackernews.com/2026/07/critical-cursor-flaws-could-let-prompt.html)
|
||||||
|
- [Agentjacking Attack – The Hacker News](https://thehackernews.com/2026/06/agentjacking-attack-tricks-ai-coding.html)
|
||||||
|
- [Friendly Fire: AI Agents Built to Catch Malicious Code – The Hacker News](https://thehackernews.com/2026/07/friendly-fire-ai-agents-built-to-catch.html)
|
||||||
|
- [Microsoft Warns Poisoned MCP Tool Descriptions – The Hacker News](https://thehackernews.com/2026/06/microsoft-warns-poisoned-mcp-tool.html)
|
||||||
|
- [MCP STDIO Command Injection Advisory – OX Security](https://www.ox.security/blog/mcp-supply-chain-advisory-rce-vulnerabilities-across-the-ai-ecosystem/)
|
||||||
|
- [LiteLLM Vulnerability Chain – The Hacker News](https://thehackernews.com/2026/06/litellm-vulnerability-chain-lets-low.html)
|
||||||
|
- [Red-Teaming the Agentic Red-Team (arXiv 2606.24496)](https://arxiv.org/pdf/2606.24496)
|
||||||
@@ -69,11 +69,12 @@ _DUMMY_HOST_KEY = (
|
|||||||
|
|
||||||
@skip_unless_docker()
|
@skip_unless_docker()
|
||||||
@unittest.skipIf(
|
@unittest.skipIf(
|
||||||
os.environ.get("GITEA_ACTIONS") == "true",
|
os.environ.get("GITEA_ACTIONS") == "true"
|
||||||
"skipped under act_runner: egress_tls_init uses a host bind mount "
|
and os.environ.get("BOT_BOTTLE_BACKEND") != "firecracker",
|
||||||
"the runner container can't see, and the network topology hides "
|
"skipped under act_runner unless BOT_BOTTLE_BACKEND=firecracker: "
|
||||||
"sibling-gateway visibility — same constraint as the other "
|
"egress_tls_init uses a host bind mount the runner container can't "
|
||||||
"bottle-bringup integration tests",
|
"see, and the network topology hides sibling-gateway visibility — "
|
||||||
|
"these constraints don't apply on the self-hosted KVM runner",
|
||||||
)
|
)
|
||||||
class TestSandboxEscape(unittest.TestCase):
|
class TestSandboxEscape(unittest.TestCase):
|
||||||
"""End-to-end attacks against a real bottle. The bottle stays
|
"""End-to-end attacks against a real bottle. The bottle stays
|
||||||
|
|||||||
Reference in New Issue
Block a user