docs(research): consolidate agent sandbox landscape

This commit is contained in:
2026-07-20 14:35:49 +00:00
parent 09debcf4f0
commit af1690ab22
10 changed files with 285 additions and 254 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0023: smolmachines bottle backend
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis
@@ -1,6 +1,6 @@
# PRD 0032: Decompose smolmachines launch and harden bringup sequencing
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-claude
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0038: smolmachines Env Contract and Secret-Safe Injection
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-codex
@@ -1,6 +1,6 @@
# PRD 0039: smolmachines Capability-Block Remediation
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-codex
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0042: smolmachines Cross-Backend Parity Tests
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis-codex
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0057: Promote smolmachines to default backend; convert Docker to example-only
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** didericis
+1 -1
View File
@@ -1,6 +1,6 @@
# PRD 0068: smolmachines backend on Linux
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
- **Status:** Superseded (2026-07-11) — was Active
- **Author:** Claude
+277 -64
View File
@@ -1,16 +1,16 @@
# Landscape: AI-agent sandbox tools
A broader survey than [`landscape-containerized-claude.md`](landscape-containerized-claude.md),
which focused on Claude-Code-specific containerizers. This one covers
general AI-agent sandbox / containment projects — some Claude-specific,
some agent-agnostic, some hosted SaaS — and contrasts them with
bot-bottle's design.
Survey of AI-agent sandbox and containment projects — including local
coding-agent wrappers, agent-agnostic runtimes, hosted platforms, and
governance layers — contrasted with bot-bottle's design. The original
Claude-Code-specific containerizer survey was folded into this note on
2026-07-20 so there is one landscape and one positioning verdict.
Research conducted 2026-05-11. CubeSandbox added 2026-07-18 (see its
per-project note and the addendum at the end). Also updated 2026-07-18:
bot-bottle no longer uses **pipelock** — outbound DLP is now bot-bottle's
own (deliberately simple) egress scanner (a mitmproxy addon with custom
detectors, PRD 0017 / 0053), and git-push secret scanning is handled by
detectors, PRD 0017 / 0052), and git-push secret scanning is handled by
**gitleaks** in the git-gate. "pipelock" below has been replaced with the
current mechanism; it survives only in older PRDs as history.
@@ -20,12 +20,24 @@ Passport); an **Agent-tailored policy** row added to the comparison table;
a separate Governance layers section added for AGT and OAP. See the
second addendum at the end.
Updated 2026-07-20: the borrowable-ideas status was reconciled with the
current implementation. In-flight credential injection and the microVM
backends have shipped, while per-use SSH confirmation was superseded by
keeping git credentials out of the agent entirely.
Also updated 2026-07-20: **E2B and Daytona added as first-class entries.**
Earlier revisions mentioned E2B only as the API and lifecycle model that
CubeSandbox implements, and omitted Daytona entirely. That was a survey gap,
not a principled scope exclusion: both are major hosted sandbox platforms and
belong in this landscape even though they target platform builders rather than
bot-bottle's local single-operator workflow.
## Summary
Fifteen projects surveyed across two categories: isolation/sandbox tools
and governance/pre-action authorization layers (the latter don't provide
VM or container isolation but do per-agent policy enforcement at the
tool-call level). None duplicate bot-bottle's combination of local
The main table compares bot-bottle against fifteen isolation/sandbox tools.
Governance/pre-action authorization and credential-only layers are covered
separately because they don't provide VM or container isolation. None
duplicate bot-bottle's combination of local
VM-per-bottle isolation, a declarative per-role manifest, per-agent
egress allowlist + outbound-content DLP, bottle/agent split, and the
composable `extends:` policy model. Three clusters stand out:
@@ -34,8 +46,9 @@ composable `extends:` policy model. Three clusters stand out:
single-user, thin wrappers over an existing OS primitive
(`sandbox-exec`, Podman + Landlock).
- **Different category (isolation)** — tilde.run (hosted SaaS), boxlite
and microsandbox (microVM libraries for platform builders), CubeSandbox
(self-hosted multi-tenant microVM service), endo-familiar
and microsandbox (microVM libraries for platform builders), E2B and Daytona
(hosted sandbox platforms), CubeSandbox (self-hosted multi-tenant microVM
service), endo-familiar
(capability-security paradigm, no OS isolation).
- **New: governance/pre-action layers** — Microsoft AGT and Open Agent
Passport (OAP): framework-embedded tool-call interceptors with
@@ -52,15 +65,17 @@ ergonomic enough that microVMs are **now bot-bottle's default backend**
only as a legacy fallback for CI / hosts without KVM or Apple Container.
That discussion has since shipped, not just been theorized.
**The one that matters most for positioning is CubeSandbox** — it is the
first surveyed project to ship bot-bottle's would-be wedge (default-deny
egress allowlist + full audit logs + in-flight credential custody so keys
never enter the sandbox) *combined with* per-sandbox microVM isolation,
**The one that matters most for positioning is CubeSandbox** — it ships
bot-bottle's bundle of default-deny egress allowlisting, full audit logs, and
in-flight credential custody *combined with* per-sandbox microVM isolation,
open-source under Apache 2.0, with Tencent Cloud behind it and 10.4k
stars. It's a self-hosted multi-tenant service for platform builders, not
a single-user declarative tool, so it doesn't collide head-on — but it
narrows the "nobody else bundles egress custody + credential injection"
claim that the monetization positioning leans on. See the addendum.
claim that the monetization positioning leans on. Daytona now also offers
domain/CIDR firewall policy plus in-flight header credential substitution and
response scrubbing, although its higher tiers are not default-deny and its
production platform is proprietary. See the addendum.
## Per-project notes
@@ -200,6 +215,80 @@ claim that the monetization positioning leans on. See the addendum.
also supported.
- **Maturity**: Active through April 2026.
### E2B *(added 2026-07-20)*
- **Source**: https://github.com/e2b-dev/e2b ; https://e2b.dev/docs
- **License**: Apache 2.0 (~12.4k stars); commercial hosted service with
self-hosting/BYOC support.
- **Isolation**: Firecracker microVM per sandbox.
- **Locality**: Cloud-hosted by default; self-hosting uses Terraform on AWS or
GCP (with other targets documented as works in progress).
- **Agent integration**: LLM-agnostic Python and JavaScript/TypeScript SDKs;
code-interpreter and desktop-sandbox products. Platform primitive rather
than a coding-agent wrapper.
- **Config**: Programmatic SDK/API plus templates. Network configuration
supports internet on/off, outbound allow/deny rules, and a custom egress
proxy.
- **Network policy**: Configurable per sandbox, but not documented as
default-deny and no built-in outbound-content DLP is documented.
- **Credentials**: Environment variables passed to the sandbox are explicitly
not private at the OS level. No built-in in-flight application-credential
injection is documented.
- **Persistence**: Full memory + filesystem pause/resume, snapshots, and
auto-resume. Continuous runtime is tier-limited, while paused sandboxes are
retained indefinitely.
- **Maturity**: Established hosted platform and the API compatibility target
used by CubeSandbox.
### Daytona *(added 2026-07-20)*
- **Source**: https://github.com/daytonaio/daytona ;
https://www.daytona.io/docs/
- **License**: Current production platform is proprietary. The former AGPL
repository remains public but is no longer maintained after Daytona moved
production development closed-source in June 2026.
- **Isolation**: Hosted container sandboxes by default, with separate Linux
and Windows VM sandbox classes for dedicated-OS workloads. Each sandbox has
its own filesystem and network stack; VM-only features include memory
pause/resume and forking.
- **Locality**: Hosted multi-tenant service, with dedicated/custom regions and
customer runners available.
- **Agent integration**: LLM/framework-agnostic SDKs (Python, TypeScript, Go,
Ruby, Java), API, and CLI; official agent-framework guides. Platform
primitive rather than a local coding-agent wrapper.
- **Config**: Programmatic per-sandbox image/snapshot, resources, lifecycle,
firewall, and secrets.
- **Network policy**: Per-sandbox IPv4/domain allowlists and block-all mode,
subordinate to organization/tier policy. Full internet access is the
default on higher tiers, so it is configurable rather than uniformly
default-deny.
- **Credentials**: First-class secret manager with the same phantom-token
pattern as bot-bottle: the sandbox environment gets an opaque placeholder,
an HTTPS proxy substitutes the real secret in headers only for allowed
hosts, and responses are scrubbed back to the placeholder.
- **Persistence**: Persistent filesystem for stopped container sandboxes;
memory + filesystem pause/resume for VM sandboxes; snapshots and configurable
auto-stop.
- **Maturity**: Production commercial platform. Notable April 2026 credential
exposure was patched; the June 2026 closed-source transition materially
changes its transparency/self-hosting posture.
### Other hosted runtimes carried forward from the earlier survey
- **Northflank Sandboxes** — hosted or customer-cloud, microVM-backed
containers with SDK-managed lifecycle, optional persistent volumes, and
sub-second claimed boot. This is a platform primitive for untrusted code and
agents, not a local agent wrapper or role-policy layer.
- **Cloudflare Sandbox SDK** — Workers/Durable Objects API over VM-isolated
Linux containers for command, file, process, and service execution. It is a
hosted TypeScript platform primitive; application authentication,
authorization, and credential-proxy patterns remain the integrator's job.
Both belong to the same “build your agent platform on this runtime” category as
E2B and Daytona. They were named but not analyzed in depth by the original
Claude-specific note, so they remain outside the main comparison table rather
than being presented with false precision.
### CubeSandbox *(added 2026-07-18)*
- **Source**: https://github.com/TencentCloud/CubeSandbox ;
HN launch https://news.ycombinator.com/item?id=47863430
@@ -316,6 +405,92 @@ claim that the monetization positioning leans on. See the addendum.
preview — APIs may change.
- **Maturity**: Early research preview.
## Claude-specific wrappers and developer environments
These projects were the focus of the original containerized-Claude survey.
They remain useful comparisons for local developer experience, but most are
templates or wrappers rather than policy-bearing sandbox platforms, so they
are grouped here instead of widening the main table further.
### claudebox
- **Source**: https://github.com/RchGrav/claudebox
- **Isolation**: Docker, with per-project images, authentication state, and
configuration.
- **Agent integration**: Claude Code wrapper with 15+ preconfigured language
and task profiles.
- **Network policy**: Per-project firewall allowlists.
- **Closest overlap**: local one-command developer workflow and project-scoped
network policy.
- **Difference**: profiles describe development toolchains, not named agent
roles. There is no bottle/agent split, composable role manifest, provider
plugin layer, or outbound-content DLP.
### Spritz / claude-code-sandbox
- **Source**: https://github.com/textcortex/claude-code-sandbox (archived;
points to its successor, Spritz).
- **Isolation**: The original project ran Claude Code in local Docker with
bypass permissions; Spritz moved toward Kubernetes-native multi-agent
infrastructure.
- **Difference**: the successor targets cluster orchestration rather than a
low-dependency local launcher. It is architecturally closer to hosted or
Kubernetes platform runtimes than to bot-bottle's single-operator CLI.
### Trail of Bits claude-code-devcontainer
- **Source**: https://github.com/trailofbits/claude-code-devcontainer
- **Isolation**: A Docker devcontainer that exposes only project files and is
designed to run Claude Code with `bypassPermissions` for security audits and
untrusted-code review.
- **Difference**: a hardened, reusable environment definition rather than an
agent launcher or fleet. It has no named-role manifest, per-role credential
custody, supervision plane, or multi-backend abstraction.
### Smaller wrappers and official templates
Projects such as `arezi/claude-sandbox`, `nkrefman/claude-sandbox`, and
`VishalJ99/claude-docker`, plus Docker/Anthropic devcontainer templates, prove
there is steady demand for “Claude in a container.” They are deliberately
small launch/build configurations. They compete on setup simplicity, not on
role-aware policy, credential custody, persistent supervision, or a fleet
model, and are better treated as a product category than as individual rows.
### SuperHQ
- **Source**: https://superhq.ai/
- **Isolation**: Apple-Silicon desktop application using local microVMs via
Virtualization.framework/libkrun-era components.
- **Agent integration**: Claude Code, Codex, and Pi in a GUI, with mobile
remote access.
- **Credentials and review**: host-side auth gateway injects credentials on
the wire; a temporary overlay stages writes for diff-and-accept review.
- **Closest overlap**: local microVM isolation, multi-provider launching, and
credential custody for security-minded individual developers.
- **Difference**: GUI desktop product on Apple Silicon rather than a
cross-platform declarative CLI/fleet layer. The July 2026 snapshot in the
original survey recorded a user request for per-run tool-call and network
audit logging; treat that as point-in-time rather than a permanent gap.
## Credential gateway without isolation
### OneCLI
[OneCLI](https://onecli.sh/) is a framework-agnostic identity gateway rather
than a sandbox. Its phantom-token design gives the agent a placeholder and
substitutes the encrypted real credential at the network layer. It therefore
matches bot-bottle closely on secret custody, and is more portable because it
can sit in front of agents launched by anything, but it supplies no container
or VM boundary, filesystem isolation, role manifest, or egress-content DLP.
The positioning consequence from the earlier survey still holds: secret
custody alone is not unique. bot-bottle's relevant combination is local
isolation + default-deny egress + payload DLP + declarative roles + credential
custody. OneCLI's managed tier also places custody with a third party, whereas
bot-bottle keeps it within operator-controlled infrastructure. See
[`agent-credential-proxy-landscape.md`](agent-credential-proxy-landscape.md)
for the detailed build-versus-adopt analysis.
## Governance / pre-action authorization layers
These two tools don't provide VM or filesystem isolation; they intercept
@@ -371,19 +546,19 @@ them.
*Isolation/sandbox tools only. AGT and OAP are governance layers — see their per-project notes above.*
| Axis | bot-bottle | endo-familiar | litterbox | agent-safehouse | matchlock | tilde.run | boxlite | microsandbox | smolmachines | CubeSandbox | Cleanroom | container-use | Docker sbx | Anthropic srt |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Isolation | MicroVM per bottle default (Firecracker/KVM on Linux, Apple Container on macOS) + own egress DLP scanner; Docker legacy fallback, gVisor there if present | Object-capability (no OS isolation) | Podman + opt. Landlock | macOS `sandbox-exec` | MicroVM (Firecracker / Virt.fw) | Hosted container (unverified) | MicroVM (KVM / Hypervisor.fw) | MicroVM (libkrun) | MicroVM (libkrun / KVM) | MicroVM (RustVMM / KVM) | MicroVM (Firecracker / Virt.fw) | Docker container + git worktree | MicroVM (proprietary) | OS-level (Seatbelt / bubblewrap / WFP) — no container |
| Local vs hosted | Local | Local | Local (Linux) | Local (macOS) | Local | Hosted SaaS | Local | Local | Local | Self-hosted (server/cluster) | Self-hosted server | Local | Local | Local |
| Open source | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | MIT | No | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Proprietary | Apache 2.0 (experimental) |
| Agent target | Claude Code | Generic (demo) | Generic | Multi-agent wrapper | Generic (+ Claude/OpenAI SDKs) | Claude focus | Generic | Claude + Cursor (MCP/Skills) | Generic (AGENTS.md) | E2B-compatible (platform builders) | CI / generic process | Claude Code, Cursor, Windsurf (MCP) | Claude Code, Codex, Gemini CLI, Copilot, Kiro | Claude Code (and any process) |
| Network policy | Default-deny via own egress scanner + per-bottle allowlist + content DLP + gitleaks on git push | Capability model only | Limited | Not addressed | Default-deny + allowlist + secret-injecting proxy | Default-deny + logging | Per-VM net (unverified) | Not documented | Off by default + allowlist | Default-deny allowlist + instant egress block + audit logs + per-sandbox tokens (eBPF) + credential vault | Default-deny + per-repo host allowlist (cleanroom.yaml) | Not addressed | Default-deny; Open / Balanced / Locked Down presets; live TUI network panel | Proxy-based allowlist/denylist (HTTP + SOCKS5); custom proxy supported |
| Parallel agents | Yes (one bottle per agent) | n/a | Not addressed | One at a time | Multiple VMs | Yes (dashboard) | SDK-level | SDK-level | Architectural | Yes (2,000+/host claimed) | Yes (server model) | Yes (per-agent containers + worktrees) | Yes | Yes |
| Long-running posture | Persistent by default (named, supervised) | n/a (demo) | Session (up while in use) | Per-invocation | Ephemeral VM per run | Per-run (versioned) | Ephemeral + snapshot/fork | Ephemeral / on-demand | Named persistent by default | Ephemeral + auto pause/resume | Per-run + suspend/resume | Per-agent container (ephemeral) | Per-session; branch mode creates git worktree in .sbx/ | Per-invocation |
| DX: run Claude yolo-style | One command → interactive yolo Claude (`start <agent>`, `--dangerously-skip-permissions` default) | n/a (lib demo) | Wizard + build, then run claude inside (Linux only) | One-command wrapper (`safehouse claude --dangerously-skip-permissions`) | CLI: run a cmd in a VM (not a Claude wrapper) | Hosted (`tilde exec`), not local-native | SDK code required (build the run yourself) | CLI/MCP: sandbox-as-a-tool for the agent, not a wrapper around it | SSH into a named machine, run claude there | Stand up a cluster + drive via E2B SDK | CI-oriented, not a Claude wrapper | MCP server: `claude mcp add container-use -- container-use stdio` | One command: `sbx` wraps claude with `--dangerously-skip-permissions` default | Library/wrapper, not a standalone CLI |
| Config | JSON manifest (bottles + agents) | Programmatic refs | CLI wizard | Profile files / shell fns | CLI / SDK | DSL + CLI + SDK | SDK | CLI / SDK / MCP | TOML Smolfile | E2B-compatible SDK | cleanroom.yaml in repo | None (no policy config) | Preset levels at launch | Programmatic per-invocation (allow/deny lists) |
| Agent-tailored policy | Yes — bottle/agent split; declarative per-role egress + credentials; composable via `extends:` | Partial — capability model scopes per-agent, but no declarative role manifest | No | Partial — per-agent profile files (Seatbelt); no egress | No | Yes — per-agent DSL RBAC (allow/deny/approve per action/repo/agent) | No | No | No | No — per-sandbox SDK config, not role-scoped | Partial — per-repo cleanroom.yaml, not per-role | No | No — network presets only | No |
| Maturity | Active July 2026 | Research (2022+) | Early (~66 ⭐) | Active (~1.8k ⭐) | Experimental (~574 ⭐) | Private preview | YC, ~4.7k ⭐ | YC, ~6k ⭐, beta | ~3.1k ⭐ | Tencent, prod, ~10.4k ⭐ | Active (Buildkite product) | Early development | GA 2026 | Early research preview |
| Axis | bot-bottle | endo-familiar | litterbox | agent-safehouse | matchlock | tilde.run | boxlite | microsandbox | smolmachines | E2B | Daytona | CubeSandbox | Cleanroom | container-use | Docker sbx | Anthropic srt |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Isolation | MicroVM per bottle default (Firecracker/KVM on Linux, Apple Container on macOS) + own egress DLP scanner; Docker legacy fallback, gVisor there if present | Object-capability (no OS isolation) | Podman + opt. Landlock | macOS `sandbox-exec` | MicroVM (Firecracker / Virt.fw) | Hosted container (unverified) | MicroVM (KVM / Hypervisor.fw) | MicroVM (libkrun) | MicroVM (libkrun / KVM) | Firecracker microVM | Container or Linux/Windows VM class | MicroVM (RustVMM / KVM) | MicroVM (Firecracker / Virt.fw) | Docker container + git worktree | MicroVM (proprietary) | OS-level (Seatbelt / bubblewrap / WFP) — no container |
| Local vs hosted | Local | Local | Local (Linux) | Local (macOS) | Local | Hosted SaaS | Local | Local | Local | Hosted; self-host/BYOC available | Hosted; dedicated/custom regions | Self-hosted (server/cluster) | Self-hosted server | Local | Local | Local |
| Open source | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | MIT | No | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Production closed-source; legacy AGPL repo unmaintained | Apache 2.0 | Apache 2.0 | Apache 2.0 | Proprietary | Apache 2.0 (experimental) |
| Agent target | Claude Code, Codex, Pi, and provider plugins | Generic (demo) | Generic | Multi-agent wrapper | Generic (+ Claude/OpenAI SDKs) | Claude focus | Generic | Claude + Cursor (MCP/Skills) | Generic (AGENTS.md) | LLM-agnostic platform builders | LLM-agnostic platform builders | E2B-compatible (platform builders) | CI / generic process | Claude Code, Cursor, Windsurf (MCP) | Claude Code, Codex, Gemini CLI, Copilot, Kiro | Claude Code (and any process) |
| Network policy | Default-deny via own egress scanner + per-bottle allowlist + content DLP + gitleaks on git push | Capability model only | Limited | Not addressed | Default-deny + allowlist + secret-injecting proxy | Default-deny + logging | Per-VM net (unverified) | Not documented | Off by default + allowlist | Per-sandbox allow/deny rules and custom egress proxy; internet configurable | Per-sandbox CIDR/domain allowlist or block-all; tier policy; secret-injecting proxy | Default-deny allowlist + instant egress block + audit logs + per-sandbox tokens (eBPF) + credential vault | Default-deny + per-repo host allowlist (cleanroom.yaml) | Not addressed | Default-deny; Open / Balanced / Locked Down presets; live TUI network panel | Proxy-based allowlist/denylist (HTTP + SOCKS5); custom proxy supported |
| Parallel agents | Yes (one bottle per agent) | n/a | Not addressed | One at a time | Multiple VMs | Yes (dashboard) | SDK-level | SDK-level | Architectural | Yes (platform service) | Yes (platform service) | Yes (2,000+/host claimed) | Yes (server model) | Yes (per-agent containers + worktrees) | Yes | Yes |
| Long-running posture | Persistent by default (named, supervised) | n/a (demo) | Session (up while in use) | Per-invocation | Ephemeral VM per run | Per-run (versioned) | Ephemeral + snapshot/fork | Ephemeral / on-demand | Named persistent by default | Runtime tier limits + indefinite pause/resume | Persistent filesystem; VM pause/resume; configurable auto-stop | Ephemeral + auto pause/resume | Per-run + suspend/resume | Per-agent container (ephemeral) | Per-session; branch mode creates git worktree in .sbx/ | Per-invocation |
| DX: run Claude yolo-style | One command → interactive yolo Claude (`start <agent>`, `--dangerously-skip-permissions` default) | n/a (lib demo) | Wizard + build, then run claude inside (Linux only) | One-command wrapper (`safehouse claude --dangerously-skip-permissions`) | CLI: run a cmd in a VM (not a Claude wrapper) | Hosted (`tilde exec`), not local-native | SDK code required (build the run yourself) | CLI/MCP: sandbox-as-a-tool for the agent, not a wrapper around it | SSH into a named machine, run claude there | SDK/CLI sandbox; wire the agent yourself | SDK/CLI sandbox; wire the agent yourself | Stand up a cluster + drive via E2B SDK | CI-oriented, not a Claude wrapper | MCP server: `claude mcp add container-use -- container-use stdio` | One command: `sbx` wraps claude with `--dangerously-skip-permissions` default | Library/wrapper, not a standalone CLI |
| Config | YAML-in-Markdown manifests (bottles + agents) | Programmatic refs | CLI wizard | Profile files / shell fns | CLI / SDK | DSL + CLI + SDK | SDK | CLI / SDK / MCP | TOML Smolfile | SDK/API + templates | SDK/API/CLI + images/snapshots | E2B-compatible SDK | cleanroom.yaml in repo | None (no policy config) | Preset levels at launch | Programmatic per-invocation (allow/deny lists) |
| Agent-tailored policy | Yes — bottle/agent split; declarative per-role egress + credentials; composable via `extends:` | Partial — capability model scopes per-agent, but no declarative role manifest | No | Partial — per-agent profile files (Seatbelt); no egress | No | Yes — per-agent DSL RBAC (allow/deny/approve per action/repo/agent) | No | No | No | No — per-sandbox SDK config | No — per-sandbox SDK config | No — per-sandbox SDK config, not role-scoped | Partial — per-repo cleanroom.yaml, not per-role | No | No — network presets only | No |
| Maturity | Active July 2026 | Research (2022+) | Early (~66 ⭐) | Active (~1.8k ⭐) | Experimental (~574 ⭐) | Private preview | YC, ~4.7k ⭐ | YC, ~6k ⭐, beta | ~3.1k ⭐ | Established hosted platform, ~12.4k ⭐ | Production commercial; closed-source since June 2026 | Tencent, prod, ~10.4k ⭐ | Active (Buildkite product) | Early development | GA 2026 | Early research preview |
## What's closest, what's different
@@ -433,47 +608,81 @@ alternative.
**Solving a different problem.** tilde.run is hosted SaaS for team /
production agent pipelines with data-versioned rollback — explicitly
opposite to bot-bottle's "infrastructure I control" goal. boxlite,
microsandbox, and CubeSandbox are infrastructure libraries/services aimed
at platform builders embedding sandboxes into agent frameworks; they
opposite to bot-bottle's "infrastructure I control" goal. E2B and Daytona
are hosted sandbox platforms, while boxlite, microsandbox, and CubeSandbox
are infrastructure libraries/services aimed at platform builders embedding
sandboxes into agent frameworks; they
would be a *backend* bot-bottle could call, not a competitor to its
manifest layer. endo-familiar is in a different paradigm entirely:
capability passing rather than kernel boundaries.
## Borrowable ideas
What bot-bottle already has that the survey suggested as
differentiators:
### Already shipped or otherwise addressed
- Default-deny egress with a per-agent allowlist (own egress scanner).
- DLP scanning of outbound traffic.
- Bottle / agent split (manifest layer above the isolation primitive).
- gVisor auto-detection on Linux.
- **In-flight secret injection** (suggested by matchlock) — **shipped.**
Real provider and git-host tokens are held outside the agent and injected
by the egress gateway on matching routes. The agent receives only proxy
URLs and, where a client requires a credential-shaped value, a placeholder;
`GITEA_TOKEN` and equivalent real tokens do not appear in the agent's
environment.
- **MicroVM backend** — **shipped.** MicroVMs are now the default:
Firecracker on KVM Linux and Apple Container on macOS. Docker is the legacy
fallback.
- **Per-use SSH key confirmation** (suggested by litterbox) — **addressed by
stronger credential custody instead.** The agent does not hold the upstream
git SSH key or an SSH-agent socket: git-gate holds the credential and gates
git operations. A confirmation wrapper inside the agent would therefore
protect a credential that is no longer there. Operator approval at the gate
remains the appropriate control point for any future per-use confirmation.
Ideas worth considering, without abandoning the Python-stdlib-first /
local, single-operator stance:
### Still worth considering
1. **Per-use SSH key confirmation** (from litterbox). Even with
KnownHostKey pinning and the egress DLP scanner, a wrapper SSH agent that
prompts on each key use (e.g. via `osascript` / `notify-send`) would
catch an agent doing something off-policy with a key it legitimately
holds. Pure-stdlib, no new deps.
2. **In-flight secret injection** (from matchlock). The egress scanner
already does allowlisting and DLP; teaching it to *inject* tokens at
proxy time so e.g. `GITEA_TOKEN` never appears in the container's
env would close the "agent reads its own env and exfiltrates" path.
Fits the existing egress-proxy architecture.
3. **MicroVM backend**~~on the radar~~ **shipped since this survey.**
microVMs are now bot-bottle's default (Firecracker on KVM Linux, Apple
Container on macOS); Docker is the legacy fallback. The libkrun / Apple
Virtualization.framework ergonomics that microsandbox, smolmachines,
and matchlock demonstrated turned out to be enough to make it the
default rather than an opt-in.
- **Live network activity in the supervisor TUI** (from Docker sbx): show
allowed and blocked connections and let the operator propose policy changes
from the existing supervision surface.
- **Tamper-evident audit records** (from OAP): sign and hash-chain egress and
supervision decisions for compliance-sensitive deployments.
- **Behaviour-informed policy downgrade** (from Microsoft AGT): use repeated
DLP alerts or supervision holds as a signal to narrow policy or request
closer review. This needs a carefully specified trust model before it can be
more than a heuristic.
Not worth borrowing: the SDK-first programmatic API style of boxlite /
microsandbox (cuts against the declarative-manifest stance), and the
hosted-SaaS dashboard model of tilde.run (cuts against the
"infrastructure I control" goal).
## Publishing and positioning verdict
Publishing remains worthwhile, but the defensible claim is the combination,
not any single primitive. Credential custody is matched by OneCLI, matchlock,
Daytona, Docker sbx, and CubeSandbox; local one-command isolation is matched by
agent-safehouse and Docker sbx; hosted microVM execution is a crowded platform
category.
bot-bottle remains unusual in combining:
- local, operator-controlled execution with persistent named bottles;
- one declarative role layer across Claude Code, Codex, Pi, and provider
plugins;
- composable agent/bottle manifests, skills, and system prompts;
- Firecracker/Apple Container isolation with a Docker fallback;
- default-deny per-role egress, payload DLP, and git-push secret scanning;
- credentials injected outside the agent process; and
- supervision and audit state suited to long-running parallel agents.
The practical wedge is “as easy as native yolo, with declarative role policy
and self-hosted custody,” including scoped access to private LAN/Tailnet
services that cloud-first runtimes cannot provide without additional network
plumbing. The main competitive risks are a local wrapper such as claudebox or
Docker sbx growing a role-manifest layer, and GUI products such as SuperHQ
adding equivalent policy and audit depth.
## Caveats
- Star counts and last-commit dates are point-in-time snapshots.
@@ -490,7 +699,8 @@ hosted-SaaS dashboard model of tilde.run (cuts against the
CubeSandbox (Tencent Cloud, Apache 2.0, ~10.4k stars, HN launch
[#47863430](https://news.ycombinator.com/item?id=47863430)) is the first
project in this survey to combine, in one open-source stack, everything
open-source, self-hostable project in this survey to combine, in one stack,
the main primitives
bot-bottle treated as its differentiator:
- **Egress custody (connection level)** — default-deny domain allowlist
@@ -558,7 +768,7 @@ instead of per-action prompts. On this axis the field splits cleanly:
network egress; bot-bottle adds VM-grade isolation, egress DLP, and
persistent/parallel bottles across macOS + Linux.
- **Libraries / services** (you build the run yourself): boxlite,
microsandbox, CubeSandbox, E2B. These hand you an SDK or a cluster and
microsandbox, CubeSandbox, E2B, Daytona. These hand you an SDK or a cluster and
expect you to wire the agent in — powerful for platform builders,
heavyweight for "just run Claude on my laptop." microsandbox's MCP/Skills
angle is *sandbox-as-a-tool the agent calls*, which is the inverse of
@@ -566,10 +776,11 @@ instead of per-action prompts. On this axis the field splits cleanly:
- **In between:** litterbox (wizard + build, Linux only), smolmachines
(SSH into a named machine), matchlock (run a command in a VM).
So DX is a genuine bot-bottle differentiator, and the only project that
matches it (agent-safehouse) does so with materially weaker isolation and
no egress story. "As easy as native yolo, but actually sandboxed" is a
defensible one-liner.
So DX is a genuine bot-bottle differentiator. agent-safehouse matches the
one-command wrapper with weaker isolation and no egress story; Docker sbx now
matches it at microVM strength but remains proprietary and preset-based. "As
easy as native yolo, with declarative role policy" is the narrower defensible
one-liner.
Why it still doesn't collide head-on:
@@ -577,7 +788,8 @@ Why it still doesn't collide head-on:
builders* (drop-in E2B replacement, SDK-driven, 2,000 sandboxes on a
box). bot-bottle is a *single-operator, declarative-manifest tool for
the infrastructure I run*. Different buyer, different ergonomics — no
JSON manifest, no bottle/agent split, no "one command on my laptop."
declarative role manifest, no bottle/agent split, no "one command on my
laptop."
2. **Backend, not competitor.** Like boxlite/microsandbox, CubeSandbox is
something bot-bottle could sit *on top of* — a `"runtime": "microvm"`
or `"runtime": "cubesandbox"` backend under the manifest layer — while
@@ -588,7 +800,8 @@ Why it matters anyway:
- The "nobody else bundles connection-level egress allowlist + audit +
in-flight credential custody" line is **no longer true for the
primitive** — a well-funded, 10k-star open-source project now ships it.
primitive** — CubeSandbox ships the open-source/self-hosted combination,
and Daytona ships a proprietary firewall + credential-substitution variant.
But **content DLP on authorized channels is still not matched** (see
above), and neither is the *layer above* the primitive (declarative
manifest, cross-vendor orchestration, operator UX, the
@@ -599,8 +812,8 @@ Why it matters anyway:
that in mind.
- Worth a closer look at **how** CubeSandbox does credential injection
and per-sandbox egress tokens (eBPF virtual switch vs. bot-bottle's
mitmproxy egress proxy) before the next iteration of bot-bottle's
in-flight-secret feature — see borrowable idea #2 above.
mitmproxy egress proxy) when hardening bot-bottle's now-shipped
credential-custody implementation.
## Addendum 2026-07-18 (second pass) — agent-tailored policy landscape
@@ -639,7 +852,7 @@ whether a permitted action is consistent with the agent's actual task.
See `hn-agent-safety-discourse-july-2026.md` for the blast-radius
analysis.
**Borrowable from new tools:**
**Open ideas from new tools (also summarized above):**
- **Microsoft AGT's trust-score decay** — privilege that reflects
observed behaviour rather than static provisioning. Applied to
@@ -1,182 +0,0 @@
# Landscape: containerized AI coding agent tools
Research into whether bot-bottle is redundant with existing projects, and
whether it's worth publishing.
## Summary
The "AI coding agents in isolated sandboxes" space is active but not saturated.
bot-bottle occupies a distinct position: no surveyed project combines all five
of its defining features. Publishing is likely worthwhile, with the main risk
being claudebox expanding to absorb the same niche.
**Updated 2026-07-09:** bot-bottle now supports three isolation backends
(Docker, Apple `container`, smolmachines/libkrun microVMs) and three built-in
agent providers (Claude Code, OpenAI Codex, Pi) with an open plugin system for
arbitrary providers. This meaningfully strengthens the differentiation against
all surveyed competitors.
## Closest competitor: claudebox
[RchGrav/claudebox](https://github.com/RchGrav/claudebox) is the most
feature-complete analog. It runs Claude Code in Docker with per-project
isolated images, 15+ pre-configured dev-language profiles, and per-project
network firewall allowlists. Actively maintained with multiple forks.
What it lacks: manifest-driven named agents, per-agent env resolution modes
(prompt / host-forward / literal), skill directory injection, per-agent system
prompts, SSH-agent forwarding without copying private keys, home+project
manifest merge.
## Other surveyed projects
- **textcortex/claude-code-sandbox → spritz** — evolved toward
Kubernetes-native multi-agent infra; not stdlib-first or local-Docker.
Original sandbox repo is archived.
- **trailofbits/claude-code-devcontainer** — devcontainer config for security
audits; not a general agent launcher.
- **Several small solo repos** (arezi/claude-sandbox, nkrefman/claude-sandbox,
VishalJ99/claude-docker) — lightweight Docker wrappers with no multi-agent
config layer.
- **Docker's official sandbox templates** — launch-and-run Dockerfiles plus an
npm-based runtime; not a manifest-driven fleet manager.
## Adjacent (different model)
- **dagger/container-use** (mid-2025) — exposes an MCP server so the *agent*
spins up its own containers with Git worktrees. Inverted model vs. bot-bottle
(agent controls container rather than being launched into one by a manifest).
Still marked early-development.
- **E2B, Northflank, Cloudflare Sandbox SDK** — cloud-hosted SaaS sandbox
runtimes; fundamentally different architecture.
- **superhq.ai / SuperHQ** (v0.4.4, April 2026) — macOS desktop app (Rust/GPUI)
that runs Claude Code, Codex, and Pi inside microVMs via Apple's
Virtualization.framework (their own shuru-sdk / libkrun). Auth gateway
injects API keys on the wire so the sandbox never sees them; tmpfs overlay
stages agent writes for diff-and-accept review; mobile remote access via
remote.superhq.ai. Early alpha, free on launch, Apple Silicon only.
Overlap: both projects cover agent isolation, credential proxying, and
multi-provider support (Claude Code / Codex / Pi). Differences: SuperHQ is a
GUI desktop app with no manifest layer; bot-bottle is a CLI fleet manager with
named agents, skills injection, per-agent system prompts, and cross-platform
backends (Docker, Apple `container`, smolmachines). SuperHQ's microVM
isolation story is now partially matched by bot-bottle's `macos_container` and
smolmachines backends. Worth watching — it targets the same security-minded
power-user audience and moves fast.
**Known gap in SuperHQ (user-requested, as of 2026-07-09):** A named user
(Brian Cheong, Founder, Dunialabs.io) explicitly called out the absence of
per-run audit logging: tool calls and network egress. Bot-bottle covers both:
network egress is logged by pipelock/mitmproxy, and per-run op-log/audit state
is persisted to SQLite.
- **OneCLI** ([onecli.sh](https://onecli.sh/)) — YC-backed, GA, open-source
(Apache-2.0, Rust) "identity gateway for AI agents": a credential/secret
broker that holds API keys and OAuth tokens out of the agent's reach and
injects them at the network layer (phantom-token — the agent sees a
placeholder, the gateway swaps in the real, AES-256-GCM-encrypted credential
at request time). Framework-agnostic and drop-in for any HTTP-calling agent,
50+ app integrations, plus a hosted cloud tier with a per-agent dashboard and
audit logs. Full technical breakdown in
[`agent-credential-proxy-landscape.md`](agent-credential-proxy-landscape.md).
**How close a competitor:** near-exact on the *single axis of agent secret
custody* — the exact thing bot-bottle sells as "the agent never sees real
credentials, even via `printenv`." OneCLI does that one job well, is mature
and funded, and is *more portable* (it sits in front of anything; bot-bottle
only helps agents launched through bot-bottle). Takeaway: bot-bottle should
stop treating secret custody as a *unique* differentiator. But OneCLI is
**not** a competitor to bot-bottle's actual product — it does no agent
sandboxing (containers/microVMs), no fleet/manifest layer, no named agents /
skills / per-agent system prompts, no multi-provider launching, no egress
firewall.
**Our edge:** (1) *Isolation is the product, not a proxy.* OneCLI keeps the
key out of reach at the network layer, but the agent itself still runs
unsandboxed — a hijacked agent behind OneCLI has full run of its host and can
exfil captured data through any allowed host. bot-bottle runs the agent inside
a kernel/VM-enforced sandbox, injects credentials across that same
out-of-process boundary, *and* clamps egress with pipelock — defense in depth
vs. a single network layer. (2) *Fleet + manifest model* with named agents,
skills, per-agent system prompts, multi-provider and multi-backend — OneCLI
has no equivalent. (3) *Trust posture:* OneCLI's managed tier reintroduces a
third-party credential custodian, whereas bot-bottle's OSS-runtime +
paid-control-plane split keeps custody inside the operator's own boundary —
the stronger story for the security-minded self-hoster. (4) *Runs inside your
network boundary — local/internal reach.* Because bot-bottle executes the
agent on your own host (homelab, corporate LAN, a Tailnet) and egress is a
manifest field, giving an agent *scoped* access to **internal** resources — a
private Gitea, a LAN database, a Tailscale node — is just another egress-route
line, not a networking project (the same move an operator already makes to
reach their Tailscale services). OneCLI's OSS core can self-host too, but it's
a credential *broker* for outbound API calls, not an agent runtime, and its
managed tier + 50+ integrations are oriented at public SaaS — it doesn't put
the agent behind your firewall for you. This is a reach advantage, distinct
from the isolation ones above, and it's a wedge cloud-first agent products
(Devin, Copilot Workspace, OneCLI Cloud) structurally can't match. **Tactical
read:**
adopt OneCLI's OSS core for the credential slice if building is undesirable
(it's mature now); don't build atop its managed tier (competitor, not
dependency); re-position bot-bottle on isolation + fleet + self-hosted custody
rather than "we hide your secrets."
## What no found project does
None combine:
1. Named-agent manifest with per-agent env resolution (prompt / host-forward / literal), supporting multiple providers (Claude Code, Codex, Pi, arbitrary plugins)
2. Skills directory injection
3. Per-agent system prompts
4. SSH-agent key forwarding without copying private keys into the container
5. Home + project manifest merge
6. Pluggable isolation backends: Docker (Linux/macOS), Apple `container` (macOS microVMs), smolmachines/libkrun microVMs
7. Per-run audit log: network egress via pipelock/mitmproxy + op-log persisted to SQLite
**In-flight directions (not yet shipped):**
- **Forge-native dispatch (issue #317):** Gitea webhook → orchestrator spins up a bottle
with the issue body as prompt → agent works → bottle freezes awaiting review comment →
rehydrates on comment → tears down on PR close. The issue-to-PR lifecycle concept is not
novel (Devin, Copilot Workspace, SWE-agent all do this as cloud services); what's
distinct is doing it self-hosted, manifest-driven, inside bot-bottle's isolation
primitives.
- **Paid web control plane (issue #327):** Browser-based multi-host agent launch and
monitoring; account-scoped bottle and agent definitions; secret custody (encrypted at
rest, injected into the sidecar at launch, never exposed to the agent or returned by any
read API). Monetization model: OSS runtime free, control plane paid — a standard split
(HashiCorp, Grafana) applied to a self-hosted agent sandbox. The principled secret
custody model (agent never sees real credentials, even via printenv) is more rigorous
than most surveyed tools but not unprecedented.
## Publishing verdict
Worth publishing. Differentiators that matter to the target audience (power
users running parallel AI coding agent sessions with distinct personas/tooling):
- The Python-stdlib-first, low-dependency design — competitors are npm-based,
Rust/GUI, or Kubernetes-native.
- Named agents with distinct skills and system prompts, not just language profiles.
- Multi-backend isolation: Docker, Apple `container` microVMs, and
smolmachines/libkrun — single manifest works across all three.
- Multi-provider: Claude Code, Codex, Pi, plus an open plugin system for
arbitrary providers.
- SSH forwarding without key copying.
- Per-run audit log (tool calls + network egress) — an explicitly requested gap
in SuperHQ as of 2026-07-09.
- Forge-native dispatch and a paid control plane (in flight) bring bot-bottle
into the same product category as cloud services like Devin and Copilot
Workspace — but self-hosted, with stronger isolation guarantees and a
manifest-driven fleet model those services don't have.
Main risk: claudebox adds manifest/agent config; SuperHQ is moving fast on the
GUI / microVM side. The space is moving fast enough that publishing sooner is
better if establishing prior art matters.
Discovery will be slow without active promotion; an Anthropic Discord post or
HN "Show HN" would do most of the work.
## Caveats
- GitHub search cannot surface private or very new repos comprehensively.
- Counts (stars, forks) were not confirmed for every project.
- Initial research conducted 2026-05-07; SuperHQ entry added 2026-07-09; the space moves fast.
+1 -1
View File
@@ -1,6 +1,6 @@
# smolmachines as a VM backend for bot-bottle
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/landscape-containerized-claude.md`.
> **Superseded (2026-07-11).** The smolmachines backend was removed — Linux now uses the Firecracker backend, macOS uses macos-container. Kept as a historical record; see the removal commit `c07ebca` and `docs/research/agent-sandbox-landscape.md`.
Evaluation of whether [smolmachines](https://smolmachines.com/) would
simplify the macOS agent-VM-isolation work spelled out in