|
|
|
@@ -1,16 +1,16 @@
|
|
|
|
|
# Landscape: AI-agent sandbox tools
|
|
|
|
|
|
|
|
|
|
A broader survey than [`landscape-containerized-claude.md`](landscape-containerized-claude.md),
|
|
|
|
|
which focused on Claude-Code-specific containerizers. This one covers
|
|
|
|
|
general AI-agent sandbox / containment projects — some Claude-specific,
|
|
|
|
|
some agent-agnostic, some hosted SaaS — and contrasts them with
|
|
|
|
|
bot-bottle's design.
|
|
|
|
|
Survey of AI-agent sandbox and containment projects — including local
|
|
|
|
|
coding-agent wrappers, agent-agnostic runtimes, hosted platforms, and
|
|
|
|
|
governance layers — contrasted with bot-bottle's design. The original
|
|
|
|
|
Claude-Code-specific containerizer survey was folded into this note on
|
|
|
|
|
2026-07-20 so there is one landscape and one positioning verdict.
|
|
|
|
|
|
|
|
|
|
Research conducted 2026-05-11. CubeSandbox added 2026-07-18 (see its
|
|
|
|
|
per-project note and the addendum at the end). Also updated 2026-07-18:
|
|
|
|
|
bot-bottle no longer uses **pipelock** — outbound DLP is now bot-bottle's
|
|
|
|
|
own (deliberately simple) egress scanner (a mitmproxy addon with custom
|
|
|
|
|
detectors, PRD 0017 / 0053), and git-push secret scanning is handled by
|
|
|
|
|
detectors, PRD 0017 / 0052), and git-push secret scanning is handled by
|
|
|
|
|
**gitleaks** in the git-gate. "pipelock" below has been replaced with the
|
|
|
|
|
current mechanism; it survives only in older PRDs as history.
|
|
|
|
|
|
|
|
|
@@ -20,12 +20,24 @@ Passport); an **Agent-tailored policy** row added to the comparison table;
|
|
|
|
|
a separate Governance layers section added for AGT and OAP. See the
|
|
|
|
|
second addendum at the end.
|
|
|
|
|
|
|
|
|
|
Updated 2026-07-20: the borrowable-ideas status was reconciled with the
|
|
|
|
|
current implementation. In-flight credential injection and the microVM
|
|
|
|
|
backends have shipped, while per-use SSH confirmation was superseded by
|
|
|
|
|
keeping git credentials out of the agent entirely.
|
|
|
|
|
|
|
|
|
|
Also updated 2026-07-20: **E2B and Daytona added as first-class entries.**
|
|
|
|
|
Earlier revisions mentioned E2B only as the API and lifecycle model that
|
|
|
|
|
CubeSandbox implements, and omitted Daytona entirely. That was a survey gap,
|
|
|
|
|
not a principled scope exclusion: both are major hosted sandbox platforms and
|
|
|
|
|
belong in this landscape even though they target platform builders rather than
|
|
|
|
|
bot-bottle's local single-operator workflow.
|
|
|
|
|
|
|
|
|
|
## Summary
|
|
|
|
|
|
|
|
|
|
Fifteen projects surveyed across two categories: isolation/sandbox tools
|
|
|
|
|
and governance/pre-action authorization layers (the latter don't provide
|
|
|
|
|
VM or container isolation but do per-agent policy enforcement at the
|
|
|
|
|
tool-call level). None duplicate bot-bottle's combination of local
|
|
|
|
|
The main table compares bot-bottle against fifteen isolation/sandbox tools.
|
|
|
|
|
Governance/pre-action authorization and credential-only layers are covered
|
|
|
|
|
separately because they don't provide VM or container isolation. None
|
|
|
|
|
duplicate bot-bottle's combination of local
|
|
|
|
|
VM-per-bottle isolation, a declarative per-role manifest, per-agent
|
|
|
|
|
egress allowlist + outbound-content DLP, bottle/agent split, and the
|
|
|
|
|
composable `extends:` policy model. Three clusters stand out:
|
|
|
|
@@ -34,8 +46,9 @@ composable `extends:` policy model. Three clusters stand out:
|
|
|
|
|
single-user, thin wrappers over an existing OS primitive
|
|
|
|
|
(`sandbox-exec`, Podman + Landlock).
|
|
|
|
|
- **Different category (isolation)** — tilde.run (hosted SaaS), boxlite
|
|
|
|
|
and microsandbox (microVM libraries for platform builders), CubeSandbox
|
|
|
|
|
(self-hosted multi-tenant microVM service), endo-familiar
|
|
|
|
|
and microsandbox (microVM libraries for platform builders), E2B and Daytona
|
|
|
|
|
(hosted sandbox platforms), CubeSandbox (self-hosted multi-tenant microVM
|
|
|
|
|
service), endo-familiar
|
|
|
|
|
(capability-security paradigm, no OS isolation).
|
|
|
|
|
- **New: governance/pre-action layers** — Microsoft AGT and Open Agent
|
|
|
|
|
Passport (OAP): framework-embedded tool-call interceptors with
|
|
|
|
@@ -52,15 +65,17 @@ ergonomic enough that microVMs are **now bot-bottle's default backend**
|
|
|
|
|
only as a legacy fallback for CI / hosts without KVM or Apple Container.
|
|
|
|
|
That discussion has since shipped, not just been theorized.
|
|
|
|
|
|
|
|
|
|
**The one that matters most for positioning is CubeSandbox** — it is the
|
|
|
|
|
first surveyed project to ship bot-bottle's would-be wedge (default-deny
|
|
|
|
|
egress allowlist + full audit logs + in-flight credential custody so keys
|
|
|
|
|
never enter the sandbox) *combined with* per-sandbox microVM isolation,
|
|
|
|
|
**The one that matters most for positioning is CubeSandbox** — it ships
|
|
|
|
|
bot-bottle's bundle of default-deny egress allowlisting, full audit logs, and
|
|
|
|
|
in-flight credential custody *combined with* per-sandbox microVM isolation,
|
|
|
|
|
open-source under Apache 2.0, with Tencent Cloud behind it and 10.4k
|
|
|
|
|
stars. It's a self-hosted multi-tenant service for platform builders, not
|
|
|
|
|
a single-user declarative tool, so it doesn't collide head-on — but it
|
|
|
|
|
narrows the "nobody else bundles egress custody + credential injection"
|
|
|
|
|
claim that the monetization positioning leans on. See the addendum.
|
|
|
|
|
claim that the monetization positioning leans on. Daytona now also offers
|
|
|
|
|
domain/CIDR firewall policy plus in-flight header credential substitution and
|
|
|
|
|
response scrubbing, although its higher tiers are not default-deny and its
|
|
|
|
|
production platform is proprietary. See the addendum.
|
|
|
|
|
|
|
|
|
|
## Per-project notes
|
|
|
|
|
|
|
|
|
@@ -200,6 +215,80 @@ claim that the monetization positioning leans on. See the addendum.
|
|
|
|
|
also supported.
|
|
|
|
|
- **Maturity**: Active through April 2026.
|
|
|
|
|
|
|
|
|
|
### E2B *(added 2026-07-20)*
|
|
|
|
|
|
|
|
|
|
- **Source**: https://github.com/e2b-dev/e2b ; https://e2b.dev/docs
|
|
|
|
|
- **License**: Apache 2.0 (~12.4k stars); commercial hosted service with
|
|
|
|
|
self-hosting/BYOC support.
|
|
|
|
|
- **Isolation**: Firecracker microVM per sandbox.
|
|
|
|
|
- **Locality**: Cloud-hosted by default; self-hosting uses Terraform on AWS or
|
|
|
|
|
GCP (with other targets documented as works in progress).
|
|
|
|
|
- **Agent integration**: LLM-agnostic Python and JavaScript/TypeScript SDKs;
|
|
|
|
|
code-interpreter and desktop-sandbox products. Platform primitive rather
|
|
|
|
|
than a coding-agent wrapper.
|
|
|
|
|
- **Config**: Programmatic SDK/API plus templates. Network configuration
|
|
|
|
|
supports internet on/off, outbound allow/deny rules, and a custom egress
|
|
|
|
|
proxy.
|
|
|
|
|
- **Network policy**: Configurable per sandbox, but not documented as
|
|
|
|
|
default-deny and no built-in outbound-content DLP is documented.
|
|
|
|
|
- **Credentials**: Environment variables passed to the sandbox are explicitly
|
|
|
|
|
not private at the OS level. No built-in in-flight application-credential
|
|
|
|
|
injection is documented.
|
|
|
|
|
- **Persistence**: Full memory + filesystem pause/resume, snapshots, and
|
|
|
|
|
auto-resume. Continuous runtime is tier-limited, while paused sandboxes are
|
|
|
|
|
retained indefinitely.
|
|
|
|
|
- **Maturity**: Established hosted platform and the API compatibility target
|
|
|
|
|
used by CubeSandbox.
|
|
|
|
|
|
|
|
|
|
### Daytona *(added 2026-07-20)*
|
|
|
|
|
|
|
|
|
|
- **Source**: https://github.com/daytonaio/daytona ;
|
|
|
|
|
https://www.daytona.io/docs/
|
|
|
|
|
- **License**: Current production platform is proprietary. The former AGPL
|
|
|
|
|
repository remains public but is no longer maintained after Daytona moved
|
|
|
|
|
production development closed-source in June 2026.
|
|
|
|
|
- **Isolation**: Hosted container sandboxes by default, with separate Linux
|
|
|
|
|
and Windows VM sandbox classes for dedicated-OS workloads. Each sandbox has
|
|
|
|
|
its own filesystem and network stack; VM-only features include memory
|
|
|
|
|
pause/resume and forking.
|
|
|
|
|
- **Locality**: Hosted multi-tenant service, with dedicated/custom regions and
|
|
|
|
|
customer runners available.
|
|
|
|
|
- **Agent integration**: LLM/framework-agnostic SDKs (Python, TypeScript, Go,
|
|
|
|
|
Ruby, Java), API, and CLI; official agent-framework guides. Platform
|
|
|
|
|
primitive rather than a local coding-agent wrapper.
|
|
|
|
|
- **Config**: Programmatic per-sandbox image/snapshot, resources, lifecycle,
|
|
|
|
|
firewall, and secrets.
|
|
|
|
|
- **Network policy**: Per-sandbox IPv4/domain allowlists and block-all mode,
|
|
|
|
|
subordinate to organization/tier policy. Full internet access is the
|
|
|
|
|
default on higher tiers, so it is configurable rather than uniformly
|
|
|
|
|
default-deny.
|
|
|
|
|
- **Credentials**: First-class secret manager with the same phantom-token
|
|
|
|
|
pattern as bot-bottle: the sandbox environment gets an opaque placeholder,
|
|
|
|
|
an HTTPS proxy substitutes the real secret in headers only for allowed
|
|
|
|
|
hosts, and responses are scrubbed back to the placeholder.
|
|
|
|
|
- **Persistence**: Persistent filesystem for stopped container sandboxes;
|
|
|
|
|
memory + filesystem pause/resume for VM sandboxes; snapshots and configurable
|
|
|
|
|
auto-stop.
|
|
|
|
|
- **Maturity**: Production commercial platform. Notable April 2026 credential
|
|
|
|
|
exposure was patched; the June 2026 closed-source transition materially
|
|
|
|
|
changes its transparency/self-hosting posture.
|
|
|
|
|
|
|
|
|
|
### Other hosted runtimes carried forward from the earlier survey
|
|
|
|
|
|
|
|
|
|
- **Northflank Sandboxes** — hosted or customer-cloud, microVM-backed
|
|
|
|
|
containers with SDK-managed lifecycle, optional persistent volumes, and
|
|
|
|
|
sub-second claimed boot. This is a platform primitive for untrusted code and
|
|
|
|
|
agents, not a local agent wrapper or role-policy layer.
|
|
|
|
|
- **Cloudflare Sandbox SDK** — Workers/Durable Objects API over VM-isolated
|
|
|
|
|
Linux containers for command, file, process, and service execution. It is a
|
|
|
|
|
hosted TypeScript platform primitive; application authentication,
|
|
|
|
|
authorization, and credential-proxy patterns remain the integrator's job.
|
|
|
|
|
|
|
|
|
|
Both belong to the same “build your agent platform on this runtime” category as
|
|
|
|
|
E2B and Daytona. They were named but not analyzed in depth by the original
|
|
|
|
|
Claude-specific note, so they remain outside the main comparison table rather
|
|
|
|
|
than being presented with false precision.
|
|
|
|
|
|
|
|
|
|
### CubeSandbox *(added 2026-07-18)*
|
|
|
|
|
- **Source**: https://github.com/TencentCloud/CubeSandbox ;
|
|
|
|
|
HN launch https://news.ycombinator.com/item?id=47863430
|
|
|
|
@@ -316,6 +405,92 @@ claim that the monetization positioning leans on. See the addendum.
|
|
|
|
|
preview — APIs may change.
|
|
|
|
|
- **Maturity**: Early research preview.
|
|
|
|
|
|
|
|
|
|
## Claude-specific wrappers and developer environments
|
|
|
|
|
|
|
|
|
|
These projects were the focus of the original containerized-Claude survey.
|
|
|
|
|
They remain useful comparisons for local developer experience, but most are
|
|
|
|
|
templates or wrappers rather than policy-bearing sandbox platforms, so they
|
|
|
|
|
are grouped here instead of widening the main table further.
|
|
|
|
|
|
|
|
|
|
### claudebox
|
|
|
|
|
|
|
|
|
|
- **Source**: https://github.com/RchGrav/claudebox
|
|
|
|
|
- **Isolation**: Docker, with per-project images, authentication state, and
|
|
|
|
|
configuration.
|
|
|
|
|
- **Agent integration**: Claude Code wrapper with 15+ preconfigured language
|
|
|
|
|
and task profiles.
|
|
|
|
|
- **Network policy**: Per-project firewall allowlists.
|
|
|
|
|
- **Closest overlap**: local one-command developer workflow and project-scoped
|
|
|
|
|
network policy.
|
|
|
|
|
- **Difference**: profiles describe development toolchains, not named agent
|
|
|
|
|
roles. There is no bottle/agent split, composable role manifest, provider
|
|
|
|
|
plugin layer, or outbound-content DLP.
|
|
|
|
|
|
|
|
|
|
### Spritz / claude-code-sandbox
|
|
|
|
|
|
|
|
|
|
- **Source**: https://github.com/textcortex/claude-code-sandbox (archived;
|
|
|
|
|
points to its successor, Spritz).
|
|
|
|
|
- **Isolation**: The original project ran Claude Code in local Docker with
|
|
|
|
|
bypass permissions; Spritz moved toward Kubernetes-native multi-agent
|
|
|
|
|
infrastructure.
|
|
|
|
|
- **Difference**: the successor targets cluster orchestration rather than a
|
|
|
|
|
low-dependency local launcher. It is architecturally closer to hosted or
|
|
|
|
|
Kubernetes platform runtimes than to bot-bottle's single-operator CLI.
|
|
|
|
|
|
|
|
|
|
### Trail of Bits claude-code-devcontainer
|
|
|
|
|
|
|
|
|
|
- **Source**: https://github.com/trailofbits/claude-code-devcontainer
|
|
|
|
|
- **Isolation**: A Docker devcontainer that exposes only project files and is
|
|
|
|
|
designed to run Claude Code with `bypassPermissions` for security audits and
|
|
|
|
|
untrusted-code review.
|
|
|
|
|
- **Difference**: a hardened, reusable environment definition rather than an
|
|
|
|
|
agent launcher or fleet. It has no named-role manifest, per-role credential
|
|
|
|
|
custody, supervision plane, or multi-backend abstraction.
|
|
|
|
|
|
|
|
|
|
### Smaller wrappers and official templates
|
|
|
|
|
|
|
|
|
|
Projects such as `arezi/claude-sandbox`, `nkrefman/claude-sandbox`, and
|
|
|
|
|
`VishalJ99/claude-docker`, plus Docker/Anthropic devcontainer templates, prove
|
|
|
|
|
there is steady demand for “Claude in a container.” They are deliberately
|
|
|
|
|
small launch/build configurations. They compete on setup simplicity, not on
|
|
|
|
|
role-aware policy, credential custody, persistent supervision, or a fleet
|
|
|
|
|
model, and are better treated as a product category than as individual rows.
|
|
|
|
|
|
|
|
|
|
### SuperHQ
|
|
|
|
|
|
|
|
|
|
- **Source**: https://superhq.ai/
|
|
|
|
|
- **Isolation**: Apple-Silicon desktop application using local microVMs via
|
|
|
|
|
Virtualization.framework/libkrun-era components.
|
|
|
|
|
- **Agent integration**: Claude Code, Codex, and Pi in a GUI, with mobile
|
|
|
|
|
remote access.
|
|
|
|
|
- **Credentials and review**: host-side auth gateway injects credentials on
|
|
|
|
|
the wire; a temporary overlay stages writes for diff-and-accept review.
|
|
|
|
|
- **Closest overlap**: local microVM isolation, multi-provider launching, and
|
|
|
|
|
credential custody for security-minded individual developers.
|
|
|
|
|
- **Difference**: GUI desktop product on Apple Silicon rather than a
|
|
|
|
|
cross-platform declarative CLI/fleet layer. The July 2026 snapshot in the
|
|
|
|
|
original survey recorded a user request for per-run tool-call and network
|
|
|
|
|
audit logging; treat that as point-in-time rather than a permanent gap.
|
|
|
|
|
|
|
|
|
|
## Credential gateway without isolation
|
|
|
|
|
|
|
|
|
|
### OneCLI
|
|
|
|
|
|
|
|
|
|
[OneCLI](https://onecli.sh/) is a framework-agnostic identity gateway rather
|
|
|
|
|
than a sandbox. Its phantom-token design gives the agent a placeholder and
|
|
|
|
|
substitutes the encrypted real credential at the network layer. It therefore
|
|
|
|
|
matches bot-bottle closely on secret custody, and is more portable because it
|
|
|
|
|
can sit in front of agents launched by anything, but it supplies no container
|
|
|
|
|
or VM boundary, filesystem isolation, role manifest, or egress-content DLP.
|
|
|
|
|
|
|
|
|
|
The positioning consequence from the earlier survey still holds: secret
|
|
|
|
|
custody alone is not unique. bot-bottle's relevant combination is local
|
|
|
|
|
isolation + default-deny egress + payload DLP + declarative roles + credential
|
|
|
|
|
custody. OneCLI's managed tier also places custody with a third party, whereas
|
|
|
|
|
bot-bottle keeps it within operator-controlled infrastructure. See
|
|
|
|
|
[`agent-credential-proxy-landscape.md`](agent-credential-proxy-landscape.md)
|
|
|
|
|
for the detailed build-versus-adopt analysis.
|
|
|
|
|
|
|
|
|
|
## Governance / pre-action authorization layers
|
|
|
|
|
|
|
|
|
|
These two tools don't provide VM or filesystem isolation; they intercept
|
|
|
|
@@ -371,19 +546,19 @@ them.
|
|
|
|
|
|
|
|
|
|
*Isolation/sandbox tools only. AGT and OAP are governance layers — see their per-project notes above.*
|
|
|
|
|
|
|
|
|
|
| Axis | bot-bottle | endo-familiar | litterbox | agent-safehouse | matchlock | tilde.run | boxlite | microsandbox | smolmachines | CubeSandbox | Cleanroom | container-use | Docker sbx | Anthropic srt |
|
|
|
|
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
|
|
|
| Isolation | MicroVM per bottle default (Firecracker/KVM on Linux, Apple Container on macOS) + own egress DLP scanner; Docker legacy fallback, gVisor there if present | Object-capability (no OS isolation) | Podman + opt. Landlock | macOS `sandbox-exec` | MicroVM (Firecracker / Virt.fw) | Hosted container (unverified) | MicroVM (KVM / Hypervisor.fw) | MicroVM (libkrun) | MicroVM (libkrun / KVM) | MicroVM (RustVMM / KVM) | MicroVM (Firecracker / Virt.fw) | Docker container + git worktree | MicroVM (proprietary) | OS-level (Seatbelt / bubblewrap / WFP) — no container |
|
|
|
|
|
| Local vs hosted | Local | Local | Local (Linux) | Local (macOS) | Local | Hosted SaaS | Local | Local | Local | Self-hosted (server/cluster) | Self-hosted server | Local | Local | Local |
|
|
|
|
|
| Open source | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | MIT | No | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Proprietary | Apache 2.0 (experimental) |
|
|
|
|
|
| Agent target | Claude Code | Generic (demo) | Generic | Multi-agent wrapper | Generic (+ Claude/OpenAI SDKs) | Claude focus | Generic | Claude + Cursor (MCP/Skills) | Generic (AGENTS.md) | E2B-compatible (platform builders) | CI / generic process | Claude Code, Cursor, Windsurf (MCP) | Claude Code, Codex, Gemini CLI, Copilot, Kiro | Claude Code (and any process) |
|
|
|
|
|
| Network policy | Default-deny via own egress scanner + per-bottle allowlist + content DLP + gitleaks on git push | Capability model only | Limited | Not addressed | Default-deny + allowlist + secret-injecting proxy | Default-deny + logging | Per-VM net (unverified) | Not documented | Off by default + allowlist | Default-deny allowlist + instant egress block + audit logs + per-sandbox tokens (eBPF) + credential vault | Default-deny + per-repo host allowlist (cleanroom.yaml) | Not addressed | Default-deny; Open / Balanced / Locked Down presets; live TUI network panel | Proxy-based allowlist/denylist (HTTP + SOCKS5); custom proxy supported |
|
|
|
|
|
| Parallel agents | Yes (one bottle per agent) | n/a | Not addressed | One at a time | Multiple VMs | Yes (dashboard) | SDK-level | SDK-level | Architectural | Yes (2,000+/host claimed) | Yes (server model) | Yes (per-agent containers + worktrees) | Yes | Yes |
|
|
|
|
|
| Long-running posture | Persistent by default (named, supervised) | n/a (demo) | Session (up while in use) | Per-invocation | Ephemeral VM per run | Per-run (versioned) | Ephemeral + snapshot/fork | Ephemeral / on-demand | Named persistent by default | Ephemeral + auto pause/resume | Per-run + suspend/resume | Per-agent container (ephemeral) | Per-session; branch mode creates git worktree in .sbx/ | Per-invocation |
|
|
|
|
|
| DX: run Claude yolo-style | One command → interactive yolo Claude (`start <agent>`, `--dangerously-skip-permissions` default) | n/a (lib demo) | Wizard + build, then run claude inside (Linux only) | One-command wrapper (`safehouse claude --dangerously-skip-permissions`) | CLI: run a cmd in a VM (not a Claude wrapper) | Hosted (`tilde exec`), not local-native | SDK code required (build the run yourself) | CLI/MCP: sandbox-as-a-tool for the agent, not a wrapper around it | SSH into a named machine, run claude there | Stand up a cluster + drive via E2B SDK | CI-oriented, not a Claude wrapper | MCP server: `claude mcp add container-use -- container-use stdio` | One command: `sbx` wraps claude with `--dangerously-skip-permissions` default | Library/wrapper, not a standalone CLI |
|
|
|
|
|
| Config | JSON manifest (bottles + agents) | Programmatic refs | CLI wizard | Profile files / shell fns | CLI / SDK | DSL + CLI + SDK | SDK | CLI / SDK / MCP | TOML Smolfile | E2B-compatible SDK | cleanroom.yaml in repo | None (no policy config) | Preset levels at launch | Programmatic per-invocation (allow/deny lists) |
|
|
|
|
|
| Agent-tailored policy | Yes — bottle/agent split; declarative per-role egress + credentials; composable via `extends:` | Partial — capability model scopes per-agent, but no declarative role manifest | No | Partial — per-agent profile files (Seatbelt); no egress | No | Yes — per-agent DSL RBAC (allow/deny/approve per action/repo/agent) | No | No | No | No — per-sandbox SDK config, not role-scoped | Partial — per-repo cleanroom.yaml, not per-role | No | No — network presets only | No |
|
|
|
|
|
| Maturity | Active July 2026 | Research (2022+) | Early (~66 ⭐) | Active (~1.8k ⭐) | Experimental (~574 ⭐) | Private preview | YC, ~4.7k ⭐ | YC, ~6k ⭐, beta | ~3.1k ⭐ | Tencent, prod, ~10.4k ⭐ | Active (Buildkite product) | Early development | GA 2026 | Early research preview |
|
|
|
|
|
| Axis | bot-bottle | endo-familiar | litterbox | agent-safehouse | matchlock | tilde.run | boxlite | microsandbox | smolmachines | E2B | Daytona | CubeSandbox | Cleanroom | container-use | Docker sbx | Anthropic srt |
|
|
|
|
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
|
|
|
| Isolation | MicroVM per bottle default (Firecracker/KVM on Linux, Apple Container on macOS) + own egress DLP scanner; Docker legacy fallback, gVisor there if present | Object-capability (no OS isolation) | Podman + opt. Landlock | macOS `sandbox-exec` | MicroVM (Firecracker / Virt.fw) | Hosted container (unverified) | MicroVM (KVM / Hypervisor.fw) | MicroVM (libkrun) | MicroVM (libkrun / KVM) | Firecracker microVM | Container or Linux/Windows VM class | MicroVM (RustVMM / KVM) | MicroVM (Firecracker / Virt.fw) | Docker container + git worktree | MicroVM (proprietary) | OS-level (Seatbelt / bubblewrap / WFP) — no container |
|
|
|
|
|
| Local vs hosted | Local | Local | Local (Linux) | Local (macOS) | Local | Hosted SaaS | Local | Local | Local | Hosted; self-host/BYOC available | Hosted; dedicated/custom regions | Self-hosted (server/cluster) | Self-hosted server | Local | Local | Local |
|
|
|
|
|
| Open source | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | MIT | No | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Production closed-source; legacy AGPL repo unmaintained | Apache 2.0 | Apache 2.0 | Apache 2.0 | Proprietary | Apache 2.0 (experimental) |
|
|
|
|
|
| Agent target | Claude Code, Codex, Pi, and provider plugins | Generic (demo) | Generic | Multi-agent wrapper | Generic (+ Claude/OpenAI SDKs) | Claude focus | Generic | Claude + Cursor (MCP/Skills) | Generic (AGENTS.md) | LLM-agnostic platform builders | LLM-agnostic platform builders | E2B-compatible (platform builders) | CI / generic process | Claude Code, Cursor, Windsurf (MCP) | Claude Code, Codex, Gemini CLI, Copilot, Kiro | Claude Code (and any process) |
|
|
|
|
|
| Network policy | Default-deny via own egress scanner + per-bottle allowlist + content DLP + gitleaks on git push | Capability model only | Limited | Not addressed | Default-deny + allowlist + secret-injecting proxy | Default-deny + logging | Per-VM net (unverified) | Not documented | Off by default + allowlist | Per-sandbox allow/deny rules and custom egress proxy; internet configurable | Per-sandbox CIDR/domain allowlist or block-all; tier policy; secret-injecting proxy | Default-deny allowlist + instant egress block + audit logs + per-sandbox tokens (eBPF) + credential vault | Default-deny + per-repo host allowlist (cleanroom.yaml) | Not addressed | Default-deny; Open / Balanced / Locked Down presets; live TUI network panel | Proxy-based allowlist/denylist (HTTP + SOCKS5); custom proxy supported |
|
|
|
|
|
| Parallel agents | Yes (one bottle per agent) | n/a | Not addressed | One at a time | Multiple VMs | Yes (dashboard) | SDK-level | SDK-level | Architectural | Yes (platform service) | Yes (platform service) | Yes (2,000+/host claimed) | Yes (server model) | Yes (per-agent containers + worktrees) | Yes | Yes |
|
|
|
|
|
| Long-running posture | Persistent by default (named, supervised) | n/a (demo) | Session (up while in use) | Per-invocation | Ephemeral VM per run | Per-run (versioned) | Ephemeral + snapshot/fork | Ephemeral / on-demand | Named persistent by default | Runtime tier limits + indefinite pause/resume | Persistent filesystem; VM pause/resume; configurable auto-stop | Ephemeral + auto pause/resume | Per-run + suspend/resume | Per-agent container (ephemeral) | Per-session; branch mode creates git worktree in .sbx/ | Per-invocation |
|
|
|
|
|
| DX: run Claude yolo-style | One command → interactive yolo Claude (`start <agent>`, `--dangerously-skip-permissions` default) | n/a (lib demo) | Wizard + build, then run claude inside (Linux only) | One-command wrapper (`safehouse claude --dangerously-skip-permissions`) | CLI: run a cmd in a VM (not a Claude wrapper) | Hosted (`tilde exec`), not local-native | SDK code required (build the run yourself) | CLI/MCP: sandbox-as-a-tool for the agent, not a wrapper around it | SSH into a named machine, run claude there | SDK/CLI sandbox; wire the agent yourself | SDK/CLI sandbox; wire the agent yourself | Stand up a cluster + drive via E2B SDK | CI-oriented, not a Claude wrapper | MCP server: `claude mcp add container-use -- container-use stdio` | One command: `sbx` wraps claude with `--dangerously-skip-permissions` default | Library/wrapper, not a standalone CLI |
|
|
|
|
|
| Config | YAML-in-Markdown manifests (bottles + agents) | Programmatic refs | CLI wizard | Profile files / shell fns | CLI / SDK | DSL + CLI + SDK | SDK | CLI / SDK / MCP | TOML Smolfile | SDK/API + templates | SDK/API/CLI + images/snapshots | E2B-compatible SDK | cleanroom.yaml in repo | None (no policy config) | Preset levels at launch | Programmatic per-invocation (allow/deny lists) |
|
|
|
|
|
| Agent-tailored policy | Yes — bottle/agent split; declarative per-role egress + credentials; composable via `extends:` | Partial — capability model scopes per-agent, but no declarative role manifest | No | Partial — per-agent profile files (Seatbelt); no egress | No | Yes — per-agent DSL RBAC (allow/deny/approve per action/repo/agent) | No | No | No | No — per-sandbox SDK config | No — per-sandbox SDK config | No — per-sandbox SDK config, not role-scoped | Partial — per-repo cleanroom.yaml, not per-role | No | No — network presets only | No |
|
|
|
|
|
| Maturity | Active July 2026 | Research (2022+) | Early (~66 ⭐) | Active (~1.8k ⭐) | Experimental (~574 ⭐) | Private preview | YC, ~4.7k ⭐ | YC, ~6k ⭐, beta | ~3.1k ⭐ | Established hosted platform, ~12.4k ⭐ | Production commercial; closed-source since June 2026 | Tencent, prod, ~10.4k ⭐ | Active (Buildkite product) | Early development | GA 2026 | Early research preview |
|
|
|
|
|
|
|
|
|
|
## What's closest, what's different
|
|
|
|
|
|
|
|
|
@@ -433,47 +608,81 @@ alternative.
|
|
|
|
|
|
|
|
|
|
**Solving a different problem.** tilde.run is hosted SaaS for team /
|
|
|
|
|
production agent pipelines with data-versioned rollback — explicitly
|
|
|
|
|
opposite to bot-bottle's "infrastructure I control" goal. boxlite,
|
|
|
|
|
microsandbox, and CubeSandbox are infrastructure libraries/services aimed
|
|
|
|
|
at platform builders embedding sandboxes into agent frameworks; they
|
|
|
|
|
opposite to bot-bottle's "infrastructure I control" goal. E2B and Daytona
|
|
|
|
|
are hosted sandbox platforms, while boxlite, microsandbox, and CubeSandbox
|
|
|
|
|
are infrastructure libraries/services aimed at platform builders embedding
|
|
|
|
|
sandboxes into agent frameworks; they
|
|
|
|
|
would be a *backend* bot-bottle could call, not a competitor to its
|
|
|
|
|
manifest layer. endo-familiar is in a different paradigm entirely:
|
|
|
|
|
capability passing rather than kernel boundaries.
|
|
|
|
|
|
|
|
|
|
## Borrowable ideas
|
|
|
|
|
|
|
|
|
|
What bot-bottle already has that the survey suggested as
|
|
|
|
|
differentiators:
|
|
|
|
|
### Already shipped or otherwise addressed
|
|
|
|
|
|
|
|
|
|
- Default-deny egress with a per-agent allowlist (own egress scanner).
|
|
|
|
|
- DLP scanning of outbound traffic.
|
|
|
|
|
- Bottle / agent split (manifest layer above the isolation primitive).
|
|
|
|
|
- gVisor auto-detection on Linux.
|
|
|
|
|
- **In-flight secret injection** (suggested by matchlock) — **shipped.**
|
|
|
|
|
Real provider and git-host tokens are held outside the agent and injected
|
|
|
|
|
by the egress gateway on matching routes. The agent receives only proxy
|
|
|
|
|
URLs and, where a client requires a credential-shaped value, a placeholder;
|
|
|
|
|
`GITEA_TOKEN` and equivalent real tokens do not appear in the agent's
|
|
|
|
|
environment.
|
|
|
|
|
- **MicroVM backend** — **shipped.** MicroVMs are now the default:
|
|
|
|
|
Firecracker on KVM Linux and Apple Container on macOS. Docker is the legacy
|
|
|
|
|
fallback.
|
|
|
|
|
- **Per-use SSH key confirmation** (suggested by litterbox) — **addressed by
|
|
|
|
|
stronger credential custody instead.** The agent does not hold the upstream
|
|
|
|
|
git SSH key or an SSH-agent socket: git-gate holds the credential and gates
|
|
|
|
|
git operations. A confirmation wrapper inside the agent would therefore
|
|
|
|
|
protect a credential that is no longer there. Operator approval at the gate
|
|
|
|
|
remains the appropriate control point for any future per-use confirmation.
|
|
|
|
|
|
|
|
|
|
Ideas worth considering, without abandoning the Python-stdlib-first /
|
|
|
|
|
local, single-operator stance:
|
|
|
|
|
### Still worth considering
|
|
|
|
|
|
|
|
|
|
1. **Per-use SSH key confirmation** (from litterbox). Even with
|
|
|
|
|
KnownHostKey pinning and the egress DLP scanner, a wrapper SSH agent that
|
|
|
|
|
prompts on each key use (e.g. via `osascript` / `notify-send`) would
|
|
|
|
|
catch an agent doing something off-policy with a key it legitimately
|
|
|
|
|
holds. Pure-stdlib, no new deps.
|
|
|
|
|
2. **In-flight secret injection** (from matchlock). The egress scanner
|
|
|
|
|
already does allowlisting and DLP; teaching it to *inject* tokens at
|
|
|
|
|
proxy time so e.g. `GITEA_TOKEN` never appears in the container's
|
|
|
|
|
env would close the "agent reads its own env and exfiltrates" path.
|
|
|
|
|
Fits the existing egress-proxy architecture.
|
|
|
|
|
3. **MicroVM backend** — ~~on the radar~~ **shipped since this survey.**
|
|
|
|
|
microVMs are now bot-bottle's default (Firecracker on KVM Linux, Apple
|
|
|
|
|
Container on macOS); Docker is the legacy fallback. The libkrun / Apple
|
|
|
|
|
Virtualization.framework ergonomics that microsandbox, smolmachines,
|
|
|
|
|
and matchlock demonstrated turned out to be enough to make it the
|
|
|
|
|
default rather than an opt-in.
|
|
|
|
|
- **Live network activity in the supervisor TUI** (from Docker sbx): show
|
|
|
|
|
allowed and blocked connections and let the operator propose policy changes
|
|
|
|
|
from the existing supervision surface.
|
|
|
|
|
- **Tamper-evident audit records** (from OAP): sign and hash-chain egress and
|
|
|
|
|
supervision decisions for compliance-sensitive deployments.
|
|
|
|
|
- **Behaviour-informed policy downgrade** (from Microsoft AGT): use repeated
|
|
|
|
|
DLP alerts or supervision holds as a signal to narrow policy or request
|
|
|
|
|
closer review. This needs a carefully specified trust model before it can be
|
|
|
|
|
more than a heuristic.
|
|
|
|
|
|
|
|
|
|
Not worth borrowing: the SDK-first programmatic API style of boxlite /
|
|
|
|
|
microsandbox (cuts against the declarative-manifest stance), and the
|
|
|
|
|
hosted-SaaS dashboard model of tilde.run (cuts against the
|
|
|
|
|
"infrastructure I control" goal).
|
|
|
|
|
|
|
|
|
|
## Publishing and positioning verdict
|
|
|
|
|
|
|
|
|
|
Publishing remains worthwhile, but the defensible claim is the combination,
|
|
|
|
|
not any single primitive. Credential custody is matched by OneCLI, matchlock,
|
|
|
|
|
Daytona, Docker sbx, and CubeSandbox; local one-command isolation is matched by
|
|
|
|
|
agent-safehouse and Docker sbx; hosted microVM execution is a crowded platform
|
|
|
|
|
category.
|
|
|
|
|
|
|
|
|
|
bot-bottle remains unusual in combining:
|
|
|
|
|
|
|
|
|
|
- local, operator-controlled execution with persistent named bottles;
|
|
|
|
|
- one declarative role layer across Claude Code, Codex, Pi, and provider
|
|
|
|
|
plugins;
|
|
|
|
|
- composable agent/bottle manifests, skills, and system prompts;
|
|
|
|
|
- Firecracker/Apple Container isolation with a Docker fallback;
|
|
|
|
|
- default-deny per-role egress, payload DLP, and git-push secret scanning;
|
|
|
|
|
- credentials injected outside the agent process; and
|
|
|
|
|
- supervision and audit state suited to long-running parallel agents.
|
|
|
|
|
|
|
|
|
|
The practical wedge is “as easy as native yolo, with declarative role policy
|
|
|
|
|
and self-hosted custody,” including scoped access to private LAN/Tailnet
|
|
|
|
|
services that cloud-first runtimes cannot provide without additional network
|
|
|
|
|
plumbing. The main competitive risks are a local wrapper such as claudebox or
|
|
|
|
|
Docker sbx growing a role-manifest layer, and GUI products such as SuperHQ
|
|
|
|
|
adding equivalent policy and audit depth.
|
|
|
|
|
|
|
|
|
|
## Caveats
|
|
|
|
|
|
|
|
|
|
- Star counts and last-commit dates are point-in-time snapshots.
|
|
|
|
@@ -490,7 +699,8 @@ hosted-SaaS dashboard model of tilde.run (cuts against the
|
|
|
|
|
|
|
|
|
|
CubeSandbox (Tencent Cloud, Apache 2.0, ~10.4k stars, HN launch
|
|
|
|
|
[#47863430](https://news.ycombinator.com/item?id=47863430)) is the first
|
|
|
|
|
project in this survey to combine, in one open-source stack, everything
|
|
|
|
|
open-source, self-hostable project in this survey to combine, in one stack,
|
|
|
|
|
the main primitives
|
|
|
|
|
bot-bottle treated as its differentiator:
|
|
|
|
|
|
|
|
|
|
- **Egress custody (connection level)** — default-deny domain allowlist
|
|
|
|
@@ -558,7 +768,7 @@ instead of per-action prompts. On this axis the field splits cleanly:
|
|
|
|
|
network egress; bot-bottle adds VM-grade isolation, egress DLP, and
|
|
|
|
|
persistent/parallel bottles across macOS + Linux.
|
|
|
|
|
- **Libraries / services** (you build the run yourself): boxlite,
|
|
|
|
|
microsandbox, CubeSandbox, E2B. These hand you an SDK or a cluster and
|
|
|
|
|
microsandbox, CubeSandbox, E2B, Daytona. These hand you an SDK or a cluster and
|
|
|
|
|
expect you to wire the agent in — powerful for platform builders,
|
|
|
|
|
heavyweight for "just run Claude on my laptop." microsandbox's MCP/Skills
|
|
|
|
|
angle is *sandbox-as-a-tool the agent calls*, which is the inverse of
|
|
|
|
@@ -566,10 +776,11 @@ instead of per-action prompts. On this axis the field splits cleanly:
|
|
|
|
|
- **In between:** litterbox (wizard + build, Linux only), smolmachines
|
|
|
|
|
(SSH into a named machine), matchlock (run a command in a VM).
|
|
|
|
|
|
|
|
|
|
So DX is a genuine bot-bottle differentiator, and the only project that
|
|
|
|
|
matches it (agent-safehouse) does so with materially weaker isolation and
|
|
|
|
|
no egress story. "As easy as native yolo, but actually sandboxed" is a
|
|
|
|
|
defensible one-liner.
|
|
|
|
|
So DX is a genuine bot-bottle differentiator. agent-safehouse matches the
|
|
|
|
|
one-command wrapper with weaker isolation and no egress story; Docker sbx now
|
|
|
|
|
matches it at microVM strength but remains proprietary and preset-based. "As
|
|
|
|
|
easy as native yolo, with declarative role policy" is the narrower defensible
|
|
|
|
|
one-liner.
|
|
|
|
|
|
|
|
|
|
Why it still doesn't collide head-on:
|
|
|
|
|
|
|
|
|
@@ -577,7 +788,8 @@ Why it still doesn't collide head-on:
|
|
|
|
|
builders* (drop-in E2B replacement, SDK-driven, 2,000 sandboxes on a
|
|
|
|
|
box). bot-bottle is a *single-operator, declarative-manifest tool for
|
|
|
|
|
the infrastructure I run*. Different buyer, different ergonomics — no
|
|
|
|
|
JSON manifest, no bottle/agent split, no "one command on my laptop."
|
|
|
|
|
declarative role manifest, no bottle/agent split, no "one command on my
|
|
|
|
|
laptop."
|
|
|
|
|
2. **Backend, not competitor.** Like boxlite/microsandbox, CubeSandbox is
|
|
|
|
|
something bot-bottle could sit *on top of* — a `"runtime": "microvm"`
|
|
|
|
|
or `"runtime": "cubesandbox"` backend under the manifest layer — while
|
|
|
|
@@ -588,7 +800,8 @@ Why it matters anyway:
|
|
|
|
|
|
|
|
|
|
- The "nobody else bundles connection-level egress allowlist + audit +
|
|
|
|
|
in-flight credential custody" line is **no longer true for the
|
|
|
|
|
primitive** — a well-funded, 10k-star open-source project now ships it.
|
|
|
|
|
primitive** — CubeSandbox ships the open-source/self-hosted combination,
|
|
|
|
|
and Daytona ships a proprietary firewall + credential-substitution variant.
|
|
|
|
|
But **content DLP on authorized channels is still not matched** (see
|
|
|
|
|
above), and neither is the *layer above* the primitive (declarative
|
|
|
|
|
manifest, cross-vendor orchestration, operator UX, the
|
|
|
|
@@ -599,8 +812,8 @@ Why it matters anyway:
|
|
|
|
|
that in mind.
|
|
|
|
|
- Worth a closer look at **how** CubeSandbox does credential injection
|
|
|
|
|
and per-sandbox egress tokens (eBPF virtual switch vs. bot-bottle's
|
|
|
|
|
mitmproxy egress proxy) before the next iteration of bot-bottle's
|
|
|
|
|
in-flight-secret feature — see borrowable idea #2 above.
|
|
|
|
|
mitmproxy egress proxy) when hardening bot-bottle's now-shipped
|
|
|
|
|
credential-custody implementation.
|
|
|
|
|
|
|
|
|
|
## Addendum 2026-07-18 (second pass) — agent-tailored policy landscape
|
|
|
|
|
|
|
|
|
@@ -639,7 +852,7 @@ whether a permitted action is consistent with the agent's actual task.
|
|
|
|
|
See `hn-agent-safety-discourse-july-2026.md` for the blast-radius
|
|
|
|
|
analysis.
|
|
|
|
|
|
|
|
|
|
**Borrowable from new tools:**
|
|
|
|
|
**Open ideas from new tools (also summarized above):**
|
|
|
|
|
|
|
|
|
|
- **Microsoft AGT's trust-score decay** — privilege that reflects
|
|
|
|
|
observed behaviour rather than static provisioning. Applied to
|
|
|
|
|