Files
bot-bottle/docs/research/agent-sandbox-landscape.md
T

51 KiB
Raw Blame History

Landscape: AI-agent sandbox tools

Survey of AI-agent sandbox and containment projects — including local coding-agent wrappers, agent-agnostic runtimes, hosted platforms, and governance layers — contrasted with bot-bottle's design. The original Claude-Code-specific containerizer survey was folded into this note on 2026-07-20 so there is one landscape and one positioning verdict.

Research conducted 2026-05-11. CubeSandbox added 2026-07-18 (see its per-project note and the addendum at the end). Also updated 2026-07-18: bot-bottle no longer uses pipelock — outbound DLP is now bot-bottle's own (deliberately simple) egress scanner (a mitmproxy addon with custom detectors, PRD 0017 / 0052), and git-push secret scanning is handled by gitleaks in the git-gate. "pipelock" below has been replaced with the current mechanism; it survives only in older PRDs as history.

Updated again 2026-07-18: six additional tools added (Cleanroom, container-use, Docker sbx, Anthropic srt, Microsoft AGT, Open Agent Passport); an Agent-tailored policy row added to the comparison table; a separate Governance layers section added for AGT and OAP. See the second addendum at the end.

Updated 2026-07-20: the borrowable-ideas status was reconciled with the current implementation. In-flight credential injection and the microVM backends have shipped, while per-use SSH confirmation was superseded by keeping git credentials out of the agent entirely.

Also updated 2026-07-20: E2B and Daytona added as first-class entries. Earlier revisions mentioned E2B only as the API and lifecycle model that CubeSandbox implements, and omitted Daytona entirely. That was a survey gap, not a principled scope exclusion: both are major hosted sandbox platforms and belong in this landscape even though they target platform builders rather than bot-bottle's local single-operator workflow.

Summary

The main table compares bot-bottle against fifteen isolation/sandbox tools. Governance/pre-action authorization and credential-only layers are covered separately because they don't provide VM or container isolation. None duplicate bot-bottle's combination of local VM-per-bottle isolation, a declarative per-role manifest, per-agent egress allowlist + outbound-content DLP, bottle/agent split, and the composable extends: policy model. Three clusters stand out:

  • Closest neighbours — agent-safehouse and litterbox: local, single-user, thin wrappers over an existing OS primitive (sandbox-exec, Podman + Landlock).
  • Different category (isolation) — tilde.run (hosted SaaS), boxlite and microsandbox (microVM libraries for platform builders), E2B and Daytona (hosted sandbox platforms), CubeSandbox (self-hosted multi-tenant microVM service), endo-familiar (capability-security paradigm, no OS isolation).
  • New: governance/pre-action layers — Microsoft AGT and Open Agent Passport (OAP): framework-embedded tool-call interceptors with per-agent declarative policy. Closest competitors on agent-tailored policy, but operate at the tool-call level rather than providing network/filesystem isolation; they complement rather than substitute.

The microVM cluster (matchlock, smolmachines, boxlite, microsandbox, CubeSandbox) is the most relevant for the v2 isolation discussion in stronger-isolation-alternatives.md: libkrun and Apple's Virtualization.framework have made local microVMs ergonomic enough that microVMs are now bot-bottle's default backend (Firecracker on KVM Linux, Apple Container on macOS), with Docker kept only as a legacy fallback for CI / hosts without KVM or Apple Container. That discussion has since shipped, not just been theorized.

The one that matters most for positioning is CubeSandbox — it ships bot-bottle's bundle of default-deny egress allowlisting, full audit logs, and in-flight credential custody combined with per-sandbox microVM isolation, open-source under Apache 2.0, with Tencent Cloud behind it and 10.4k stars. It's a self-hosted multi-tenant service for platform builders, not a single-user declarative tool, so it doesn't collide head-on — but it narrows the "nobody else bundles egress custody + credential injection" claim that the monetization positioning leans on. Daytona now also offers domain/CIDR firewall policy plus in-flight header credential substitution and response scrubbing, although its higher tiers are not default-deny and its production platform is proprietary. See the addendum.

Per-project notes

endo-familiar

  • Source: https://dcfoundation.io/containing-ai-agents-the-endo-familiar-demo/ ; https://github.com/endojs/endo
  • License: Apache 2.0
  • Isolation: Object-capability runtime in Hardened JavaScript. Not OS-level — agents simply cannot reference resources they were not handed.
  • Locality: Local / decentralized; WebSocket relay for capability sharing across machines.
  • Agent integration: Agent-agnostic, demo only.
  • Config: Programmatic capability passing; "pet name" system for human-readable capability handles.
  • Network policy: Capability model is the policy; no allowlist or firewall.
  • Maturity: Research demo, Foresight Institute grant. Production use of endo is via Agoric and MetaMask, not as a containment tool.

litterbox

  • Source: https://litterbox.work/ ; https://github.com/Gerharddc/litterbox
  • License: Apache 2.0 (~66 stars)
  • Isolation: Podman container on Linux + Wayland socket forwarding; optional Landlock LSM for filesystem restriction.
  • Locality: Local, Linux only.
  • Agent integration: Generic dev sandbox; works with any agent that runs inside the container.
  • Config: Interactive CLI wizard — define (Dockerfile template), build (prompts), start (launch).
  • Network policy: "Limited isolation by default" — no strict allowlist documented.
  • Notable: Per-key SSH agent confirmation dialogs.
  • Maturity: Early-stage, ~66 stars.

agent-safehouse

  • Source: https://agent-safehouse.dev/ ; https://github.com/eugene1g/agent-safehouse
  • HN launch: #47301085 (March 12 2026) — 823 points
  • License: Apache 2.0 (~1,781 stars at launch)
  • Isolation: macOS sandbox-exec (Seatbelt) profiles — kernel-level syscall interception, no container.
  • Locality: Local, macOS only.
  • Agent integration: Explicit multi-agent wrapper — Claude Code, OpenAI Codex, Gemini CLI, Cline, Aider. Usage: safehouse claude --dangerously-skip-permissions.
  • Config: Shell functions or custom sandbox-exec profile files; LLM-assisted profile generation supported.
  • Network policy: Not addressed.
  • Notable from HN thread: Creator acknowledged the project is "just a policy-generator for sandbox-exec — no dependencies, no daemons, no subscription; I did put in many hours to identify the minimum required permissions for agents to continue working." Simon Willison noted that evaluating whether a sandboxing tool actually works as intended is hard. Top community sentiment: "I honestly think that sandboxing is currently THE major challenge that needs to be solved for the tech to fully realise its potential." The macOS Docker gap (Docker for Mac runs inside a Linux VM, so sandbox-exec is the only native primitive for bare-metal macOS processes) was the stated motivation.
  • Maturity: Active through March 2026.

matchlock

  • Source: https://github.com/jingkaihe/matchlock
  • License: MIT (~574 stars, v0.2.10)
  • Isolation: MicroVMs — Firecracker on Linux, Apple Virtualization.framework on macOS. Transparent proxy via nftables DNAT (Linux) or gVisor userspace TCP/IP (macOS).
  • Locality: Local (Homebrew, .deb, .rpm).
  • Agent integration: Agent-agnostic; SDK examples for Anthropic Claude API and OpenAI. Go, Python, TypeScript SDKs.
  • Config: CLI flags (--allow-host, --secret, --no-network) or SDK builder pattern. No manifest file.
  • Network policy: Default-deny + per-host allowlist.
  • Notable: Secrets injected in-flight by the host proxy — they never enter the VM.
  • Maturity: Marked experimental.

tilde.run

  • Source: https://tilde.run/
  • License: Proprietary, hosted SaaS.
  • Isolation: Cloud-hosted containers; underlying mechanism not publicly stated (unverified whether OCI containers or microVMs).
  • Locality: Hosted only.
  • Agent integration: Claude orchestration explicit; CLI (tilde exec) and Python SDK; plain-English agent instructions.
  • Config: DSL for RBAC policies (allow / deny / require human approval per action, per repo, per agent).
  • Network policy: Default-deny with per-request logging; cloud metadata endpoints and private networks blocked.
  • Persistence: All changes versioned and rollback-able via lakeFS; atomic commits per run.
  • Maturity: Private preview, © 2025, built by the lakeFS team.

boxlite

  • Source: https://boxlite.ai/ ; https://github.com/boxlite-ai/boxlite
  • License: Apache 2.0 (~4,700 stars, YC-backed)
  • Isolation: MicroVMs with dedicated Linux kernel per box — KVM on Linux, Hypervisor.framework on macOS. Not containers/namespaces.
  • Locality: Local, no daemon.
  • Agent integration: Explicitly targets AI agents; MCP server companion (boxlite-ai/boxlite-mcp). Pivoted from dev environments in 2025.
  • Config: SDK only — Python, Node.js, Rust, C; Go pending. No declarative manifest.
  • Network policy: "Isolated Network per VM" — details not public (unverified).
  • Notable: Sub-50ms boot, snapshot / fork / clone of VM state. Self description: "the SQLite of sandboxing".
  • Maturity: Active, YC.

microsandbox

  • Source: https://github.com/microsandbox/microsandbox (the superradcompany/microsandbox URL redirects to the same project).
  • License: Apache 2.0 (~6,000 stars, YC-backed)
  • Isolation: MicroVMs via libkrun, OCI-compatible images. Sub-100ms boot, rootless, no daemon, embeddable as a library.
  • Locality: Local.
  • Agent integration: Explicit Claude Code + Cursor targeting via "Agent Skills" packages and an MCP server. Agents can create their own sandboxes programmatically.
  • Config: CLI (msb), SDKs (Rust, Python, TypeScript), MCP server.
  • Network policy: Not detailed in public docs.
  • Maturity: Beta, breaking changes expected; most-starred project in this set.

smolmachines

  • Source: https://smolmachines.com/ ; https://github.com/smol-machines/smolvm
  • License: Apache 2.0 (~3,100 stars)
  • Isolation: MicroVMs via libkrun — Hypervisor.framework on macOS, KVM on Linux. No shared kernel.
  • Locality: Local, no daemon.
  • Agent integration: Includes an AGENTS.md; designed with coding agents in mind but no MCP/Skills turnkey integration.
  • Config: TOML Smolfiles declaring image, networking, volumes, SSH agent access, GPU acceleration. Portable .smolmachine files.
  • Network policy: Off by default; per-host allowlist via --allow-host.
  • Persistence: Named machines persistent by default; ephemeral runs also supported.
  • Maturity: Active through April 2026.

E2B (added 2026-07-20)

  • Source: https://github.com/e2b-dev/e2b ; https://e2b.dev/docs
  • License: Apache 2.0 (~12.4k stars); commercial hosted service with self-hosting/BYOC support.
  • Isolation: Firecracker microVM per sandbox.
  • Locality: Cloud-hosted by default; self-hosting uses Terraform on AWS or GCP (with other targets documented as works in progress).
  • Agent integration: LLM-agnostic Python and JavaScript/TypeScript SDKs; code-interpreter and desktop-sandbox products. Platform primitive rather than a coding-agent wrapper.
  • Config: Programmatic SDK/API plus templates. Network configuration supports internet on/off, outbound allow/deny rules, and a custom egress proxy.
  • Network policy: Configurable per sandbox, but not documented as default-deny and no built-in outbound-content DLP is documented.
  • Credentials: Environment variables passed to the sandbox are explicitly not private at the OS level. No built-in in-flight application-credential injection is documented.
  • Persistence: Full memory + filesystem pause/resume, snapshots, and auto-resume. Continuous runtime is tier-limited, while paused sandboxes are retained indefinitely.
  • Maturity: Established hosted platform and the API compatibility target used by CubeSandbox.

Daytona (added 2026-07-20)

  • Source: https://github.com/daytonaio/daytona ; https://www.daytona.io/docs/
  • License: Current production platform is proprietary. The former AGPL repository remains public but is no longer maintained after Daytona moved production development closed-source in June 2026.
  • Isolation: Hosted container sandboxes by default, with separate Linux and Windows VM sandbox classes for dedicated-OS workloads. Each sandbox has its own filesystem and network stack; VM-only features include memory pause/resume and forking.
  • Locality: Hosted multi-tenant service, with dedicated/custom regions and customer runners available.
  • Agent integration: LLM/framework-agnostic SDKs (Python, TypeScript, Go, Ruby, Java), API, and CLI; official agent-framework guides. Platform primitive rather than a local coding-agent wrapper.
  • Config: Programmatic per-sandbox image/snapshot, resources, lifecycle, firewall, and secrets.
  • Network policy: Per-sandbox IPv4/domain allowlists and block-all mode, subordinate to organization/tier policy. Full internet access is the default on higher tiers, so it is configurable rather than uniformly default-deny.
  • Credentials: First-class secret manager with the same phantom-token pattern as bot-bottle: the sandbox environment gets an opaque placeholder, an HTTPS proxy substitutes the real secret in headers only for allowed hosts, and responses are scrubbed back to the placeholder.
  • Persistence: Persistent filesystem for stopped container sandboxes; memory + filesystem pause/resume for VM sandboxes; snapshots and configurable auto-stop.
  • Maturity: Production commercial platform. Notable April 2026 credential exposure was patched; the June 2026 closed-source transition materially changes its transparency/self-hosting posture.

Other hosted runtimes carried forward from the earlier survey

  • Northflank Sandboxes — hosted or customer-cloud, microVM-backed containers with SDK-managed lifecycle, optional persistent volumes, and sub-second claimed boot. This is a platform primitive for untrusted code and agents, not a local agent wrapper or role-policy layer.
  • Cloudflare Sandbox SDK — Workers/Durable Objects API over VM-isolated Linux containers for command, file, process, and service execution. It is a hosted TypeScript platform primitive; application authentication, authorization, and credential-proxy patterns remain the integrator's job.

Both belong to the same “build your agent platform on this runtime” category as E2B and Daytona. They were named but not analyzed in depth by the original Claude-specific note, so they remain outside the main comparison table rather than being presented with false precision.

CubeSandbox (added 2026-07-18)

  • Source: https://github.com/TencentCloud/CubeSandbox ; HN launch https://news.ycombinator.com/item?id=47863430
  • License: Apache 2.0 (~10.4k stars). By Tencent Cloud; described as "battle-tested, production-ready" infra already running in Tencent Cloud. Rust / Go / C.
  • Isolation: MicroVMs via RustVMM + KVM — "each sandbox gets its own Guest OS kernel, no Docker shared-kernel escapes." Hardware-level isolation, dedicated kernel per instance.
  • Locality: Self-hosted, but server/cluster-oriented, not a single-user local CLI. Deploy guides target PVM cloud VMs, bare metal, and dev. A single 96-vCPU host is claimed to run 2,000+ concurrent sandboxes.
  • Agent integration: Drop-in E2B SDK replacement (single env-var change) — the headline compatibility claim. OpenClaw assistant integration; general LLM-code execution. Aimed at platform builders, not one developer's laptop.
  • Config: Programmatic via the E2B-compatible SDK. No declarative manifest.
  • Network policy: This is the striking part — domain allowlists, instant block on unauthorized egress, full audit logs, per-sandbox traffic tokens, policy-routing egress, enforced by an eBPF-based virtual switch giving kernel-level network isolation. Closest match yet to bot-bottle's own default-deny + per-bottle allowlist egress model.
  • Credentials: Credential vault — agents call external APIs / LLMs while "keys never enter the sandbox, model context, or logs." Same in-flight-injection idea as matchlock, but productized as a vault.
  • Performance: <60ms cold start (claimed 2.550× faster than alternatives), <5MB memory per instance; millisecond snapshot rollback is upcoming.
  • Maturity: Open-sourced July 2026 off production Tencent Cloud use; most-starred project in this set (~10.4k).

Cleanroom (added 2026-07-18)

  • Source: https://github.com/buildkite/cleanroom
  • License: Apache 2.0
  • Isolation: MicroVM — Firecracker on Linux, Virtualization.framework on macOS. Digest-pinned OCI images.
  • Locality: Self-hosted server (CI-oriented).
  • Agent integration: Generic process sandbox; CI-first, not a Claude/agent wrapper.
  • Config: cleanroom.yaml in the repo being sandboxed defines egress rules, resources, and network policy. Cleanroom resolves this from the commit being run.
  • Network policy: Default-deny + per-repo hostname allowlist (resolved from DNS answers + destination IP:port). Co-hosted services on the same IP:port are not distinguished. OIDC-backed auth for remote servers.
  • Credentials: Host-side only; not injected in-flight but not present in the VM.
  • Notable: Policy lives in the repo being sandboxed, not in an agent-role definition — closer to per-repo scoping than per-role. Supports Docker-inside-sandbox (services.docker.required: true), OIDC authorization, suspend/resume lifecycle.
  • Maturity: Active Buildkite product.

container-use (added 2026-07-18)

  • Source: https://github.com/dagger/container-use
  • License: Apache 2.0
  • Isolation: Docker container per agent + git worktree per agent. Containers share the host kernel; stronger than bare host but weaker than microVM.
  • Locality: Local.
  • Agent integration: MCP stdio server — Claude Code, Cursor, Windsurf. claude mcp add container-use -- container-use stdio.
  • Config: None for security policy. Environments are provisioned on demand; no allowlist or credential config.
  • Network policy: Not addressed.
  • Notable: Per-agent git branches (container-use/<env_name>); parallel agents without filesystem conflict; real-time log visibility and terminal attach for intervention; git-based review workflow. Oriented toward parallel development safety, not security containment.
  • Maturity: Early development, active.

Docker sbx (added 2026-07-18)

  • Source: Docker proprietary (sbx CLI, separate from docker).
  • License: Proprietary.
  • Isolation: MicroVM (Docker's own implementation) — each session gets its own kernel, Docker daemon inside the VM, and filesystem.
  • Locality: Local (macOS and Windows; does not require Docker Desktop).
  • Agent integration: Explicit wrapper — Claude Code, Codex, Gemini CLI, Copilot CLI, Kiro. Launches agent inside the VM with --dangerously-skip-permissions by default.
  • Config: Open / Balanced / Locked Down network presets at launch. No per-role manifest.
  • Network policy: Default-deny; preset levels control strictness. TUI dashboard shows a live log of every outbound connection (allowed and blocked) with point-and-click allow/block for hosts.
  • Credentials: OS keychain + host-side proxy injection — API keys never enter the VM.
  • Notable: Best DX among microVM tools (one command, works like native yolo Claude but inside a VM); branch mode creates a git worktree in .sbx/. Network policy is preset-based, not role-declarative.
  • Maturity: GA 2026.

Anthropic srt (added 2026-07-18)

  • Source: https://github.com/anthropic-experimental/sandbox-runtime (@anthropic-ai/sandbox-runtime on npm, sandbox-runtime on PyPI)
  • License: Apache 2.0 (experimental).
  • Isolation: OS-level only — Seatbelt (sandbox-exec) on macOS, bubblewrap on Linux, WFP (Windows Filtering Platform) account-fenced on Windows. No container or VM. Lowest overhead in the set.
  • Locality: Local.
  • Agent integration: Claude Code's sandboxed bash tool uses this internally. Can wrap any arbitrary process (srt <command>). Cloud Claude Code sessions use full microVMs instead.
  • Config: Programmatic per-invocation — allow/deny path lists for filesystem; allow/denylist for network (HTTP proxy + SOCKS5).
  • Network policy: Proxy-based filtering (HTTP + SOCKS5); domain allowlist/denylist enforced at proxy layer. Custom proxy supported (e.g. mitmproxy for inspection + audit). Processes that ignore proxy env vars may bypass filtering on some platforms.
  • Notable: Cross-platform (macOS/Linux/Windows); wraps any process, not just agents; no role/manifest concept. Annotated as a research preview — APIs may change.
  • Maturity: Early research preview.

Claude-specific wrappers and developer environments

These projects were the focus of the original containerized-Claude survey. They remain useful comparisons for local developer experience, but most are templates or wrappers rather than policy-bearing sandbox platforms, so they are grouped here instead of widening the main table further.

claudebox

  • Source: https://github.com/RchGrav/claudebox
  • Isolation: Docker, with per-project images, authentication state, and configuration.
  • Agent integration: Claude Code wrapper with 15+ preconfigured language and task profiles.
  • Network policy: Per-project firewall allowlists.
  • Closest overlap: local one-command developer workflow and project-scoped network policy.
  • Difference: profiles describe development toolchains, not named agent roles. There is no bottle/agent split, composable role manifest, provider plugin layer, or outbound-content DLP.

Spritz / claude-code-sandbox

  • Source: https://github.com/textcortex/claude-code-sandbox (archived; points to its successor, Spritz).
  • Isolation: The original project ran Claude Code in local Docker with bypass permissions; Spritz moved toward Kubernetes-native multi-agent infrastructure.
  • Difference: the successor targets cluster orchestration rather than a low-dependency local launcher. It is architecturally closer to hosted or Kubernetes platform runtimes than to bot-bottle's single-operator CLI.

Trail of Bits claude-code-devcontainer

  • Source: https://github.com/trailofbits/claude-code-devcontainer
  • Isolation: A Docker devcontainer that exposes only project files and is designed to run Claude Code with bypassPermissions for security audits and untrusted-code review.
  • Difference: a hardened, reusable environment definition rather than an agent launcher or fleet. It has no named-role manifest, per-role credential custody, supervision plane, or multi-backend abstraction.

Smaller wrappers and official templates

Projects such as arezi/claude-sandbox, nkrefman/claude-sandbox, and VishalJ99/claude-docker, plus Docker/Anthropic devcontainer templates, prove there is steady demand for “Claude in a container.” They are deliberately small launch/build configurations. They compete on setup simplicity, not on role-aware policy, credential custody, persistent supervision, or a fleet model, and are better treated as a product category than as individual rows.

SuperHQ

  • Source: https://superhq.ai/
  • Isolation: Apple-Silicon desktop application using local microVMs via Virtualization.framework/libkrun-era components.
  • Agent integration: Claude Code, Codex, and Pi in a GUI, with mobile remote access.
  • Credentials and review: host-side auth gateway injects credentials on the wire; a temporary overlay stages writes for diff-and-accept review.
  • Closest overlap: local microVM isolation, multi-provider launching, and credential custody for security-minded individual developers.
  • Difference: GUI desktop product on Apple Silicon rather than a cross-platform declarative CLI/fleet layer. The July 2026 snapshot in the original survey recorded a user request for per-run tool-call and network audit logging; treat that as point-in-time rather than a permanent gap.

Credential gateway without isolation

OneCLI

OneCLI is a framework-agnostic identity gateway rather than a sandbox. Its phantom-token design gives the agent a placeholder and substitutes the encrypted real credential at the network layer. It therefore matches bot-bottle closely on secret custody, and is more portable because it can sit in front of agents launched by anything, but it supplies no container or VM boundary, filesystem isolation, role manifest, or egress-content DLP.

The positioning consequence from the earlier survey still holds: secret custody alone is not unique. bot-bottle's relevant combination is local isolation + default-deny egress + payload DLP + declarative roles + credential custody. OneCLI's managed tier also places custody with a third party, whereas bot-bottle keeps it within operator-controlled infrastructure. See agent-credential-proxy-landscape.md for the detailed build-versus-adopt analysis.

Governance / pre-action authorization layers

These two tools don't provide VM or filesystem isolation; they intercept tool calls before execution and evaluate them against a per-agent declarative policy. They are the closest competitors on agent-tailored policy and complement isolation sandboxes rather than substituting for them.

Microsoft Agent Governance Toolkit (AGT) (added 2026-07-18)

  • Source: https://github.com/microsoft/agent-governance-toolkit
  • License: MIT (~3.3k stars, open-sourced April 2, 2026).
  • Isolation: None (OS/VM). Execution rings (03, inspired by CPU privilege levels) control what an agent can do at the framework layer. MCP security gateway treats MCP traffic as an untrusted boundary.
  • Locality: Embedded in the agent framework (Python, TypeScript, .NET, Rust, Go; 20+ framework adapters).
  • Agent integration: Framework-agnostic. Plugs into Semantic Kernel, AutoGen, and others as a middleware layer.
  • Config: YAML policy per agent — tools can be allowed, denied, sandboxed, or routed through an approval step. Every action passes through a governance gate checking: agent DID, trust score, risk tier, requested tool, action type, and policy rules.
  • Network policy: Not directly — operates at tool-call level.
  • Credentials: Per-agent DID (Ed25519 decentralized identifier); agent does not borrow a human's credentials.
  • Notable: Dynamic trust score (01,000, behavioral decay) — privilege follows observed behaviour, not just provisioning. Covers all 10 OWASP Agentic Top 10 risks. Kill switch + SLO monitoring. Sub-ms policy enforcement.
  • Maturity: MIT, ~3.3k , v3.7.0 May 2026.

Open Agent Passport (OAP) (added 2026-07-18)

  • Source: https://github.com/aporthq/aport-spec ; spec at https://api.aport.io/spec/spec/oap/oap-spec.md/ ; arXiv 2603.20953
  • License: Open specification.
  • Isolation: None. Pre-action hook only — intercepts tool calls synchronously before execution, evaluates against a cloud-registry declarative policy, fails closed.
  • Locality: Local hook + cloud policy registry.
  • Agent integration: Framework-agnostic; hook pattern.
  • Config: Declarative policy rules in a cloud registry (evaluated in order; first failing rule denies). Ed25519-signed, hash-chained audit records per decision.
  • Network policy: Not directly.
  • Notable: 53ms median authorization decision (N=1,000). In an adversarial testbed ($5,000 bounty, 1,151 sessions), social engineering succeeded 74.6% of the time under a permissive policy; under a restrictive OAP policy, 0% success across 879 attempts. Assumes framework runtime is not compromised.
  • Maturity: Specification + reference implementation, 2026.

Comparison table

Isolation/sandbox tools only. AGT and OAP are governance layers — see their per-project notes above.

Axis bot-bottle endo-familiar litterbox agent-safehouse matchlock tilde.run boxlite microsandbox smolmachines E2B Daytona CubeSandbox Cleanroom container-use Docker sbx Anthropic srt
Isolation MicroVM per bottle default (Firecracker/KVM on Linux, Apple Container on macOS) + own egress DLP scanner; Docker legacy fallback, gVisor there if present Object-capability (no OS isolation) Podman + opt. Landlock macOS sandbox-exec MicroVM (Firecracker / Virt.fw) Hosted container (unverified) MicroVM (KVM / Hypervisor.fw) MicroVM (libkrun) MicroVM (libkrun / KVM) Firecracker microVM Container or Linux/Windows VM class MicroVM (RustVMM / KVM) MicroVM (Firecracker / Virt.fw) Docker container + git worktree MicroVM (proprietary) OS-level (Seatbelt / bubblewrap / WFP) — no container
Local vs hosted Local Local Local (Linux) Local (macOS) Local Hosted SaaS Local Local Local Hosted; self-host/BYOC available Hosted; dedicated/custom regions Self-hosted (server/cluster) Self-hosted server Local Local Local
Open source Apache 2.0 Apache 2.0 Apache 2.0 Apache 2.0 MIT No Apache 2.0 Apache 2.0 Apache 2.0 Apache 2.0 Production closed-source; legacy AGPL repo unmaintained Apache 2.0 Apache 2.0 Apache 2.0 Proprietary Apache 2.0 (experimental)
Agent target Claude Code, Codex, Pi, and provider plugins Generic (demo) Generic Multi-agent wrapper Generic (+ Claude/OpenAI SDKs) Claude focus Generic Claude + Cursor (MCP/Skills) Generic (AGENTS.md) LLM-agnostic platform builders LLM-agnostic platform builders E2B-compatible (platform builders) CI / generic process Claude Code, Cursor, Windsurf (MCP) Claude Code, Codex, Gemini CLI, Copilot, Kiro Claude Code (and any process)
Network policy Default-deny via own egress scanner + per-bottle allowlist + content DLP + gitleaks on git push Capability model only Limited Not addressed Default-deny + allowlist + secret-injecting proxy Default-deny + logging Per-VM net (unverified) Not documented Off by default + allowlist Per-sandbox allow/deny rules and custom egress proxy; internet configurable Per-sandbox CIDR/domain allowlist or block-all; tier policy; secret-injecting proxy Default-deny allowlist + instant egress block + audit logs + per-sandbox tokens (eBPF) + credential vault Default-deny + per-repo host allowlist (cleanroom.yaml) Not addressed Default-deny; Open / Balanced / Locked Down presets; live TUI network panel Proxy-based allowlist/denylist (HTTP + SOCKS5); custom proxy supported
Parallel agents Yes (one bottle per agent) n/a Not addressed One at a time Multiple VMs Yes (dashboard) SDK-level SDK-level Architectural Yes (platform service) Yes (platform service) Yes (2,000+/host claimed) Yes (server model) Yes (per-agent containers + worktrees) Yes Yes
Long-running posture Persistent by default (named, supervised) n/a (demo) Session (up while in use) Per-invocation Ephemeral VM per run Per-run (versioned) Ephemeral + snapshot/fork Ephemeral / on-demand Named persistent by default Runtime tier limits + indefinite pause/resume Persistent filesystem; VM pause/resume; configurable auto-stop Ephemeral + auto pause/resume Per-run + suspend/resume Per-agent container (ephemeral) Per-session; branch mode creates git worktree in .sbx/ Per-invocation
DX: run Claude yolo-style One command → interactive yolo Claude (start <agent>, --dangerously-skip-permissions default) n/a (lib demo) Wizard + build, then run claude inside (Linux only) One-command wrapper (safehouse claude --dangerously-skip-permissions) CLI: run a cmd in a VM (not a Claude wrapper) Hosted (tilde exec), not local-native SDK code required (build the run yourself) CLI/MCP: sandbox-as-a-tool for the agent, not a wrapper around it SSH into a named machine, run claude there SDK/CLI sandbox; wire the agent yourself SDK/CLI sandbox; wire the agent yourself Stand up a cluster + drive via E2B SDK CI-oriented, not a Claude wrapper MCP server: claude mcp add container-use -- container-use stdio One command: sbx wraps claude with --dangerously-skip-permissions default Library/wrapper, not a standalone CLI
Config YAML-in-Markdown manifests (bottles + agents) Programmatic refs CLI wizard Profile files / shell fns CLI / SDK DSL + CLI + SDK SDK CLI / SDK / MCP TOML Smolfile SDK/API + templates SDK/API/CLI + images/snapshots E2B-compatible SDK cleanroom.yaml in repo None (no policy config) Preset levels at launch Programmatic per-invocation (allow/deny lists)
Agent-tailored policy Yes — bottle/agent split; declarative per-role egress + credentials; composable via extends: Partial — capability model scopes per-agent, but no declarative role manifest No Partial — per-agent profile files (Seatbelt); no egress No Yes — per-agent DSL RBAC (allow/deny/approve per action/repo/agent) No No No No — per-sandbox SDK config No — per-sandbox SDK config No — per-sandbox SDK config, not role-scoped Partial — per-repo cleanroom.yaml, not per-role No No — network presets only No
Maturity Active July 2026 Research (2022+) Early (~66 ) Active (~1.8k ) Experimental (~574 ) Private preview YC, ~4.7k YC, ~6k , beta ~3.1k Established hosted platform, ~12.4k Production commercial; closed-source since June 2026 Tencent, prod, ~10.4k Active (Buildkite product) Early development GA 2026 Early research preview

What's closest, what's different

Closest in design and scope. agent-safehouse and litterbox sit nearest bot-bottle: local, single-user, thin wrappers over an existing OS primitive, low-dep. The split is the isolation primitive — bot-bottle now defaults to a VM per bottle (Firecracker microVM on KVM Linux, Apple Container on macOS) with its own DLP-scanning egress proxy, keeping Docker only as a legacy fallback; agent-safehouse uses sandbox-exec; litterbox uses Podman + Landlock. matchlock and smolmachines are close on both the policy side (default-deny net, per-host allowlist) and — now that bot-bottle has moved off containers-by-default — the microVM isolation primitive. Note: Apple Container 1.0 stable shipped June 9 2026 (frozen CLI and APIs), which makes the macOS backend stable surface area rather than a moving target.

New closest on agent-tailored policy. Two governance tools are the direct competitors on the "coarse-grained sandbox" axis. tilde.run has had per-agent DSL RBAC since its launch (though it's hosted SaaS). Microsoft AGT is the most serious new entrant: per-agent DID identity, YAML policy that can allow/deny/sandbox/approve individual tool calls per agent, and a dynamic behavioural trust score. It operates at the framework tool-call layer, not the network layer — so it's complementary to bot-bottle's network/filesystem isolation rather than a direct substitute, but on the "does this sandbox know what this agent is for?" question it is the most complete answer in the field. OAP's pre-action hook pattern achieves similar goals with cryptographic audit and a 0% adversarial-attack success rate under a restrictive policy.

New closest on DX. Docker sbx is the first tool in this set that matches bot-bottle on the "one command, dangerously-skip-permissions safe by default" DX bar, at microVM isolation strength, with host-side credential injection. It is proprietary, preset-based (not role- declarative), and cloud-agent-specific, but it directly competes on the UX proposition. agent-safehouse was the previous DX peer; Docker sbx materially raises the bar.

New closest on repo-scoped policy. Cleanroom (Buildkite) is the first tool to combine microVM isolation with a declarative egress policy file — though the policy lives in the repo being sandboxed (cleanroom.yaml), not in an agent-role manifest. That makes it per- repo rather than per-role: the same Cleanroom config applies to any agent running in that repo. The distinction matters for bot-bottle's use case (one developer running multiple agent roles with different egress footprints), but for CI/CD use cases Cleanroom is a direct alternative.

Solving a different problem. tilde.run is hosted SaaS for team / production agent pipelines with data-versioned rollback — explicitly opposite to bot-bottle's "infrastructure I control" goal. E2B and Daytona are hosted sandbox platforms, while boxlite, microsandbox, and CubeSandbox are infrastructure libraries/services aimed at platform builders embedding sandboxes into agent frameworks; they would be a backend bot-bottle could call, not a competitor to its manifest layer. endo-familiar is in a different paradigm entirely: capability passing rather than kernel boundaries.

Borrowable ideas

Already shipped or otherwise addressed

  • Default-deny egress with a per-agent allowlist (own egress scanner).
  • DLP scanning of outbound traffic.
  • Bottle / agent split (manifest layer above the isolation primitive).
  • gVisor auto-detection on Linux.
  • In-flight secret injection (suggested by matchlock) — shipped. Real provider and git-host tokens are held outside the agent and injected by the egress gateway on matching routes. The agent receives only proxy URLs and, where a client requires a credential-shaped value, a placeholder; GITEA_TOKEN and equivalent real tokens do not appear in the agent's environment.
  • MicroVM backendshipped. MicroVMs are now the default: Firecracker on KVM Linux and Apple Container on macOS. Docker is the legacy fallback.
  • Per-use SSH key confirmation (suggested by litterbox) — addressed by stronger credential custody instead. The agent does not hold the upstream git SSH key or an SSH-agent socket: git-gate holds the credential and gates git operations. A confirmation wrapper inside the agent would therefore protect a credential that is no longer there. Operator approval at the gate remains the appropriate control point for any future per-use confirmation.

Still worth considering

  • Live network activity in the supervisor TUI (from Docker sbx): show allowed and blocked connections and let the operator propose policy changes from the existing supervision surface.
  • Tamper-evident audit records (from OAP): sign and hash-chain egress and supervision decisions for compliance-sensitive deployments.
  • Behaviour-informed policy downgrade (from Microsoft AGT): use repeated DLP alerts or supervision holds as a signal to narrow policy or request closer review. This needs a carefully specified trust model before it can be more than a heuristic.

Not worth borrowing: the SDK-first programmatic API style of boxlite / microsandbox (cuts against the declarative-manifest stance), and the hosted-SaaS dashboard model of tilde.run (cuts against the "infrastructure I control" goal).

Publishing and positioning verdict

Publishing remains worthwhile, but the defensible claim is the combination, not any single primitive. Credential custody is matched by OneCLI, matchlock, Daytona, Docker sbx, and CubeSandbox; local one-command isolation is matched by agent-safehouse and Docker sbx; hosted microVM execution is a crowded platform category.

bot-bottle remains unusual in combining:

  • local, operator-controlled execution with persistent named bottles;
  • one declarative role layer across Claude Code, Codex, Pi, and provider plugins;
  • composable agent/bottle manifests, skills, and system prompts;
  • Firecracker/Apple Container isolation with a Docker fallback;
  • default-deny per-role egress, payload DLP, and git-push secret scanning;
  • credentials injected outside the agent process; and
  • supervision and audit state suited to long-running parallel agents.

The practical wedge is “as easy as native yolo, with declarative role policy and self-hosted custody,” including scoped access to private LAN/Tailnet services that cloud-first runtimes cannot provide without additional network plumbing. The main competitive risks are a local wrapper such as claudebox or Docker sbx growing a role-manifest layer, and GUI products such as SuperHQ adding equivalent policy and audit depth.

Caveats

  • Star counts and last-commit dates are point-in-time snapshots.
  • Several projects' network and persistence behaviour is not documented publicly; items so derived are marked (unverified).
  • The superradcompany/microsandbox URL in the original prompt redirects to microsandbox/microsandbox; the surveyed project is the same.
  • CubeSandbox performance/scale numbers (<60ms cold start, <5MB/instance, 2,000+ sandboxes per 96-vCPU host) are the project's own launch claims, not independently verified here.

Addendum 2026-07-18 — CubeSandbox and the positioning read

CubeSandbox (Tencent Cloud, Apache 2.0, ~10.4k stars, HN launch #47863430) is the first open-source, self-hostable project in this survey to combine, in one stack, the main primitives bot-bottle treated as its differentiator:

  • Egress custody (connection level) — default-deny domain allowlist (L7 domain/SNI filtering), instant block on unauthorized egress, per-sandbox traffic tokens, full audit logs of destinations (eBPF virtual switch, "CubeVS"). This matches bot-bottle's egress scanner at the connection level, productized — see the one thing it does not match, below.
  • Credential custody — a vault where keys "never enter the sandbox, model context, or logs." This is the in-flight-injection idea from matchlock, but as a first-class feature, and it's exactly the cross-vendor "egress audit + custody" wedge the monetization positioning treats as the one defensible moat.
  • Isolation on par with bot-bottle's current default — a dedicated guest kernel per sandbox (RustVMM/KVM). bot-bottle now defaults to the same class of boundary (Firecracker microVM / Apple Container), so this is parity, not an edge; CubeSandbox's remaining edge is running that per-kernel isolation multi-tenant at scale on one host.

The one axis CubeSandbox does not cover — and where bot-bottle stays distinctive:

  • Content DLP on authorized channels. CubeSandbox's egress control is connection-level: it decides whether a destination is allowed and logs it, and its vault keeps injected credentials out of the sandbox entirely. Neither inspects the payload of traffic to an allowed destination. So an agent that exfiltrates over a permitted channel — pasting a repo's contents, an agent-derived secret, or PHI into an allowed API/domain — is not caught by CubeSandbox. bot-bottle's own egress DLP scanner does scan that: response + websocket content against the resolved per-flow config, with per-bottle token redaction (see recent egress commits). The vault approach is arguably stronger for the specific case of pre-known injected credentials (they can't leak if they were never present), but it is not a substitute for content inspection of everything else.

Long-running posture — a sharper axis than raw isolation. E2B and CubeSandbox are ephemeral-per-task by design; a long-running agent is an architected pattern on top, not the default. E2B: 5-minute default timeout, continuous runtime tier-capped (~1h Hobby / ~24h Pro), duration achieved via pause/resume (preserves filesystem + memory + processes; reconnect by sandbox ID via Sandbox.connect(); resume resets the timeout to 5 min; auto-pause via on_timeout: "pause"). CubeSandbox mirrors this (E2B drop-in) with first-class auto pause/resume and hundred-ms checkpoint/fork — and, self-hosted, sets its own timeout policy with no vendor tier caps. bot-bottle inverts the model: a bottle is persistent, named, and supervised by default — long-running is the default, not a session-management loop over pause/resume. smolmachines is the other persistent-by-default project in this set. For anyone building agents that run for hours/days, this posture difference matters more than the isolation primitive.

DX — the "run Claude yolo-style" bar. The reason claude --dangerously-skip-permissions is so widely used is DX: it's one command and the agent just goes. The bottle thesis is to make a sandboxed run that easy — start <agent> builds the image on first run and drops you into an interactive Claude session that already has --dangerously-skip-permissions on by default (contrib/claude/agent_provider.py), with the sandbox as the guardrail instead of per-action prompts. On this axis the field splits cleanly:

  • Wrappers around the agent (as-easy-as-native): bot-bottle and agent-safehouse (safehouse claude --dangerously-skip-permissions). These are the run-Claude experience. agent-safehouse is the real DX peer — but it's macOS-only Seatbelt, single-run, and doesn't address network egress; bot-bottle adds VM-grade isolation, egress DLP, and persistent/parallel bottles across macOS + Linux.
  • Libraries / services (you build the run yourself): boxlite, microsandbox, CubeSandbox, E2B, Daytona. These hand you an SDK or a cluster and expect you to wire the agent in — powerful for platform builders, heavyweight for "just run Claude on my laptop." microsandbox's MCP/Skills angle is sandbox-as-a-tool the agent calls, which is the inverse of wrapping the agent.
  • In between: litterbox (wizard + build, Linux only), smolmachines (SSH into a named machine), matchlock (run a command in a VM).

So DX is a genuine bot-bottle differentiator. agent-safehouse matches the one-command wrapper with weaker isolation and no egress story; Docker sbx now matches it at microVM strength but remains proprietary and preset-based. "As easy as native yolo, with declarative role policy" is the narrower defensible one-liner.

Why it still doesn't collide head-on:

  1. Shape. CubeSandbox is a multi-tenant service for platform builders (drop-in E2B replacement, SDK-driven, 2,000 sandboxes on a box). bot-bottle is a single-operator, declarative-manifest tool for the infrastructure I run. Different buyer, different ergonomics — no declarative role manifest, no bottle/agent split, no "one command on my laptop."
  2. Backend, not competitor. Like boxlite/microsandbox, CubeSandbox is something bot-bottle could sit on top of — a "runtime": "microvm" or "runtime": "cubesandbox" backend under the manifest layer — while keeping the manifest, the bottle/agent split, and the local, single-operator default.

Why it matters anyway:

  • The "nobody else bundles connection-level egress allowlist + audit + in-flight credential custody" line is no longer true for the primitive — CubeSandbox ships the open-source/self-hosted combination, and Daytona ships a proprietary firewall + credential-substitution variant. But content DLP on authorized channels is still not matched (see above), and neither is the layer above the primitive (declarative manifest, cross-vendor orchestration, operator UX, the phone-control/dashboard north star). Those two — outbound-payload DLP and the orchestration layer — are where the defensible ground now sits; the connection-level allowlist + vault mechanism, on its own, is no longer differentiating. Revisit the monetization open/paid line with that in mind.
  • Worth a closer look at how CubeSandbox does credential injection and per-sandbox egress tokens (eBPF virtual switch vs. bot-bottle's mitmproxy egress proxy) when hardening bot-bottle's now-shipped credential-custody implementation.

Addendum 2026-07-18 (second pass) — agent-tailored policy landscape

The second-pass question was: how novel is bot-bottle's per-agent, role-tailored sandbox relative to the expanded field?

The short answer: on the isolation + network + role-tailoring combination, bot-bottle remains the only tool in this set. On role-tailored policy at the tool-call level, Microsoft AGT and OAP are the most complete answers, but they don't provide isolation; they complement rather than substitute.

The competitive picture by axis:

  • Agent-tailored egress (declarative, per-role) — bot-bottle and tilde.run. Cleanroom is per-repo, not per-role. Everyone else is per-session or not addressed.
  • Agent-tailored tool-call policy (declarative, per-agent identity) — Microsoft AGT (YAML policy + DID identity + trust score), OAP (declarative policy rules + cryptographic audit). Neither provides network/filesystem isolation.
  • Composable policy (role overlays) — bot-bottle (extends:). No other tool surveyed supports composable role-policy inheritance.
  • Isolation + DX (one-command safe yolo) — bot-bottle and Docker sbx. Docker sbx is proprietary, preset-based, and cloud-agent-specific; it's the first DX-class competitor at microVM isolation strength.

What the HN "coarse-grained" complaint maps to: The complaint is that a VM isolates the filesystem but doesn't know if the agent should be sending an email. bot-bottle's bottle/agent split is a structural answer to this: the bottle manifest declares exactly what the role can reach, and the sandbox enforces it at the network layer. Microsoft AGT is the most complete answer at the semantic/tool-call layer. The gap both leave open is intent classification — knowing whether a permitted action is consistent with the agent's actual task. See hn-agent-safety-discourse-july-2026.md for the blast-radius analysis.

Open ideas from new tools (also summarized above):

  • Microsoft AGT's trust-score decay — privilege that reflects observed behaviour rather than static provisioning. Applied to bot-bottle: a bottle that has triggered DLP alerts or supervise holds could auto-downgrade its network preset, or flag the session for closer review. Fits the existing supervise-server architecture.
  • Docker sbx's live network TUI — real-time per-session view of allowed and blocked outbound connections with point-and-click allow/block. cli.py supervise is the right surface; adding a live-connections panel would directly address the "I can't see what the agent is doing" gap without any backend changes.
  • OAP's cryptographic audit chain — Ed25519-signed, hash-chained audit records. Currently bot-bottle logs egress decisions but doesn't chain them. A tamper-evident audit record per session would be useful for the compliance use case the CubeSandbox positioning targets.