871 lines
51 KiB
Markdown
871 lines
51 KiB
Markdown
# Landscape: AI-agent sandbox tools
|
||
|
||
Survey of AI-agent sandbox and containment projects — including local
|
||
coding-agent wrappers, agent-agnostic runtimes, hosted platforms, and
|
||
governance layers — contrasted with bot-bottle's design. The original
|
||
Claude-Code-specific containerizer survey was folded into this note on
|
||
2026-07-20 so there is one landscape and one positioning verdict.
|
||
|
||
Research conducted 2026-05-11. CubeSandbox added 2026-07-18 (see its
|
||
per-project note and the addendum at the end). Also updated 2026-07-18:
|
||
bot-bottle no longer uses **pipelock** — outbound DLP is now bot-bottle's
|
||
own (deliberately simple) egress scanner (a mitmproxy addon with custom
|
||
detectors, PRD 0017 / 0052), and git-push secret scanning is handled by
|
||
**gitleaks** in the git-gate. "pipelock" below has been replaced with the
|
||
current mechanism; it survives only in older PRDs as history.
|
||
|
||
Updated again 2026-07-18: six additional tools added (Cleanroom,
|
||
container-use, Docker sbx, Anthropic srt, Microsoft AGT, Open Agent
|
||
Passport); an **Agent-tailored policy** row added to the comparison table;
|
||
a separate Governance layers section added for AGT and OAP. See the
|
||
second addendum at the end.
|
||
|
||
Updated 2026-07-20: the borrowable-ideas status was reconciled with the
|
||
current implementation. In-flight credential injection and the microVM
|
||
backends have shipped, while per-use SSH confirmation was superseded by
|
||
keeping git credentials out of the agent entirely.
|
||
|
||
Also updated 2026-07-20: **E2B and Daytona added as first-class entries.**
|
||
Earlier revisions mentioned E2B only as the API and lifecycle model that
|
||
CubeSandbox implements, and omitted Daytona entirely. That was a survey gap,
|
||
not a principled scope exclusion: both are major hosted sandbox platforms and
|
||
belong in this landscape even though they target platform builders rather than
|
||
bot-bottle's local single-operator workflow.
|
||
|
||
## Summary
|
||
|
||
The main table compares bot-bottle against fifteen isolation/sandbox tools.
|
||
Governance/pre-action authorization and credential-only layers are covered
|
||
separately because they don't provide VM or container isolation. None
|
||
duplicate bot-bottle's combination of local
|
||
VM-per-bottle isolation, a declarative per-role manifest, per-agent
|
||
egress allowlist + outbound-content DLP, bottle/agent split, and the
|
||
composable `extends:` policy model. Three clusters stand out:
|
||
|
||
- **Closest neighbours** — agent-safehouse and litterbox: local,
|
||
single-user, thin wrappers over an existing OS primitive
|
||
(`sandbox-exec`, Podman + Landlock).
|
||
- **Different category (isolation)** — tilde.run (hosted SaaS), boxlite
|
||
and microsandbox (microVM libraries for platform builders), E2B and Daytona
|
||
(hosted sandbox platforms), CubeSandbox (self-hosted multi-tenant microVM
|
||
service), endo-familiar
|
||
(capability-security paradigm, no OS isolation).
|
||
- **New: governance/pre-action layers** — Microsoft AGT and Open Agent
|
||
Passport (OAP): framework-embedded tool-call interceptors with
|
||
per-agent declarative policy. Closest competitors on agent-tailored
|
||
policy, but operate at the tool-call level rather than providing
|
||
network/filesystem isolation; they complement rather than substitute.
|
||
|
||
The microVM cluster (matchlock, smolmachines, boxlite, microsandbox,
|
||
CubeSandbox) is the most relevant for the v2 isolation discussion in
|
||
[`stronger-isolation-alternatives.md`](stronger-isolation-alternatives.md):
|
||
libkrun and Apple's Virtualization.framework have made local microVMs
|
||
ergonomic enough that microVMs are **now bot-bottle's default backend**
|
||
(Firecracker on KVM Linux, Apple Container on macOS), with Docker kept
|
||
only as a legacy fallback for CI / hosts without KVM or Apple Container.
|
||
That discussion has since shipped, not just been theorized.
|
||
|
||
**The one that matters most for positioning is CubeSandbox** — it ships
|
||
bot-bottle's bundle of default-deny egress allowlisting, full audit logs, and
|
||
in-flight credential custody *combined with* per-sandbox microVM isolation,
|
||
open-source under Apache 2.0, with Tencent Cloud behind it and 10.4k
|
||
stars. It's a self-hosted multi-tenant service for platform builders, not
|
||
a single-user declarative tool, so it doesn't collide head-on — but it
|
||
narrows the "nobody else bundles egress custody + credential injection"
|
||
claim that the monetization positioning leans on. Daytona now also offers
|
||
domain/CIDR firewall policy plus in-flight header credential substitution and
|
||
response scrubbing, although its higher tiers are not default-deny and its
|
||
production platform is proprietary. See the addendum.
|
||
|
||
## Per-project notes
|
||
|
||
### endo-familiar
|
||
- **Source**: https://dcfoundation.io/containing-ai-agents-the-endo-familiar-demo/ ; https://github.com/endojs/endo
|
||
- **License**: Apache 2.0
|
||
- **Isolation**: Object-capability runtime in Hardened JavaScript. Not
|
||
OS-level — agents simply cannot reference resources they were not
|
||
handed.
|
||
- **Locality**: Local / decentralized; WebSocket relay for capability
|
||
sharing across machines.
|
||
- **Agent integration**: Agent-agnostic, demo only.
|
||
- **Config**: Programmatic capability passing; "pet name" system for
|
||
human-readable capability handles.
|
||
- **Network policy**: Capability model is the policy; no allowlist or
|
||
firewall.
|
||
- **Maturity**: Research demo, Foresight Institute grant. Production use
|
||
of `endo` is via Agoric and MetaMask, not as a containment tool.
|
||
|
||
### litterbox
|
||
- **Source**: https://litterbox.work/ ; https://github.com/Gerharddc/litterbox
|
||
- **License**: Apache 2.0 (~66 stars)
|
||
- **Isolation**: Podman container on Linux + Wayland socket forwarding;
|
||
optional Landlock LSM for filesystem restriction.
|
||
- **Locality**: Local, Linux only.
|
||
- **Agent integration**: Generic dev sandbox; works with any agent that
|
||
runs inside the container.
|
||
- **Config**: Interactive CLI wizard — `define` (Dockerfile template),
|
||
`build` (prompts), `start` (launch).
|
||
- **Network policy**: "Limited isolation by default" — no strict
|
||
allowlist documented.
|
||
- **Notable**: Per-key SSH agent confirmation dialogs.
|
||
- **Maturity**: Early-stage, ~66 stars.
|
||
|
||
### agent-safehouse
|
||
- **Source**: https://agent-safehouse.dev/ ; https://github.com/eugene1g/agent-safehouse
|
||
- **HN launch**: [#47301085](https://news.ycombinator.com/item?id=47301085) (March 12 2026) — 823 points
|
||
- **License**: Apache 2.0 (~1,781 stars at launch)
|
||
- **Isolation**: macOS `sandbox-exec` (Seatbelt) profiles — kernel-level
|
||
syscall interception, no container.
|
||
- **Locality**: Local, macOS only.
|
||
- **Agent integration**: Explicit multi-agent wrapper — Claude Code,
|
||
OpenAI Codex, Gemini CLI, Cline, Aider. Usage:
|
||
`safehouse claude --dangerously-skip-permissions`.
|
||
- **Config**: Shell functions or custom `sandbox-exec` profile files;
|
||
LLM-assisted profile generation supported.
|
||
- **Network policy**: Not addressed.
|
||
- **Notable from HN thread**: Creator acknowledged the project is "just a
|
||
policy-generator for `sandbox-exec` — no dependencies, no daemons, no
|
||
subscription; I did put in many hours to identify the minimum required
|
||
permissions for agents to continue working." Simon Willison noted that
|
||
evaluating whether a sandboxing tool actually works as intended is hard.
|
||
Top community sentiment: *"I honestly think that sandboxing is currently
|
||
THE major challenge that needs to be solved for the tech to fully realise
|
||
its potential."* The macOS Docker gap (Docker for Mac runs inside a Linux
|
||
VM, so `sandbox-exec` is the only native primitive for bare-metal macOS
|
||
processes) was the stated motivation.
|
||
- **Maturity**: Active through March 2026.
|
||
|
||
### matchlock
|
||
- **Source**: https://github.com/jingkaihe/matchlock
|
||
- **License**: MIT (~574 stars, v0.2.10)
|
||
- **Isolation**: MicroVMs — Firecracker on Linux, Apple
|
||
Virtualization.framework on macOS. Transparent proxy via nftables DNAT
|
||
(Linux) or gVisor userspace TCP/IP (macOS).
|
||
- **Locality**: Local (Homebrew, .deb, .rpm).
|
||
- **Agent integration**: Agent-agnostic; SDK examples for Anthropic
|
||
Claude API and OpenAI. Go, Python, TypeScript SDKs.
|
||
- **Config**: CLI flags (`--allow-host`, `--secret`, `--no-network`) or
|
||
SDK builder pattern. No manifest file.
|
||
- **Network policy**: Default-deny + per-host allowlist.
|
||
- **Notable**: Secrets injected in-flight by the host proxy — they never
|
||
enter the VM.
|
||
- **Maturity**: Marked experimental.
|
||
|
||
### tilde.run
|
||
- **Source**: https://tilde.run/
|
||
- **License**: Proprietary, hosted SaaS.
|
||
- **Isolation**: Cloud-hosted containers; underlying mechanism not
|
||
publicly stated (unverified whether OCI containers or microVMs).
|
||
- **Locality**: Hosted only.
|
||
- **Agent integration**: Claude orchestration explicit; CLI
|
||
(`tilde exec`) and Python SDK; plain-English agent instructions.
|
||
- **Config**: DSL for RBAC policies (allow / deny / require human
|
||
approval per action, per repo, per agent).
|
||
- **Network policy**: Default-deny with per-request logging; cloud
|
||
metadata endpoints and private networks blocked.
|
||
- **Persistence**: All changes versioned and rollback-able via lakeFS;
|
||
atomic commits per run.
|
||
- **Maturity**: Private preview, © 2025, built by the lakeFS team.
|
||
|
||
### boxlite
|
||
- **Source**: https://boxlite.ai/ ; https://github.com/boxlite-ai/boxlite
|
||
- **License**: Apache 2.0 (~4,700 stars, YC-backed)
|
||
- **Isolation**: MicroVMs with dedicated Linux kernel per box — KVM on
|
||
Linux, Hypervisor.framework on macOS. Not containers/namespaces.
|
||
- **Locality**: Local, no daemon.
|
||
- **Agent integration**: Explicitly targets AI agents; MCP server
|
||
companion (boxlite-ai/boxlite-mcp). Pivoted from dev environments in
|
||
2025.
|
||
- **Config**: SDK only — Python, Node.js, Rust, C; Go pending. No
|
||
declarative manifest.
|
||
- **Network policy**: "Isolated Network per VM" — details not public
|
||
*(unverified)*.
|
||
- **Notable**: Sub-50ms boot, snapshot / fork / clone of VM state. Self
|
||
description: "the SQLite of sandboxing".
|
||
- **Maturity**: Active, YC.
|
||
|
||
### microsandbox
|
||
- **Source**: https://github.com/microsandbox/microsandbox (the
|
||
`superradcompany/microsandbox` URL redirects to the same project).
|
||
- **License**: Apache 2.0 (~6,000 stars, YC-backed)
|
||
- **Isolation**: MicroVMs via libkrun, OCI-compatible images.
|
||
Sub-100ms boot, rootless, no daemon, embeddable as a library.
|
||
- **Locality**: Local.
|
||
- **Agent integration**: Explicit Claude Code + Cursor targeting via
|
||
"Agent Skills" packages and an MCP server. Agents can create their own
|
||
sandboxes programmatically.
|
||
- **Config**: CLI (`msb`), SDKs (Rust, Python, TypeScript), MCP server.
|
||
- **Network policy**: Not detailed in public docs.
|
||
- **Maturity**: Beta, breaking changes expected; most-starred project in
|
||
this set.
|
||
|
||
### smolmachines
|
||
- **Source**: https://smolmachines.com/ ; https://github.com/smol-machines/smolvm
|
||
- **License**: Apache 2.0 (~3,100 stars)
|
||
- **Isolation**: MicroVMs via libkrun — Hypervisor.framework on macOS,
|
||
KVM on Linux. No shared kernel.
|
||
- **Locality**: Local, no daemon.
|
||
- **Agent integration**: Includes an `AGENTS.md`; designed with coding
|
||
agents in mind but no MCP/Skills turnkey integration.
|
||
- **Config**: TOML Smolfiles declaring image, networking, volumes, SSH
|
||
agent access, GPU acceleration. Portable `.smolmachine` files.
|
||
- **Network policy**: Off by default; per-host allowlist via
|
||
`--allow-host`.
|
||
- **Persistence**: Named machines persistent by default; ephemeral runs
|
||
also supported.
|
||
- **Maturity**: Active through April 2026.
|
||
|
||
### E2B *(added 2026-07-20)*
|
||
|
||
- **Source**: https://github.com/e2b-dev/e2b ; https://e2b.dev/docs
|
||
- **License**: Apache 2.0 (~12.4k stars); commercial hosted service with
|
||
self-hosting/BYOC support.
|
||
- **Isolation**: Firecracker microVM per sandbox.
|
||
- **Locality**: Cloud-hosted by default; self-hosting uses Terraform on AWS or
|
||
GCP (with other targets documented as works in progress).
|
||
- **Agent integration**: LLM-agnostic Python and JavaScript/TypeScript SDKs;
|
||
code-interpreter and desktop-sandbox products. Platform primitive rather
|
||
than a coding-agent wrapper.
|
||
- **Config**: Programmatic SDK/API plus templates. Network configuration
|
||
supports internet on/off, outbound allow/deny rules, and a custom egress
|
||
proxy.
|
||
- **Network policy**: Configurable per sandbox, but not documented as
|
||
default-deny and no built-in outbound-content DLP is documented.
|
||
- **Credentials**: Environment variables passed to the sandbox are explicitly
|
||
not private at the OS level. No built-in in-flight application-credential
|
||
injection is documented.
|
||
- **Persistence**: Full memory + filesystem pause/resume, snapshots, and
|
||
auto-resume. Continuous runtime is tier-limited, while paused sandboxes are
|
||
retained indefinitely.
|
||
- **Maturity**: Established hosted platform and the API compatibility target
|
||
used by CubeSandbox.
|
||
|
||
### Daytona *(added 2026-07-20)*
|
||
|
||
- **Source**: https://github.com/daytonaio/daytona ;
|
||
https://www.daytona.io/docs/
|
||
- **License**: Current production platform is proprietary. The former AGPL
|
||
repository remains public but is no longer maintained after Daytona moved
|
||
production development closed-source in June 2026.
|
||
- **Isolation**: Hosted container sandboxes by default, with separate Linux
|
||
and Windows VM sandbox classes for dedicated-OS workloads. Each sandbox has
|
||
its own filesystem and network stack; VM-only features include memory
|
||
pause/resume and forking.
|
||
- **Locality**: Hosted multi-tenant service, with dedicated/custom regions and
|
||
customer runners available.
|
||
- **Agent integration**: LLM/framework-agnostic SDKs (Python, TypeScript, Go,
|
||
Ruby, Java), API, and CLI; official agent-framework guides. Platform
|
||
primitive rather than a local coding-agent wrapper.
|
||
- **Config**: Programmatic per-sandbox image/snapshot, resources, lifecycle,
|
||
firewall, and secrets.
|
||
- **Network policy**: Per-sandbox IPv4/domain allowlists and block-all mode,
|
||
subordinate to organization/tier policy. Full internet access is the
|
||
default on higher tiers, so it is configurable rather than uniformly
|
||
default-deny.
|
||
- **Credentials**: First-class secret manager with the same phantom-token
|
||
pattern as bot-bottle: the sandbox environment gets an opaque placeholder,
|
||
an HTTPS proxy substitutes the real secret in headers only for allowed
|
||
hosts, and responses are scrubbed back to the placeholder.
|
||
- **Persistence**: Persistent filesystem for stopped container sandboxes;
|
||
memory + filesystem pause/resume for VM sandboxes; snapshots and configurable
|
||
auto-stop.
|
||
- **Maturity**: Production commercial platform. Notable April 2026 credential
|
||
exposure was patched; the June 2026 closed-source transition materially
|
||
changes its transparency/self-hosting posture.
|
||
|
||
### Other hosted runtimes carried forward from the earlier survey
|
||
|
||
- **Northflank Sandboxes** — hosted or customer-cloud, microVM-backed
|
||
containers with SDK-managed lifecycle, optional persistent volumes, and
|
||
sub-second claimed boot. This is a platform primitive for untrusted code and
|
||
agents, not a local agent wrapper or role-policy layer.
|
||
- **Cloudflare Sandbox SDK** — Workers/Durable Objects API over VM-isolated
|
||
Linux containers for command, file, process, and service execution. It is a
|
||
hosted TypeScript platform primitive; application authentication,
|
||
authorization, and credential-proxy patterns remain the integrator's job.
|
||
|
||
Both belong to the same “build your agent platform on this runtime” category as
|
||
E2B and Daytona. They were named but not analyzed in depth by the original
|
||
Claude-specific note, so they remain outside the main comparison table rather
|
||
than being presented with false precision.
|
||
|
||
### CubeSandbox *(added 2026-07-18)*
|
||
- **Source**: https://github.com/TencentCloud/CubeSandbox ;
|
||
HN launch https://news.ycombinator.com/item?id=47863430
|
||
- **License**: Apache 2.0 (~10.4k stars). By Tencent Cloud; described as
|
||
"battle-tested, production-ready" infra already running in Tencent
|
||
Cloud. Rust / Go / C.
|
||
- **Isolation**: MicroVMs via RustVMM + KVM — "each sandbox gets its own
|
||
Guest OS kernel, no Docker shared-kernel escapes." Hardware-level
|
||
isolation, dedicated kernel per instance.
|
||
- **Locality**: Self-hosted, but **server/cluster-oriented**, not a
|
||
single-user local CLI. Deploy guides target PVM cloud VMs, bare metal,
|
||
and dev. A single 96-vCPU host is claimed to run 2,000+ concurrent
|
||
sandboxes.
|
||
- **Agent integration**: **Drop-in E2B SDK replacement** (single env-var
|
||
change) — the headline compatibility claim. OpenClaw assistant
|
||
integration; general LLM-code execution. Aimed at platform builders,
|
||
not one developer's laptop.
|
||
- **Config**: Programmatic via the E2B-compatible SDK. No declarative
|
||
manifest.
|
||
- **Network policy**: This is the striking part — **domain allowlists,
|
||
instant block on unauthorized egress, full audit logs, per-sandbox
|
||
traffic tokens, policy-routing egress**, enforced by an eBPF-based
|
||
virtual switch giving kernel-level network isolation. Closest match yet
|
||
to bot-bottle's own default-deny + per-bottle allowlist egress model.
|
||
- **Credentials**: **Credential vault** — agents call external APIs / LLMs
|
||
while "keys never enter the sandbox, model context, or logs." Same
|
||
in-flight-injection idea as matchlock, but productized as a vault.
|
||
- **Performance**: <60ms cold start (claimed 2.5–50× faster than
|
||
alternatives), <5MB memory per instance; millisecond snapshot rollback
|
||
is upcoming.
|
||
- **Maturity**: Open-sourced July 2026 off production Tencent Cloud use;
|
||
most-starred project in this set (~10.4k).
|
||
|
||
### Cleanroom *(added 2026-07-18)*
|
||
- **Source**: https://github.com/buildkite/cleanroom
|
||
- **License**: Apache 2.0
|
||
- **Isolation**: MicroVM — Firecracker on Linux, Virtualization.framework
|
||
on macOS. Digest-pinned OCI images.
|
||
- **Locality**: Self-hosted server (CI-oriented).
|
||
- **Agent integration**: Generic process sandbox; CI-first, not a
|
||
Claude/agent wrapper.
|
||
- **Config**: `cleanroom.yaml` in the repo being sandboxed defines egress
|
||
rules, resources, and network policy. Cleanroom resolves this from the
|
||
commit being run.
|
||
- **Network policy**: Default-deny + per-repo hostname allowlist (resolved
|
||
from DNS answers + destination IP:port). Co-hosted services on the same
|
||
IP:port are not distinguished. OIDC-backed auth for remote servers.
|
||
- **Credentials**: Host-side only; not injected in-flight but not present
|
||
in the VM.
|
||
- **Notable**: Policy lives in the *repo being sandboxed*, not in an
|
||
agent-role definition — closer to per-repo scoping than per-role.
|
||
Supports Docker-inside-sandbox (`services.docker.required: true`), OIDC
|
||
authorization, suspend/resume lifecycle.
|
||
- **Maturity**: Active Buildkite product.
|
||
|
||
### container-use *(added 2026-07-18)*
|
||
- **Source**: https://github.com/dagger/container-use
|
||
- **License**: Apache 2.0
|
||
- **Isolation**: Docker container per agent + git worktree per agent.
|
||
Containers share the host kernel; stronger than bare host but weaker
|
||
than microVM.
|
||
- **Locality**: Local.
|
||
- **Agent integration**: MCP stdio server — Claude Code, Cursor, Windsurf.
|
||
`claude mcp add container-use -- container-use stdio`.
|
||
- **Config**: None for security policy. Environments are provisioned on
|
||
demand; no allowlist or credential config.
|
||
- **Network policy**: Not addressed.
|
||
- **Notable**: Per-agent git branches (`container-use/<env_name>`);
|
||
parallel agents without filesystem conflict; real-time log visibility
|
||
and terminal attach for intervention; git-based review workflow.
|
||
Oriented toward parallel development safety, not security containment.
|
||
- **Maturity**: Early development, active.
|
||
|
||
### Docker sbx *(added 2026-07-18)*
|
||
- **Source**: Docker proprietary (`sbx` CLI, separate from `docker`).
|
||
- **License**: Proprietary.
|
||
- **Isolation**: MicroVM (Docker's own implementation) — each session gets
|
||
its own kernel, Docker daemon inside the VM, and filesystem.
|
||
- **Locality**: Local (macOS and Windows; does not require Docker Desktop).
|
||
- **Agent integration**: Explicit wrapper — Claude Code, Codex, Gemini
|
||
CLI, Copilot CLI, Kiro. Launches agent inside the VM with
|
||
`--dangerously-skip-permissions` by default.
|
||
- **Config**: Open / Balanced / Locked Down network presets at launch. No
|
||
per-role manifest.
|
||
- **Network policy**: Default-deny; preset levels control strictness. TUI
|
||
dashboard shows a live log of every outbound connection (allowed and
|
||
blocked) with point-and-click allow/block for hosts.
|
||
- **Credentials**: OS keychain + host-side proxy injection — API keys
|
||
never enter the VM.
|
||
- **Notable**: Best DX among microVM tools (one command, works like native
|
||
yolo Claude but inside a VM); branch mode creates a git worktree in
|
||
`.sbx/`. Network policy is preset-based, not role-declarative.
|
||
- **Maturity**: GA 2026.
|
||
|
||
### Anthropic srt *(added 2026-07-18)*
|
||
- **Source**: https://github.com/anthropic-experimental/sandbox-runtime
|
||
(`@anthropic-ai/sandbox-runtime` on npm, `sandbox-runtime` on PyPI)
|
||
- **License**: Apache 2.0 (experimental).
|
||
- **Isolation**: OS-level only — Seatbelt (`sandbox-exec`) on macOS,
|
||
bubblewrap on Linux, WFP (Windows Filtering Platform) account-fenced on
|
||
Windows. **No container or VM.** Lowest overhead in the set.
|
||
- **Locality**: Local.
|
||
- **Agent integration**: Claude Code's sandboxed bash tool uses this
|
||
internally. Can wrap any arbitrary process (`srt <command>`). Cloud
|
||
Claude Code sessions use full microVMs instead.
|
||
- **Config**: Programmatic per-invocation — allow/deny path lists for
|
||
filesystem; allow/denylist for network (HTTP proxy + SOCKS5).
|
||
- **Network policy**: Proxy-based filtering (HTTP + SOCKS5); domain
|
||
allowlist/denylist enforced at proxy layer. Custom proxy supported
|
||
(e.g. mitmproxy for inspection + audit). Processes that ignore proxy
|
||
env vars may bypass filtering on some platforms.
|
||
- **Notable**: Cross-platform (macOS/Linux/Windows); wraps any process,
|
||
not just agents; no role/manifest concept. Annotated as a research
|
||
preview — APIs may change.
|
||
- **Maturity**: Early research preview.
|
||
|
||
## Claude-specific wrappers and developer environments
|
||
|
||
These projects were the focus of the original containerized-Claude survey.
|
||
They remain useful comparisons for local developer experience, but most are
|
||
templates or wrappers rather than policy-bearing sandbox platforms, so they
|
||
are grouped here instead of widening the main table further.
|
||
|
||
### claudebox
|
||
|
||
- **Source**: https://github.com/RchGrav/claudebox
|
||
- **Isolation**: Docker, with per-project images, authentication state, and
|
||
configuration.
|
||
- **Agent integration**: Claude Code wrapper with 15+ preconfigured language
|
||
and task profiles.
|
||
- **Network policy**: Per-project firewall allowlists.
|
||
- **Closest overlap**: local one-command developer workflow and project-scoped
|
||
network policy.
|
||
- **Difference**: profiles describe development toolchains, not named agent
|
||
roles. There is no bottle/agent split, composable role manifest, provider
|
||
plugin layer, or outbound-content DLP.
|
||
|
||
### Spritz / claude-code-sandbox
|
||
|
||
- **Source**: https://github.com/textcortex/claude-code-sandbox (archived;
|
||
points to its successor, Spritz).
|
||
- **Isolation**: The original project ran Claude Code in local Docker with
|
||
bypass permissions; Spritz moved toward Kubernetes-native multi-agent
|
||
infrastructure.
|
||
- **Difference**: the successor targets cluster orchestration rather than a
|
||
low-dependency local launcher. It is architecturally closer to hosted or
|
||
Kubernetes platform runtimes than to bot-bottle's single-operator CLI.
|
||
|
||
### Trail of Bits claude-code-devcontainer
|
||
|
||
- **Source**: https://github.com/trailofbits/claude-code-devcontainer
|
||
- **Isolation**: A Docker devcontainer that exposes only project files and is
|
||
designed to run Claude Code with `bypassPermissions` for security audits and
|
||
untrusted-code review.
|
||
- **Difference**: a hardened, reusable environment definition rather than an
|
||
agent launcher or fleet. It has no named-role manifest, per-role credential
|
||
custody, supervision plane, or multi-backend abstraction.
|
||
|
||
### Smaller wrappers and official templates
|
||
|
||
Projects such as `arezi/claude-sandbox`, `nkrefman/claude-sandbox`, and
|
||
`VishalJ99/claude-docker`, plus Docker/Anthropic devcontainer templates, prove
|
||
there is steady demand for “Claude in a container.” They are deliberately
|
||
small launch/build configurations. They compete on setup simplicity, not on
|
||
role-aware policy, credential custody, persistent supervision, or a fleet
|
||
model, and are better treated as a product category than as individual rows.
|
||
|
||
### SuperHQ
|
||
|
||
- **Source**: https://superhq.ai/
|
||
- **Isolation**: Apple-Silicon desktop application using local microVMs via
|
||
Virtualization.framework/libkrun-era components.
|
||
- **Agent integration**: Claude Code, Codex, and Pi in a GUI, with mobile
|
||
remote access.
|
||
- **Credentials and review**: host-side auth gateway injects credentials on
|
||
the wire; a temporary overlay stages writes for diff-and-accept review.
|
||
- **Closest overlap**: local microVM isolation, multi-provider launching, and
|
||
credential custody for security-minded individual developers.
|
||
- **Difference**: GUI desktop product on Apple Silicon rather than a
|
||
cross-platform declarative CLI/fleet layer. The July 2026 snapshot in the
|
||
original survey recorded a user request for per-run tool-call and network
|
||
audit logging; treat that as point-in-time rather than a permanent gap.
|
||
|
||
## Credential gateway without isolation
|
||
|
||
### OneCLI
|
||
|
||
[OneCLI](https://onecli.sh/) is a framework-agnostic identity gateway rather
|
||
than a sandbox. Its phantom-token design gives the agent a placeholder and
|
||
substitutes the encrypted real credential at the network layer. It therefore
|
||
matches bot-bottle closely on secret custody, and is more portable because it
|
||
can sit in front of agents launched by anything, but it supplies no container
|
||
or VM boundary, filesystem isolation, role manifest, or egress-content DLP.
|
||
|
||
The positioning consequence from the earlier survey still holds: secret
|
||
custody alone is not unique. bot-bottle's relevant combination is local
|
||
isolation + default-deny egress + payload DLP + declarative roles + credential
|
||
custody. OneCLI's managed tier also places custody with a third party, whereas
|
||
bot-bottle keeps it within operator-controlled infrastructure. See
|
||
[`agent-credential-proxy-landscape.md`](agent-credential-proxy-landscape.md)
|
||
for the detailed build-versus-adopt analysis.
|
||
|
||
## Governance / pre-action authorization layers
|
||
|
||
These two tools don't provide VM or filesystem isolation; they intercept
|
||
tool calls before execution and evaluate them against a per-agent
|
||
declarative policy. They are the closest competitors on **agent-tailored
|
||
policy** and complement isolation sandboxes rather than substituting for
|
||
them.
|
||
|
||
### Microsoft Agent Governance Toolkit (AGT) *(added 2026-07-18)*
|
||
- **Source**: https://github.com/microsoft/agent-governance-toolkit
|
||
- **License**: MIT (~3.3k stars, open-sourced April 2, 2026).
|
||
- **Isolation**: None (OS/VM). Execution rings (0–3, inspired by CPU
|
||
privilege levels) control what an agent can do at the framework layer.
|
||
MCP security gateway treats MCP traffic as an untrusted boundary.
|
||
- **Locality**: Embedded in the agent framework (Python, TypeScript, .NET,
|
||
Rust, Go; 20+ framework adapters).
|
||
- **Agent integration**: Framework-agnostic. Plugs into Semantic Kernel,
|
||
AutoGen, and others as a middleware layer.
|
||
- **Config**: YAML policy per agent — tools can be `allowed`, `denied`,
|
||
`sandboxed`, or routed through an `approval` step. Every action passes
|
||
through a governance gate checking: agent DID, trust score, risk tier,
|
||
requested tool, action type, and policy rules.
|
||
- **Network policy**: Not directly — operates at tool-call level.
|
||
- **Credentials**: Per-agent DID (Ed25519 decentralized identifier); agent
|
||
does not borrow a human's credentials.
|
||
- **Notable**: Dynamic trust score (0–1,000, behavioral decay) —
|
||
privilege follows observed behaviour, not just provisioning. Covers all
|
||
10 OWASP Agentic Top 10 risks. Kill switch + SLO monitoring. Sub-ms
|
||
policy enforcement.
|
||
- **Maturity**: MIT, ~3.3k ⭐, v3.7.0 May 2026.
|
||
|
||
### Open Agent Passport (OAP) *(added 2026-07-18)*
|
||
- **Source**: https://github.com/aporthq/aport-spec ; spec at
|
||
https://api.aport.io/spec/spec/oap/oap-spec.md/ ; arXiv 2603.20953
|
||
- **License**: Open specification.
|
||
- **Isolation**: None. Pre-action hook only — intercepts tool calls
|
||
synchronously before execution, evaluates against a cloud-registry
|
||
declarative policy, fails closed.
|
||
- **Locality**: Local hook + cloud policy registry.
|
||
- **Agent integration**: Framework-agnostic; hook pattern.
|
||
- **Config**: Declarative policy rules in a cloud registry (evaluated in
|
||
order; first failing rule denies). Ed25519-signed, hash-chained audit
|
||
records per decision.
|
||
- **Network policy**: Not directly.
|
||
- **Notable**: 53ms median authorization decision (N=1,000). In an
|
||
adversarial testbed ($5,000 bounty, 1,151 sessions), social engineering
|
||
succeeded 74.6% of the time under a permissive policy; under a
|
||
restrictive OAP policy, 0% success across 879 attempts. Assumes
|
||
framework runtime is not compromised.
|
||
- **Maturity**: Specification + reference implementation, 2026.
|
||
|
||
## Comparison table
|
||
|
||
*Isolation/sandbox tools only. AGT and OAP are governance layers — see their per-project notes above.*
|
||
|
||
| Axis | bot-bottle | endo-familiar | litterbox | agent-safehouse | matchlock | tilde.run | boxlite | microsandbox | smolmachines | E2B | Daytona | CubeSandbox | Cleanroom | container-use | Docker sbx | Anthropic srt |
|
||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
||
| Isolation | MicroVM per bottle default (Firecracker/KVM on Linux, Apple Container on macOS) + own egress DLP scanner; Docker legacy fallback, gVisor there if present | Object-capability (no OS isolation) | Podman + opt. Landlock | macOS `sandbox-exec` | MicroVM (Firecracker / Virt.fw) | Hosted container (unverified) | MicroVM (KVM / Hypervisor.fw) | MicroVM (libkrun) | MicroVM (libkrun / KVM) | Firecracker microVM | Container or Linux/Windows VM class | MicroVM (RustVMM / KVM) | MicroVM (Firecracker / Virt.fw) | Docker container + git worktree | MicroVM (proprietary) | OS-level (Seatbelt / bubblewrap / WFP) — no container |
|
||
| Local vs hosted | Local | Local | Local (Linux) | Local (macOS) | Local | Hosted SaaS | Local | Local | Local | Hosted; self-host/BYOC available | Hosted; dedicated/custom regions | Self-hosted (server/cluster) | Self-hosted server | Local | Local | Local |
|
||
| Open source | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | MIT | No | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 | Production closed-source; legacy AGPL repo unmaintained | Apache 2.0 | Apache 2.0 | Apache 2.0 | Proprietary | Apache 2.0 (experimental) |
|
||
| Agent target | Claude Code, Codex, Pi, and provider plugins | Generic (demo) | Generic | Multi-agent wrapper | Generic (+ Claude/OpenAI SDKs) | Claude focus | Generic | Claude + Cursor (MCP/Skills) | Generic (AGENTS.md) | LLM-agnostic platform builders | LLM-agnostic platform builders | E2B-compatible (platform builders) | CI / generic process | Claude Code, Cursor, Windsurf (MCP) | Claude Code, Codex, Gemini CLI, Copilot, Kiro | Claude Code (and any process) |
|
||
| Network policy | Default-deny via own egress scanner + per-bottle allowlist + content DLP + gitleaks on git push | Capability model only | Limited | Not addressed | Default-deny + allowlist + secret-injecting proxy | Default-deny + logging | Per-VM net (unverified) | Not documented | Off by default + allowlist | Per-sandbox allow/deny rules and custom egress proxy; internet configurable | Per-sandbox CIDR/domain allowlist or block-all; tier policy; secret-injecting proxy | Default-deny allowlist + instant egress block + audit logs + per-sandbox tokens (eBPF) + credential vault | Default-deny + per-repo host allowlist (cleanroom.yaml) | Not addressed | Default-deny; Open / Balanced / Locked Down presets; live TUI network panel | Proxy-based allowlist/denylist (HTTP + SOCKS5); custom proxy supported |
|
||
| Parallel agents | Yes (one bottle per agent) | n/a | Not addressed | One at a time | Multiple VMs | Yes (dashboard) | SDK-level | SDK-level | Architectural | Yes (platform service) | Yes (platform service) | Yes (2,000+/host claimed) | Yes (server model) | Yes (per-agent containers + worktrees) | Yes | Yes |
|
||
| Long-running posture | Persistent by default (named, supervised) | n/a (demo) | Session (up while in use) | Per-invocation | Ephemeral VM per run | Per-run (versioned) | Ephemeral + snapshot/fork | Ephemeral / on-demand | Named persistent by default | Runtime tier limits + indefinite pause/resume | Persistent filesystem; VM pause/resume; configurable auto-stop | Ephemeral + auto pause/resume | Per-run + suspend/resume | Per-agent container (ephemeral) | Per-session; branch mode creates git worktree in .sbx/ | Per-invocation |
|
||
| DX: run Claude yolo-style | One command → interactive yolo Claude (`start <agent>`, `--dangerously-skip-permissions` default) | n/a (lib demo) | Wizard + build, then run claude inside (Linux only) | One-command wrapper (`safehouse claude --dangerously-skip-permissions`) | CLI: run a cmd in a VM (not a Claude wrapper) | Hosted (`tilde exec`), not local-native | SDK code required (build the run yourself) | CLI/MCP: sandbox-as-a-tool for the agent, not a wrapper around it | SSH into a named machine, run claude there | SDK/CLI sandbox; wire the agent yourself | SDK/CLI sandbox; wire the agent yourself | Stand up a cluster + drive via E2B SDK | CI-oriented, not a Claude wrapper | MCP server: `claude mcp add container-use -- container-use stdio` | One command: `sbx` wraps claude with `--dangerously-skip-permissions` default | Library/wrapper, not a standalone CLI |
|
||
| Config | YAML-in-Markdown manifests (bottles + agents) | Programmatic refs | CLI wizard | Profile files / shell fns | CLI / SDK | DSL + CLI + SDK | SDK | CLI / SDK / MCP | TOML Smolfile | SDK/API + templates | SDK/API/CLI + images/snapshots | E2B-compatible SDK | cleanroom.yaml in repo | None (no policy config) | Preset levels at launch | Programmatic per-invocation (allow/deny lists) |
|
||
| Agent-tailored policy | Yes — bottle/agent split; declarative per-role egress + credentials; composable via `extends:` | Partial — capability model scopes per-agent, but no declarative role manifest | No | Partial — per-agent profile files (Seatbelt); no egress | No | Yes — per-agent DSL RBAC (allow/deny/approve per action/repo/agent) | No | No | No | No — per-sandbox SDK config | No — per-sandbox SDK config | No — per-sandbox SDK config, not role-scoped | Partial — per-repo cleanroom.yaml, not per-role | No | No — network presets only | No |
|
||
| Maturity | Active July 2026 | Research (2022+) | Early (~66 ⭐) | Active (~1.8k ⭐) | Experimental (~574 ⭐) | Private preview | YC, ~4.7k ⭐ | YC, ~6k ⭐, beta | ~3.1k ⭐ | Established hosted platform, ~12.4k ⭐ | Production commercial; closed-source since June 2026 | Tencent, prod, ~10.4k ⭐ | Active (Buildkite product) | Early development | GA 2026 | Early research preview |
|
||
|
||
## What's closest, what's different
|
||
|
||
**Closest in design and scope.** agent-safehouse and litterbox sit
|
||
nearest bot-bottle: local, single-user, thin wrappers over an
|
||
existing OS primitive, low-dep. The split is the isolation primitive —
|
||
bot-bottle now defaults to a VM per bottle (Firecracker microVM on KVM
|
||
Linux, Apple Container on macOS) with its own DLP-scanning egress proxy,
|
||
keeping Docker only as a legacy fallback; agent-safehouse uses
|
||
`sandbox-exec`; litterbox uses Podman + Landlock. matchlock and
|
||
smolmachines are close on *both* the policy side (default-deny net,
|
||
per-host allowlist) and — now that bot-bottle has moved off
|
||
containers-by-default — the microVM isolation primitive. Note: Apple
|
||
Container 1.0 stable shipped June 9 2026 (frozen CLI and APIs), which
|
||
makes the macOS backend stable surface area rather than a moving target.
|
||
|
||
**New closest on agent-tailored policy.** Two governance tools are the
|
||
direct competitors on the "coarse-grained sandbox" axis. **tilde.run**
|
||
has had per-agent DSL RBAC since its launch (though it's hosted SaaS).
|
||
**Microsoft AGT** is the most serious new entrant: per-agent DID
|
||
identity, YAML policy that can allow/deny/sandbox/approve individual tool
|
||
calls per agent, and a dynamic behavioural trust score. It operates at
|
||
the framework tool-call layer, not the network layer — so it's
|
||
complementary to bot-bottle's network/filesystem isolation rather than a
|
||
direct substitute, but on the "does this sandbox know what this agent is
|
||
for?" question it is the most complete answer in the field. OAP's
|
||
pre-action hook pattern achieves similar goals with cryptographic audit
|
||
and a 0% adversarial-attack success rate under a restrictive policy.
|
||
|
||
**New closest on DX.** **Docker sbx** is the first tool in this set that
|
||
matches bot-bottle on the "one command, dangerously-skip-permissions safe
|
||
by default" DX bar, at microVM isolation strength, with host-side
|
||
credential injection. It is proprietary, preset-based (not role-
|
||
declarative), and cloud-agent-specific, but it directly competes on the
|
||
UX proposition. agent-safehouse was the previous DX peer; Docker sbx
|
||
materially raises the bar.
|
||
|
||
**New closest on repo-scoped policy.** **Cleanroom** (Buildkite) is the
|
||
first tool to combine microVM isolation with a declarative egress policy
|
||
file — though the policy lives in the repo being sandboxed
|
||
(`cleanroom.yaml`), not in an agent-role manifest. That makes it per-
|
||
repo rather than per-role: the same Cleanroom config applies to any
|
||
agent running in that repo. The distinction matters for bot-bottle's
|
||
use case (one developer running multiple agent *roles* with different
|
||
egress footprints), but for CI/CD use cases Cleanroom is a direct
|
||
alternative.
|
||
|
||
**Solving a different problem.** tilde.run is hosted SaaS for team /
|
||
production agent pipelines with data-versioned rollback — explicitly
|
||
opposite to bot-bottle's "infrastructure I control" goal. E2B and Daytona
|
||
are hosted sandbox platforms, while boxlite, microsandbox, and CubeSandbox
|
||
are infrastructure libraries/services aimed at platform builders embedding
|
||
sandboxes into agent frameworks; they
|
||
would be a *backend* bot-bottle could call, not a competitor to its
|
||
manifest layer. endo-familiar is in a different paradigm entirely:
|
||
capability passing rather than kernel boundaries.
|
||
|
||
## Borrowable ideas
|
||
|
||
### Already shipped or otherwise addressed
|
||
|
||
- Default-deny egress with a per-agent allowlist (own egress scanner).
|
||
- DLP scanning of outbound traffic.
|
||
- Bottle / agent split (manifest layer above the isolation primitive).
|
||
- gVisor auto-detection on Linux.
|
||
- **In-flight secret injection** (suggested by matchlock) — **shipped.**
|
||
Real provider and git-host tokens are held outside the agent and injected
|
||
by the egress gateway on matching routes. The agent receives only proxy
|
||
URLs and, where a client requires a credential-shaped value, a placeholder;
|
||
`GITEA_TOKEN` and equivalent real tokens do not appear in the agent's
|
||
environment.
|
||
- **MicroVM backend** — **shipped.** MicroVMs are now the default:
|
||
Firecracker on KVM Linux and Apple Container on macOS. Docker is the legacy
|
||
fallback.
|
||
- **Per-use SSH key confirmation** (suggested by litterbox) — **addressed by
|
||
stronger credential custody instead.** The agent does not hold the upstream
|
||
git SSH key or an SSH-agent socket: git-gate holds the credential and gates
|
||
git operations. A confirmation wrapper inside the agent would therefore
|
||
protect a credential that is no longer there. Operator approval at the gate
|
||
remains the appropriate control point for any future per-use confirmation.
|
||
|
||
### Still worth considering
|
||
|
||
- **Live network activity in the supervisor TUI** (from Docker sbx): show
|
||
allowed and blocked connections and let the operator propose policy changes
|
||
from the existing supervision surface.
|
||
- **Tamper-evident audit records** (from OAP): sign and hash-chain egress and
|
||
supervision decisions for compliance-sensitive deployments.
|
||
- **Behaviour-informed policy downgrade** (from Microsoft AGT): use repeated
|
||
DLP alerts or supervision holds as a signal to narrow policy or request
|
||
closer review. This needs a carefully specified trust model before it can be
|
||
more than a heuristic.
|
||
|
||
Not worth borrowing: the SDK-first programmatic API style of boxlite /
|
||
microsandbox (cuts against the declarative-manifest stance), and the
|
||
hosted-SaaS dashboard model of tilde.run (cuts against the
|
||
"infrastructure I control" goal).
|
||
|
||
## Publishing and positioning verdict
|
||
|
||
Publishing remains worthwhile, but the defensible claim is the combination,
|
||
not any single primitive. Credential custody is matched by OneCLI, matchlock,
|
||
Daytona, Docker sbx, and CubeSandbox; local one-command isolation is matched by
|
||
agent-safehouse and Docker sbx; hosted microVM execution is a crowded platform
|
||
category.
|
||
|
||
bot-bottle remains unusual in combining:
|
||
|
||
- local, operator-controlled execution with persistent named bottles;
|
||
- one declarative role layer across Claude Code, Codex, Pi, and provider
|
||
plugins;
|
||
- composable agent/bottle manifests, skills, and system prompts;
|
||
- Firecracker/Apple Container isolation with a Docker fallback;
|
||
- default-deny per-role egress, payload DLP, and git-push secret scanning;
|
||
- credentials injected outside the agent process; and
|
||
- supervision and audit state suited to long-running parallel agents.
|
||
|
||
The practical wedge is “as easy as native yolo, with declarative role policy
|
||
and self-hosted custody,” including scoped access to private LAN/Tailnet
|
||
services that cloud-first runtimes cannot provide without additional network
|
||
plumbing. The main competitive risks are a local wrapper such as claudebox or
|
||
Docker sbx growing a role-manifest layer, and GUI products such as SuperHQ
|
||
adding equivalent policy and audit depth.
|
||
|
||
## Caveats
|
||
|
||
- Star counts and last-commit dates are point-in-time snapshots.
|
||
- Several projects' network and persistence behaviour is not
|
||
documented publicly; items so derived are marked *(unverified)*.
|
||
- The `superradcompany/microsandbox` URL in the original prompt
|
||
redirects to `microsandbox/microsandbox`; the surveyed project is the
|
||
same.
|
||
- CubeSandbox performance/scale numbers (<60ms cold start, <5MB/instance,
|
||
2,000+ sandboxes per 96-vCPU host) are the project's own launch claims,
|
||
not independently verified here.
|
||
|
||
## Addendum 2026-07-18 — CubeSandbox and the positioning read
|
||
|
||
CubeSandbox (Tencent Cloud, Apache 2.0, ~10.4k stars, HN launch
|
||
[#47863430](https://news.ycombinator.com/item?id=47863430)) is the first
|
||
open-source, self-hostable project in this survey to combine, in one stack,
|
||
the main primitives
|
||
bot-bottle treated as its differentiator:
|
||
|
||
- **Egress custody (connection level)** — default-deny domain allowlist
|
||
(L7 domain/SNI filtering), instant block on unauthorized egress,
|
||
per-sandbox traffic tokens, full audit logs of destinations (eBPF
|
||
virtual switch, "CubeVS"). This matches bot-bottle's egress scanner at
|
||
the *connection level*, productized — see the one thing it does **not**
|
||
match, below.
|
||
- **Credential custody** — a vault where keys "never enter the sandbox,
|
||
model context, or logs." This is the in-flight-injection idea from
|
||
matchlock, but as a first-class feature, and it's exactly the
|
||
cross-vendor "egress audit + custody" wedge the monetization
|
||
positioning treats as the one defensible moat.
|
||
- **Isolation on par with bot-bottle's current default** — a dedicated
|
||
guest kernel per sandbox (RustVMM/KVM). bot-bottle now defaults to the
|
||
same class of boundary (Firecracker microVM / Apple Container), so this
|
||
is parity, not an edge; CubeSandbox's remaining edge is running that
|
||
per-kernel isolation multi-tenant at scale on one host.
|
||
|
||
The one axis CubeSandbox does **not** cover — and where bot-bottle stays
|
||
distinctive:
|
||
|
||
- **Content DLP on *authorized* channels.** CubeSandbox's egress control
|
||
is connection-level: it decides *whether* a destination is allowed and
|
||
logs it, and its vault keeps *injected* credentials out of the sandbox
|
||
entirely. Neither inspects the *payload* of traffic to an allowed
|
||
destination. So an agent that exfiltrates over a permitted channel —
|
||
pasting a repo's contents, an agent-derived secret, or PHI into an
|
||
allowed API/domain — is not caught by CubeSandbox. bot-bottle's own
|
||
egress DLP scanner does scan that: response + websocket content against
|
||
the resolved per-flow config, with per-bottle token redaction (see
|
||
recent egress commits). The vault
|
||
approach is arguably *stronger* for the specific case of pre-known
|
||
injected credentials (they can't leak if they were never present), but
|
||
it is not a substitute for content inspection of everything else.
|
||
|
||
**Long-running posture — a sharper axis than raw isolation.** E2B and
|
||
CubeSandbox are *ephemeral-per-task* by design; a long-running agent is an
|
||
architected pattern on top, not the default. E2B: 5-minute default
|
||
timeout, continuous runtime tier-capped (~1h Hobby / ~24h Pro), duration
|
||
achieved via **pause/resume** (preserves filesystem + memory + processes;
|
||
reconnect by sandbox ID via `Sandbox.connect()`; resume resets the timeout
|
||
to 5 min; auto-pause via `on_timeout: "pause"`). CubeSandbox mirrors this
|
||
(E2B drop-in) with first-class auto pause/resume and hundred-ms
|
||
checkpoint/fork — and, self-hosted, sets its own timeout policy with no
|
||
vendor tier caps. bot-bottle inverts the model: a bottle is **persistent,
|
||
named, and supervised by default** — long-running *is* the default, not a
|
||
session-management loop over pause/resume. smolmachines is the other
|
||
persistent-by-default project in this set. For anyone building agents that
|
||
run for hours/days, this posture difference matters more than the
|
||
isolation primitive.
|
||
|
||
**DX — the "run Claude yolo-style" bar.** The reason `claude
|
||
--dangerously-skip-permissions` is so widely used is DX: it's one command
|
||
and the agent just goes. The bottle thesis is to make a *sandboxed* run
|
||
that easy — `start <agent>` builds the image on first run and drops you
|
||
into an interactive Claude session that already has
|
||
`--dangerously-skip-permissions` on by default
|
||
(`contrib/claude/agent_provider.py`), with the sandbox as the guardrail
|
||
instead of per-action prompts. On this axis the field splits cleanly:
|
||
- **Wrappers around the agent** (as-easy-as-native): bot-bottle and
|
||
**agent-safehouse** (`safehouse claude --dangerously-skip-permissions`).
|
||
These *are* the run-Claude experience. agent-safehouse is the real DX
|
||
peer — but it's macOS-only Seatbelt, single-run, and doesn't address
|
||
network egress; bot-bottle adds VM-grade isolation, egress DLP, and
|
||
persistent/parallel bottles across macOS + Linux.
|
||
- **Libraries / services** (you build the run yourself): boxlite,
|
||
microsandbox, CubeSandbox, E2B, Daytona. These hand you an SDK or a cluster and
|
||
expect you to wire the agent in — powerful for platform builders,
|
||
heavyweight for "just run Claude on my laptop." microsandbox's MCP/Skills
|
||
angle is *sandbox-as-a-tool the agent calls*, which is the inverse of
|
||
wrapping the agent.
|
||
- **In between:** litterbox (wizard + build, Linux only), smolmachines
|
||
(SSH into a named machine), matchlock (run a command in a VM).
|
||
|
||
So DX is a genuine bot-bottle differentiator. agent-safehouse matches the
|
||
one-command wrapper with weaker isolation and no egress story; Docker sbx now
|
||
matches it at microVM strength but remains proprietary and preset-based. "As
|
||
easy as native yolo, with declarative role policy" is the narrower defensible
|
||
one-liner.
|
||
|
||
Why it still doesn't collide head-on:
|
||
|
||
1. **Shape.** CubeSandbox is a *multi-tenant service for platform
|
||
builders* (drop-in E2B replacement, SDK-driven, 2,000 sandboxes on a
|
||
box). bot-bottle is a *single-operator, declarative-manifest tool for
|
||
the infrastructure I run*. Different buyer, different ergonomics — no
|
||
declarative role manifest, no bottle/agent split, no "one command on my
|
||
laptop."
|
||
2. **Backend, not competitor.** Like boxlite/microsandbox, CubeSandbox is
|
||
something bot-bottle could sit *on top of* — a `"runtime": "microvm"`
|
||
or `"runtime": "cubesandbox"` backend under the manifest layer — while
|
||
keeping the manifest, the bottle/agent split, and the local,
|
||
single-operator default.
|
||
|
||
Why it matters anyway:
|
||
|
||
- The "nobody else bundles connection-level egress allowlist + audit +
|
||
in-flight credential custody" line is **no longer true for the
|
||
primitive** — CubeSandbox ships the open-source/self-hosted combination,
|
||
and Daytona ships a proprietary firewall + credential-substitution variant.
|
||
But **content DLP on authorized channels is still not matched** (see
|
||
above), and neither is the *layer above* the primitive (declarative
|
||
manifest, cross-vendor orchestration, operator UX, the
|
||
phone-control/dashboard north star). Those two — outbound-payload DLP
|
||
and the orchestration layer — are where the defensible ground now sits;
|
||
the connection-level allowlist + vault mechanism, on its own, is no
|
||
longer differentiating. Revisit the monetization open/paid line with
|
||
that in mind.
|
||
- Worth a closer look at **how** CubeSandbox does credential injection
|
||
and per-sandbox egress tokens (eBPF virtual switch vs. bot-bottle's
|
||
mitmproxy egress proxy) when hardening bot-bottle's now-shipped
|
||
credential-custody implementation.
|
||
|
||
## Addendum 2026-07-18 (second pass) — agent-tailored policy landscape
|
||
|
||
The second-pass question was: how novel is bot-bottle's per-agent,
|
||
role-tailored sandbox relative to the expanded field?
|
||
|
||
**The short answer:** on the isolation + network + role-tailoring
|
||
combination, bot-bottle remains the only tool in this set. On
|
||
role-tailored *policy at the tool-call level*, Microsoft AGT and OAP are
|
||
the most complete answers, but they don't provide isolation; they
|
||
complement rather than substitute.
|
||
|
||
**The competitive picture by axis:**
|
||
|
||
- *Agent-tailored egress (declarative, per-role)* — bot-bottle and
|
||
tilde.run. Cleanroom is per-repo, not per-role. Everyone else is
|
||
per-session or not addressed.
|
||
- *Agent-tailored tool-call policy (declarative, per-agent identity)* —
|
||
Microsoft AGT (YAML policy + DID identity + trust score), OAP
|
||
(declarative policy rules + cryptographic audit). Neither provides
|
||
network/filesystem isolation.
|
||
- *Composable policy (role overlays)* — bot-bottle (`extends:`). No
|
||
other tool surveyed supports composable role-policy inheritance.
|
||
- *Isolation + DX (one-command safe yolo)* — bot-bottle and Docker sbx.
|
||
Docker sbx is proprietary, preset-based, and cloud-agent-specific;
|
||
it's the first DX-class competitor at microVM isolation strength.
|
||
|
||
**What the HN "coarse-grained" complaint maps to:** The complaint is
|
||
that a VM isolates the filesystem but doesn't know if the agent
|
||
*should* be sending an email. bot-bottle's bottle/agent split is a
|
||
structural answer to this: the bottle manifest declares exactly what
|
||
the role can reach, and the sandbox enforces it at the network layer.
|
||
Microsoft AGT is the most complete answer at the semantic/tool-call
|
||
layer. The gap both leave open is *intent classification* — knowing
|
||
whether a permitted action is consistent with the agent's actual task.
|
||
See `hn-agent-safety-discourse-july-2026.md` for the blast-radius
|
||
analysis.
|
||
|
||
**Open ideas from new tools (also summarized above):**
|
||
|
||
- **Microsoft AGT's trust-score decay** — privilege that reflects
|
||
observed behaviour rather than static provisioning. Applied to
|
||
bot-bottle: a bottle that has triggered DLP alerts or supervise holds
|
||
could auto-downgrade its network preset, or flag the session for
|
||
closer review. Fits the existing supervise-server architecture.
|
||
- **Docker sbx's live network TUI** — real-time per-session view of
|
||
allowed and blocked outbound connections with point-and-click
|
||
allow/block. `cli.py supervise` is the right surface; adding a
|
||
live-connections panel would directly address the "I can't see what
|
||
the agent is doing" gap without any backend changes.
|
||
- **OAP's cryptographic audit chain** — Ed25519-signed, hash-chained
|
||
audit records. Currently bot-bottle logs egress decisions but doesn't
|
||
chain them. A tamper-evident audit record per session would be useful
|
||
for the compliance use case the CubeSandbox positioning targets.
|