feat(macos): run containers inside a bottle via guest-local podman
tracker-policy-pr / check-pr (pull_request) Successful in 12s
lint / lint (push) Successful in 51s
test / integration-docker (pull_request) Successful in 40s
test / unit (pull_request) Successful in 46s
test / integration-firecracker (pull_request) Successful in 3m35s
test / coverage (pull_request) Successful in 43s
test / publish-infra (pull_request) Has been skipped

Adds the `nested_containers` bottle flag. On the macOS backend it starts
a rootless podman service inside the bottle and exposes its
Docker-compatible API socket, so the agent still runs `docker` and
`docker compose`. No host daemon socket is mounted and the guest gains
no capabilities; backends that cannot run a guest-local engine reject
the flag in the shared prepare template rather than ignoring it.

Podman rather than rootless Docker because Apple Container's capability
bounding set omits CAP_SYS_ADMIN, which the kernel requires to write a
multi-range uid_map via newuidmap. With no subordinate UID range podman
falls back to a single-UID self-mapping an unprivileged process may
write itself, so the image build strips /etc/subuid and /etc/subgid
entries rather than adding them.

That mapping is also why nested containers are not an isolation layer:
root inside one is the agent user outside it. They are a build/test
convenience; the bottle remains the boundary.

Ports the spike branch onto main, renaming docker_access — it implied
Docker and granted access to nothing on the host — and drops podman
from the derived layer now that every built-in image ships it (#451).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-07-21 15:39:11 -04:00
parent ef89ed084f
commit ba25845886
20 changed files with 1219 additions and 9 deletions
+106
View File
@@ -0,0 +1,106 @@
# PRD prd-new: Containers inside a bottle
- **Status:** Draft
- **Author:** Claude
- **Created:** 2026-07-21
- **Issue:** #392
## Summary
Let an agent run `docker` and `docker compose` *inside* its bottle by starting
a guest-local rootless podman service that exposes a Docker-compatible API
socket. Gated per bottle by `nested_containers: true`. No host daemon socket is
mounted and no capability is added to the guest, on any backend.
## Problem
Agent tasks routinely involve containers: `docker compose up` to verify a
scaffolded stack, a throwaway database for a test run, building an image to
check a Dockerfile works. Bottles have no way to do any of that, so those tasks
either fail or get done outside the bottle — which defeats the point of the
bottle.
## Goals / success criteria
- A bottle with `nested_containers: true` on the `macos-container` backend can
run `docker compose up -d --wait` against a stock compose file, pulling from
a routed registry and serving from the workspace.
- The agent's habits do not change: `docker` and `docker compose`, not
`podman`.
- No host Docker socket is mounted, on any backend.
- No capability is added to the Apple Container guest.
- Backends without a guest-local engine reject the flag with a clear error
rather than ignoring it.
- Bottles that do not set the flag carry none of its cost.
## Non-goals
- Nested containers as an isolation layer. The single-UID mapping means `root`
inside a nested container is the agent user outside it. This is for build and
test workloads; the bottle stays the security boundary.
- The `docker` and `firecracker` backends. Both reject the flag for now.
- Reaching registries that are not routed through the bottle's egress proxy.
## Design
### Why podman, not rootless Docker
Apple Container's capability bounding set omits `CAP_SYS_ADMIN`, which the
kernel requires in order to write a multi-range `uid_map` via `newuidmap`.
Rootless Docker has no path that avoids that write, so it cannot run in a
bottle without granting `CAP_SYS_ADMIN` — which is close to root in practical
terms and gives up most of what the bottle is for.
Podman does have a path: with **no** subordinate UID range configured it falls
back to a single-UID self-mapping, which an unprivileged process may write
itself. The image build therefore *removes* the agent user's `/etc/subuid` and
`/etc/subgid` entries rather than adding them — their presence is exactly what
would send podman down the `newuidmap` path. Full negative result in
[`docs/research/rootless-docker-in-apple-container-spike.md`](../research/rootless-docker-in-apple-container-spike.md).
### The flag, and its name
The flag is kept because it gates real costs, not because podman needs a
privilege grant: ~100MB of derived image, a resident service per bottle, and
relaxed modes on `/dev/fuse` and `/dev/net/tun` that the majority of bottles
should not get.
It is named `nested_containers`, not `docker_access`. It implies no Docker and
grants access to nothing on the host — naming it after Docker access would
describe the one thing the design refuses to do.
### Shape
- `bot_bottle/backend/macos_container/nested_containers.py` — derived-image
build, guest env, device preparation, service start/readiness.
- `bot_bottle/backend/macos_container/nested-containers-init.sh` — the
unprivileged bootstrap that runs inside the bottle. Fails closed on every
prerequisite it needs (podman, the storage/network helpers, writable device
nodes, an *empty* subordinate range) rather than degrading.
- The derived image `…-nested-containers` layers the Docker CLI, the compose
plugin, `fuse-overlayfs`, `slirp4netns`, and `uidmap` onto the agent image.
Podman itself already ships in every built-in image (#451).
- `BottleBackend.supports_nested_containers` — false by default, so a backend
that cannot honor the flag fails in the shared `prepare` template.
Guest configuration the single-UID mapping forces:
`ignore_chown_errors` (no second UID for layers to be chowned to),
`cgroups="disabled"` and `cgroup_manager="cgroupfs"` (no cgroup delegation
reaches the guest), `events_logger="file"` (no journald socket).
### Registry reach
Image pulls egress through the bottle's proxy like everything else, so each
registry needs a route. Docker Hub and GHCR need `preserve_auth: true` — their
token dance uses a client-fetched per-scope bearer token that the proxy would
otherwise strip. Layer pulls should also set `dlp.outbound_detectors: false`
and `dlp.inbound_detectors: false`: body scanning on a multi-hundred-MB layer
is what triggers #455, and a registry route's scanned bodies are compressed
layer blobs rather than anything a detector can read.
## Open questions
- Whether the `docker` and `firecracker` backends want an equivalent, or
whether guest-local containers stay macOS-only.
- Outbound networking *from* a nested container (DNS, proxy env, CA) beyond
what the compat socket gives the bottle itself.
@@ -0,0 +1,86 @@
# Egress proxy OOMs on large downloads
Found on 2026-07-21 while running the rootless-podman spike
(`docs/research/rootless-docker-in-apple-container-spike.md`). Recorded
rather than fixed — the fix is a security-relevant decision, not a
mechanical patch.
## Summary
A single large HTTPS download through the gateway kills the egress
proxy. `mitmdump` buffers whole response bodies so the DLP detectors can
scan them, grows past the gateway container's memory limit, and is
OOM-killed by the cgroup. Nothing restarts it.
Two properties make this worse than a failed download:
- **The gateway is a per-host singleton.** Every bottle shares it, so
one bottle's download takes egress away from all of them.
- **There is no restart on death.** The gateway supervisor is
`while : ; do wait ; done`; a killed daemon stays dead until the infra
container is recreated.
So ordinary agent activity — pulling a container image, downloading a
model or dataset, fetching a large tarball — is a denial of service
against every other bottle on the host. No malice required, though it is
trivially reachable on purpose.
## Evidence
Triggered by `docker compose up` pulling `quay.io/fedora/python-312`
(two layers, ~82MB and ~83MB) inside a bottle. The pull itself
succeeded; the *next* request failed:
```
initializing source docker://quay.io/fedora/python-312:latest:
pinging container registry quay.io: Get "https://quay.io/v2/":
proxyconnect tcp: dial tcp 192.168.128.39:9099: connect: connection refused
```
From the gateway's `dmesg`:
```
python3 invoked oom-killer: gfp_mask=0x100cca(GFP_HIGHUSER_MOVABLE), order=0
oom-kill:constraint=CONSTRAINT_MEMCG,
oom_memcg=/container/bot-bottle-mac-infra,
task_memcg=/container/bot-bottle-mac-infra,task=mitmdump,pid=118
Memory cgroup out of memory: Killed process 118 (mitmdump)
total-vm:1391936kB, anon-rss:997768kB
```
~1GB RSS against a 1024MB container. Note the amplification: ~165MB of
layers produced ~1GB of resident memory, so the buffering is several
copies deep (encoded body, decoded body, and the text conversion the
regex detectors scan).
Afterwards the gateway container was still running and healthy-looking —
orchestrator, supervise, and git-http all alive — with no `mitmdump`
process at all, and it stayed that way until the container was
recreated. A liveness check on the container would not have caught this.
## Reproduction
1. Launch any bottle with an egress route to a host serving a large file.
2. Download >~150MB over HTTPS through the proxy.
3. `dmesg | grep -i oom` inside `bot-bottle-mac-infra`, and note that no
`mitmdump` process remains.
Beware a false negative when checking: truncating the process listing
(`cut -c1-45`) cuts before the binary name, because `mitmdump` runs as
`/usr/local/bin/python3.12 /usr/local/bin/mitmdump …`.
## Fix options, not yet chosen
1. **Restart dead daemons.** Smallest change and strictly an
improvement: an OOM then degrades one download instead of removing
egress for every bottle. Does not stop the OOM.
2. **Cap the scanned body size.** Above a threshold, stop buffering —
either skip the scan or stream it. This is the root-cause fix and a
security decision: a size threshold is exactly the hole an exfiltrator
would aim for, so "skip above N" trades a DoS for a covert channel.
Streaming with a bounded window keeps coverage, at more complexity.
3. **Raise the gateway's memory limit.** Moves the threshold; does not
remove it.
Worth noting that (1) and (2) are complementary — the restart gap is
worth closing regardless of how the memory behaviour is resolved.
@@ -0,0 +1,359 @@
# Rootless Docker inside Apple Container bottles
Spike branch: `spike/rootless-docker-macos` (`a4d8461`)
**Outcome:** the podman recommendation below shipped as the
`nested_containers` bottle flag — see
[`docs/prds/prd-new-nested-containers.md`](../prds/prd-new-nested-containers.md).
The `docker_access` name used throughout the spike text was renamed on the
way in; it granted no access to anything on the host.
## Summary
**Negative result.** Rootless Docker cannot run inside an Apple
Container bottle without granting the bottle `CAP_SYS_ADMIN`. This is a
kernel constraint on writing multi-range `uid_map`, not a packaging gap
we can close with a better init script, a different base image, or more
careful `/etc/subuid` handling.
The spike was built on the premise — stated in
`bot_bottle/backend/macos_container/rootless_docker.py` — that it would
*"deliberately refuse to compensate for missing prerequisites with outer
capabilities, a privileged container, or a host Docker socket."* That
premise is exactly what the experiment falsified. The two ways forward
are to abandon the premise (add `CAP_SYS_ADMIN` to the bottle, and with
it most of the isolation the bottle exists to provide) or to abandon
rootless Docker.
Recommendation: abandon rootless Docker. Podman does not have this
problem — see [Podman is not blocked by
this](#podman-is-not-blocked-by-this) below.
## Local environment
Tested on 2026-07-21:
```console
$ sw_vers
ProductName: macOS
ProductVersion: 26.5.1
BuildVersion: 25F80
$ container --version
container CLI version 1.0.0 (build: release, commit: ee848e3)
$ uname -a # inside the bottle
Linux ... 6.18.15 #1 SMP Tue Mar 17 01:36:53 UTC 2026 aarch64 GNU/Linux
```
## The failure
`tests/integration/test_macos_rootless_docker_spike.py` builds the
image, launches the bottle, and dies in `rootless_docker.start`:
```
+ exec rootlesskit --net=slirp4netns --mtu=65520 ... dockerd-rootless.sh
[rootlesskit:parent] error: failed to setup UID/GID map:
newuidmap 1100 [0 1000 1 1 100000 65536] failed:
newuidmap: write to uid_map failed: Operation not permitted
```
## Why it fails
Every prerequisite you would normally suspect is present and correct in
the guest:
| Check | Result |
| --- | --- |
| `/usr/bin/newuidmap` | `-rwsr-xr-x root root` — setuid bit intact, survived the OCI export |
| `/` mount options | `rw,relatime`**not** `nosuid` |
| `NoNewPrivs` | `0` |
| `Seccomp` | `0`, no filters |
| `/etc/subuid`, `/etc/subgid` | `node:100000:65536` in both |
| user namespace | `user:[4026531837]`, identical to pid 1 — the *initial* userns |
| `unshare -U -r true` | succeeds |
| `/proc/sys/user/max_user_namespaces` | `4505` |
The one thing that is missing is in the capability bounding set that
Apple Container gives the container:
```
CapBnd: 00000000a80425fb
= chown, dac_override, fowner, fsetid, kill, setgid, setuid, setpcap,
net_bind_service, net_raw, sys_chroot, mknod, audit_write, setfcap
```
No `CAP_SYS_ADMIN`. That is the whole story, and the chain is:
1. The kernel's `map_write()` gates writing a `uid_map` on
`file_ns_capable(file, ns, CAP_SYS_ADMIN)` — capability over the
**new** user namespace, evaluated against the credentials that opened
`/proc/<pid>/uid_map`.
2. `newuidmap` is setuid-root, so it runs with euid 0 — but its
capability sets are clamped by the bounding set, which has no
`CAP_SYS_ADMIN`.
3. `cap_capable()` has a shortcut that grants *all* capabilities when
the caller's userns is the new namespace's parent **and**
`ns->owner == cred->euid`. It does not apply: the namespace was
created by `node` (uid 1000) while `newuidmap` runs as euid 0.
4. So the check falls through to the effective-set test in the initial
userns, which fails. `EPERM`.
Note that the single-line unprivileged path (`unshare -U -r`) works
precisely because it does not go through `newuidmap` and does not need
`CAP_SYS_ADMIN`. Only the multi-range subuid mapping that rootless
Docker requires does.
This is the same constraint that makes upstream's `dind-rootless` image
require `--privileged`. It is not specific to Apple Container, except
that Apple Container gives us no bounding set that includes
`CAP_SYS_ADMIN` by default.
## It does work with the capability — which is the point
Adding the capability clears the failure immediately, and exposes one
further, much smaller blocker: `/dev/net/tun` exists (the kernel has
tun; `/proc/misc` lists `200 tun`) but Apple Container creates it
`crw------- root root`, so uid 1000 cannot open it and `slirp4netns`
fails with `open: Permission denied`. A `chmod 0666 /dev/net/tun` as
root inside the bottle fixes that, and needs no capability beyond what
the bottle already has.
With both applied by hand, the daemon comes up completely:
```console
$ container run --rm -u root --cap-add CAP_SYS_ADMIN \
bot-bottle-claude:latest-rootless-docker sh -c '...'
Server Version: 20.10.24+dfsg1
Storage Driver: fuse-overlayfs
Cgroup Driver: none
Cgroup Version: 2
API listen on /tmp/rt/docker.sock
```
So `rootless-docker-init.sh` and `rootless_docker.py` are *correct*.
The spike did not fail on a bug. It failed on its own premise.
Two secondary findings from that run, relevant if anyone revisits this:
- Debian's `docker.io` package pins Docker **20.10** (EOL), not the 28.x
implied by the `docker:28-cli` compose plugin the image copies in.
- `Cgroup Driver: none` — no resource limits on nested containers.
## Why we should not just add the capability
`CAP_SYS_ADMIN` is close to a superset of "root" in practical terms —
mount, `pivot_root`, namespace manipulation, and a long tail of
subsystem-specific powers. Granting it to the agent bottle would
undercut the containment argument the rest of the backend is built
around, including the deliberately narrow choices immediately adjacent
to it in `launch.py` (`--cap-drop CAP_NET_RAW`, no `NET_ADMIN`, a
host-only agent network). Trading all of that for nested `docker
compose` is a bad exchange.
## Podman is not blocked by this
Sanity-checked on the same host, same kernel, same runtime, so the
comparison is apples to apples:
| Scenario | Result |
| --- | --- |
| Podman rootless, `/etc/subuid` populated | **Fails identically**`newuidmap: write to uid_map failed: Operation not permitted` |
| Podman rootless, no subuid ranges, `--network=host` | **Works**, no added capabilities |
| Podman rootless, no subuid ranges, default netns, `/dev/net/tun` at `0600` | Fails — `slirp4netns: open("/dev/net/tun"): Permission denied` |
| Podman rootless, no subuid ranges, default netns, `/dev/net/tun` at `0666` | **Works**, no added capabilities |
The difference is that podman degrades gracefully when no subuid range
is available: it falls back to a single-UID self-mapping, which an
unprivileged process may write itself, so `newuidmap` is never invoked
and `CAP_SYS_ADMIN` is never needed. Docker's rootless mode has no
equivalent fallback.
The cost of that fallback is real and should be weighed before building
on it: with a single-UID mapping, every UID inside a nested container
collapses onto the bottle's own uid 1000. There is no UID separation
between the agent and anything it runs — `root` in a nested container is
the agent user outside it. It also requires `ignore_chown_errors` on the
storage driver. Whether that is acceptable depends on whether the bottle
boundary (which is unchanged) or the nested-container boundary (which is
effectively nil) is the one we are relying on.
## What the podman spike then needed
The podman implementation that replaced the Docker one on this branch
turned up two more device-node blockers of the same shape as
`/dev/net/tun` — Apple Container creates the node, but 0600 root:root:
- **`/dev/fuse`** — blocks the `fuse-overlayfs` storage driver
(`fuse: failed to open /dev/fuse: Permission denied`). Without it the
only working driver is `vfs`, which copies whole layers per container.
- **`/dev/net/tun`** — blocks `slirp4netns`, which rootless podman uses
for the default bridge network.
Both are fixed by `chmod 0666` as root inside the bottle, which needs no
capability the bottle does not already hold. This is categorically
different from the `CAP_SYS_ADMIN` requirement: it is a permission on a
node that already exists, not an outer privilege grant.
One design note worth recording: the agent-facing surface stays `docker`
and `docker compose`, pointed at podman's Docker-compatible API socket
via `DOCKER_HOST`. Setting `netns="host"` in `containers.conf` does *not*
propagate through that compat API — stock `docker run` and compose files
request bridge networking explicitly — so slirp4netns (and therefore the
`/dev/net/tun` chmod) is required for ordinary compose files to work at
all. Host networking remains available per-workload via
`--network=host`.
Verified working in a bottle with zero added capabilities: fuse-overlayfs
storage, the compat API socket, `docker run` on both bridge and host
networking, and published ports.
### Nested pulls collide with our own egress DLP
The first live run got podman up and `docker compose` running, then
failed on the image pull:
```
web Pulling
initializing source docker://python:3.12-alpine: reading manifest ...
StatusCode: 403, egress DLP: Generic Bearer JWT found in body
```
This is bot-bottle's own egress scanner, not a podman problem. The
Docker registry auth flow carries a bearer JWT *by protocol*, and the
`token_patterns` detector's `Generic Bearer JWT` rule
(`Bearer\s+[A-Za-z0-9._\-]{50,}`) matches it on every pull. Any bottle
that pulls images will hit this.
The fix is per-route detector scoping, which the egress config already
supports — drop `token_patterns` on the registry hosts and keep
`known_secrets`:
```json
{"host": "registry-1.docker.io",
"dlp": {"outbound_detectors": ["known_secrets"]}}
```
That is the right trade rather than a grudging one: `known_secrets`
matches the bottle's *actual* credential values, so real exfil through a
registry host is still caught. `token_patterns` on a registry route only
ever produces protocol noise.
Worth generalising later: any manifest enabling `nested_containers` needs
this on its registry routes, so it probably belongs in a shared
registry-route snippet rather than being copy-pasted per bottle.
### And then registry auth collides with the Authorization strip
With DLP scoped, the pull failed differently: `unauthorized:
authentication required`. This one is architectural.
`egress_addon.py` strips agent-set `Authorization` unconditionally
before forwarding — deliberately, so an agent cannot smuggle a
credential out in a header the DLP detectors don't recognise. A route
may carry gateway-injected auth instead, but only from a *static* token
in an env var (`auth_scheme` + `token_env`).
Docker registry auth doesn't fit that shape. The client fetches a
short-lived, per-repository-scope bearer token from `auth.docker.io` and
presents it to `registry-1.docker.io`. There is no static token to
inject, and the token the client legitimately obtained is stripped.
Measured inside a bottle, by hand:
| Step | Result |
| --- | --- |
| Fetch token from `auth.docker.io` | 200, 5409-byte token body |
| Manifest request **with** that valid token | 401 |
| Manifest request with **no** Authorization | 401 — identical |
A valid token behaves exactly like sending none, which is direct
evidence the header never arrives. Any nested-container workflow that
pulls from a registry is blocked on this, so it is not a detail that can
be deferred: pulling base images is most of what nested containers are
for.
### Registries that skip the token dance work today
Not every registry needs the stripped header. Measured directly:
| Registry | Manifest request with no `Authorization` |
| --- | --- |
| `quay.io` | 200 |
| `mcr.microsoft.com` | 200 |
| `registry.k8s.io` | 307 (redirect, no auth) |
| `ghcr.io` | 401 |
| `registry-1.docker.io` | 401 |
So "just add the registry to the bottle config" genuinely works — for
quay, MCR, registry.k8s.io, or any unauthenticated internal registry.
Docker Hub and GHCR are the ones that need the strip resolved. The
acceptance test uses quay for exactly this reason.
Resolving it for Docker Hub means picking one of:
1. **Per-route opt-in to preserve client Authorization.** Smallest
change. Note the compounding effect on exactly these routes: the DLP
scoping above already removed `token_patterns` there, so a
preserved-auth registry route is one where the agent may send bearer
tokens that neither the strip nor the pattern detector inspects.
`known_secrets` still applies, so the bottle's real credentials are
still caught.
2. **A registry-aware gateway** that performs the token dance itself and
injects the result. Preserves the invariant fully; materially more
work, and it makes the gateway speak a specific registry protocol.
3. **Pre-seed images at provision time** (host-side `container image
save` into podman storage), so bottles never pull at runtime.
Preserves the invariant, and limits nested containers to
pre-approved images — which fits the custody positioning, at the cost
of no ad-hoc `docker pull`.
4. **Stop.** Nested containers are not supported on this backend.
### Podman 4.3.1 silently swallows container exit codes
Debian bookworm — which the current agent base image is built on —
ships podman 4.3.1. Through its Docker-compatible API, `docker run`
returns 0 no matter what the container did:
| Command | podman 4.3.1 | podman 5.4.2 |
| --- | --- | --- |
| `docker run … sh -c 'exit 7'` (compat API) | **0** | 7 |
| `docker run … sh -c 'exit 0'` (compat API) | 0 | 0 |
| `podman run … sh -c 'exit 7'` (native) | 7 | 7 |
This is worse than a broken feature: every failing command an agent runs
via `docker run` reports success. A test suite, a build step, or a CI
script inside a bottle would pass while failing. It also silently
defeated the acceptance test's egress-containment assertion, which is
why that assertion now checks an in-band marker rather than an exit
code.
Podman 5.4.2 (Debian trixie) fixes it, but needs two packages that
bookworm's podman does not: `passt` (podman 5's default network tool)
and `nftables` (netavark shells out to `nft`; without it every run fails
with `unable to upgrade to tcp, received 500`). With both installed,
exit codes propagate correctly and the compat API behaves.
The open question this leaves is where podman 5 comes from, since the
agent base is bookworm-based:
1. **Move the agent images to Debian trixie.** Trixie is current stable.
Correct, and the blast radius is every bottle, not just this feature.
2. **Drop the compat socket and use podman natively** (`podman-docker`
provides a `docker` shim; compose comes from `podman-compose`).
Native podman propagates exit codes correctly even on 4.3.1. Contained
to this feature, at the cost of `docker compose` becoming
`docker-compose`/`podman-compose`.
3. **Ship bookworm's podman 4.3.1 with the compat socket** — not viable.
Silent false success is a correctness bug agents cannot see.
## Recommendation
1. Do not revive rootless Docker on this backend. This document is the
record of why.
2. Nested containers, if wanted, come from podman under the
single-mapping constraint — with the explicit understanding that the
nested-container boundary carries no security weight. `root` in a
nested container is the agent user outside it.
3. Nested containers are therefore a build/test convenience. The bottle
remains the security boundary, exactly as it was.