Completes the matrix: all five distros now pass both `test` and `test-ready`,
validated end-to-end on the delphi KVM host (QEMU 11.0.2).
Alpine — the generic_ cloud image silently ignores a NoCloud seed, so cloud-init
never ran. Fixes, all confirmed by booting:
- switch to the nocloud_ image variant (built for a local seed),
- start sshd from a runcmd (OpenRC doesn't auto-start it after key injection),
- give the account a throwaway password — Alpine's non-PAM sshd refuses pubkey
auth for a cloud-init-locked ('!') account, unlike the UsePAM=yes distros,
- install sudo via cloud-init packages: (the minimal image has none, so the
test-ready `sudo apk add` prereq failed).
Also add instance-id to the NoCloud meta-data, which the nocloud image requires.
NixOS — publishes no downloadable cloud qcow2 (Hydra builds AMIs), so the harness
now BUILDS one with nixos-generators (new scripts/linux-install-test-nixos.nix:
cloud-init for the key, flakes enabled, deliberately no python/git/pipx).
ensure_base_image branches to build_nixos_image for the nixos distro. The
test-ready prereq uses `nix profile install` pinned to nixpkgs/nixos-24.11,
because the guest's default unstable registry builds pipx from source and its
test suite currently fails to build.
Validation: 10/10 cells pass, clean teardown, no leaked VMs or run dirs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
First full run on the KVM host surfaced three harness bugs and two stale
image URLs; all fixed here.
Harness bugs:
- Liveness probe used `bot-bottle --version`, which the CLI does not implement
(unknown args die non-zero) — so every SUCCESSFUL install was misreported as
"no runnable entry point". Switched to `bot-bottle --help`, which exits 0
before any DB/migration/network work.
- cmd_down never removed serial.log, so its rmdir failed and every run left an
orphan scratch dir behind. Added serial.log to the cleanup.
- The teardown trap was armed AFTER cmd_up, but cmd_up's wait_for_ssh can fail
with QEMU already running (a guest that never opens SSH) — leaking the VM.
Arm the trap before cmd_up.
Stale image URLs:
- Fedora 41 is EOL and 404s; bumped to Fedora 44 (44-1.7).
- Alpine bumped to 3.21.7; its cloud images ship a .sha512 only (this verifier
is sha256), so SUM_URL is now empty (skip) with a note.
Validated on delphi (QEMU 11.0.2): ubuntu/fedora/arch pass BOTH test and
test-ready. Alpine (cloud-init seed not applied → no SSH) and NixOS (no
upstream cloud qcow2) remain blocked on image provisioning, documented in the
research note as follow-ups.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replaces the BB_TEST_PREREQS env toggle with the two subcommands the macOS
harness established, so the two repos read the same:
test A bare host — prerequisites NOT set up (the default state of a
stock cloud image). Exercises install.sh's own prerequisite-guard
logic. SOUND (PASS) when install.sh either installs cleanly or
declines with one of its own recognized, actionable errors; a
crash or unrecognized failure fails.
test-ready Prerequisites satisfied — the harness installs python3/git/pipx
first, then runs install.sh, which must actually land: entry point
runnable, doctor reporting a usable python and config.
Structure now mirrors scripts/macos-install-test.sh: a shared _test_cycle
driving up -> (prereqs) -> run -> verdict -> down, thin cmd_test / cmd_test_ready
wrappers setting _STEPS / _PASS_CLAIM / _REQUIRE_INSTALL, a standalone `prereqs`
subcommand, and a _test_teardown that prints the PASS/FAIL claim.
The doctor check is now the macOS-style classifier rather than a bare exit-code
read: a Traceback is an install defect (fail), a missing `ok: python:` /
`ok: config:` line is a fail, and a not-ready backend is reported, not fatal
(BB_TEST_REQUIRE_BACKEND=1 makes it fatal for a nested-virt host). This is where
Linux diverges from macOS test-ready — a plain VM has no nested KVM for
Firecracker and the harness doesn't provision the Docker backend, so backend
readiness is never the Linux criterion.
test-all now runs every distro × {test, test-ready}. Research note updated.
Validated with `bash -n` and `shellcheck`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Splits the Linux clean-install harness into two variants, selected with
BB_TEST_PREREQS:
with (default) — install python3 + git + pipx first, then install.sh.
The ready-host happy path; PASS = install.sh exits 0 and
leaves a runnable bot-bottle entry point.
without — run install.sh on the BARE cloud image. Exercises
install.sh's own prerequisite-guard logic (python gate,
git-for-git-specs gate, pipx/pip PEP-668 handling). PASS =
install.sh either fully succeeds OR declines with one of
its own recognized, actionable prerequisite errors; a
crash or unrecognized failure is a FAIL.
cmd_run now captures install.sh's exit code + full output (install.rc /
install.log) instead of aborting on non-zero, so the verdict step applies the
variant's criterion. The old assert_installed is replaced by a quiet
entry_point_runnable helper plus classify_outcome, which matches the bare-host
declines against the exact die() messages install.sh prints. doctor still runs
for visibility, but only when an entry point exists.
test-all now runs the full matrix -- every distro x both variants -- each cell
in its own throwaway VM, and can be pinned to one variant via BB_TEST_PREREQS.
Research note updated to document both variants and their pass criteria.
Validated with `bash -n` and `shellcheck`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds scripts/linux-install-test.sh, a throwaway-VM harness that exercises
install.sh the way a brand-new user would across a Linux distro matrix
(Ubuntu, Fedora, Arch, Alpine, NixOS), plus the research note
docs/research/testing-clean-install-on-linux.md that motivates the approach.
This is the Linux counterpart to the macOS clean-install harness. On Linux
the boundary of choice flips from a throwaway user account to a disposable
KVM VM: it's a genuine kernel + userland + package-manager boundary that
wipes to nothing on teardown, and a single harness can swap distro cloud
images to cover the several package-management regimes Linux fragments into.
The host already needs KVM for the Firecracker backend, so a per-run,
copy-on-write VM (qemu-img create -b base) is cheap here -- the equivalent
of `docker run --rm`, for a whole machine.
Per run the harness caches one read-only base image, boots a throwaway
overlay via QEMU/KVM with user-mode networking (no root, no bridge), injects
an ephemeral SSH key + passwordless login through a cloud-init seed ISO,
installs the distro's prerequisites (python3 + git + pipx), pipes THIS
checkout's install.sh into the guest exactly as `curl ... | sh` would, and
asserts the CLI installed cleanly. The overlay is deleted on teardown, so
even the OS-level prerequisites are wiped -- unlike a throwaway user, the
reset is total.
Subcommands mirror the macOS harness (up/run/status/down/test) plus test-all
(the matrix) and ssh (an interactive guest shell). test arms an EXIT/INT/TERM
trap the moment the VM exists, so a failure or Ctrl-C still tears it down.
Scope is installer correctness, not runtime: there is no nested KVM/Docker in
the VM, so `bot-bottle doctor` correctly reports every backend not-ready and
exits non-zero by design. The pass criterion is therefore install.sh exiting
0, the bot-bottle entry point being present on the fresh user's PATH, and
`bot-bottle --version` running -- not a green doctor. doctor's output is still
printed so a real installer regression (broken shim, import error) stays
visible.
Validated with `bash -n` and `shellcheck`. Runtime is host-only (needs
/dev/kvm, qemu, cloud-localds), so like the macOS harness it isn't exercised
by the Linux PR CI. The cloud-image URLs in the DISTRO table are the one
place to bump when a distro cuts a newer build.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
`doctor` told users to run `./cli.py backend setup`. cli.py is a four-line
wrapper at the repo root that calls bot_bottle.cli:main, and it does not ship —
[tool.setuptools.packages.find] includes only bot_bottle*, so anyone who
installed rather than cloned has no such file. Confirmed against a real
install: the user's tree contains exactly one executable, `bot-bottle`, and no
cli.py anywhere, while doctor recommended `./cli.py` three times. Invisible in
development, where ./cli.py works fine from a checkout, which is why only a
clean-install test surfaced it.
The two entry points are the same code, so the fix is to name the one that
always exists. 141 replacements across 45 files: the runtime messages that
caused this, plus README, docs, PRDs, research notes and test prose, so nothing
teaches the invocation a user cannot run.
scripts/demo.sh is deliberately untouched — it *executes* ./cli.py from a
checkout, where that is the correct and available path.
Verified end to end: a sandbox install now reports
"Run: bot-bottle backend setup --backend=firecracker", and bot-bottle is on
that user's PATH. Unit suite unchanged against the pre-existing baseline.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEfZZhakx13bxTfXcZCoS5
Both variants now pass, and the with-prerequisites run settled several things
that were guesses before:
* The Apple `container` service is per-user, confirmed directly rather than
inferred: the throwaway account's appRoot is its own
(/Users/bbtest/Library/Application Support/com.apple.container/), and
starting it left the admin's service running and its agents intact.
* Setting up a new account is two steps, not one. The guest kernel lives in
that same per-user app root, so a fresh account has none and
`container system start` prompts to download it — and only prompts, since
the flags default to asking. Headless callers need
--enable-kernel-install.
* Neither step is performed by `bot-bottle backend setup
--backend=macos-container`, which checks and then defers to `container
system start`. So the real path for a new account is
`container system start --enable-kernel-install`.
* Any of it requires entering the user's launchd domain with `launchctl
asuser`, because the apiserver is a per-user agent reached over XPC.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEfZZhakx13bxTfXcZCoS5
"Does the installer work" and "can a new user actually run a bottle" are
different questions, and collapsing them is what made the previous run
ambiguous. Split them into two cycles over the same throwaway account:
test a macOS system WITHOUT the prerequisites set up for this user,
which is the default state of every new account since the Apple
`container` service is per-user. Asserts the install is sound;
reports backend readiness without failing on it, because
install.sh provides no backend and cannot regress one.
test-ready the same system WITH them. Runs `container system start` for the
throwaway user, then demands doctor go fully green, backend
included — the end-to-end claim.
The new `prereqs` step runs `container system start` rather than
`bot-bottle backend setup --backend=macos-container`, because that subcommand
does not start anything: it checks, then tells you to run `container system
start` yourself (backend/macos_container/setup.py, "no privileged host setup
required"). The harness runs what the product actually asks for.
test-ready sets BB_TEST_REQUIRE_BACKEND, so it demands exactly what `test`
merely reports. Both share one cycle function; the step count and the PASS
claim are the only differences beyond the extra step, so neither variant can
drift from the other's teardown or dirty-account guarantees.
Rig covers the distinguishing case directly: identical host state where `test`
passes on install soundness and `test-ready` starts the service and reaches a
green backend, plus the service-start failure and an install failure under
test-ready. 22 scenarios, all passing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEfZZhakx13bxTfXcZCoS5
The first run against the branch's own code confirmed the doctor fix — no
traceback, the firecracker probes now report "TAP pool: 0/8" instead of dying
on PermissionError — and then failed for a reason the installer cannot fix.
The Apple `container` service is per-USER. `container system status` reports an
appRoot under ~/Library/Application Support/com.apple.container/, and bbtest
sees "container system service: NOT running" while the creating admin's is
running. A brand-new account therefore has no backend until it runs
`container system start` once, doctor rightly fails, and `test` would be
permanently red for host state that install.sh neither creates nor can regress.
So stop treating doctor's exit code as one verdict. `test` now fails when the
install is broken — no entry point, doctor crashed, or doctor could not report
a usable python and config dir — and reports backend readiness as a note.
BB_TEST_REQUIRE_BACKEND=1 restores the strict behaviour.
A traceback is checked for explicitly rather than inferred from the exit code,
because those are the same value. The PermissionError bug exited non-zero
exactly like a missing prerequisite does; only the traceback distinguishes a
defect from an environment fact, and that distinction is the whole point of
this change.
The research doc's footprint table had the service as a host-wide launchd
service that survives user deletion. It does not. The state lives in the home
and goes with it, so the reset is more complete than claimed — but the
corollary is that a fresh account cannot run a bottle until the service is
started for it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEfZZhakx13bxTfXcZCoS5
The harness's second run hit the wall the first one predicted: a fresh account
has no pipx, so install.sh fell to `pip install --user`, and every Python a Mac
offers — Homebrew and python.org alike — is externally managed, so PEP 668
blocked it. That fallback was never a fallback on macOS; it was a dead end that
printed instructions.
Replace it with a venv at ~/.bot-bottle/venv (BOT_BOTTLE_VENV to move it),
with the console script symlinked into ~/.local/bin. PEP 668 does not apply
inside a venv, and venv is stdlib, so unlike pipx there is nothing to bootstrap
first. pipx stays the preferred path when present, so anyone already managing
their Python apps that way is unaffected — and the post-install PATH check now
asks pipx for PIPX_BIN_DIR instead of assuming ~/.local/bin.
Keeping the venv under ~/.bot-bottle rather than ~/.local/share means the whole
footprint stays in one directory, which is what lets the throwaway-account
teardown remain a complete reset.
This removes the PEP 668 pre-flight and the sysconfig user-scheme lookup, both
of which existed only to serve the --user path. Their tests go with them:
* `detects_externally_managed_python` asserted the check that is now moot;
replaced by one asserting pipx is still preferred when present.
* `checks_pip_usable_before_fallback` pinned a pip probe that no longer runs;
replaced by one asserting the venv's own pip does the install, since using
the base interpreter's would install outside the venv.
* `resolves_user_scripts_dir_not_hardcoded` and
`macos_user_scheme_is_not_dot_local_bin` guarded the ~/Library/Python
scripts-dir lookup. Nothing installs there now. The surviving "don't
hardcode" concern is pipx's bin dir, which has its own test.
Five tests are added for the new path: the venv fallback exists, no --user path
survives, the venv is under the config dir, venv creation failure names
python3-venv (Debian ships it separately), and the entry point is exposed
outside the venv.
Verified end to end in a sandbox HOME with a fresh-account PATH and no pipx:
venv built, package installed, symlink created, `doctor` reached and green
(python 3.14.5, macos-container ready), exit 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEfZZhakx13bxTfXcZCoS5
`test` chains up -> run -> status -> down, which is the loop you actually want
when verifying a clean install. Three things make it more than a convenience
wrapper:
* It refuses to start against an existing account. A reused home is not a clean
install, so testing one silently would defeat the harness.
* It tears the account down from an EXIT/INT trap armed the moment the account
exists, so a failed run — or a Ctrl-C mid-install — still leaves the machine
clean. BB_TEST_KEEP=1 opts out to poke at a failure.
* Its verdict is stricter than the installer's. install.sh exits 0 when it
finishes but `doctor` reports unmet prerequisites, so "the installer
succeeded" is not a useful assertion; `test` fails if the install fails, if
bot-bottle never reached the new user's PATH, or if doctor is unhappy. That
meant giving cmd_status a real exit status instead of swallowing doctor's.
Also fixes two bugs in `run`'s installer staging, by removing the staging
entirely and feeding install.sh in on stdin:
* `mktemp /tmp/bb-install.XXXXXX.sh` did not do what it looks like. BSD mktemp
only substitutes trailing Xs, so every run wrote the *same* predictable path,
as root, mode 644, in a world-writable directory.
* The cleanup only ran on the normal and failure returns, so an interrupted run
leaked the file.
Piping on stdin sidesteps both: root opens the redirect before sudo drops
privileges, so the mode-700 home that motivated the staging is a non-issue,
there is no file to leak, and `sh -s` matches the documented `curl … | sh`
shape more closely than executing a staged copy did.
The command-level `exit 1`s become `return 1` so the steps compose under the
trap, and the "next, run this" hints are suppressed inside `test`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WEfZZhakx13bxTfXcZCoS5
Add scripts/macos-install-test.sh, a throwaway-user harness for exercising
install.sh the way a brand-new user would on macOS, plus the research note
that motivates the approach.
The harness has up/run/status/down/deep-reset subcommands. Because install.sh
writes only to the user home (pipx venv, ~/.bot-bottle, a PATH line) and never
installs the backend, deleting the account is a complete, deterministic reset
of the install surface. A disposable macOS VM can't stand in on M1/M2: the
Apple `container` backend needs Virtualization.framework, and running it inside
a guest VM requires nested virtualization (M3+ only), so a throwaway user is
the only way to reach the real host backend from a clean $HOME.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Lower the diff-coverage gate from 90% to 80% and the critical-module
target from 90% to 85%. The 90% diff gate forced back-fill tests on
nearly every changed line; 80% keeps new code honest without the churn.
Global coverage stays informational per ADR 0004 (no new gate added).
Updates scripts/diff_coverage.py, scripts/coverage.sh,
scripts/critical-modules.txt, .gitea/workflows/test.yml, and records the
change as a dated revision in ADR 0004.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Addresses the third review round on PR #481.
- `bot-bottle doctor` now checks `is_backend_ready()` (a full backend
status() probe: daemon reachable, network pool present, KVM usable)
instead of the cheap PATH-only `is_backend_available()`. A host with a
stopped Docker daemon or half-configured Firecracker no longer reports
`ok: backend` / exit 0 when `start` can't actually work; each not-ready
backend prints its own diagnostics, and doctor passes only if at least
one backend is ready.
- `install.sh` resolves the pip `--user` scripts directory from the
interpreter (`sysconfig.get_path("scripts", get_preferred_scheme("user"))`)
instead of hardcoding `~/.local/bin`, which is wrong on a python.org
macOS interpreter (`~/Library/Python/<X.Y>/bin`). The PATH guidance now
prints the actual directory.
Tests: doctor tests mock `is_backend_ready` (the readiness contract) and
cover the not-ready → fail path; a new install-script test drives the
macOS `osx_framework_user` scheme and asserts it resolves a
non-~/.local/bin directory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Addresses the review on PR #481.
Self-contained wheel (review point 1): the gateway/infra/orchestrator
images build from a context that must hold bot_bottle/, pyproject.toml,
and the root-level Dockerfiles. Modules previously located these by
walking __file__ to the repo root, so an installed wheel (package in
site-packages, no repo root) passed `doctor` but failed `start`.
- Add bot_bottle/resources.py: build_root() returns the repo root in a
checkout (unchanged) or a staged copy from the wheel's bundled
_resources/ otherwise; dockerfile()/nix_netpool_module()/
netpool_script() derive from it.
- setup.py bundles the root Dockerfiles, nix module, netpool script, and
pyproject.toml into bot_bottle/_resources/ at build; MANIFEST.in ships
them in the sdist.
- Route every _REPO_ROOT/_REPO_DIR call site (docker/macos launch, macos
infra, firecracker infra_vm/infra_artifact/setup, orchestrator
lifecycle/gateway) through resources. Checkout behavior is unchanged.
install.sh prerequisites (review point 2): check for git when installing
a git+ spec, and — before the pip fallback — that pip is usable and the
interpreter isn't externally managed (PEP 668), pointing at pipx.
Tests: test_resources covers checkout + staged-wheel layouts;
test_wheel_install builds the wheel, installs it into an isolated venv,
and asserts `doctor` runs and build_root() yields a valid context.
Running `start` end-to-end still needs a Docker/KVM host (CI).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>