Replaces the BB_TEST_PREREQS env toggle with the two subcommands the macOS
harness established, so the two repos read the same:
test A bare host — prerequisites NOT set up (the default state of a
stock cloud image). Exercises install.sh's own prerequisite-guard
logic. SOUND (PASS) when install.sh either installs cleanly or
declines with one of its own recognized, actionable errors; a
crash or unrecognized failure fails.
test-ready Prerequisites satisfied — the harness installs python3/git/pipx
first, then runs install.sh, which must actually land: entry point
runnable, doctor reporting a usable python and config.
Structure now mirrors scripts/macos-install-test.sh: a shared _test_cycle
driving up -> (prereqs) -> run -> verdict -> down, thin cmd_test / cmd_test_ready
wrappers setting _STEPS / _PASS_CLAIM / _REQUIRE_INSTALL, a standalone `prereqs`
subcommand, and a _test_teardown that prints the PASS/FAIL claim.
The doctor check is now the macOS-style classifier rather than a bare exit-code
read: a Traceback is an install defect (fail), a missing `ok: python:` /
`ok: config:` line is a fail, and a not-ready backend is reported, not fatal
(BB_TEST_REQUIRE_BACKEND=1 makes it fatal for a nested-virt host). This is where
Linux diverges from macOS test-ready — a plain VM has no nested KVM for
Firecracker and the harness doesn't provision the Docker backend, so backend
readiness is never the Linux criterion.
test-all now runs every distro × {test, test-ready}. Research note updated.
Validated with `bash -n` and `shellcheck`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
8.8 KiB
Testing a clean bot-bottle install on Linux
How do you exercise install.sh the way a brand-new user would — on a
pristine Linux environment you can throw away afterward — without
polluting your daily-driver host, and across the several package-management
regimes Linux fragments into? This is the Linux counterpart to
testing-clean-install-on-macos.md;
the conclusion is different because Linux gives us a boundary macOS doesn't.
Summary
On macOS the honest options were a throwaway user or a VM, and the throwaway user won on pragmatics (nested virtualization is gated to M3+). On Linux the calculus flips: a disposable KVM virtual machine, booted from a distro cloud image and deleted per run, is both the cleanest boundary and the one that lets a single harness cover Ubuntu, Fedora, Arch, Alpine, and NixOS. The host already requires KVM for the Firecracker backend, so the VM is cheap here.
The harness lives at scripts/linux-install-test.sh.
Per run it caches one read-only base image, boots a throwaway copy-on-write
overlay (qemu-img create -b base), installs the distro's prerequisites,
pipes this checkout's install.sh into the guest exactly as curl … | sh
would, asserts the CLI installed, and deletes the overlay — the Linux
equivalent of docker run --rm, for a whole machine.
Why a VM, not a container or a throwaway user
| Mechanism | Why it's the wrong boundary here |
|---|---|
Container (docker run --rm) |
Shares the host kernel and ships a deliberately minimal userland — no systemd, a stubbed-out package manager story, and (crucially) it doesn't reproduce the externally-managed Python (PEP 668) that real desktop/server installs put in front of the user. It tests "does install.sh run in a container," not "does it run on a real distro." |
Throwaway user (useradd/userdel) |
The macOS pick, but weaker on Linux: it reaches the real host, yet every system package it installs (python, pipx, git via apt/dnf/…) stays behind, and it can only ever test the one distro the host runs. The whole Linux-specific value is the cross-distro matrix. |
| Disposable KVM VM (this harness) | A genuine kernel + userland + package-manager boundary that wipes to nothing on teardown, and swaps freely between distro cloud images. The one real cost — nested virtualization for the backend — doesn't apply, because we gate the installer, not the runtime (below). |
Two variants: test (bare host) and test-ready (prepared host)
install.sh never installs a backend, and never installs its own toolchain
prerequisites (python3, git, pipx) — it installs the bot-bottle package and
runs doctor, which reports what's missing
(install.sh header,
bot_bottle/cli/commands/doctor.py).
That leaves two distinct things worth testing, split into two subcommands that
mirror the macOS harness's test / test-ready convention (there the split is
the backend service; here it is the toolchain the installer needs):
test—install.shruns on the bare cloud image, prerequisites and all left as the vendor ships them. This exercisesinstall.sh's own prerequisite-guard logic — the entire first half of the script (python version gate, git-for-git-specs gate, pipx/pip PEP-668 handling).test-ready— the harness installs python3 + git + pipx first (theprereqsstep), then runsinstall.sh. This is the prepared-host happy path: does a clean install actually land and produce a working CLI?
Pass criteria differ by variant:
| Variant | PASS when |
|---|---|
test |
install.sh either installs cleanly (the image already carried enough) or declines with one of its own recognized, actionable prerequisite errors (missing python3/git, no usable pip, PEP 668). A crash or an unrecognized failure is a FAIL. |
test-ready |
install.sh actually lands: the bot-bottle entry point is present and runs, and doctor reports a usable python and config without crashing. A graceful decline is no longer good enough. |
Neither variant requires a green doctor: inside the VM there is no nested KVM
or Docker, so the backend is correctly reported not-ready — install.sh does
not install a backend and cannot regress one, and this harness does not
provision the Docker backend. This is where Linux necessarily diverges from the
macOS test-ready, which reaches the host backend; BB_TEST_REQUIRE_BACKEND=1
makes readiness fatal anyway, for a nested-virt host that can satisfy it. The
verdict instead classifies doctor's output the way the macOS harness does — a
Traceback is an install defect (fail), a missing python/config line is a
fail, a not-ready backend is reported — so a genuine installer regression (a
broken shim, an import error, a botched PATH) stays visible.
test-all runs the full matrix — every distro × both variants — each cell in
its own throwaway VM, and prints a per-cell PASS/FAIL summary.
The distro matrix is the point
Each distro exercises a different corner of the installer:
| Distro | Cloud image | What it stresses |
|---|---|---|
| Ubuntu (noble) | cloud-images.ubuntu.com |
The common case; apt's pipx, externally-managed Python (PEP 668) → install.sh's pipx path. |
| Fedora | Fedora Cloud Base Generic | dnf packaging, a different default Python, BSD-style checksum file. |
| Arch | geo.mirror.pkgbuild.com/images/latest |
Rolling / newest Python; python-pipx. |
| Alpine | Alpine "cloud" (cloudinit) image | musl libc + BusyBox sh — the harshest POSIX-sh host for a #!/bin/sh installer. |
| NixOS | channels.nixos.org OpenStack image |
No FHS ~/.local on PATH by default; nix-env user-profile prereqs; pipx laying a self-contained venv on a non-FHS host. |
In test-ready the harness installs python3 + git + pipx first on each
distro (install.sh installs none of them), so all five drive the recommended
pipx path. test then removes that scaffolding and lets each distro's bare
image collide with install.sh's guards — on most cloud images python3 is
present (cloud-init needs it) but git and pipx are not, so install.sh is
expected to decline at the git-for-git-specs gate or the PEP-668 pip check with
an actionable message. Both are legitimate, and the two variants together cover
the whole first half of the installer as well as the happy path.
What a clean install touches (the footprint that decides "wipeable")
| Artifact | Location | In the guest's $HOME? |
Survives VM teardown? |
|---|---|---|---|
| Config / state / db | ~/.bot-bottle/{agents,bottles,contrib,…} (install.sh) |
✅ | ❌ overlay deleted |
| pipx venv + shim | ~/.local/pipx/venvs/bot-bottle, shim in ~/.local/bin |
✅ | ❌ overlay deleted |
pip --user fallback |
~/.local/lib + ~/.local/bin |
✅ | ❌ overlay deleted |
Distro prerequisites (python/git/pipx, test-ready only) |
system paths via apt/dnf/pacman/apk/nix-env |
❌ | ❌ overlay deleted |
Unlike the macOS throwaway user (whose Homebrew / Apple-Container / Rosetta footprint survives), every row here dies with the overlay — that is the VM's whole advantage. The cached base image is read-only backing and is the only thing that persists between runs, on purpose.
Design notes baked into the harness
- User-mode networking (
-netdev user,hostfwd=tcp:127.0.0.1:PORT-:22): no root, no bridge, no host network state touched. Only SSH is forwarded. - cloud-init seed ISO (
cloud-localds) injects an ephemeral SSH keypair and a passwordless-sudo login. The keypair is generated per run and deleted on teardown; the guest can't be logged into after it's gone. - Copy-on-write overlay: the cached base is never mutated, so a corrupt or
interrupted run can't poison the cache; downloads land at
*.partialand are renamed only after checksum verification. - Checksums: verified against each vendor's published sums file at
download time (GNU
hash file, bare-hash, and Fedora's BSDSHA256 (file) = hashformats are all handled). The NixOS channel image ships no stable sums file, so that one needsBB_TEST_SKIP_VERIFY=1. test-allruns every distro × both variants (testandtest-ready), each cell in its own subshell on its own forwarded port, so one cell's failure (or teardown trap) can't abort the matrix; it prints a per-cell PASS/FAIL summary.
Not wired into PR CI
Like the macOS harness, the runtime is host-specific (needs /dev/kvm,
qemu, and cloud-localds) and is not exercised by the Linux pull-request
runner. It is validated statically (bash -n, shellcheck) and run by hand
on a KVM-capable host. The cloud-image URLs in the DISTRO table are the one
place to bump when a distro cuts a newer build.