Files
bot-bottle/docs/research/testing-clean-install-on-linux.md
T
didericis 1518f73de5 feat: add bare-host install variant to the Linux harness
Splits the Linux clean-install harness into two variants, selected with
BB_TEST_PREREQS:

  with    (default) — install python3 + git + pipx first, then install.sh.
                      The ready-host happy path; PASS = install.sh exits 0 and
                      leaves a runnable bot-bottle entry point.
  without           — run install.sh on the BARE cloud image. Exercises
                      install.sh's own prerequisite-guard logic (python gate,
                      git-for-git-specs gate, pipx/pip PEP-668 handling). PASS =
                      install.sh either fully succeeds OR declines with one of
                      its own recognized, actionable prerequisite errors; a
                      crash or unrecognized failure is a FAIL.

cmd_run now captures install.sh's exit code + full output (install.rc /
install.log) instead of aborting on non-zero, so the verdict step applies the
variant's criterion. The old assert_installed is replaced by a quiet
entry_point_runnable helper plus classify_outcome, which matches the bare-host
declines against the exact die() messages install.sh prints. doctor still runs
for visibility, but only when an entry point exists.

test-all now runs the full matrix -- every distro x both variants -- each cell
in its own throwaway VM, and can be pinned to one variant via BB_TEST_PREREQS.
Research note updated to document both variants and their pass criteria.

Validated with `bash -n` and `shellcheck`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-27 12:24:31 -04:00

8.3 KiB
Raw Blame History

Testing a clean bot-bottle install on Linux

How do you exercise install.sh the way a brand-new user would — on a pristine Linux environment you can throw away afterward — without polluting your daily-driver host, and across the several package-management regimes Linux fragments into? This is the Linux counterpart to testing-clean-install-on-macos.md; the conclusion is different because Linux gives us a boundary macOS doesn't.

Summary

On macOS the honest options were a throwaway user or a VM, and the throwaway user won on pragmatics (nested virtualization is gated to M3+). On Linux the calculus flips: a disposable KVM virtual machine, booted from a distro cloud image and deleted per run, is both the cleanest boundary and the one that lets a single harness cover Ubuntu, Fedora, Arch, Alpine, and NixOS. The host already requires KVM for the Firecracker backend, so the VM is cheap here.

The harness lives at scripts/linux-install-test.sh. Per run it caches one read-only base image, boots a throwaway copy-on-write overlay (qemu-img create -b base), installs the distro's prerequisites, pipes this checkout's install.sh into the guest exactly as curl … | sh would, asserts the CLI installed, and deletes the overlay — the Linux equivalent of docker run --rm, for a whole machine.

Why a VM, not a container or a throwaway user

Mechanism Why it's the wrong boundary here
Container (docker run --rm) Shares the host kernel and ships a deliberately minimal userland — no systemd, a stubbed-out package manager story, and (crucially) it doesn't reproduce the externally-managed Python (PEP 668) that real desktop/server installs put in front of the user. It tests "does install.sh run in a container," not "does it run on a real distro."
Throwaway user (useradd/userdel) The macOS pick, but weaker on Linux: it reaches the real host, yet every system package it installs (python, pipx, git via apt/dnf/…) stays behind, and it can only ever test the one distro the host runs. The whole Linux-specific value is the cross-distro matrix.
Disposable KVM VM (this harness) A genuine kernel + userland + package-manager boundary that wipes to nothing on teardown, and swaps freely between distro cloud images. The one real cost — nested virtualization for the backend — doesn't apply, because we gate the installer, not the runtime (below).

Two variants: a ready host and a bare host

install.sh never installs a backend, and never installs its own toolchain prerequisites (python3, git, pipx) — it installs the bot-bottle package and runs doctor, which reports what's missing (install.sh header, bot_bottle/cli/commands/doctor.py). That leaves two distinct things worth testing, selected with BB_TEST_PREREQS:

  • with (default) — the harness installs python3 + git + pipx first, then runs install.sh. This is the ready-host happy path: does a clean install actually land and produce a working CLI?
  • withoutinstall.sh runs on the bare cloud image, prerequisites and all left as the vendor ships them. This exercises install.sh's own prerequisite-guard logic — the entire first half of the script (python version gate, git-for-git-specs gate, pipx/pip PEP-668 handling).

Pass criteria differ by variant:

Variant PASS when
with install.sh exits 0 and a bot-bottle entry point is present and bot-bottle --version runs.
without install.sh either fully succeeds (the image already carried enough) or declines with one of its own recognized, actionable prerequisite errors (missing python3/git, no usable pip, PEP 668). A crash or an unrecognized failure is a FAIL.

Neither variant requires a green doctor: inside the VM there is no nested KVM or Docker, so every backend is correctly reported not-ready and doctor exits non-zero by design — and install.sh swallows that (its trailing if doctor; then … else … fi leaves the script's exit code at 0). doctor's full output is still printed whenever an entry point exists, so a genuine installer regression (a broken shim, an import error, a botched PATH) stays visible.

test-all runs the full matrix — every distro × both variants — each cell in its own throwaway VM, and prints a per-cell PASS/FAIL summary.

The distro matrix is the point

Each distro exercises a different corner of the installer:

Distro Cloud image What it stresses
Ubuntu (noble) cloud-images.ubuntu.com The common case; apt's pipx, externally-managed Python (PEP 668) → install.sh's pipx path.
Fedora Fedora Cloud Base Generic dnf packaging, a different default Python, BSD-style checksum file.
Arch geo.mirror.pkgbuild.com/images/latest Rolling / newest Python; python-pipx.
Alpine Alpine "cloud" (cloudinit) image musl libc + BusyBox sh — the harshest POSIX-sh host for a #!/bin/sh installer.
NixOS channels.nixos.org OpenStack image No FHS ~/.local on PATH by default; nix-env user-profile prereqs; pipx laying a self-contained venv on a non-FHS host.

In the with variant the harness installs python3 + git + pipx first on each distro (install.sh installs none of them), so all five drive the recommended pipx path. The without variant then removes that scaffolding and lets each distro's bare image collide with install.sh's guards — on most cloud images python3 is present (cloud-init needs it) but git and pipx are not, so install.sh is expected to decline at the git-for-git-specs gate or the PEP-668 pip check with an actionable message. Both are legitimate, and the two variants together cover the whole first half of the installer as well as the happy path.

What a clean install touches (the footprint that decides "wipeable")

Artifact Location In the guest's $HOME? Survives VM teardown?
Config / state / db ~/.bot-bottle/{agents,bottles,contrib,…} (install.sh) overlay deleted
pipx venv + shim ~/.local/pipx/venvs/bot-bottle, shim in ~/.local/bin overlay deleted
pip --user fallback ~/.local/lib + ~/.local/bin overlay deleted
Distro prerequisites (python/git/pipx) system paths via apt/dnf/pacman/apk/nix-env overlay deleted

Unlike the macOS throwaway user (whose Homebrew / Apple-Container / Rosetta footprint survives), every row here dies with the overlay — that is the VM's whole advantage. The cached base image is read-only backing and is the only thing that persists between runs, on purpose.

Design notes baked into the harness

  • User-mode networking (-netdev user,hostfwd=tcp:127.0.0.1:PORT-:22): no root, no bridge, no host network state touched. Only SSH is forwarded.
  • cloud-init seed ISO (cloud-localds) injects an ephemeral SSH keypair and a passwordless-sudo login. The keypair is generated per run and deleted on teardown; the guest can't be logged into after it's gone.
  • Copy-on-write overlay: the cached base is never mutated, so a corrupt or interrupted run can't poison the cache; downloads land at *.partial and are renamed only after checksum verification.
  • Checksums: verified against each vendor's published sums file at download time (GNU hash file, bare-hash, and Fedora's BSD SHA256 (file) = hash formats are all handled). The NixOS channel image ships no stable sums file, so that one needs BB_TEST_SKIP_VERIFY=1.
  • test-all runs every distro × both variants, each cell in its own subshell on its own forwarded port, so one cell's failure (or teardown trap) can't abort the matrix; it prints a per-cell PASS/FAIL summary. Pin BB_TEST_PREREQS to run just one variant's row.

Not wired into PR CI

Like the macOS harness, the runtime is host-specific (needs /dev/kvm, qemu, and cloud-localds) and is not exercised by the Linux pull-request runner. It is validated statically (bash -n, shellcheck) and run by hand on a KVM-capable host. The cloud-image URLs in the DISTRO table are the one place to bump when a distro cuts a newer build.