Repository navigation
Conversation
The target install is the same ~940 packages for every machine; only the kernel, CPU microcode, audio firmware and Tailscale differ. pacman extracts that set single-threaded at ~40 packages/s, about 34s on any machine with four or more cores, and no amount of parallelism reaches it. So build-root-image.sh pacstraps the invariant set once at ISO build time into a btrfs subvolume mounted compress=zstd:3 (the level the installer mounts the target with) and ships it as a `btrfs send --compressed-data` stream. The orchestrator receives it at the target filesystem's top level right after archinstall mounts the layout, snapshots it writable in place of the empty @ subvolume, replays the mount table, and then lets archinstall finish with the per-machine delta (install_base_delta mirrors minimal_installation minus the bulk pacstrap), users and fstab. The application installers strap only what the target lacks, since the image carries their package sets and the mirror no longer does. Measured in the same 16-vCPU VM: the package phase drops from 35.7s to 24.5s and the whole install from 41.8s to 30.6s, with an identical set of 942 packages installed; the installed system boots. The receive itself is ~17s and independent of CPU count. The offline mirror keeps only what is still pacstrapped at install time: the live ISO's own packages, the per-machine packages, and the omarchy-other.packages extras omarchy-apply-system may pull in, resolved against the offline repo itself so the keep-set can only name files the mirror holds. The ISO grows from 6.2GB to 9.1GB, the live closure and the extras now sitting beside the 5.2GB image. Build details: the container needs loop device nodes made by hand (Docker fills /dev once, at start), pacman-key's gpg-agent must be stopped before the image unmounts, and stale copies of locally rebuilt omarchy packages are evicted from the shared pacman cache so mkarchiso's pacstrap does not hit a checksum mismatch. The dashboard gets a phase_progress signal from the unpack so the bar moves while the local pacman db is still empty, and the live ISO prefetches the leading bytes of the stream during the wizard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P9b7oP8j2GA6sZZ9e8aJYv
btrfs's incompressibility heuristic skips a lot of data in this tree that zstd compresses fine, and the level only costs build time: btrfs receive stores the extents as they arrive. Against plain compress=zstd:3 the send stream goes from 5.2GB to 3.3GB and the installed root from 5.4GB to 4.0GB, at the same ~17s receive time; level 15 adds under two minutes to the ISO build. btrfs clamps anything above 15. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P9b7oP8j2GA6sZZ9e8aJYv
--free-space installs into the unallocated tail of a synthetic Windows-style disk: a FAT32 ESP and an ext4 data partition with a marker file, then ~76GiB of free space. The fixture is built without root (parted on the raw file, mkfs at the partition offsets), so unlike omarchy-iso-test-windows-disk it carries no EFI/Microsoft directory; the configurator's free-space mechanics are the same either way. The wizard is driven through the mode picker and the free-space confirm, and once the installed system is up the harness checks from inside it that both pre-existing partitions and the marker survived. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P9b7oP8j2GA6sZZ9e8aJYv
mkarchiso pacstraps the live root from the complete mirror at build time, but at install time only packages the root image lacks can ever be downloaded from it. The live root's customize_airootfs.sh removes every package file the image already holds at the same version from its copy of the mirror, keeping the repo db complete so `pacman -S --needed` over the hardware scripts' mixed package lists still resolves every name. The build cache goes back to keeping the whole download closure, so a rebuild downloads nothing; the shipped selection is decided per build from the image's local db. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P9b7oP8j2GA6sZZ9e8aJYv
arch_install_system partitions, formats and encrypts as its first step, and only then did _install_root_image check that the stream exists and that the layout puts the root on a btrfs @ subvolume, with the LVM guard later still. Any of those failing left a wiped disk (encrypted, on that path) with no system on it. All three are predicates on the ISO and the configurator JSON, so run them in prepare_install_target, the phase before anything destructive. The protected path checks the real mounts, which exist already. Existence does not cover a truncated stream, which on a badly flashed USB is the likelier failure: btrfs receive's per-command checksums catch that too, but after the disk is gone. build-iso.sh now writes a sha256 next to the stream and the same pre-flight verifies it, which also warms the page cache for the unpack. Unit tests cover the pre-flight checks and the subvolume swap in _install_root_image, asserted on the subprocess sequence the way create_factory_snapshot already is. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
The root image was pacstrapped from a repo db built over the unpruned cache. That cache persists across builds and can hold several versions of the same package; repo-add keeps whichever file it processes last (warning on downgrade, hidden by -q) and the glob orders by name, so foo-1.9 beats foo-1.10. The mirror db was rebuilt correctly after the prune, leaving an image that could carry an older package than the mirror beside it advertises. The resolve/prune/repo-add block does not depend on the image, so run it first and drop the early repo-add: the image now resolves against exactly the files this build ships. image.packages joins the download list so the pruned mirror always holds the image's own packages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
0cefda3 hardcoded -smp 16 while benchmarking the image receive and never mentioned it; it oversubscribes the VM on anything with fewer cores. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
The resolved count is the image's package count plus the kernel closure, so since the image replaced the full pacstrap it is the one build-time signal that would catch a short root image; a WARNING that ships no denominator and lets the build continue is not enough for that. The shipped-mirror selection right above already exits on a bad count. Also let the zero case of that selection reach its error message: grep -c exits 1 on no match, which under set -e killed the script before the "looks wrong: 0 of N" line could print. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
The hook masking moved whatever was at the hook path to .omarchy-backup and put a /dev/null symlink in its place, with no check for a mask left behind by a run that died before its cleanup. A second run would then move the symlink over the real backup and mask the host's hook for good. Inert for the ISO build (fresh container every time), but the script documents itself as runnable standalone; skip such hooks the way the orchestrator's _is_devnull_symlink does. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
Seed the send stream into mkarchiso's work/iso tree next to airootfs.sfs instead of the live squashfs: mkarchiso packs that directory as is, with the boot records intact, and the live system reads the stream straight off the boot medium at /run/archiso/bootmnt/arch/x86_64/omarchy-root.btrfs. Measured against the squashfs location, same build cache, back-to-back: the ISO is byte-for-byte the same size, mkarchiso is 5s quicker (no 3GB copy into the squashfs), and the install's package phase drops from 18.7s to 16.2s — reading through squashfs costs a copy per 1MiB block even with no decompression. airootfs.sfs shrinks to the live root, and the image can be pulled out of the ISO with any ISO9660 tool. The orchestrator and the wizard-time prefetch look at the ISO path first and fall back to the squashfs path, so mixed old/new pieces still work. Builds before this left the stream in the persistent build cache, where it would ship a second time; the build removes it. The build also logs timestamps around the image step and mkarchiso. The pre-flight checksum follows the stream: build-iso.sh writes the sha256 next to it in the ISO tree, and the orchestrator derives the checksum path from whichever stream location it finds. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
The image was pacstrapped with -K, so it carried /etc/pacman.d/gnupg with a master key (secring.gpg) that every install would share. A signing key must never be distributed; pacstrap the image with -G and remove anything a scriptlet might have seeded regardless, as #108 does. On a target pacstrapped directly, as on quattro, pacstrap -K initialised a per-machine keyring and the keyring packages' scriptlets populated it in the same run. Here those packages come from the image, where their scriptlets ran with no keyring to populate, so the orchestrator does it: after the last pacstrap (each one runs its own pacman-key --init on the target), pacman-key --init, idempotent for the key the delta pacstrap already generated, then --populate archlinux omarchy from the target's own keyring files. Chroot-free via --gpgdir and --populate-from, and the gpg daemons are killed on every path so the target can be unmounted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
pacman-key --init/--populate took the critical path synchronously and left a gpg-agent and dirmngr to be killed by name afterwards. Start it with systemd-run --wait --pipe right after the last pacstrap instead, and join it in create_factory_snapshot: nothing in between reads the keyring (the offline repo is SigLevel = Never) or writes it, so the Limine, user and finalizer phases hide its few seconds, and the snapshot waits so @factory never captures it half-written. A unit rather than a detached child: systemd kills the gpg daemons with the rest of the cgroup the moment pacman-key exits, so no sockets under the target's gnupg dir survive to block the unmount; the dashboard's process-group kill does not reach it while systemctl stop still does, which main() runs on every exit path; and its output lands in the journal whatever happens to the orchestrator. --wait --pipe give a child to join with the unit's exit status and output; --collect releases the name after a failure. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XdQdQ7YZmBdbvaZNQ2xjKN
The pre-flight sha256 of the root image ran inline in the orchestrator's prepare_install_target phase: 1.6s on an NVMe-backed VM, tens of seconds from a USB stick, all of it after the user had pressed Install. The medium sits idle while the user works through the configurator, so move the read there: omarchy-root-image-verify.service runs `sha256sum -c` on the ISO copy of the stream as a oneshot at boot, niced and at idle I/O class, with RemainAfterExit so the verdict persists. prepare_install_target now only collects it: done → go on, failed → the corrupt-medium error with the unit's journal tail, still running → wait with progress read from the hasher's /proc/PID/io, never started → start it and wait. Measured in the install harness, the phase drops from 1.6s to 0.0s. The unit is the only verifier: the inline hashlib loop goes, and with it the squashfs fallback location for the stream, which only existed so a live root and an orchestrator from either side of the move to the plain ISO file could still pair up. Every ISO now ships the stream and its checksum at /run/archiso/bootmnt/arch/x86_64, which both conditions of the unit require; an orchestrator that finds no unit to ask fails the install instead of hashing quietly. The wizard-time prefetch in .automated_script.sh waits for the unit before reading the image: two sequential readers on one USB stick seek against each other, and the unit's pass is the warm-up anyway. Its head read afterwards is a cache hit where the image fit, and re-warms the leading bytes where a small budget let the kernel drop them behind the hash. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DvydT2sYFgBvji2nk9QbDp
A corrupt install medium failed the install with the right error and the
disk untouched, but the dashboard centres each line of the failure in ~80
columns and clipped the one sentence that mattered: "root image stream is
corrupt: omarchy-ro…" — the advice to re-flash never reached the screen,
and four of the five "last log lines" were systemd's exit/failed/consumed
boilerplate, which had pushed sha256sum's own "FAILED" line out of the
tail. Seen on a throttled-cdrom run with one digit of the recorded sha256
flipped.
Lead with the action on its own short line ("install medium is corrupt:
re-flash it"), put the detail on the next, and append only what sha256sum
wrote (journalctl -u <unit> _COMM=sha256sum) instead of the last five
journal lines. The dashboard folds the failed phase's lines at word
boundaries rather than truncating them.
The corrupt-image integration scenario keeps it that way: copy the ISO,
flip one hex digit of the checksum in place (ISO9660 has no per-file
integrity data; reflink makes the copy free), autoinstall from it, and
assert the verify unit failed, the install halted in the pre-flight phase
with the re-flash advice and sha256sum's verdict, nothing later ran, the
target disk still has no partition table, and both the stop screen and the
advice are visible. It boots the ISO itself rather than the installed base
and reaches the live root over SSH through a tty3 console login, now a
base-test.sh helper (bootstrap_live_root_ssh / ssh_live_root).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DvydT2sYFgBvji2nk9QbDp
The scenario flipped a digit of the recorded checksum, which exercises the same failure but is not what a bad medium looks like: on a badly flashed stick the checksum file is fine and the bytes under it are not. Damage the image instead: find the stream on the ISO copy by the NUL-terminated magic every btrfs send stream starts with, and overwrite 16 bytes a third of the way in, deep in extent data. The checksum file is untouched. The fixture no longer re-hashes the damaged stream to check its own work; the assertions on the install's behaviour are the test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DvydT2sYFgBvji2nk9QbDp
The archiso hook defaults to copytoram=auto, which copies the airootfs to RAM and then unmounts /run/archiso/bootmnt when the boot device is not an optical drive, the airootfs is under 4 GiB (ours is 3.18 GB) and MemAvailable exceeds the image size plus 2 GiB. Every USB-booted laptop with 6 GB or more trips it. The installer then fails in prepare_install_target with "root image stream missing", because the root image is deliberately kept out of the airootfs and read straight off the medium. The QEMU integration tests attach the ISO as an IDE CD-ROM, which is the one case the auto rule excludes, so this only showed up on real hardware (ThinkPad X200s, 8 GB; X200, 6 GB). Add copytoram=n to every live-boot entry (BIOS syslinux, GRUB, systemd-boot), guard that with a unit test, and make the missing-stream error say when the medium was released by copytoram so a hand-edited cmdline fails with a useful message. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJj2oLbxWnJn3Cvqdmz1HS
omarchy-root-image-verify.service hashes the multi-GB root image from the boot medium at boot with IOSchedulingClass=idle, while the live system pages its airootfs in lazily from the same medium (copytoram is off). The idle class only means anything under BFQ; the default mq-deadline ignores I/O priority, so on a slow USB stick the hash competes as an equal with every squashfs page-in and the boot crawls. A throttled QEMU boot (usb-storage capped at 33 MB/s, 6 GB RAM, the ISO booted as a real USB stick under SeaBIOS) confirms it: with a buffered sha256sum hog running idle-class, interactive random reads complete at ~50/s under mq-deadline versus ~130/s under BFQ -- roughly 2.5-3x more of the device handed to the live system. The configurator is interactive by ~45s either way while the hash runs to ~120s in the background. Add a udev rule that puts USB disks, SD cards and optical drives on BFQ (internal SATA/NVMe install targets keep their default), a boot-medium helper the verify unit runs as ExecStartPre to log the device and its scheduler next to the verify result, and a unit test over the shipped udev rules. The elevator= kernel parameter cannot do this: it was tied to the legacy single-queue block layer and became a no-op when blk-mq landed in 5.0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJj2oLbxWnJn3Cvqdmz1HS
…helper The full-disk path refuses a corrupt medium before archinstall formats (prepare_install_target runs the verify before phase 3). The free-space path formats in the configurator — parted, wipefs, luksFormat, mkfs — before the orchestrator ever starts, so a corrupt medium there halted only after two partitions had been created and LUKS-formatted with the user's passphrase, with no rollback from the orchestrator's failure path. Found by an automated review of #113; the cidata-based corrupt-image test never reached it because autoinstall skips the configurator. Fold the verdict collection and the boot-medium/scheduler logging into one script, omarchy-wait-root-image-verify (replacing omarchy-iso-boot-medium): it logs the boot device and its scheduler, then collects the boot-time hasher's verdict, waiting for the unit if it is still running and starting it if it never did. The configurator runs it before run_partition_execute on the free-space path; the orchestrator's verify_root_image_stream now shells out to the same script instead of reimplementing the systemd handoff in Python. One source of truth for both disk-touching paths; whoever reaches it first pays the wait. Drops the unit's ExecStartPre (the script logs the medium now) and the Python _systemctl_show/_process_read_bytes/_journal_tail helpers (the per-byte verify progress bar goes with them; the hash is almost always done before either caller reaches the gate). The copytoram-released-medium message moves into the script. New wait-root-image-verify-test.sh drives the gate with stubbed systemctl/findmnt; the Python verify tests now cover the shell-out. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJj2oLbxWnJn3Cvqdmz1HS
The QEMU integration harness only ever booted UEFI (OVMF), so the legacy BIOS install path — archinstall's layout, Limine's MBR install, and booting that MBR — was covered by real hardware alone. Make the harness firmware-aware: OMARCHY_INTEGRATION_FIRMWARE=uefi|bios (or the runner's --bios flag) picks OVMF or QEMU's built-in SeaBIOS. Under BIOS the VMs run with no pflash, base images cache under a separate -bios dir, and the install boots the ISO from an ide-cd (SeaBIOS boots that reliably; it will not boot the isohybrid image off a fallback USB device). The install phase's wait-for-SSH-after-reboot then already proves the MBR boots. Add firmware-boot-test.sh, which boots the installed base and asserts it came up in the expected mode and carries the matching bootloader: under UEFI an EFI runtime, a Limine EFI binary on the ESP, and an efibootmgr entry; under BIOS no EFI runtime, Limine's BIOS stage under /boot, and Limine's boot code in the disk MBR. factory-reset is UEFI-only (shared-ESP dual boot) and skips under BIOS. Validated both ways: a full BIOS install reboots into the installed system and passes all four BIOS assertions; the UEFI assertions pass against the existing UEFI base. (Noted in passing: a BIOS install also drops inert EFI artifacts under /boot/EFI as a UEFI-machine fallback; harmless, not asserted either way.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJj2oLbxWnJn3Cvqdmz1HS
The live-ISO console SSH bootstrap switches to a spare TTY with ctrl-alt-f3 before logging in. On BIOS that first keystroke lands on the ISOLINUX menu, which cancels its auto-boot countdown on any key — so the ISO never booted and the bootstrap timed out waiting for a login prompt (corrupt-image under --bios). Press Enter a few times first to commit the highlighted default (the install medium); it boots ISOLINUX, and once booted the presses are harmless newlines, well before the installer dashboard appears. BIOS-only, so the UEFI path is untouched. corrupt-image now passes 8/8 under --bios. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJj2oLbxWnJn3Cvqdmz1HS
…t equal The header said the image is compressed "at the same zstd level the installer mounts the target with," but IMAGE_COMPRESS is compress-force=zstd:15 while the installed system mounts compress=zstd (level 3) — a deliberately higher level, since it only costs build time and btrfs receive stores the extents as-is. Point the header at IMAGE_COMPRESS and its existing explanation instead of restating a wrong equivalence. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YJj2oLbxWnJn3Cvqdmz1HS
mkarchiso ships configs/grub/loopback.cfg at /boot/grub/loopback.cfg on the ISO (_make_common_bootmode_grub_cfg copies every profile grub/*.cfg into isofs), and that is the file Ventoy and a hand-written GRUB entry use to boot the image as a file on a disk. archiso_loop_mnt sets archisodevice to the loop device it creates, so the auto rule's one exclusion -- an image on /dev/sr* -- never applies there, and copytoram=auto turns on for exactly the reasons it did on the ThinkPads. /run/archiso/bootmnt is then unmounted and the install aborts at the root-image gate. boot-cmdline-test.sh could not catch it: loopback.cfg was not in its file list, and it finds the medium by img_dev/img_loop rather than archisosearchuuid, so match on archisobasedir instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Type=oneshot units are exempt from systemd's default start timeout, so a stick that stalls reads instead of erroring would hang the boot-time hash -- and the install waiting on it -- forever, with nothing on screen. build-iso.sh now writes a drop-in next to the unit sized to the image it just built: a 2 MiB/s floor over the stream size plus ten minutes of slack. On timeout systemd sets Result=timeout, and the wait helper turns that into 'install medium is too slow: try another USB stick or port' instead of the corrupt-medium re-flash advice, which would not help. The 2 MiB/s floor is a first cut; old sticks on the X200's USB 2.0 ports will calibrate it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015y69myH5K7wPRNt6p21GGp
The NBD and NFS hooks force copytoram=y unless the cmdline says exactly n, but both keep the server's image tree mounted, so pinned they can still stream the root image and LAN installs keep working. HTTP cannot be fixed by pinning: its hook downloads only the airootfs (plus optional checksums) into a tmpfs and never mounts an image tree, so the root image does not exist on that path and every install from it would die at the pre-flight gate. The entry is removed, and the boot-cmdline guard now covers the PXE file and refuses any reintroduced archiso_http_srv entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015y69myH5K7wPRNt6p21GGp
The refactor to the boot-time verifier dropped the byte progress the old in-phase hash reported, leaving 'Preparing install target' on the dashboard's time-driven band. While the unit is still activating, the orchestrator now mirrors the hasher's read offset (the unit's MainPID is sha256sum; /proc fdinfo pos says how far it has read) into phase_progress, then collects the verdict through the helper as before. Best effort throughout: a missed sample can never fail an install. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015y69myH5K7wPRNt6p21GGp
corrupt-image-test.sh was the one file the firmware guard missed: it copied OVMF_VARS_TEMPLATE unconditionally, keeping an edk2 dependency on a --bios run even though start_vm ignores it there. Gated on uefi like base-test.sh. The BFQ rule's comment now says what the rule actually matches -- every USB whole disk (install targets and cidata drives included), optical drives, and every mmcblk node down to eMMC boot0/rpmb -- and why that wider-than-boot-medium scope is fine: the consequence is only the scheduler choice. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015y69myH5K7wPRNt6p21GGp
The free-space configurator gate swallowed the wait helper's output on purpose (stdout carries the scheduler line, stderr the failure message), so a wait on a slow medium showed a frozen 'Verifying the install medium' header with nothing moving -- the one place the boot-time hash was still invisible. The helper now draws a \r-updating percent line when OMARCHY_VERIFY_PROGRESS names a sink (the configurator passes /dev/tty), computed from the hasher's fdinfo read position -- the same source the orchestrator mirrors into the dashboard, so no pv dependency and sha256sum -c stays exactly as it is. Success finishes the line at 100%; a failure, which includes a timed-out unfinished read, just ends it. Default behavior without the variable is byte-identical. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015y69myH5K7wPRNt6p21GGp
The full-disk path clears the install-medium verify inside the orchestrator's 'Preparing install target' phase, whose dashboard band is 15 per-mille wide -- sized for the subsecond case where the boot-time hash already finished. When install is pressed while the hash is still running, minutes of real progress move the bar about one cell and the UI reads as hung; the free-space path got a live percent line in f63e0e8, but the dashboard never did. The dashboard now maps the phase's progress into an honest percent and prints 'verifying the install medium: N%' on the row between the bar and the tip line -- blank in every other state, so the layout height never changes. The unpack band is wide enough that its progress stays in the bar itself. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016C4YbdRMV55n8FxMxwu8Bn
Brings in #118 (UTF-8-safe command capture), #119 (release checksums), and #120 (install-media diagnosis on pacstrap rejection). One semantic reconciliation: the new command-capture test builds a stub ctx for create_factory_snapshot, which on this branch first joins the target-keyring unit via ctx.state -- the stub now carries state={} so the join is a no-op and the test reaches the mangled-device check it is about. ./test/all green (88 Python tests + all shell suites). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016C4YbdRMV55n8FxMxwu8Bn
The free-space install formats inside the configurator, so the gate added in 30596b0 is the only thing standing between a corrupt install medium and parted, wipefs, luksFormat and mkfs on a disk that already holds someone else's OS. Nothing tests that it is still there: corrupt-image-test.sh autoinstalls from cidata, which skips the configurator entirely, and wait-root-image-verify-test.sh exercises the helper rather than its call site. Replacing the whole gate with `if false` leaves ./test/all green -- which is how the gate came to be missing in the first place. A static ordering check costs nothing and catches that: the verify call has to exist inside the free_space branch and precede run_partition_execute. Verified by mutation -- removing the gate and moving the format above it each fail the new test, and both passed without it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…passed scenarios drop their disk overlays corrupt-image looked for omarchy-root.btrfs.zst; a block-copy ISO ships omarchy-root.img.zst, so xorriso found nothing and the scenario exited without an assertion on every run. It resolves the name from the ISO and removes its ISO copy on exit (7 GB wherever cp cannot reflink). With the name resolved the scenario makes its 8 assertions and passes on the encrypted ISO. finish() deletes a passed scenario's qcow2 overlays, which are reproducible from the base image and are what fills a ramdisk run directory (OMARCHY_INTEGRATION_KEEP_DISKS=1 keeps them).
…stall, no date forks Measured in a traced install: - snapper create-config took 2.86 s closing every fd up to RLIMIT_NOFILE in each of four forks, and takes 0.04 s with the limit at 65536. prlimit --nofile=65536 wraps the finalizer chroots. - archinstall's arch-chroot -S wraps each command in a transient systemd unit: a chpasswd took 0.15 s for 5 ms of work. Installer._chroot_argv and Installer.run_command, its string-form twin, use plain arch-chroot. - The image's install log helper forked date 131 times, about 0.5 s; an image-build sed makes bash format the timestamps.
…s for the console to answer a size query About 1 boot in 10 of an installed system stops at "Starting Switch Root" until a key is pressed: systemd 257+ queries the console size with ANSI sequences at PID 1 start, and with Plymouth holding the VT the reply sometimes never arrives. Measured on the pre-built UKI (3/30, 1/40) and on the stock busybox initramfs (7/40); every hang resumed on Enter; 0/40 with a serial console, 0/40 with plymouth.enable=0, 0/80 with this parameter. Users reported it on real hardware as a black screen after the passphrase (omacom/omarchy#2931, #2665, discussions/6267). systemd fixed two variants in v261 and Arch's 261.2 still shows it. The parameter goes into the Limine entry (both initramfs flavours) and the UKI's embedded cmdline (the Secure Boot path); nothing changes visibly under quiet splash.
…he archive already has every locale DELTA.locale_and_keyboard took 0.77 s in a traced install: locale-gen rebuilds the whole locale archive for the one locale the installer uncomments. The image ships the archive with en_US.UTF-8, and _run_command skips locale-gen when localedef --list-archive already covers every uncommented locale.gen entry.
…d as "did not run" The wait helper polls the verify unit's ActiveState every half second and treats anything outside active/failed/activating/deactivating as "the unit did not run". An empty string lands there too, and an empty string is not a unit state: it is systemctl failing, which happens when PID 1 does not answer within D-Bus's 25 s, and PID 1 can stall on the same dying medium while it stops the hasher. A run of the slow-medium scenario (a medium choked to 64 KB/s while the hash is mid-read) ended with "did not run" 54 s into the gate; the same ISO otherwise passes the scenario, with the stop landing 17 s after the timeout. The helper retries an empty answer through up to a minute of silence, waits through reloading, refreshing and maintenance, and names the state in the "did not run" message.
…on a real NVMe) An install on real hardware (12600K, Samsung 980 PRO, LUKS partition) spent 86.6 s of its 92.7 s in `zstdcat | dd bs=64M oflag=direct`. The same command on the same partition, with the same luksFormat and open flags, takes 91.1 s cold and 86.8 s with the stream in the page cache. dd read 64 KiB per read() from the pipe (102,667 partial records for 6.26 GiB), and for a partial block coreutils dd drops O_DIRECT, so every write was buffered. Traced with bpftrace: dd's write() returns in 16-32 us with buffer-head allocations in the stack, and the NVMe receives 2,038,376 write requests per GiB of 512-1024 bytes, because writeback to a block device works in units of its logical block size and the LUKS mapper has 512-byte sectors. A buffered dd (no oflag=direct) measures the same 89 s. QEMU hides all of it: 4.9 s there. With iflag=fullblock the same trace shows 252 writes of 4 MiB per GiB and NVMe requests of 128-256 KiB. On the same partition, stream in RAM: bs=1M 2.86 s, 2M 2.68 s, 4M 3.05 s, 16M 3.97 s; cold from a 430 MB/s stick 9.3 s, which is the stick. None of these is the limit: dm-crypt workqueue flags (all three settings within noise), 4096-byte LUKS sectors (2.88 s), encryption at all (2.98 s with no LUKS; AES-XTS runs at 7.9 GB/s per thread), a relay dd (3.5 s) or an overlapped reader/writer (no gain). The drive itself takes 2.6-3.2 GB/s. The write logs dd's record counts and warns when more than one block was partial, since no VM run can show this regression.
An encrypted install opened its volume three times and built its initramfs on the target. Measured on a 12600K, each open is a 2.2 s key derivation at the configurator's 2 s iter-time, and limine-mkinitcpio takes 3.8 s. - Open the volume once and keep it open. archinstall closed and reopened it between formatting, creating the subvolumes and mounting the layout. OMARCHY_LUKS_KEEP_OPEN=0 restores that. - Open with --allow-discards and --perf-no_read_workqueue, written to the LUKS2 header with --persistent so every later open applies them. Discards are new for whole-disk installs (archinstall's cryptdevice= has no options): TRIM reaches the SSD, and which blocks are unused becomes visible on the raw device, as cryptsetup-open(8) warns. no_write_workqueue is left out: in the install VM it slowed the image write from 5.0 s to 6.7-7.5 s. - Use the pre-built UKI for encrypted installs too. The UKI is built from its own config with the systemd hooks plus sd-encrypt and plymouth. - Put rd.luks.name= and rd.luks.options= on the cmdline next to cryptdevice=. Limine passes the entry's cmdline to the UKI, and the systemd initramfs ignores cryptdevice= and would wait for the mapper forever. The busybox initramfs a later rebuild produces ignores rd.luks.*, so one cmdline serves both. - test/integration: wait_for_ssh answers systemd-cryptsetup's prompt as well as the busybox hook's, so scenarios can boot an encrypted install that uses the pre-built UKI. The key derivation is unchanged: archinstall formats with argon2id at the configurator's iter-time, benchmarked on the machine.
… the whole mirror download arch-mact2 replaced apple-bcm-firmware with apple-bcm-firmware-fetcher (conflicts with the old name); the published omarchy package list still names the old one, so pacman -Syw fails with 'target not found' on every fresh build. A warm builder cache hides it. build-iso syncs the databases first, drops the names no repository offers with a warning, and lists them in unresolved-packages.txt on the ISO.
…ory; failed scenarios can discard their disks Two knobs for a host that keeps run directories on a tmpfs sized for base images and overlays. OMARCHY_INTEGRATION_SCRATCH_DIR moves corrupt-image's 7 GB ISO copy (read once at boot, once by the hasher) off it. OMARCHY_INTEGRATION_DISCARD_DISKS=1 deletes a failed scenario's overlays too: a CI host never inspects them, and factory-reset, whose three shared-ESP assertions are known failures, otherwise parks a full re-install's overlay for every scenario after it. On an 18 GB tmpfs the base image, that overlay and the ISO copy fill it exactly, the next screendump comes back empty and corrupt-image's two OCR assertions fail with the right text on screen.
…staller greeter getty@.service is Type=idle and waits, up to 5 s, for every active job to be dispatched before running agetty. The live ISO's boot transaction never drains: systemd-time-wait-sync.service waits for a clock sync that an offline installer never gets, with other start jobs queued behind it. So the autologin on tty1, and with it the installer's first screen, paid the full 5 s cap on every boot. The journal shows it directly: "Started Getty on tty1" at 6.9 s, the root session opened at 12.1 s, nothing in between; an agetty and login straced as a Type=simple transient unit on the same tty finish in under a second, and once the pending jobs drain a getty restart takes 20 ms. Measured by booting the ISO's own kernel and initramfs under KVM from a USB stick, with the drop-in injected as a systemd credential, two boots each: greeter at 15.0 and 14.9 s stock, 9.1 and 9.6 s with Type=simple. Plymouth is not involved: a boot without "splash" takes the same 14.6 to 15.1 s.
Speed up encrypted installs with a single unlock and the prebuilt UKI
Speed up the finalize phase of root-image installs
Fix the boot hang at Starting Switch Root and test hibernation
…ine is written The pre-built UKI unlocks the root with the systemd hooks, which read rd.luks.name= and ignore cryptdevice=. with_cmdline_options() adds those parameters, but it was called from the archinstall-derived path only. The pre-mounted path (an install into free space, where the configurator partitions and formats) builds its cmdline in _build_pre_mounted_cmdline and went straight to _write_limine_defaults without them. On a real-hardware free-space install the Limine entry had cryptdevice=UUID=<the right uuid> and root=/dev/mapper/omarchy_root but no rd.luks.name=, so the initramfs made no cryptsetup job, never asked for the passphrase, and printed "A start job is running for /dev/mapper/omarchy_root (... / no limit)" forever. Whole-disk installs, the only kind a VM scenario performs, were fine. The call moves into _write_limine_defaults, the one writer every variant passes through (it adds the console parameter for the same reason), and the writer refuses an encrypted cmdline that still has no rd.luks.name=. with_cmdline_options() is idempotent, since the archinstall-derived cmdline reaches the writer having been through it, and for a UUID= spec it falls back to the UUID itself when the by-uuid link is not there yet or blkid returns nothing, so a lookup hiccup cannot silently drop the unlock. test/unit/test_encrypted_cmdline.py checks both, because the free-space path partitions interactively and no integration scenario can reach it. test/unit/archinstall_fakes.py gives tests a stand-in archinstall, so they import the orchestrator modules they test.
Arch's 35-systemd-update pacman hook touches /usr after every transaction, which arms ldconfig.service, systemd-hwdb-update, systemd-journal-catalog-update and systemd-sysusers for the next boot. pacman has already done all four at transaction time (it runs ldconfig itself; the 20-/25-systemd-* hooks run sysusers, hwdb and the catalog), so the first boot only repeats the work. On a hardware install (12600K, 980 PRO) "Rebuild Dynamic Linker Cache" alone was 1.32 s of a 4.54 s userspace, and ld.so.cache was rewritten at first boot four minutes after the install had written it. The installer runs systemd-update-done in the target right before the factory snapshot, after the per-machine packages (the last transaction, which re-arms the condition), so the snapshot a factory reset restores boots as fast as the install. The harness records systemd-analyze from the one first boot it sees and fails the install if any of the four units ran.
…times, no core allowed On a real-hardware install (12600K, UHD 770) Chromium died with SIGSEGV 2.2 s after its first launch, before drawing a window. It had been started the way the session opens any link: a transient user unit running `uwsm-app -- /usr/bin/chromium <url>`, on a profile that had never been used, 33 s into the first boot. 24 launches of exactly that shape in VMs, on both the default and the encrypted variant, were all clean, and the package versions match a system on the same hardware that has never recorded a Chromium crash, so this does not reproduce the crash. It guards the path: the scenario already opened Chromium once on an existing profile for the clipboard check; it also opens a link through uwsm-app on a fresh profile three times (OMARCHY_BROWSER_LAUNCHES) and fails if a window is missing or coredumpctl records a new chromium core.
…ce TRIM, and the filesystem step logs its pieces archinstall partitions, formats LUKS, makes a btrfs with subvolumes and mounts it; the root image is then written over that filesystem. The step was one line in the log: 2.0 s in a VM, 4.4 s on a Dell XPS 16 (2026) and 22.2 s on a ThinkPad E14, and neither hardware log could say which piece cost what. mkfs.btrfs TRIMs the entire device before formatting unless given -K. On a virtual disk that is instant; through dm-crypt on a 1 TB laptop SSD it is seconds, and far more on a DRAM-less drive, which fits all three numbers. For a filesystem that lives two seconds the TRIM only delays the install, so when a root image is present mkfs.btrfs gets -K through archinstall's own extra-options parameter. Nothing else in the step changes; the installed system keeps fstrim.timer enabled. A filesystem that stays (no root image) keeps its TRIM. Each piece logs a [step] line: FS.partition, FS.mkfs.<type>, FS.luksFormat, FS.luks_open, FS.create_btrfs_subvolumes, and the total of the udevadm settle calls. The wrappers are best effort (an archinstall that moved a name runs the step exactly as before) and put every patched name back, also when the step raises; test_filesystem_step_tweaks.py checks the flag, the timing and the restoration against fakes, on 3.12 through 3.14.
The cases compare the configurator's layout list against the system's, so they need a systemd-booted Arch: localectl reads /usr/share/kbd/keymaps and refuses to run when PID 1 is not systemd. On the live ISO and on an Arch host they run as before. In a container or on an Ubuntu CI runner they raised RuntimeError and failed the whole discovery; they are skipped there, so test/unit runs on any host and a real failure stands out.
Assertions that grep the installer dashboard (corrupt-image's "installation stopped" and "re-flash", slow-medium's advice) failed on CI although the text was on the screen. The kernel blanks a virtual terminal after ten idle minutes, and a scenario that waits for an installer to stop takes that long on a busy runner, so the screendump comes back black and OCR reads nothing. Locally the same scenarios finish before the console blanks. ocr_screen retries once after a Shift press, which unblanks the VT and which no dashboard reads. Where the console is not blank the retry never fires.
A single frame is not a reliable sample of the dashboard: a console that
blanked needs the unblank to land, a repaint can be caught half-done, and
OCR is not perfectly repeatable. corrupt-image's "re-flash advice" check
failed on CI on a screen that showed the text.
screen_shows in corrupt-image and slow-medium polls for up to 20 s, and its
pattern is an extended regex so a hyphenated word can be matched loosely
("re.?flash"). A genuine failure still fails, 20 s later.
archiso has no autodetect on purpose, because one image must boot any machine, so the kms hook there makes mkinitcpio add every DRM driver and every DRM driver's firmware. With it, initramfs-linux-t2.img is 257,184,084 bytes, 229,529,384 of it an uncompressed early cpio holding 152,755,560 bytes of usr/lib/firmware, and 105,251,014 of that is the same four nouveau GSP blobs the pre-built UKI carries. The ISO shipped them twice, and a third time inside omarchy-root.img.zst. kms is not what makes graphics work on the live ISO. This initramfs only has to find and mount airootfs.sfs; every GPU driver loads from the live root afterwards, as it does on an installed system. The hook only moves them earlier so Plymouth can splash before the switch. simpledrm, drm and drm_kms_helper are in the kernel's modules.builtin, so a UEFI machine keeps a DRM device on the EFI framebuffer with nothing in the initramfs, and a BIOS boot keeps vgacon. The greeter is a TUI on either. The bytes are latency as well as size: GRUB reads the whole file off the medium before the kernel starts. 257 MB is 0.6 s on a 430 MB/s stick and 6.6 s on a USB 2 port measured at 39 MB/s.
pacstrap leaves /etc newer than the ld.so.cache, journal catalog and sysusers stamp it just wrote, so ConditionNeedsUpdate= fires on every live boot and systemd re-runs ldconfig.service, systemd-journal-catalog-update.service and systemd-sysusers.service against a read-only squashfs that cannot have changed since it was built. Measured on the live ISO: "Rebuild Dynamic Linker Cache" takes 403 ms, plus the catalog and sysusers runs. customize_airootfs.sh runs systemd-update-done, as the installer does for the installed system. The call sits above the script's early exit, so it runs whether or not the mirror prune does, and uses the absolute path: systemd ships the binary under /usr/lib/systemd, outside PATH. If it is missing the build warns and continues, since a live boot that redoes the work is slower but not broken.
cloud-init's final stage prints each SSH host key's fingerprint and randomart to the console, about 50 lines. On a choked medium that stage runs late (216 s of uptime in the slow-medium scenario), after the installer has stopped with "install medium is too slow: try another USB stick or port", and the advice scrolls off the screen. slow-medium's "dashboard shows the slow-medium advice" check failed on it every run, once the installer reached its gate before cloud-init finished. cloud.cfg.d turns off the two console modules, and a dash-prefix drop-in keeps every cloud-*.service in the journal instead of journal+console. test/integration.d: check() captures the screen when an assertion fails. A failed screen assertion left nothing to look at, because the last capture predated it.
Fix the slow image write and boot failures found on real hardware
Harden screen-reading tests against blank and mid-repaint frames
Shrink the live ISO initramfs to 91 MiB and skip repeated boot work
|
Have you considered basing the changes on #132 ? |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The landing branch for the F1 install work. (making the installer feel closer to F1 pit stop 🏎️)
It starts as @hegjon's #113 brought up to today's quattro, with his 69 commits intact.
What
In @hegjon's base, the installer unpacks a prebuilt root image and adds this machine's packages on top, instead of pacstrapping ~940 packages.
The rest of the series is opened as small pull requests against this branch, one topic each, and merged here in order.
Next
This PR goes into quattro once all subsequent PRs are all in.
All credit for the root-image design and the work in #113 goes to @hegjon.
Merging this with a merge commit also merges #113.