Booting via BMC virtual CD reads the ~2.8 GB filesystem squashfs sequentially during the live-boot toram copy; a mid-read drop of the redirected medium loses the whole copy and fails the boot (v14). Split the rootfs into self-contained semantic layers so a retry re-reads at most one ~500-700 MiB layer, not everything. This is a resilience / reduced-re-read mechanism, not a fix for the virtual-media instability. NVIDIA variants now ship 7 layers (00-base, 05-firmware, 08-desktop, 10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs, 40-nvidia-dcgm-cuda) plus an explicit live/filesystem.module that fixes their OverlayFS order; amd/nogpu keep a single squashfs. - lib/squashfs-layers.sh: deterministic classifier (dpkg file ownership plus explicit rules for build.sh-injected files, never a path substring), per-layer mksquashfs, 800 MiB hard ceiling, unsquashfs -s plus strict extraction of every layer, merged-rootfs bootability check. - build.sh: split the monolith after the full lb build, verify and merge, write the module file, delete the monolith only then; abort before ISO assembly on any failure. Runs the builder test suites up front. - fast-path: force a full build for a multi-layer medium; fast_path_repack_squashfs hard-refuses (it would drop layers). - iso-validation.sh: validate_iso_squashfs_layers (module vs layer set match, size ceiling, no lone giant squashfs) and validate_iso_media_integrity (xorriso -check_media). - bee-install: honour filesystem.module order, abort on any layer failure. - 9013-toram-retry: record the real rsync exit code (it printed a false rc=0) and correct the "resumes the tail" comment (rsync without --partial keeps only fully-copied layers). No unsafe partial resume. - tests: test-squashfs-layers.sh plus a multi-layer guard in test-build-libs.sh; both run at the top of every build. - docs: bible-local architecture and decision, iso/README, iso-build-rules. Verified by a full nvidia build: 7 layers 622/199/256/466/37/567/562 MiB, every validator passes, xorriso -check_media good, merged rootfs bootable. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
7.6 KiB
Semantic SquashFS Layers
Module: squashfs-layers (v1.0) - iso/builder/lib/squashfs-layers.sh
Why
The live medium used to ship one filesystem-v<ver>.squashfs of about 2.8 GB.
Booting through a BMC / IPMI virtual CD reads that single file sequentially
during the live-boot toram copy. If the redirected medium drops off the bus
part-way through (observed on v14), the whole copy is lost and the boot fails
with No supported filesystem images found at /live.
Splitting the root filesystem into several self-contained squashfs layers means a mid-copy failure only costs the layer in flight, and the retry loop re-reads far less data.
This is a resilience / reduced-re-read mechanism. It is not a fix for
the underlying BMC virtual-media instability - see
decisions/2026-09-04-squashfs-semantic-layers.md and
decisions/2026-09-04-supported-systems-minimum-16gb-ram.md.
Layers (NVIDIA variants: nvidia, nvidia-legacy)
| Layer | Contents | Ownership |
|---|---|---|
filesystem-v<ver>-00-base.squashfs |
Debian rootfs, kernel + modules, systemd, Bee, networking, CLI diagnostic tools, dpkg database. Boot-critical, read first. | catch-all: any file not claimed by a higher layer |
filesystem-v<ver>-05-firmware.squashfs |
Device firmware for NICs / wifi / non-NVIDIA GPUs | dpkg: firmware-* (but not firmware-nvidia-*, which rides with the driver; not *-microcode, which stays in base) |
filesystem-v<ver>-08-desktop.squashfs |
Local-console GUI: X.org, lightdm, openbox, mesa/llvm, GTK, chromium, mupdf, fonts. SSH / headless never touches this layer. |
dpkg: xserver-xorg*, xfonts-*, x11-*, lightdm*, openbox, chromium*, mupdf*, libllvm*, libgl1-mesa-dri, libglx-mesa0, glx-alternative-mesa, libgtk-3-* / libgtk2.0-* / libgtkmm-* / libgtksourceview-*, libpango-* / libcairo2 / libgdk-pixbuf-*, fonts-*, feh, scrot |
filesystem-v<ver>-10-nvidia-driver.squashfs |
NVIDIA kernel modules (*.ko), driver userspace (libnvidia-*, libcuda.so*), nvidia-smi, GSP firmware, OpenCL ICD, modprobe config, alternatives, bee-nvidia-* units |
dpkg: nvidia-modprobe, nvidia-kernel-common, glx-alternative-nvidia, nvidia-tesla-*, firmware-nvidia-*, libnvidia-*, nvtop, clinfo, ocl-icd-libopencl1. injected: usr/local/lib/nvidia/*.ko, usr/local/bin/nvidia-smi, usr/lib/libcuda.so*, usr/lib/libnvidia-*, usr/lib/firmware/nvidia/*, etc/OpenCL/vendors/nvidia.icd, bee-nvidia.service, bee-nvidia-load/-recover, bee-check-nvswitch |
filesystem-v<ver>-20-nvidia-platform.squashfs |
Fabric Manager, libnvidia-nscq, nvlsm, DCGM core + non-CUDA proprietary DCGM, related units |
dpkg: nvidia-fabricmanager, libnvidia-nscq, nvlsm, libibumad3, datacenter-gpu-manager-4-core, datacenter-gpu-manager-4-proprietary. injected: nvidia-fabricmanager.service.d/* |
filesystem-v<ver>-30-nvidia-cuda-libs.squashfs |
CUDA userspace runtime (cuBLAS / cuBLASLt / cudart), NCCL, nccl-tests, bee GPU stress worker | injected: usr/lib/libnccl.so*, usr/lib/libcublas.so*, usr/lib/libcublasLt.so*, usr/lib/libcudart.so*, all_reduce_perf, bee-gpu-burn-worker, bee-gpu-burn, bee-nccl-gpu-stress, bee-john-gpu-stress |
filesystem-v<ver>-40-nvidia-dcgm-cuda.squashfs |
CUDA-linked DCGM components (dcgmproftester CUDA kernels) - split out from 30 because they are large |
dpkg: datacenter-gpu-manager-4-cuda13, datacenter-gpu-manager-4-proprietary-cuda13. injected: bee-dcgmproftester-staggered |
amd, nogpu: single filesystem-v<ver>.squashfs, no filesystem.module. The
builder leaves the monolith untouched for these variants.
Classification rules
- dpkg file ownership (
var/lib/dpkg/info/*.list), never a path substring. A file is not moved to an NVIDIA layer merely because its path containsnvidia. - Files that
build.shinjects directly from the build cache and the project overlay belong to no.deb; they are classified by the explicitbee_layer_injected_rulestable (extended-regex on the merged-usr-canonical relative path). Injected rules override dpkg ownership. - Pre-merged-usr paths dpkg records (
/bin,/sbin,/lib,/lib64) are canonicalised to/usr/...before matching. - Anything unclaimed falls to
00-base. bin/sbin/lib/lib64are force-kept in00-baseregardless of dpkg records (somefirmware-*.listfiles carry a bare/libentry).- The classifier fails the build if the partition is not complete and disjoint,
if a merged-usr compat symlink is not in base, or if any upper layer would
carry a real top-level
bin/sbin/lib/lib64path. - Output is deterministic: every list is
LC_ALL=Csorted,mksquashfsruns with-processors 1and fixed options. It does not depend onfindorder or locale. - Published on the medium:
live/filesystem.layers.txt(per-layer file counts).
Ordering (must stay consistent everywhere)
Module order: 00-base, 05-firmware, 08-desktop, 10-nvidia-driver,
20-nvidia-platform, 30-nvidia-cuda-libs, 40-nvidia-dcgm-cuda.
live-boot 1:20230131 with MODULE=filesystem (the default) reads
live/filesystem.module verbatim - an ordered list of image names, one per
line - and stacks them so the last listed image has the highest OverlayFS
priority. If the module file were absent it would fall back to a
locale-sensitive lexical sort; the NN- numeric prefixes make that
equivalent, but the module file is authoritative.
The same order is used by:
bee-install: readsfilesystem.module, unpacks each layer withunsquashfs -f(last write wins), aborts install on any layer failure.- The builder fast path: currently disabled for a multi-layer medium; it
forces a full
lb build+ deterministic re-split (needs_full_buildreturns true,fast_path_repack_squashfshard-refuses).
So a higher-numbered layer always wins a conflict, in live-boot, in a disk install, and in a rebuild.
Size budget
Target 500-700 MiB compressed per layer (BEE_LAYER_TARGET_MIB, soft warning).
Hard ceiling 800 MiB (BEE_LAYER_MAX_MIB): a layer over the ceiling fails the
build so a layer cannot silently grow back toward the monolith. If a semantic
layer is genuinely larger, split it further by purpose or package family. This
is why 30-nvidia-cuda-libs / 40-nvidia-dcgm-cuda are separate, and why the
~1 GB base rootfs was split into 00-base + 05-firmware + 08-desktop
(2026-09-04): device firmware and the local-console GUI are self-contained and
not on the SSH / headless path.
Verification (in build.sh, both build paths)
unsquashfs -s+ fullunsquashfs -strict-errorsextraction of every layer.- Reconstruct the merged rootfs in module order; assert
/usr/sbin/init,/usr/lib/systemd/systemd,/usr/local/bin/bee, the merged-usr symlinks, and the NVIDIA sentinels (nvidia-smi,nvidia.ko, GSP firmware,libnvidia-ml,dcgmi,nv-hostengine,dcgmproftester,libcudart/libcublas/libnccl,bee-nvidia.service). validate_iso_squashfs_layers: the final ISO carriesfilesystem.module, the module list and thelive/*.squashfsset match exactly, every layer is under the size ceiling, and a multi-layer variant never ships a single giant squashfs.validate_iso_rootfs_layout/validate_iso_nvidia_runtimealready iterate everylive/*.squashfs.
Regression tests: iso/builder/test-squashfs-layers.sh (classification,
completeness, per-layer validity, merged bootability, module order, missing
layer, size ceiling) and the multi-layer guard in
iso/builder/test-build-libs.sh. Both run at the top of every build.sh.