Files
bee/bible-local/architecture/squashfs-layers.md
T
Mikhail ChusavitinandClaude Sonnet 5 b8c45d54c1 feat(iso): split the live medium into semantic SquashFS layers
Booting via BMC virtual CD reads the ~2.8 GB filesystem squashfs
sequentially during the live-boot toram copy; a mid-read drop of the
redirected medium loses the whole copy and fails the boot (v14). Split
the rootfs into self-contained semantic layers so a retry re-reads at
most one ~500-700 MiB layer, not everything. This is a resilience /
reduced-re-read mechanism, not a fix for the virtual-media instability.

NVIDIA variants now ship 7 layers (00-base, 05-firmware, 08-desktop,
10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs,
40-nvidia-dcgm-cuda) plus an explicit live/filesystem.module that fixes
their OverlayFS order; amd/nogpu keep a single squashfs.

- lib/squashfs-layers.sh: deterministic classifier (dpkg file ownership
  plus explicit rules for build.sh-injected files, never a path
  substring), per-layer mksquashfs, 800 MiB hard ceiling, unsquashfs -s
  plus strict extraction of every layer, merged-rootfs bootability check.
- build.sh: split the monolith after the full lb build, verify and merge,
  write the module file, delete the monolith only then; abort before ISO
  assembly on any failure. Runs the builder test suites up front.
- fast-path: force a full build for a multi-layer medium;
  fast_path_repack_squashfs hard-refuses (it would drop layers).
- iso-validation.sh: validate_iso_squashfs_layers (module vs layer set
  match, size ceiling, no lone giant squashfs) and
  validate_iso_media_integrity (xorriso -check_media).
- bee-install: honour filesystem.module order, abort on any layer failure.
- 9013-toram-retry: record the real rsync exit code (it printed a false
  rc=0) and correct the "resumes the tail" comment (rsync without
  --partial keeps only fully-copied layers). No unsafe partial resume.
- tests: test-squashfs-layers.sh plus a multi-layer guard in
  test-build-libs.sh; both run at the top of every build.
- docs: bible-local architecture and decision, iso/README, iso-build-rules.

Verified by a full nvidia build: 7 layers 622/199/256/466/37/567/562 MiB,
every validator passes, xorriso -check_media good, merged rootfs bootable.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 15:38:14 +03:00

7.6 KiB

Semantic SquashFS Layers

Module: squashfs-layers (v1.0) - iso/builder/lib/squashfs-layers.sh

Why

The live medium used to ship one filesystem-v<ver>.squashfs of about 2.8 GB. Booting through a BMC / IPMI virtual CD reads that single file sequentially during the live-boot toram copy. If the redirected medium drops off the bus part-way through (observed on v14), the whole copy is lost and the boot fails with No supported filesystem images found at /live.

Splitting the root filesystem into several self-contained squashfs layers means a mid-copy failure only costs the layer in flight, and the retry loop re-reads far less data.

This is a resilience / reduced-re-read mechanism. It is not a fix for the underlying BMC virtual-media instability - see decisions/2026-09-04-squashfs-semantic-layers.md and decisions/2026-09-04-supported-systems-minimum-16gb-ram.md.

Layers (NVIDIA variants: nvidia, nvidia-legacy)

Layer Contents Ownership
filesystem-v<ver>-00-base.squashfs Debian rootfs, kernel + modules, systemd, Bee, networking, CLI diagnostic tools, dpkg database. Boot-critical, read first. catch-all: any file not claimed by a higher layer
filesystem-v<ver>-05-firmware.squashfs Device firmware for NICs / wifi / non-NVIDIA GPUs dpkg: firmware-* (but not firmware-nvidia-*, which rides with the driver; not *-microcode, which stays in base)
filesystem-v<ver>-08-desktop.squashfs Local-console GUI: X.org, lightdm, openbox, mesa/llvm, GTK, chromium, mupdf, fonts. SSH / headless never touches this layer. dpkg: xserver-xorg*, xfonts-*, x11-*, lightdm*, openbox, chromium*, mupdf*, libllvm*, libgl1-mesa-dri, libglx-mesa0, glx-alternative-mesa, libgtk-3-* / libgtk2.0-* / libgtkmm-* / libgtksourceview-*, libpango-* / libcairo2 / libgdk-pixbuf-*, fonts-*, feh, scrot
filesystem-v<ver>-10-nvidia-driver.squashfs NVIDIA kernel modules (*.ko), driver userspace (libnvidia-*, libcuda.so*), nvidia-smi, GSP firmware, OpenCL ICD, modprobe config, alternatives, bee-nvidia-* units dpkg: nvidia-modprobe, nvidia-kernel-common, glx-alternative-nvidia, nvidia-tesla-*, firmware-nvidia-*, libnvidia-*, nvtop, clinfo, ocl-icd-libopencl1. injected: usr/local/lib/nvidia/*.ko, usr/local/bin/nvidia-smi, usr/lib/libcuda.so*, usr/lib/libnvidia-*, usr/lib/firmware/nvidia/*, etc/OpenCL/vendors/nvidia.icd, bee-nvidia.service, bee-nvidia-load/-recover, bee-check-nvswitch
filesystem-v<ver>-20-nvidia-platform.squashfs Fabric Manager, libnvidia-nscq, nvlsm, DCGM core + non-CUDA proprietary DCGM, related units dpkg: nvidia-fabricmanager, libnvidia-nscq, nvlsm, libibumad3, datacenter-gpu-manager-4-core, datacenter-gpu-manager-4-proprietary. injected: nvidia-fabricmanager.service.d/*
filesystem-v<ver>-30-nvidia-cuda-libs.squashfs CUDA userspace runtime (cuBLAS / cuBLASLt / cudart), NCCL, nccl-tests, bee GPU stress worker injected: usr/lib/libnccl.so*, usr/lib/libcublas.so*, usr/lib/libcublasLt.so*, usr/lib/libcudart.so*, all_reduce_perf, bee-gpu-burn-worker, bee-gpu-burn, bee-nccl-gpu-stress, bee-john-gpu-stress
filesystem-v<ver>-40-nvidia-dcgm-cuda.squashfs CUDA-linked DCGM components (dcgmproftester CUDA kernels) - split out from 30 because they are large dpkg: datacenter-gpu-manager-4-cuda13, datacenter-gpu-manager-4-proprietary-cuda13. injected: bee-dcgmproftester-staggered

amd, nogpu: single filesystem-v<ver>.squashfs, no filesystem.module. The builder leaves the monolith untouched for these variants.

Classification rules

  • dpkg file ownership (var/lib/dpkg/info/*.list), never a path substring. A file is not moved to an NVIDIA layer merely because its path contains nvidia.
  • Files that build.sh injects directly from the build cache and the project overlay belong to no .deb; they are classified by the explicit bee_layer_injected_rules table (extended-regex on the merged-usr-canonical relative path). Injected rules override dpkg ownership.
  • Pre-merged-usr paths dpkg records (/bin, /sbin, /lib, /lib64) are canonicalised to /usr/... before matching.
  • Anything unclaimed falls to 00-base.
  • bin/sbin/lib/lib64 are force-kept in 00-base regardless of dpkg records (some firmware-* .list files carry a bare /lib entry).
  • The classifier fails the build if the partition is not complete and disjoint, if a merged-usr compat symlink is not in base, or if any upper layer would carry a real top-level bin/sbin/lib/lib64 path.
  • Output is deterministic: every list is LC_ALL=C sorted, mksquashfs runs with -processors 1 and fixed options. It does not depend on find order or locale.
  • Published on the medium: live/filesystem.layers.txt (per-layer file counts).

Ordering (must stay consistent everywhere)

Module order: 00-base, 05-firmware, 08-desktop, 10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs, 40-nvidia-dcgm-cuda.

live-boot 1:20230131 with MODULE=filesystem (the default) reads live/filesystem.module verbatim - an ordered list of image names, one per line - and stacks them so the last listed image has the highest OverlayFS priority. If the module file were absent it would fall back to a locale-sensitive lexical sort; the NN- numeric prefixes make that equivalent, but the module file is authoritative.

The same order is used by:

  • bee-install: reads filesystem.module, unpacks each layer with unsquashfs -f (last write wins), aborts install on any layer failure.
  • The builder fast path: currently disabled for a multi-layer medium; it forces a full lb build + deterministic re-split (needs_full_build returns true, fast_path_repack_squashfs hard-refuses).

So a higher-numbered layer always wins a conflict, in live-boot, in a disk install, and in a rebuild.

Size budget

Target 500-700 MiB compressed per layer (BEE_LAYER_TARGET_MIB, soft warning). Hard ceiling 800 MiB (BEE_LAYER_MAX_MIB): a layer over the ceiling fails the build so a layer cannot silently grow back toward the monolith. If a semantic layer is genuinely larger, split it further by purpose or package family. This is why 30-nvidia-cuda-libs / 40-nvidia-dcgm-cuda are separate, and why the ~1 GB base rootfs was split into 00-base + 05-firmware + 08-desktop (2026-09-04): device firmware and the local-console GUI are self-contained and not on the SSH / headless path.

Verification (in build.sh, both build paths)

  1. unsquashfs -s + full unsquashfs -strict-errors extraction of every layer.
  2. Reconstruct the merged rootfs in module order; assert /usr/sbin/init, /usr/lib/systemd/systemd, /usr/local/bin/bee, the merged-usr symlinks, and the NVIDIA sentinels (nvidia-smi, nvidia.ko, GSP firmware, libnvidia-ml, dcgmi, nv-hostengine, dcgmproftester, libcudart/libcublas/libnccl, bee-nvidia.service).
  3. validate_iso_squashfs_layers: the final ISO carries filesystem.module, the module list and the live/*.squashfs set match exactly, every layer is under the size ceiling, and a multi-layer variant never ships a single giant squashfs.
  4. validate_iso_rootfs_layout / validate_iso_nvidia_runtime already iterate every live/*.squashfs.

Regression tests: iso/builder/test-squashfs-layers.sh (classification, completeness, per-layer validity, merged bootability, module order, missing layer, size ceiling) and the multi-layer guard in iso/builder/test-build-libs.sh. Both run at the top of every build.sh.