# Semantic SquashFS Layers **Module:** `squashfs-layers` (v1.0) - `iso/builder/lib/squashfs-layers.sh` ## Why The live medium used to ship one `filesystem-v.squashfs` of about 2.8 GB. Booting through a BMC / IPMI virtual CD reads that single file sequentially during the live-boot `toram` copy. If the redirected medium drops off the bus part-way through (observed on v14), the whole copy is lost and the boot fails with `No supported filesystem images found at /live`. Splitting the root filesystem into several self-contained squashfs layers means a mid-copy failure only costs the layer in flight, and the retry loop re-reads far less data. This is a **resilience / reduced-re-read mechanism**. It is **not** a fix for the underlying BMC virtual-media instability - see `decisions/2026-09-04-squashfs-semantic-layers.md` and `decisions/2026-09-04-supported-systems-minimum-16gb-ram.md`. ## Layers (NVIDIA variants: `nvidia`, `nvidia-legacy`) | Layer | Contents | Ownership | |---|---|---| | `filesystem-v-00-base.squashfs` | Debian rootfs, kernel + modules, systemd, Bee, networking, CLI diagnostic tools, dpkg database. Boot-critical, read first. | catch-all: any file not claimed by a higher layer | | `filesystem-v-05-firmware.squashfs` | Device firmware for NICs / wifi / non-NVIDIA GPUs | dpkg: `firmware-*` (but not `firmware-nvidia-*`, which rides with the driver; not `*-microcode`, which stays in base) | | `filesystem-v-08-desktop.squashfs` | Local-console GUI: X.org, `lightdm`, `openbox`, mesa/llvm, GTK, chromium, mupdf, fonts. SSH / headless never touches this layer. | dpkg: `xserver-xorg*`, `xfonts-*`, `x11-*`, `lightdm*`, `openbox`, `chromium*`, `mupdf*`, `libllvm*`, `libgl1-mesa-dri`, `libglx-mesa0`, `glx-alternative-mesa`, `libgtk-3-*` / `libgtk2.0-*` / `libgtkmm-*` / `libgtksourceview-*`, `libpango-*` / `libcairo2` / `libgdk-pixbuf-*`, `fonts-*`, `feh`, `scrot` | | `filesystem-v-10-nvidia-driver.squashfs` | NVIDIA kernel modules (`*.ko`), driver userspace (`libnvidia-*`, `libcuda.so*`), `nvidia-smi`, GSP firmware, OpenCL ICD, modprobe config, alternatives, `bee-nvidia-*` units | dpkg: `nvidia-modprobe`, `nvidia-kernel-common`, `glx-alternative-nvidia`, `nvidia-tesla-*`, `firmware-nvidia-*`, `libnvidia-*`, `nvtop`, `clinfo`, `ocl-icd-libopencl1`. injected: `usr/local/lib/nvidia/*.ko`, `usr/local/bin/nvidia-smi`, `usr/lib/libcuda.so*`, `usr/lib/libnvidia-*`, `usr/lib/firmware/nvidia/*`, `etc/OpenCL/vendors/nvidia.icd`, `bee-nvidia.service`, `bee-nvidia-load`/`-recover`, `bee-check-nvswitch` | | `filesystem-v-20-nvidia-platform.squashfs` | Fabric Manager, `libnvidia-nscq`, `nvlsm`, DCGM core + non-CUDA proprietary DCGM, related units | dpkg: `nvidia-fabricmanager`, `libnvidia-nscq`, `nvlsm`, `libibumad3`, `datacenter-gpu-manager-4-core`, `datacenter-gpu-manager-4-proprietary`. injected: `nvidia-fabricmanager.service.d/*` | | `filesystem-v-30-nvidia-cuda-libs.squashfs` | CUDA userspace runtime (cuBLAS / cuBLASLt / cudart), NCCL, nccl-tests, bee GPU stress worker | injected: `usr/lib/libnccl.so*`, `usr/lib/libcublas.so*`, `usr/lib/libcublasLt.so*`, `usr/lib/libcudart.so*`, `all_reduce_perf`, `bee-gpu-burn-worker`, `bee-gpu-burn`, `bee-nccl-gpu-stress`, `bee-john-gpu-stress` | | `filesystem-v-40-nvidia-dcgm-cuda.squashfs` | CUDA-linked DCGM components (`dcgmproftester` CUDA kernels) - split out from 30 because they are large | dpkg: `datacenter-gpu-manager-4-cuda13`, `datacenter-gpu-manager-4-proprietary-cuda13`. injected: `bee-dcgmproftester-staggered` | `amd`, `nogpu`: single `filesystem-v.squashfs`, no `filesystem.module`. The builder leaves the monolith untouched for these variants. ## Classification rules - **dpkg file ownership** (`var/lib/dpkg/info/*.list`), never a path substring. A file is not moved to an NVIDIA layer merely because its path contains `nvidia`. - Files that `build.sh` injects directly from the build cache and the project overlay belong to no `.deb`; they are classified by the explicit `bee_layer_injected_rules` table (extended-regex on the merged-usr-canonical relative path). Injected rules override dpkg ownership. - Pre-merged-usr paths dpkg records (`/bin`, `/sbin`, `/lib`, `/lib64`) are canonicalised to `/usr/...` before matching. - Anything unclaimed falls to `00-base`. - `bin`/`sbin`/`lib`/`lib64` are force-kept in `00-base` regardless of dpkg records (some `firmware-*` `.list` files carry a bare `/lib` entry). - The classifier fails the build if the partition is not complete and disjoint, if a merged-usr compat symlink is not in base, or if any upper layer would carry a real top-level `bin`/`sbin`/`lib`/`lib64` path. - Output is deterministic: every list is `LC_ALL=C` sorted, `mksquashfs` runs with `-processors 1` and fixed options. It does not depend on `find` order or locale. - Published on the medium: `live/filesystem.layers.txt` (per-layer file counts). ## Ordering (must stay consistent everywhere) Module order: `00-base`, `05-firmware`, `08-desktop`, `10-nvidia-driver`, `20-nvidia-platform`, `30-nvidia-cuda-libs`, `40-nvidia-dcgm-cuda`. `live-boot 1:20230131` with `MODULE=filesystem` (the default) reads `live/filesystem.module` verbatim - an ordered list of image names, one per line - and stacks them so the **last listed image has the highest OverlayFS priority**. If the module file were absent it would fall back to a locale-sensitive lexical `sort`; the `NN-` numeric prefixes make that equivalent, but the module file is authoritative. The same order is used by: - `bee-install`: reads `filesystem.module`, unpacks each layer with `unsquashfs -f` (last write wins), aborts install on any layer failure. - The builder fast path: currently **disabled** for a multi-layer medium; it forces a full `lb build` + deterministic re-split (`needs_full_build` returns true, `fast_path_repack_squashfs` hard-refuses). So a higher-numbered layer always wins a conflict, in live-boot, in a disk install, and in a rebuild. ## Size budget Target 500-700 MiB compressed per layer (`BEE_LAYER_TARGET_MIB`, soft warning). Hard ceiling 800 MiB (`BEE_LAYER_MAX_MIB`): a layer over the ceiling fails the build so a layer cannot silently grow back toward the monolith. If a semantic layer is genuinely larger, split it further by purpose or package family. This is why `30-nvidia-cuda-libs` / `40-nvidia-dcgm-cuda` are separate, and why the ~1 GB base rootfs was split into `00-base` + `05-firmware` + `08-desktop` (2026-09-04): device firmware and the local-console GUI are self-contained and not on the SSH / headless path. ## Verification (in `build.sh`, both build paths) 1. `unsquashfs -s` + full `unsquashfs -strict-errors` extraction of every layer. 2. Reconstruct the merged rootfs in module order; assert `/usr/sbin/init`, `/usr/lib/systemd/systemd`, `/usr/local/bin/bee`, the merged-usr symlinks, and the NVIDIA sentinels (`nvidia-smi`, `nvidia.ko`, GSP firmware, `libnvidia-ml`, `dcgmi`, `nv-hostengine`, `dcgmproftester`, `libcudart`/`libcublas`/`libnccl`, `bee-nvidia.service`). 3. `validate_iso_squashfs_layers`: the final ISO carries `filesystem.module`, the module list and the `live/*.squashfs` set match exactly, every layer is under the size ceiling, and a multi-layer variant never ships a single giant squashfs. 4. `validate_iso_rootfs_layout` / `validate_iso_nvidia_runtime` already iterate every `live/*.squashfs`. Regression tests: `iso/builder/test-squashfs-layers.sh` (classification, completeness, per-layer validity, merged bootability, module order, missing layer, size ceiling) and the multi-layer guard in `iso/builder/test-build-libs.sh`. Both run at the top of every `build.sh`.