feat(iso): split the live medium into semantic SquashFS layers
Booting via BMC virtual CD reads the ~2.8 GB filesystem squashfs sequentially during the live-boot toram copy; a mid-read drop of the redirected medium loses the whole copy and fails the boot (v14). Split the rootfs into self-contained semantic layers so a retry re-reads at most one ~500-700 MiB layer, not everything. This is a resilience / reduced-re-read mechanism, not a fix for the virtual-media instability. NVIDIA variants now ship 7 layers (00-base, 05-firmware, 08-desktop, 10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs, 40-nvidia-dcgm-cuda) plus an explicit live/filesystem.module that fixes their OverlayFS order; amd/nogpu keep a single squashfs. - lib/squashfs-layers.sh: deterministic classifier (dpkg file ownership plus explicit rules for build.sh-injected files, never a path substring), per-layer mksquashfs, 800 MiB hard ceiling, unsquashfs -s plus strict extraction of every layer, merged-rootfs bootability check. - build.sh: split the monolith after the full lb build, verify and merge, write the module file, delete the monolith only then; abort before ISO assembly on any failure. Runs the builder test suites up front. - fast-path: force a full build for a multi-layer medium; fast_path_repack_squashfs hard-refuses (it would drop layers). - iso-validation.sh: validate_iso_squashfs_layers (module vs layer set match, size ceiling, no lone giant squashfs) and validate_iso_media_integrity (xorriso -check_media). - bee-install: honour filesystem.module order, abort on any layer failure. - 9013-toram-retry: record the real rsync exit code (it printed a false rc=0) and correct the "resumes the tail" comment (rsync without --partial keeps only fully-copied layers). No unsafe partial resume. - tests: test-squashfs-layers.sh plus a multi-layer guard in test-build-libs.sh; both run at the top of every build. - docs: bible-local architecture and decision, iso/README, iso-build-rules. Verified by a full nvidia build: 7 layers 622/199/256/466/37/567/562 MiB, every validator passes, xorriso -check_media good, merged rootfs bootable. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
bc57b85d3f
commit
b8c45d54c1
@@ -0,0 +1,111 @@
|
||||
# Semantic SquashFS Layers
|
||||
|
||||
**Module:** `squashfs-layers` (v1.0) - `iso/builder/lib/squashfs-layers.sh`
|
||||
|
||||
## Why
|
||||
|
||||
The live medium used to ship one `filesystem-v<ver>.squashfs` of about 2.8 GB.
|
||||
Booting through a BMC / IPMI virtual CD reads that single file sequentially
|
||||
during the live-boot `toram` copy. If the redirected medium drops off the bus
|
||||
part-way through (observed on v14), the whole copy is lost and the boot fails
|
||||
with `No supported filesystem images found at /live`.
|
||||
|
||||
Splitting the root filesystem into several self-contained squashfs layers means
|
||||
a mid-copy failure only costs the layer in flight, and the retry loop re-reads
|
||||
far less data.
|
||||
|
||||
This is a **resilience / reduced-re-read mechanism**. It is **not** a fix for
|
||||
the underlying BMC virtual-media instability - see
|
||||
`decisions/2026-09-04-squashfs-semantic-layers.md` and
|
||||
`decisions/2026-09-04-supported-systems-minimum-16gb-ram.md`.
|
||||
|
||||
## Layers (NVIDIA variants: `nvidia`, `nvidia-legacy`)
|
||||
|
||||
| Layer | Contents | Ownership |
|
||||
|---|---|---|
|
||||
| `filesystem-v<ver>-00-base.squashfs` | Debian rootfs, kernel + modules, systemd, Bee, networking, CLI diagnostic tools, dpkg database. Boot-critical, read first. | catch-all: any file not claimed by a higher layer |
|
||||
| `filesystem-v<ver>-05-firmware.squashfs` | Device firmware for NICs / wifi / non-NVIDIA GPUs | dpkg: `firmware-*` (but not `firmware-nvidia-*`, which rides with the driver; not `*-microcode`, which stays in base) |
|
||||
| `filesystem-v<ver>-08-desktop.squashfs` | Local-console GUI: X.org, `lightdm`, `openbox`, mesa/llvm, GTK, chromium, mupdf, fonts. SSH / headless never touches this layer. | dpkg: `xserver-xorg*`, `xfonts-*`, `x11-*`, `lightdm*`, `openbox`, `chromium*`, `mupdf*`, `libllvm*`, `libgl1-mesa-dri`, `libglx-mesa0`, `glx-alternative-mesa`, `libgtk-3-*` / `libgtk2.0-*` / `libgtkmm-*` / `libgtksourceview-*`, `libpango-*` / `libcairo2` / `libgdk-pixbuf-*`, `fonts-*`, `feh`, `scrot` |
|
||||
| `filesystem-v<ver>-10-nvidia-driver.squashfs` | NVIDIA kernel modules (`*.ko`), driver userspace (`libnvidia-*`, `libcuda.so*`), `nvidia-smi`, GSP firmware, OpenCL ICD, modprobe config, alternatives, `bee-nvidia-*` units | dpkg: `nvidia-modprobe`, `nvidia-kernel-common`, `glx-alternative-nvidia`, `nvidia-tesla-*`, `firmware-nvidia-*`, `libnvidia-*`, `nvtop`, `clinfo`, `ocl-icd-libopencl1`. injected: `usr/local/lib/nvidia/*.ko`, `usr/local/bin/nvidia-smi`, `usr/lib/libcuda.so*`, `usr/lib/libnvidia-*`, `usr/lib/firmware/nvidia/*`, `etc/OpenCL/vendors/nvidia.icd`, `bee-nvidia.service`, `bee-nvidia-load`/`-recover`, `bee-check-nvswitch` |
|
||||
| `filesystem-v<ver>-20-nvidia-platform.squashfs` | Fabric Manager, `libnvidia-nscq`, `nvlsm`, DCGM core + non-CUDA proprietary DCGM, related units | dpkg: `nvidia-fabricmanager`, `libnvidia-nscq`, `nvlsm`, `libibumad3`, `datacenter-gpu-manager-4-core`, `datacenter-gpu-manager-4-proprietary`. injected: `nvidia-fabricmanager.service.d/*` |
|
||||
| `filesystem-v<ver>-30-nvidia-cuda-libs.squashfs` | CUDA userspace runtime (cuBLAS / cuBLASLt / cudart), NCCL, nccl-tests, bee GPU stress worker | injected: `usr/lib/libnccl.so*`, `usr/lib/libcublas.so*`, `usr/lib/libcublasLt.so*`, `usr/lib/libcudart.so*`, `all_reduce_perf`, `bee-gpu-burn-worker`, `bee-gpu-burn`, `bee-nccl-gpu-stress`, `bee-john-gpu-stress` |
|
||||
| `filesystem-v<ver>-40-nvidia-dcgm-cuda.squashfs` | CUDA-linked DCGM components (`dcgmproftester` CUDA kernels) - split out from 30 because they are large | dpkg: `datacenter-gpu-manager-4-cuda13`, `datacenter-gpu-manager-4-proprietary-cuda13`. injected: `bee-dcgmproftester-staggered` |
|
||||
|
||||
`amd`, `nogpu`: single `filesystem-v<ver>.squashfs`, no `filesystem.module`. The
|
||||
builder leaves the monolith untouched for these variants.
|
||||
|
||||
## Classification rules
|
||||
|
||||
- **dpkg file ownership** (`var/lib/dpkg/info/*.list`), never a path substring.
|
||||
A file is not moved to an NVIDIA layer merely because its path contains
|
||||
`nvidia`.
|
||||
- Files that `build.sh` injects directly from the build cache and the project
|
||||
overlay belong to no `.deb`; they are classified by the explicit
|
||||
`bee_layer_injected_rules` table (extended-regex on the merged-usr-canonical
|
||||
relative path). Injected rules override dpkg ownership.
|
||||
- Pre-merged-usr paths dpkg records (`/bin`, `/sbin`, `/lib`, `/lib64`) are
|
||||
canonicalised to `/usr/...` before matching.
|
||||
- Anything unclaimed falls to `00-base`.
|
||||
- `bin`/`sbin`/`lib`/`lib64` are force-kept in `00-base` regardless of dpkg
|
||||
records (some `firmware-*` `.list` files carry a bare `/lib` entry).
|
||||
- The classifier fails the build if the partition is not complete and disjoint,
|
||||
if a merged-usr compat symlink is not in base, or if any upper layer would
|
||||
carry a real top-level `bin`/`sbin`/`lib`/`lib64` path.
|
||||
- Output is deterministic: every list is `LC_ALL=C` sorted, `mksquashfs` runs
|
||||
with `-processors 1` and fixed options. It does not depend on `find` order or
|
||||
locale.
|
||||
- Published on the medium: `live/filesystem.layers.txt` (per-layer file counts).
|
||||
|
||||
## Ordering (must stay consistent everywhere)
|
||||
|
||||
Module order: `00-base`, `05-firmware`, `08-desktop`, `10-nvidia-driver`,
|
||||
`20-nvidia-platform`, `30-nvidia-cuda-libs`, `40-nvidia-dcgm-cuda`.
|
||||
|
||||
`live-boot 1:20230131` with `MODULE=filesystem` (the default) reads
|
||||
`live/filesystem.module` verbatim - an ordered list of image names, one per
|
||||
line - and stacks them so the **last listed image has the highest OverlayFS
|
||||
priority**. If the module file were absent it would fall back to a
|
||||
locale-sensitive lexical `sort`; the `NN-` numeric prefixes make that
|
||||
equivalent, but the module file is authoritative.
|
||||
|
||||
The same order is used by:
|
||||
|
||||
- `bee-install`: reads `filesystem.module`, unpacks each layer with
|
||||
`unsquashfs -f` (last write wins), aborts install on any layer failure.
|
||||
- The builder fast path: currently **disabled** for a multi-layer medium; it
|
||||
forces a full `lb build` + deterministic re-split
|
||||
(`needs_full_build` returns true, `fast_path_repack_squashfs` hard-refuses).
|
||||
|
||||
So a higher-numbered layer always wins a conflict, in live-boot, in a disk
|
||||
install, and in a rebuild.
|
||||
|
||||
## Size budget
|
||||
|
||||
Target 500-700 MiB compressed per layer (`BEE_LAYER_TARGET_MIB`, soft warning).
|
||||
Hard ceiling 800 MiB (`BEE_LAYER_MAX_MIB`): a layer over the ceiling fails the
|
||||
build so a layer cannot silently grow back toward the monolith. If a semantic
|
||||
layer is genuinely larger, split it further by purpose or package family. This
|
||||
is why `30-nvidia-cuda-libs` / `40-nvidia-dcgm-cuda` are separate, and why the
|
||||
~1 GB base rootfs was split into `00-base` + `05-firmware` + `08-desktop`
|
||||
(2026-09-04): device firmware and the local-console GUI are self-contained and
|
||||
not on the SSH / headless path.
|
||||
|
||||
## Verification (in `build.sh`, both build paths)
|
||||
|
||||
1. `unsquashfs -s` + full `unsquashfs -strict-errors` extraction of every layer.
|
||||
2. Reconstruct the merged rootfs in module order; assert `/usr/sbin/init`,
|
||||
`/usr/lib/systemd/systemd`, `/usr/local/bin/bee`, the merged-usr symlinks, and the
|
||||
NVIDIA sentinels (`nvidia-smi`, `nvidia.ko`, GSP firmware, `libnvidia-ml`,
|
||||
`dcgmi`, `nv-hostengine`, `dcgmproftester`, `libcudart`/`libcublas`/`libnccl`,
|
||||
`bee-nvidia.service`).
|
||||
3. `validate_iso_squashfs_layers`: the final ISO carries `filesystem.module`,
|
||||
the module list and the `live/*.squashfs` set match exactly, every layer is
|
||||
under the size ceiling, and a multi-layer variant never ships a single
|
||||
giant squashfs.
|
||||
4. `validate_iso_rootfs_layout` / `validate_iso_nvidia_runtime` already iterate
|
||||
every `live/*.squashfs`.
|
||||
|
||||
Regression tests: `iso/builder/test-squashfs-layers.sh` (classification,
|
||||
completeness, per-layer validity, merged bootability, module order, missing
|
||||
layer, size ceiling) and the multi-layer guard in
|
||||
`iso/builder/test-build-libs.sh`. Both run at the top of every `build.sh`.
|
||||
@@ -0,0 +1,67 @@
|
||||
# Split the live medium into semantic SquashFS layers
|
||||
|
||||
**Date:** 2026-09-04
|
||||
**Status:** active
|
||||
|
||||
## Context
|
||||
|
||||
EASY-BEE shipped one `filesystem-v<ver>.squashfs` of about 2.8 GB. v13 booted
|
||||
through the BMC virtual CD; v14 did not. During the live-boot `toram` copy,
|
||||
`rsync` reads that single file sequentially and the read stopped part-way
|
||||
through (around the middle), after which live-boot moved the partial RAM copy
|
||||
over the medium and the boot failed.
|
||||
|
||||
The squashfs inside the build artifact was verified intact
|
||||
(`unsquashfs -strict-errors`). Supported systems have at least 16 GB of RAM
|
||||
(`2026-09-04-supported-systems-minimum-16gb-ram.md`), so this is not memory
|
||||
exhaustion. The failure is in the virtual-media read path.
|
||||
|
||||
live-boot v14.03 already supports several `/live/*.squashfs` merged via
|
||||
OverlayFS, and `bee-install` already unpacks multiple squashfs in lexical
|
||||
order. The old builder fast path assumed exactly one squashfs (it picked the
|
||||
first and deleted the rest).
|
||||
|
||||
## Decision
|
||||
|
||||
- After a full `lb build`, the builder deterministically splits the monolith
|
||||
into semantic layers (see `architecture/squashfs-layers.md`): `00-base`,
|
||||
`05-firmware`, `08-desktop`, `10-nvidia-driver`, `20-nvidia-platform`,
|
||||
`30-nvidia-cuda-libs`, `40-nvidia-dcgm-cuda` for NVIDIA variants; a single
|
||||
untouched squashfs for `amd` / `nogpu`. The base rootfs was ~1 GB compressed,
|
||||
so device firmware (`firmware-*`) and the local-console GUI stack (X.org,
|
||||
lightdm, mesa, chromium, mupdf, fonts) - neither on the SSH / headless path -
|
||||
were carved out to keep every layer near the 500-700 MiB target.
|
||||
- Layer boundaries come from **dpkg file ownership** plus an explicit,
|
||||
code-described rule table for the non-`.deb` files `build.sh` injects. A path
|
||||
is never moved to an NVIDIA layer just because it contains `nvidia`.
|
||||
- Layer names carry the project version (`filesystem-v<ver>-NN-slug.squashfs`),
|
||||
matching the existing `filesystem-v<ver>.squashfs` scheme.
|
||||
- An explicit `live/filesystem.module` fixes the order live-boot uses. The
|
||||
same order governs `bee-install` and any rebuild. A higher-numbered layer
|
||||
always wins a conflict.
|
||||
- Each layer stays a self-contained valid squashfs, kept under an 800 MiB
|
||||
compressed ceiling (build fails otherwise). merged-usr (`/bin`, `/sbin`,
|
||||
`/lib`, `/lib64` as symlinks) is preserved; upper layers never carry a real
|
||||
top-level compat directory and never use opaque/whiteout semantics.
|
||||
- The monolith is deleted only after every layer is built, individually
|
||||
verified, and successfully re-merged into a bootable rootfs. Any failure
|
||||
aborts the build before the ISO is assembled - a partial layer set is never
|
||||
published.
|
||||
- The builder fast path is **disabled for a multi-layer medium**: it forces a
|
||||
full `lb build` + re-split. `fast_path_repack_squashfs` hard-refuses to run.
|
||||
A layer-aware repack may come later.
|
||||
- `9013-toram-retry` now records the real `rsync` exit code (it printed a false
|
||||
`rc=0`) and no longer claims it resumes the tail of the current file
|
||||
(`rsync` without `--partial` keeps only fully-copied files). No unsafe
|
||||
partial-file resume is introduced.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A mid-copy virtual-media drop now costs one layer (target 500-700 MiB, hard
|
||||
ceiling 800 MiB), not the whole 2.8 GB rootfs; the retry re-reads only that
|
||||
layer.
|
||||
- Every NVIDIA ISO build now runs the full path (no fast path) until a
|
||||
layer-aware repack exists. Build time is unchanged for full builds;
|
||||
overlay-only iterations lose the fast path for NVIDIA variants.
|
||||
- The split does not address the root BMC virtual-media instability; it reduces
|
||||
the blast radius and the repeatedly-read volume.
|
||||
@@ -19,3 +19,4 @@ One file per decision, named `YYYY-MM-DD-short-topic.md`.
|
||||
| 2026-09-03 | PCIe link verdict comes only from the existing real-traffic GPU bandwidth SAT | active |
|
||||
| 2026-09-03 | Runtime Copy to RAM succeeds only after active loop devices move to tmpfs | active |
|
||||
| 2026-09-04 | Supported systems have at least 16 GB of RAM | active |
|
||||
| 2026-09-04 | Split the live medium into semantic SquashFS layers | active |
|
||||
|
||||
@@ -15,6 +15,33 @@ This applies to:
|
||||
- `iso/builder/config/package-lists/*.list.chroot`
|
||||
- Any package referenced in bootloader configs, hooks, or overlay scripts
|
||||
|
||||
## SquashFS layer rule
|
||||
|
||||
The NVIDIA live medium is not one squashfs. `build.sh` splits the monolith into
|
||||
semantic layers (`filesystem-v<ver>-NN-*.squashfs`) after the full `lb build`,
|
||||
driven by `iso/builder/lib/squashfs-layers.sh`. See
|
||||
`bible-local/architecture/squashfs-layers.md` for the layer model and
|
||||
`decisions/2026-09-04-squashfs-semantic-layers.md` for why.
|
||||
|
||||
Rules:
|
||||
|
||||
- Classify by dpkg file ownership, not path substring. New non-`.deb` files
|
||||
injected by `build.sh` must get an explicit entry in
|
||||
`bee_layer_injected_rules` (or they fall to `00-base`).
|
||||
- `live/filesystem.module` is authoritative for layer order; keep it in sync
|
||||
with the `NN-` prefixes. Later layer wins a conflict, in live-boot,
|
||||
`bee-install`, and any rebuild.
|
||||
- Keep each layer a self-contained valid squashfs under the 800 MiB ceiling.
|
||||
Split a semantic layer further rather than blindly slicing bytes.
|
||||
- Never let an upper layer carry a real top-level `bin`/`sbin`/`lib`/`lib64`
|
||||
path - that breaks merged-usr the same way the fast-path bug did.
|
||||
- The builder deletes the monolith only after every layer verifies and
|
||||
re-merges into a bootable rootfs. Any failure must abort before ISO assembly.
|
||||
- The fast path stays disabled for a multi-layer medium until a layer-aware
|
||||
repack exists. Do not "take the first squashfs".
|
||||
- `test-squashfs-layers.sh` and `test-build-libs.sh` run at the top of every
|
||||
build; keep them green.
|
||||
|
||||
## Bootloader sync rule
|
||||
|
||||
The ISO has two canonical bootloader templates whose live entries must remain
|
||||
|
||||
Reference in New Issue
Block a user