feat(iso): split the live medium into semantic SquashFS layers

Booting via BMC virtual CD reads the ~2.8 GB filesystem squashfs
sequentially during the live-boot toram copy; a mid-read drop of the
redirected medium loses the whole copy and fails the boot (v14). Split
the rootfs into self-contained semantic layers so a retry re-reads at
most one ~500-700 MiB layer, not everything. This is a resilience /
reduced-re-read mechanism, not a fix for the virtual-media instability.

NVIDIA variants now ship 7 layers (00-base, 05-firmware, 08-desktop,
10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs,
40-nvidia-dcgm-cuda) plus an explicit live/filesystem.module that fixes
their OverlayFS order; amd/nogpu keep a single squashfs.

- lib/squashfs-layers.sh: deterministic classifier (dpkg file ownership
  plus explicit rules for build.sh-injected files, never a path
  substring), per-layer mksquashfs, 800 MiB hard ceiling, unsquashfs -s
  plus strict extraction of every layer, merged-rootfs bootability check.
- build.sh: split the monolith after the full lb build, verify and merge,
  write the module file, delete the monolith only then; abort before ISO
  assembly on any failure. Runs the builder test suites up front.
- fast-path: force a full build for a multi-layer medium;
  fast_path_repack_squashfs hard-refuses (it would drop layers).
- iso-validation.sh: validate_iso_squashfs_layers (module vs layer set
  match, size ceiling, no lone giant squashfs) and
  validate_iso_media_integrity (xorriso -check_media).
- bee-install: honour filesystem.module order, abort on any layer failure.
- 9013-toram-retry: record the real rsync exit code (it printed a false
  rc=0) and correct the "resumes the tail" comment (rsync without
  --partial keeps only fully-copied layers). No unsafe partial resume.
- tests: test-squashfs-layers.sh plus a multi-layer guard in
  test-build-libs.sh; both run at the top of every build.
- docs: bible-local architecture and decision, iso/README, iso-build-rules.

Verified by a full nvidia build: 7 layers 622/199/256/466/37/567/562 MiB,
every validator passes, xorriso -check_media good, merged rootfs bootable.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Mikhail Chusavitin
2026-09-04 15:38:14 +03:00
co-authored by Claude Sonnet 5
parent bc57b85d3f
commit b8c45d54c1
13 changed files with 1056 additions and 13 deletions
+27
View File
@@ -15,6 +15,33 @@ This applies to:
- `iso/builder/config/package-lists/*.list.chroot`
- Any package referenced in bootloader configs, hooks, or overlay scripts
## SquashFS layer rule
The NVIDIA live medium is not one squashfs. `build.sh` splits the monolith into
semantic layers (`filesystem-v<ver>-NN-*.squashfs`) after the full `lb build`,
driven by `iso/builder/lib/squashfs-layers.sh`. See
`bible-local/architecture/squashfs-layers.md` for the layer model and
`decisions/2026-09-04-squashfs-semantic-layers.md` for why.
Rules:
- Classify by dpkg file ownership, not path substring. New non-`.deb` files
injected by `build.sh` must get an explicit entry in
`bee_layer_injected_rules` (or they fall to `00-base`).
- `live/filesystem.module` is authoritative for layer order; keep it in sync
with the `NN-` prefixes. Later layer wins a conflict, in live-boot,
`bee-install`, and any rebuild.
- Keep each layer a self-contained valid squashfs under the 800 MiB ceiling.
Split a semantic layer further rather than blindly slicing bytes.
- Never let an upper layer carry a real top-level `bin`/`sbin`/`lib`/`lib64`
path - that breaks merged-usr the same way the fast-path bug did.
- The builder deletes the monolith only after every layer verifies and
re-merges into a bootable rootfs. Any failure must abort before ISO assembly.
- The fast path stays disabled for a multi-layer medium until a layer-aware
repack exists. Do not "take the first squashfs".
- `test-squashfs-layers.sh` and `test-build-libs.sh` run at the top of every
build; keep them green.
## Bootloader sync rule
The ISO has two canonical bootloader templates whose live entries must remain