Files
bee/bible-local/decisions/2026-09-04-squashfs-semantic-layers.md
T
Mikhail ChusavitinandClaude Sonnet 5 b8c45d54c1 feat(iso): split the live medium into semantic SquashFS layers
Booting via BMC virtual CD reads the ~2.8 GB filesystem squashfs
sequentially during the live-boot toram copy; a mid-read drop of the
redirected medium loses the whole copy and fails the boot (v14). Split
the rootfs into self-contained semantic layers so a retry re-reads at
most one ~500-700 MiB layer, not everything. This is a resilience /
reduced-re-read mechanism, not a fix for the virtual-media instability.

NVIDIA variants now ship 7 layers (00-base, 05-firmware, 08-desktop,
10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs,
40-nvidia-dcgm-cuda) plus an explicit live/filesystem.module that fixes
their OverlayFS order; amd/nogpu keep a single squashfs.

- lib/squashfs-layers.sh: deterministic classifier (dpkg file ownership
  plus explicit rules for build.sh-injected files, never a path
  substring), per-layer mksquashfs, 800 MiB hard ceiling, unsquashfs -s
  plus strict extraction of every layer, merged-rootfs bootability check.
- build.sh: split the monolith after the full lb build, verify and merge,
  write the module file, delete the monolith only then; abort before ISO
  assembly on any failure. Runs the builder test suites up front.
- fast-path: force a full build for a multi-layer medium;
  fast_path_repack_squashfs hard-refuses (it would drop layers).
- iso-validation.sh: validate_iso_squashfs_layers (module vs layer set
  match, size ceiling, no lone giant squashfs) and
  validate_iso_media_integrity (xorriso -check_media).
- bee-install: honour filesystem.module order, abort on any layer failure.
- 9013-toram-retry: record the real rsync exit code (it printed a false
  rc=0) and correct the "resumes the tail" comment (rsync without
  --partial keeps only fully-copied layers). No unsafe partial resume.
- tests: test-squashfs-layers.sh plus a multi-layer guard in
  test-build-libs.sh; both run at the top of every build.
- docs: bible-local architecture and decision, iso/README, iso-build-rules.

Verified by a full nvidia build: 7 layers 622/199/256/466/37/567/562 MiB,
every validator passes, xorriso -check_media good, merged rootfs bootable.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 15:38:14 +03:00

3.6 KiB

Split the live medium into semantic SquashFS layers

Date: 2026-09-04 Status: active

Context

EASY-BEE shipped one filesystem-v<ver>.squashfs of about 2.8 GB. v13 booted through the BMC virtual CD; v14 did not. During the live-boot toram copy, rsync reads that single file sequentially and the read stopped part-way through (around the middle), after which live-boot moved the partial RAM copy over the medium and the boot failed.

The squashfs inside the build artifact was verified intact (unsquashfs -strict-errors). Supported systems have at least 16 GB of RAM (2026-09-04-supported-systems-minimum-16gb-ram.md), so this is not memory exhaustion. The failure is in the virtual-media read path.

live-boot v14.03 already supports several /live/*.squashfs merged via OverlayFS, and bee-install already unpacks multiple squashfs in lexical order. The old builder fast path assumed exactly one squashfs (it picked the first and deleted the rest).

Decision

  • After a full lb build, the builder deterministically splits the monolith into semantic layers (see architecture/squashfs-layers.md): 00-base, 05-firmware, 08-desktop, 10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs, 40-nvidia-dcgm-cuda for NVIDIA variants; a single untouched squashfs for amd / nogpu. The base rootfs was ~1 GB compressed, so device firmware (firmware-*) and the local-console GUI stack (X.org, lightdm, mesa, chromium, mupdf, fonts) - neither on the SSH / headless path - were carved out to keep every layer near the 500-700 MiB target.
  • Layer boundaries come from dpkg file ownership plus an explicit, code-described rule table for the non-.deb files build.sh injects. A path is never moved to an NVIDIA layer just because it contains nvidia.
  • Layer names carry the project version (filesystem-v<ver>-NN-slug.squashfs), matching the existing filesystem-v<ver>.squashfs scheme.
  • An explicit live/filesystem.module fixes the order live-boot uses. The same order governs bee-install and any rebuild. A higher-numbered layer always wins a conflict.
  • Each layer stays a self-contained valid squashfs, kept under an 800 MiB compressed ceiling (build fails otherwise). merged-usr (/bin, /sbin, /lib, /lib64 as symlinks) is preserved; upper layers never carry a real top-level compat directory and never use opaque/whiteout semantics.
  • The monolith is deleted only after every layer is built, individually verified, and successfully re-merged into a bootable rootfs. Any failure aborts the build before the ISO is assembled - a partial layer set is never published.
  • The builder fast path is disabled for a multi-layer medium: it forces a full lb build + re-split. fast_path_repack_squashfs hard-refuses to run. A layer-aware repack may come later.
  • 9013-toram-retry now records the real rsync exit code (it printed a false rc=0) and no longer claims it resumes the tail of the current file (rsync without --partial keeps only fully-copied files). No unsafe partial-file resume is introduced.

Consequences

  • A mid-copy virtual-media drop now costs one layer (target 500-700 MiB, hard ceiling 800 MiB), not the whole 2.8 GB rootfs; the retry re-reads only that layer.
  • Every NVIDIA ISO build now runs the full path (no fast path) until a layer-aware repack exists. Build time is unchanged for full builds; overlay-only iterations lose the fast path for NVIDIA variants.
  • The split does not address the root BMC virtual-media instability; it reduces the blast radius and the repeatedly-read volume.