Commit Graph
2 Commits
Author SHA1 Message Date
mchusandClaude Sonnet 5 9c5c29239c fix(iso): let empty directories survive the squashfs layer split
bee_layer_classify only ever tracked regular files and symlinks
(`find ... -type f -o -type l`), so any directory that is empty at
build time — like /var/log/nvidia-dcgm, correctly created and chowned
by the datacenter-gpu-manager postinst — was silently dropped from
every layer's rsync --files-from list and never reached the built
ISO. This is why bbc6fb1's 78d1b9b follow-up (seeding a marker file
just for that one path) kept the directory alive: it was a targeted
workaround for a general gap in the classifier, not a fix of it.

Replace that workaround with the general mechanism: classify also
walks every directory, computes the subset that is empty all the way
down (no file or symlink anywhere in its subtree — a directory that
does hold files needs no entry, rsync already recreates it as an
implied parent), and assigns each one to a layer via the same
dpkg-ownership / injected-rule precedence used for files. Add an
injected rule routing /var/log/nvidia-dcgm to 20-nvidia-platform,
alongside the DCGM binaries that actually use it, instead of letting
it fall through to base by default.

bee_layer_build folds each layer's empty-dir list into the same
rsync --files-from call; recursion into a directory that is
by-construction empty copies nothing extra. Revert the 9000/9999 hook
changes from 78d1b9b now that they're redundant, and cover the new
path with test-squashfs-layers.sh (ruled, unruled, and nested-empty
directories, asserted present in the merged rootfs after a real
mksquashfs/unsquashfs round-trip).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-12 16:19:36 +03:00
Mikhail ChusavitinandClaude Sonnet 5 b8c45d54c1 feat(iso): split the live medium into semantic SquashFS layers
Booting via BMC virtual CD reads the ~2.8 GB filesystem squashfs
sequentially during the live-boot toram copy; a mid-read drop of the
redirected medium loses the whole copy and fails the boot (v14). Split
the rootfs into self-contained semantic layers so a retry re-reads at
most one ~500-700 MiB layer, not everything. This is a resilience /
reduced-re-read mechanism, not a fix for the virtual-media instability.

NVIDIA variants now ship 7 layers (00-base, 05-firmware, 08-desktop,
10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs,
40-nvidia-dcgm-cuda) plus an explicit live/filesystem.module that fixes
their OverlayFS order; amd/nogpu keep a single squashfs.

- lib/squashfs-layers.sh: deterministic classifier (dpkg file ownership
  plus explicit rules for build.sh-injected files, never a path
  substring), per-layer mksquashfs, 800 MiB hard ceiling, unsquashfs -s
  plus strict extraction of every layer, merged-rootfs bootability check.
- build.sh: split the monolith after the full lb build, verify and merge,
  write the module file, delete the monolith only then; abort before ISO
  assembly on any failure. Runs the builder test suites up front.
- fast-path: force a full build for a multi-layer medium;
  fast_path_repack_squashfs hard-refuses (it would drop layers).
- iso-validation.sh: validate_iso_squashfs_layers (module vs layer set
  match, size ceiling, no lone giant squashfs) and
  validate_iso_media_integrity (xorriso -check_media).
- bee-install: honour filesystem.module order, abort on any layer failure.
- 9013-toram-retry: record the real rsync exit code (it printed a false
  rc=0) and correct the "resumes the tail" comment (rsync without
  --partial keeps only fully-copied layers). No unsafe partial resume.
- tests: test-squashfs-layers.sh plus a multi-layer guard in
  test-build-libs.sh; both run at the top of every build.
- docs: bible-local architecture and decision, iso/README, iso-build-rules.

Verified by a full nvidia build: 7 layers 622/199/256/466/37/567/562 MiB,
every validator passes, xorriso -check_media good, merged rootfs bootable.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-04 15:38:14 +03:00