Booting via BMC virtual CD reads the ~2.8 GB filesystem squashfs sequentially during the live-boot toram copy; a mid-read drop of the redirected medium loses the whole copy and fails the boot (v14). Split the rootfs into self-contained semantic layers so a retry re-reads at most one ~500-700 MiB layer, not everything. This is a resilience / reduced-re-read mechanism, not a fix for the virtual-media instability. NVIDIA variants now ship 7 layers (00-base, 05-firmware, 08-desktop, 10-nvidia-driver, 20-nvidia-platform, 30-nvidia-cuda-libs, 40-nvidia-dcgm-cuda) plus an explicit live/filesystem.module that fixes their OverlayFS order; amd/nogpu keep a single squashfs. - lib/squashfs-layers.sh: deterministic classifier (dpkg file ownership plus explicit rules for build.sh-injected files, never a path substring), per-layer mksquashfs, 800 MiB hard ceiling, unsquashfs -s plus strict extraction of every layer, merged-rootfs bootability check. - build.sh: split the monolith after the full lb build, verify and merge, write the module file, delete the monolith only then; abort before ISO assembly on any failure. Runs the builder test suites up front. - fast-path: force a full build for a multi-layer medium; fast_path_repack_squashfs hard-refuses (it would drop layers). - iso-validation.sh: validate_iso_squashfs_layers (module vs layer set match, size ceiling, no lone giant squashfs) and validate_iso_media_integrity (xorriso -check_media). - bee-install: honour filesystem.module order, abort on any layer failure. - 9013-toram-retry: record the real rsync exit code (it printed a false rc=0) and correct the "resumes the tail" comment (rsync without --partial keeps only fully-copied layers). No unsafe partial resume. - tests: test-squashfs-layers.sh plus a multi-layer guard in test-build-libs.sh; both run at the top of every build. - docs: bible-local architecture and decision, iso/README, iso-build-rules. Verified by a full nvidia build: 7 layers 622/199/256/466/37/567/562 MiB, every validator passes, xorriso -check_media good, merged rootfs bootable. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
133 lines
4.6 KiB
Bash
Executable File
133 lines
4.6 KiB
Bash
Executable File
#!/bin/sh
|
|
# 9013-toram-retry.hook.chroot — wrap live-boot's "copy whole medium to RAM"
|
|
# rsync in a retry loop with exponential backoff.
|
|
#
|
|
# Why: on some hosts the boot medium (notably BMC / IPMI virtual CD redirection)
|
|
# drops off the bus part-way through the ~2.8 GB squashfs read. rsync then dies
|
|
# with "Input/output error (5)" and every file "has vanished", live-boot moves
|
|
# the partial RAM copy over the medium, and the boot ends with
|
|
# "No supported filesystem images found at /live" -> BOOT FAILED.
|
|
#
|
|
# live-boot's own retry is just rsync's single in-run delta retry. This patch
|
|
# adds several full attempts with a geometrically growing pause between them
|
|
# (15s, 30s, 60s, 120s, 240s, 480s, 900s), giving the virtual media time to
|
|
# re-enumerate. Between attempts the medium is unmounted, waited on, and
|
|
# remounted.
|
|
#
|
|
# rsync here runs without --partial, so it keeps only files it copied in full;
|
|
# a file interrupted mid-transfer is discarded and re-read from the start on the
|
|
# next attempt. The medium is split into semantic squashfs layers
|
|
# (filesystem-v<ver>-NN-*.squashfs, see lib/squashfs-layers.sh), so a retry only
|
|
# re-reads the layer that was in flight, not the whole ~2.8 GB rootfs. We do
|
|
# NOT enable --partial / --append: resuming a partial squashfs without a
|
|
# post-copy integrity check would risk booting a truncated layer.
|
|
set -e
|
|
|
|
TORAM_SCRIPT="/usr/lib/live/boot/9990-toram-todisk.sh"
|
|
|
|
if [ ! -f "${TORAM_SCRIPT}" ]; then
|
|
echo "9013-toram-retry: ${TORAM_SCRIPT} not found, skipping"
|
|
exit 0
|
|
fi
|
|
|
|
if grep -q "bee_rsync_retry" "${TORAM_SCRIPT}"; then
|
|
echo "9013-toram-retry: already patched, skipping"
|
|
exit 0
|
|
fi
|
|
|
|
RSYNC_LINE='rsync -a --progress ${copyfrom}/\* ${copyto} 1>/dev/console'
|
|
if ! grep -qF 'rsync -a --progress ${copyfrom}/* ${copyto} 1>/dev/console' "${TORAM_SCRIPT}"; then
|
|
echo "9013-toram-retry: expected rsync line not found in ${TORAM_SCRIPT}" >&2
|
|
echo "9013-toram-retry: live-boot layout changed -- refusing to patch blindly" >&2
|
|
exit 1
|
|
fi
|
|
|
|
echo "9013-toram-retry: patching ${TORAM_SCRIPT}"
|
|
|
|
# 1. Inject the retry helper right after the shebang.
|
|
cat > /tmp/bee_rsync_retry.func <<'FUNC'
|
|
|
|
# --- bee: retry the whole-medium RAM copy with exponential backoff ----------
|
|
bee_rsync_retry ()
|
|
{
|
|
_src="${1}"
|
|
_dst="${2}"
|
|
_dev=$(awk -v m="${_src}" '$2 == m {print $1; exit}' /proc/mounts)
|
|
_fst=$(awk -v m="${_src}" '$2 == m {print $3; exit}' /proc/mounts)
|
|
_delay=15
|
|
_maxdelay=900
|
|
_try=1
|
|
_maxtry=8
|
|
|
|
while : ; do
|
|
echo " * bee: toram copy attempt ${_try}/${_maxtry} (medium ${_dev:-?} ${_fst:-?})" 1>/dev/console
|
|
|
|
if rsync -a --progress --timeout=180 ${_src}/* ${_dst} 1>/dev/console
|
|
then
|
|
_rc=0
|
|
else
|
|
_rc=$?
|
|
fi
|
|
|
|
if [ "${_rc}" -eq 0 ]
|
|
then
|
|
echo " * bee: toram copy completed on attempt ${_try}" 1>/dev/console
|
|
return 0
|
|
fi
|
|
|
|
echo " * bee: toram copy FAILED (rsync rc=${_rc}) on attempt ${_try}" 1>/dev/console
|
|
|
|
if [ "${_try}" -ge "${_maxtry}" ]
|
|
then
|
|
echo " * bee: giving up after ${_try} attempts" 1>/dev/console
|
|
return 1
|
|
fi
|
|
|
|
echo " * bee: waiting ${_delay}s for the medium to re-settle..." 1>/dev/console
|
|
sleep "${_delay}"
|
|
|
|
# Recover the medium: unmount, wait for the device node, remount.
|
|
if [ -n "${_dev}" ]
|
|
then
|
|
umount "${_src}" 2>/dev/null || umount -l "${_src}" 2>/dev/null || true
|
|
_w=0
|
|
while [ ! -b "${_dev}" ] && [ "${_w}" -lt 120 ]
|
|
do
|
|
sleep 3
|
|
_w=$((_w + 3))
|
|
done
|
|
mount -r ${_fst:+-t ${_fst}} "${_dev}" "${_src}" 2>/dev/null \
|
|
|| mount -r "${_dev}" "${_src}" 2>/dev/null \
|
|
|| echo " * bee: remount of ${_dev} failed, will retry anyway" 1>/dev/console
|
|
fi
|
|
|
|
_try=$((_try + 1))
|
|
_delay=$((_delay * 2))
|
|
[ "${_delay}" -gt "${_maxdelay}" ] && _delay="${_maxdelay}"
|
|
done
|
|
}
|
|
FUNC
|
|
|
|
sed -i "2r /tmp/bee_rsync_retry.func" "${TORAM_SCRIPT}"
|
|
rm -f /tmp/bee_rsync_retry.func
|
|
|
|
# 2. Route the whole-medium copy through the helper and honour its exit code
|
|
# (upstream ignores rsync's exit status and moves the partial copy anyway).
|
|
sed -i \
|
|
's#\([[:space:]]*\)rsync -a --progress ${copyfrom}/[*] ${copyto} 1>/dev/console.*#\1bee_rsync_retry "${copyfrom}" "${copyto}" || { log_warning_msg "bee: toram copy failed after all retries"; rmdir "${copyto}" 2>/dev/null || true; return 1; }#' \
|
|
"${TORAM_SCRIPT}"
|
|
|
|
if ! grep -q "bee_rsync_retry \"\${copyfrom}\"" "${TORAM_SCRIPT}"; then
|
|
echo "9013-toram-retry: call-site substitution failed" >&2
|
|
exit 1
|
|
fi
|
|
|
|
echo "9013-toram-retry: patch applied"
|
|
grep -n "bee_rsync_retry" "${TORAM_SCRIPT}"
|
|
|
|
# 3. Rebuild the initramfs so the patched script lands in initrd.img.
|
|
KVER=$(ls /lib/modules | sort -V | tail -1)
|
|
echo "9013-toram-retry: rebuilding initramfs for kernel ${KVER}"
|
|
update-initramfs -u -k "${KVER}"
|
|
echo "9013-toram-retry: done"
|