fix(nvidia): load ib_umad and install nvlsm/libnvidia-nscq for NVL5 fabric
The NVSwitch-detection fix (983f41a) was necessary but not sufficient on
HGX B200 (Kaytus KR9288-X3): ExecCondition now correctly runs fabricmanager,
but nvidia-fabricmanager-start.sh then aborts with "Kernel module ib_umad
has not been loaded" on this NVL5+ board, leaving the fabric stuck in
"In Progress" and cascading into failed dcgmi diag/NVBandwidth/NCCL/Validate
GPU SAT tasks (confirmed via support bundle 20260729-165331).
Neither nvidia-fabricmanager nor DCGM declare the NVLink5 support packages
as apt dependencies (checked against the cuda-repo Packages index directly):
ib_umad must be modprobed before FM starts, nvlsm (InfiniBand Subnet Manager)
is launched internally by FM's own start script once present, and DCGM
dlopens libnvidia-nscq at runtime for NVSwitch health monitoring.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
983f41a1f0
commit
6c8be629d2
@@ -7,3 +7,10 @@ After=bee-nvidia.service
|
||||
# Skip fabricmanager on systems without NVSwitch hardware.
|
||||
# ExecCondition exits 1-254 → unit is silently skipped (inactive, not failed).
|
||||
ExecCondition=/usr/local/bin/bee-check-nvswitch
|
||||
|
||||
# On NVL5+ systems (e.g. HGX B200) nvidia-fabricmanager-start.sh polls for an
|
||||
# Infiniband device before starting FM and refuses to start if "ib_umad" isn't
|
||||
# loaded, even though mlx5_ib/ib_core/ib_uverbs already are — nothing else on
|
||||
# the live ISO triggers its autoload. Load it explicitly; "-" keeps this from
|
||||
# failing the unit on older HGX generations without an ib_umad module.
|
||||
ExecStartPre=-/sbin/modprobe ib_umad
|
||||
|
||||
Reference in New Issue
Block a user