dcgmi discovery -l is a preflight/metadata step ahead of the real DCGM diag jobs; a transient failure racing nv-hostengine startup shouldn't flip the whole pack's status, so it's now marked informational with a couple of retries. Separately, bound the fabricmanager/nvidia-dcgm systemctl restart/start calls in bee-nvidia-load with a timeout so a wedged unit (e.g. fabric training stuck on a bad NVSwitch fabric) can't hang bee-nvidia.service forever and block dcgm from ever starting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>