nvidia-interconnect (NCCL all_reduce_perf) and nvidia-bandwidth (NVBandwidth) verify fabric connectivity and bandwidth — they are not sustained burn loads. Move both from the Burn section to the Validate section under the stress-mode toggle, alongside the other DCGM diagnostic tests moved in the previous commit. - Add sat-card-nvidia-interconnect and sat-card-nvidia-bandwidth validate cards (stress-only, all selected GPUs at once) - Add runNvidiaFabricValidate() for all-GPU-at-once dispatch - Add nvidiaAllGPUTargets handling in expandSATTarget/runAllSAT - Remove Interconnect / Bandwidth card from Burn section - Remove nvidia-interconnect and nvidia-bandwidth from runAllBurnTasks and the gpu/tools availability map Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
146 KiB
146 KiB