Files
bee/bible-local/decisions/2026-08-31-nvbandwidth-validate-single-deep-matrix.md

2.3 KiB

nvbandwidth: single all-GPU pass in Validate, per-NUMA-node matrix only in the deep tier

Date: 2026-08-31 Status: active Amends: 2026-07-27-nvbandwidth-per-socket-split.md

Context

2026-07-27 made RunNvidiaBandwidthPack split dcgmi diag -r nvbandwidth using Linux PCI numa_node locality, followed by an all-GPU pass. A NUMA node is not assumed to be identical to a physical CPU socket.

Two problems showed up on an 8x H200 NVL / dual-socket EPYC box (210619KUGGXGS2000008):

  1. SATEstimatedNvidiaBandwidthSec assigns 2700 seconds to one all-GPU pass. Three such invocations are assigned 8100 seconds, which is outside the intended Validate duration.
  2. The split silently never engaged anyway: normalizeNvidiaBDF returned nvidia-smi's upper-case PCI BDF (0000:CB:00.0) while /sys/bus/pci/devices entries are lower-case, so readPCINumaNode failed for every GPU on a bus with a hex letter and gpuBandwidthSocketGroups fell back to one group.

Decision

  • normalizeNvidiaBDF now lower-cases (and trims) the BDF, so sysfs numa_node / link-speed reads actually resolve.
  • RunNvidiaBandwidthPack takes a fullMatrix bool. The only thing it changes is which GPU set each nvbandwidth invocation gets via -i - the command itself is untouched (no extra flags, no testcase filtering).
    • Validate (stress_mode=false) -> fullMatrix=false: one pass, -i <all selected GPUs>. No NUMA-locality split.
    • Stress / deep (stress_mode=true) -> fullMatrix=true: the 2026-07-27 behaviour - one pass per resolved NUMA-node group, then one all-GPU pass. If any selected GPU has no resolved NUMA node, the code does not guess a group and falls back to the single all-GPU pass.
  • task_runner passes t.params.StressMode through.

Consequences

  • Validate schedules exactly one nvbandwidth invocation regardless of NUMA-node count.
  • The NUMA-locality isolation from 2026-07-27 is preserved, just moved to the tier where a 2-3x runtime is acceptable. DO NOT re-add the split to the Validate path.
  • SATEstimatedNvidiaBandwidthSec still reflects a single pass; the deep tier's multi-pass duration is not modelled - revisit once real multi-node deep-run logs exist (same open item as 2026-07-27).