platform/app: fix nvbandwidth split fallback, blackbox discovery churn, and add blocking sync brackets around load steps
- gpuBandwidthSocketGroups: a single GPU whose NUMA node fails to resolve no longer collapses the whole per-socket nvbandwidth split into one fallback pass — it now folds into the last resolved group instead, preserving isolation for the sockets that did resolve. - blackbox discoverMarkedTargets: skip mounting/unmounting devices that already have a running worker on every 2s discovery tick. This was observed hammering the same USB target continuously (mount+unmount every ~2s for the whole session) and contending with the worker's own sync cycle, plausibly explaining multi-minute sync cycles seen on a real crash bundle. - syncFilesystem now calls syscall.Sync() directly instead of spawning /bin/sync per copied file; blackbox mounts removable targets with -o sync so writes are durable without relying on the app-level sync as the primary mechanism. - New platform.SetSyncBracketHook / satJob.syncBracket: blocks (with a bounded timeout) on blackbox actually reaching removable media right before and right after a diagnostic's real load step (nvbandwidth, memtester, stress-ng, dcgmi diag, nccl, smartctl/nvme self-test...), instead of only firing a fire-and-forget kick after the job's own log file is written. A crash mid-load now has durable evidence the load started, not just whatever streamed to the RAM-backed export dir before blackbox's next scheduled cycle. Found investigating a real support bundle where blackbox's last successful sync (19:55:25) predated both the previous job finishing and the crashing nvbandwidth job starting (19:56:57) — none of the crash window ever reached durable media. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
6d13c17d36
commit
e03267a72f
@@ -91,6 +91,33 @@ func TestGPUBandwidthSocketGroupsSplitsBySocket(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestGPUBandwidthSocketGroupsFoldsUnresolvedIntoLastGroup(t *testing.T) {
|
||||
// GPU 4's NUMA node fails to resolve (e.g. a flaky sysfs read), but the
|
||||
// other 5 GPUs still clearly span two sockets — the split should survive
|
||||
// and GPU 4 should ride along with the last group rather than being
|
||||
// tested alone or collapsing the whole thing to one pass.
|
||||
fakeNvidiaSmiBusIDs(t, "0, 00000000:05:00.0\n1, 00000000:06:00.0\n2, 00000000:76:00.0\n3, 00000000:77:00.0\n4, 00000000:F4:00.0\n5, 00000000:F5:00.0\n")
|
||||
fakeNUMANodes(t, map[string]string{
|
||||
"0000:05:00.0": "0\n",
|
||||
"0000:06:00.0": "0\n",
|
||||
"0000:76:00.0": "0\n",
|
||||
"0000:77:00.0": "0\n",
|
||||
// GPU 4 (F4:00.0) deliberately missing.
|
||||
"0000:F5:00.0": "1\n",
|
||||
})
|
||||
|
||||
groups := gpuBandwidthSocketGroups([]int{0, 1, 2, 3, 4, 5}, nil)
|
||||
if len(groups) != 2 {
|
||||
t.Fatalf("groups=%v want 2 groups", groups)
|
||||
}
|
||||
if joinIndexList(groups[0]) != "0,1,2,3" {
|
||||
t.Fatalf("groups[0]=%v want 0,1,2,3", groups[0])
|
||||
}
|
||||
if joinIndexList(groups[1]) != "4,5" {
|
||||
t.Fatalf("groups[1]=%v want 4,5 (unresolved GPU 4 folded into last group)", groups[1])
|
||||
}
|
||||
}
|
||||
|
||||
func TestGPUBandwidthSocketGroupsFallsBackToSingleGroup(t *testing.T) {
|
||||
t.Run("single NUMA node", func(t *testing.T) {
|
||||
fakeNvidiaSmiBusIDs(t, "0, 00000000:05:00.0\n1, 00000000:06:00.0\n")
|
||||
|
||||
Reference in New Issue
Block a user