-
platform/app: fix nvbandwidth split fallback, blackbox discovery churn, and add blocking sync brackets around load steps
released this
2026-07-28 14:22:05 +03:00 - gpuBandwidthSocketGroups: a single GPU whose NUMA node fails to resolve
no longer collapses the whole per-socket nvbandwidth split into one
fallback pass — it now folds into the last resolved group instead,
preserving isolation for the sockets that did resolve. - blackbox discoverMarkedTargets: skip mounting/unmounting devices that
already have a running worker on every 2s discovery tick. This was
observed hammering the same USB target continuously (mount+unmount
every ~2s for the whole session) and contending with the worker's own
sync cycle, plausibly explaining multi-minute sync cycles seen on a
real crash bundle. - syncFilesystem now calls syscall.Sync() directly instead of spawning
/bin/sync per copied file; blackbox mounts removable targets with
-o sync so writes are durable without relying on the app-level sync as
the primary mechanism. - New platform.SetSyncBracketHook / satJob.syncBracket: blocks (with a
bounded timeout) on blackbox actually reaching removable media right
before and right after a diagnostic's real load step (nvbandwidth,
memtester, stress-ng, dcgmi diag, nccl, smartctl/nvme self-test...),
instead of only firing a fire-and-forget kick after the job's own log
file is written. A crash mid-load now has durable evidence the load
started, not just whatever streamed to the RAM-backed export dir before
blackbox's next scheduled cycle.
Found investigating a real support bundle where blackbox's last
successful sync (19:55:25) predated both the previous job finishing and
the crashing nvbandwidth job starting (19:56:57) — none of the crash
window ever reached durable media.Co-Authored-By: Claude Sonnet 5 noreply@anthropic.com
Downloads
- gpuBandwidthSocketGroups: a single GPU whose NUMA node fails to resolve