platform/app: fix nvbandwidth split fallback, blackbox discovery churn, and add blocking sync brackets around load steps
- gpuBandwidthSocketGroups: a single GPU whose NUMA node fails to resolve no longer collapses the whole per-socket nvbandwidth split into one fallback pass — it now folds into the last resolved group instead, preserving isolation for the sockets that did resolve. - blackbox discoverMarkedTargets: skip mounting/unmounting devices that already have a running worker on every 2s discovery tick. This was observed hammering the same USB target continuously (mount+unmount every ~2s for the whole session) and contending with the worker's own sync cycle, plausibly explaining multi-minute sync cycles seen on a real crash bundle. - syncFilesystem now calls syscall.Sync() directly instead of spawning /bin/sync per copied file; blackbox mounts removable targets with -o sync so writes are durable without relying on the app-level sync as the primary mechanism. - New platform.SetSyncBracketHook / satJob.syncBracket: blocks (with a bounded timeout) on blackbox actually reaching removable media right before and right after a diagnostic's real load step (nvbandwidth, memtester, stress-ng, dcgmi diag, nccl, smartctl/nvme self-test...), instead of only firing a fire-and-forget kick after the job's own log file is written. A crash mid-load now has durable evidence the load started, not just whatever streamed to the RAM-backed export dir before blackbox's next scheduled cycle. Found investigating a real support bundle where blackbox's last successful sync (19:55:25) predated both the previous job finishing and the crashing nvbandwidth job starting (19:56:57) — none of the crash window ever reached durable media. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
6d13c17d36
commit
e03267a72f
@@ -180,6 +180,16 @@ func New(sys *platform.System) *App {
|
||||
platform.SetJobBoundaryHook(func(string) {
|
||||
requestBlackboxSync(DefaultExportDir)
|
||||
})
|
||||
// For the actual load step of a diagnostic (nvbandwidth, memtester,
|
||||
// stress-ng, dcgmi diag...) — as opposed to the cheap discovery/inventory
|
||||
// steps around it — block until blackbox has actually copied the
|
||||
// evidence that the load is about to start onto removable media, and
|
||||
// again once it finishes. Best-effort: a stuck blackbox target logs a
|
||||
// warning (see runSyncBracketHook in sat.go) rather than blocking the
|
||||
// diagnostic itself.
|
||||
platform.SetSyncBracketHook(func(jobName, phase string) error {
|
||||
return requestBlackboxSyncAndWait(DefaultExportDir, DefaultBlackboxStatePath, blackboxSyncBracketTimeout)
|
||||
})
|
||||
return a
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user