• v12.69 e03267a72f

    platform/app: fix nvbandwidth split fallback, blackbox discovery churn, and add blocking sync brackets around load steps

    mchus released this 2026-07-28 14:22:05 +03:00

    • gpuBandwidthSocketGroups: a single GPU whose NUMA node fails to resolve
      no longer collapses the whole per-socket nvbandwidth split into one
      fallback pass — it now folds into the last resolved group instead,
      preserving isolation for the sockets that did resolve.
    • blackbox discoverMarkedTargets: skip mounting/unmounting devices that
      already have a running worker on every 2s discovery tick. This was
      observed hammering the same USB target continuously (mount+unmount
      every ~2s for the whole session) and contending with the worker's own
      sync cycle, plausibly explaining multi-minute sync cycles seen on a
      real crash bundle.
    • syncFilesystem now calls syscall.Sync() directly instead of spawning
      /bin/sync per copied file; blackbox mounts removable targets with
      -o sync so writes are durable without relying on the app-level sync as
      the primary mechanism.
    • New platform.SetSyncBracketHook / satJob.syncBracket: blocks (with a
      bounded timeout) on blackbox actually reaching removable media right
      before and right after a diagnostic's real load step (nvbandwidth,
      memtester, stress-ng, dcgmi diag, nccl, smartctl/nvme self-test...),
      instead of only firing a fire-and-forget kick after the job's own log
      file is written. A crash mid-load now has durable evidence the load
      started, not just whatever streamed to the RAM-backed export dir before
      blackbox's next scheduled cycle.

    Found investigating a real support bundle where blackbox's last
    successful sync (19:55:25) predated both the previous job finishing and
    the crashing nvbandwidth job starting (19:56:57) — none of the crash
    window ever reached durable media.

    Co-Authored-By: Claude Sonnet 5 noreply@anthropic.com

    Downloads