fix(nvidia): surface Xid 79/154 GPU bus-fall-off as a physical-reboot-required signal
nvidia-bug-report.sh appends .gz to --output-file when gzip is available, which silently produced empty nvidia-bug-report.txt in support bundles (cat looked for the uncompressed name that never existed). Also: a GPU that falls off the PCIe/NVLink bus (Xid 79) or gets flagged for Node Reboot Required (Xid 154) mid-SAT-run left every downstream test failing with generic, unrelated-looking errors (CUDA "unknown error", "unable to determine device handle") with no indication the GPU needed a physical power-cycle to recover. Detect these codes from SAT run logs and surface a plain-English "physical reboot required" message in the task's failure detail, the persisted component-status DB, a dashboard banner on the Hardware Summary card, and topology diagram GPU-node severity. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
a34e823f82
commit
b2b3f86c8d
@@ -28,9 +28,10 @@ func TestXidSeverity(t *testing.T) {
|
||||
wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "xid 79 unknown code falls back to caller default",
|
||||
line: "NVRM: Xid (PCI:0000:65:00): 79, pid=1234, GPU has fallen off the bus",
|
||||
wantOK: false,
|
||||
name: "xid 79 GPU fell off the bus is critical",
|
||||
line: "NVRM: Xid (PCI:0000:65:00): 79, pid=1234, GPU has fallen off the bus",
|
||||
wantSev: "critical",
|
||||
wantOK: true,
|
||||
},
|
||||
{
|
||||
name: "no xid in line",
|
||||
|
||||
Reference in New Issue
Block a user