Files
logpile/bible-local/docs/nvidia-bug-report-analysis.md
Mikhail ChusavitinandClaude Sonnet 5 a430a8c46f feat(nvidia_bug_report): extract Xid/SXid GPU error events
Parse Xid and SXid entries from the kernel log into GPU error events with
severity and the decoded failure reason.

See ADL-055.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-27 17:05:52 +03:00

4.5 KiB

Analyzing nvidia-bug-report.log.gz — agent instructions

Short checklist for an AI agent (or a human) given a nvidia-bug-report-*.log.gz and asked "what's wrong with this box."

1. Structured inventory pass (LOGPile)

Run it through LOGPile's nvidia_bug_report parser first — it gives board/CPU/memory/ GPU/NIC/PSU inventory as structured JSON (internal/parser/vendors/nvidia_bug_report/). As of this session the decompressor no longer hard-caps at 50MB (streams with a gzip-bomb ratio guard instead — see internal/parser/archive.go), so large dumps come through whole. Known gap: PSU present flag was buggy (fixed), sensors/FRU are legitimately empty for this format (no ipmitool sdr/fru in host-side bug reports).

2. Grep passes on the raw decompressed log

In priority order:

  1. Xid (PCI: — GPU error reports. Note all PCI:xxxx addresses and Xid codes, count occurrences, get first/last timestamp per GPU.
  2. SXid — NVSwitch (fabric manager) errors on HGX/NVSwitch systems.
  3. GPU recovery action changed — tracks a GPU through None → Drain and Reset → GPU Reset Required — a GPU stuck at "Reset Required" for a long span is a live incident, not a transient blip.
  4. fell off the bus, Uncorrectable, Double Bit ECC, Row Remap — hardware memory/PCIe failures.
  5. Thermal/power: thermal slowdown, power slowdown, Reason: SW\|HW.
  6. nvidia-smi -q section per GPU: ECC Errors, Retired Pages, Remapped Rows, GPU Recovery Action — cross-check current (not historical) state at capture time.
  7. segfault, kernel panics, Call Trace, MCE/machine-check in dmesg/journalctl excerpts.

Cross-reference: an Xid burst that stops before the nvidia-smi -q snapshot and shows GPU Recovery Action: None + clean ECC counters at the end means the box already self-recovered (e.g. via reboot) — still worth flagging, but not an active fire.

3. Classify each Xid, don't just list it

Look it up in NVIDIA's Xid catalog (link below) before writing a verdict. Same code can mean different things depending on class:

  • App-level GPU exception (bad kernel launch, illegal memory access) — usually not hardware-actionable, correlate with the offending pid/process name in the log.
  • NVLink/fabric fault (MSE Degraded, link training failures) — check knvlink* messages, Fabric Manager logs — often requires GPU/node reset, can recur → RMA candidate if repeated.
  • ECC/memory retirement (uncorrectable DRAM, row-remap failures) — RMA candidate, check Retired Pages / Remapped Rows counts.
  • XID 79 (GPU fell off the bus) — drain the node, treat as hardware failure.

4. Escalation bar

Recommend RMA/hardware escalation when: repeated fatal Xid/SXid on the same GPU across multiple days, non-zero uncorrectable ECC or remap failures, or a GPU that required manual reset more than once. Otherwise: flag as informational / correlate with workload.

What already exists online (recipes found 2026-08-26)

No public "AI agent prompt/recipe" specifically for feeding this file to an LLM was found — the closest things are the shell-script scanners above; an agent should treat their grep pattern lists as a starting checklist (§2 above already folds them in).