Files
logpile/bible-local/docs/nvidia-bug-report-analysis.md
Mikhail ChusavitinandClaude Sonnet 5 a430a8c46f feat(nvidia_bug_report): extract Xid/SXid GPU error events
Parse Xid and SXid entries from the kernel log into GPU error events with
severity and the decoded failure reason.

See ADL-055.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-27 17:05:52 +03:00

78 lines
4.5 KiB
Markdown

# Analyzing nvidia-bug-report.log.gz — agent instructions
Short checklist for an AI agent (or a human) given a `nvidia-bug-report-*.log.gz` and
asked "what's wrong with this box."
## 1. Structured inventory pass (LOGPile)
Run it through LOGPile's `nvidia_bug_report` parser first — it gives board/CPU/memory/
GPU/NIC/PSU inventory as structured JSON (`internal/parser/vendors/nvidia_bug_report/`).
As of this session the decompressor no longer hard-caps at 50MB (streams with a
gzip-bomb ratio guard instead — see `internal/parser/archive.go`), so large dumps
come through whole. Known gap: PSU `present` flag was buggy (fixed), sensors/FRU are
legitimately empty for this format (no `ipmitool sdr`/`fru` in host-side bug reports).
## 2. Grep passes on the raw decompressed log
In priority order:
1. `Xid (PCI:` — GPU error reports. Note **all** `PCI:xxxx` addresses and Xid codes,
count occurrences, get first/last timestamp per GPU.
2. `SXid` — NVSwitch (fabric manager) errors on HGX/NVSwitch systems.
3. `GPU recovery action changed` — tracks a GPU through `None → Drain and Reset →
GPU Reset Required` — a GPU stuck at "Reset Required" for a long span is a live
incident, not a transient blip.
4. `fell off the bus`, `Uncorrectable`, `Double Bit ECC`, `Row Remap` — hardware
memory/PCIe failures.
5. Thermal/power: `thermal slowdown`, `power slowdown`, `Reason: SW\|HW`.
6. `nvidia-smi -q` section per GPU: `ECC Errors`, `Retired Pages`, `Remapped Rows`,
`GPU Recovery Action` — cross-check current (not historical) state at capture time.
7. `segfault`, kernel panics, `Call Trace`, `MCE`/machine-check in dmesg/journalctl
excerpts.
Cross-reference: an Xid burst that stops before the `nvidia-smi -q` snapshot and shows
`GPU Recovery Action: None` + clean ECC counters at the end means the box already
self-recovered (e.g. via reboot) — still worth flagging, but not an active fire.
## 3. Classify each Xid, don't just list it
Look it up in NVIDIA's Xid catalog (link below) before writing a verdict. Same code
can mean different things depending on class:
- App-level GPU exception (bad kernel launch, illegal memory access) — usually not
hardware-actionable, correlate with the offending `pid`/process name in the log.
- NVLink/fabric fault (`MSE Degraded`, link training failures) — check
`knvlink*` messages, Fabric Manager logs — often requires GPU/node reset, can
recur → RMA candidate if repeated.
- ECC/memory retirement (uncorrectable DRAM, row-remap failures) — RMA candidate,
check `Retired Pages` / `Remapped Rows` counts.
- `XID 79` (GPU fell off the bus) — drain the node, treat as hardware failure.
## 4. Escalation bar
Recommend RMA/hardware escalation when: repeated fatal Xid/SXid on the same GPU across
multiple days, non-zero uncorrectable ECC or remap failures, or a GPU that required
manual reset more than once. Otherwise: flag as informational / correlate with workload.
## What already exists online (recipes found 2026-08-26)
- **Lambda Labs — `check-nvidia-bug-report.sh`**: an actual automated scanner for this
exact file, open source. Clones and runs against the extracted report; flags Xid
errors, thermal/power slowdown, segfaults, CPU throttling, "fell off the bus",
RmInit failures. This is the closest thing to a ready-made "recipe" — worth using
directly or mining its grep patterns.
https://github.com/lambdal-support/lambda-public-tools (see `check-nvidia-bug-report/`)
Guide: [Lambda Docs — Using the nvidia-bug-report.log file to troubleshoot your system](https://docs.lambda.ai/education/linux-usage/using-the-nvidia-bug-report.log-file-to-troubleshoot-your-system/)
- **NVIDIA — GPU Debug Guidelines**: official step-by-step (collect → `dcgmi diag` →
classify Xid → escalate/RMA criteria), the authoritative human-oriented recipe.
https://docs.nvidia.com/deploy/gpu-debug-guidelines/index.html
- **NVIDIA — Xid Errors reference**: the Xid code catalog itself (meanings, no fixed
hardware/software verdict per code — context-dependent).
https://docs.nvidia.com/deploy/xid-errors/introduction.html
- **AWS re:Post — Troubleshoot Xid errors in NVIDIA GPU-accelerated instances**:
cloud-instance-flavored version of the same workflow.
https://repost.aws/knowledge-center/ec2-linux-troubleshoot-xid-errors
No public "AI agent prompt/recipe" specifically for feeding this file to an LLM was
found — the closest things are the shell-script scanners above; an agent should treat
their grep pattern lists as a starting checklist (§2 above already folds them in).