19 KiB
PCIe Gen1-at-idle GPU warning: history of attempts, and the fix
Date: 2026-08-24
Status: superseded in part by 2026-09-03-pcie-link-verdict-under-real-traffic.md
Symptom
On NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, reanimator.json /
status/component-status.json (pcie:gpu:nvidia) report Warning:
"PCIe link speed degraded: running at Gen1, capable of Gen5" for every GPU in
the system, for the entire diagnostic session — even though every functional
test that actually exercises the GPUs (dcgmi diag targeted-stress/targeted-power/
pulse_test, nvbandwidth, NCCL all_reduce_perf, GPU config check) passes clean,
link width stays x16/x16 throughout, and there is no AER/Xid activity in
dmesg. Concretely observed in the blackbox 2026-08-13 (BEE-SP v12.84) NF5468-M7-A0-R0-00 28CC05483.
This is not a one-off — five separate commits over five months have targeted this exact false-positive-vs-real-fault ambiguity, and it is still not fully solved. This doc is the record of what was tried and why each attempt only closed part of the gap, so we stop re-discovering the same dead ends.
Timeline
| Date | Commit | What it did | Gap it left |
|---|---|---|---|
| 2026-04-01 | eb60100 |
NVIDIA collector switched LinkSpeed/MaxLinkSpeed from raw sysfs (current_link_speed) to nvidia-smi --query-gpu=pcie.link.gen.current,..., on the theory that "the driver knows the negotiated speed regardless of current power state" and sysfs reflects only instantaneous physical state. |
False premise, confirmed wrong by this very bundle. export/gpu/nvidia-smi-q.txt (the actual nvidia-smi -q dump, not sysfs) shows PCIe Generation / Device Current: 1 for every GPU. nvidia-smi's own query reflects the same ASPM/power-managed downshift as sysfs on driver 580.159.03 + Blackwell — switching the data source changed nothing for this failure mode. |
| 2026-04-02 | 99cece5 |
Added lspci -vvv and per-GPU sysfs link files to the on-demand support bundle (export/gpu/) so a human could manually cross-check LnkCap/LnkSta. |
Diagnostic aid only, no automated verdict. Also — see below — this file only ships when someone clicks "Download Support Bundle" in the web UI; it never reaches a blackbox capture. |
| 2026-04-12 | 05c1fde |
Added applyPCIeLinkSpeedWarning: sets Warning (or Critical for NVLink bridges) whenever LinkSpeed < MaxLinkSpeed, using whatever LinkSpeed was populated by (at that point) eb60100's nvidia-smi values. This is the commit that actually introduced the warning we're chasing. |
Compares "current" to "max" as an absolute rule with no allowance for idle power states — this is the root of the false-positive on any idle GPU. |
| 2026-04-12 | 4f94ebc |
Same day: added pcie_aspm=off, intel_idle.max_cstate=1, processor.max_cstate=1 to the boot kernel command line, as an OS-level attempt to stop links from ever downshifting in the first place. |
Confirmed present but ineffective for this failure mode. This exact bundle's dmesg.txt shows pcie_aspm=off ... intel_idle.max_cstate=1 processor.max_cstate=1 active in the kernel command line — and the GPUs still trained down to Gen1 at idle anyway. pcie_aspm=off disables the platform's standard PCIe ASPM (L0s/L1) link states; it does not touch the GPU driver's own independent runtime power management, which downclocks the PCIe link as part of the GPU's P-state machine. Two different mechanisms, only one of which this flag controls. |
| 2026-06-12 | 2320925 |
Excluded permanently-disabled PCIe devices (e.g. Switchtec fabric-management endpoints on HGX H100 baseboards) from the warning entirely, via sysfs enable==0. Documented in 2026-06-12-pcie-disabled-device-link-warning.md. |
Fixed a different false-positive (management-plane chips, not GPUs at idle). Doesn't touch this bug. |
| 2026-08-04 | a34e823 |
Fixed enrichPCIeWithNVIDIAData unconditionally clobbering dev.Status after applyPCIeLinkSpeedWarning had already set Warning — a later NVIDIA-data enrichment pass was silently downgrading it back to OK while leaving the stale ErrorDescription behind. Added severity-ordered merge (OK/Unknown < Warning < Critical). |
Made the Warning stick reliably in the hardware snapshot (reanimator.json) for the first time — before this, it could randomly vanish depending on collector pass ordering. Necessary fix, but it also means the false-positive from 05c1fde now survives more reliably than before. |
| 2026-08-06 | b1f165e |
Two things: (1) writePCIeGPUStatusesToDB — pushes the collector's PCIe status into component-status.json every audit cycle, because none of the SAT jobs (nvidia, nvidia-config, nvidia-interconnect, nvidia-bandwidth) ever checked link speed, so the DB-backed status (what the web UI "Hardware Summary" chip and status/component-status.json in every bundle read) had been silently showing OK regardless of the collector's own Warning. (2) Added export/gpu/pcie-nvidia-link-under-load.txt to the support bundle: runs bee-gpu-burn --seconds 8 and resamples the same sysfs link attributes mid-load, so a human/agent can tell idle power-saving from a real degraded slot/riser by comparing it against the idle pcie-nvidia-link.txt. |
(1) closed the "SAT says OK" mismatch — this is why status/component-status.json in the 08-13 bundle correctly shows Warning and doesn't get silently overwritten by the string of sat:nvidia*: OK entries in its history. (2) is the one piece of tooling actually designed to answer "is this real or just idle" — and it is absent from this bundle (see next section). Also: component_status_db.go's Record() merge (newSev > curSev) has no downgrade path — once Warning is recorded, nothing (not even a later audit:pcie poll reporting the link back at Gen5) can lower it back to OK within that DB file's lifetime. Not new in this commit, but this is the mechanism that pins the warning for the rest of the session even if the link genuinely retrains under load later. |
| 2026-08-17 | e3697c0 |
Fixed matchesGPUVendor/isGPUDevice matching any PCIe device with "Controller" in its class string or same-vendor ID as a GPU — same-vendor NICs/NVSwitch bridges could trip the pcie:gpu:<vendor> alarm. |
Correctness fix for which devices get judged, not for the idle-vs-real ambiguity itself. |
| 2026-08-24 (today) | 9add561 |
Bounded every support_bundle.go subprocess (including the bee-gpu-burn-driven under-load capture from b1f165e) with a timeout, because a genuinely wedged GPU could hang bundle generation forever. |
Unrelated to the false-positive itself, but relevant context: confirms the under-load capture is a live, still-evolving code path, not dead code. |
Why this specific bundle still shows the confusing state
Two independent things line up to explain exactly what you're looking at:
-
The blackbox (USB auto-sync) capture path never runs
b1f165e's disambiguation tooling at all, regardless of build freshness.pcie-nvidia-link.txt/pcie-nvidia-link-under-load.txtlive insupport_bundle.go'ssupportBundleCommands, which only executes insideBuildSupportBundle— triggered by a human clicking "Download Support Bundle" in the web UI. The blackbox worker (blackbox.gosyncCycle/captureSnapshots) instead callscategorizeExportTreeon whatever the periodic collector already staged in the live export directory, plus journalctl/dmesg/status snapshots — it does not invokesupportBundleCommands. That's also why this bundle has nomanifest.txt:blackbox.go:604notes the blackbox path intentionally skips it, shipping onlyREADME.mdas the "how to read this" doc. Net effect: the one artifact designed to answer "idle or real fault" for a GPU PCIe warning is structurally unreachable from a blackbox pull. You would only get it by opening the web UI and downloading a support bundle from that same host while the warning is still active — not useful for a server that already shipped or is already offline. -
Even where the fixes did land, they don't address the actual mechanism. This bundle's own
dmesg.txtprovespcie_aspm=offand both max_cstate=1 flags from4f94ebcwere active at boot, and the GPU still trained to Gen1 at idle — andnvidia-smi-q.txtproveseb60100's switch to nvidia-smi's own query (instead of sysfs) reports the exact same Gen1 reading. Both mitigations were built on an assumption (ASPM is the mechanism / nvidia-smi reports negotiated capability, not current power state) that this hardware+driver combination (RTX PRO 6000 Blackwell Server Edition, driver 580.159.03) disproves. The actual mechanism is the NVIDIA driver's own runtime power management downclocking the link independent of platform ASPM — nothing in the current codebase distinguishes "GPU driver decided to save power" from "riser/slot is actually degraded." -
Separately,
component_status_db.go'sRecord()has no downgrade path (seeb1f165erow above): once any source writesWarningforpcie:gpu:nvidia, that status cannot go back toOKwithin the same session even if a later poll of the same source reports the link back at full speed. Combined with #2, a single idle-time sample taken at 14:20:05 (before any load test ran) permanently pins the whole GPU subsystem's status for the rest of that ~9-day diagnostic session (14:21 → 08-17 10:03), regardless of how many load tests pass clean in between.
Options going forward
Pick one (or combine):
- Make the under-load resample part of the periodic collector /
blackbox path, not support-bundle-only — e.g. run it once per boot
from
bee-audit/bee-nvidiaright after the SAT GPU load tests (nvidia,nvidia-bandwidth,nvidia-targeted-stressalready load the GPUs for minutes at a time — piggyback the sysfs resample on one of those instead of a dedicated 8sbee-gpu-burnrun) and write it into the live export tree so blackbox mirrors it automatically. - Let a real load test clear the Warning. Have one of the GPU SAT
jobs that's already running a sustained load (
nvidia-bandwidth,nvidia-targeted-stress) resample link speed at the end of its run and callwritePCIeGPUStatusesToDBwith the fresh reading, allowing an explicit downgrade path for this one key when the same mechanism that raised the warning reports it resolved under controlled load — rather than opening upRecord()'s merge logic in general (which is deliberately sticky for a reason: don't want a real intermittent PSU/ECC fault to be silently forgotten because one later poll came back clean). - Stop comparing against idle sysfs/nvidia-smi state at all for NVIDIA GPUs specifically, and instead rely on the SAT bandwidth numbers themselves (nvbandwidth's measured GB/s vs. expected-for-Gen5-x16 threshold) as the actual link-health signal — this sidesteps the idle-vs-load ambiguity entirely since it measures the thing you actually care about (does the link deliver Gen5 throughput when asked), not a point-in-time speed field.
- Leave detection as-is, fix only the messaging: keep flagging Gen1 at
idle as
Warning, but make theerror_descriptionexplicit that this is unconfirmed/idle-sampled ("Warning: idle PCIe link at Gen1 (max Gen5, unconfirmed under load)") so a human readingreanimator.jsoncold isn't misled into thinking it's a proven hardware fault — closest to a documentation-only fix, cheapest, but keeps the ambiguity forever.
Resolution (2026-08-24)
Historical note: the two SAT targets described below were the resolution at the time. The 2026-09-03 follow-up removes the forced-retrain target and folds the GPU link verdict into the existing bandwidth SAT after field evidence showed that retraining without traffic can remain at Gen1.
Landed a variant of options 1–3 that turned out simpler than any of them individually once we stopped trying to make the idle reading recoverable:
Stop writing a status from the idle reading at all. parseLspciDevice
(collector/pcie.go) no longer calls applyPCIeLinkSpeedWarning on every
idle collector pass — LinkSpeed/MaxLinkSpeed stay populated as plain
descriptive fields in reanimator.json, but nothing sets Status to
Warning from them anymore. This makes Record()'s one-way severity merge
(no downgrade path, see the b1f165e row above) a non-issue by
construction: if nothing ever writes an unverified Warning, there is
nothing that later needs downgrading. applyPCIeLinkSpeedWarning itself is
kept, now documented as reserved for a caller that already put the device
under real traffic.
Two new verified-load SAT targets replace the idle signal:
pcie-link(platform/pcie_link_check.go) — forces every enabled PCIe device (not just GPUs) to retrain via the PCIe spec's Link Control "Retrain Link" bit (setpci … CAP_EXP+0x10.w, poll Link Status bit 11 until training clears), then compares the post-retrain negotiated speed against the device's own max. This is the generic answer to "what about non-GPU PCIe cards" from the prior discussion: retraining is a mechanism every PCIe endpoint supports, so one check now covers NICs/HBAs/switches that have nobee-gpu-burn-equivalent load tool. Classifies devices by PCI class code (0x03= Display) + vendor ID (0x10de/0x1002), not name substrings, per the existingno-hardcoded-vendorscontract. Disabled devices (enable==0) are left alone, same carve-out as the 2026-06-12 decision. Writesgpu_nvidia_status/gpu_amd_status/other_statusintosummary.txt;ApplySATResultToDBroutes each into its own component key (pcie:gpu:nvidia,pcie:gpu:amd,pcie:link:other) so a degraded NIC is never reported as a GPU fault.nvidia-pcie-bandwidth(platform/nvidia_pcie_bandwidth.go) — GPU-only, drivesdcgmi diag -r nvbandwidth(the same toolnvidia-bandwidthalready uses for P2P throughput) and resamples each involved GPU's link speed via sysfs immediately after, deliberately independent of nvbandwidth's own GB/s pass/fail — this SAT's verdict is purely "did the link train up to max under real traffic." Feedspcie:gpu:nvidiaalongsidepcie-link.
Both are wired into the existing task queue/webui exactly like
nvidia-config (routes in server.go, dispatch case in task_runner.go,
priority in api.go's defaultTaskPriority, cards on the Validate "Check"
page, included in "Run All Check SAT").
Open follow-up, not yet built: neither target runs automatically today
— a Warning only clears/confirms when an operator runs one of these two
SATs (or "Run All Check SAT", which now includes pcie-link). If that
turns out to matter in practice, wire pcie-link (it's fast — a retrain is
milliseconds per device, not a sustained burn) into bee-audit's periodic
cycle.
Follow-up (2026-08-24, same day): the original artifact never reached blackbox at all
Separately from the SAT work above, re-examined why pcie-nvidia-link.txt /
pcie-nvidia-link-under-load.txt (the b1f165e diagnostic pair) were
missing from the blackbox analyzed earlier in this doc, even on a build new
enough to have that code. Root cause, confirmed against the actual pipeline
(not build/version drift as first guessed): those two files only ever lived
in supportBundleCommands (app/support_bundle.go), which exclusively
runs inside BuildSupportBundle — the on-demand "Download Support Bundle"
web UI action. The blackbox USB auto-sync worker (blackbox.go
syncCycle) never calls that function; it only mirrors whatever
platform.CaptureTechnicalDump already wrote into the live export tree at
boot (bee-audit.service, Type=oneshot) via categorizeExportTree. So
the artifact was structurally unreachable from a blackbox pull no matter
how fresh the ISO was — this was the real explanation for the earlier
"missing file" mystery, not the stale-build hypothesis floated above.
Fix: moved both scripts (as shared exported constants,
platform.PCIeNvidiaLinkScript / PCIeNvidiaLinkUnderLoadScript, so
support-bundle and techdump run byte-identical scripts instead of two
copies that can drift) into platform.techDumpNvidiaCommands, so
CaptureTechnicalDump captures them once at boot alongside the existing
nvidia-smi-* dumps. bee-audit is oneshot, so the one-time ~8s
bee-gpu-burn cost for the under-load sample is a boot-time cost, not a
per-blackbox-sync-cycle one. Added both filenames to techdumpBucketFor
(→ export/gpu/) so categorizeExportTree buckets them correctly.
supportBundleCommands still re-runs both on-demand for a fresher sample
than the boot-time one — that's intentional, not a duplicate to clean up.
Same-pattern audit of the rest of supportBundleCommands: while in
there, checked every other export/* entry (the ones the README
tags as hardware-facing, i.e. everything except livecd/*) for the same
"only reachable via on-demand support bundle, never blackbox" gap. Also
fixed export/gpu/kernel-aer-nvidia.txt (dmesg filtered for
AER/NVRM/Xid lines) the same way — same file, same fix, now
platform.KernelAERNvidiaScript shared between both paths. This one stood
out because it's the exact file I reached for by hand (grepping raw
dmesg.txt myself) when first triaging the blackbox that started this
whole investigation — its absence wasn't hypothetical.
Found but not changed — flagging for a decision, not fixed unilaterally, since promoting more of these turns boot-time audit into a slower fixed cost for every boot, not just support-bundle downloads:
export/platform/lspci-nn.txt(lspci -nn) — cheap, static, would be a trivial add.export/gpu/lspci-video-vv.txt,export/gpu/lspci-nvidia-bridges-vv.txt— cheap-ish, somewhat redundant withlspci-vvv.txt(already in techdump) but with NVIDIA-specific bridge-chain framing that's genuinely easier to read.export/gpu/systemctl-nvidia-units.txt,export/gpu/dcgmi-nvlink-status.txt,export/gpu/fabric-manager-paths.txt— cheap, static.export/gpu/nvidia-smi-topo-fresh.txt/-nvlink-status-fresh.txt/-nvlink-errors-fresh.txt— deliberately not a gap: their entire purpose is being a fresher resample than the boot-timenvidia-smi-topo.txtetc. that techdump already captures and blackbox already mirrors; only useful as an on-demand recapture.export/gpu/nvidia-bug-report.txt(30–120s vianvidia-bug-report.sh) and the networkexport/network/ethtool-*/mstflint-query.txtentries — left alone; meaningfully heavier or more device-count-dependent than the others, worth a deliberate cost/benefit call rather than folding in by default.