# GPU PCIe Test Methodology Which of the checks below run is decided by the backend (`POST /api/sat/run-all` -> `handler.planSATRunAll`), not the web page: it enumerates the hardware, waits for the NVIDIA driver to finish enumerating GPUs, and enqueues only the applicable checks. The TPM check is planned only when sysfs reports a TPM 2.x device (`tpm_version_major=2`). See `bible-local/decisions/2026-08-31-backend-driven-sat-planning.md`. ## Validate - CPU check - `lscpu` - `sensors` - `stress-ng` - Memory check - `free` - `timeout memtester` - `free` - NVMe storage check - `nvme id-ctrl` - `nvme smart-log` - `nvme device-self-test` - SATA/SAS storage check - `smartctl -H -A` - `smartctl -t short` - Basic NVIDIA GPU check - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dmidecode -t baseboard` - `dmidecode -t system` - `dcgmi diag -r 2` - Inter-GPU communication check - `all_reduce_perf` - GPU bandwidth check - `dcgmi diag -r nvbandwidth` - **one pass across all GPUs** (`-i `). No NUMA-locality split in Validate. ## Validate -> Stress - Extended NVIDIA GPU check - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dmidecode -t baseboard` - `dmidecode -t system` - `dcgmi diag -r 3` - NVIDIA targeted stress - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dcgmi diag -r targeted_stress` - NVIDIA targeted power - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dcgmi diag -r targeted_power` - NVIDIA pulse test - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dcgmi diag -r pulse_test` - Inter-GPU communication check - `all_reduce_perf` - GPU bandwidth check - `dcgmi diag -r nvbandwidth` - **per-NUMA-node matrix**: one pass per completely resolved Linux PCI NUMA group, then one all-GPU pass. If any selected GPU has no resolved `numa_node`, no group is guessed and the test falls back to one all-GPU pass. See `bible-local/decisions/2026-07-27-nvbandwidth-per-socket-split.md` and `bible-local/decisions/2026-08-31-nvbandwidth-validate-single-deep-matrix.md`.