- GPU load switches from bee-gpu-burn (compute burn, ~88% TDP) to dcgmproftester -t 1004 / targeted_power via resolveBenchmarkPowerLoadCommand — the same engine Power/Thermal Fit uses and the hottest sustained NVIDIA load we have, so fans are actually pushed toward their ceiling. - Sample loop is now IPMI-hang-proof: every ipmitool read is time-boxed in an abandonable goroutine, and the poll interval backs off geometrically (1s→30s) when reads are slow, tightening again on recovery. A plateau is only trusted while telemetry is healthy; degraded runs ride out to MaxLoadSec. Summary gains fan_samples / telemetry_degraded. Drops the per-second nvidia-smi+power+cpu-temp sampling from the hot loop. - Dead code removed: FanStressRow, GPUStressMetric, sampleFanStressRow, sampleGPUStressMetrics, WriteFanStressCSV/WriteFanSensorsCSV, analyzeMaxTemp, sampleSystemPowerResolved. Topology fan tiles: - size encodes the fan's ceiling RPM (its class), not current speed; coloured fill rising from the bottom encodes live duty cycle (current / ceiling), shown only when the ceiling was measured. - glyph spin rate now maps absolute RPM into a human-perceptible band (fanSpinPeriodSec: 2.2s/turn at <=1000 RPM, 0.35s at >=13000). - the "N fans · N OK · tile size ∝ …" caption line is gone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019VHG21rgTUiR1G3qFHTVmN
83 lines
3.2 KiB
Markdown
83 lines
3.2 KiB
Markdown
# bee — Project Bible
|
|
|
|
Project-specific architecture, decisions, and runtime contracts.
|
|
Generic engineering rules live in `bible/rules/patterns/`.
|
|
|
|
## Files
|
|
|
|
| File | Contents |
|
|
|---|---|
|
|
| `architecture/system-overview.md` | What bee does, scope, tech stack |
|
|
| `architecture/runtime-flows.md` | Boot sequence, audit flow, service order |
|
|
| `architecture/squashfs-layers.md` | Semantic SquashFS layer model, ownership, order, verification |
|
|
| `docs/customer-gpu-test-methodology.md` | Customer-facing GPU PCIe Validate / Validate -> Stress test list |
|
|
| `docs/hardware-ingest-contract.md` | Current Reanimator hardware ingest JSON contract |
|
|
| `docs/validate-vs-burn.md` | Validate and Validate -> Stress hardware test policy |
|
|
| `decisions/` | Architectural decision log, including read-only submodule policy |
|
|
| `proposals/` | RFCs and contract change proposals for Reanimator Core |
|
|
|
|
## Validate Test Matrix
|
|
|
|
### Validate
|
|
|
|
- CPU check
|
|
- `lscpu`
|
|
- `sensors`
|
|
- `stress-ng`
|
|
- TPM check (read-only)
|
|
- `tpm2_getcap properties-fixed`
|
|
- `tpm2_getcap pcrs`
|
|
- `tpm2_pcrread`
|
|
- `tpm2_gettestresult` (reads the existing result; does not start `TPM2_SelfTest`)
|
|
- Memory check
|
|
- `free`
|
|
- `timeout <timeout_sec> memtester`
|
|
- `free`
|
|
- NVMe storage check
|
|
- `nvme id-ctrl`
|
|
- `nvme smart-log`
|
|
- `nvme device-self-test`
|
|
- SATA/SAS storage check
|
|
- `smartctl -H -A`
|
|
- `smartctl -t short`
|
|
- Basic NVIDIA GPU check
|
|
- `nvidia-smi -pm 1`
|
|
- `nvidia-smi -q`
|
|
- `dmidecode -t baseboard`
|
|
- `dmidecode -t system`
|
|
- `dcgmi diag -r 2`
|
|
- Inter-GPU communication check
|
|
- `all_reduce_perf`
|
|
- GPU bandwidth check
|
|
- `dcgmi diag -r nvbandwidth` (per CPU socket, then all selected GPUs, on multi-socket systems -- see `decisions/2026-07-27-nvbandwidth-per-socket-split.md`)
|
|
|
|
### Validate -> Stress
|
|
|
|
- Extended NVIDIA GPU check
|
|
- `nvidia-smi -pm 1`
|
|
- `nvidia-smi -q`
|
|
- `dmidecode -t baseboard`
|
|
- `dmidecode -t system`
|
|
- `dcgmi diag -r 3`
|
|
- NVIDIA targeted stress
|
|
- `nvidia-smi -pm 1`
|
|
- `nvidia-smi -q`
|
|
- `dcgmi diag -r targeted_stress`
|
|
- NVIDIA targeted power
|
|
- `nvidia-smi -pm 1`
|
|
- `nvidia-smi -q`
|
|
- `dcgmi diag -r targeted_power`
|
|
- NVIDIA pulse test
|
|
- `nvidia-smi -pm 1`
|
|
- `nvidia-smi -q`
|
|
- `dcgmi diag -r pulse_test`
|
|
- Inter-GPU communication check
|
|
- `all_reduce_perf`
|
|
- GPU bandwidth check
|
|
- `dcgmi diag -r nvbandwidth` (per CPU socket, then all selected GPUs, on multi-socket systems -- see `decisions/2026-07-27-nvbandwidth-per-socket-split.md`)
|
|
- Fan ceiling check (Load tier / `3. Load` only)
|
|
- `stressapptest` (or `stress-ng`) + the hottest GPU load — `dcgmproftester -t 1004` / `targeted_power` (Power/Thermal Fit engine), **not** `bee-gpu-burn` — run simultaneously
|
|
- `ipmitool sdr type Fan` on an adaptive interval (1 s floor, backs off to 30 s when the BMC gets slow, every read time-boxed) until every fan plateaus (~1 min flat, only trusted while telemetry is healthy) or the 15 min cap
|
|
- records each fan's observed peak RPM to the fan-observation store (used by the Topology fan tiles: size ∝ ceiling, fill ∝ duty cycle); FAIL only on a fan at 0 RPM / IPMI cr-nr under load; **cancelled ("not applicable")**, never failed, if the host cannot be loaded or has no fan sensors
|
|
- see `decisions/2026-09-04-fan-ceiling-check.md`
|