# bee — Project Bible Project-specific architecture, decisions, and runtime contracts. Generic engineering rules live in `bible/rules/patterns/`. ## Files | File | Contents | |---|---| | `architecture/system-overview.md` | What bee does, scope, tech stack | | `architecture/runtime-flows.md` | Boot sequence, audit flow, service order | | `architecture/squashfs-layers.md` | Semantic SquashFS layer model, ownership, order, verification | | `docs/customer-gpu-test-methodology.md` | Customer-facing GPU PCIe Validate / Validate -> Stress test list | | `docs/hardware-ingest-contract.md` | Current Reanimator hardware ingest JSON contract | | `docs/validate-vs-burn.md` | Validate and Validate -> Stress hardware test policy | | `decisions/` | Architectural decision log, including read-only submodule policy | | `proposals/` | RFCs and contract change proposals for Reanimator Core | ## Validate Test Matrix ### Validate - CPU check - `lscpu` - `sensors` - `stress-ng` - TPM check (read-only) - `tpm2_getcap properties-fixed` - `tpm2_getcap pcrs` - `tpm2_pcrread` - `tpm2_gettestresult` (reads the existing result; does not start `TPM2_SelfTest`) - Memory check - `free` - `timeout memtester` - `free` - NVMe storage check - `nvme id-ctrl` - `nvme smart-log` - `nvme device-self-test` - SATA/SAS storage check - `smartctl -H -A` - `smartctl -t short` - Basic NVIDIA GPU check - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dmidecode -t baseboard` - `dmidecode -t system` - `dcgmi diag -r 2` - Inter-GPU communication check - `all_reduce_perf` - GPU bandwidth check - `dcgmi diag -r nvbandwidth` (per CPU socket, then all selected GPUs, on multi-socket systems -- see `decisions/2026-07-27-nvbandwidth-per-socket-split.md`) ### Validate -> Stress - Extended NVIDIA GPU check - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dmidecode -t baseboard` - `dmidecode -t system` - `dcgmi diag -r 3` - NVIDIA targeted stress - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dcgmi diag -r targeted_stress` - NVIDIA targeted power - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dcgmi diag -r targeted_power` - NVIDIA pulse test - `nvidia-smi -pm 1` - `nvidia-smi -q` - `dcgmi diag -r pulse_test` - Inter-GPU communication check - `all_reduce_perf` - GPU bandwidth check - `dcgmi diag -r nvbandwidth` (per CPU socket, then all selected GPUs, on multi-socket systems -- see `decisions/2026-07-27-nvbandwidth-per-socket-split.md`) - Fan ceiling check (Load tier / `3. Load` only) - `stressapptest` (or `stress-ng`) + the hottest GPU load — `dcgmproftester -t 1004` / `targeted_power` (Power/Thermal Fit engine), **not** `bee-gpu-burn` — run simultaneously - `ipmitool sdr type Fan` on an adaptive interval (1 s floor, backs off to 30 s when the BMC gets slow, every read time-boxed) until every fan plateaus (~1 min flat, only trusted while telemetry is healthy) or the 15 min cap - records each fan's observed peak RPM to the fan-observation store (used by the Topology fan tiles: size ∝ ceiling, fill ∝ duty cycle); FAIL only on a fan at 0 RPM / IPMI cr-nr under load; **cancelled ("not applicable")**, never failed, if the host cannot be loaded or has no fan sensors - see `decisions/2026-09-04-fan-ceiling-check.md`