# README — how to read this bundle **If you are an AI agent that just listed or was pointed at this directory: read this file before opening any other file in it.** It tells you what's here and exactly where to look, so you don't burn time grepping through random logs to answer something this file already answers directly. You are looking at a diagnostic export from **bee** (BEE-SP), a hardware audit / acceptance-testing / burn-in tool for GPU servers. It runs on the server itself (as a live-CD or an installed service) and produces this export by: collecting a full hardware inventory, running SAT (System Acceptance Test) checks against CPU/memory/storage/GPU/RAID, capturing raw diagnostic command output, and logging its own systemd services. The bundle is organized around **what kind of thing each file is**, not which bee component happened to produce it: - **`export/`** — raw, vendor-native tool output, grouped by physical device type. Nothing in here needs bee context to read; a NIC vendor can be handed `export/network/` and it means the same thing it would from any other collector. Also holds `reanimator.json`, a ready-to-ingest hardware snapshot. - **`status/`** — *computed* diagnosis (component health verdicts, runtime status, metrics history), not raw tool output and not a task log. - **`tasks/`** — bee's own bookkeeping: full per-run logs, orchestration state, service logs. Expect duplication with `export/` here (e.g. a storage SAT run's narrative disk report repeats data already in `export/storage/`) — that's intentional, this is the "show your work" tier. - **`livecd/`** — only relevant to *this particular live-boot/kiosk session* (GUI stack, DHCP-assigned IPs, kernel ring buffer for this boot). Not hardware facts — ignore this entirely when troubleshooting the physical machine. ## Fastest answers to common questions - **"What hardware is in this server?"** → `export/reanimator.json`. The master inventory: board, firmware, CPUs, memory, storage, PCIe devices (GPUs/NICs/RAID controllers), power supplies, sensors, event logs. Every component has a `status` field (`OK`/`Warning`/`Critical`/`Unknown`). Same document a `POST /ingest/hardware` to Reanimator expects (see `bible-local/docs/hardware-ingest-contract.md` in the bee repo). - **"Did the acceptance tests pass?"** → grep `overall_status=` across every `tasks/bee-sat/*/summary.txt`. That key is always present and is the fastest single-line verdict for that run (`OK`, `FAILED`, `PARTIAL`, or `UNSUPPORTED`). - **"What's the *current* rolled-up health per component?"** → `status/component-status.json`. One record per component key (e.g. `pcie:gpu:nvidia`, `cpu:all`, `psu:0`), each with a `status` and a `history` array. This is what feeds `reanimator.json`'s `status` fields — it's the source, not a duplicate. **The history is a transition log, not a per-check journal** — a component whose status never changed only has one entry, even if it was checked hundreds of times. A gap in timestamps does not mean it stopped being monitored. - **"Is a specific bee service healthy, or did it crash/restart-loop?"** → `tasks/_services/.status.txt` (current `systemctl status`) and `tasks/_services/.journal.log` (`journalctl -u ` for that service's window). `tasks/_services/combined.journal.log` has everything, chronological, if you need cross-service correlation. NVIDIA DCGM/fabric-manager daemon status lives in `export/gpu/` instead (see below); the display stack's lives in `livecd/gui/`. - **"What's the RAID/drive state?"** → `export/storage/storcli64-drives.json` and `storcli2-show-all.json` (LSI controllers) — per-slot state like `JBOD`, `UGood`, `Onln`, `UBad`. A drive in `JBOD`/`UBad` state cannot join a new virtual disk without first being converted (`set good force`). - **"GPU topology / NVLink health?"** → `export/gpu/nvidia-smi-topo.txt` (`NV#` = bonded NVLink pair, `PIX`/ `NODE`/`SYS` = PCIe-only, no NVLink), `nvidia-smi-nvlink-status.txt` (per-link active/inactive), `nvidia-smi-nvlink-errors.txt` (replay/ recovery/CRC counters, should be zero), and `tasks/bee-sat/nvidia-config-*/summary.txt`'s `nvlink_pairs_checked` (how many pairs it found to check, not how many passed). If `export/gpu/` also has `nvidia-smi-topo-fresh.txt` / `-nvlink-status-fresh.txt` / `-nvlink-errors-fresh.txt`, those were recaptured live at bundle-build time — prefer them over the non-`-fresh` ones if the two disagree, since the non-`-fresh` files are only as current as the last audit cycle. **Not every GPU config has NVLink bridges — check whether this SKU/order is supposed to have them before calling their absence a fault** (all-`PIX` topology, empty nvlink-status/errors files, and `nvlink_pairs_checked=0` together mean "none detected," which is only a problem if the config calls for NVLink). If NVLink is expected, one `` lane next to active ones on an otherwise-bonded pair is the real fault signature — a degraded link, not an absent one. - **"What tasks were run from the web UI, in what order, with what result?"** → `tasks/_state/tasks-state.json` is the index (id, target, status, timestamps, paths). Each task also has its own directory `tasks/__/` with `task.log` (live output), `report.json`/`report.html` (rendered result + charts). - **"What build/version is this, and when was it captured?"** → `manifest.txt` (`bee_version=`, `generated_at_utc=`) at the top level. If the bundle's own folder/file name embeds `(BEE-SP vXX.YY)`, that's the same version, useful for cross-referencing against the bee git repo's release tags if you're checking whether a specific fix shipped in this build. - **"Was there a CPU/memory/power utilization spike around time T?"** → `status/metrics.db`, a SQLite database (`sys_metrics` table: `ts`, `cpu_load_pct`, `mem_load_pct`, `power_w`, `power_source`, `power_mode`, `power_reason`). Open with any SQLite client. ## Signatures - `export/platform/ipmitool-sel-time.txt` disagrees with `manifest.txt`'s `generated_at_utc` by more than minutes → BMC RTC drifted → all `ipmitool-sel.txt` timestamps unreliable; use record-ID order (hex counter, column 1), not printed date. - `storcli64` → `"Status": "Failure", "Description": "No Controller found"` → no legacy MegaRAID controller present, not a drive/RAID fault. Cross-ref `storcli2-show-all.json`'s `"Number of Controllers"` and `lspci`/ `nvme-list.json` for actual storage hardware. Same pattern as `nvlink_pairs_checked=0` above. - `systemctl restart ... timed out` in a `tasks/_services/bee-*.log` → implicates a SAT run only if its timestamp is inside that run's `run_at_utc`/`started_at`–`done_at` window in `tasks/_state/tasks-state.json`; outside that window (e.g. boot bring-up) it's unrelated. - SEL sequence `Power Supply Failure detected`/`AC lost` (all PSUs) → `ACPI Legacy OFF` → `Chassis intrusion` → `Power Button pressed` → `Legacy ON` → full power-cycle, not a failing PSU. Isolated failure assertion with no OFF/ON bracket, or one PSU failing while siblings stay healthy → real PSU fault. - Every conclusion needs a file/line/key citation. `Failure`/`Critical`/ `timeout`/`Error` matched without reading the surrounding context is not a citation. ## Diagnosing "it rebooted/crashed during test X" 1. `tasks/_state/tasks-state.json` → task stuck `pending`/`running`, or last `tasks/bee-sat/-*/` with no `summary.txt` = check running at crash time. 2. That directory's `verbose.log` → last subprocess with no matching `finish`/exit-code line = the trigger. 3. `livecd/host/dmesg.txt` starting at uptime 0, ending after a few hundred seconds = fresh-boot log captured after an unclean reset (confirms unclean reboot; does not contain the crash itself). 4. Diff against a step that passed cleanly → isolates what's specific to the failing workload (e.g. GPU-to-GPU/NVLink traffic vs. per-GPU compute-only) instead of a generic cause (thermal, power) that would also hit the passing step. ## Top-level layout ``` manifest.txt bee_version, host, generated_at_utc, export_dir, then a flat "pathsize" listing of every file in this bundle — a quick inventory/sanity check. README.md this file — read this first. export/ raw, vendor-native hardware data, by device type — see "export/" below. status/ computed diagnosis, not raw data — see below. tasks/ bee's own task-run bookkeeping — see below. livecd/ this boot/kiosk session only, not hardware — see below. ``` ## `export/` in detail Raw tool/command output, grouped by the physical device it describes. Vendor-neutral: nothing here assumes familiarity with bee. | Path | What it is | |---|---| | `reanimator.json` | The master hardware snapshot (see "Fastest answers" above) — ready to `POST /ingest/hardware`. | | `cpu/` | `lscpu.txt`, `dmidecode-type4.txt` (Processor Information). | | `memory/` | `dmidecode-type17.txt` (Memory Device, one entry per DIMM). | | `storage/` | `smartctl-.json`, `nvme-id-ctrl-.json`, `nvme-smart-log-.json`, `nvme-list.json`, `storcli64-drives.json`, `storcli2-show-all.json`, `lsblk.json`. | | `gpu/` | `nvidia-smi-*.txt/.csv` (state/topology/NVLink; `-fresh` variants are recaptured live at bundle-build time — see "GPU topology" above), `dcgmi-nvlink-status.txt`, `nvidia-bug-report.txt`, `nvidia-dcgm.service`/`nvidia-fabricmanager.service` status+journal, `fabric-manager-paths.txt`, `fabricmanager.log`/`nvlsm.log`, `lspci-nvidia-bridges-vv.txt`, `pcie-nvidia-link.txt` (idle sample — a Gen1 reading here can be normal driver power management, not a fault), `pcie-nvidia-link-under-load.txt` (same sysfs attributes resampled while `bee-gpu-burn` is actively loading the GPUs; if the link trains up to full speed here, the idle Gen1 reading was power saving — if it stays at Gen1 under load, that's a real link/riser/slot degradation), `pcie-aer-sysfs.txt`, `kernel-aer-nvidia.txt` (AER/NVRM/Xid-filtered dmesg), `lspci-video-vv.txt`, `systemctl-nvidia-units.txt`. AMD systems get `rocm-smi*.txt` here instead. | | `network/` | `ethtool-{info,link,module}.txt` (per-NIC), `mstflint-query.txt` (Mellanox/NVIDIA NICs). | | `platform/` | `dmidecode-type{0,1,2}.txt` (BIOS/System/Baseboard), `ipmitool-{fru,sdr,sensor,sel,sel-time}.txt` (BMC), `sensors.json`, `lspci-{nn,vmm,vvv}.txt`. | ## `status/` in detail Computed diagnosis, not raw data — see "Fastest answers" above for `component-status.json` and `metrics.db`. | Path | What it is | |---|---| | `component-status.json` | Current rolled-up component health, written by bee's live `/dev/kmsg` watcher and by SAT run results. Source of `reanimator.json`'s `status`/`status_history` fields. | | `runtime-health.json` / `.log` | A lighter, more frequent health snapshot than the full audit — good for "was it fine 5 minutes ago" without the cost of a full re-audit. | | `metrics.db` | SQLite time-series (CPU/mem load, power draw/source/mode) — see "Fastest answers" above. | ## `tasks/` in detail bee's own bookkeeping. Duplication with `export/` is expected here — this tier keeps the full story of a specific run, `export/` keeps the clean current-state answer. | Path | What it is | |---|---| | `bee-sat/-/` | One directory per acceptance-test run — `target` is what was tested (`gpu-nvidia`, `cpu`, `memory`, `storage`, `nccl-tests`, `gpu-nvidia-bandwidth`, `nvidia-config` covering GPU config/NVLink/Confidential-Computing readiness, etc). Each contains `summary.txt` (key=value, always has `overall_status`), a full human-readable report (e.g. `storage`'s per-drive `disk-NN--report.txt` — a narrative verdict, not raw tool output; the raw data behind it lives in `export/storage/`), numbered per-job logs, and `verbose.log` (every subprocess invocation + exit code — the place to look when a summary doesn't explain *why* something failed). | | `bee-bench/` | Benchmark run artifacts, if any have been run. | | `__/` | One directory per task launched from the web UI's task queue — see "Fastest answers" above. | | `_state/bee-audit.json` | The full hardware-audit document as bee last generated it — same content as `export/reanimator.json`, kept here as bee's own internal record. | | `_state/tasks-state.json` | Index of every task (id, target, status, timestamps, artifact paths). | | `_state/blackbox-state.json` | Present if continuous blackbox capture (mirroring the export dir to removable media on a schedule) is or was active; tracks the sync target and last successful sync. | | `_services/bee-*.log` | Per-service startup/runtime logs for each bee-owned systemd unit. | | `_services/.status.txt`, `.journal.log`, `combined.journal.log` | Captured fresh at bundle-build time, for every bee-owned service (`bee-audit`, `bee-web`, `bee-network`, `bee-nvidia`, `bee-preflight`, `bee-selfheal`, `bee-sshsetup`, `bee-blackbox`). | ## Why a bee service/check might have failed Before treating a failed/inactive unit as a hardware fault, check whether its own logic is *designed* to skip or fail under specific benign conditions — several of these look like failures in `systemctl status` but are expected behavior on that particular machine. | Service / check | What normal failure/skip looks like and why | |---|---| | `bee-blackbox.service` | Restarts forever (`Restart=always`) mirroring `export/` to removable media; only a real problem if `_state/blackbox-state.json`'s last-sync timestamp is stale — no USB media present is a common, harmless cause. | | `bee-nvidia.service` | Exits 0 and skips entirely if no `10de:0300`/`0302` (NVIDIA display/3D) PCI function is present, or if `/usr/local/bin/bee-nvidia-load` is absent (non-NVIDIA build) — not a failure. A real failure is `insmod` erroring, or (SXM→PCIe conversion boards) GSP firmware entering a crash/reload loop that never populates `/proc/devices` with `nvidiactl` within 90s; the script then tries `rmmod` to fall back to `NVreg_EnableGpuFirmware=0` — if `rmmod` also fails (`EBUSY`, module stuck), it logs `gsp-stuck` to `/run/bee-nvidia-mode` and the machine must be rebooted with `bee.nvidia.mode=gsp-off` (or `nomsi` if the dmesg tail shows `Failed to enable MSI-X`) at the boot menu. See `export/gpu/systemctl-nvidia-units.txt` and this service's `dmesg:` lines in its log for the exact symptom. | | `nvidia-fabricmanager.service` (`bee-check-nvswitch` `ExecCondition`) | **Not a failure when inactive.** `ExecCondition=/usr/local/bin/bee-check-nvswitch` decides whether this SKU has an NVSwitch fabric, in order: (1) `nvidia-smi -q`'s `GPU Fabric GUID` field has a real value (not `N/A`) — the primary signal, since on platforms like HGX B200 (Kaytus KR9288-X3) NVSwitch is managed via HMC/ERoT and never appears on host-PCIe at all; (2) `/proc/driver/nvidia-nvswitch/devices` is non-empty; (3) legacy `lspci -Dn` scan for a vendor `10de` class `0680` (Bridge, Other) device, for older HGX generations where NVSwitch *is* host-PCIe-visible. If all three say no (plain PCIe GPUs / direct NVLink bridges with no switch), the check exits 1 and systemd marks the unit `inactive (dead)` with condition unmet — never `failed`. Confirm via `systemctl status nvidia-fabricmanager.service` showing `ConditionResult=no`, plus `export/gpu/nvidia-smi-topo.txt` showing all-`PIX` topology (see "GPU topology" above) — same signature as `nvlink_pairs_checked=0`. **A stale build with the pre-driver-level lspci-only check will misfire on host-PCIe-invisible NVSwitch platforms** (e.g. HGX B200): condition (3) alone always says no, FM is always skipped, and `cuInit`/NCCL/`dcgmi diag` then fail with "system not yet initialized" on otherwise-healthy hardware — check `bee_version` in `manifest.txt` against the fix if this pattern shows up. An actual fabric-manager fault (real NVSwitch present, condition passes, but `ExecStart` errors) instead shows up as `failed`/`activating (auto-restart)` with training errors in `export/gpu/fabricmanager.log`/`nvlsm.log`. | | `nvidia-dcgm.service` | Ordered `After=bee-nvidia.service`; if it starts before device nodes exist, `nv-hostengine` enumerates zero GPUs and never rescans — `bee-nvidia-load` works around this by force-restarting it once modules are confirmed loaded, so a stale empty inventory here usually means that restart itself failed or timed out (see `bee-nvidia.log`'s `WARN: failed to restart nvidia-dcgm.service` / `timed out after 60s` lines) rather than a DCGM bug. | | `bee-audit.service` / `bee-preflight.service` | `bee-preflight` deliberately never fails the unit (`; exit 0` regardless of the underlying `bee preflight` return code) — check `runtime-health.log`'s own `WARN: preflight exited with rc=N` line, not unit status, for a real preflight problem. `bee-audit` failing outright usually means the underlying `bee audit` binary hit an unrecoverable collection error; check `export/bee-audit.log`, not just `systemctl status`. | | `bee-network.service` | Also structured to exit 0 unconditionally — `list_interfaces` filters out virtual/loopback interfaces, then DHCP is attempted per-NIC with a 300s timeout; a per-interface DHCP failure or timeout is logged (`DHCP timed out on ` / `DHCP failed on `) but does not fail the unit, since a lease-less NIC (no cable, no DHCP server on that segment) is common and not itself a hardware defect. Check `bee-network.log`'s `leases acquired: N` vs `interfaces scanned: N` line for the real outcome. | | `bee-sshsetup.service` | Fails only on `useradd`/`chpasswd`/file-write errors (e.g. read-only `/etc`), which would be a filesystem problem, not an SSH-specific one. Whether it configures key-only auth or the `bee`/`eeb` password fallback is a deliberate branch on `/etc/bee-ssh-password-fallback`'s presence, not a fault either way — check which branch its log took before assuming misconfiguration. | | `bee-selfheal.service`/`.timer` | A oneshot that runs on a timer and restarts other bee services when it detects specific bad states — its own "failure" is rare (only if the lock dir can't be created, or the restart calls themselves fail) but its *log content* (`_services/bee-selfheal.log`) is often the fastest place to see which downstream service was unhealthy and had to be kicked: `/dev/nvidia0` missing while a GPU is present, an interrupted `.tmp` write for `bee-audit.json`/`runtime-health.json`, `bee-web.service` inactive, or `bee-web`'s own `/healthz` endpoint not answering on `127.0.0.1:80`. | | `bee-web.service` | `Restart=always`, no `MemoryMax` (deliberately — it spawns multi-GB GPU stress subprocesses like `dcgmproftester` that a cgroup limit would OOM-kill); a genuine crash-loop here is unusual and worth its full `journal.log`, not just the latest restart. | | `bee-hpc-tuning.service` | Best-effort CPU governor/C-state tuning ordered before `bee-nvidia`/`bee-audit`; failure here (e.g. governor sysfs knobs absent on this CPU/BIOS combination) does not block the rest of boot but can explain unexpectedly poor benchmark numbers if not caught. | ## `livecd/` in detail Only relevant to this particular live-boot/kiosk session — not the physical hardware. Safe to ignore entirely when troubleshooting hardware. | Path | What it is | |---|---| | `gui/` | `Xorg.0.log`, `lightdm/`, `xinput-list.txt`, `libinput-list-devices.txt`, `drm-cards.txt`, `proc-fb.txt`, `input-devices.txt`, `udevadm-input.txt`, `loginctl-{sessions,seats}.txt`, `ps-gui.txt`, `systemctl-gui-units.txt`, `home-bee-xsession-errors.log`, `home-bee-chromium-debug.log`, `dmesg-gui-video-input.txt` (GUI/video/input-filtered dmesg), `display-manager.service`/`lightdm.service` status+journal. | | `host/` | `uname.txt`, `cmdline.txt`, `lsmod.txt`, `mount.txt`, `df-h.txt`, `dmesg.txt` (full kernel ring buffer for this boot), `ip-{addr,link,link-stats,route}.txt` (this session's network config, e.g. DHCP-assigned addresses — not a hardware fact about the NIC itself; NIC hardware/firmware state is in `export/network/`). | ## Timestamps Everything is UTC unless a filename or field name says otherwise (`*_local`, etc. — rare).