Files
bee/audit/internal/app/assets/README.md
T
Mikhail Chusavitin ced2175fb0 platform: split nvbandwidth SAT into per-socket + all-GPU passes
On multi-socket systems, run the NVIDIA bandwidth diagnostic once per
CPU socket before the all-GPU pass, so a crash confined to the all-GPU
run (with clean per-socket passes preceding it) isolates a cross-socket
peer-to-peer fault instead of leaving it conflated with a general
GPU/PCIe issue. Single-socket systems keep the original one-pass shape.

Also expand the support-bundle README with reference notes distilled
from a real analysis pass (BMC clock drift, "0/empty" tool output
meaning absent hardware rather than a fault, timestamp-matching before
assigning causality, and a normal-power-cycle SEL signature), plus a
step-by-step recipe for diagnosing an unexpected reboot/crash during a
specific test.
2026-07-27 17:01:39 +03:00

153 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# README — how to read this bundle
**If you are an AI agent that just listed or was pointed at this directory:
read this file before opening any other file in it.** It tells you what's
here and exactly where to look, so you don't burn time grepping through
random logs to answer something this file already answers directly.
You are looking at a diagnostic export from **bee** (BEE-SP), a hardware
audit / acceptance-testing / burn-in tool for GPU servers. It runs on the
server itself (as a live-CD or an installed service) and produces this
export by: collecting a full hardware inventory, running SAT (System
Acceptance Test) checks against CPU/memory/storage/GPU/RAID, capturing raw
diagnostic command output, and logging its own systemd services.
## Fastest answers to common questions
- **"What hardware is in this server?"** → `export/bee-audit.json`. The
master inventory: board, firmware, CPUs, memory, storage, PCIe devices
(GPUs/NICs/RAID controllers), power supplies, sensors, event logs. Every
component has a `status` field (`OK`/`Warning`/`Critical`/`Unknown`).
- **"Did the acceptance tests pass?"** → grep `overall_status=` across every
`export/bee-sat/*/summary.txt`. That key is always present and is the
fastest single-line verdict for that run (`OK`, `FAILED`, `PARTIAL`, or
`UNSUPPORTED`).
- **"What's the *current* rolled-up health per component?"** →
`export/component-status.json`. One record per component key (e.g.
`pcie:gpu:nvidia`, `cpu:all`, `psu:0`), each with a `status` and a
`history` array. **The history is a transition log, not a per-check
journal** — a component whose status never changed only has one entry,
even if it was checked hundreds of times. A gap in timestamps does not
mean it stopped being monitored.
- **"Is a specific service healthy, or did it crash/restart-loop?"** →
`systemd/<service>.status.txt` (current `systemctl status`) and
`systemd/<service>.journal.log` (`journalctl -u <service>` for that
service's window). `systemd/combined.journal.log` has everything,
chronological, if you need cross-service correlation.
- **"What's the RAID/drive state?"** →
`export/techdump/storcli64-drives.json` and `storcli2-show-all.json` (LSI
controllers) — per-slot state like `JBOD`, `UGood`, `Onln`, `UBad`. A
drive in `JBOD`/`UBad` state cannot join a new virtual disk without first
being converted (`set good force`).
- **"GPU topology / NVLink health?"** →
`export/techdump/nvidia-smi-topo.txt` (`NV#` = bonded NVLink pair, `PIX`/
`NODE`/`SYS` = PCIe-only, no NVLink), `nvidia-smi-nvlink-status.txt`
(per-link active/inactive), `nvidia-smi-nvlink-errors.txt` (replay/
recovery/CRC counters, should be zero), and `bee-sat/nvidia-config-*/summary.txt`'s
`nvlink_pairs_checked` (how many pairs it found to check, not how many
passed). **Not every GPU config has NVLink bridges — check whether this
SKU/order is supposed to have them before calling their absence a fault**
(all-`PIX` topology, empty nvlink-status/errors files, and
`nvlink_pairs_checked=0` together mean "none detected," which is only a
problem if the config calls for NVLink). If NVLink is expected, one
`<inactive>` lane next to active ones on an otherwise-bonded pair is the
real fault signature — a degraded link, not an absent one.
- **"What tasks were run from the web UI, in what order, with what
result?"** → `export/tasks-state.json` is the index (id, target, status,
timestamps, paths). Each task also has its own directory
`export/tasks/<NNN>_<slug>_<done|failed>/` with `task.log` (live output),
`report.json`/`report.html` (rendered result + charts).
- **"What build/version is this, and when was it captured?"** →
`manifest.txt` (`bee_version=`, `generated_at_utc=`) at the top level. If
the bundle's own folder/file name embeds `(BEE-SP vXX.YY)`, that's the
same version, useful for cross-referencing against the bee git repo's
release tags if you're checking whether a specific fix shipped in this
build.
## Signatures
- `ipmitool-sel-time.txt` disagrees with `manifest.txt`'s `generated_at_utc`
by more than minutes → BMC RTC drifted → all `ipmitool-sel.txt` timestamps
unreliable; use record-ID order (hex counter, column 1), not printed date.
- `storcli64``"Status": "Failure", "Description": "No Controller found"`
→ no legacy MegaRAID controller present, not a drive/RAID fault. Cross-ref
`storcli2-show-all.json`'s `"Number of Controllers"` and `lspci`/
`nvme-list.json` for actual storage hardware. Same pattern as
`nvlink_pairs_checked=0` above.
- `systemctl restart ... timed out` in a `bee-*.log` → implicates a SAT run
only if its timestamp is inside that run's `run_at_utc`/`started_at`
`done_at` window in `tasks-state.json`; outside that window (e.g. boot
bring-up) it's unrelated.
- SEL sequence `Power Supply Failure detected`/`AC lost` (all PSUs) →
`ACPI Legacy OFF``Chassis intrusion``Power Button pressed`
`Legacy ON` → full power-cycle, not a failing PSU. Isolated failure
assertion with no OFF/ON bracket, or one PSU failing while siblings stay
healthy → real PSU fault.
- Every conclusion needs a file/line/key citation. `Failure`/`Critical`/
`timeout`/`Error` matched without reading the surrounding context is not
a citation.
## Diagnosing "it rebooted/crashed during test X"
1. `tasks-state.json` → task stuck `pending`/`running`, or last
`bee-sat/<target>-*/` with no `summary.txt` = check running at crash time.
2. That directory's `verbose.log` → last subprocess with no matching
`finish`/exit-code line = the trigger.
3. `system/dmesg.txt` starting at uptime 0, ending after a few hundred
seconds = fresh-boot log captured after an unclean reset (confirms
unclean reboot; does not contain the crash itself).
4. Diff against a step that passed cleanly → isolates what's specific to
the failing workload (e.g. GPU-to-GPU/NVLink traffic vs. per-GPU
compute-only) instead of a generic cause (thermal, power) that would
also hit the passing step.
## Top-level layout
```
manifest.txt bee_version, host, generated_at_utc, export_dir,
then a flat "path<TAB>size" listing of every file
in this bundle — a quick inventory/sanity check.
README.md this file — read this first.
export/ mirror of the live /appdata/bee/export directory —
see "export/" below.
systemd/ <service>.status.txt + <service>.journal.log per
monitored systemd unit, plus combined.journal.log.
techdump/, system/ raw command output not tied to a specific bee-*
service — see below.
```
## `export/` in detail
This is a straight mirror of the live server's export directory, so
everything below also applies when reading a raw `bee export` output
directly (not wrapped in a support-bundle archive).
| Path | What it is |
|---|---|
| `bee-audit.json` | The master hardware snapshot (see above). |
| `bee-audit.log` | Log of the hardware-audit collector itself: what it queried, what it skipped and why (e.g. a field unsupported by the current driver — not necessarily an error). |
| `bee-web.log` | Web UI service log. |
| `bee-network.log`, `bee-nvidia.log`, `bee-sshsetup.log`, `bee-selfheal.log`, `bee-blackbox.log`, `bee-hpc-tuning.log` | Per-service startup/runtime logs for the correspondingly-named systemd unit. |
| `runtime-health.json` / `.log` | A lighter, more frequent health snapshot than the full audit — good for "was it fine 5 minutes ago" without the cost of a full re-audit. |
| `component-status.json` | Current rolled-up component health — see "Fastest answers" above. |
| `blackbox-state.json` | Present if continuous blackbox capture (mirroring this export dir to removable media on a schedule) is or was active; tracks the sync target and last successful sync. |
| `techdump/` | Raw diagnostic command output, captured once per audit cycle. This is ground truth: `nvidia-smi-*.txt/.csv` (GPU state/topology/NVLink), `storcli64-drives.json` / `storcli2-show-all.json` (RAID), `lspci-*.txt`, `lscpu.txt`, `lsblk.json`, `dmidecode-*.txt`, `ipmitool-*.txt` (BMC sensors/SEL/FRU), `smartctl-*.json`, `nvme-list.json`, `sensors.json`. Higher-level views (the audit JSON, the web UI's topology page) are built from these. |
| `bee-sat/<target>-<timestamp>/` | One directory per acceptance-test run — `target` is what was tested (`gpu-nvidia`, `cpu`, `memory`, `storage`, `nccl-tests`, `gpu-nvidia-bandwidth`, `nvidia-config` covering GPU config/NVLink/Confidential-Computing readiness, etc). Each contains `summary.txt` (key=value, always has `overall_status`), a full human-readable report, numbered per-job logs, and `verbose.log` (every subprocess invocation + exit code — the place to look when a summary doesn't explain *why* something failed). |
| `tasks/<NNN>_<slug>_<done\|failed>/` | One directory per task launched from the web UI's task queue — see "Fastest answers" above. |
| `tasks-state.json` | Index of every task (id, target, status, timestamps, artifact paths). |
| `systemd/` (nested) | Historical per-service snapshots captured as part of an audit cycle, same shape as the top-level `systemd/` described below. |
## `systemd/` and `system/` (top level, sibling of `export/`)
- `systemd/<service>.status.txt`, `systemd/<service>.journal.log`,
`systemd/combined.journal.log` — captured fresh at bundle-build time (not
mirrored from `export/`), so this is the most current service state.
- `system/` — general OS-level diagnostics not specific to bee:
`dmesg.txt`, X server / display-manager logs. Mostly relevant to
physical/console access issues, not hardware health.
## Timestamps
Everything is UTC unless a filename or field name says otherwise
(`*_local`, etc. — rare).