309 lines
17 KiB
Markdown
309 lines
17 KiB
Markdown
# Runtime Flows — bee
|
|
|
|
## Network isolation — CRITICAL
|
|
|
|
**The live CD runs in an isolated network segment with no internet access.**
|
|
All binaries, kernel modules, and tools must be baked into the ISO at build time.
|
|
No package installation, no downloads, and no package manager calls are allowed at boot.
|
|
DHCP is used only for LAN (operator SSH access). Internet is NOT available.
|
|
|
|
## Boot sequence (single ISO)
|
|
|
|
The live system is expected to boot with `toram`, so `live-boot` copies the full read-only medium into RAM before mounting the root filesystem. After that point, runtime must not depend on the original USB/BMC virtual media staying readable.
|
|
|
|
`systemd` boot order:
|
|
|
|
```
|
|
local-fs.target
|
|
├── bee-sshsetup.service (enables SSH key auth; password fallback only if marker exists)
|
|
│ └── ssh.service (OpenSSH on port 22 — starts without network)
|
|
├── bee-network.service (starts `dhclient -nw` on all physical interfaces, non-blocking)
|
|
├── bee-nvidia.service (insmod nvidia*.ko from /usr/local/lib/nvidia/,
|
|
│ creates /dev/nvidia* nodes)
|
|
├── bee-audit.service (runs `bee audit` → /var/log/bee-audit.json,
|
|
│ never blocks boot on partial collector failures)
|
|
├── bee-web.service (runs `bee web` on :80 — full interactive web UI)
|
|
└── bee-desktop.service (startx → openbox + chromium http://localhost/)
|
|
```
|
|
|
|
**Critical invariants:**
|
|
- The live ISO boots with `boot=live toram`. Runtime binaries must continue working even if the original boot media disappears after early boot.
|
|
- OpenSSH MUST start without network. `bee-sshsetup.service` runs before `ssh.service`.
|
|
- `bee-network.service` uses `dhclient -nw` (background) — network bring-up is best effort and non-blocking.
|
|
- `bee-nvidia.service` loads modules via `insmod` with absolute paths — NOT `modprobe`.
|
|
Reason: the modules are shipped in the ISO overlay under `/usr/local/lib/nvidia/`, not in the host module tree.
|
|
- `bee-nvidia-load` refreshes `nvidia-fabricmanager.service` / `nvidia-dcgm.service`
|
|
with `systemctl --no-block try-restart` only. DO NOT make it a blocking
|
|
`systemctl {start,restart}`: `bee-nvidia.service` is `Type=oneshot` and
|
|
`Before=` both units, so a synchronous call deadlocks against that ordering
|
|
and previously reached a 60-second wrapper timeout for each unit. See
|
|
`decisions/2026-08-31-bee-nvidia-restart-deadlock.md`.
|
|
- `bee-audit.service` does not wait for `network-online.target`; audit is local and must run even if DHCP is broken.
|
|
- `bee-audit.service` logs audit failures but does not turn partial collector problems into a boot blocker.
|
|
- `bee-web.service` binds `0.0.0.0:80` and always renders the current `/var/log/bee-audit.json` contents.
|
|
- Audit JSON now includes a `hardware.summary` block with overall verdict and warning/failure counts.
|
|
|
|
## Console and login flow
|
|
|
|
Local-console behavior:
|
|
|
|
```text
|
|
tty1
|
|
└── live-config autologin → bee
|
|
└── /home/bee/.profile (prints web UI URLs)
|
|
|
|
display :0
|
|
└── bee-desktop.service (User=bee)
|
|
└── startx /usr/local/bin/bee-openbox-session -- :0
|
|
├── tint2 (taskbar)
|
|
├── chromium http://localhost/
|
|
└── openbox (WM)
|
|
```
|
|
|
|
Rules:
|
|
- local `tty1` lands in user `bee`, not directly in `root`
|
|
- `bee-desktop.service` starts X11 + openbox + Chromium automatically after `bee-web.service`
|
|
- Chromium opens `http://localhost/` — the full interactive web UI
|
|
- SSH is independent from the desktop path
|
|
- serial console support is enabled for VM boot debugging
|
|
- Default boot keeps the server-safe graphics path (`nomodeset` + forced `fbdev`) for IPMI/BMC consoles
|
|
- Higher-resolution mode selection is expected only when booting through an explicit `bee.display=kms` menu entry, which disables the forced `fbdev` Xorg config before `lightdm`
|
|
|
|
## ISO build sequence
|
|
|
|
```
|
|
build-in-container.sh [--authorized-keys /path/to/keys]
|
|
1. compile `bee` binary (always; version/git state is part of the artifact)
|
|
2. create a temporary overlay staging dir under `dist/`
|
|
3. inject authorized_keys into staged `root/.ssh/` (or set password fallback marker)
|
|
4. copy `bee` binary → staged `/usr/local/bin/bee`
|
|
5. copy vendor binaries from `iso/vendor/` → staged `/usr/local/bin/`
|
|
(`storcli64`, `sas2ircu`, `sas3ircu`, `arcconf`, `ssacli` — optional; `mstflint` comes from the Debian package set)
|
|
6. `build-nvidia-module.sh`:
|
|
a. install Debian kernel headers if missing
|
|
b. download NVIDIA `.run` installer (sha256 verified, cached in `dist/`)
|
|
c. extract installer
|
|
d. build kernel modules against Debian headers
|
|
e. create `libnvidia-ml.so.1` / `libcuda.so.1` symlinks in cache
|
|
f. cache in `dist/nvidia-<version>-<kver>/`
|
|
7. `build-cublas.sh`:
|
|
a. download `libcublas`, `libcublasLt`, `libcudart` runtime + dev packages from the NVIDIA CUDA Debian repo
|
|
b. verify packages against repo `Packages.gz`
|
|
c. extract headers for `bee-gpu-burn` worker build
|
|
d. cache userspace libs in `dist/cublas-<version>+cuda<series>/`
|
|
8. build `bee-gpu-burn` worker against extracted cuBLASLt/cudart headers
|
|
9. inject NVIDIA `.ko` → staged `/usr/local/lib/nvidia/`
|
|
10. inject `nvidia-smi` → staged `/usr/local/bin/nvidia-smi`
|
|
11. inject `libnvidia-ml` + `libcuda` + `libcublas` + `libcublasLt` + `libcudart` → staged `/usr/lib/`
|
|
12. write staged `/etc/bee-release` (versions + git commit)
|
|
13. patch staged `motd` with build metadata
|
|
14. copy `iso/builder/` into a temporary live-build workdir under `dist/`
|
|
15. sync staged overlay into workdir `config/includes.chroot/`
|
|
16. choose the build path from persisted content/ABI/overlay state:
|
|
a. full: run `lb clean --all && lb config && lb build`
|
|
b. fast: unpack the last squashfs, sync the staged overlay, repack it,
|
|
then rebuild checksums, bootloader assets, ISO, and zsync
|
|
17. validate the final ISO boot menus, volume label, memtest, GRUB assets,
|
|
and variant runtime before publishing it
|
|
```
|
|
|
|
Build host notes:
|
|
- `build-in-container.sh` targets `linux/amd64` builder containers by default, including Docker Desktop on macOS / Apple Silicon.
|
|
- Override with `BEE_BUILDER_PLATFORM=<os/arch>` only if you intentionally need a different container platform.
|
|
- If the local builder image under the same tag was previously built for the wrong architecture, the script rebuilds it automatically.
|
|
|
|
**Critical invariants:**
|
|
- `DEBIAN_KERNEL_ABI` in `iso/builder/VERSIONS` pins the exact kernel ABI used in BOTH places:
|
|
1. `build-in-container.sh` / `build-nvidia-module.sh` — Debian kernel headers for module build
|
|
2. `auto/config` — `linux-image-${DEBIAN_KERNEL_ABI}` in the ISO
|
|
- NVIDIA modules go to staged `usr/local/lib/nvidia/` — NOT to `/lib/modules/<kver>/extra/`.
|
|
- `bee-gpu-burn` worker must be built against cached CUDA userspace headers from `build-cublas.sh`, not against random host-installed CUDA headers.
|
|
- The live ISO must ship `libcublas`, `libcublasLt`, and `libcudart` together with `libcuda` so tensor-core stress works without internet or package installs at boot.
|
|
- The source overlay in `iso/overlay/` is treated as immutable source. Build-time files are injected only into the staged overlay.
|
|
- Fast-path state lives outside the rsync-managed live-build workdir and is
|
|
accepted only when the heavy-input content hash and resolved kernel ABI
|
|
match the last successful full build. A failed full build never leaves a
|
|
valid completion marker. The workdir's `binary/` tree is preserved because
|
|
it is the source artifact for squashfs reuse.
|
|
- Bootloader menu text has two canonical sources only:
|
|
`config/bootloaders/grub-efi/grub.cfg` and
|
|
`config/bootloaders/isolinux/live.cfg.in`. `lib/bootloader.sh` renders those
|
|
templates after both full and fast builds; hooks do not append duplicate
|
|
menu entries.
|
|
- Build orchestration stays in `build.sh`; ISO validation, bootloader
|
|
rendering, fast-path/memtest recovery, and logging helpers live under
|
|
`iso/builder/lib/`. Run `iso/builder/test-build-libs.sh` after changing
|
|
those helpers or the canonical boot parameters.
|
|
- ISO filename, squashfs filename, ISO volume label, and the live system's hostname all derive from the same `easy-bee-<variant>-v<version>` scheme (`ISO_BASENAME`/`SQUASHFS_FILENAME`/`BEE_ISO_VOLUME`/`BEE_HOSTNAME` in `build.sh`) instead of the live-build default (`debian`). Keep new naming derived from `PROJECT_VERSION_EFFECTIVE`/`BUILD_VARIANT` in sync with this set rather than hardcoding a new scheme.
|
|
- Every live boot entry carries `udev.children_max=1`,
|
|
`intel_iommu=on`, `iommu.passthrough=0`, and
|
|
`efi=disable_early_pci_dma`. Only the single failsafe entry additionally
|
|
carries `pci=realloc iommu.strict=1`; `iommu=pt` is forbidden. The final-ISO
|
|
validator enforces this for both GRUB and isolinux.
|
|
- The live-build workdir under `dist/` is disposable; source files under `iso/builder/` stay clean.
|
|
- Container build requires `--privileged` because `live-build` uses mounts/chroots/loop devices during ISO assembly.
|
|
- On macOS / Docker Desktop, the builder still must run as `linux/amd64` so the shipped ISO binaries remain `amd64`.
|
|
- Operators must provision enough RAM to hold the full compressed live medium plus normal runtime overhead, because `toram` copies the entire read-only ISO payload into memory before the system reaches steady state.
|
|
|
|
## Post-boot smoke test
|
|
|
|
After booting a live ISO, run to verify all critical components:
|
|
|
|
```sh
|
|
ssh root@<ip> 'sh -s' < iso/builder/smoketest.sh
|
|
```
|
|
|
|
Exit code 0 = all required checks pass. All `FAIL` lines must be zero before shipping.
|
|
|
|
Key checks: NVIDIA modules loaded, `nvidia-smi` sees all GPUs, lib symlinks present,
|
|
systemd services running, audit completed with NVIDIA enrichment, LAN reachability.
|
|
|
|
Current validation state:
|
|
- local/libvirt VM boot path is validated for `systemd`, SSH, `bee audit`, `bee-network`, and Web UI startup
|
|
- real hardware validation is still required before treating the ISO as release-ready
|
|
|
|
## Overlay mechanism
|
|
|
|
`live-build` copies files from `config/includes.chroot/` into the ISO filesystem.
|
|
`build.sh` prepares a staged overlay, then syncs it into a temporary workdir's
|
|
`config/includes.chroot/` before running `lb build`.
|
|
|
|
## Collector flow
|
|
|
|
```
|
|
`bee audit` start
|
|
1. board collector (dmidecode -t 0,1,2)
|
|
2. cpu collector (dmidecode -t 4)
|
|
3. memory collector (dmidecode -t 17)
|
|
4. storage collector (lsblk -J, smartctl -j, nvme id-ctrl, nvme smart-log)
|
|
5. pcie collector (lspci -vmm -D, /sys/bus/pci/devices/)
|
|
6. psu collector (ipmitool fru + sdr — silent if no /dev/ipmi0)
|
|
7. nvidia enrichment (nvidia-smi — skipped if binary absent or driver not loaded)
|
|
8. TPM inventory (sysfs presence + `tpm2_getcap properties-fixed`; no state changes)
|
|
9. output JSON → /var/log/bee-audit.json
|
|
```
|
|
|
|
Every collector returns `nil, nil` on tool-not-found. Errors are logged, never fatal.
|
|
|
|
Acceptance flows:
|
|
- `bee sat nvidia` → diagnostic archive with `nvidia-smi -q` + `nvidia-bug-report` + lightweight `bee-gpu-burn`
|
|
- NVIDIA GPU burn-in can use either `bee-gpu-burn` or `bee-john-gpu-stress` (John the Ripper jumbo via OpenCL)
|
|
- `bee sat memory` → `memtester` archive
|
|
- `bee sat storage` → SMART/NVMe diagnostic archive and short self-test trigger where supported
|
|
- `bee` TPM Validate → read-only capabilities, PCR values, and existing self-test result; never starts `TPM2_SelfTest`
|
|
- SAT `summary.txt` now includes `overall_status` and per-job `*_status` values (`OK`, `FAILED`, `UNSUPPORTED`)
|
|
- `bee-gpu-burn` should prefer cuBLASLt GEMM load over the old integer/PTX burn path:
|
|
- Ampere: `fp16` + `fp32`/TF32 tensor-core load
|
|
- Ada / Hopper: add `fp8`
|
|
- Blackwell+: add `fp4`
|
|
- PTX fallback is only for missing cuBLASLt/userspace or unsupported narrow datatypes
|
|
- Runtime overrides:
|
|
- `BEE_MEMTESTER_SIZE_MB`
|
|
- `BEE_MEMTESTER_PASSES`
|
|
- NVIDIA Bandwidth SAT (`RunNvidiaBandwidthPack`, `dcgmi diag -r nvbandwidth`) in
|
|
Stress mode runs per resolved PCI NUMA node first, then all selected GPUs together
|
|
(`03-dcgmi-nvbandwidth-socket0.log`, `...-socket1.log`, `...-all.log`) --
|
|
see `decisions/2026-07-27-nvbandwidth-per-socket-split.md`. Single-node
|
|
systems (or systems where any GPU's NUMA node cannot be resolved) keep the
|
|
original single `NN-dcgmi-nvbandwidth.log` shape.
|
|
|
|
## SAT job output durability
|
|
|
|
```
|
|
runAcceptancePackCtx job loop (per satJob)
|
|
1. run the job's command; streamExecOutput writes each output line to the
|
|
job's log file as it arrives (not only when the process exits)
|
|
2. write the job's final log file (same content the live stream already
|
|
wrote, plus any health-check suffix)
|
|
3. call satJobBoundaryHook(jobName) if set
|
|
4. append run_at_utc / *_status to summary.txt
|
|
```
|
|
|
|
**Critical invariants:**
|
|
- DO NOT change `streamExecOutput` back to buffering output in memory and
|
|
writing the job's log file only once, after the command exits -- see
|
|
`decisions/2026-07-27-sat-live-output-and-blackbox-kick.md`. The live ISO's
|
|
export directory sits on a `toram` RAM-backed overlay (see Boot sequence
|
|
above), so a crash mid-command is otherwise unrecoverable for that job.
|
|
- `platform` never imports `app`; `SetJobBoundaryHook` is a plain
|
|
`func(string)` seam (same pattern as `satExecCommand`/`satStat`), not a
|
|
direct call into blackbox internals.
|
|
|
|
## Blackbox sync flow
|
|
|
|
```
|
|
bee-blackbox.service (separate process from bee-web/bee-audit)
|
|
1. discover enrolled removable-media targets every blackboxDiscoverInterval (2s)
|
|
2. per enrolled target, blackboxWorker.run():
|
|
a. syncCycle(): mount target, mirror exportDir -> removable media,
|
|
write README.md/manifest docs, fsync
|
|
b. record lastKickSeen = mtime of exportDir/.blackbox-kick
|
|
c. wait for: the adaptive flushPeriod timer (1-30s), OR
|
|
a stop signal, OR
|
|
a poll tick (blackboxKickPollInterval, 250ms) showing the kick
|
|
file is newer than lastKickSeen -- whichever comes first
|
|
3. on kick-triggered wake: sync immediately, do not wait out the
|
|
remaining flushPeriod
|
|
```
|
|
|
|
**Critical invariants:**
|
|
- `app.New()` wires `platform.SetJobBoundaryHook` to touch
|
|
`exportDir/.blackbox-kick` after every SAT job finishes. If a new SAT
|
|
execution path bypasses `runAcceptancePackCtx`'s job loop, it will not
|
|
trigger this kick -- data from that path still eventually reaches
|
|
blackbox via the adaptive timer, just not promptly.
|
|
- The adaptive `flushPeriod` logic (`adjustFlushPeriod`) is unchanged by the
|
|
kick mechanism -- the kick only short-circuits the *wait*, it does not
|
|
reset `flushPeriod` itself.
|
|
- DO NOT assume a local write under the live ISO's export directory is
|
|
durable on its own (RAM-backed overlay) -- blackbox's mirror to removable
|
|
media is the only real persistence boundary across a hard reset.
|
|
- `BuildSupportBundle` stages into a private `os.MkdirTemp` parent, not a
|
|
shared `os.TempDir()/bee-support-stage-<host>-<ts>` path. DO NOT go back to
|
|
a time-derived staging path: two builds in the same wall-clock second (two
|
|
operators, or an on-demand build racing the blackbox worker) then share one
|
|
tree and one's deferred `os.RemoveAll` truncates the other's archive.
|
|
|
|
## NVIDIA SAT Web UI flow
|
|
|
|
```
|
|
Web UI: Acceptance Tests page -> Run Test button
|
|
1. POST /api/sat/nvidia/run -> returns job_id
|
|
2. GET /api/sat/stream?job_id=... (SSE): streams stdout/stderr lines live
|
|
3. After completion: archive written to /appdata/bee/export/bee-sat/
|
|
summary.txt contains overall_status (OK / FAILED / UNSUPPORTED) and per-job status
|
|
```
|
|
|
|
## Run All (validate / check) flow
|
|
|
|
```
|
|
Web UI: "Run All" button -> POST /api/sat/run-all
|
|
body: operator intent only { stress_mode, amd_targets[], nvidia_gpu_indices[] }
|
|
server (handler.planSATRunAll):
|
|
1. always: cpu, memory, storage, pcie-link
|
|
2. tpm - only if App.TPMPresent() finds tpm_version_major=2
|
|
3. nvidia-config - if DetectGPUPresence().Nvidia || NvidiaInitializing
|
|
4. wait for NVIDIA enumeration: repeat fresh ListNvidiaGPUs queries until
|
|
at least one GPU is returned, NvidiaGSPMode=="gsp-stuck", or 75s
|
|
5. nvidia / nvidia-interconnect / nvidia-bandwidth / nvidia-pcie-bandwidth
|
|
(+ targeted-stress/power/pulse when stress_mode) - only once ready,
|
|
-i = App.ListNvidiaGPUs() indices (intersected with the requested subset)
|
|
6. amd / amd-mem / amd-bandwidth - if DetectGPUPresence().AMD and selected
|
|
response: { task_ids[], task_count, notes[] } (notes = what was skipped and why)
|
|
```
|
|
|
|
**Critical invariants:**
|
|
- Hardware presence, readiness, and which tasks to run are decided server-side.
|
|
DO NOT move this back into page JS (`satSelectedGPUIndices().length` gating):
|
|
a browser-cached empty GPU list then silently drops every GPU test. See
|
|
`decisions/2026-08-31-backend-driven-sat-planning.md`.
|
|
- `DetectGPUPresence` is the shared detection source (existing operational
|
|
vendor detection plus an lspci display-class fallback).
|
|
`/api/gpu/presence`, `/api/gpu/tools` and the planner all use it.
|
|
- `bee-gpu-burn` / `bee-john-gpu-stress` use `exec.CommandContext`: killed on job context cancel.
|
|
- Metric goroutine uses stopCh/doneCh pattern; main goroutine waits `<-doneCh` before reading rows (no mutex needed).
|
|
- SVG chart is fully offline: no JS, no external CSS, pure inline SVG.
|
|
- `RunNvidiaBandwidthPack` runs one all-GPU `nvbandwidth` pass in Validate; the
|
|
per-NUMA-node matrix is Stress-tier only (`fullMatrix` arg). See
|
|
`decisions/2026-08-31-nvbandwidth-validate-single-deep-matrix.md`.
|