A second NF5280M6 BMC-dump/live-CD pair showed RESTful version info
firmware versions carrying a build stamp ("08.05.01 (02/21/2024
16:51:27)") the live-CD does not, which churns a FIRMWARE_CHANGED event
on every source switch. cleanFirmwareVersion strips a trailing " (...)"
so BIOS/BMC match the bare live-CD version. See ADL-064 amendment.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011LffAvostt3uMkiUbVUiyM
109 KiB
10 — Architectural Decision Log (ADL)
Rule: Every significant architectural decision must be recorded here before or alongside the code change. This applies to humans and AI assistants alike.
Format: date · title · context · decision · consequences
ADL-001 — In-memory only state (no database)
Date: project start
Context: LOGPile is designed as a standalone diagnostic tool, not a persistent service.
Decision: All parsed/collected data lives in Server.result (in-memory). No database, no files written.
Consequences:
- Data is lost on process restart — intentional.
- Simple deployment: single binary, no setup required.
- JSON export is the persistence mechanism for users who want to save results.
ADL-002 — Vendor parser auto-registration via init()
Date: project start
Context: Need an extensible parser registry without a central factory function.
Decision: Each vendor parser registers itself in its package's init() function.
vendors/vendors.go holds blank imports to trigger registration.
Consequences:
- Adding a new parser requires only: implement interface + add one blank import.
- No central list to maintain (other than the import file).
go test ./...will include new parsers automatically.
ADL-003 — Highest-confidence parser wins
Date: project start
Context: Multiple parsers may partially match an archive (e.g. generic + specific vendor).
Decision: Run all parsers' Detect(), select the one returning the highest score (0–100).
Consequences:
- Generic fallback (score 15) only activates when no vendor parser scores higher.
- Parsers must be conservative with high scores (70+) to avoid false positives.
ADL-004 — Canonical hardware.devices as single source of truth
Date: v1.5.0
Context: UI tabs and Reanimator exporter were reading from different sub-fields of
AnalysisResult, causing potential drift.
Decision: Introduce hardware.devices as the canonical inventory repository.
All UI tabs and all exporters must read exclusively from this repository.
Consequences:
- Any UI vs Reanimator discrepancy is classified as a bug, not a "known difference".
- Deduplication logic runs once in the repository builder (serial → bdf → distinct).
- New hardware attributes must be added to canonical schema first, then mapped to consumers.
ADL-005 — No hardcoded PCI model strings; use pci.ids
Date: v1.5.0
Context: NVIDIA and other vendors release new GPU models frequently; hardcoded maps
required code changes for each new model ID.
Decision: Use the pciutils/pciids database (git submodule, embedded at build time).
PCI vendor/device ID → human-readable model name via lookup.
Consequences:
- New GPU models can be supported by updating
pci.idswithout code changes. make buildauto-syncspci.idsfrom submodule before compilation.- External override via
LOGPILE_PCI_IDS_PATHenv var.
ADL-006 — Reanimator export uses canonical hardware.devices (not raw sub-fields)
Date: v1.5.0
Context: Early Reanimator exporter read from Hardware.GPUs, Hardware.NICs, etc.
directly, diverging from UI data.
Decision: Reanimator exporter must use hardware.devices — the same source as the UI.
Exporter groups/filters canonical records by section; does not rebuild from sub-fields.
Consequences:
- Guarantees UI and export consistency.
- Exporter code is simpler — mainly a filter+map, not a data reconstruction.
ADL-007 — Documentation language is English
Date: 2026-02-20
Context: Codebase documentation was mixed Russian/English, reducing clarity for
international contributors and AI assistants.
Decision: All maintained project documentation (docs/bible/, README.md,
CLAUDE.md, and new technical docs) must be written in English.
Consequences:
- Bible is authoritative in English.
- AI assistants get consistent, unambiguous context.
ADL-008 — Bible is the single source of truth for architecture docs
Date: 2026-02-23
Context: Architecture information was duplicated across README.md, CLAUDE.md,
and the Bible, creating drift risk and stale guidance for humans and AI agents.
Decision: Keep architecture and technical design documentation only in docs/bible/.
Top-level README.md and CLAUDE.md must remain minimal pointers/instructions.
Consequences:
- Reduces documentation drift and duplicate updates.
- AI assistants are directed to one authoritative source before making changes.
- Documentation updates that affect architecture must include Bible changes (and ADL entries when significant).
ADL-009 — Redfish analysis is performed from raw snapshot replay (unified tunnel)
Date: 2026-02-24
Context: Live Redfish collection and raw export re-analysis used different parsing paths,
which caused drift and made bug fixes difficult to validate consistently.
Decision: Redfish live collection must produce a raw_payloads.redfish_tree snapshot first,
then run the same replay analyzer used for imported raw exports.
Consequences:
- Same
redfish_treeinput produces the same parsed result in live and offline modes. - Debugging parser issues can be done against exported raw bundles without live BMC access.
- Snapshot completeness becomes critical; collector seeds/limits are part of analyzer correctness.
ADL-010 — Raw export is a self-contained re-analysis package (not a final result dump)
Date: 2026-02-24
Context: Exporting only normalized AnalysisResult loses raw source fidelity and prevents
future parser improvements from being applied to already collected data.
Decision: Export Raw Data produces a self-contained raw package (JSON or ZIP bundle)
that the application can reopen and re-analyze. Parsed data in the package is optional and not
the source of truth on import.
Consequences:
- Re-opening an export always re-runs analysis from raw source (
redfish_treeor uploaded file bytes). - Raw bundles include collection context and diagnostics for debugging (
collect.log,parser_fields.json). - Endpoint compatibility is preserved (
/api/export/json) while actual payload format may be a bundle.
ADL-011 — Redfish snapshot crawler is bounded, prioritized, and failure-tolerant
Date: 2026-02-24 Context: Full Redfish trees on modern GPU systems are large, noisy, and contain many vendor-specific or non-fetchable links. Unbounded crawling and naive queue design caused hangs and incomplete snapshots. Decision: Use a bounded snapshot crawler with:
- explicit document cap (
LOGPILE_REDFISH_SNAPSHOT_MAX_DOCS) - priority seed paths (PCIe/Fabrics/Firmware/Storage/PowerSubsystem/ThermalSubsystem)
- normalized
@odata.idpaths (strip#fragment) - noisy expected error filtering (404/405/410/501 hidden from UI)
- queue capacity sized to crawl cap to avoid producer/consumer deadlock Consequences:
- Snapshot collection remains stable on large BMC trees.
- Most high-value inventory paths are reached before the cap.
- UI progress remains useful while debug logs retain low-level fetch failures.
ADL-012 — Vendor-specific storage inventory probing is allowed as fallback
Date: 2026-02-24
Context: Some Supermicro BMCs expose empty standard Storage/.../Drives collections while
real disk inventory exists under vendor-specific Disk.Bay endpoints and enclosure links.
Decision: When standard drive collections are empty, collector/replay may probe vendor-style
.../Drives/Disk.Bay.* endpoints and follow Storage.Links.Enclosures[*] to recover physical drives.
Consequences:
- Higher storage inventory coverage on Supermicro HBA/HA-RAID/MRVL/NVMe backplane implementations.
- Replay must mirror the same probing behavior to preserve deterministic results.
- Probing remains bounded (finite candidate set) to avoid runaway requests.
ADL-013 — PowerSubsystem is preferred over legacy Power on newer Redfish implementations
Date: 2026-02-24
Context: X14+/newer Redfish implementations increasingly expose authoritative PSU data in
PowerSubsystem/PowerSupplies, while legacy /Power may be incomplete or schema-shifted.
Decision: Prefer Chassis/*/PowerSubsystem/PowerSupplies as the primary PSU source and use
legacy Chassis/*/Power as fallback.
Consequences:
- Better compatibility with newer BMC firmware generations.
- Legacy systems remain supported without special-case collector selection.
- Snapshot priority seeds must include
PowerSubsystemresources.
ADL-014 — Threshold logic lives on the server; UI reflects status only
Date: 2026-02-24 Context: Duplicating threshold math in frontend and backend creates drift and inconsistent highlighting (e.g. PSU mains voltage range checks). Decision: Business threshold evaluation (e.g. PSU voltage nominal range) must be computed on the server; frontend only renders status/flags returned by the API. Consequences:
- Single source of truth for threshold policies.
- UI can evolve visually without re-implementing domain logic.
- API payloads may carry richer status semantics over time.
ADL-015 — Supermicro crashdump archive parser removed from active registry
Date: 2026-03-01
Context: The Supermicro crashdump parser (SMC Crash Dump Parser) produced low-value
results for current workflows and was explicitly rejected as a supported archive path.
Decision: Remove supermicro vendor parser from active registration and project source.
Do not include it in /api/parsers output or parser documentation matrix.
Consequences:
- Supermicro crashdump archives (
CDump.txtformat) are no longer parsed by a dedicated vendor parser. - Such archives fall back to other matching parsers (typically
generic) unless a new replacement parser is added. - Reintroduction requires a new parser package and an explicit registry import in
vendors/vendors.go.
ADL-016 — Device-bound firmware must not appear in hardware.firmware
Date: 2026-03-01
Context: Dell TSR DCIM_SoftwareIdentity lists firmware for every component (NICs,
PSUs, disks, backplanes) in addition to system-level firmware. Naively importing all entries
into Hardware.Firmware caused device firmware to appear twice in Reanimator: once in the
device's own record and again in the top-level firmware list.
Decision:
Hardware.Firmwarecontains only system-level firmware (BIOS, BMC/iDRAC, CPLD, Lifecycle Controller, storage controllers, BOSS).- Device-bound entries (NIC, PSU, Disk, Backplane, GPU) must not be added to
Hardware.Firmware. - Parsers must store the FQDD (or equivalent slot identifier) in
FirmwareInfo.Descriptionso the Reanimator exporter can filter by FQDD prefix. - The exporter's
isDeviceBoundFirmwareFQDD()function performs this filter. Consequences: - Any new parser that ingests a per-device firmware inventory must follow the same rule.
- Device firmware is accessible only via the device's own record, not the firmware list.
ADL-017 — Vendor-embedded MAC addresses must be stripped from model name fields
Date: 2026-03-01
Context: Dell TSR embeds MAC addresses directly in ProductName and ElementName
fields (e.g. "NVIDIA ConnectX-6 Lx 2x 25G SFP28 OCP3.0 SFF - C4:70:BD:DB:56:08").
This caused model names to contain MAC addresses in NIC model, NIC firmware device name,
and potentially other fields.
Decision: Strip any - XX:XX:XX:XX:XX:XX suffix from all model/name string fields
at parse time before storing in any model struct. Use the regex
\s+-\s+([0-9A-Fa-f]{2}:){5}[0-9A-Fa-f]{2}$.
Consequences:
- Model names are clean and consistent across all devices.
- All parsers must apply this stripping to any field used as a device name or model.
- Confirmed affected fields in Dell:
DCIM_NICView.ProductName,DCIM_SoftwareIdentity.ElementName.
ADL-018 — NVMe bay probe must be restricted to storage-capable chassis types
Date: 2026-03-12
Context: shouldAdaptiveNVMeProbe was introduced in 2fa4a12 to recover NVMe drives on
Supermicro BMCs that expose empty Drives collections but serve disks at direct Disk.Bay.N
paths. The function returns true for any chassis with an empty Members array. On
Supermicro HGX systems (SYS-A21GE-NBRT and similar) ~35 sub-chassis (GPU, NVSwitch,
PCIeRetimer, ERoT, IRoT, BMC, FPGA) all carry ChassisType=Module/Component/Zone and
expose empty /Drives collections. Without filtering, each triggered 384 HTTP requests →
13 440 requests ≈ 22 minutes of pure I/O waste per collection.
Decision: Before probing Disk.Bay.N candidates for a chassis, check its ChassisType
via chassisTypeCanHaveNVMe. Skip if type is Module, Component, or Zone. Keep probing
for Enclosure, RackMount, and any unrecognised type (fail-safe).
Consequences:
- On HGX systems post-probe NVMe goes from ~22 min to effectively zero.
- NVMe backplane recovery (
Enclosuretype) is unaffected. - Any new chassis type that hosts NVMe storage is covered by the default
truepath. chassisTypeCanHaveNVMeand the candidate-selection loop must have unit tests covering both the excluded types and the storage-capable types (seeTestChassisTypeCanHaveNVMeandTestNVMePostProbeSkipsNonStorageChassis).
ADL-019 — Redfish post-probe recovery is profile-owned acquisition policy
Date: 2026-03-18
Context: Numeric collection post-probe and direct NVMe Disk.Bay recovery were still
controlled by collector-core heuristics, which kept platform-specific acquisition behavior in
redfish.go and made vendor/topology refactoring incomplete.
Decision: Move expensive Redfish post-probe enablement into profile-owned acquisition policy.
The collector core may execute bounded post-probe loops, but profiles must explicitly enable:
- numeric collection post-probe
- direct NVMe
Disk.Bayrecovery - sensor collection post-probe Consequences:
- Generic collector flow no longer implicitly turns on storage/NVMe recovery for every platform.
- Supermicro-specific direct NVMe recovery and generic numeric collection recovery are now regression-tested through profile fixtures.
- Future platform storage/post-probe behavior must be added through profile tuning, not new
vendor-shaped
ifbranches in collector core.
ADL-020 — Redfish critical plan-B activation is profile-owned recovery policy
Date: 2026-03-18
Context: critical plan-B and profile plan-B were still effectively always-on collector
behavior once paths were present, including critical collection member retry and slow numeric
child probing. That kept acquisition recovery semantics in redfish.go instead of the profile
layer.
Decision: Move plan-B activation into profile-owned recovery policy. Profiles must explicitly
enable:
- critical collection member retry
- slow numeric probing during critical plan-B
- profile-specific plan-B pass Consequences:
- Recovery behavior is now observable in raw Redfish diagnostics alongside other tuning.
- Generic/fallback recovery remains available through profile policy instead of implicit collector defaults.
- Future platform-specific plan-B behavior must be introduced through profile tuning and tests, not through new unconditional collector branches.
ADL-021 — Extra discovered-path storage seeds must be profile-scoped, not core-baseline
Date: 2026-03-18
Context: The collector core baseline seed list still contained storage-specific discovered-path
suffixes such as SimpleStorage and Storage/IntelVROC/*. These are useful on some platforms,
but they are acquisition extensions layered on top of discovered Systems/* resources, not part
of the minimal vendor-neutral Redfish baseline.
Decision: Move such discovered-path expansions into profile-owned scoped path policy. The
collector core keeps the vendor-neutral baseline; profiles may add extra system/chassis/manager
suffixes that are expanded over discovered members during acquisition planning.
Consequences:
- Platform-shaped storage discovery no longer lives in
redfish.gobaseline seed construction. - Extra discovered-path branches are visible in plan diagnostics and fixture regression tests.
- Future model/vendor storage path expansions must be added through scoped profile policy instead of editing the shared baseline seed list.
ADL-022 — Adaptive prefetch eligibility is profile-owned policy
Date: 2026-03-18
Context: The adaptive prefetch executor was still driven by hardcoded include/exclude path
rules in redfish.go. That made GPU/storage/network prefetch shaping part of collector-core
knowledge rather than profile-owned acquisition policy.
Decision: Move prefetch eligibility rules into profile tuning. The collector core still runs
adaptive prefetch, but profiles provide:
IncludeSuffixesfor critical paths eligible for prefetchExcludeContainsfor path shapes that must never be prefetched Consequences:- Prefetch behavior is now visible in raw Redfish diagnostics and test fixtures.
- Platform- or topology-specific prefetch shaping no longer requires editing collector-core string lists.
- Future prefetch tuning must be introduced through profiles and regression tests.
ADL-023 — Core critical baseline is roots-only; critical shaping is profile-owned
Date: 2026-03-18
Context: redfishCriticalEndpoints(...) still encoded a broad set of system/chassis/manager
critical branches directly in collector core. This mixed minimal crawl invariants with profile-
specific acquisition shaping.
Decision: Reduce collector-core critical baseline to vendor-neutral roots only:
/redfish/v1- discovered
Systems/* - discovered
Chassis/* - discovered
Managers/*
Profiles now own additional critical shaping through:
- scoped critical suffix policy for discovered resources
- explicit top-level
CriticalPathsConsequences: - Critical inventory breadth is now explained by the acquisition plan, not hidden in collector helper defaults.
- Generic profile still provides the previous broad critical coverage, so behavior stays stable.
- Future critical-path tuning must be implemented in profiles and regression-tested there.
ADL-024 — Live Redfish execution plans are resolved inside redfishprofile
Date: 2026-03-18
Context: Even after moving seeds, scoped paths, critical shaping, recovery, and prefetch
policy into profiles, redfish.go still manually merged discovered resources with those policy
fragments. That left acquisition-plan resolution logic in collector core.
Decision: Introduce redfishprofile.ResolveAcquisitionPlan(...) as the boundary between
profile planning and collector execution. redfishprofile now resolves:
- baseline seeds
- baseline critical roots
- scoped path expansions
- explicit profile seed/critical/plan-B paths
The collector core consumes the resolved plan and executes it. Consequences:
- Acquisition planning logic is now testable in
redfishprofilewithout going through the live collector. redfish.gono longer owns path-resolution helpers for seeds/critical planning.- This creates a clean next step toward true per-profile acquisition hooks beyond static policy fragments.
ADL-025 — Post-discovery acquisition refinement belongs to profile hooks
Date: 2026-03-18
Context: Some acquisition behavior depends not only on vendor/model hints, but on what the
lightweight Redfish discovery actually returned. Static absolute path lists in profile plans are
too rigid for such cases and reintroduce guessed platform knowledge.
Decision: Add a post-discovery acquisition refinement hook to Redfish profiles. Profiles may
mutate the resolved execution plan after discovered Systems/*, Chassis/*, and Managers/*
are known.
First concrete use:
- MSI now derives GPU chassis seeds and
.../Sensorscritical/plan-B paths from discoveredChassis/GPU*resources instead of hardcodedGPU1..GPU4absolute paths in the static plan. Additional use: - Supermicro now derives
UpdateService/Oem/Supermicro/FirmwareInventorycritical/plan-B paths from resource hints instead of carrying that absolute path in the static plan. Additional use: - Dell now derives
Managers/iDRAC.Embedded.*acquisition paths from discovered manager resources instead of carryingManagers/iDRAC.Embedded.1as a static absolute path. Consequences: - Profile modules can react to actual discovery results without pushing conditional logic back
into
redfish.go. - Diagnostics still show the final refined plan because the collector stores the refined plan, not only the pre-refinement template.
- Future vendor-specific discovery-dependent acquisition behavior should be implemented through this hook rather than new collector-core branches.
ADL-026 — Replay analysis uses a resolved profile plan, not ad-hoc directives only
Date: 2026-03-18
Context: Replay still relied on a flat AnalysisDirectives struct assembled centrally,
while vendor-specific conditions often depended on the actual snapshot shape. That made analysis
behavior harder to explain and kept too much vendor logic in generic replay collectors.
Decision: Introduce redfishprofile.ResolveAnalysisPlan(...) for replay. The resolved
analysis plan contains:
- active match result
- resolved analysis directives
- analysis notes explaining snapshot-aware hook activation
Profiles may refine this plan using the snapshot and discovered resources before replay collectors run.
First concrete uses:
- MSI enables processor-GPU fallback and MSI chassis lookup only when the snapshot actually
contains GPU processors and
Chassis/GPU* - HGX enables processor-GPU alias fallback from actual HGX/GPU_SXM topology signals in the snapshot
- Supermicro enables NVMe backplane and known-controller recovery from actual snapshot paths Consequences:
- Replay behavior is now closer to the acquisition architecture: a resolved profile plan feeds the executor.
redfish_analysis_planis stored in raw payload metadata for offline debugging.- Future analysis-side vendor logic should move into profile refinement hooks instead of growing the central directive builder.
ADL-027 — Replay GPU/storage executors consume resolved analysis plans
Date: 2026-03-18
Context: Even after introducing ResolveAnalysisPlan(...), replay GPU/storage collectors still
accepted a raw AnalysisDirectives struct. That preserved an implicit shortcut from the old design
and weakened the plan/executor boundary.
Decision: Replay GPU/storage executors now accept redfishprofile.ResolvedAnalysisPlan
directly. The executor reads resolved directives from the plan instead of being passed a standalone
directive bundle.
Consequences:
- GPU and storage replay execution now follows the same architectural pattern as acquisition: resolve plan first, execute second.
- Future profile-owned execution helpers can use plan notes or additional resolved fields without changing the executor API again.
- Remaining replay areas should migrate the same way instead of continuing to accept raw directive structs.
ADL-019 — isDeviceBoundFirmwareName must cover vendor-specific naming patterns per vendor
Date: 2026-03-12
Context: isDeviceBoundFirmwareName was written to filter Dell-style device firmware names
("GPU SomeDevice", "NIC OnboardLAN"). When Supermicro Redfish FirmwareInventory was added
(6c19a58), no Supermicro-specific patterns were added. Supermicro names a NIC entry
"NIC1 System Slot0 AOM-DP805-IO" — a digit follows the type prefix directly, bypassing the
"nic " (space-terminated) check. 29 device-bound entries leaked into hardware.firmware on
SYS-A21GE-NBRT (HGX B200). Commit 9c5512d attempted a fix by adding _fw_gpu_ patterns,
but checked DeviceName which contains "Software Inventory" (from the Redfish Name field),
not the firmware inventory ID. The patterns were dead code from the moment they were committed.
Decision:
isDeviceBoundFirmwareNamemust be extended for each new vendor whose FirmwareInventory naming convention differs from the existing patterns.- When adding HGX/Supermicro patterns, check that the pattern matches the field value that
collectFirmwareInventoryactually stores — trace the data path from Redfish doc toFirmwareInfo.DeviceNamebefore writing the condition. TestIsDeviceBoundFirmwareNamemust contain at least one case per vendor format. Consequences:- New vendors with FirmwareInventory support require a test covering both device-bound names (must return true) and system-level names (must return false) before the code ships.
- The dead
_fw_gpu_/_fw_nvswitch_/_inforom_gpu_patterns were replaced with correct prefix+digit checks ("gpu" + digit,"nic" + digit) and explicit string checks ("nvmecontroller","power supply","software inventory").
ADL-020 — Dell TSR device-bound firmware filtered via FQDD; InfiniBand routed to NetworkAdapters
Date: 2026-03-15
Context: Dell TSR sysinfo_DCIM_SoftwareIdentity.xml lists firmware for every installed
component. parseSoftwareIdentityXML dumped all of these into hardware.firmware without
filtering, so device-bound entries such as "Mellanox Network Adapter" (FQDD InfiniBand.Slot.1-1)
and "PERC H755 Front" (FQDD RAID.SL.3-1) appeared in the reanimator export alongside system
firmware like BIOS and iDRAC. Confirmed on PowerEdge R6625 (8VS2LG4).
Additionally, DCIM_InfiniBandView was not handled in the parser switch, so Mellanox ConnectX-6
appeared only as a PCIe device with model: "16x or x16" (from DataBusWidth fallback).
parseControllerView called addFirmware with description "storage controller" instead of the
FQDD, so the FQDD-based filter in the exporter could not remove it.
Decision:
isDeviceBoundFirmwareFQDDextended with"infiniband."and"fc."prefixes;"raid.backplane."broadened to"raid."to coverRAID.SL.*,RAID.Integrated.*, etc.DCIM_InfiniBandViewrouted toparseNICView→ device appears asNetworkAdapterwith correct firmware, MAC address, and VendorID/DeviceID."InfiniBand."added topcieFQDDNoisePrefixto suppress the duplicateDCIM_PCIDeviceViewentry (DataBusWidth-only, no useful data).parseControllerViewnow passesfqddas theaddFirmwaredescription so the FQDD filter removes the entry in the exporter.parsePCIeDeviceViewnow prioritisesprops["description"](chip model, e.g."MT28908 Family [ConnectX-6]") overprops["devicedescription"](location string) forpcie.Description.convertPCIeDevicesmodel fallback order:PartNumber → Description → DeviceClass.
Consequences:
hardware.firmwarecontains only system-level entries; NIC/RAID/storage-controller firmware lives on the respective device record.TestParseDellInfiniBandViewandTestIsDeviceBoundFirmwareFQDDguard the regression.- Any future Dell TSR device class whose FQDD prefix is not yet in the prefix list may still leak;
extend
isDeviceBoundFirmwareFQDDand add a test case when encountered.
ADL-021 — pci.ids enrichment: chip model and vendor resolved from PCI IDs when source data is generic or missing
Date: 2026-03-15
Context:
Dell TSR DCIM_InfiniBandView.ProductName reports a generic marketing name ("Mellanox Network
Adapter") instead of the precise chip identifier ("MT28908 Family [ConnectX-6]"). The actual
chip model is available in pci.ids by VendorID:DeviceID (15B3:101B). Vendor name may also be
absent when no VendorName / Manufacturer property is present.
The general rule was established: if model is not found in source data but PCI IDs are known,
resolve model from pci.ids. This rule applies broadly across all export paths.
Decision (two-layer enrichment):
- Parser layer (Dell,
parseNICView): WhenVendorID != 0 && DeviceID != 0, preferpciids.DeviceName(vendorID, deviceID)over the product name from logs. This makes the chip identifier the primary model for NIC/InfiniBand adapters (more specific than marketing name). FillVendorfrompciids.VendorName(vendorID)when the vendor field is otherwise empty. Same fallback applied inparsePCIeDeviceViewfor emptyDescription. - Exporter layer (
convertPCIeFromDevices): General rule — whend.Model == ""after all legacy fallbacks andVendorID != 0 && DeviceID != 0, setmodel = pciids.DeviceName(...). Also fill emptymanufacturerfrompciids.VendorName(...). This covers all parsers/sources.
Consequences:
- Mellanox InfiniBand slot now reports
model: "MT28908 Family [ConnectX-6]"andmanufacturer: "Mellanox Technologies"in the reanimator export. - For NICs where pci.ids has no entry, the original product name is kept (pci.ids returns "").
TestParseDellInfiniBandViewasserts the model and vendor from pci.ids.
ADL-022 — CPUAffinity parsed into NUMANode for PCIe, NIC, and controller devices
Date: 2026-03-15
Context:
Dell TSR DCIM view classes report CPUAffinity for NIC, InfiniBand, PCIe, and controller
devices. Values are "1", "2" (NUMA node index), or "Not Applicable" (for devices that bridge
both CPUs or have no CPU affinity). This data is needed for topology-aware diagnostics.
Decision:
- Add
NUMANode int(JSON:"numa_node,omitempty") tomodels.PCIeDevice,models.NetworkAdapter,models.HardwareDevice, andReanimatorPCIe. - Parse from
props["cpuaffinity"]usingparseIntLoose: numeric values ("1", "2") map directly; "Not Applicable" returns 0 (omitted viaomitempty). - Thread through
buildDevicesFromLegacy(PCIe and NIC sections) andconvertPCIeFromDevices. parseControllerViewalso parses CPUAffinity since RAID controllers have NUMA affinity.
Consequences:
numa_node: 1or2appears in reanimator export for devices with known affinity.- Value 0 / absent means "not reported" — covers both "Not Applicable" and sources that don't provide CPUAffinity at all.
TestParseDellCPUAffinityverifies numeric values parsed correctly and "Not Applicable"→0.
ADL-023 — Reanimator export must match ingest contract exactly
Date: 2026-03-15
Context:
LOGPile's Reanimator export had drifted from the strict ingest contract. It emitted fields that
Reanimator does not currently accept (status_at_collection, numa_node),
while missing fields and sections now present in the contract (hardware.sensors,
pcie_devices[].mac_addresses). Memory export rules also diverged from the ingest side: empty or
serial-less DIMMs were still exported.
Decision:
- Treat the Reanimator ingest contract as the authoritative schema for
GET /api/export/reanimator. - Emit only fields present in the current upstream contract revision.
- Add
hardware.sensors,pcie_devices[].mac_addresses,pcie_devices[].numa_node, and upstream-approved component telemetry/health fields. - Leave out fields that are still not part of the upstream contract.
- Map internal
source_type=archiveto externalsource_type=logfile. - Skip memory entries that are empty, not present, or missing serial numbers.
- Generate CPU and PCIe serials only in the forms allowed by the contract.
- Mirror the applied contract in
bible-local/docs/hardware-ingest-contract.md.
Consequences:
- Some previously exported diagnostic fields are intentionally dropped from the Reanimator payload until the upstream contract adds them.
- Internal models may retain richer fields than the current export schema.
hardware.devicesis canonical only after merge with legacy hardware slices; partial parser-owned canonical records must not hide CPUs, memory, storage, NICs, or PSUs still stored in legacy fields.- CSV and Reanimator exports must use the same merged canonical inventory to avoid divergent export contents across surfaces.
- Future exporter changes must update both the code and the mirrored contract document together.
ADL-024 — Component presence is implicit; Redfish linked metrics are part of replay correctness
Date: 2026-03-15
Context:
The upstream ingest contract allows present, but current export semantics do not need to send
present=true for populated components. At the same time, several important Redfish component
telemetry fields were only available through linked metric resources such as ProcessorMetrics,
MemoryMetrics, and DriveMetrics. Without collecting and replaying these linked documents,
live collection and raw snapshot replay still underreported component health fields.
Decision:
- Do not serialize
present=truein Reanimator export. Presence is represented by the presence of the component record itself. - Do not export component records marked
present=false. - Interpret CPU
firmwarein Reanimator payload as CPU microcode. - Treat Redfish linked metric resources
ProcessorMetrics,MemoryMetrics,DriveMetrics,EnvironmentMetrics, and genericMetricsas part of analyzer correctness when they are linked from component resources. - Replay logic must merge these linked metric resources back into CPU, memory, storage, PCIe, GPU,
NIC, and PSU component
Detailsthe same way live collection expects them to be used.
Consequences:
- Reanimator payloads are smaller and avoid redundant
present=truenoise while still excluding empty slots and absent components. - Any future exporter change that reintroduces serialized component presence needs an explicit contract review.
- Raw Redfish snapshot completeness now includes linked per-component metric resources, not only top-level inventory members.
- CPU microcode is no longer expected in top-level
hardware.firmware; it belongs on the CPU component record.
ADL-025 — Missing serial numbers must remain absent in Reanimator export
Date: 2026-03-15 Context: LOGPile previously generated synthetic serial numbers for components that had no real serial in source data, especially CPUs and PCIe-class devices. This made the payload look richer, but the serials were not authoritative and could mislead downstream consumers. Reanimator can already accept missing serials and generate its own internal fallback identifiers when needed.
Decision:
- Do not synthesize fake serial numbers in LOGPile's Reanimator export.
- If a component has no real serial in parsed source data, export the serial field as absent.
- This applies to CPUs, PCIe devices, GPUs, NICs, and any other component class unless an upstream contract explicitly requires a deterministic exporter-generated identifier.
- Any fallback serial generation defined by the upstream contract is ingest-side Reanimator behavior, not LOGPile exporter behavior.
Consequences:
- Exported payloads carry only source-backed serial numbers.
- Fake identifiers such as
BOARD-...-CPU-...or synthetic PCIe serials are no longer considered acceptable exporter behavior. - Any future attempt to reintroduce generated serials requires an explicit contract review and a new ADL entry.
ADL-026 — Live Redfish collection uses explicit preflight host-power confirmation
Date: 2026-03-15 Context: Live Redfish inventory can be incomplete when the managed host is powered off. At the same time, LOGPile must not silently power on a host without explicit user choice. The collection workflow therefore needs a preflight step that verifies connectivity, shows current host power state to the user, and only powers on the host when the user explicitly chose that path.
Decision:
- Add a dedicated live preflight API step before collection starts.
- UI first runs connectivity and power-state check, then offers:
- collect as-is
- power on and collect
- if the host is off and the user does not answer within 5 seconds, default to collecting without powering the host on
- Redfish collection may power on the host only when the request explicitly sets
power_on_if_host_off=true - when LOGPile powers on the host for collection, it must try to power the host back off after collection completes
- if LOGPile did not power the host on itself, it must never power the host off
- all preflight and power-control steps must be logged into the collection log and therefore into the raw-export bundle
Consequences:
- Live collection becomes a two-step UX: probe first, collect second.
- Raw bundles preserve operator-visible evidence of power-state decisions and power-control attempts.
- Power-on failures do not block collection entirely; they only downgrade completeness expectations.
ADL-027 — Sensors without numeric readings are not exported
Date: 2026-03-15 Context: Some parsed sensor records carry only a name, unit, or status, but no actual numeric reading. Such records are not useful as telemetry in Reanimator export and create noisy, low-value sensor lists.
Decision:
- Do not export temperature, power, fan, or other sensor records unless they carry a real numeric measurement value.
- Presence of a sensor name or health/status alone is not sufficient for export.
Consequences:
- Exported sensor groups contain only actionable telemetry.
- Parsers and collectors may still keep non-numeric sensor artifacts internally for diagnostics, but Reanimator export must filter them out.
ADL-028 — Reanimator PCIe export excludes storage endpoints and synthetic serials
Date: 2026-03-15
Context:
Some Redfish and archive sources expose NVMe drives both as storage inventory and as PCIe-visible
endpoints. Exporting such drives in both hardware.storage and hardware.pcie_devices creates
duplicates without adding useful topology value. At the same time, PCIe-class export still had old
fallback behavior that generated synthetic serial numbers when source serials were absent.
Decision:
- Export disks and NVMe drives only through
hardware.storage. - Do not export storage endpoints as
hardware.pcie_devices, even if the source inventory exposes them as PCIe/NVMe devices. - Keep real PCIe storage controllers such as RAID and HBA adapters in
hardware.pcie_devices. - Do not synthesize PCIe/GPU/NIC serial numbers in LOGPile; missing serials stay absent.
- Treat placeholder names such as
Network Device Viewas non-authoritative and prefer resolved device names when stronger data exists.
Consequences:
- Reanimator payloads no longer duplicate NVMe drives between storage and PCIe sections.
- PCIe export remains topology-focused while storage export remains component-focused.
- Missing PCIe-class serials no longer produce fake
BOARD-...-PCIE-...identifiers.
ADL-029 — Local exporter guidance tracks upstream contract v2.7 terminology
Date: 2026-03-15
Context:
The upstream Reanimator hardware ingest contract moved to v2.7 and clarified several points that
matter for LOGPile documentation: ingest-side serial fallback rules, canonical PCIe addressing via
slot, the optional event_logs section, and the shared manufactured_year_week field.
Decision:
- Keep the local mirrored contract file as an exact copy of the upstream
v2.7document. - Describe CPU/PCIe serial fallback as Reanimator ingest behavior, not LOGPile exporter behavior.
- Treat
pcie_devices.slotas the canonical address on the LOGPile side as well;bdfmay remain an internal fallback/dedupe key but is not serialized in the payload. - Export
event_logsonly from normalized parser/collector events that can be mapped to contract sourceshost/bmc/redfishwithout synthesizing message content. - Export
manufactured_year_weekonly as a reliable passthrough when a parser/collector already extracted a validYYYY-Wwwvalue.
Consequences:
- Local bible wording no longer conflicts with upstream contract terminology.
- Reanimator payloads use contract-native PCIe addressing and no longer expose
bdfas a parallel coordinate. - LOGPile event export remains strictly source-derived; internal warnings such as LOGPile analysis
notes do not leak into Reanimator
event_logs.
ADL-030 — Audit result rendering is delegated to embedded reanimator/chart
Date: 2026-03-16 Context: LOGPile already owns file upload, Redfish collection, archive parsing, normalization, and Reanimator export. Maintaining a second host-side audit renderer for the same data created presentation drift and duplicated UI logic.
Decision:
- Use vendored
reanimator/chartas the only audit result viewer. - Keep LOGPile responsible for service flows: upload, live collection, batch convert, raw export, Reanimator export, and parse-error reporting.
- Render the current dataset by converting it to Reanimator JSON and passing that snapshot to
embedded
chartunder/chart/current.
Consequences:
- Reanimator JSON becomes the single presentation contract for the audit surface.
- The host UI becomes a service shell around the viewer instead of maintaining its own field-by-field tabs.
internal/chartmust be updated explicitly as a git submodule when the viewer changes.
ADL-031 — Redfish uses profile-driven acquisition and unified ingest entrypoints
Date: 2026-03-17 Context: Redfish collection had accumulated platform-specific probing in the shared collector path, while upload and raw-export replay still entered analysis through direct handler branches. This made vendor/model tuning harder to contain and increased regression risk when one topology needed a special acquisition strategy.
Decision:
- Introduce
internal/ingest.Serviceas the internal source-family entrypoint for archive parsing and Redfish raw replay. - Introduce
internal/collector/redfishprofile/for Redfish profile matching and modular hooks. - Split Redfish behavior into coordinated phases:
- acquisition planning during live collection
- analysis hooks during snapshot replay
- Use score-based profile matching. If confidence is low, enter fallback acquisition mode and aggregate only safe additive profile probes.
- Allow profile modules to provide bounded acquisition tuning hints such as crawl cap, prefetch behavior, and expensive post-probe toggles.
- Allow profile modules to own model-specific
CriticalPathsand boundedPlanBPathsso vendor recovery targets stop leaking into the collector core. - Expose Redfish profile matching as structured diagnostics during live collection: logs must contain all module scores, and collect job status must expose active modules for the UI.
Consequences:
- Server handlers stop owning parser-vs-replay branching details directly.
- Vendor/model-specific Redfish logic gets an explicit module boundary.
- Unknown-vendor Redfish collection becomes slower but more complete by design.
- Tactical Redfish fixes should move into profile modules instead of widening generic replay logic.
- Repo-owned compact fixtures under
internal/collector/redfishprofile/testdata/, derived from representative raw-export snapshots, are used to lock profile matching and acquisition tuning for known MSI and Supermicro-family shapes.
ADL-032 — MSI ghost GPU filter: exclude GPUs with temperature=0 on powered-on host
Date: 2026-03-18
Context:
MSI/AMI BMC caches GPU inventory from the host via Host Interface (in-band). When GPUs are
removed without a reboot the old entries remain in Chassis/GPU* and
Systems/Self/Processors/GPU* with Status.Health: OK, State: Enabled. The BMC has no
out-of-band mechanism to detect physical absence. A physically present GPU always reports
an ambient temperature (>0°C) even when idle; a stale cached entry returns Reading: 0.
Decision:
- Add
EnableMSIGhostGPUFilterdirective (enabled by MSI profile'srefineAnalysisalongsideEnableProcessorGPUFallback). - In
collectGPUsFromProcessors: for each processor GPU, resolve its chassis path and readChassis/GPU{n}/Sensors/GPU{n}_Temperature. IfPowerState=OnandReading=0→ skip. - Filter only applies when host is powered on; when host is off all temperatures are 0 and the signal is ambiguous.
Consequences:
- Ghost GPUs from previous hardware configurations no longer appear in the inventory.
- Filter is MSI-profile-owned and does not affect HGX, Supermicro, or generic paths.
- Any new MSI GPU chassis that uses a different temperature sensor path will bypass the filter (safe default: include rather than wrongly exclude).
ADL-033 — Reanimator export collected_at uses inventory LastModifiedTime with 30-day fallback
Date: 2026-03-18
Context:
For Redfish sources the BMC Manager DateTime reflects when the BMC clock read the time, not
when the hardware inventory was last known-good. InventoryData/Status.LastModifiedTime
(AMI/MSI OEM endpoint) records the actual timestamp of the last successful host-pushed
inventory cycle and is a better proxy for "when was this hardware configuration last confirmed".
Decision:
inferInventoryLastModifiedTimereadsLastModifiedTimefrom the snapshot and setsAnalysisResult.InventoryLastModifiedAt.reanimatorCollectedAt()in the exporter selectsInventoryLastModifiedAtwhen it is set and no older than 30 days; otherwise falls back toCollectedAt.- Fallback rationale: inventory older than 30 days is likely from a long-running server with no recent reboot; using the actual collection date is more useful for the downstream consumer.
- The inventory timestamp is also logged during replay and live collection for diagnostics.
Consequences:
- Reanimator export
collected_atreflects the last confirmed inventory cycle on AMI/MSI BMCs. - On non-AMI BMCs or when
InventoryData/Statusis absent, behavior is unchanged. - If inventory is stale (>30 days), collection date is used as before.
ADL-034 — Redfish inventory invalidated before host power-on
Date: 2026-03-18
Context:
When a host is powered on by the collector (power_on_if_host_off=true), the BMC still holds
inventory from the previous boot. If hardware changed between shutdowns, the new boot will push
fresh inventory — but only if the BMC accepts it (CRC mismatch triggers re-population). Without
explicit invalidation, unchanged CRCs can cause the BMC to skip re-processing even after a
hardware change.
Decision:
- Before any power-on attempt,
invalidateRedfishInventoryPOSTs to{systemPath}/Oem/Ami/Inventory/Crcwith all groups zeroed (CPU,DIMM,PCIE,CERTIFICATES,SECUREBOOT). - Best-effort: a 404/405 response (non-AMI BMC) is logged and silently ignored.
- The invalidation is logged at
INFOlevel and surfaced as a collect progress message.
Consequences:
- On AMI/MSI BMCs: the next boot will push a full fresh inventory regardless of whether CRCs appear unchanged, eliminating ghost components from prior hardware configurations.
- On non-AMI BMCs: the POST fails immediately (endpoint does not exist), nothing changes.
- Invalidation runs only when
power_on_if_host_off=trueand host is confirmed off.
ADL-035 — Redfish hardware event log collection from Systems LogServices
Date: 2026-03-18
Context: Redfish BMCs expose event logs via LogServices/{svc}/Entries. On MSI/AMI this includes the IPMI SEL with hardware events (temperature, power, drive failures, etc.). Live collection previously collected only inventory/sensor snapshots; event history was unavailable in Reanimator.
Decision:
- After tree-walk, fetch hardware log entries separately via
collectRedfishLogEntries()(not part of tree-walk to avoid bloat). - Only
Systems/{sys}/LogServicesis queried — Managers LogServices (BMC audit/journal) are excluded. - Log services with Id/Name containing "audit", "journal", "bmc", "security", "manager", "debug" are skipped.
- Entries older than 7 days (client-side filter) are discarded. Pages are followed until an out-of-window entry is found (assumes newest-first ordering, typical for BMCs).
- Entries with
EntryType: "Oem"orMessageIdcontaining user/auth/login keywords are filtered as non-hardware. - Raw entries stored in
rawPayloads["redfish_log_entries"]as[]map[string]interface{}. - Parsed to
models.EventinparseRedfishLogEntries()during replay — same path for live and offline. - Max 200 entries per log service, 500 total to limit BMC load. Consequences:
- Hardware event history (last 7 days) visible in Reanimator
EventLogssection. - No impact on existing inventory pipeline or offline archive replay (archives without
redfish_log_entrieskey silently skip parsing). - Adds extra HTTP requests during live collection (sequential, after tree-walk completes).
ADL-036 — Redfish profile matching may use platform grammar hints beyond vendor strings
Date: 2026-03-25
Context:
Some BMCs expose unusable Manufacturer / Model values (NULL, placeholders, or generic SoC
names) while still exposing a stable platform-specific Redfish grammar: repeated member names,
firmware inventory IDs, OEM action names, and target-path quirks. Matching only on vendor
strings forced such systems into fallback mode even when the platform shape was consistent.
Decision:
- Extend
redfishprofile.MatchSignalswith doc-derived hint tokens collected from discovery docs and replay snapshots. - Allow profile matchers to score on stable platform grammar such as:
- collection member naming (
outboardPCIeCard*, drive slot grammars) - firmware inventory member IDs
- OEM action/type markers and linked target paths
- collection member naming (
- During live collection, gather only lightweight extra hint collections needed for matching
(
NetworkInterfaces,NetworkAdapters,Drives,UpdateService/FirmwareInventory), not slow deep inventory branches. - Keep such profiles out of fallback aggregation unless they are proven safe as broad additive hints.
Consequences:
- Platform-family profiles can activate even when vendor strings are absent or set to
NULL. - Matching logic becomes more robust for OEM BMC implementations that differ mainly by Redfish grammar rather than by explicit vendor strings.
- Live collection gains a small amount of extra discovery I/O to harvest stable member IDs, but
avoids slow deep probes such as
Assemblyjust for profile selection.
ADL-037 — easy-bee archives are parsed from the embedded bee-audit snapshot
Date: 2026-03-25
Context:
reanimator-easy-bee support bundles already contain a normalized hardware snapshot in
export/bee-audit.json plus supporting logs and techdump files. Rebuilding the same inventory
from raw techdump/ files inside LOGPile would duplicate parser logic and create drift between
the producer utility and archive importer.
Decision:
- Add a dedicated
easy_beevendor parser forbee-support-*.tar.gzbundles. - Detect the bundle by
manifest.txt(bee_version=...) plusexport/bee-audit.json. - Parse the archive from the embedded snapshot first; treat
techdump/and runtime files as secondary context only. - Normalize snapshot-only fields needed by LOGPile, notably:
- flatten
hardware.sensorsgroups into[]SensorReading - turn runtime issues/status into
[]Event - synthesize a board FRU entry when the snapshot does not include FRU data
- flatten
Consequences:
- LOGPile stays aligned with the schema emitted by
reanimator-easy-bee. - Adding support required only a thin archive adapter instead of a full hardware parser.
- If the upstream utility changes the embedded snapshot schema, the
easy_beeadapter is the only place that must be updated.
ADL-038 — HPE AHS parser uses hybrid extraction instead of full zbb schema decoding
Date: 2026-03-30
Context: HPE iLO Active Health System exports (.ahs) are proprietary ABJR containers
with gzip-compressed zbb payloads. The sample inventory data contains two practical signal
families: printable SMBIOS/FRU-style strings and embedded Redfish JSON subtrees, especially for
storage controllers and drives. Full zbb binary schema decoding is not documented and would add
significant complexity before proving user value.
Decision: Support HPE AHS with a hybrid parser:
- decode the outer
ABJRcontainer - gunzip embedded members when applicable
- extract inventory from printable SMBIOS/FRU payloads
- extract storage/controller/backplane details from embedded Redfish JSON objects
- enrich firmware and PSU inventory from auxiliary package payloads such as
bcert.pkg - do not attempt complete semantic decoding of the internal
zbbrecord format Consequences: - Parser reaches inventory-grade usefulness quickly for HPE
.ahsuploads. - Storage inventory is stronger than text-only parsing because it reuses structured Redfish data when present.
- Auxiliary package payloads can supply missing firmware/PSU fields even when the main SMBIOS-like blob is incomplete.
- Future deeper
zbbdecoding can be added incrementally without replacing the current parser contract.
ADL-039 — Canonical inventory keeps DIMMs with unknown capacity when identity is known
Date: 2026-03-30
Context: Some sources, notably HPE iLO AHS SMBIOS-like blobs, expose installed DIMM identity
(slot, serial, part number, manufacturer) but do not include capacity. The parser already extracts
those modules into Hardware.Memory, but canonical device building and export previously dropped
them because size_mb == 0.
Decision: Treat a DIMM as installed inventory when present=true and it has identifying
memory fields such as serial number or part number, even if size_mb is unknown.
Consequences:
- HPE AHS uploads now show real installed memory modules instead of hiding them.
- Empty slots still stay filtered because they lack inventory identity or are marked absent.
- Specification/export can include "size unknown" memory entries without inventing capacity data.
ADL-040 — HPE Redfish normalization prefers chassis Devices/* over generic PCIe topology labels
Date: 2026-03-30
Context: HPE ProLiant Gen11 Redfish snapshots expose parallel inventory trees. Chassis/*/PCIeDevices/*
is good for topology presence, but often reports only generic DeviceType values such as
SingleFunction. Chassis/*/Devices/* carries the concrete slot label, richer device type, and
product-vs-spare part identifiers for the same physical NIC/controller. Replay fallback over empty
storage volume collections can also discover Volumes/Capabilities children, which are not real
logical volumes.
Decision:
- Treat Redfish
SKUas a valid fallback forhardware.board.part_numberwhenPartNumberis empty. - Ignore
Volumes/Capabilitiesdocuments during logical-volume parsing. - Enrich
Chassis/*/PCIeDevices/*entries with matchingChassis/*/Devices/*documents by serial/name/part identity. - Keep
pcie.device_classsemantic; do not replace it with model or part-number strings when Redfish exposes only generic topology labels.
Consequences:
- HPE Redfish imports now keep the server SKU in
hardware.board.part_number. - Empty volume collections no longer produce fake
Capabilitiesvolume records. - HPE PCIe inventory gets better slot labels like
OCP 3.0 Slot 15plus concrete classes such asLOM/NICorSAS/SATA Storage Controller. part_numberremains available separately for model identity, without polluting the class field.
ADL-041 — Redfish replay drops topology-only PCIe noise classes from canonical inventory
Date: 2026-04-01
Context: Some Redfish BMCs, especially MSI/AMI GPU systems, expose a very wide PCIe topology
tree under Chassis/*/PCIeDevices/*. Besides real endpoint devices, the replay sees bridge stages,
CPU-side helper functions, IMC/mesh signal-processing nodes, USB/SPI side controllers, and GPU
display-function duplicates reported as generic Display Device. Keeping all of them in
hardware.pcie_devices pollutes downstream exports such as Reanimator and hides the actual
endpoint inventory signal.
Decision:
- Filter topology-only PCIe records during Redfish replay, not in the UI layer.
- Drop PCIe entries with replay-resolved classes:
BridgeProcessorSignalProcessingControllerSerialBusController
- Drop
DisplayControllerentries when the source Redfish PCIe document is the generic MSI-styleDescription: "Display Device"duplicate. - Drop PCIe network endpoints when their PCIe functions already link to
NetworkDeviceFunctions, because those devices are represented canonically inhardware.network_adapters. - When
Systems/*/NetworkInterfaces/*links back to a chassisNetworkAdapter, match against the fully enriched chassis NIC identity to avoid creating a second ghost NIC row with the rawNetworkAdapter_*slot/name. - Treat generic Redfish object names such as
NetworkAdapter_*andPCIeDevice_*as placeholder models and replace them from PCI IDs when a concrete vendor/device match exists. - Drop MSI-style storage service PCIe endpoints whose resolved device names are only
Volume Management Device NVMe RAID ControllerorPCIe Switch management endpoint; storage inventory already comes from the Redfish storage tree. - Normalize Ethernet-class NICs into the single exported class
NetworkController; do not splitEthernetControllerinto a separate top-level inventory section. - Keep endpoint classes such as
NetworkController,MassStorageController, and dedicated GPU inventory coming fromhardware.gpus.
Consequences:
hardware.pcie_devicesbecomes closer to real endpoint inventory instead of raw PCIe topology.- Reanimator exports stop showing MSI bridge/processor/display duplicate noise.
- Reanimator exports no longer duplicate the same MSI NIC as both
PCIeDevice_*andNetworkAdapter_*. - Replay no longer creates extra NIC rows from
Systems/NetworkInterfaceswhen the same adapter was already normalized fromChassis/NetworkAdapters. - MSI VMD / PCIe switch storage service endpoints no longer pollute PCIe inventory.
- UI/Reanimator group all Ethernet NICs under the same
NETWORKCONTROLLERsection. - Canonical NIC inventory prefers resolved PCI product names over generic Redfish placeholder names.
- The raw Redfish snapshot still remains available in
raw_payloads.redfish_treefor low-level troubleshooting if topology details are ever needed.
ADL-042 — xFusion file-export archives merge AppDump inventory with RTOS/Log snapshots
Date: 2026-04-04
Context: xFusion iBMC tar.gz exports expose the base inventory in AppDump/, but the most
useful NIC and firmware details live elsewhere: NIC firmware/MAC snapshots in
LogDump/netcard/netcard_info.txt and system firmware versions in
RTOSDump/versioninfo/app_revision.txt. Parsing only AppDump/ left xFusion uploads detectable but
incomplete for UI and Reanimator consumers.
Decision:
- Treat xFusion file-export
tar.gzbundles as a first-class archive parser input. - Merge OCP NIC identity from
AppDump/card_manage/card_infowith the latest per-slot snapshot fromLogDump/netcard/netcard_info.txtto producehardware.network_adapters. - Import system-level firmware from
RTOSDump/versioninfo/app_revision.txtintohardware.firmware. - Allow FRU fallback from
RTOSDump/versioninfo/fruinfo.txtwhenAppDump/FruData/fruinfo.txtis absent.
Consequences:
- xFusion uploads now preserve NIC BDF, MAC, firmware, and serial identity in normalized output.
- System firmware such as BIOS and iBMC versions survives xFusion file exports.
- xFusion archives participate more reliably in canonical device/export flows without special UI cases.
ADL-043 — Extended HGX diagnostic plan-B is opt-in from the live collect form
Date: 2026-04-13
Context: Some Supermicro HGX Redfish targets expose slow or hanging component-chassis inventory
collections during critical plan-B, especially under Chassis/HGX_* for Assembly,
Accelerators, Drives, NetworkAdapters, and PCIeDevices. Default collection should not
block operators on deep diagnostic retries that are useful mainly for troubleshooting.
Decision: Keep the normal snapshot/replay path unchanged, but gate those heavy HGX
component-chassis critical plan-B retries behind the existing live-collect debug_payloads flag,
presented in the UI as "Сбор расширенных данных для диагностики".
Consequences:
- Default live collection skips those heavy diagnostic plan-B retries and reaches replay faster.
- Operators can explicitly opt into the slower diagnostic path when they need deeper collection.
- The same user-facing toggle continues to enable extra debug payload capture for troubleshooting.
ADL-044 — LOGPile project release tags use vN.M
Date: 2026-04-13
Context: The repository accumulated release tags in vN.M.P form, while the shared module
versioning contract in bible/rules/patterns/module-versioning/contract.md standardizes version
shape as N.M. Release tooling reads the git tag verbatim into build metadata and release
artifacts, so inconsistent tag shape leaks directly into packaged versions.
Decision: Use vN.M for LOGPile project release tags going forward. Do not create new
vN.M.P tags for repository releases. Build metadata, release directory names, and release notes
continue to inherit the exact git tag string from git describe --tags.
Consequences:
- Future project releases have a two-component version string such as
v1.12. - Release artifacts and
--versionoutput stay aligned with the tag shape without extra mapping. - Existing historical
vN.M.Ptags remain as-is unless explicitly rewritten.
ADL-045 — Generic live IPMI collector is deferred; Redfish remains the only production live path
Date: 2026-04-22
Context: Sprint issue #12 proposed a generic IPMI collector for SEL/FRU/sensors. By this
point LOGPile already has a production Redfish pipeline with replayable raw snapshots, profile-
driven acquisition, and normalized event/sensor/inventory extraction. Redfish also already covers
the current product goals better than IPMI for live collection: richer inventory, structured
resource relationships, and vendor log access via LogServices, including SEL-style logs on many
implementations.
Decision: Do not build a generic live IPMI collector now. Keep ipmi_mock.go only as a
protocol placeholder in the registry and UI/API contract. Treat Redfish as the only production
live collection path. Revisit IPMI only if real field evidence shows that a specific target class
cannot provide required data over Redfish. If revisited, prefer a narrow fallback scope such as
IPMI SEL fallback, IPMI FRU fallback, or IPMI sensor fallback rather than a second full
collector architecture.
Consequences:
- Issue
#12is closed as deferred/not planned, not as implemented. - Live collection architecture stays centered on replayable
raw_payloads.redfish_tree. - The codebase avoids introducing a second generic live-ingest/replay contract for IPMI data.
- Future IPMI work must be justified by concrete Redfish gaps on real hardware, not by protocol symmetry alone.
ADL-046 — The web shell delegates report rendering to internal/chart
Date: 2026-04-22
Context: The frontend had two competing report paths: the embedded internal/chart viewer and
an older client-side renderer in web/static/js/app.js for config, firmware, sensors, serials,
events, and parse errors. That duplication left dead controls in the shell and made the report
source of truth ambiguous.
Decision: The web/ frontend shell is responsible only for data intake, job control, and
top-level actions. The report itself must be rendered exclusively through internal/chart.
Do not keep parallel report sections, filters, or table renderers in shell JavaScript.
Consequences:
- The browser UI has a single report rendering path:
/chart/currentinside the embedded viewer. - Report-level filtering or extra report sections must be implemented in
internal/chart, not inweb/static/js/app.js. - Removing legacy DOM renderers from the shell is a correctness fix, not a behavior regression.
ADL-047 — Inspur onekeylog per-file component/*.txt D-Bus layout is parsed alongside component.log
Date: 2026-07-23
Context: A field dump (dump_29E201150_20260723-1418.tar.gz) came from an Inspur/Kaytus
onekeylog BMC that does not produce the combined component/component.log the inspur parser
expected. Instead it dumps each component separately under component/*.txt as raw D-Bus GetAll
transcripts (GETALL <object> OBJect blocks, tab-separated "field" "type":"x" "data":value
triples, not valid JSON). The archive root is also named dump_<serial>_<timestamp>/ rather than
onekeylog/, so path-based detection did not recognize the layout either. As a result PSU and fan
data were silently empty on this archive class even though the source files carried full data.
Decision: Treat the per-file component/*.txt D-Bus transcript as a second supported onekeylog
layout, not a new vendor. Added internal/parser/vendors/inspur/component_dbus.go with a
line-oriented GETALL block parser (parseDBusGetAllObjects), wired as a fallback in parser.go
when component.log is absent: ParseComponentDirPowerSupply from PowerSupplyInfo.txt and
ParseComponentDirFan from FanInfo.txt. Extended Detect() with onekeylog_dreport.log and
component/powersupplyinfo.txt markers so this layout is recognized without relying solely on
asset.json content markers. GETALL object names can recur across command sections with different
field subsets (e.g. Pwm_N under FanPWM has a Value reading, the same name under FanControl has
only a Target setpoint); fields are unioned across occurrences with first-seen-wins per key so a
later content-free duplicate cannot blank out an earlier real reading.
Consequences:
- PSU and fan RPM/PWM telemetry now parse correctly for this onekeylog layout.
component/NetworkAdapter.txt,component/HDDBpListInfo.txt(busctl--verboseobject-tree dump) andcomponent/RAID.txt(mixed per-controller formats) remain unparsed for this layout; NIC data is not lost since PCIe inventory fromasset.jsonalready carries NIC model/MAC. Revisit only if a real archive needs that specific data and asset.json does not cover it.- Any future component/*.txt GETALL consumer should reuse
parseDBusGetAllObjectsrather than re-implementing block splitting or field extraction.
ADL-048 — Inspur HGX dump_<serial>_<timestamp>/ layout: ipmitool-output FRU/sensors, diagnose-json PCIe, commer-comp events
Date: 2026-07-29
Context: A field dump from an HGX B200 / KR9288-X3 (dump_23DB01633_20260727-1359.tar.gz)
uses yet another onekeylog variant that has no devicefrusdr.log at all (unlike ADL-047's
per-file D-Bus layout, which is missing only component.log). FRU and sensors come from raw
ipmitool fru print / ipmitool sensor list / ipmitool sdr elist output under
component/{fru,sensor,sdr}.txt. Without a fallback the parser returned fru: 0, sensors: 0 on
an otherwise well-detected archive (Detect() still scored via onekeylog_dreport.log), silently
dropping inventory a diagnostic case (intermittent GPU PCIe dropout) depended on.
Decision:
component/fru.txtuses the exact sameFRU Device Description :block format as the FRU section ofdevicefrusdr.log, so it's parsed with the existingParseFRUunchanged.component/sensor.txt(ipmitool sensor list, has value+unit+status per line) is the primary sensor fallback;component/sdr.txt(ipmitool sdr elist, no value on disabled sensors) is used only ifsensor.txtis absent. New parsers incomponent_fallback.go(ParseSensorList,ParseSDRElist).log/bmc/diagnose/OtrdDiagnoseComponent.json's"Pcie Device Info"array is a structural, point-in-time PCIe snapshot (independent of SEL/IDL alarm history) — parsed byParseOtrdDiagnosePCIeindiagnose.gointomodels.PCIeDevice(using the existingPresent *boolfield) and merged via the existingMergePCIeDevices. A device present but running below its negotiated max link speed/width getsStatus = "Link Degraded"and a matching Warning event (BuildPCIeLinkDegradationEvents), independent of GPU-fault SEL/IDL parsing ingpu_status.go.log/sel.csvis textually identical toselelist.csv(ipmitool sel elist -c -Zoutput with the same 6-column CSV body) and is now a fallback source whenselelist.csvis absent, reusingParseSELListWithLocationunchanged.log/bmc/commer-comp/<commerslot|commerhmc|commerswvr|commerswcpld>/component logs (plus rotated*.tar.gz.Nparts, decoded in-memory) share one line format ([timestamp][file, line][level] message);ParseCommerCompEventsincommercomp.goemits an event per line that looks like a real failure, after dropping a fixed noise list (Invalid Comp Data,Error in update redis,MutexTimeout = 0).- Fixed two long-standing
ParseFRUbugs surfaced by this archive:Board Serial/Board Part Numbervalues were overwritten by a later, placeholderProduct Serial : 0/Product Part Number : NULLline within the same FRU block (e.g.SCM_FRU, which has no product-level FRU fields). Placeholder values ("0","NULL") no longer overwrite an already-set real value. - Since duplicate SEL entries can now surface through more than one source file (
sel.csv+ IDL/BMC logs) for the same moment,Parse()runs a finaldedupSELEventspass collapsingSource == "SEL"events with identical(timestamp, event_type, description). This is deliberately narrower than IDL's own timestamp-inclusive dedup inidl.go, which must keep distinct occurrences of a recurring alarm. - When
Parse()still ends up withlen(FRU) == 0 || len(Sensors) == 0after all fallbacks, it appends aCollectionError{Section: "inventory"}so the gap is visible in the API response instead of silently returning an empty inventory. Consequences: stats.fru/stats.sensorsare non-zero on this dump class; GPU presence/link state is available from a source independent of SEL/IDL alarm parsing.- BIOS-change-settings context (
biosChangedSettings/Bios_change_settings*.json) and BIOS POST codes (log/server/progress_log/host0/*) from the same issue are not yet parsed — deferred, no acceptance case depended on them yet. - Any future ipmitool-text-output fallback (FRU/sensor/sdr) should extend
component_fallback.gorather than adding another one-off parser.
ADL-049 — HGX tray/baseboard identity is a distinct entity from the vendor mechanical-carrier FRU
Date: 2026-07-29
Context: Issue #21 (Inspur/Kaytus HGX B200 dumps). hw.BoardInfo/vendor FRU (component/fru.txt
Board Product : CA, a YZCA-* part) identifies the mechanical carrier/tray shipped by Inspur —
it's bolted to the chassis and does not change when the actual NVIDIA HGX baseboard ("delta
board", SXM+NVSwitch) is swapped. Three dumps from the same server showed the CA carrier serial
constant across all three while the NVIDIA-assigned tray (699-26612-*) and baseboard
(935-26287-*) serials — read from log/bmc/oem-commer-log/HGX_HWInfo_FWVersion.log — stayed
identical between dumps A/B (~4 weeks apart) and both changed together in dump C, i.e. only
HGX_HWInfo_FWVersion.log actually reflects a baseboard swap.
Decision:
- Added
models.HGXIdentity(Tray,Baseboard, each aModel/PartNumber/SerialNumbertriple) asHardwareConfig.HGX, populated byparseHGXIdentityininternal/parser/vendors/inspur/hgx_hwinfo.gofrom the sameHGX_HWInfo_FWVersion.logfile already used for GPU assembly/firmware enrichment. - The log is a sequence of
# curl ... <redfish-path>comment lines each followed by that request's JSON response;splitCurlBlockschunks on the comment lines so fields are attributed to the path that produced them (classified bytray/baseboardsubstring, order-independent field regexes — unlike the existing fixed-orderreHGXGPUBlockregex for GPU assembly). Paths containinggpu_sxmor/processors/are explicitly excluded so a per-GPU triple can never be misattributed to the tray/baseboard entity. hgxValue()normalizes Redfish'sNA/N/Aplaceholder (seen when GPUs are unpowered but the baseboard itself still responds) to empty string, applied to both the new identity parser and the existing per-GPU assembly parser, so"NA"never leaks into a serial/model/part field.- Deliberately left
hw.BoardInfo/vendor FRU parsing (fru.go) unchanged — it is not wrong, just a different entity (mechanical carrier). Consumers that need "did the actual GPU board change" must compareHardwareConfig.HGX, notBoardInfo. Consequences: - Dump-to-dump baseboard/tray swap detection (issue #21's P2) is now possible by comparing two
HardwareConfig.HGXvalues; not yet wired into any diff/comparison UI. - GPU-status "baseboard responds, GPU not readable" surfacing (P1) is a natural follow-on now that
GPU fields normalize through the same
NA-aware path, but no dedicated event/diagnostic was added yet — deferred, no acceptance case depended on it.
ADL-050 — Dell iDRAC10 TSR bundles ship inventory as a raw Redfish walk, not DCIM-XML
Date: 2026-08-11
Context: A TSR from a PowerEdge R7715 (TSR20260721231613_1TVFYL4.zip, iDRAC10-generation
firmware) parsed as dell and showed events but no hardware inventory. The dell parser
(internal/parser/vendors/dell/parser.go) built the entire Hardware tree from three files inside
the .pl.zip: sysinfo_dcim_view.xml, sysinfo_dcim_softwareidentity.xml, sysinfo_cim_sensor.xml.
This iDRAC generation does not produce those files at all; tsr/hardware/sysinfo/inventory/ instead
contains redfishidracwalk.tar.gz — a captured crawl of the iDRAC's own Redfish tree (one JSON
document per resource, stored as <url-path>/index.json, e.g.
redfish/v1/Systems/System.Embedded.1/index.json). Hardware therefore stayed effectively empty
(only BoardInfo/one iDRAC firmware entry from metadata.json), while curr_lclog.xml still
populated Events normally — hence "shows only logs".
Decision: Added internal/parser/vendors/redfishtree (shared, non-registering helper package)
that reconstructs a path -> document tree from a tar.gz- or zip-packaged Redfish walk and feeds it
to the existing collector.ReplayRedfishFromRawPayloads replay machinery (previously only used for
live BMC collection and reanimator import). Detection is two-step and vendor-independent: (1) a path
hint — an archive member path containing redfish and ending in .tar.gz/.tgz/.zip; (2)
structural confirmation — the unpacked tree must contain a document whose own @odata.id is exactly
/redfish/v1, plus a /redfish/v1/Systems or /redfish/v1/Chassis collection. The tree is keyed by
each document's own @odata.id (not the on-disk directory name), since some resource names are
URL-encoded on disk (e.g. Assembly%23) but not in the JSON payload.
internal/parser/vendors/dell/redfish_walk.go uses this to enrich the DCIM-XML-derived result:
append-only merge (mergeRedfishReplay) into the same slices the existing dedupeX calls already
resolve, so DCIM-derived entries (appended first) win over duplicates from the replay, and the
replay only fills in what DCIM-XML didn't provide. Also registered a new low-confidence fallback
vendor parser, internal/parser/vendors/redfishwalk (Vendor() = redfish_walk, confidence 35 —
above the generic fallback's 15, below every dedicated vendor parser), using the same
redfishtree helpers directly as its Parse(), so any other vendor that starts shipping this kind
of raw Redfish walk is picked up automatically without a dedicated parser.
Consequences:
- Dell iDRAC10 TSR bundles now populate full hardware inventory (CPUs, memory, storage, PCIe, NICs,
PSUs, firmware) from
redfishidracwalk.tar.gzwhen DCIM-XML is absent, verified end-to-end on the R7715 archive above (1 CPU, 2 DIMMs, 2 storage, 2 PCIe, 1 NIC, 2 PSU, 19 firmware entries). - Any future vendor parser that wants to consume a captured Redfish walk as enrichment should reuse
redfishtree.FindCandidateArchives+redfishtree.Build, not re-implement tar/zip walking. redfishwalkis a genuine fallback: it only winsDetect()when no dedicated vendor parser scores higher on the same archive, per the registry's highest-confidence-wins rule.
ADL-051 — Reanimator export/re-import round trip silently dropped Memory and PSUs
Date: 2026-08-11
Context: After ADL-050 fixed Dell iDRAC10 inventory parsing, the same PowerEdge R7715 archive
still showed no Memory or Power Supplies sections in the /chart/current web view, even though
internal/exporter.ConvertToReanimator produced both correctly from a fresh TSR upload (verified via
direct API calls against the running server). The web view did break identically after re-uploading
a previously downloaded reanimator.json export back into LOGPile (the "Reanimator" round trip: export,
then re-import the same file to inspect/verify it) — reproduced via POST /api/upload with that file.
Root cause: ReanimatorMemory.Present and ReanimatorPSU.Present (internal/exporter/reanimator_models.go)
were already declared as *bool json:"present,omitempty", matching ReanimatorStorage.Present, but
unlike convertStorageFromDevices (which sets Present: &presentValue on every emitted item),
convertMemoryFromDevices and convertPSUsFromDevices never populated that field on the structs they
built — a plain omission, not a deliberate contract choice. So exported JSON always had "present": true
for storage but no present key at all for memory/PSU items. parseUploadedSnapshot (handlers.go)
re-imports a reanimator export via a direct json.Unmarshal straight into models.AnalysisResult
(the internal shape, not a dedicated import mapper); with no present key to unmarshal, Go's zero
value left MemoryDIMM.Present / PSU.Present as false. On the next ConvertToReanimator call (every
/chart/current render re-converts from the current in-memory result), MemoryDIMM.IsInstalledInventory()
requires Present == true, and convertPSUsFromDevices has an explicit if !present { continue } — so
every memory/PSU entry, despite carrying full data, got filtered out as "not installed".
Decision: Populate Present on ReanimatorMemory/ReanimatorPSU in convertMemoryFromDevices/
convertPSUsFromDevices, mirroring what convertStorageFromDevices already does. No changes needed to
the import path or the Present-based filtering — the filtering itself is correct, it was just never
given the field it needs from a re-imported export.
Consequences:
- Re-uploading a LOGPile-exported
reanimator.jsonnow preserves Memory and Power Supplies through the round trip, verified against the live server: upload TSR → export reanimator.json → re-upload it →/chart/currentstill shows both sections. - Regression test:
TestConvertToReanimator_MemoryAndPSURoundTripSurvivesReimportininternal/exporter/reanimator_converter_test.go— converts a minimal result, marshals it, unmarshals back intomodels.AnalysisResult(simulatingparseUploadedSnapshot), and asserts the reconverted output still contains both entries. - General lesson for this converter: any
Reanimator*struct field meant to round-trip throughparseUploadedSnapshotmust actually be populated by itsconvert*FromDevicesfunction — the struct declaring the field is not enough.Storagewas the only category doing this correctly before this fix; worth auditing PCIe/NIC/GPU conversion the same way if a similar round-trip gap is reported for them.
ADL-052 — Added hardware.licenses[] collection/export (contract v2.12), Dell iDRAC10 first
Date: 2026-08-11
Context: Reanimator's hardware ingest contract added an optional hardware.licenses[] section in
v2.12 (2026-08-11) for software/firmware licenses and feature-on-demand activations (BMC advanced
licenses, vGPU, CPU FoD, RAID feature unlocks, etc). LOGPile's local copy of
bible-local/docs/hardware-ingest-contract.md was still v2.11 and was refreshed from
reanimator/core/bible-local/docs/hardware-ingest-contract.md. Scoped the first implementation to the
Dell iDRAC10 Redfish-walk path (vendors/dell + vendors/redfishwalk, see ADL-050/051), since that's
the only resolver currently producing a full captured Redfish tree with a LicenseService collection
in it — other vendor parsers have no comparable source for this data yet.
Decision:
models.License(internal/models/models.go) added, mirroring the contract's field set 1:1 (name,license_key,vendor,type,feature,component_ref,activated_at,expires_at,present, status fields).HardwareConfig.Licenses []Licenseadded.internal/collector/redfish_replay_licenses.go:collectLicenses()reads the standard DMTF/redfish/v1/LicenseService/Licensescollection (License.v1_xschema) via the existingredfishSnapshotReader.getCollectionMembershelper — no Dell-specific parsing needed, this is a generic DMTF resource, so any future vendor whose Redfish walk includes it gets license collection for free throughReplayRedfishFromRawPayloads.AuthorizationScope: "Device"licenses getcomponent_reffromLinks.AuthorizedDevices[0];Service-scoped licenses stay system-level (nocomponent_ref).LicenseOrigin != "Installed"maps toPresent: false.internal/parser/vendors/dell/redfish_walk.go'smergeRedfishReplayappends replayedLicensesintoresult.Hardware.Licenseslike every other category.internal/exporter/reanimator_converter.go:convertLicenses+dedupeLicenses(dedup key:license_key, falling back tocomponent_ref|name) added, wired intoConvertToReanimatordirectly fromhw.Licenses— licenses don't go through the canonical-devices merge/dedup pipeline used for PCIe/GPU/NIC, since they carry no physical identity to merge on.ReanimatorLicense.Presentis set on every emitted record (learned from ADL-051 — declaring the field isn't enough, it must actually be populated for the reanimator round-trip to survive).internal/chart/viewer/render.go: added alicensessection (betweenpower_suppliesandsensors) so licenses show up in the/chart/currentweb view like every other hardware category. Consequences:- Verified end-to-end on the PowerEdge R7715 (1TVFYL4) TSR archive: 3 licenses extracted from
redfishidracwalk.tar.gz("Secure Enterprise Key Manager", "Secured Component Verification", "iDRAC10 17G Enterprise License"), all system-level (AuthorizationScope: "Service"), correctly exported withpresent: trueand visible in the live/chart/currentpage. - No other vendor parser populates
Hardware.Licensesyet. If a future TSR/log source carries license data outside a Redfish walk (e.g. embedded in a vendor-specific XML/JSON file), it needs its own parsing —collectLicenses()only covers the generic RedfishLicenseServicepath.
ADL-053 — Added hardware.storage[].vendor_id/device_id (contract v2.13)
Date: 2026-08-11
Context: Contract v2.13 added optional vendor_id/device_id (PCI Vendor ID / Device ID, decimal)
to hardware.storage[], modeled on the existing pcie_devices[].vendor_id/device_id fields, for
NVMe drives whose numeric PCI IDs are known. Reanimator uses these to merge differently-worded model
strings for the same physical device into one registry entry, same as it already does for
pcie_devices. bible-local/docs/hardware-ingest-contract.md refreshed 2.12 → 2.13 from
reanimator/core/bible-local/docs/hardware-ingest-contract.md.
Decision:
models.Storage.VendorID/DeviceID(int,vendor_id/device_id,omitempty) added (internal/models/models.go), mirroringmodels.PCIeDevice/models.HardwareDevice.exporter.ReanimatorStorage.VendorID/DeviceIDadded; wired through bothconvertStorage(legacy direct path) andconvertStorageFromDevices(canonical-devices path, the one actually used byConvertToReanimator). The canonical-devices path required a second fix:canonicalDevicesForExport'shw.Storage→HardwareDeviceconversion loop (reanimator_converter.go, storageappendDevicecall) wasn't copyingVendorID/DeviceIDonto theHardwareDeviceeither — both hops needed the field, same failure shape as ADL-051'sPresentbug.- Populated at collection time in two places that actually have PCI IDs for storage today:
- Redfish collector (
internal/collector/redfish.golive path andinternal/collector/redfish_replay_storage.gosnapshot-replay path, the one Dell iDRAC10 TSR walks go through):parseDriveWithSupplementalDocsnow readsVendorId/DeviceIdoff the drive doc itself, then off any doc insupplementalDocs— everycollectStorage/collectStorage(replay) call site was updated to append the drive's linkedPCIeFunctionsdoc(s) (getLinkedPCIeFunctions) intosupplementalDocs, same helper already used for GPU/NIC/PCIeDevice. Reuses the existing genericLinks.PCIeFunctionsmechanism — no Drive-specific fetch added. - Inspur (
internal/parser/vendors/inspur/asset.go,ParseAssetJSON):asset.json's ownPcieInfo[]array already carriesVendorId/DeviceId/PcieSlotper entry; built aPcieSlot → (VendorID, DeviceID)map from it and look it up byhdd.PcieSlotwhen building eachStorageentry — no new data source, same join key (PcieSlot) already used to enrich NVMe model name/serial fromdevicefrusdr.log/audit.log.
- Redfish collector (
- Other vendor parsers (Dell WSMAN/DCIM-XML path, h3c, lenovo_xcc, xigmanas) have no numeric PCI ID
source for storage today and were left unchanged. Unraid (
lspci -nnoutput) and xfusion (PCIe Card Infotable) have numeric IDs in a sibling structure but no reliable join key back to the specific storage entry yet — deferred, would need a new BDF↔block-device correlation mechanism. Consequences: go build ./...andgo test ./...clean. Added coverage:TestParseComponentDetails_UseLinkedSupplementalMetrics(redfish_test.go, linked-PCIeFunction vendor/device ID extraction),TestParseAssetJSON_HddEnrichedWithPcieVendorDeviceID(inspur), andTestConvertToReanimator_StorageVendorDeviceIDSurvivesRoundTrip(exporter, full export→marshal→ reimport→re-export round trip).
ADL-054 — Single-file gzip decompression is no longer capped at a fixed byte count
Date: 2026-08-26
Context: A real nvidia-bug-report-*.log.gz (HGX B200, 53.25MB decompressed) was silently
truncated by extractTarGzFromReader's non-tar branch, which hard-capped decompressed content at
maxGzipDecompressedSize (50MB) via io.LimitReader. The tail lost in this specific case was
low-value sar-style telemetry, but the truncation was silent-by-default (only surfaced as a
generic warning event) and the fixed cap has no relationship to what a legitimate bug-report bundle
can contain — larger dumps would lose structured sections (nvidia-smi -q, dmesg tail) outright.
Decision: Replaced the fixed cap with streaming decompression plus a decompression-ratio
guard (internal/parser/archive.go): a countingReader tracks bytes actually consumed from the
compressed source, and readGzipWithBombGuard reads the decompressed stream in 256KB chunks,
aborting only if the ratio exceeds gzipBombRatio (300:1, evaluated only after
gzipBombMinCompressed = 4KB of compressed input to avoid false positives from gzip header
overhead) or an absolute gzipAbsoluteCeiling (1GB, matching the existing maxSingleFileSizeLarge
used for large tar-in-gz) is hit. This only touches the plain-single-gzipped-file branch — tar.gz
archives already use the generous 1GB maxSingleFileSizeLarge limit and were unaffected.
Consequences:
- Legitimate large single-file logs (any size, low compression ratio) now come through whole instead
of being silently cut off; the extraction is still fully buffered into memory afterward (
[]byte), matching every other extractor in this file — noVendorParserinterface change. - Zip/gzip-bomb protection is now ratio-based instead of absolute-size-based, catching bombs of any target size while still allowing large legitimate files through.
- Added
TestExtractArchiveFromReaderGZ_NoLongerCapsAt50MBandTestExtractArchiveFromReaderGZ_AbortsOnDecompressionBomb(archive_test.go).
ADL-055 — nvidia_bug_report parser now extracts Xid/SXid GPU error events
Date: 2026-08-26
Context: The parser only ever extracted hardware inventory (board/CPU/memory/GPU/NIC/PSU) plus
three fixed informational events (memory config, GPU detection, driver version). All GPU/NVLink
fault triage (Xid/SXid errors, "GPU recovery action changed", "fell off the bus") had to be done by
hand with grep against the raw decompressed log — see bible-local/docs/nvidia-bug-report-analysis.md
for the manual checklist this was based on, itself informed by NVIDIA's GPU Debug Guidelines and
Lambda Labs' check-nvidia-bug-report.sh.
Decision: Added internal/parser/vendors/nvidia_bug_report/errors.go (parseXidAndFaultEvents),
wired into Parser.Parse. It scans syslog-style kernel lines for NVRM: Xid (PCI:...) /
NVRM: SXid (PCI:...) and emits one models.Event per occurrence, classified critical vs
warning via a hardcoded set of hardware-leaning Xid codes (hardwareLeaningXidCodes: 48, 63, 64,
74, 79, 92, 94, 95, 119, 120, 122, 123, 134, 150, 154, 155 — ECC/NVLink/bus-fault-associated per
NVIDIA's Xid catalog) or a "GPU Reset Required" recovery-action transition, everything else
warning. A separate fellOffBusRegex catches the same fault when logged outside an Xid line.
Timestamps are syslog-style with no year; parseSyslogTimestamp assumes CollectedAt's year and
rolls back one year if that guess would land implausibly after the report's collection time (>48h),
mirroring the existing year-inference pattern in unraid/xigmanas.
One specific NVLink topology-discovery retry message
(knvlink*RxDetect*: Failed to update Rx Detect Link mask) was observed flooding a real dump
114,440 times in a single burst — logging one event per line would blow up the event list, so it's
aggregated into a single summary event (count + first/last timestamp), severity critical above
100 occurrences.
Consequences:
- Real HGX B200 dump used for validation now surfaces 8 Xid-31 warnings + 6 Xid-150/154 criticals + 1 aggregated NVLink-flood critical, matching the manual grep-based analysis from this session exactly, with no event-count explosion.
- Severity classification is a heuristic triage hint, not a verdict — same caveat as documented in
bible-local/docs/nvidia-bug-report-analysis.md; Xid meaning is context-dependent (app bug vs hardware) even for codes not in the hardware-leaning set. - Bumped
parserVersionto1.3. Addederrors_test.go: code-based severity classification, SXid event-type distinction, flood aggregation (500 synthetic lines → 1 event), fell-off-bus detection, and year-rollover timestamp handling.
ADL-056 — Inspur event ingestion uses source PRI and an explicit-offset timezone timeline
Date: 2026-08-27
Context: Manual validation of an NF5688M7 one-key log (S/N 23E102624) found that inventory and
the primary NVMe fault were correct, but crit.log* was not ingested (dropping 346 MCTP timeouts and
one ME self-test failure), severity was inferred only from rotated filename, and naive SEL timestamps
were parsed with stale timezone.conf=Asia/Shanghai. The retained history actually changes from
+08:00 to +03:00; the July SEL copy of an IDL event was consequently five hours wrong.
Decision: Inspur parser v2.4 now:
- ingests emerg/alert/crit/error/warning/notice/info syslog families and maps severity from RFC 5424 PRI, with narrow content exceptions for four known-benign AMI/KVM status strings;
- omits fabricated 1970 syslog records whose only usable timestamp is seconds-since-boot;
- resolves offset-less SEL rows against a bounded timeline sampled from explicit-offset IDL records,
using
timezone.confonly as fallback and preferring the embedded IDL timestamp over its outer syslog envelope; - treats deassert records as informational, deduplicates equivalent IDL/SEL copies at the same
normalized instant (preferring IDL), and stable-sorts the combined event stream by timestamp.
Consequences: The real NF5688M7 dump now exports all 346 MCTP timeouts and the ME self-test code
195, the July
Sys_Healthevent appears once at2026-07-23T23:20:38Z, known benign driver/KVM strings no longer inflate Critical/Warning, no 1970 event remains, andevent_logsis chronological. Synthetic regression coverage was added toevent_logs_test.goandsel_test.go.
ADL-057 — Inspur CPU PPIN is exported as source-backed serial; active NVMe faults update storage status
Date: 2026-08-27
Context: NF5688M7 S/N 23E102624 exposes stable per-socket CPU identifiers only as PPIN in
both asset.json and RESTful CPU info. The same values and socket mapping were confirmed in the
older 2026-03-26 dump. Lenovo paper LP1890 describes PPIN as the processor-assigned serial/physical
CPU identifier; Intel documents it as a 64-bit physical-processor inventory identifier. LOGPile
parsed PPIN but left CPU.SerialNumber empty, so Reanimator dropped it. The same case also had an
active NvmeIndex:6 drive fault while every exported storage record remained Unknown.
Decision: Inspur parser v2.5:
- copies non-empty BMC PPIN into both
CPU.PPINandCPU.SerialNumber; this is source-backed identity, not a generated fallback, and preserves CPU0/CPU1 socket mapping; - projects IDL NVMe drive-fault Assert/Deassert transitions onto
hardware.storageby the observed zero-basedNvmeIndex:N→ one-basedOB(N+1)mapping; - uses
dev_status.logas the authoritative active-fault snapshot when present, records transition time/history/details, and exports an active fault asCritical; - leaves unaffected drives
Unknownunless the source provides an affirmative health result; mere presence is not treated as proof ofOK. Consequences: On the real dump CPU0 exports serialD46E5D6B1D3E40E1, CPU1 exportsD44F8D6B9155EE0E, and onlyOB07exportsCriticalwith theStatus_Flags errordescription; OB01–OB06 remainUnknown. Regression tests cover asset/component PPIN mapping, targeted NVMe fault, deassert clearing, and rejection of unrelated NVMe inventory-change events.
ADL-058 — PPIN is the common source-backed CPU serial fallback
Date: 2026-08-27
Context: The NF5688M7 fix initially copied PPIN into CPU.SerialNumber only inside the Inspur
parser. Dell TSR, H3C INI/XML, and Redfish also expose PPIN (including OEM public processor serial),
but their parsed CPU serial could remain empty. Canonical devices and Reanimator could additionally
drop a PPIN already present in the model or device details.
Decision: CPU identity resolution is centralized in models.ResolveCPUSerialNumber. A valid,
explicit source serial has priority; otherwise a valid source PPIN is used. Empty values and known
placeholders (N/A, Unknown, and equivalents) are rejected. Dell, H3C, Inspur, and Redfish apply
the resolver while parsing, and API/canonical-device and Reanimator conversion apply it again at the
output boundary. No identifier is generated from socket, model, board serial, or another component.
Consequences: All current parsers that expose PPIN now return it consistently as CPU
serial_number; CSV, API, and Reanimator retain the same identity. Parsers with a genuine separate
CPU serial preserve that value. Regression tests cover vendor inputs, common output paths, priority,
fallback, and placeholder rejection.
ADL-059 — CPU identity regression audits use an offline batch CLI
Date: 2026-08-27
Context: The PPIN-to-CPU-serial fix must be checked once against a large directory of historical
dumps. The existing HTTP batch converter emits complete Reanimator payloads and requires multipart
uploads, while this audit needs small per-dump CPU-only JSON records and must continue past malformed
or unsupported archives.
Decision: Add cmd/logpile-cpu-audit as a non-server entry point over the existing parser
registry. It recursively discovers supported dump files, parses them with bounded concurrency, and
writes one path-mirrored CPU identity report per input plus a deterministic summary. Each CPU report
contains the parsed PPIN and serial, resolver expectation, promotion flag, and regression finding.
Failures are isolated and represented as JSON. Plain .txt/.log inputs are opt-in for directory
scans to avoid treating files inside unpacked dumps as independent dumps.
Consequences: Historical audit runs need no HTTP server or full hardware export, retain source
path and parser/version provenance, and can be reviewed or diffed as ordinary JSON. The tool shares
the production parser and CPU resolver, so its result reflects the same behavior as API and export
paths without duplicating vendor parsing logic.
Amendment (2026-08-27): The helper also supports -reanimator mode for a flat, directly
ingestible CPU-only dataset. Audit reports are not Reanimator payloads and must not be sent to
POST /ingest/hardware. Reanimator mode converts through the production exporter, retains mandatory
collected_at and hardware.board, removes non-CPU hardware sections, skips empty inputs, and writes
no non-ingest JSON summaries alongside payload files.
ADL-060 — HPE AHS DIMM inventory dedupes per slot, newest snapshot wins
Date: 2026-08-27
Context: The AHS blackbox stores a full DIMM inventory snapshot for every POST
cycle, and parseDIMMs flattens every snapshot from every blackbox record into one
token stream. dedupeMemory keyed on SerialNumber (falling back to slot|part),
so when a module was physically swapped between captures the same physical slot
survived once per historical occupant. On a real bench dump where PROC 1/2 DIMM 7
was swapped mid-session this reported 26 modules for a 24-slot board and made the
"same P/N" and "mixed P/N" configurations look identical.
Decision: dedupeMemory collapses to one entry per non-empty Slot, and the
last occurrence for a slot wins. Token order follows blackbox-record order, which is
chronological, so the last occurrence is the most recent POST snapshot. Entries with
no slot label keep the previous serial/part fallback dedupe.
Consequences: AHS memory inventory always matches the physical slot count and
reflects the latest observed population. Historical swaps are no longer surfaced as
extra modules (if a per-slot history is wanted later it must be built explicitly,
not inferred from dedupe residue). Regression test TestDedupeMemorySlotSwap.
ADL-061 — xFusion mem_info rows survive an embedded newline in the SPD BOM field
Date: 2026-08-31
Context: The iBMC AppDump/CpuMem/mem_info file copies the raw SPD "bom number"
(col 14) verbatim, and it contains arbitrary binary bytes. On a G5500 V7 dump
(S/N 210619KUGGXGS2000017) one DIMM's BOM field held a 0x0A, splitting that
record across two physical lines. parseMemInfo split the file on "\n" and
treated each line as a record, so the head line lost every field after the BOM
column (type, speed, part number) and the tail fragment
s, M321R8GA0EB2-CWMXH, ..., OK passed the len(parts) >= 9 guard and became a
phantom 33rd "DIMM" with slot s, manufacturer N/A, serial OK. The BEE-SP
export of the same host (dmidecode-based) correctly showed 32 healthy DIMMs.
Decision: parseMemInfo now feeds reassembleMemInfoRows, which folds any
line not starting with Memory<digit> (or the slot(col 1) header) back into the
previous logical row before comma-splitting. looksLikeMemInfoRowStart gates this.
Consequences: The phantom DIMM is gone and the split record is fully recovered.
Both collectors now yield the same 32-module inventory for this host — identical
serial/size/type/speed/part/status per module; only slot labels (Memory140 vs
DIMM140(E)), the location field, and status_checked_at still differ, which is
inherent to the two sources. Regression test TestParseMemInfo_EmbeddedNewlineInBOM.
ADL-062 — BMC-dump and live-CD exports of the same server must agree component-for-component
Date: 2026-08-31
Context: The audit tool that ingests Reanimator exports treats any per-component
field change between imports as a component replacement. Importing a BMC dump
(xfusion) and a BEE-SP live-CD bundle (easy_bee) for the same server produced
large spurious diffs: a device present in one bundle and absent in the other, and
the same physical part carrying different slot labels / model strings / vendor
names depending on the collector.
Decision: Normalize toward a single cross-collector representation, splitting
the work between the two vendor parsers and the shared exporter:
- Exporter (
reanimator_converter.go) — cross-vendor rules that any two collectors must converge on:canonicalMemorySlot(drop dmidecode channel tag(J), rewrite BMCMemory111→DIMM111);canonicalGPUModel(NVIDIA data-center GPUs reduce to the bare chip token, "NVIDIA H200 NVL"→"H200");canonicalStorageMediaAndInterface("NVMe" is a bus not a medium → type SSD / interface NVMe; bare PCIe drive → NVMe);manufacturerFromStorageModel;isRemovableUSBStorageDevice(a live-CD boot stick is not inventory);isOnboardControllerPCIeDevice(SATA/NVMe/MegaRAID/PCIe-switch controller functions that only an lspci scan ever reports — an add-in HBA/RAID card with its own serial or part number is kept). - xfusion — DIMM slot from the
mem_info"dimm name" column,locationdropped; GPUslot= BDF (Reanimator contract); NIC emitted per PCI function fromnetcard_info.txt(BDF + per-port MAC, shared card serial), replacing the single card-level adapter, so it matches an lspci view and no longer collides with a GPU on a wrong BDF; NIC manufacturer left blank when it is the system OEM so the exporter resolves the silicon vendor from pci.ids;(U6216)chip designator stripped from firmware versions. - easy_bee — PSU bay numbers rebased 0→1 (
normalizePSUSlots), bare single-letter PSU "firmware" (a leaked FRU version) cleared,board.part_numbertaken from the bundle'sipmitool-fru.txtchassis "Product Part Number" to match the BMC value. Consequences: For the reference server every physical component now appears in both exports keyed identically (serial, or BDF for PCIe). Tests:TestCanonicalMemorySlot,TestCanonicalGPUModel,TestCanonicalStorageMediaAndInterface,TestIsOnboardControllerPCIeDevice,TestNormalizePSUSlots, updatedTestParse_ServerFileExport_NetworkAdaptersAndFirmware.
Verified against reanimator/core (the ingesting audit tool), 2026-08-31.
Reanimator persists only a narrow slice of what the export carries, so most
residual field diffs cannot generate a change event:
- Component identity =
vendor_serialonly (partstable). Fallback chain perflattenPCIe: real serial → MAC →{board}-PCIE-{slot}. Memory/storage/PSU: serial → synthetic-by-slot.partsstores only vendor, model, serial — no size, speed, link width, block size, clocks, wattage, location, numa/iommu, so a null→value transition on any of those is a no-op. vendor/modelare re-canonicalized on ingest via the alias resolver usingvendor_id/device_id("prefer the authoritative pci.ids name"). Both bundles carry the same PCI IDs, so GPU model and NIC vendor already converge server-side regardless of the raw strings — the logpile-sidecanonicalGPUModel/ OEM-vendor blanking are belt-and-suspenders for non-reanimator consumers.slot_nameis written once at install and only re-patched when the install itself changes (res.Changed); a component that stays continuously present never gets a slot-change event, so differing slot labels are harmless once presence is stable.- Asset firmware =
machine_firmware_stateskeyed by(machine_id, device_name); an event fires only on a version change for the same name. A name absent from an import is left untouched (no downgrade), so the 6 BMC-only firmware entries are inert.iBMC(BMC dump) vsBMC(live CD) stay two independent, non-churning rows — deliberately not normalized, because unifying the name would then churn on the version (3.08.05.85vs3.08). - Reanimator has its own
IsDeviceIgnored(vendor, model)alias-driven ignore list for "onboard controllers, service USB flash drives" — the logpile-sideisRemovableUSBStorageDevice/isOnboardControllerPCIeDevicefilters overlap it but remove the dependency on the alias dict being configured.
The one real cross-collector reinstall risk was the OCP NIC: with the card serial on it, reanimator collapses both ports to one serial-keyed component, while a live-CD bundle (no serial) stays two MAC-keyed components → remove + 2×install on every source switch. Fixed by emitting the xFusion NIC as per-port entries with no serial, so both sources key on the port MAC and produce the same two components. Net: after this work the only fields that still differ between the two bundles are ones reanimator does not track, so alternating collection methods for one server produces no spurious install/remove/firmware events.
ADL-063 — xFusion disk_info "Capacity" is binary-unit; normalize to decimal GB
Date: 2026-09-01
Context: Diffing a BMC dump (xfusion) against the BEE-SP live-CD bundle
(easy_bee) for the same G5500 V7 (S/N 210619KUGGXGS2000017) showed every storage
device in the BMC export with size_gb absent while the live-CD export had real
sizes (KIOXIA 7681, INTEL 3840). Root cause: parseDiskInfo
(internal/parser/vendors/xfusion/hardware.go) read capacity with
fmt.Sscanf(fields["Capacity"], "%f GB", &capFloat). The iBMC file writes
Capacity : 6.986 TB / 3.492 TB for NVMe drives, so the literal GB never
matched and sizeGB stayed 0. Even for the 446.625 GB boot-SSD case the old
code truncated the binary GiB value (446) instead of the vendor's decimal spec.
Decision: Added parseDiskCapacityGB: parse <number> <unit> (TB/GB/MB,
case-insensitive), treat the number as binary (iBMC reports GiB/TiB mislabeled as
GB/TB), convert to decimal GB (round(value * 2^n / 1e9)). This matches both the
drive's marketed decimal capacity and the BEE-SP live-CD size_gb. A few GB of
rounding slack vs the live-CD's exact byte count is accepted (size_gb is
display-only and not persisted by Reanimator).
Consequences:
- BMC-dump storage now carries
size_gbmatching the live-CD export component-for-component (7681/7681, 3839/3840). Drive Temperature, byte-swapped systemGUIDinOptPme/pram/per_power_off.ini, and PCIe link speed/width remain unparsed for xFusion: temperature has nomodels.Storagefield yet, the GUID is fragile wire-order, and card_info carries no link fields (that data only exists in the live-CD's lspci view, not the dump).- Regression test
TestParseDiskInfo_CapacityUnits.
ADL-064 — Inspur combined component.log: parse RESTful FRU info: and float fans_power
Date: 2026-09-01
Context: Diffing a BMC dump (inspur) against the BEE-SP live-CD bundle
(easy_bee) for the same NF5280M6 (S/N 24C319579) surfaced two blind spots in the
combined-component.log onekeylog layout (classic onekeylog/ root, has
component/component.log, but no devicefrusdr.log and no asset.json):
- Board
manufacturer/product_name/part_number/uuidall empty andstats.fru: 0(with a "still missing: FRU" collection error).component.logcarries aRESTful FRU info:JSON array (BMC_FRU + PSU/backplane/riser FRUs withdevice.system_uuid,board,productareas) that no parser read. - Zero fan sensors (live-CD had 8).
FanRESTInfo.FansPowerwas typedintbut this firmware writes"fans_power": 12.000000;json.Unmarshalof the whole fan block failed, soparseFanSensors/parseFanEventsreturned nil. Decision:
internal/parser/vendors/inspur/component_fru.go(ParseComponentLogFRU) parses theRESTful FRU info:array into[]models.FRUInfo, preferring the product area (operator-facing system serial / asset tag) over the board area (PCB serial) for identity, and setshw.BoardInfo.UUIDfromsystem_uuiddirectly. Placeholder area values (0,NULL,N/A, blank) are skipped. Wired inparser.goas a fallback only whenresult.FRUis still empty after the devicefrusdr.log path, so dumps that already have real FRU data are untouched.FanRESTInfo.FansPowerchanged tofloat64. Consequences:- This dump class now yields full board identity + UUID matching the live-CD
export, 8 fan sensors,
stats.fru: 7, and no collection error. - Real motherboard
part_number(YZMB-01642-102) is exported where the live-CD had only the"0"product-area placeholder - a deliberate improvement, not a regression (Reanimator treats"0"/"NULL"board part as absent). - Remaining live-CD-vs-BMC diffs for this host are not parser gaps: the combined
component.logcarries no full IPMI SDR (so ~16 sensors vs 72), the live-CD0000:00:1f.2PCH power-management function is lspci-only noise, and the live-CD actually missed one DIMM (CPU1_C1D0) that the BMC dump reports. - Tests:
TestParseComponentLogFRU_BoardIdentityAndUUID,TestParseComponentLogFRU_AbsentSection,TestParseComponentLogSensors_FloatFansPower.
Amendment (2026-09-01): a second NF5280M6 pair showed iBMC RESTful version info firmware versions carry a build-timestamp suffix
(08.05.01 (02/21/2024 16:51:27)) that the live-CD does not, churning a
FIRMWARE_CHANGED event per import. cleanFirmwareVersion now strips a trailing
(...) from those entries (BIOS then matches the live-CD exactly; BMC
7.11.02 vs the live-CD's IPMI-truncated 7.11 is inherent and left as-is).
Test TestExtractComponentFirmware_StripsBuildStamp.