Files
bee/bible-local/decisions/2026-08-31-bee-nvidia-restart-deadlock.md
T

2.8 KiB

bee-nvidia.service: never call blocking systemctl restart on units ordered After= itself

Date: 2026-08-31 Status: active

Symptom

Every affected NVIDIA boot reaches two 60-second wrapper timeouts in bee-nvidia.service. bee-nvidia.log:

[bee-nvidia] restarting nvidia-fabricmanager.service (timeout 60s)
[bee-nvidia] WARN: systemctl restart nvidia-fabricmanager.service timed out after 60s
[bee-nvidia] restarting nvidia-dcgm.service (timeout 60s)
[bee-nvidia] WARN: systemctl restart nvidia-dcgm.service timed out after 60s
[bee-nvidia] done

Seen on both 210619KUGGXGS2000008 and ...017 (8x H200 NVL). That 120 s window overlaps exactly with when an operator opens the web UI and clicks "Run All" - during it nvidia-smi may not yet enumerate all GPUs, which is how bundle ...017 ended up running the check set with no GPU tests (see 2026-08-31-backend-driven-sat-planning.md).

Root cause

bee-nvidia.service is Type=oneshot and Before=nvidia-fabricmanager.service nvidia-dcgm.service. Its ExecStart (bee-nvidia-load) then ran, synchronously:

timeout 60 systemctl restart nvidia-fabricmanager.service
timeout 60 systemctl restart nvidia-dcgm.service

A oneshot unit is not "active" until ExecStart returns. Both target units are ordered After=bee-nvidia.service, so systemd queues them behind bee-nvidia and will not run them while bee-nvidia-load is still executing. bee-nvidia-load blocks on systemctl restart waiting for exactly that job to complete -> deadlock -> broken only when timeout 60 fires. Twice.

nvidia-smi -q inside bee-check-nvswitch (the fabricmanager ExecCondition) is not part of this ordering cycle; the deadlock is structural.

Decision

bee-nvidia-load no longer blocks on those units. It calls systemctl --no-block try-restart for each:

  • --no-block returns immediately; systemd runs the job after bee-nvidia exits, via the existing Before= ordering.
  • try-restart only acts if the unit is already running (the "stale instance from a reload / re-run" case the old code worried about). If it is inactive, this command does nothing; an enabled unit already queued by the normal boot transaction starts after bee-nvidia.service via the declared ordering.

The SYSTEMCTL_TIMEOUT / timeout_systemctl wrapper and the fallback systemctl start / systemctl status branches are gone. --no-block means systemctl does not wait for the queued unit job to finish.

Consequences

  • NVIDIA boot no longer waits for those two 60-second wrapper deadlines.
  • DO NOT reintroduce a synchronous systemctl {start,restart} of any unit that is After=bee-nvidia.service from inside bee-nvidia-load. If a unit genuinely must be up before the script returns, invert the ordering instead.