feat(display): systemd watchdog and heartbeat for a frozen render loop (#687)

If the render loop gets stuck inside a plugin's display(), ledmatrix.service
stays active and the panel stays frozen. This adds a way to detect that.

- src/display_watchdog.py (standard library only) sends sd_notify over
  $NOTIFY_SOCKET and writes /run/ledmatrix/display-heartbeat.json. Only the
  render thread counts: beats from other threads are ignored.
- ledmatrix.service: WatchdogSec=120, NotifyAccess=main,
  RuntimeDirectory=ledmatrix (0755), RestartSteps=4 and
  RestartMaxDelaySec=2min. It stays Type=simple. run.py widens the watchdog
  to 15 min for start-up, and load_plugin() does the same on the render
  thread. The loop arms after its first frame.
- /api/v3/health adds checks.display_loop: running, stalled (no heartbeat
  for over 60s, which makes the status degraded) or not_reported. With web
  login on, a caller who is not logged in still gets only healthy/degraded,
  and a stall degrades that answer.
- The update verifier requires a fresh heartbeat from the restarted display
  when the display it replaced was writing one. A frozen panel is rolled
  back.
- Existing installs get the systemd watchdog only after install_service.sh
  is re-run. The heartbeat works right away.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Chuck
2026-09-30 11:15:31 -04:00
committed by GitHub
co-authored by Claude Opus 5.5
parent b09434a418
commit 64c7289593
20 changed files with 1635 additions and 12 deletions
+44 -2
View File
@@ -52,6 +52,7 @@ each other. They share three things:
| Preview frame | `/tmp/led_matrix_preview.png` | display: `DisplayManager`, gated by [`snapshot_policy`](../src/common/snapshot_policy.py) | web: display SSE stream, `/api/v3/health` (file age) |
| Preview viewer marker | `/tmp/led_matrix_preview_viewer` | web, while a preview is open | display: writes full-rate snapshots only while it is fresh |
| Hardware init status | `/tmp/led_matrix_hw_status.json` | display | web: `/api/v3/hardware/status` |
| Render-loop heartbeat | `/run/ledmatrix/display-heartbeat.json` (tmpfs) | display: the render thread, via [`display_watchdog`](../src/display_watchdog.py) | web: `/api/v3/health` (`checks.display_loop`); the update health check |
The on-demand start route starts `ledmatrix.service` when it is not running
(`start_service`, on by default) but never restarts a running one: the display
@@ -228,6 +229,43 @@ then normal rotation.
`sync.role`: a leader sends a follower its share of each frame over UDP
(port 5765).
### Liveness
A render thread stuck inside a plugin leaves the service "active" and the
panel frozen, so liveness is reported by the render thread itself
([`src/display_watchdog.py`](../src/display_watchdog.py), standard library
only). `beat()` from any other thread is ignored: the update worker, Vegas's
tick thread and the prefetcher keep running while the render thread is stuck,
and must not vouch for it.
- **Check-in points.** The top of `run()`'s loop (`loop_pass()`), every
dwell second (`_sleep_with_plugin_updates`), every frame of the per-screen
loops (`_display_once`), every frame of Vegas's own loop and static pause
(`coordinator.run_iteration`), each plugin fetched for a Vegas cycle
(`StreamManager._fetch_plugin_content`), each update on the
`synchronous_updates` path, and every frame pushed
(`DisplayManager.update_display` -> `note_frame()`). Beats are
rate-limited to one ping and one heartbeat write every 5 s.
- **systemd watchdog.** `ledmatrix.service` is `Type=simple` with
`WatchdogSec=120` and `NotifyAccess=main`. `run.py` sends
`WATCHDOG_USEC` = 15 minutes before importing anything heavy (start-up loads
plugins and runs the 20 s update budget, and the watchdog clock starts with
the process). After the first frame -- or the first full pass, when there is
nothing to draw -- the loop sends `READY=1`, restores the unit's 120 s and
pings. `PluginManager.load_plugin()` on the render thread (a plugin enabled
from the web UI, or loaded for on-demand) gets 15 minutes again, since it
can run pip. A missed deadline is a SIGABRT; faulthandler, enabled on
arming, dumps every thread's stack to the journal.
- **Heartbeat.** `/run/ledmatrix/display-heartbeat.json`
(`{"pid", "mono", "wall"}`; `RuntimeDirectory=ledmatrix`, 0755, file 0644 so
the web user can read it). Readers compare `mono` with their own
`time.monotonic()` -- CLOCK_MONOTONIC is shared by every process and does not
jump when NTP first sets an RTC-less Pi's clock. `/api/v3/health` calls it
`stalled` past 60 s; no file is `not_reported` and changes nothing. A clean
stop removes it. Without `RuntimeDirectory=` (an older unit) the display,
as root, creates the directory itself; off Linux, or without root, there
is no heartbeat.
## Plugin system
[`src/plugin_system/`](../src/plugin_system/):
@@ -310,7 +348,9 @@ everything else through `_reinstall_with_rollback()`.
`ledmatrix-update-verify.path`, which runs the verifier as a separate unit
(so restarting the web service does not kill it). The verifier restarts
both services, waits for the web API to answer and the display service to
stay up, and on failure resets to the previous commit and restarts again.
stay up -- and, when the display wrote a heartbeat before the update, to
keep one fresh from the restarted process (see Liveness) -- and on failure
resets to the previous commit and restarts again.
Plugin updates run only after a verified core update. State is in
`data/auto_update_state.json` and `data/auto_update_pending.json`.
- **Startup validator.** `StartupValidator`
@@ -318,7 +358,9 @@ everything else through `_reinstall_with_rollback()`.
`DisplayController.__init__`: config and cache directory first, then
enabled plugins once the plugin manager exists. It also warns when an
installed systemd unit differs from its template in `systemd/`. Results
are logged; startup continues either way.
are logged; startup continues either way. Nothing rewrites installed units
on update: a unit change such as the watchdog reaches an existing install
only when `install_service.sh` is re-run.
## Where to start reading
+9 -1
View File
@@ -2101,9 +2101,17 @@ Health of the web interface, display service, config file, plugin system and
display snapshot. `data.status` is `healthy` or `degraded`, with
`data.services` and `data.checks`.
`data.checks.display_loop` is the display's render-loop heartbeat: `running`
(with `heartbeat_age_seconds`), `stalled` (no heartbeat for 60s: the panel is
frozen even if the service is active; the status turns `degraded`), or
`not_reported` when the display writes none (not started yet, the dev server,
Windows), which does not affect the status.
Open even when the web login is on, for uptime monitors; a caller that is not
logged in (and has no token) then gets only `{"status": "success", "data":
{"status": "healthy" | "degraded"}}`.
{"status": "healthy" | "degraded"}}`. A stalled render loop still shows there
as `degraded`; the `checks` detail is only for logged-in callers, tokens and
requests from the Pi itself.
### Hardware Status
+58
View File
@@ -516,6 +516,64 @@ sudo systemctl cat ledmatrix-web | grep User
python3 scripts/check_plugin.py --plugin plugin-id
```
#### Panel Frozen, or the Display Restarts Every Few Minutes
**Symptoms:**
- The panel stops changing while `systemctl status ledmatrix` says `active`
- The display restarts on its own, a couple of minutes after it froze
- `/api/v3/health` shows `checks.display_loop.status` as `stalled`
The display's render loop checks in with systemd every few seconds
(`WatchdogSec=120` in `ledmatrix.service`) and writes a heartbeat to
`/run/ledmatrix/display-heartbeat.json`. When the loop gets stuck -- almost
always inside one plugin's `display()` -- the check-ins stop, and after two
minutes systemd kills and restarts the display. The kill dumps every thread's
stack into the log, so it says which plugin was stuck.
**Solutions:**
1. **Find the stuck plugin.** Look for the watchdog kill and the stack dump
after it. The render loop is the thread whose stack runs through
`display_controller.py` in `run` (usually the `Current thread` block);
the first `plugin-repos/...` file in it is the plugin:
```bash
sudo journalctl -u ledmatrix --since "1 hour ago" | grep -A40 "Watchdog timeout"
```
2. **Check the heartbeat by hand.** Its age should stay under about ten
seconds while the display runs:
```bash
cat /run/ledmatrix/display-heartbeat.json
curl -s http://localhost:5000/api/v3/health | python3 -m json.tool | grep -A3 display_loop
```
`not_reported` means the display writes no heartbeat: it has not drawn
its first frame yet, or it runs an older version.
3. **Disable the plugin** in the web UI and report it to its author with the
stack dump. Restarts that repeat back off from 10 seconds to two minutes
apart, so a plugin that hangs on every start does not restart the display
hundreds of times an hour.
4. **Is the watchdog installed?** Installs from before it keep their old unit
until the installer is re-run (a startup warning says the unit differs
from its template):
```bash
systemctl show -p WatchdogUSec ledmatrix # 2min once running; 0 = not installed
sudo ./scripts/install/install_service.sh
```
`WatchdogUSec` reads `15min` for the first minutes after a start: that is
the start-up allowance, narrowed to two minutes once the first frame is on
the panel.
5. **A plugin that legitimately blocks longer** than two minutes (it should
not; `display()` runs on the render thread) can be given more time with a
drop-in, `sudo systemctl edit ledmatrix`:
```ini
[Service]
WatchdogSec=300
```
`WatchdogSec=0` turns the watchdog off.
#### Stale Cache Data
**Symptoms:**