mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 06:15:09 +00:00
feat(display): systemd watchdog and heartbeat for a frozen render loop (#687)
If the render loop gets stuck inside a plugin's display(), ledmatrix.service stays active and the panel stays frozen. This adds a way to detect that. - src/display_watchdog.py (standard library only) sends sd_notify over $NOTIFY_SOCKET and writes /run/ledmatrix/display-heartbeat.json. Only the render thread counts: beats from other threads are ignored. - ledmatrix.service: WatchdogSec=120, NotifyAccess=main, RuntimeDirectory=ledmatrix (0755), RestartSteps=4 and RestartMaxDelaySec=2min. It stays Type=simple. run.py widens the watchdog to 15 min for start-up, and load_plugin() does the same on the render thread. The loop arms after its first frame. - /api/v3/health adds checks.display_loop: running, stalled (no heartbeat for over 60s, which makes the status degraded) or not_reported. With web login on, a caller who is not logged in still gets only healthy/degraded, and a stall degrades that answer. - The update verifier requires a fresh heartbeat from the restarted display when the display it replaced was writing one. A frozen panel is rolled back. - Existing installs get the systemd watchdog only after install_service.sh is re-run. The heartbeat works right away. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
+44
-2
@@ -52,6 +52,7 @@ each other. They share three things:
|
||||
| Preview frame | `/tmp/led_matrix_preview.png` | display: `DisplayManager`, gated by [`snapshot_policy`](../src/common/snapshot_policy.py) | web: display SSE stream, `/api/v3/health` (file age) |
|
||||
| Preview viewer marker | `/tmp/led_matrix_preview_viewer` | web, while a preview is open | display: writes full-rate snapshots only while it is fresh |
|
||||
| Hardware init status | `/tmp/led_matrix_hw_status.json` | display | web: `/api/v3/hardware/status` |
|
||||
| Render-loop heartbeat | `/run/ledmatrix/display-heartbeat.json` (tmpfs) | display: the render thread, via [`display_watchdog`](../src/display_watchdog.py) | web: `/api/v3/health` (`checks.display_loop`); the update health check |
|
||||
|
||||
The on-demand start route starts `ledmatrix.service` when it is not running
|
||||
(`start_service`, on by default) but never restarts a running one: the display
|
||||
@@ -228,6 +229,43 @@ then normal rotation.
|
||||
`sync.role`: a leader sends a follower its share of each frame over UDP
|
||||
(port 5765).
|
||||
|
||||
### Liveness
|
||||
|
||||
A render thread stuck inside a plugin leaves the service "active" and the
|
||||
panel frozen, so liveness is reported by the render thread itself
|
||||
([`src/display_watchdog.py`](../src/display_watchdog.py), standard library
|
||||
only). `beat()` from any other thread is ignored: the update worker, Vegas's
|
||||
tick thread and the prefetcher keep running while the render thread is stuck,
|
||||
and must not vouch for it.
|
||||
|
||||
- **Check-in points.** The top of `run()`'s loop (`loop_pass()`), every
|
||||
dwell second (`_sleep_with_plugin_updates`), every frame of the per-screen
|
||||
loops (`_display_once`), every frame of Vegas's own loop and static pause
|
||||
(`coordinator.run_iteration`), each plugin fetched for a Vegas cycle
|
||||
(`StreamManager._fetch_plugin_content`), each update on the
|
||||
`synchronous_updates` path, and every frame pushed
|
||||
(`DisplayManager.update_display` -> `note_frame()`). Beats are
|
||||
rate-limited to one ping and one heartbeat write every 5 s.
|
||||
- **systemd watchdog.** `ledmatrix.service` is `Type=simple` with
|
||||
`WatchdogSec=120` and `NotifyAccess=main`. `run.py` sends
|
||||
`WATCHDOG_USEC` = 15 minutes before importing anything heavy (start-up loads
|
||||
plugins and runs the 20 s update budget, and the watchdog clock starts with
|
||||
the process). After the first frame -- or the first full pass, when there is
|
||||
nothing to draw -- the loop sends `READY=1`, restores the unit's 120 s and
|
||||
pings. `PluginManager.load_plugin()` on the render thread (a plugin enabled
|
||||
from the web UI, or loaded for on-demand) gets 15 minutes again, since it
|
||||
can run pip. A missed deadline is a SIGABRT; faulthandler, enabled on
|
||||
arming, dumps every thread's stack to the journal.
|
||||
- **Heartbeat.** `/run/ledmatrix/display-heartbeat.json`
|
||||
(`{"pid", "mono", "wall"}`; `RuntimeDirectory=ledmatrix`, 0755, file 0644 so
|
||||
the web user can read it). Readers compare `mono` with their own
|
||||
`time.monotonic()` -- CLOCK_MONOTONIC is shared by every process and does not
|
||||
jump when NTP first sets an RTC-less Pi's clock. `/api/v3/health` calls it
|
||||
`stalled` past 60 s; no file is `not_reported` and changes nothing. A clean
|
||||
stop removes it. Without `RuntimeDirectory=` (an older unit) the display,
|
||||
as root, creates the directory itself; off Linux, or without root, there
|
||||
is no heartbeat.
|
||||
|
||||
## Plugin system
|
||||
|
||||
[`src/plugin_system/`](../src/plugin_system/):
|
||||
@@ -310,7 +348,9 @@ everything else through `_reinstall_with_rollback()`.
|
||||
`ledmatrix-update-verify.path`, which runs the verifier as a separate unit
|
||||
(so restarting the web service does not kill it). The verifier restarts
|
||||
both services, waits for the web API to answer and the display service to
|
||||
stay up, and on failure resets to the previous commit and restarts again.
|
||||
stay up -- and, when the display wrote a heartbeat before the update, to
|
||||
keep one fresh from the restarted process (see Liveness) -- and on failure
|
||||
resets to the previous commit and restarts again.
|
||||
Plugin updates run only after a verified core update. State is in
|
||||
`data/auto_update_state.json` and `data/auto_update_pending.json`.
|
||||
- **Startup validator.** `StartupValidator`
|
||||
@@ -318,7 +358,9 @@ everything else through `_reinstall_with_rollback()`.
|
||||
`DisplayController.__init__`: config and cache directory first, then
|
||||
enabled plugins once the plugin manager exists. It also warns when an
|
||||
installed systemd unit differs from its template in `systemd/`. Results
|
||||
are logged; startup continues either way.
|
||||
are logged; startup continues either way. Nothing rewrites installed units
|
||||
on update: a unit change such as the watchdog reaches an existing install
|
||||
only when `install_service.sh` is re-run.
|
||||
|
||||
## Where to start reading
|
||||
|
||||
|
||||
@@ -2101,9 +2101,17 @@ Health of the web interface, display service, config file, plugin system and
|
||||
display snapshot. `data.status` is `healthy` or `degraded`, with
|
||||
`data.services` and `data.checks`.
|
||||
|
||||
`data.checks.display_loop` is the display's render-loop heartbeat: `running`
|
||||
(with `heartbeat_age_seconds`), `stalled` (no heartbeat for 60s: the panel is
|
||||
frozen even if the service is active; the status turns `degraded`), or
|
||||
`not_reported` when the display writes none (not started yet, the dev server,
|
||||
Windows), which does not affect the status.
|
||||
|
||||
Open even when the web login is on, for uptime monitors; a caller that is not
|
||||
logged in (and has no token) then gets only `{"status": "success", "data":
|
||||
{"status": "healthy" | "degraded"}}`.
|
||||
{"status": "healthy" | "degraded"}}`. A stalled render loop still shows there
|
||||
as `degraded`; the `checks` detail is only for logged-in callers, tokens and
|
||||
requests from the Pi itself.
|
||||
|
||||
### Hardware Status
|
||||
|
||||
|
||||
@@ -516,6 +516,64 @@ sudo systemctl cat ledmatrix-web | grep User
|
||||
python3 scripts/check_plugin.py --plugin plugin-id
|
||||
```
|
||||
|
||||
#### Panel Frozen, or the Display Restarts Every Few Minutes
|
||||
|
||||
**Symptoms:**
|
||||
- The panel stops changing while `systemctl status ledmatrix` says `active`
|
||||
- The display restarts on its own, a couple of minutes after it froze
|
||||
- `/api/v3/health` shows `checks.display_loop.status` as `stalled`
|
||||
|
||||
The display's render loop checks in with systemd every few seconds
|
||||
(`WatchdogSec=120` in `ledmatrix.service`) and writes a heartbeat to
|
||||
`/run/ledmatrix/display-heartbeat.json`. When the loop gets stuck -- almost
|
||||
always inside one plugin's `display()` -- the check-ins stop, and after two
|
||||
minutes systemd kills and restarts the display. The kill dumps every thread's
|
||||
stack into the log, so it says which plugin was stuck.
|
||||
|
||||
**Solutions:**
|
||||
|
||||
1. **Find the stuck plugin.** Look for the watchdog kill and the stack dump
|
||||
after it. The render loop is the thread whose stack runs through
|
||||
`display_controller.py` in `run` (usually the `Current thread` block);
|
||||
the first `plugin-repos/...` file in it is the plugin:
|
||||
```bash
|
||||
sudo journalctl -u ledmatrix --since "1 hour ago" | grep -A40 "Watchdog timeout"
|
||||
```
|
||||
|
||||
2. **Check the heartbeat by hand.** Its age should stay under about ten
|
||||
seconds while the display runs:
|
||||
```bash
|
||||
cat /run/ledmatrix/display-heartbeat.json
|
||||
curl -s http://localhost:5000/api/v3/health | python3 -m json.tool | grep -A3 display_loop
|
||||
```
|
||||
`not_reported` means the display writes no heartbeat: it has not drawn
|
||||
its first frame yet, or it runs an older version.
|
||||
|
||||
3. **Disable the plugin** in the web UI and report it to its author with the
|
||||
stack dump. Restarts that repeat back off from 10 seconds to two minutes
|
||||
apart, so a plugin that hangs on every start does not restart the display
|
||||
hundreds of times an hour.
|
||||
|
||||
4. **Is the watchdog installed?** Installs from before it keep their old unit
|
||||
until the installer is re-run (a startup warning says the unit differs
|
||||
from its template):
|
||||
```bash
|
||||
systemctl show -p WatchdogUSec ledmatrix # 2min once running; 0 = not installed
|
||||
sudo ./scripts/install/install_service.sh
|
||||
```
|
||||
`WatchdogUSec` reads `15min` for the first minutes after a start: that is
|
||||
the start-up allowance, narrowed to two minutes once the first frame is on
|
||||
the panel.
|
||||
|
||||
5. **A plugin that legitimately blocks longer** than two minutes (it should
|
||||
not; `display()` runs on the render thread) can be given more time with a
|
||||
drop-in, `sudo systemctl edit ledmatrix`:
|
||||
```ini
|
||||
[Service]
|
||||
WatchdogSec=300
|
||||
```
|
||||
`WatchdogSec=0` turns the watchdog off.
|
||||
|
||||
#### Stale Cache Data
|
||||
|
||||
**Symptoms:**
|
||||
|
||||
Reference in New Issue
Block a user