mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 06:15:09 +00:00
If the render loop gets stuck inside a plugin's display(), ledmatrix.service stays active and the panel stays frozen. This adds a way to detect that. - src/display_watchdog.py (standard library only) sends sd_notify over $NOTIFY_SOCKET and writes /run/ledmatrix/display-heartbeat.json. Only the render thread counts: beats from other threads are ignored. - ledmatrix.service: WatchdogSec=120, NotifyAccess=main, RuntimeDirectory=ledmatrix (0755), RestartSteps=4 and RestartMaxDelaySec=2min. It stays Type=simple. run.py widens the watchdog to 15 min for start-up, and load_plugin() does the same on the render thread. The loop arms after its first frame. - /api/v3/health adds checks.display_loop: running, stalled (no heartbeat for over 60s, which makes the status degraded) or not_reported. With web login on, a caller who is not logged in still gets only healthy/degraded, and a stall degrades that answer. - The update verifier requires a fresh heartbeat from the restarted display when the display it replaced was writing one. A frozen panel is rolled back. - Existing installs get the systemd watchdog only after install_service.sh is re-run. The heartbeat works right away. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
91 lines
4.8 KiB
Desktop File
91 lines
4.8 KiB
Desktop File
[Unit]
|
|
Description=LED Matrix Display Service
|
|
After=network-online.target
|
|
Wants=network-online.target
|
|
|
|
[Service]
|
|
Type=simple
|
|
User=root
|
|
WorkingDirectory=__PROJECT_ROOT_DIR__
|
|
Environment=PYTHONDONTWRITEBYTECODE=1
|
|
# glibc gives each allocating thread its own malloc arena, up to 8 x CPU count,
|
|
# and an arena that has grown is never handed back to the OS. This process runs
|
|
# 9 threads on a 3-core Pi, so the ceiling is 24 arenas -- and a rig measured at
|
|
# 1030 MB resident held 23 large anonymous mappings on 64 MB-aligned addresses,
|
|
# 920 MB of them, while the live data it was actually holding (widest scroll
|
|
# strip seen: 35,746 x 64) accounts for roughly 15 MB. That gap is arena bloat,
|
|
# not leaked objects: RSS was flat across repeated sampling, not climbing.
|
|
#
|
|
# Capping the arenas trades a little allocator concurrency for a large amount of
|
|
# resident memory on a device that has neither to spare. 2 is the usual value;
|
|
# raise it if frame times regress.
|
|
Environment=MALLOC_ARENA_MAX=2
|
|
ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/run.py
|
|
# Restart=always, not on-failure: run.py exiting 0 (a clean shutdown path taken
|
|
# for a reason that no longer applies, e.g. a config reload) would otherwise leave
|
|
# the service stopped and the panel dark indefinitely, with systemd considering
|
|
# that a successful outcome and never bringing it back.
|
|
Restart=always
|
|
RestartSec=10
|
|
# Back off when it keeps failing: 10s after the first failure, growing to two
|
|
# minutes by the fourth, so a plugin that crashes or hangs the display on
|
|
# every start retries a couple of dozen times an hour instead of hundreds.
|
|
# Deliberately not StartLimitBurst=: once that trips the unit stays failed --
|
|
# the panel dark until someone reboots -- and every start is refused until the
|
|
# interval passes, including the web UI's Start button and the automatic
|
|
# update's rollback, neither of which may run "systemctl reset-failed".
|
|
# systemd before 254 (Debian Bookworm has 252) ignores these two lines with an
|
|
# "Unknown key name" warning and keeps the flat RestartSec.
|
|
RestartSteps=4
|
|
RestartMaxDelaySec=2min
|
|
# Render-loop watchdog (src/display_watchdog.py). The render thread itself
|
|
# pings systemd every few seconds, so a render loop stuck inside a plugin --
|
|
# service still "active", panel frozen -- stops the pings, and systemd kills
|
|
# the process (SIGABRT: faulthandler writes every thread's stack to the
|
|
# journal) and restarts it.
|
|
#
|
|
# Type=simple, not Type=notify. Type=notify would hold "systemctl start" and
|
|
# "restart" until READY=1, i.e. until plugins have loaded -- minutes on a slow
|
|
# board -- and the web interface, the installer and the update health check all
|
|
# call those with timeouts well short of that. The process still sends READY=1
|
|
# (harmless here); NotifyAccess=main is what lets systemd hear it at all, and
|
|
# only from run.py itself, not from pip or anything else it starts.
|
|
#
|
|
# 120s is the steady-state limit, four times the longest gap a healthy loop
|
|
# has: a screen's first display() call runs under PluginExecutor's 30s
|
|
# timeout. Everything else the loop blocks on checks in between steps (Vegas
|
|
# frames, each plugin fetched for a Vegas cycle, each dwell second). Start-up
|
|
# is longer than this and happens before the loop exists, so the process
|
|
# widens the limit to 15 minutes as it starts and narrows it back to this
|
|
# value after its first frame; loading a plugin enabled from the web UI, which
|
|
# can run pip on the render thread, gets the same 15 minutes. Raise it with a
|
|
# drop-in (systemctl edit ledmatrix) if a plugin legitimately needs longer;
|
|
# WatchdogSec=0 turns it off.
|
|
WatchdogSec=120
|
|
NotifyAccess=main
|
|
# /run/ledmatrix, for the render loop's heartbeat (display-heartbeat.json),
|
|
# which /api/v3/health and the update health check read. /run is tmpfs, so a
|
|
# write every few seconds never touches the SD card. 0755 and root-owned: the
|
|
# web interface runs as another user and only needs to read it. Removed when
|
|
# the service stops, so a stopped display leaves no stale heartbeat behind.
|
|
RuntimeDirectory=ledmatrix
|
|
RuntimeDirectoryMode=0755
|
|
# Memory ceiling as a share of physical RAM, so one unit file suits a 512 MB
|
|
# Pi Zero 2 W and an 8 GB Pi 5 alike. This is a backstop, not a tuning knob: it
|
|
# turns "the board runs out of memory, stops being able to fork, and takes sshd
|
|
# and the panel down together until someone pulls the plug" into "this one
|
|
# service restarts".
|
|
#
|
|
# NOTE: Raspberry Pi firmware boots the kernel with cgroup_disable=memory, and
|
|
# systemd accepts this setting and then silently ignores it. Verify with:
|
|
# grep memory /sys/fs/cgroup/cgroup.controllers
|
|
# If that prints nothing, add "cgroup_enable=memory cgroup_memory=1" to
|
|
# /boot/firmware/cmdline.txt (all on line 1) and reboot. first_time_install.sh
|
|
# does this for you.
|
|
MemoryMax=85%
|
|
StandardOutput=journal
|
|
StandardError=journal
|
|
SyslogIdentifier=ledmatrix
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target |