Nine CodeRabbit findings, five in code. **Health state (the one that matters).** The non-dict guard did not cover a dict missing fields the callers index directly, which is the shape actually seen in the wild: a record carrying only circuit_state produced `plugin clock-simple operation failed: 'circuit_state'` about fifty times a minute with the panel frozen. The record is now completed against the defaults per field rather than trusted or discarded wholesale. Per field matters: a first pass rejected any incomplete record outright, which reset a tripped breaker and real failure counts to healthy because one optional field was absent -- an existing test caught it. Values of the wrong type (a counter persisted as a string, an unknown circuit_state) fall back individually, valid neighbours survive, and newer fields the schema has grown since (degraded, degraded_reason) are carried through untouched. **Cache ceiling.** MemoryCache.set() accepted entries without bound between cleanup sweeps, which run every 300s by default, so a burst could take the cache far past max_size -- the unbounded growth the limit exists to stop. Eviction now runs under the same lock on every write, sharing one helper with the periodic sweep so the two cannot drift. **Installer, cgroups.** Only cgroup_enable=memory was checked, so a board carrying that without cgroup_memory=1 reported success and got no change, leaving MemoryMax= inert. Each parameter is now checked and appended independently; verified against all four combinations, single line preserved. **Installer, journald.** Persistence was inferred from /var/log/journal being non-empty, which proves neither Storage=persistent nor a size cap -- the directory survives a switch back to volatile. The effective configuration is read instead (systemd-analyze cat-config, falling back to the conf files), and an explicitly configured SystemMaxUse is preserved rather than overwritten. Verified across volatile, persistent-without-cap, persistent-with-user-cap, cap-without-storage, and commented-only configs. **Dependency extras.** _extras_are_satisfied stopped at one level, so a gated dependency that itself requests an extra (requests[socks]) passed on the base distribution's version while the extra's own dependency was missing, and pip was skipped. It now recurses, with a visited (distribution, extras) set so a cycle terminates. Docs: both kernel command-line paths documented (the installer falls back to /boot/cmdline.txt), daemon-reload and restart added after the systemd override example, memory exhaustion added to the SSH summary with its power-cycle-only recovery, and a language on the fenced block for MD040. Tests: five for the health-state repair including the exact wild shape and that record_failure/record_success no longer raise against it, and one for the cache ceiling. Both mutation-checked. Full suite 2927 passed, with the one pre-existing tmpfs failure that also fails on main. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
4.2 KiB
Running on Low-Memory Boards
Applies to the Pi Zero 2 W (512 MB), Pi 3 / 3B+ (1 GB), and the 1 GB Pi 4. If your board has 2 GB or more you can skip this document.
The failure this prevents
The display process is the largest thing on the board. On a 1 GB Pi 3B+ with around 20 plugins enabled it settles near 600 MB of 905 MB usable, leaving under 200 MB of headroom for everything else.
When that headroom runs out, the board does not crash cleanly. fork() starts
failing, and because a new process is needed to do almost anything, the
symptoms look nothing like "out of memory":
| What you see | Why |
|---|---|
| SSH accepts the connection then closes it instantly, before any banner | sshd forks a session per connection; the fork fails |
| The web UI still responds quickly | Already running, serves from existing threads, forks nothing |
| Ping is perfect, 0% loss | Handled entirely in the kernel |
| The panel is dark | The display process was killed and cannot be respawned |
| The clock is wrong after the next boot | fake-hwclock's periodic save is a scheduled job, and it cannot fork either |
The board looks healthy from the outside and cannot be logged into. Only a power cycle clears it. If you are here because SSH stopped working, also see SSH_UNAVAILABLE_AFTER_INSTALL.md, which covers the more common cause (AP mode).
Check your headroom
free -m
ps -eo rss,comm --sort=-rss | head -5
If MemAvailable is under ~150 MB while the display is running, you are close
to the edge. To watch it over time:
watch -n 30 'free -m | head -2'
Available memory that falls steadily rather than holding flat means you will reach the wall; it is a question of when.
What to do
1. Enable the memory cgroup controller. Without it, the MemoryMax=85% in
systemd/ledmatrix.service is accepted by systemd and silently ignored, so the
service has no ceiling and a runaway takes the whole board down instead of just
restarting. Raspberry Pi firmware disables this controller by default.
first_time_install.sh does this for you. To check it took effect:
grep memory /sys/fs/cgroup/cgroup.controllers
If that prints nothing, add cgroup_enable=memory cgroup_memory=1 to the
kernel command line and reboot. Edit whichever file your image uses —
/boot/firmware/cmdline.txt on current Raspberry Pi OS, /boot/cmdline.txt on
older layouts (the installer checks the first and falls back to the second).
Everything must stay on a single line.
This changes the failure mode from "the board becomes unreachable" to "the display service restarts". It is a safety net, not a fix.
2. Run fewer plugins. This is the actual remedy. Every enabled plugin costs memory permanently — its module, its parsed config, and its cached API responses. On a 512 MB or 1 GB board, keep the enabled set small and prefer plugins that poll infrequently.
3. Lower the cache ceiling. The in-memory cache is sized from total RAM (150 entries at 1 GB and below, up to 1500 at 8 GB). To go lower still:
# /etc/systemd/system/ledmatrix.service.d/override.conf
[Service]
Environment=LEDMATRIX_CACHE_MAX_ENTRIES=75
Writing the file does not change the running service. Reload systemd and restart it:
sudo systemctl daemon-reload
sudo systemctl restart ledmatrix
Fewer entries means more API calls, so lower this only while you are actually short of memory.
4. Consider MemoryHigh. MemoryMax kills and restarts. MemoryHigh
throttles and reclaims instead, which is gentler — but on a board where the
process genuinely wants more than the limit, sustained reclaim can stall the
render loop and show as visible stutter on the panel. Add it only if you prefer
degraded output to a restart:
[Service]
MemoryHigh=70%
Keep your logs
These images default to volatile journald storage, so every reboot destroys the
logs — including the ones explaining why the board rebooted. first_time_install.sh
enables persistent storage capped at 64 MB. To confirm:
journalctl --list-boots
More than one boot listed means logs are surviving reboots. If only one is
listed, journald is still writing to /run (tmpfs).