Compare commits

...
Author SHA1 Message Date
ChuckandClaude Opus 5 ef1e9e0eee docs: guidance for 512MB and 1GB boards
Documents the memory ceiling on small boards and, more usefully, what
running into it actually looks like: sshd accepting connections and
closing them before the banner, the web UI still responding normally,
clean ping, a dark panel, and a wrong clock after the next boot. None of
those read as "out of memory", which makes the failure hard to identify
from the symptoms.

Cross-referenced from SSH_UNAVAILABLE_AFTER_INSTALL.md, since "I can't
SSH in any more" is how most people will first meet this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:20:49 -04:00
ChuckandClaude Opus 5 8927a1b6b1 perf(memory): size the cache to the board and stop reinstalling deps
On a 1GB Pi 3B+ the display process settles around 600MB RSS of 905MB
total. When the remaining headroom runs out the failure is not a clean
crash: fork() starts returning ENOMEM, so sshd accepts connections and
closes them before its banner, timer jobs stop running, and the panel
goes dark, while already-resident processes keep serving normally. The
board looks healthy from outside and cannot be logged into. Only a power
cycle clears it.

Three contributing causes:

- MemoryCache had a fixed 1000-entry ceiling. Entries are parsed API
  payloads of tens of KB, so one ceiling cannot serve both a 512MB Zero
  2 W and an 8GB Pi 5. Now scaled from MemTotal (150 entries at <=1GB,
  1500 at >=8GB), overridable with LEDMATRIX_CACHE_MAX_ENTRIES.

- requirements_are_satisfied() returned False for any requirement with
  extras, so a plugin depending on python-socketio[client] re-ran pip on
  every single start: ~8s, a network dependency, and a 100-200MB spike
  at the least convenient moment. During a restart loop it repeats for
  each restart. Extras are now resolved one level deep against installed
  metadata, keeping the conservative "anything unverifiable falls
  through to pip" contract.

- ledmatrix.service had no memory ceiling. MemoryMax=85% expressed as a
  percentage so one unit file suits every board. Note this needs the
  memory cgroup controller, which Pi firmware disables by default;
  first_time_install.sh now adds cgroup_enable=memory to cmdline.txt,
  and the unit file documents how to verify it took effect.

first_time_install.sh also enables persistent journald storage (capped
at 64M). Default storage is volatile, so every reboot destroys the logs
that would explain why the board rebooted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:13:30 -04:00
8 changed files with 278 additions and 7 deletions
+104
View File
@@ -0,0 +1,104 @@
# Running on Low-Memory Boards
Applies to the Pi Zero 2 W (512 MB), Pi 3 / 3B+ (1 GB), and the 1 GB Pi 4.
If your board has 2 GB or more you can skip this document.
## The failure this prevents
The display process is the largest thing on the board. On a 1 GB Pi 3B+ with
around 20 plugins enabled it settles near **600 MB of 905 MB usable**, leaving
under 200 MB of headroom for everything else.
When that headroom runs out, the board does not crash cleanly. `fork()` starts
failing, and because a new process is needed to do almost anything, the
symptoms look nothing like "out of memory":
| What you see | Why |
|---|---|
| SSH accepts the connection then closes it instantly, before any banner | `sshd` forks a session per connection; the fork fails |
| The web UI still responds quickly | Already running, serves from existing threads, forks nothing |
| Ping is perfect, 0% loss | Handled entirely in the kernel |
| The panel is dark | The display process was killed and cannot be respawned |
| The clock is wrong after the next boot | `fake-hwclock`'s periodic save is a scheduled job, and it cannot fork either |
The board looks healthy from the outside and cannot be logged into. Only a
power cycle clears it. If you are here because SSH stopped working, also see
[SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md), which
covers the more common cause (AP mode).
## Check your headroom
```bash
free -m
ps -eo rss,comm --sort=-rss | head -5
```
If `MemAvailable` is under ~150 MB while the display is running, you are close
to the edge. To watch it over time:
```bash
watch -n 30 'free -m | head -2'
```
Available memory that falls steadily rather than holding flat means you will
reach the wall; it is a question of when.
## What to do
**1. Enable the memory cgroup controller.** Without it, the `MemoryMax=85%` in
`systemd/ledmatrix.service` is accepted by systemd and silently ignored, so the
service has no ceiling and a runaway takes the whole board down instead of just
restarting. Raspberry Pi firmware disables this controller by default.
`first_time_install.sh` does this for you. To check it took effect:
```bash
grep memory /sys/fs/cgroup/cgroup.controllers
```
If that prints nothing, add `cgroup_enable=memory cgroup_memory=1` to the
single line in `/boot/firmware/cmdline.txt` and reboot.
This changes the failure mode from "the board becomes unreachable" to "the
display service restarts". It is a safety net, not a fix.
**2. Run fewer plugins.** This is the actual remedy. Every enabled plugin costs
memory permanently — its module, its parsed config, and its cached API
responses. On a 512 MB or 1 GB board, keep the enabled set small and prefer
plugins that poll infrequently.
**3. Lower the cache ceiling.** The in-memory cache is sized from total RAM
(150 entries at 1 GB and below, up to 1500 at 8 GB). To go lower still:
```ini
# /etc/systemd/system/ledmatrix.service.d/override.conf
[Service]
Environment=LEDMATRIX_CACHE_MAX_ENTRIES=75
```
Fewer entries means more API calls, so lower this only while you are actually
short of memory.
**4. Consider `MemoryHigh`.** `MemoryMax` kills and restarts. `MemoryHigh`
throttles and reclaims instead, which is gentler — but on a board where the
process genuinely wants more than the limit, sustained reclaim can stall the
render loop and show as visible stutter on the panel. Add it only if you prefer
degraded output to a restart:
```ini
[Service]
MemoryHigh=70%
```
## Keep your logs
These images default to volatile journald storage, so every reboot destroys the
logs — including the ones explaining why the board rebooted. `first_time_install.sh`
enables persistent storage capped at 64 MB. To confirm:
```bash
journalctl --list-boots
```
More than one boot listed means logs are surviving reboots. If only one is
listed, journald is still writing to `/run` (tmpfs).
+1
View File
@@ -14,6 +14,7 @@ the one-shot installer. The pages here go deeper.
5. [TROUBLESHOOTING.md](TROUBLESHOOTING.md) — common issues and fixes
6. [SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md) — recovering SSH after install
7. [CONFIG_DEBUGGING.md](CONFIG_DEBUGGING.md) — diagnosing config problems
8. [LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md) — Pi Zero 2 W / 3B+ / 1GB Pi 4 memory limits
## I want to write a plugin
+16 -1
View File
@@ -20,7 +20,22 @@ The installation script:
- Installs and configures `dnsmasq` (DHCP server for AP mode)
- These services can interfere with normal WiFi client mode
### 3. Reboot After Installation
### 3. The Board Ran Out of Memory
On a 512MB or 1GB board, memory exhaustion stops `sshd` being able to fork a
session process. The connection is accepted and then closed immediately, before
any banner:
```
kex_exchange_identification: Connection closed by remote host
```
The giveaway is that the board is otherwise healthy — ping is clean and the web
UI still responds — but nothing that needs to start a new process works, and
the panel is usually dark. Only a power cycle clears it. See
[LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md).
### 4. Reboot After Installation
If the script reboots the Pi (which it recommends), network services may restart in a different state, potentially triggering AP mode.
+37
View File
@@ -1688,6 +1688,43 @@ else
echo "$CMDLINE_FILE not found; skipping isolcpus optimization"
fi
# Enable the memory cgroup controller (idempotent).
# The Pi firmware boots with cgroup_disable=memory, so systemd's MemoryMax= is
# accepted and silently ignored — the display service then has no ceiling, and
# a runaway takes the whole board down (sshd can no longer fork, the panel goes
# dark) rather than just restarting the one service.
if [ "$SKIP_PERF" != "1" ] && [ -f "$CMDLINE_FILE" ]; then
if grep -q 'cgroup_enable=memory' "$CMDLINE_FILE"; then
echo "cgroup_enable=memory already present in $CMDLINE_FILE"
else
echo "Adding cgroup_enable=memory to $CMDLINE_FILE..."
cp "$CMDLINE_FILE" "$CMDLINE_FILE.bak" 2>/dev/null || true
sed -i '1 s/$/ cgroup_enable=memory cgroup_memory=1/' "$CMDLINE_FILE"
echo " Takes effect after reboot. Verify with:"
echo " grep memory /sys/fs/cgroup/cgroup.controllers"
fi
fi
# Persist the journal (idempotent).
# These images default to volatile storage: journald keeps everything in /run
# (tmpfs), so every reboot destroys the logs — including the ones that would
# explain why the board rebooted. Capped so an SD card is not worn out by logs.
if [ -d /var/log/journal ] && [ -n "$(ls -A /var/log/journal 2>/dev/null)" ]; then
echo "Persistent journald storage already enabled"
else
echo "Enabling persistent journald storage..."
mkdir -p /etc/systemd/journald.conf.d
cat > /etc/systemd/journald.conf.d/ledmatrix-persistent.conf <<'JOURNALD'
# Installed by LEDMatrix first_time_install.sh
[Journal]
Storage=persistent
SystemMaxUse=64M
JOURNALD
mkdir -p /var/log/journal
systemd-tmpfiles --create --prefix /var/log/journal >/dev/null 2>&1 || true
systemctl restart systemd-journald >/dev/null 2>&1 || true
fi
# Ensure dtparam=audio=off in config.txt (idempotent)
if [ "$SKIP_PERF" = "1" ]; then
: # skipped
+47
View File
@@ -4,11 +4,58 @@ Memory Cache
Handles in-memory caching with TTL support, size limits, and automatic cleanup.
"""
import os
import time
import threading
import logging
from typing import Dict, Any, Optional
# Historical fixed ceiling, kept as the fallback when RAM cannot be read.
DEFAULT_MAX_SIZE = 1000
def _total_memory_mb() -> Optional[float]:
"""Physical RAM in MB, or None where /proc/meminfo is unavailable."""
try:
with open('/proc/meminfo', 'r', encoding='utf-8') as fh:
for line in fh:
if line.startswith('MemTotal:'):
return int(line.split()[1]) / 1024
except (OSError, ValueError, IndexError):
return None
return None
def default_max_size() -> int:
"""Entry ceiling scaled to this machine's RAM.
One fixed ceiling cannot serve both a 512 MB Pi Zero 2 W and an 8 GB Pi 5.
Entries here are parsed API payloads that routinely run tens of kilobytes
each, so a thousand of them is a comfortable cache on a large board and a
substantial fraction of total RAM on a small one — where the process
competing for that RAM is also driving the panel. Set
LEDMATRIX_CACHE_MAX_ENTRIES to override.
"""
override = os.environ.get('LEDMATRIX_CACHE_MAX_ENTRIES')
if override:
try:
value = int(override)
if value > 0:
return value
except ValueError:
pass
total_mb = _total_memory_mb()
if total_mb is None:
return DEFAULT_MAX_SIZE
if total_mb < 1536: # 512 MB and 1 GB boards
return 150
if total_mb < 3072: # 2 GB
return 400
if total_mb < 6144: # 4 GB
return 800
return 1500 # 8 GB and up
class MemoryCache:
"""Manages in-memory cache with TTL and size limits."""
+4 -2
View File
@@ -33,7 +33,7 @@ import logging
import threading
import tempfile
from src.exceptions import CacheError
from src.cache.memory_cache import MemoryCache
from src.cache.memory_cache import MemoryCache, default_max_size
from src.cache.disk_cache import DiskCache
from src.cache.cache_strategy import CacheStrategy
from src.cache.cache_metrics import CacheMetrics
@@ -84,7 +84,9 @@ class CacheManager:
self.logger.warning("ConfigManager not available, using default cache intervals")
# Initialize cache components using composition
self._memory_cache_component = MemoryCache(max_size=1000, cleanup_interval=300.0)
self._memory_cache_component = MemoryCache(
max_size=default_max_size(), cleanup_interval=300.0
)
self._disk_cache_component = DiskCache(cache_dir=self.cache_dir, logger=self.logger)
self._strategy_component = CacheStrategy(config_manager=self.config_manager, logger=self.logger)
self._metrics_component = CacheMetrics(logger=self.logger)
+56 -4
View File
@@ -14,7 +14,7 @@ import sys
import subprocess
import threading
from pathlib import Path
from typing import Dict, Any, Optional, Tuple, Type
from typing import Dict, Any, List, Optional, Tuple, Type
import logging
from packaging.requirements import InvalidRequirement, Requirement
@@ -45,6 +45,58 @@ def requirements_has_real_deps(requirements_file: str) -> bool:
return False
def _extra_dependencies(dist_name: str, extras) -> Optional[List[Requirement]]:
"""Dependencies a distribution declares *only* behind the given extras.
Returns None when the installed metadata cannot be read or parsed, so the
caller can fall back to running pip rather than assuming anything.
"""
try:
meta = importlib.metadata.metadata(dist_name)
except importlib.metadata.PackageNotFoundError:
return None
gated: List[Requirement] = []
for raw in meta.get_all('Requires-Dist') or []:
try:
dep = Requirement(raw)
except InvalidRequirement:
return None
if dep.marker is None:
continue
# Keep only what the distribution gates behind an extra we asked for:
# satisfied when `extra` is that name, but not when no extra is
# requested. A marker that holds either way (python_version, sys_platform)
# belongs to the base install and is already covered by the version check.
if dep.marker.evaluate({'extra': ''}):
continue
if any(dep.marker.evaluate({'extra': extra}) for extra in extras):
gated.append(dep)
return gated
def _extras_are_satisfied(req: Requirement) -> bool:
"""Check the dependencies pulled in by req's extras are installed.
One level deep, not transitive: enough to tell "the extra was installed"
from "the extra was never installed", which is all the caller needs to
decide whether pip has work to do. Anything unreadable returns False, so
the caller still falls through to pip.
"""
gated = _extra_dependencies(req.name, req.extras)
if gated is None:
return False
for dep in gated:
try:
dep_version = importlib.metadata.version(dep.name)
except importlib.metadata.PackageNotFoundError:
return False
if dep.specifier and not dep.specifier.contains(dep_version, prereleases=True):
return False
return True
def requirements_are_satisfied(requirements_file: str) -> bool:
"""
Check whether every real requirement line in requirements.txt is already
@@ -76,9 +128,6 @@ def requirements_are_satisfied(requirements_file: str) -> bool:
except InvalidRequirement:
return False
if req.extras:
return False # verifying extras' sub-dependencies isn't worth it here
if req.marker is not None and not req.marker.evaluate():
continue # not applicable on this platform/interpreter
@@ -90,6 +139,9 @@ def requirements_are_satisfied(requirements_file: str) -> bool:
if req.specifier and not req.specifier.contains(installed_version, prereleases=True):
return False
if req.extras and not _extras_are_satisfied(req):
return False
return True
+13
View File
@@ -15,6 +15,19 @@ ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/run.py
# that a successful outcome and never bringing it back.
Restart=always
RestartSec=10
# Memory ceiling as a share of physical RAM, so one unit file suits a 512 MB
# Pi Zero 2 W and an 8 GB Pi 5 alike. This is a backstop, not a tuning knob: it
# turns "the board runs out of memory, stops being able to fork, and takes sshd
# and the panel down together until someone pulls the plug" into "this one
# service restarts".
#
# NOTE: Raspberry Pi firmware boots the kernel with cgroup_disable=memory, and
# systemd accepts this setting and then silently ignores it. Verify with:
# grep memory /sys/fs/cgroup/cgroup.controllers
# If that prints nothing, add "cgroup_enable=memory cgroup_memory=1" to
# /boot/firmware/cmdline.txt (all on line 1) and reboot. first_time_install.sh
# does this for you.
MemoryMax=85%
StandardOutput=journal
StandardError=journal
SyslogIdentifier=ledmatrix