Compare commits

...
Author SHA1 Message Date
ChuckandClaude Opus 5 ef1e9e0eee docs: guidance for 512MB and 1GB boards
Documents the memory ceiling on small boards and, more usefully, what
running into it actually looks like: sshd accepting connections and
closing them before the banner, the web UI still responding normally,
clean ping, a dark panel, and a wrong clock after the next boot. None of
those read as "out of memory", which makes the failure hard to identify
from the symptoms.

Cross-referenced from SSH_UNAVAILABLE_AFTER_INSTALL.md, since "I can't
SSH in any more" is how most people will first meet this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:20:49 -04:00
ChuckandClaude Opus 5 8927a1b6b1 perf(memory): size the cache to the board and stop reinstalling deps
On a 1GB Pi 3B+ the display process settles around 600MB RSS of 905MB
total. When the remaining headroom runs out the failure is not a clean
crash: fork() starts returning ENOMEM, so sshd accepts connections and
closes them before its banner, timer jobs stop running, and the panel
goes dark, while already-resident processes keep serving normally. The
board looks healthy from outside and cannot be logged into. Only a power
cycle clears it.

Three contributing causes:

- MemoryCache had a fixed 1000-entry ceiling. Entries are parsed API
  payloads of tens of KB, so one ceiling cannot serve both a 512MB Zero
  2 W and an 8GB Pi 5. Now scaled from MemTotal (150 entries at <=1GB,
  1500 at >=8GB), overridable with LEDMATRIX_CACHE_MAX_ENTRIES.

- requirements_are_satisfied() returned False for any requirement with
  extras, so a plugin depending on python-socketio[client] re-ran pip on
  every single start: ~8s, a network dependency, and a 100-200MB spike
  at the least convenient moment. During a restart loop it repeats for
  each restart. Extras are now resolved one level deep against installed
  metadata, keeping the conservative "anything unverifiable falls
  through to pip" contract.

- ledmatrix.service had no memory ceiling. MemoryMax=85% expressed as a
  percentage so one unit file suits every board. Note this needs the
  memory cgroup controller, which Pi firmware disables by default;
  first_time_install.sh now adds cgroup_enable=memory to cmdline.txt,
  and the unit file documents how to verify it took effect.

first_time_install.sh also enables persistent journald storage (capped
at 64M). Default storage is volatile, so every reboot destroys the logs
that would explain why the board rebooted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:13:30 -04:00
ChuckandClaude Opus 5 e6249dcc7e fix(service): survive corrupt health cache and clean exits
Three independent failure modes that each end with a dark panel and no
automatic recovery.

1. PluginHealthTracker._load_health_state returned the cached value
   verbatim. If that value is not a dict, every caller raises
   AttributeError: 'list' object has no attribute 'get' — during
   DisplayController.__init__, so the process dies before the display
   loop starts. systemd restarts it, the same bad entry is read back
   from disk, and it dies again: an unattended restart loop that
   survives reboots because the cause is persisted. Observed in the
   field with plugin_health:<id> holding an unrelated plugin's list
   payload. Now non-dict entries are discarded with a warning and the
   defaults are rebuilt.

2. ledmatrix.service used Restart=on-failure, so any exit with status 0
   left the unit stopped and the panel dark indefinitely — systemd
   treats it as success and never brings it back. Restart=always.

3. ledmatrix-wifi-monitor.service used StandardOutput=syslog, which
   systemd has marked obsolete; it warns and rewrites it to journal on
   every load.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 19:04:57 -04:00
9083df9f5c fix(pixlet): resolve the release tag correctly when downloading (#461)
* fix(pixlet): resolve the release tag correctly when downloading

Starlark apps render through the pixlet binary, and the installer that
fetches it silently produced nothing, so every app failed with "Pixlet
not available - Starlark apps will not work".

Two compounding defects:

The version lookup parsed the wrong token. GitHub returns the release
JSON on a single line, so `grep '"tag_name"'` matches the whole document
and the greedy `sed 's/.*"([^"]+)".*/\1/'` captures the LAST quoted
string in it. That resolved to "mentions_count", giving a download URL
for a release that does not exist. The `[ -z "$PIXLET_VERSION" ]`
fallback never fired, because the value was not empty -- just wrong.

And `curl -L -o` without `-f` writes a 404 body to the file and exits 0,
so the download was reported as successful and the first sign of trouble
was tar complaining "not in gzip format" about a page of HTML:

    → Downloading linux-arm64...
      Extracting...
    gzip: stdin: not in gzip format
    ✗ Failed to extract archive: .../pixlet_mentions_count_linux-arm64.tar.gz
    Download complete: 0/1 succeeded

Now the tag field is matched directly and the value taken from it, and
the result is checked for a version shape rather than merely being
non-empty -- a wrong-but-non-empty value is exactly what made this
silent. curl gets -f so an HTTP error is a failure, and the archive is
gzip-tested before extraction, since a proxy can return 200 with an
error page.

Verified on an arm64 rig: v0.53.1 resolved, 1/1 downloaded, the binary
runs, and the plugin's own detection finds it at
bin/pixlet/pixlet-linux-arm64.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(pixlet): anchor the version check, and don't echo response bytes raw

Both CodeRabbit findings were valid.

The shape check accepted partial matches, so "v0.53garbage", "0.53" and
"v0.5" passed it and built a download URL for a release that cannot exist
-- the failure the check was added to stop, just one step later. Anchored
at both ends now. Every tronbyt/pixlet release to date is vX.Y.Z (all 38
verified against the API), with an optional suffix left for a future -rc.1
or +build tag.

The invalid-response diagnostic printed bytes straight from whatever
answered the request. NUL and newline were filtered but escape, carriage
return and backspace were not, so an error page could rewrite the output
or bury it in a CI log. Non-printable bytes are stripped and it goes
through printf. CodeRabbit suggested hex-encoding the lot; printable
characters are kept instead, because "<!DOCTYPE html>" is the diagnostic
-- hex would make the line safe and useless.

Also corrected the comment above the parse. It asserted GitHub returns
this JSON on a single line; the API is pretty-printed by default, and I
could not get a single-line response from two machines across five header
variants. The single-line case is real (it is what produces
"mentions_count", and the failing device's error named
pixlet_mentions_count_linux-arm64.tar.gz), but it is a shape to be robust
against, not a constant. As written the comment invites the next reader to
check by hand, see pretty JSON, and conclude the fix was unnecessary.

Tests drive the real script with a stubbed curl: the tag resolves from
both response shapes, non-release values fall back, an HTTP error is
reported as a download failure rather than surfacing later as a tar error,
a non-archive body is rejected before extraction, and the diagnostic
cannot carry control bytes. The stub honours -f the way real curl does --
without that, the HTTP-error test passed against the old script too, since
both end at 0/1 and only the reporting layer differs.

Mutation-checked: 10 of the 16 fail against the pre-fix script, and the 5
covering these two findings fail against this branch's previous state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:35:02 -04:00
0901d044d3 fix(plugins): report updates that completed, not ones that were queued (#460)
run_scheduled_updates_with_changes() snapshotted plugin_last_update,
called run_scheduled_updates(), and diffed the two to answer "whose data
just changed".

But run_scheduled_updates() only enqueues. The work runs on the update
worker and stamps plugin_last_update there, after this method has already
returned, so the two snapshots were always identical and the result was
always an empty list. The only path that ever worked was the synchronous
kill-switch, where update() runs inline.

Vegas is the caller. That empty list is what feeds mark_plugin_updated(),
which drops the cached content for a plugin whose data moved -- so a
segment kept scrolling whatever it was first built from. It is the
failure the coordinator's own comments describe: last night's live game
still drawn as live the next morning. On a live rig: zero update ticks in
twenty minutes, with weather, stocks and news all updating on schedule.

The worker now records each completed update in a ledger and the call
drains it, reporting what has finished since the previous poll rather
than what this call enqueued. That costs one tick of latency -- Vegas
polls every ~4s -- and is correct whichever side of the queue the work
lands on. Failure paths are excluded: they stamp the timestamp too, to
space out retries, but no fresh data exists.

Verified on the rig it was found on: 0 update ticks before, 208 in
twenty-five minutes after, naming real plugins.

The behavioural tests here would pass with both production call sites
deleted, which mutation testing caught -- they drive the ledger directly.
So there is also a structural test asserting the invariant at the source:
wherever a successful update stamps plugin_last_update, it must record
the completion. Writing it immediately caught that _record_update_failure
stamps the same field and must not be included.

Mutation-checked: removing either call site, removing both, and dropping
the drain's clear are all caught.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:28:40 -04:00
14 changed files with 743 additions and 34 deletions
+104
View File
@@ -0,0 +1,104 @@
# Running on Low-Memory Boards
Applies to the Pi Zero 2 W (512 MB), Pi 3 / 3B+ (1 GB), and the 1 GB Pi 4.
If your board has 2 GB or more you can skip this document.
## The failure this prevents
The display process is the largest thing on the board. On a 1 GB Pi 3B+ with
around 20 plugins enabled it settles near **600 MB of 905 MB usable**, leaving
under 200 MB of headroom for everything else.
When that headroom runs out, the board does not crash cleanly. `fork()` starts
failing, and because a new process is needed to do almost anything, the
symptoms look nothing like "out of memory":
| What you see | Why |
|---|---|
| SSH accepts the connection then closes it instantly, before any banner | `sshd` forks a session per connection; the fork fails |
| The web UI still responds quickly | Already running, serves from existing threads, forks nothing |
| Ping is perfect, 0% loss | Handled entirely in the kernel |
| The panel is dark | The display process was killed and cannot be respawned |
| The clock is wrong after the next boot | `fake-hwclock`'s periodic save is a scheduled job, and it cannot fork either |
The board looks healthy from the outside and cannot be logged into. Only a
power cycle clears it. If you are here because SSH stopped working, also see
[SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md), which
covers the more common cause (AP mode).
## Check your headroom
```bash
free -m
ps -eo rss,comm --sort=-rss | head -5
```
If `MemAvailable` is under ~150 MB while the display is running, you are close
to the edge. To watch it over time:
```bash
watch -n 30 'free -m | head -2'
```
Available memory that falls steadily rather than holding flat means you will
reach the wall; it is a question of when.
## What to do
**1. Enable the memory cgroup controller.** Without it, the `MemoryMax=85%` in
`systemd/ledmatrix.service` is accepted by systemd and silently ignored, so the
service has no ceiling and a runaway takes the whole board down instead of just
restarting. Raspberry Pi firmware disables this controller by default.
`first_time_install.sh` does this for you. To check it took effect:
```bash
grep memory /sys/fs/cgroup/cgroup.controllers
```
If that prints nothing, add `cgroup_enable=memory cgroup_memory=1` to the
single line in `/boot/firmware/cmdline.txt` and reboot.
This changes the failure mode from "the board becomes unreachable" to "the
display service restarts". It is a safety net, not a fix.
**2. Run fewer plugins.** This is the actual remedy. Every enabled plugin costs
memory permanently — its module, its parsed config, and its cached API
responses. On a 512 MB or 1 GB board, keep the enabled set small and prefer
plugins that poll infrequently.
**3. Lower the cache ceiling.** The in-memory cache is sized from total RAM
(150 entries at 1 GB and below, up to 1500 at 8 GB). To go lower still:
```ini
# /etc/systemd/system/ledmatrix.service.d/override.conf
[Service]
Environment=LEDMATRIX_CACHE_MAX_ENTRIES=75
```
Fewer entries means more API calls, so lower this only while you are actually
short of memory.
**4. Consider `MemoryHigh`.** `MemoryMax` kills and restarts. `MemoryHigh`
throttles and reclaims instead, which is gentler — but on a board where the
process genuinely wants more than the limit, sustained reclaim can stall the
render loop and show as visible stutter on the panel. Add it only if you prefer
degraded output to a restart:
```ini
[Service]
MemoryHigh=70%
```
## Keep your logs
These images default to volatile journald storage, so every reboot destroys the
logs — including the ones explaining why the board rebooted. `first_time_install.sh`
enables persistent storage capped at 64 MB. To confirm:
```bash
journalctl --list-boots
```
More than one boot listed means logs are surviving reboots. If only one is
listed, journald is still writing to `/run` (tmpfs).
+1
View File
@@ -14,6 +14,7 @@ the one-shot installer. The pages here go deeper.
5. [TROUBLESHOOTING.md](TROUBLESHOOTING.md) — common issues and fixes 5. [TROUBLESHOOTING.md](TROUBLESHOOTING.md) — common issues and fixes
6. [SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md) — recovering SSH after install 6. [SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md) — recovering SSH after install
7. [CONFIG_DEBUGGING.md](CONFIG_DEBUGGING.md) — diagnosing config problems 7. [CONFIG_DEBUGGING.md](CONFIG_DEBUGGING.md) — diagnosing config problems
8. [LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md) — Pi Zero 2 W / 3B+ / 1GB Pi 4 memory limits
## I want to write a plugin ## I want to write a plugin
+16 -1
View File
@@ -20,7 +20,22 @@ The installation script:
- Installs and configures `dnsmasq` (DHCP server for AP mode) - Installs and configures `dnsmasq` (DHCP server for AP mode)
- These services can interfere with normal WiFi client mode - These services can interfere with normal WiFi client mode
### 3. Reboot After Installation ### 3. The Board Ran Out of Memory
On a 512MB or 1GB board, memory exhaustion stops `sshd` being able to fork a
session process. The connection is accepted and then closed immediately, before
any banner:
```
kex_exchange_identification: Connection closed by remote host
```
The giveaway is that the board is otherwise healthy — ping is clean and the web
UI still responds — but nothing that needs to start a new process works, and
the panel is usually dark. Only a power cycle clears it. See
[LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md).
### 4. Reboot After Installation
If the script reboots the Pi (which it recommends), network services may restart in a different state, potentially triggering AP mode. If the script reboots the Pi (which it recommends), network services may restart in a different state, potentially triggering AP mode.
+37
View File
@@ -1688,6 +1688,43 @@ else
echo "$CMDLINE_FILE not found; skipping isolcpus optimization" echo "$CMDLINE_FILE not found; skipping isolcpus optimization"
fi fi
# Enable the memory cgroup controller (idempotent).
# The Pi firmware boots with cgroup_disable=memory, so systemd's MemoryMax= is
# accepted and silently ignored — the display service then has no ceiling, and
# a runaway takes the whole board down (sshd can no longer fork, the panel goes
# dark) rather than just restarting the one service.
if [ "$SKIP_PERF" != "1" ] && [ -f "$CMDLINE_FILE" ]; then
if grep -q 'cgroup_enable=memory' "$CMDLINE_FILE"; then
echo "cgroup_enable=memory already present in $CMDLINE_FILE"
else
echo "Adding cgroup_enable=memory to $CMDLINE_FILE..."
cp "$CMDLINE_FILE" "$CMDLINE_FILE.bak" 2>/dev/null || true
sed -i '1 s/$/ cgroup_enable=memory cgroup_memory=1/' "$CMDLINE_FILE"
echo " Takes effect after reboot. Verify with:"
echo " grep memory /sys/fs/cgroup/cgroup.controllers"
fi
fi
# Persist the journal (idempotent).
# These images default to volatile storage: journald keeps everything in /run
# (tmpfs), so every reboot destroys the logs — including the ones that would
# explain why the board rebooted. Capped so an SD card is not worn out by logs.
if [ -d /var/log/journal ] && [ -n "$(ls -A /var/log/journal 2>/dev/null)" ]; then
echo "Persistent journald storage already enabled"
else
echo "Enabling persistent journald storage..."
mkdir -p /etc/systemd/journald.conf.d
cat > /etc/systemd/journald.conf.d/ledmatrix-persistent.conf <<'JOURNALD'
# Installed by LEDMatrix first_time_install.sh
[Journal]
Storage=persistent
SystemMaxUse=64M
JOURNALD
mkdir -p /var/log/journal
systemd-tmpfiles --create --prefix /var/log/journal >/dev/null 2>&1 || true
systemctl restart systemd-journald >/dev/null 2>&1 || true
fi
# Ensure dtparam=audio=off in config.txt (idempotent) # Ensure dtparam=audio=off in config.txt (idempotent)
if [ "$SKIP_PERF" = "1" ]; then if [ "$SKIP_PERF" = "1" ]; then
: # skipped : # skipped
+43 -5
View File
@@ -24,9 +24,29 @@ echo "========================================"
# Auto-detect latest version if needed # Auto-detect latest version if needed
if [ "$PIXLET_VERSION" = "latest" ]; then if [ "$PIXLET_VERSION" = "latest" ]; then
echo "Detecting latest version..." echo "Detecting latest version..."
PIXLET_VERSION=$(curl -s "https://api.github.com/repos/${REPO}/releases/latest" | grep '"tag_name"' | sed -E 's/.*"([^"]+)".*/\1/') # When this response arrives on a single line -- as it did on the device
if [ -z "$PIXLET_VERSION" ]; then # where Starlark apps were failing -- `grep '"tag_name"'` matches the whole
echo "Failed to detect latest version, using fallback" # document and a greedy `sed 's/.*"([^"]+)".*/\1/'` captures the LAST
# quoted token in it rather than the tag. That resolved to
# "mentions_count", which built a download URL for a release that does not
# exist. (The API is pretty-printed by default, which is why the old
# command looks correct when you try it by hand -- but the formatting is
# not something to depend on.) Match the field itself and take the value
# after it, which is right for either shape.
PIXLET_VERSION=$(curl -fsSL "https://api.github.com/repos/${REPO}/releases/latest" \
| grep -o '"tag_name"[[:space:]]*:[[:space:]]*"[^"]*"' \
| head -n1 \
| sed -E 's/.*:[[:space:]]*"([^"]*)".*/\1/')
# A wrong-but-non-empty value is what made the old bug silent, so check the
# shape rather than just that something came back. Anchored at both ends: a
# partial match would accept "v0.53garbage" or "0.53" and build a URL for a
# release that cannot exist, which is the failure this check is here to
# stop. Every tronbyt/pixlet release to date is vX.Y.Z; the optional suffix
# leaves room for a future -rc.1 or +build tag.
if ! printf '%s' "$PIXLET_VERSION" \
| grep -qE '^v[0-9]+\.[0-9]+\.[0-9]+([-+][0-9A-Za-z.-]+)?$'; then
echo "Could not detect the latest version (got: '${PIXLET_VERSION:-<empty>}'), using fallback"
PIXLET_VERSION="v0.50.2" PIXLET_VERSION="v0.50.2"
fi fi
fi fi
@@ -67,8 +87,26 @@ download_binary() {
temp_dir=$(mktemp -d -p "$PROJECT_ROOT" -t pixlet_download.XXXXXXXXXX) temp_dir=$(mktemp -d -p "$PROJECT_ROOT" -t pixlet_download.XXXXXXXXXX)
local temp_file="$temp_dir/$archive_name" local temp_file="$temp_dir/$archive_name"
if ! curl -L -o "$temp_file" "$url" 2>/dev/null; then # -f so an HTTP error is a failure. Without it curl writes the 404 body
echo "✗ Failed to download $arch" # to the file and exits 0, and the first sign of trouble is tar saying
# "not in gzip format" about what is actually a page of HTML.
if ! curl -fL -o "$temp_file" "$url" 2>/dev/null; then
echo "✗ Failed to download $arch from $url"
rm -rf "$temp_dir"
return 1
fi
# Belt and braces: a mirror or proxy can return 200 with an error page.
if ! gzip -t "$temp_file" 2>/dev/null; then
echo "✗ Downloaded file is not a gzip archive: $url"
# These bytes come from whatever answered the request, so strip
# everything non-printable before echoing them: an error page carrying
# terminal escapes would otherwise be able to rewrite this output or
# bury it in a CI log. Printable characters are kept rather than
# hex-encoding the lot, because "<!DOCTYPE html>" is the diagnostic.
local first_bytes
first_bytes=$(head -c 60 "$temp_file" | tr -cd '[:print:]')
printf ' (first bytes: %s)\n' "$first_bytes"
rm -rf "$temp_dir" rm -rf "$temp_dir"
return 1 return 1
fi fi
+47
View File
@@ -4,11 +4,58 @@ Memory Cache
Handles in-memory caching with TTL support, size limits, and automatic cleanup. Handles in-memory caching with TTL support, size limits, and automatic cleanup.
""" """
import os
import time import time
import threading import threading
import logging import logging
from typing import Dict, Any, Optional from typing import Dict, Any, Optional
# Historical fixed ceiling, kept as the fallback when RAM cannot be read.
DEFAULT_MAX_SIZE = 1000
def _total_memory_mb() -> Optional[float]:
"""Physical RAM in MB, or None where /proc/meminfo is unavailable."""
try:
with open('/proc/meminfo', 'r', encoding='utf-8') as fh:
for line in fh:
if line.startswith('MemTotal:'):
return int(line.split()[1]) / 1024
except (OSError, ValueError, IndexError):
return None
return None
def default_max_size() -> int:
"""Entry ceiling scaled to this machine's RAM.
One fixed ceiling cannot serve both a 512 MB Pi Zero 2 W and an 8 GB Pi 5.
Entries here are parsed API payloads that routinely run tens of kilobytes
each, so a thousand of them is a comfortable cache on a large board and a
substantial fraction of total RAM on a small one — where the process
competing for that RAM is also driving the panel. Set
LEDMATRIX_CACHE_MAX_ENTRIES to override.
"""
override = os.environ.get('LEDMATRIX_CACHE_MAX_ENTRIES')
if override:
try:
value = int(override)
if value > 0:
return value
except ValueError:
pass
total_mb = _total_memory_mb()
if total_mb is None:
return DEFAULT_MAX_SIZE
if total_mb < 1536: # 512 MB and 1 GB boards
return 150
if total_mb < 3072: # 2 GB
return 400
if total_mb < 6144: # 4 GB
return 800
return 1500 # 8 GB and up
class MemoryCache: class MemoryCache:
"""Manages in-memory cache with TTL and size limits.""" """Manages in-memory cache with TTL and size limits."""
+4 -2
View File
@@ -33,7 +33,7 @@ import logging
import threading import threading
import tempfile import tempfile
from src.exceptions import CacheError from src.exceptions import CacheError
from src.cache.memory_cache import MemoryCache from src.cache.memory_cache import MemoryCache, default_max_size
from src.cache.disk_cache import DiskCache from src.cache.disk_cache import DiskCache
from src.cache.cache_strategy import CacheStrategy from src.cache.cache_strategy import CacheStrategy
from src.cache.cache_metrics import CacheMetrics from src.cache.cache_metrics import CacheMetrics
@@ -84,7 +84,9 @@ class CacheManager:
self.logger.warning("ConfigManager not available, using default cache intervals") self.logger.warning("ConfigManager not available, using default cache intervals")
# Initialize cache components using composition # Initialize cache components using composition
self._memory_cache_component = MemoryCache(max_size=1000, cleanup_interval=300.0) self._memory_cache_component = MemoryCache(
max_size=default_max_size(), cleanup_interval=300.0
)
self._disk_cache_component = DiskCache(cache_dir=self.cache_dir, logger=self.logger) self._disk_cache_component = DiskCache(cache_dir=self.cache_dir, logger=self.logger)
self._strategy_component = CacheStrategy(config_manager=self.config_manager, logger=self.logger) self._strategy_component = CacheStrategy(config_manager=self.config_manager, logger=self.logger)
self._metrics_component = CacheMetrics(logger=self.logger) self._metrics_component = CacheMetrics(logger=self.logger)
+12 -1
View File
@@ -64,9 +64,20 @@ class PluginHealthTracker:
cache_key, max_age=None, memory_ttl=0 if force_reload else None cache_key, max_age=None, memory_ttl=0 if force_reload else None
) )
if cached: if isinstance(cached, dict) and cached:
return cached return cached
# A cache entry that is not a dict means the persisted state was written
# by something other than _save_health_state (a key collision, a partial
# write, a restored backup). Returning it verbatim makes every caller
# blow up on .get(), which takes the display down in a restart loop that
# survives reboots because the bad entry is on disk. Discard and rebuild.
if cached is not None and not isinstance(cached, dict):
self.logger.warning(
f"Discarding malformed health state for {plugin_id}: expected "
f"dict, got {type(cached).__name__}. Falling back to defaults."
)
# Default state # Default state
return { return {
'consecutive_failures': 0, 'consecutive_failures': 0,
+56 -4
View File
@@ -14,7 +14,7 @@ import sys
import subprocess import subprocess
import threading import threading
from pathlib import Path from pathlib import Path
from typing import Dict, Any, Optional, Tuple, Type from typing import Dict, Any, List, Optional, Tuple, Type
import logging import logging
from packaging.requirements import InvalidRequirement, Requirement from packaging.requirements import InvalidRequirement, Requirement
@@ -45,6 +45,58 @@ def requirements_has_real_deps(requirements_file: str) -> bool:
return False return False
def _extra_dependencies(dist_name: str, extras) -> Optional[List[Requirement]]:
"""Dependencies a distribution declares *only* behind the given extras.
Returns None when the installed metadata cannot be read or parsed, so the
caller can fall back to running pip rather than assuming anything.
"""
try:
meta = importlib.metadata.metadata(dist_name)
except importlib.metadata.PackageNotFoundError:
return None
gated: List[Requirement] = []
for raw in meta.get_all('Requires-Dist') or []:
try:
dep = Requirement(raw)
except InvalidRequirement:
return None
if dep.marker is None:
continue
# Keep only what the distribution gates behind an extra we asked for:
# satisfied when `extra` is that name, but not when no extra is
# requested. A marker that holds either way (python_version, sys_platform)
# belongs to the base install and is already covered by the version check.
if dep.marker.evaluate({'extra': ''}):
continue
if any(dep.marker.evaluate({'extra': extra}) for extra in extras):
gated.append(dep)
return gated
def _extras_are_satisfied(req: Requirement) -> bool:
"""Check the dependencies pulled in by req's extras are installed.
One level deep, not transitive: enough to tell "the extra was installed"
from "the extra was never installed", which is all the caller needs to
decide whether pip has work to do. Anything unreadable returns False, so
the caller still falls through to pip.
"""
gated = _extra_dependencies(req.name, req.extras)
if gated is None:
return False
for dep in gated:
try:
dep_version = importlib.metadata.version(dep.name)
except importlib.metadata.PackageNotFoundError:
return False
if dep.specifier and not dep.specifier.contains(dep_version, prereleases=True):
return False
return True
def requirements_are_satisfied(requirements_file: str) -> bool: def requirements_are_satisfied(requirements_file: str) -> bool:
""" """
Check whether every real requirement line in requirements.txt is already Check whether every real requirement line in requirements.txt is already
@@ -76,9 +128,6 @@ def requirements_are_satisfied(requirements_file: str) -> bool:
except InvalidRequirement: except InvalidRequirement:
return False return False
if req.extras:
return False # verifying extras' sub-dependencies isn't worth it here
if req.marker is not None and not req.marker.evaluate(): if req.marker is not None and not req.marker.evaluate():
continue # not applicable on this platform/interpreter continue # not applicable on this platform/interpreter
@@ -90,6 +139,9 @@ def requirements_are_satisfied(requirements_file: str) -> bool:
if req.specifier and not req.specifier.contains(installed_version, prereleases=True): if req.specifier and not req.specifier.contains(installed_version, prereleases=True):
return False return False
if req.extras and not _extras_are_satisfied(req):
return False
return True return True
+41 -18
View File
@@ -125,6 +125,14 @@ class PluginManager:
self._plugin_locks: Dict[str, threading.Lock] = {} self._plugin_locks: Dict[str, threading.Lock] = {}
self._plugin_locks_guard = threading.Lock() self._plugin_locks_guard = threading.Lock()
self._update_worker: Optional[threading.Thread] = None self._update_worker: Optional[threading.Thread] = None
# Plugin ids whose update() has finished since the last time anyone
# asked. Updates are dispatched to a worker thread, so a caller that
# wants to know "whose data just changed" cannot learn it by diffing
# plugin_last_update around run_scheduled_updates() -- that call only
# enqueues, and the timestamp is stamped later, on the worker. See
# run_scheduled_updates_with_changes().
self._completed_updates: set = set()
self._completed_updates_lock = threading.Lock()
self._synchronous_updates = False self._synchronous_updates = False
if self.config_manager is not None: if self.config_manager is not None:
try: try:
@@ -1025,6 +1033,7 @@ class PluginManager:
if success: if success:
with self._plugin_last_update_lock: with self._plugin_last_update_lock:
self.plugin_last_update[plugin_id] = scheduled_time self.plugin_last_update[plugin_id] = scheduled_time
self._note_update_completed(plugin_id)
self.state_manager.record_update(plugin_id) self.state_manager.record_update(plugin_id)
self.state_manager.set_state(plugin_id, PluginState.ENABLED) self.state_manager.set_state(plugin_id, PluginState.ENABLED)
if self.health_tracker: if self.health_tracker:
@@ -1089,28 +1098,41 @@ class PluginManager:
def run_scheduled_updates_with_changes(self, current_time: Optional[float] = None) -> List[str]: def run_scheduled_updates_with_changes(self, current_time: Optional[float] = None) -> List[str]:
""" """
Like run_scheduled_updates(), but also returns the plugin_ids whose Like run_scheduled_updates(), but also reports which plugins have
plugin_last_update timestamp actually advanced during this call. fresh data -- the ids whose update() has finished since the last
call, not necessarily the ones enqueued by this one.
The before/after snapshots and the update pass itself are each That distinction is the whole point. This used to snapshot
individually lock-protected against concurrent plugin_last_update plugin_last_update, call run_scheduled_updates(), and diff. But
mutation (Vegas mode calls this from its own background run_scheduled_updates() only *enqueues*: the work runs on the
update-tick thread, racing the main render loop's plugin updates), update worker and the timestamp is stamped there, after this method
so callers get an atomic "who got fresh data" answer without has already returned. The two snapshots were therefore always
reaching into plugin_last_update themselves. The lock is not held identical and the result was always empty, so Vegas never learned
across the update pass so slow/blocking plugin update() calls don't that any plugin's data had changed and kept scrolling whatever a
serialize against other plugin_last_update readers. segment was first built from -- last night's live game still drawn
as live the next morning. The only path that ever worked was the
synchronous kill-switch, where update() runs inline.
Reporting completions instead of enqueues costs a poll's worth of
latency (the Vegas tick runs every ~4s) and is correct regardless of
which side of the queue the work lands on.
""" """
with self._plugin_last_update_lock:
old_times = dict(self.plugin_last_update)
self.run_scheduled_updates(current_time) self.run_scheduled_updates(current_time)
return self.drain_completed_updates()
with self._plugin_last_update_lock: def _note_update_completed(self, plugin_id: str) -> None:
return [ """Record that a plugin's update() finished, for the next poll."""
plugin_id for plugin_id, new_time in self.plugin_last_update.items() with self._completed_updates_lock:
if new_time > old_times.get(plugin_id, 0.0) self._completed_updates.add(plugin_id)
]
def drain_completed_updates(self) -> List[str]:
"""Return and clear the plugin ids whose update() has since finished."""
with self._completed_updates_lock:
if not self._completed_updates:
return []
done = sorted(self._completed_updates)
self._completed_updates.clear()
return done
def update_all_plugins(self) -> None: def update_all_plugins(self) -> None:
""" """
@@ -1135,6 +1157,7 @@ class PluginManager:
if success: if success:
with self._plugin_last_update_lock: with self._plugin_last_update_lock:
self.plugin_last_update[plugin_id] = time.time() self.plugin_last_update[plugin_id] = time.time()
self._note_update_completed(plugin_id)
self.state_manager.record_update(plugin_id) self.state_manager.record_update(plugin_id)
self.state_manager.set_state(plugin_id, PluginState.ENABLED) self.state_manager.set_state(plugin_id, PluginState.ENABLED)
else: else:
+2 -2
View File
@@ -10,8 +10,8 @@ WorkingDirectory=__PROJECT_ROOT_DIR__
ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/scripts/utils/wifi_monitor_daemon.py --interval 30 ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/scripts/utils/wifi_monitor_daemon.py --interval 30
Restart=on-failure Restart=on-failure
RestartSec=10 RestartSec=10
StandardOutput=syslog StandardOutput=journal
StandardError=syslog StandardError=journal
SyslogIdentifier=ledmatrix-wifi-monitor SyslogIdentifier=ledmatrix-wifi-monitor
[Install] [Install]
+18 -1
View File
@@ -9,8 +9,25 @@ User=root
WorkingDirectory=__PROJECT_ROOT_DIR__ WorkingDirectory=__PROJECT_ROOT_DIR__
Environment=PYTHONDONTWRITEBYTECODE=1 Environment=PYTHONDONTWRITEBYTECODE=1
ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/run.py ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/run.py
Restart=on-failure # Restart=always, not on-failure: run.py exiting 0 (a clean shutdown path taken
# for a reason that no longer applies, e.g. a config reload) would otherwise leave
# the service stopped and the panel dark indefinitely, with systemd considering
# that a successful outcome and never bringing it back.
Restart=always
RestartSec=10 RestartSec=10
# Memory ceiling as a share of physical RAM, so one unit file suits a 512 MB
# Pi Zero 2 W and an 8 GB Pi 5 alike. This is a backstop, not a tuning knob: it
# turns "the board runs out of memory, stops being able to fork, and takes sshd
# and the panel down together until someone pulls the plug" into "this one
# service restarts".
#
# NOTE: Raspberry Pi firmware boots the kernel with cgroup_disable=memory, and
# systemd accepts this setting and then silently ignores it. Verify with:
# grep memory /sys/fs/cgroup/cgroup.controllers
# If that prints nothing, add "cgroup_enable=memory cgroup_memory=1" to
# /boot/firmware/cmdline.txt (all on line 1) and reboot. first_time_install.sh
# does this for you.
MemoryMax=85%
StandardOutput=journal StandardOutput=journal
StandardError=journal StandardError=journal
SyslogIdentifier=ledmatrix SyslogIdentifier=ledmatrix
+180
View File
@@ -0,0 +1,180 @@
"""
Tests for scripts/download_pixlet.sh -- release-tag resolution and download guards.
Background: Starlark apps render through the pixlet binary, and the installer
that fetches it failed silently. It resolved the release tag by grepping the
GitHub API response for '"tag_name"' and taking the last quoted token on the
match with a greedy sed. When the response arrives on one line that token is
"mentions_count", not the tag, so the script built a URL for a release that
cannot exist -- and `curl -L -o` without -f wrote the 404 body to the file and
exited 0, so the first sign of trouble was tar reporting "not in gzip format"
about a page of HTML.
The API is pretty-printed by default, which is exactly why this needs a test:
by hand the old command looks correct, and the failure only appears when the
formatting changes. These drive the real script with a stubbed curl on PATH, so
both response shapes are covered without touching the network.
"""
import re
import shutil
import subprocess
from pathlib import Path
import pytest
SCRIPT = Path(__file__).resolve().parent.parent / "scripts" / "download_pixlet.sh"
PRETTY = """{
"url": "https://api.github.com/repos/tronbyt/pixlet/releases/12345",
"id": 12345,
"tag_name": "v0.53.1",
"name": "v0.53.1",
"draft": false,
"prerelease": false,
"mentions_count": 3
}
"""
# The shape that broke it: one line, and the last quoted token is not the tag.
MINIFIED = (
'{"url":"https://api.github.com/repos/tronbyt/pixlet/releases/12345",'
'"id":12345,"tag_name":"v0.53.1","name":"v0.53.1","draft":false,'
'"prerelease":false,"mentions_count":3}'
)
def run_script(tmp_path, api_body, download=None):
"""Run the real script against a stubbed curl.
Args:
api_body: what the stub returns for the api.github.com request.
download: bytes to write for a release-asset request, or None to make
that request fail the way `curl -f` does on an HTTP error.
"""
root = tmp_path / "project"
(root / "scripts").mkdir(parents=True)
shutil.copy(SCRIPT, root / "scripts" / "download_pixlet.sh")
api_file = tmp_path / "api.json"
api_file.write_text(api_body)
stub_dir = tmp_path / "stub"
stub_dir.mkdir()
asset_file = tmp_path / "asset.bin"
if download is not None:
asset_file.write_bytes(download)
# Stands in for curl, including the -f semantics the fix turns on: without
# -f, real curl writes the error body to the output file and exits 0, which
# is what let a 404 masquerade as a successful download. The stub has to
# honour that or a test of the fix would pass against the old script too.
(stub_dir / "curl").write_text(f"""#!/bin/bash
out=""
url=""
fail_on_error=0
while [ $# -gt 0 ]; do
case "$1" in
-o) out="$2"; shift 2 ;;
-*f*) fail_on_error=1; shift ;;
-*) shift ;;
*) url="$1"; shift ;;
esac
done
if [[ "$url" == *api.github.com* ]]; then
cat {api_file}
exit 0
fi
if [ -f "{asset_file}" ]; then
cp "{asset_file}" "$out"
exit 0
fi
# No asset: stand in for an HTTP 404.
if [ "$fail_on_error" = "1" ]; then
exit 22
fi
printf '<!DOCTYPE html><html>404 Not Found</html>' > "$out"
exit 0
""")
(stub_dir / "curl").chmod(0o755)
return subprocess.run(
["bash", str(root / "scripts" / "download_pixlet.sh")],
capture_output=True, text=True,
env={"PATH": f"{stub_dir}:/usr/bin:/bin:/usr/sbin:/sbin",
"PIXLET_VERSION": "latest"},
)
def resolved_version(result):
match = re.search(r"^Version: (.+)$", result.stdout, re.M)
assert match, f"no version line in output:\n{result.stdout}"
return match.group(1).strip()
def test_script_is_syntactically_valid():
result = subprocess.run(["bash", "-n", str(SCRIPT)], capture_output=True, text=True)
assert result.returncode == 0, result.stderr
@pytest.mark.parametrize("body,label", [(PRETTY, "pretty"), (MINIFIED, "minified")])
def test_tag_is_resolved_from_either_response_shape(tmp_path, body, label):
"""The minified case is the regression: the last quoted token there is
"mentions_count", which is what the old greedy sed captured."""
result = run_script(tmp_path, body)
assert resolved_version(result) == "v0.53.1", f"{label}: {result.stdout}"
assert "mentions_count" not in result.stdout
@pytest.mark.parametrize(
"tag",
["mentions_count", "v0.53garbage", "0.53", "v0.5", "", "v0.53.1 ; echo pwned"],
)
def test_a_tag_that_is_not_a_release_falls_back(tmp_path, tag):
"""A wrong-but-non-empty value is what made the original bug silent, so the
check is on the shape. Partial matches must not pass: "v0.53garbage" and
"0.53" would build a URL for a release that cannot exist."""
result = run_script(tmp_path, '{"tag_name": "%s"}' % tag)
assert resolved_version(result) == "v0.50.2", result.stdout
assert "using fallback" in result.stdout
@pytest.mark.parametrize("tag", ["v0.53.1", "v1.0.0", "v0.54.0-rc.1", "v1.2.3+build.4"])
def test_real_release_tag_shapes_are_accepted(tmp_path, tag):
assert resolved_version(run_script(tmp_path, '{"tag_name": "%s"}' % tag)) == tag
def test_an_http_error_is_reported_as_a_failed_download(tmp_path):
"""Without curl -f the 404 body lands in the file and curl exits 0, so the
failure surfaced two steps later as tar complaining about gzip -- about
what was really a page of HTML. It has to be reported where it happened.
Both versions end at 0/1, so asserting only on the count would pass against
the old script; the discriminating part is which layer reports it.
"""
result = run_script(tmp_path, PRETTY, download=None)
assert "Download complete: 0/1 succeeded" in result.stdout
assert "✓ Downloaded" not in result.stdout
assert "Failed to download" in result.stdout
assert "Failed to extract" not in result.stdout, (
"an HTTP error should not surface as an extraction failure")
def test_a_non_archive_response_is_rejected_before_extraction(tmp_path):
result = run_script(tmp_path, PRETTY, download=b"<!DOCTYPE html><html>502 Bad Gateway")
assert "not a gzip archive" in result.stdout
assert "Download complete: 0/1 succeeded" in result.stdout
def test_the_diagnostic_cannot_smuggle_terminal_escapes(tmp_path):
"""Those bytes come from whatever answered the request. An error page
carrying escapes must not be able to rewrite the output or bury it."""
hostile = b"<!DOCTYPE html>\x1b[2J\x1b[31mgone\x1b[0m\rHTTP 200 OK\x08\x08"
result = run_script(tmp_path, PRETTY, download=hostile)
assert "not a gzip archive" in result.stdout
printed = re.search(r"^\s*\(first bytes: (.*)\)$", result.stdout, re.M)
assert printed, f"no diagnostic line:\n{result.stdout}"
assert "DOCTYPE" in printed.group(1), "the useful part of the page was dropped"
for forbidden in ("\x1b", "\r", "\x08", "\x00"):
assert forbidden not in printed.group(1), (
f"control byte {forbidden!r} reached the terminal")
+182
View File
@@ -0,0 +1,182 @@
#!/usr/bin/env python3
"""
Tests that "which plugins have fresh data" survives the async update worker.
Regression under test: run_scheduled_updates_with_changes() snapshotted
plugin_last_update, called run_scheduled_updates(), and diffed the two. But
run_scheduled_updates() only *enqueues* -- the work runs on the update worker
and stamps the timestamp there, after the method has already returned. The
snapshots were therefore always identical and the result always empty.
Vegas depends on that result: it is what calls mark_plugin_updated(), which
drops the cached content for a plugin whose data changed. With it always
empty, a segment kept scrolling whatever it was first built from -- the
"last night's live game still drawn as live the next morning" failure the
coordinator comments describe. Observed on a live rig: zero update ticks in
twenty minutes, with weather, stocks and news all updating.
Run: python -m pytest test/test_update_change_reporting.py -v
"""
import ast
import inspect
import sys
import threading
import unittest
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from src.plugin_system.plugin_manager import PluginManager # noqa: E402
def _manager():
"""A PluginManager with only the update-reporting state initialised."""
manager = PluginManager.__new__(PluginManager)
manager._completed_updates = set()
manager._completed_updates_lock = threading.Lock()
return manager
class DrainCompletedUpdates(unittest.TestCase):
def setUp(self):
self.manager = _manager()
def test_nothing_completed_reports_nothing(self):
self.assertEqual(self.manager.drain_completed_updates(), [])
def test_a_completed_update_is_reported(self):
self.manager._note_update_completed("news")
self.assertEqual(self.manager.drain_completed_updates(), ["news"])
def test_draining_clears_so_the_next_poll_is_empty(self):
self.manager._note_update_completed("news")
self.manager.drain_completed_updates()
self.assertEqual(
self.manager.drain_completed_updates(), [],
"a plugin must be reported once per update, not on every poll, "
"or Vegas would drop its cached content every few seconds")
def test_repeated_completions_between_polls_collapse(self):
for _ in range(5):
self.manager._note_update_completed("weather")
self.assertEqual(self.manager.drain_completed_updates(), ["weather"])
def test_multiple_plugins_are_all_reported(self):
for plugin_id in ("news", "weather", "ledmatrix-stocks"):
self.manager._note_update_completed(plugin_id)
self.assertEqual(self.manager.drain_completed_updates(),
["ledmatrix-stocks", "news", "weather"])
class CompletionReportingIsAsyncSafe(unittest.TestCase):
"""The point of the change: completion may land after the call returns."""
def setUp(self):
self.manager = _manager()
def test_an_update_completing_after_the_call_is_still_reported(self):
"""The exact shape of the bug.
The enqueueing call sees nothing, because the worker has not run yet.
The next poll must report it -- under the old diff it was lost, since
the second snapshot was taken before the worker ever stamped.
"""
first = self.manager.drain_completed_updates()
self.assertEqual(first, [], "nothing has finished yet")
# The worker finishes some time later, on its own thread.
worker = threading.Thread(
target=self.manager._note_update_completed, args=("news",))
worker.start()
worker.join()
self.assertEqual(
self.manager.drain_completed_updates(), ["news"],
"an update that finishes between polls must still be reported")
def test_concurrent_completions_are_not_lost(self):
ids = ["plugin-%02d" % i for i in range(40)]
threads = [threading.Thread(target=self.manager._note_update_completed,
args=(pid,)) for pid in ids]
for thread in threads:
thread.start()
for thread in threads:
thread.join()
self.assertEqual(self.manager.drain_completed_updates(), sorted(ids))
def test_a_completion_during_a_drain_is_not_swallowed(self):
"""A drain must not clear an entry it did not report."""
self.manager._note_update_completed("news")
reported = self.manager.drain_completed_updates()
# ...worker finishes another one immediately afterwards
self.manager._note_update_completed("weather")
self.assertEqual(reported, ["news"])
self.assertEqual(self.manager.drain_completed_updates(), ["weather"])
class EveryStampRecordsACompletion(unittest.TestCase):
"""The ledger is only correct if the production paths actually fill it.
Asserting on the mechanics alone passes even when nothing calls
_note_update_completed -- verified by deleting the call sites, which the
behavioural tests above did not notice. This checks the invariant at the
source: wherever a successful update stamps plugin_last_update, it must
also record the completion, or Vegas silently stops being told.
"""
def test_success_paths_record_the_completion(self):
import src.plugin_system.plugin_manager as pm
tree = ast.parse(inspect.getsource(pm))
stamps = []
for node in ast.walk(tree):
if not isinstance(node, ast.With):
continue
# `with self._plugin_last_update_lock:` blocks that stamp a real
# time on success. Two stamps are deliberately excluded: the 0.0
# written at registration, and the failure path, which backs the
# timestamp off to space out retries -- neither means fresh data.
assigns_time = any(
isinstance(stmt, ast.Assign)
and any(isinstance(t, ast.Subscript)
and getattr(t.value, "attr", None) == "plugin_last_update"
for t in stmt.targets)
and not (isinstance(stmt.value, ast.Constant)
and stmt.value.value == 0.0)
and "failure" not in ast.dump(stmt.value)
for stmt in node.body
)
if assigns_time:
stamps.append(node)
self.assertGreaterEqual(
len(stamps), 2,
"expected the worker and inline success paths to stamp the time; "
"if this drops, the search below is looking at the wrong thing")
for stamp in stamps:
enclosing = self._enclosing_function(tree, stamp)
calls = [n for n in ast.walk(enclosing)
if isinstance(n, ast.Call)
and getattr(n.func, "attr", None) == "_note_update_completed"]
self.assertTrue(
calls,
"%s stamps plugin_last_update on success but never calls "
"_note_update_completed, so a plugin's fresh data would never "
"be reported and Vegas would keep its stale cached content"
% enclosing.name)
@staticmethod
def _enclosing_function(tree, target):
best = None
for node in ast.walk(tree):
if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)):
if node.lineno <= target.lineno <= (node.end_lineno or node.lineno):
if best is None or node.lineno > best.lineno:
best = node
return best
if __name__ == "__main__":
unittest.main(verbosity=2)