Commit Graph
84 Commits
Author SHA1 Message Date
ChuckandClaude Opus 5.5 64c7289593 feat(display): systemd watchdog and heartbeat for a frozen render loop (#687)
If the render loop gets stuck inside a plugin's display(), ledmatrix.service
stays active and the panel stays frozen. This adds a way to detect that.

- src/display_watchdog.py (standard library only) sends sd_notify over
  $NOTIFY_SOCKET and writes /run/ledmatrix/display-heartbeat.json. Only the
  render thread counts: beats from other threads are ignored.
- ledmatrix.service: WatchdogSec=120, NotifyAccess=main,
  RuntimeDirectory=ledmatrix (0755), RestartSteps=4 and
  RestartMaxDelaySec=2min. It stays Type=simple. run.py widens the watchdog
  to 15 min for start-up, and load_plugin() does the same on the render
  thread. The loop arms after its first frame.
- /api/v3/health adds checks.display_loop: running, stalled (no heartbeat
  for over 60s, which makes the status degraded) or not_reported. With web
  login on, a caller who is not logged in still gets only healthy/degraded,
  and a stall degrades that answer.
- The update verifier requires a fresh heartbeat from the restarted display
  when the display it replaced was writing one. A frozen panel is rolled
  back.
- Existing installs get the systemd watchdog only after install_service.sh
  is re-run. The heartbeat works right away.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 11:15:31 -04:00
ChuckandClaude Opus 5.5 ba6eccb489 build(web): generate the UI's Tailwind CSS with the pinned standalone CLI (#685)
Replaces the hand-written Tailwind subset in app.css with a real, purged
Tailwind build: scripts/build_css.py runs the pinned, SHA-256-checked
standalone Tailwind CLI (no Node), the generated tailwind.css and
plugin-frame.css are committed, and CI fails when they are stale. The Pi
never builds anything. The login page (#683) now links tailwind.css too,
and the load-order test covers every template that links app.css.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:38:16 -04:00
ChuckandClaude Opus 5.5 c8a0ddcf7b chore(deprecation): retarget the 35 deprecations to 3.8.0 and add a usage scan (#681)
Moves the 35 @deprecated markers from 3.7.0 (already shipped with them in
place) to 3.8.0, and adds scripts/plugin_api_usage.py plus the generated
docs/DEPRECATIONS_3.8.md: who still calls or overrides each deprecated
method across core, the monorepo and third-party plugins. Removes nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:11:36 -04:00
ChuckandClaude Opus 5.5 9fe23af432 docs(sports): reconcile-then-promote roadmap, drift report and report-only CI job (#680)
Rewrites the roadmap in docs/SPORTS_UNIFICATION.md for the
reconcile-then-promote decision (stages 0-3 recorded as done), and adds the
sports drift report script with a report-only CI job.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:11:02 -04:00
ChuckandClaude Opus 5.5 e3c85cece6 feat(web): optional web login and API tokens, off by default (stacked on #674) (#683)
Optional web login, off by default: a device that sets no password behaves
exactly as before. Set under General > Security; then every page and API
route needs a session login or an API token (Authorization: Bearer).
Loopback, the Wi-Fi setup flow in AP mode, static files, captive-portal
probes and a reduced /api/v3/health stay open. Secrets live in the web_auth
section of config_secrets.json and no API returns them.
scripts/reset_web_password.py turns login off. Stacked on #674.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:09:40 -04:00
ChuckandClaude Opus 5.5 b8c01c69fb ci: mypy ratchet -- keep type-clean modules clean (71 modules, 536 -> 442 errors) (#661)
* ci: mypy ratchet -- keep type-clean modules clean

mypy-clean.txt lists the 71 modules under src/ that type-check clean;
scripts/check_types.py runs mypy (--follow-imports=silent) on exactly
those files and fails on any error or a missing/unsorted/duplicate entry.
A new "Type check (mypy ratchet)" CI job runs it with mypy 1.20.2 and
pinned stubs; the manual pre-commit mypy hook now runs the same script
(a local hook, so mypy sees the installed requirements like CI does).

35 modules were made clean with annotation-only fixes: hints, typing.cast,
TYPE_CHECKING imports, implicit-Optional defaults made explicit, and
annotations widened (never guards removed) where mypy called a defensive
isinstance check unreachable. No runtime behaviour change.

mypy.ini: numpy and orjson are treated as Any (follow_imports=skip, also
for stubs). numpy 2.3+ stubs use 3.12 `type` statements that mypy won't
parse at python_version 3.10, and orjson is optional, so seeing its stubs
made the result depend on whether it was installed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: annotate check_types.py's list-form mypy subprocess

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 15:01:35 -04:00
ChuckandClaude Opus 5.5 6cfcf2e384 fix(web): plugin dir resolver in routes, nmcli AP detection, daemon config reload, upload safety, BDF preview (#655)
* fix(web): plugin dir resolver in routes, nmcli AP detection, daemon config reload, upload safety

- Route plugin lookups (installed list, update, recorded version, config
  form, web UI pages) through the plugin manager's resolver so plugins in
  ledmatrix-<id> directories work.
- Captive-portal detection also sees the nmcli fallback AP (cached).
- WiFi monitor daemon re-reads wifi_config.json when its mtime changes.
- Drop the AP check in disconnect_from_network that could never fire.
- LED status file per WiFiManager; config path falls back to this checkout.
- BDF font preview via src.common.bdf_font.
- Asset uploads validate every file before saving; metadata and calendar
  credentials written atomically; no absolute path in the response;
  asset delete answers 400 for a missing body.
- Coerce string booleans in plugin toggle, on-demand start and AP force.
- SSE broadcaster clears its thread handle before exiting.
- start.py log filter handles every exc_info form.
- Cleanups: unused plugins/fonts partial work, duplicate backup catch-alls,
  raw-config error helper, update-route tidy, redundant imports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): request BDF font previews now that the server renders them

The Fonts tab skipped the preview request for .bdf files because the server
used to refuse them; /fonts/preview now draws BDF with the shared loader.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): take the update route's plugin directory from a directory listing

CodeQL flagged the path built from the request's plugin_id (the id was
already validated with safe_path_component, which CodeQL doesn't model; the
same flow on main is alerts 738/739). The directory is now the entry of
plugins_dir matched by name, so nothing built from user input reaches the
filesystem; an id with nothing installed goes to the store manager, which
reports it not found as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): read the blueprint's plugin_manager defensively in _plugin_directory

_get_plugin_version now goes through _plugin_directory, which read
api_v3.plugin_manager directly; the attribute exists only once the app sets
it, so test_path_traversal_guards::test_a_real_manifest_is_read failed
when run on its own (order-dependent in the full suite).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:42:07 -04:00
ChuckandClaude Opus 5.5 aeaeaa4e94 chore: make contributor tooling work; fix drifted docs (#648)
- mypy.ini parses again (multi-line exclude and inline value comments made
  mypy reject the file); the mypy pre-commit hook is manual-only until the
  ~500 existing errors in src/ are paid down, and CONTRIBUTING says so
- .gitignore: ignore all of config/ except the templates (ytm_auth.json and
  others weren't ignored)
- .gitattributes: LF for .sh and .service
- claude-code-review: skip fork PRs, which have no secrets
- check_system_compatibility.sh: 3.13 supported, <3.10 an error
- docs/scripts drift: emulator guide, README API Metrics, route count,
  docs index, scripts README; pyflakes nits in dev scripts

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:26:55 -04:00
ChuckandClaude Opus 5.5 da5937da3d fix: six bugs found testing main on a real Pi (#641)
* fix: six bugs found testing main on a real Pi (ledpi)

- Stopping the service now runs cleanup. systemd stops ledmatrix.service
  with SIGTERM, whose default action ended Python before run()'s finally
  block, so the update worker, Vegas and the panel were never torn down.
  main() now turns SIGTERM into KeyboardInterrupt, the Ctrl-C path.
- "Now showing" no longer turns into "unknown". display_current_state was
  only written on a mode change and the web UI reads it with max_age=120,
  so a live game or a single plugin on screen for longer read as unknown.
  It is republished every 30 s while unchanged.
- Switching Vegas on in the web UI works when it was off at startup. The
  coordinator was only created at startup; the config watcher now flags it
  and the render thread creates it.
- configure_web_sudo.sh finds reboot and poweroff in /usr/sbin. Run as the
  web user it could not, silently dropped their rules and still said it
  granted them, so the web UI's Reboot/Shutdown stopped working.
- check_system_compatibility.sh reports installed packages as installed.
  `dpkg -l | grep -q` under pipefail failed when grep exited early.
- A network failure fetching GitHub repo info logs a WARNING, not ERROR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): fixes found testing on a Pi

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: run the Linux-only script tests correctly

The sbin-lookup test set PATH=/nonexistent and then could not find bash
itself; call it by absolute path. The dpkg-query stub read $4, but the
package name is the third (last) argument. Both now pass on a Pi.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): cover the follower, long-render and startup cases

Review follow-ups on the ledpi fixes:
- The pending Vegas start is applied before the sync-follower branch too
  (_apply_pending_vegas_init), which skips _is_vegas_mode_active() while a
  follower is connected but needs the coordinator for the leader's image.
- _service_pending_changes(), which runs inside Vegas iterations and long
  screens, republishes a stale display_current_state as well; the main
  loop alone could be away for a 240 s Vegas iteration.
- The SIGTERM handler is installed after DisplayController() is built, so a
  stop during parallel plugin loading keeps the default immediate exit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 18:13:49 -04:00
ChuckandClaude Opus 5.5 8ad9d191a7 feat(perf): frame timing for every presented frame, a soak tool, a render bench and a stall watchdog (#629)
src/common/frame_timing.py times every frame the display presents, whoever drew it, and writes cumulative counters to /dev/shm. scripts/frame_soak.py grades a running service (late frames, freezes, where the time goes) and scripts/render_bench.py the hardware and render path alone. A stall watchdog logs the stacks behind any scroll held up for 250 ms or more (LEDMATRIX_STALL_WATCHDOG_MS lowers that). See docs/SCROLL_PERFORMANCE.md, "Soaking a rig".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 19:38:49 -04:00
ChuckandClaude Opus 5.5 3967a6cffc fix(security): re-harden root sudo helpers; installer fixes; ARCHITECTURE and PERMISSIONS docs (#640)
* docs: add ARCHITECTURE and PERMISSIONS guides

ARCHITECTURE.md maps the processes, the state the display and web
services share through the cache, the display loop, the plugin system,
the web UI and the update path, with links into the code and a
where-to-start table.

PERMISSIONS.md lists who owns what after install, both sudoers files
(and why iptables is not granted), the polkit rule, and which
scripts/fix_perms script to run as which user.

Both are linked from the docs index, along with the MQTT bridge README
and src/common/README.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: correct stale setup, service and troubleshooting claims

- README: quick actions run systemctl on ledmatrix.service (run.py), not
  display_controller.py; use_short_date_format has no effect; the
  installer uses system pip with --break-system-packages, not a venv.
- CONFIG_DEBUGGING: LEDMATRIX_DEBUG must be "true"; logs are in journald.
- GETTING_STARTED, WEB_INTERFACE_GUIDE, TROUBLESHOOTING: enabling a
  plugin, plugin settings, brightness and Vegas settings apply without a
  restart; matrix hardware settings still need one.
- TROUBLESHOOTING: install dependencies with sudo so the root service
  sees them; point permission problems at PERMISSIONS.md instead of a
  project-wide chown.
- ADVANCED_FEATURES: real BackgroundDataService stats keys; Vegas hooks
  return VegasDisplayMode and None falls back to capture; cache files
  are 0660; fix_web_permissions.sh runs as the web user and does not
  touch sudoers.
- STARLARK_APPS_GUIDE: only the linux-arm64 pixlet binary is downloaded.
- HOW_TO_RUN_TESTS: test class examples that exist.
- CLAUDE.md: PluginStoreManager, plugin_dirs.py, monorepo installs via
  the Trees API with ZIP fallback, requirements.txt is optional.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: mark deprecated plugin APIs and state manifest fields once

Methods @deprecated("3.7.0") (the set pinned in test_deprecation.py)
were shown as current API in the quick reference, API reference,
advanced guide, development guide and FONT_MANAGER. Each is now marked
deprecated with its replacement. FONT_MANAGER is rewritten around the
current API; the override editor is gone and override methods are
deprecated.

Required manifest fields were stated three different ways. The API
reference now has one section: the 7 schema-required fields, the 4 the
store refuses without, class_name for the loader, and the 8 to set.
The other guides link to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document every src/common module and every widget

- src/common/README.md covered 7 of 17 modules. It now has a table of
  all of them (purpose, whether plugins import it, release to floor
  on), a short entry each, and logging advice that matches the code.
- SPORTS_UNIFICATION listed two shared modules and called
  sports_helpers the first; it now lists all six.
- The widgets README lists all 28 registered widgets plus the support
  files, and absorbs the parts that only docs/widget-guide.md had
  (x-options.labels, x-advanced, x-display hidden, plugin-file-manager).
  docs/widget-guide.md is now a pointer to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(security): fix_web_permissions.sh re-hardens the root sudo helpers

The script chowns the whole project to the web user. That included
scripts/fix_perms/safe_plugin_rm.sh and safe_pip_install.sh -- the two
helpers /etc/sudoers.d/ledmatrix_web lets the web user run as root -- so
running it turned both into a root shell for whoever can edit them. It
also re-grouped config_secrets.json away from ledmatrix.

After the chown it now does what first_time_install.sh's Steps 11 and
11.1 do: helpers back to root:root 755, and config_secrets.json back to
the web unit's User=:ledmatrix 640. Each step is non-fatal and prints the
manual command if it fails.

Also fixes what the script and its docs claimed: it never configured
sudoers, its closing hint pointed at ./configure_web_sudo.sh (wrong
path), and the README and ADVANCED_FEATURES.md said to run it with sudo,
which it refuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(security): validate and harden every sudoers drop-in the scripts write

configure_wifi_permissions.sh copied its rules into
/etc/sudoers.d/ledmatrix_wifi without `visudo -c`. A malformed drop-in
makes sudo refuse every command for every user, which on a headless Pi
leaves no way back in. It now checks first and leaves the installed file
alone when the rules do not parse, as the other two writers do. (It
already used mktemp, so that part of the review did not apply.)

It also grants the two literal commands wifi_manager.py runs for
NetworkManager's shared-mode dnsmasq drop-in -- `cp
/tmp/ledmatrix-nm-dnsmasq.conf .../dnsmasq-shared.d/ledmatrix-captive.conf`
and `rm -f` of that file. The directory's mkdir was granted, the file was
not. Both are pinned in test_sudo_allowlist_covers_calls.py.

configure_web_sudo.sh wrote its rules to /tmp/ledmatrix_web_sudoers_$$,
a predictable name in a world-writable directory; it now uses mktemp with
an EXIT trap, as first_time_install.sh does. It sets mode 440 on the
installed file instead of leaving the temp file's mode, and finds visudo
in /usr/sbin when that is not on the user's PATH, which skipped the
check silently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): escape the project path in the DNS-fix and MQTT unit renderers

install_dns_fix.sh and install_mqtt_bridge.sh substituted
__PROJECT_ROOT_DIR__ with the raw path, while the other three renderers
go through sed_escape_replacement from lib_systemd_render.sh. A checkout
under a path containing `&`, `\` or `|` rendered a corrupted unit from
these two only. Both now source the helper and use it, and a test checks
that every placeholder substitution in scripts/install uses an escaped
value.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): stop the installer scripts reporting things that are not true

- first_time_install.sh printed "Password: ledmatrix123" for the setup
  access point. wifi_manager creates it as an open network ("No
  password" on the panel), so it now says so.
- Step 10.1 printed "✓ WiFi management permissions configured" straight
  after its own failure message; install_wifi_monitor.sh printed
  "✓ Package installation completed" after a failed apt install. The
  tick now only follows success.
- Step 7 printed "Web dependencies already installed ... in Step 5" in
  the one branch that runs because Step 5 did not install them, then
  created .web_deps_installed on that basis. It now warns and leaves the
  marker off so the next run retries, as the comment below it intends.
- check_system_compatibility.sh called Debian 12 Bookworm "full
  compatibility confirmed" while first_time_install.sh refuses anything
  but Debian 13. Bookworm, older Debian and non-Debian systems are now
  errors. Its counters used ((X++)), which under `set -e` exits the
  script at the first warning or error (the expression is 0), so the
  check never reached its summary on any system with one.
- configure_web_sudo.sh and configure_wifi_permissions.sh finished by
  testing `sudo -n test -f ...` and `sudo -n nmcli device status`,
  neither of which is granted, so they always reported a failure. They
  now ask `sudo -n -l` about commands the new rules do grant, which
  checks the rule without running anything.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): print the completion summary before rebooting

With -y -- and so for every one-shot `curl | bash` install, which always
passes -y -- first_time_install.sh ran `reboot` about 180 lines before
its "Installation Complete / Web UI Access" summary. reboot returns at
once, so the summary printed while the Pi was going down and the SSH
session usually dropped before the web UI address could be read.

The reboot block moves, unchanged, to the very end of the script. The
interactive prompt now also follows the summary. Because the summary now
runs before the -y reboot, its one command that could fail under
`set -Eeuo pipefail` (the SSID lookup, when nmcli reports a connected
device but no active network line) gets `|| true`; a missing SSID was
already handled as "SSID unknown".

one-shot-install.sh prints its "Next steps" after the installer returns,
by which time the reboot is under way, so it now says so, and README's
Quick Install mentions the automatic reboot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(scripts): correct wrong comments and messages, drop dead code

No behaviour change except the output text noted below.

- 2775 is setgid, not the sticky bit (first_time_install.sh Step 3.1,
  fix_plugin_permissions.sh), and root needs no "PWM hardware access"
  to plugin files.
- The 777 comments in first_time_install.sh Step 3's fallback and
  fix_assets_permissions.sh said root needs it to write. Root ignores
  mode bits; the comments now say what 777 actually opens. The 777
  itself is unchanged.
- apt_remove ends in `|| true`, so Step 12's "Some packages could not be
  removed" branch could never run; it is gone and the helper stays
  non-fatal.
- detect_web_service_user's comment named Step 8 for the web unit
  (install_service.sh installs it in Step 7.5) and now says which
  branch actually runs.
- Step 5 described an "already installed" check that does not exist;
  the ACTUAL_USER comment described the re-exec backwards.
- on_error printed a literal "\n" before "Common fixes:".
- Dead code: one-shot-install.sh's uncalled fix_tmp_permissions,
  LEDMATRIX_ELEVATED=1 (never read) on the sudo re-exec, and
  configure_web_sudo.sh's unused PYTHON_PATH, which also made a missing
  python3 fatal for rules that never mention it.
- start_display.sh / stop_display.sh said "for user: <you>"; the
  service runs as root.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(fix_perms): fix_cache_permissions.sh uses setup_cache.sh's model

There were two models for /var/cache/ledmatrix. setup_cache.sh (the
installer's Step 2) and install_web_service.sh share it through the
ledmatrix group: root:ledmatrix, 2775, files 660, which is also what
DiskCache relies on to give files the directory's group.
fix_cache_permissions.sh instead made it 777 and re-grouped it to the
invoking user's group, undoing that.

It now runs setup_cache.sh for /var/cache/ledmatrix and keeps its own
handling of ~/.ledmatrix_cache. Dropped: /var/cache/ledmatrix/
placeholder_logos (nothing reads it) and the checks against the
`daemon` user (no service runs as daemon).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: pin actions/checkout in the Claude workflows, drop template comments

claude.yml and claude-code-review.yml used actions/checkout@v4 while
test.yml and release-version-check.yml pin the v4.2.2 commit SHA; they
now pin the same SHA. The commented-out starter-template settings
(prompt, claude_args, paths, author filter) are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(scripts): index every script and list removal candidates

New scripts/README.md gives one line per top-level script and scripts
directory, marked keep, dev-only or diagnostic, and lists the eight
scripts nothing in the repo refers to as candidates for removal (kept
for now). The install, utils and dev READMEs now list the files they
were missing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: tighten two checks that mutation testing showed were too loose

- The wifi sudoers check matched `visudo -c -f "$TEMP_SUDOERS"` in the
  error report too, so replacing the check with `if false` still passed.
  It now requires the command as the condition.
- The summary test never had the setup access point up, so reinstating
  the bogus "Password: ledmatrix123" line went unnoticed. A case with
  hostapd active now checks the AP is described as open.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(permissions): describe the repaired fix_perms scripts and new WiFi grants

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): docs-scripts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:31:41 -04:00
ChuckandClaude Opus 5.5 1fe7237799 refactor(install): generate the web sudoers rules in one place (#622)
* refactor(install): generate the web sudoers rules in one place

/etc/sudoers.d/ledmatrix_web was written by two copies of the same
allow-list: a heredoc in first_time_install.sh Step 10 and a block of
echo lines in scripts/install/configure_web_sudo.sh. They drifted before
(safe_pip_install.sh was granted by one only), and a test existed just
to catch that.

Both now call web_sudoers_rules() from the new
scripts/install/lib_sudoers.sh and keep their own validate (visudo -c),
install and confirm flows.

- first_time_install.sh output is byte-for-byte unchanged, so a device
  re-running the installer gets "already up to date". If the library is
  missing, Step 10 keeps the installed file and carries on, the same way
  it handles rules that fail visudo (an empty file would pass visudo).
- configure_web_sudo.sh now writes the installer's layout: same 18 rules,
  different comments and order. It still leaves out reboot, poweroff and
  journalctl when they are missing; the library does that for both.

The drift test now pins the generator's grants, checks that neither
installer writes rules of its own, and runs each installer's call line
to check the argument order. Tests that read the rule text now read the
library.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(install): detect the web service user in one function

first_time_install.sh pasted the same WEB_SERVICE_USER detection block
three times (Step 3.1's fallback, the plugin-repos setup and Step 11).
The copies were identical apart from comments; they now call
detect_web_service_user(), whose body is that block unchanged.

Behaviour is the same: the function sets the same global and always
returns 0, as the inline if-chain did. Checked on Linux against all
three original copies across 13 layouts (installed unit with and without
User=, the repo as shipped, each grep branch, template placeholders).

The comment notes that the install_web_service.sh / install_service.sh
greps no longer match anything, so until Step 8 installs the unit the
result is "root". That behaviour is left as it was.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:51:14 -04:00
ChuckandClaude Opus 5.5 61e462c635 refactor: remove the skin system and the unused src/base_classes package (#615)
* refactor: remove the skin system

Skins never rendered with the current scoreboard plugins: the only hook was
SportsCore._render_game in src/base_classes, which no plugin builds on, so
the UI and store already treated them as unsupported. The owner decided on
2026-09-23 to remove them outright.

Removed src/skin_system/ (runtime, base class, fixtures), skins/,
scripts/validate_skin.py and their tests; the store's "type": "skin"
installer, uninstaller and hide/refuse filters (the official registry lists
no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins.

Stored skin/skin_options config values are handled in the next commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(config): drop retired skin/skin_options keys instead of validating them

A config.json written while the skin system existed can carry skin and
skin_options in any plugin section, and most plugin schemas set
additionalProperties: false. They are no longer core plugin properties;
RETIRED_PLUGIN_KEYS in schema_manager lists them and
drop_retired_plugin_keys removes them (unless the plugin's own schema
declares the name) in prepare_plugin_config, which loading, hot reload,
GET /plugins/config and both web saves already share, and in
validate_config_against_schema for callers that validate a raw section.
POST /plugins/config and /config/main also drop them from the stored
section they merge into, so they leave config.json on the next save.

Tests cover the load path (real PluginManager.load_plugin: no schema
warning, not degraded), raw and prepared validation,
validate_all_plugin_configs, and the JSON, form and /config/main saves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: remove the unused src/base_classes package

No scoreboard plugin builds on src.base_classes: the nine monorepo
scoreboards ship their own sports.py and share code through src/common
(docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins
imports it. The one import anywhere, baseball-scoreboard's
rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing
instantiates.

Removed the package and the eight test files that only tested it
(test_api_extractors, test_data_sources, test_sports_base_characterization,
test_sports_capabilities, test_sports_core_promotions,
test_sports_logo_cache_bounded, test_sports_modes_promotions,
test_sports_odds_fanout). test_common_is_hardware_free no longer lists
src.base_classes as a forbidden import, and comments in sports_helpers.py
and base_odds_manager.py stop pointing at it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: drop the skin system and src/base_classes from the docs

Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to
them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins
section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the
term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was
removed and shared code lives in src/common, in the Layering section and
the view-model-contract rule. Other docs stop pointing at the removed
package. CHANGELOG records both removals under Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(store): hide and refuse registry entries that aren't plugins

The skin filters went with the skin system, but a custom registry can still
list "type": "skin" entries, and installing one as a plugin would unpack it
into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing
type means plugin) now hides non-plugin entries from the store and
custom-registry listings, and install refuses them, in the route with a
clear 400 and in _install_plugin_impl for any other caller.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:33:24 -04:00
ChuckandClaude Opus 5.5 a231d4dbc7 chore: delete unreferenced scripts and archived docs; fix stale doc claims (#607)
* chore(scripts): delete unreferenced helper scripts

None of these is referenced by an installer, systemd unit, CI workflow,
test, the web UI or src/:

- utils/cleanup_venv.sh removes venv_web_v2, which nothing creates
- utils/clear_python_cache.sh hardcodes ~/LEDMatrix and a .webassets-cache
  nothing uses
- install/migrate_config.sh only copies the template, which the installer
  and ConfigManager already do
- install/debug_install.sh, debug/debug_web_manual.py
- diagnose_web_ui.sh and verify_web_ui.sh overlap diagnose_web_interface.sh,
  which the docs point to
- fix_internet_connectivity.sh is iptables-only (stale on nftables)
- diagnose_plugin_permissions.sh, dev/validate_python.py
- download_nba_logos.py + README_NBA_LOGOS.md: logo_downloader fetches
  logos on demand
- setup_plugin_repos.py linked into the production plugin-repos/ dir; the
  dev workflow is scripts/dev/dev_plugin_setup.sh, and
  MULTI_ROOT_WORKSPACE_SETUP.md now uses it

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(config): drop unused plugin_system flags and a dead unit comment

- config.template.json: remove plugin_system.auto_discover,
  auto_load_enabled and development_mode. Nothing reads them; the web UI
  only stores them when a client sends them. ConfigManager's migration
  only adds template keys, so existing configs keep theirs unchanged.
- config.template.json: re-indent vegas_scroll's live_* keys.
- systemd/ledmatrix.service: remove the comment documenting
  LEDMATRIX_ON_DEMAND_PLUGIN / on_demand_env.conf; nothing reads either.
- CONFIG_REFERENCE.md: say the legacy keys are no longer in the template.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: delete docs/archive and PLUGIN_IMPLEMENTATION_SUMMARY.md

- docs/archive/: superseded guides; the repository history keeps them
  and no live doc links into the directory. The one open document in it,
  WEB_UI_AUDIT_2026-09.md, moves to docs/audits/ and is linked from the
  docs index.
- PLUGIN_IMPLEMENTATION_SUMMARY.md invented usage statistics, called
  v2.0.0 current, listed shipped auto-updates as future work and
  documented a BasePlugin.get_config() that does not exist.
- docs/README.md: drop both, and stop telling contributors to archive
  obsolete pages instead of deleting them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(plugin-api): fix extra_small_font size, cache metric key and scroll pacing example

- PLUGIN_API_REFERENCE: extra_small_font loads at 7, not 6 (crisp_size
  snaps it, src/display_manager.py); get_cache_metrics() returns
  cache_hit_rate, not hit_rate (src/cache/cache_metrics.py).
- ADVANCED_PLUGIN_DEVELOPMENT: the basic scrolling example slept in a loop
  and never passed frame_hold; use ScrollHelper + scroll_config.configure()
  and set_scrolling_state(True, frame_hold=...) as PLUGIN_API_REFERENCE does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(plugin-config): match the config tab, icon and web-action docs to the code

- PLUGIN_CONFIG_QUICK_START / PLUGIN_CONFIGURATION_TABS /
  PLUGIN_CONFIGURATION_GUIDE: there is no "Reset to Defaults" button (the
  tab has Refresh, Update, Uninstall, Save Configuration); plugin config
  hot-reloads (ConfigService + on_config_change), so no restart; the
  schema is found by the fixed name config_schema.json, not a manifest
  config_schema field; the tab row is "Plugin Manager", not "Plugins";
  forms are server-rendered from /v3/partials/plugin-config/<id>; the
  duration hook is get_display_duration()/display_duration; a class_name
  mismatch raises PluginError; the store requires id, name, class_name and
  display_modes (not version); plugin_system.debug/log_level do not exist
  (use run.py -d / LEDMATRIX_DEBUG). Drop "future" features that shipped.
- PLUGIN_CONFIG_CORE_PROPERTIES: list all of CORE_PLUGIN_PROPERTIES,
  including skin, skin_options and the vegas_* tuning keys.
- PLUGIN_CUSTOM_ICONS: icon is only a Font Awesome class (fallback
  fa-puzzle-piece); emoji/URL icons and getPluginIcon() never existed in
  v3. Note that /api/v3/plugins/installed currently omits icon.
- PLUGIN_WEB_UI_ACTIONS (+ example JSON): success_message, error_message
  and step1_message are never read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(store): describe the monorepo registry and the store UI as they are

- PLUGIN_STORE_GUIDE: the Plugin Store is a section of the Plugin
  Manager tab; URL installs are "Install from GitHub" -> "Install Single
  Plugin"; bulk update exists (Check & Update All) plus opt-in weekly
  auto-update; PluginStoreManager() defaults to plugins/, so the Python
  examples pass plugin-repos; registry plugins are downloaded (GitHub API,
  ZIP fallback), not cloned; updates compare version with latest_version.
- PLUGIN_REGISTRY_SETUP_GUIDE: replace the per-plugin-repo + tag
  walkthrough with a short page on the monorepo registry (plugin_path,
  latest_version, update_registry.py) that points at the monorepo's own
  SUBMISSION.md. Drops the reference to the deleted
  PLUGIN_IMPLEMENTATION_SUMMARY.md and setup_plugin_repos.py.
- plugin_registry_template.json: use the real entry shape.
- PLUGIN_QUICK_REFERENCE: automatic background updates exist (opt-in);
  registry example and publishing steps use the monorepo, not tags.
- PLUGIN_DEVELOPMENT_GUIDE: tags/releases are not read by the store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(readme): fix the Triple Bonnet mapping, install prerequisites and backup names

- README: the Adafruit Triple Bonnet uses `regular` (3 outputs), not
  `regular-pi1` (1 output) -- src/matrix_support.py MAPPING_OUTPUTS, and
  the README's own hardware_mapping section; the template default mapping
  is adafruit-hat, the PWM mod switches it to adafruit-hat-pwm; manual
  install only needs git up front (first_time_install.sh installs
  python-dev-is-python3, cmake, ninja-build etc.; cython3/scons are not
  used); the Pi Zero 2 W is a supported low-memory board, consistent with
  PRODUCT.md, LOW_MEMORY_BOARDS.md and the installer's low-memory build;
  fix the "First_time_install.sh" spelling, an orphan "2." list item and
  the hello-world starter link (it lives in the plugins monorepo).
- CONFIG_DEBUGGING: automatic backups are
  config/backups/config.json.backup.<YYYYMMDD_HHMMSS_ffffff> (five kept),
  not config_YYYYMMDD_HHMMSS.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(dev): correct the test-running and rgbmatrix build instructions

- HOW_TO_RUN_TESTS: coverage is not collected by a plain pytest run and
  pytest.ini has no threshold; the only one is --cov-fail-under=52 in the
  core unit-test job of .github/workflows/test.yml, which runs the whole
  test/ tree (not an allowlist). Almost no tests carry markers, so
  -m integration / -m slow select nothing; drop them and -m unit as the
  quick check. Replace the hardcoded /home/chuck path.
- DEVELOPMENT: the rgbmatrix package is built with pip install . from
  the submodule root (scikit-build-core + CMake + Ninja), as
  first_time_install.sh does; there is no make build-python /
  bindings/python step, and the build deps are python-dev-is-python3,
  cmake and ninja-build, not cython3/scons.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(wifi): the setup AP is open; auto-enable can be turned off without code changes

- WIFI_NETWORK_SETUP / SSH_UNAVAILABLE_AFTER_INSTALL: both AP paths in
  src/wifi_manager.py create an open network and nothing reads
  ap_password, so drop the "ledmatrix123" password and the ap_password
  key/advice.
- SSH_UNAVAILABLE_AFTER_INSTALL: disabling automatic AP mode does not
  need code changes -- auto_enable_ap_mode is a WiFi-tab toggle and
  POST /api/v3/wifi/ap/auto-enable; note the monitor daemon reads
  wifi_config.json at start, so restart it after changing the setting.
  Use the ledpi username and a relative install path like the other docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(reference): add auto_update, drop drifted line numbers, fix UI and service details

- CONFIG_REFERENCE: document the top-level auto_update.enabled key (read
  by web_interface/auto_update.py and src/auto_update_setup.py); replace
  drifted file:line references with function names; the template's
  dim_schedule mode is "global".
- ADVANCED_FEATURES: core does not read a per-plugin background_service
  block (the sports plugins read their own), and priority is "higher
  number = higher priority" on FetchRequest but not used for ordering.
- WEB_INTERFACE_GUIDE: the General tab toggle is "Web Display Autostart"
  (web interface service), brightness is 1-100, and config paths are
  relative to the LEDMatrix folder, not /config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: drop references to code removed in #608

get_installed_plugin_info, WiFiManager's saved_networks and the six
always-skipping plugin test files are deleted there. NetworkManager already
remembers joined networks; LEDMatrix no longer stores WiFi passwords.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: don't link SKIN_SYSTEM.md from the core-properties page

#615 deletes SKIN_SYSTEM.md; with this link, whichever of the two merged
second would break test_doc_links. The skin/skin_options entries go when
#615 removes the keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:36:38 -04:00
ChuckandClaude Opus 5.5 e1ce7189f1 fix(install): make one-shot retry() retry, and drop root grants on user files (#606)
retry() in one-shot-install.sh used `if ! "$@"; then status=$?`, where $? is
the status of the negation -- always 0. A failed command was never retried
and retry() reported success, so a failed `git clone` carried on until a
later check noticed the missing checkout. It now retries (3 attempts) and
returns the command's status. The two apt steps stay non-fatal: warning and
continuing is what they effectively did before, and making them fatal would
stop installs that work today. A clone that keeps failing stops the install,
as it already did, just sooner and with the one-shot's own error message.

Both installers granted the web user NOPASSWD root on display_controller.py,
start_display.sh and stop_display.sh. Those files are owned by the user after
Step 11's chown, so the grant let the web user rewrite them and run them as
root, and nothing ever ran them through sudo. Removed from both installers,
with a test that every project file granted as root is a root-owned
fix_perms helper.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:34:20 -04:00
claude[bot]andClaude 967f3a0567 fix(install): parse the sudoers rules before installing them (#602)
Both installers generated the ledmatrix_web rules and copied them straight
into /etc/sudoers.d without ever parsing them. Every rule is built from
`which` lookups, so an empty or surprising path produces a malformed
drop-in -- and a malformed file in /etc/sudoers.d makes sudo refuse every
command for every user. On a headless Pi that is unrecoverable over SSH.

first_time_install.sh now runs `visudo -c` on the generated file and, if it
does not parse, prints what visudo said and leaves the installed file
untouched rather than replacing it with a broken one. configure_web_sudo.sh
does the same before it offers the rules for confirmation.

first_time_install.sh also built the file at a fixed /tmp path as root;
mktemp now picks the name.

test/test_sudoers_is_validated.py renders the installer's own sudoers
heredoc and checks the result with visudo -- the check neither installer
had -- and asserts the install stays gated on it.


Claude-Session: https://claude.ai/code/session_01Dby94z9PV3zVM25fqGNXTt

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-21 16:08:29 -04:00
ChuckandClaude Opus 5 116abb0daa fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service

BackgroundDataService always sent a season range first and, on a 400,
fell back to chunks without recording the rejection, so every background
season fetch spent a doomed request and live scoreboards learned nothing
from it (or it from them). The worker now consults and sets the same
6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight
to month/day chunks, and if every chunk fails the range is asked once for
a real error without re-spending the chunks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep plugin asset and action routes inside their directories

POST /plugins/assets/upload, GET /plugins/assets/list and POST
/plugins/assets/delete joined the request's plugin_id onto assets/plugins
unchecked, so '../../config' created, wrote, listed and deleted outside
it. #561 guarded only the route that serves the files. All three now go
through path_safety.resolve_under and answer 400 for anything but a
plain name, and delete only unlinks a metadata path that resolves into
that plugin's uploads directory.

PluginManager.get_plugin_directory refuses ids that are not one plain
path segment, so /plugins/action (which runs a manifest script from the
returned directory) and every other caller get the guard; the action
route also rejects such ids up front, covering its no-manager fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): report a no-op plugin update as already up to date

update_plugin() returns True both for a real update and for "nothing to
do" (a ZIP-installed monorepo plugin already at the registry version, a
bundled plugin). With no git commit to compare, POST /plugins/update
called every such success "updated successfully", so Check & Update All
counted most official plugins as updated on every run.

The route now reads what changed off the plugin itself (commit, else
manifest version, else last_updated) and returns data.update_status
(updated / up_to_date / local_only). The update-all toast is summarised
by PluginInstallManager.summarizeUpdateResults from that status, falling
back to the message for older servers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): scoreboard scroll speed no longer follows target_fps

sports_scroll computed the crisp speed ladder against the global
target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since
frame-locked presentation (#545) the helper steps a fixed number of whole
pixels per presented frame and the panel presents at its real refresh, so
the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a
50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s.

The ladder now uses the display manager's refresh_hz, then
display.hardware.limit_refresh_rate_hz, then the default. target_fps is
not consulted. Docstrings now say scroll_delay is ignored for pacing (no
behaviour change there) and describe the fixed-step model.

Tests: replace the tests that pinned target_fps as the ladder refresh and
described time-based stepping; assert speed independence from target_fps
(unit and end-to-end presented px/s against the real helper), that the
fixed per-frame step is applied, and that scroll_delay does not change
speed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): escape registry and upload values in plugin manager inline handlers

The store, saved-repository and custom-registry buttons built
onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone,
so a custom registry entry whose id contained ' closed the attribute and
added its own handler. One helper, jsStringAttr(), now HTML-escapes the
JSON literal for every one of those handlers, and the store View button
opens only http(s) repo links.

The live window.updateImageList (plugins_manager.js loads last, so its
copy wins over the file-upload widget's) wrote the uploaded file's
original name, path and ids into markup raw; they are escaped now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note plugin asset, action and inline handler guards

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): let the root pip wrapper install web_interface/requirements.txt

Update Code, the automatic update's health check and Install Base
Requirements install web_interface/requirements.txt through
safe_pip_install.sh, which only allowed the root requirements.txt. The
first commit changing that file would fail its dependency install, and
the automatic updater rolls back any update whose dependencies did not
install -- on every device, for every newer commit.

The wrapper now lists both core requirement files. Only their folders
are resolved, so a requirements.txt symlinked out of the project is
compared by its target and refused (previously the root file's own
symlink target was what got allowed). The updater's file list is a
named constant, and a test runs the real wrapper (pip stubbed) on
every file Update Code and the rollback install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): do not retry plugin requests that got an HTTP answer

PluginAPI.request wrapped everything that was not a structured error as
NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a
JSON error without error_code included. Check & Update All retries
NETWORK_ERROR, so those updates were re-sent five more times with
backoff, contrary to the #587 contract that an HTTP error response is
the server's answer.

NETWORK_ERROR now means only that fetch() rejected. Any HTTP response
without an error_code, or with a body that is not JSON, is API_ERROR
with the HTTP status attached. Tested against the shipped api_client.js.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): restart the stats window when an idle gap is dropped by size

#582 dropped an idle gap from the frame stats two ways: the reset_scroll()
sentinel, which also restarts the 5s window timer, and a size guard for
scrollers that never call reset_scroll(), which did not. On that path the
first real frame after the gap found the boundary overdue and logged a
stats line for a one-frame window. Both paths now share one seeding helper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): leave plugins alone when update_core's own rollback fails

update_core returns rollback_failed directly when a partial pull or an
update whose health check never started cannot be rolled back. run()
only held plugins back for 'verifying', so those devices still got new
plugin versions and a display restart on top of a core in an unknown
state -- the opposite of what the health-check path does, and of the
3.4.0 changelog (plugins are left alone if the rollback fails).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(api): make the REST reference match the api_v3 package

Every documented request body, query parameter and response shape was
re-checked against the handlers in web_interface/blueprints/api_v3/.
Fixes calls that failed as documented (repo_url, action_id/params,
files/image_id, font_file+font_family, ?font=, cache key,
auto_enable_ap_mode, plugin limit keys), removes the font-override
endpoints dropped in #566, corrects response shapes (plugins/config,
plugins/schema, health, metrics, operation history, github-status,
fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds
the 26 routes it omitted (backup, system auto-update/git, wifi radio,
starlark editor, MQTT bridge, status endpoints, skins).

Documents the merge semantics of partial JSON saves to /config/main and
/plugins/config and the dim-schedule POST accepting GET's days shape,
which land in the same change set. Replaces app.py line numbers and the
removed api_v3.py path with file and function names.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): remove the General-tab plugin system toggles that did nothing

plugin_system.auto_discover, auto_load_enabled and development_mode had
General-tab toggles whose help tips promised dormant plugins and verbose
logging, but nothing reads them: every enabled plugin is discovered and
loaded regardless. Remove the three toggles.

The keys stay tolerated in stored configs. The save handler now stores
a flag only when a client sends it; treating a missing key as an
unchecked box would otherwise rewrite all three to false on every
General-tab save, which still posts plugins_directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(scroll): remove dead code left by #523/#570

- Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read
  them since the numpy blend replaced the scipy path.
- Drop ScrollHelper._last_integer_position and frame_time_target, which
  were written but never read.
- Keep target_fps and set_target_fps() but document them as
  informational: nothing paces off them, yet ledmatrix-elections'
  test_scroll_pacing.py reads helper.target_fps back and third-party
  plugins may call the setter.
- Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to
  throttle", set_sub_pixel_scrolling's "default: True", and
  set_frame_based_scrolling's claim that it steps.

The plugins monorepo was grepped for every removed name; none is used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI

FONT_MANAGER.md told plugins to read display_manager.font_manager, which
does not exist, so a plugin following it failed to load with
AttributeError. The shared FontManager lives on the PluginManager and
BasePlugin._get_font_manager() returns it (with a fallback for harnesses).

Also removes the Fonts-tab override workflow and element-override panels
that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls

/plugins/store/search does not exist (404) and the list endpoint reads
query, not q. The registry guide's curl examples omitted the JSON
Content-Type, so the handlers saw an empty body and answered 400.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): use the shared core-key list in the last three private copies

StartupValidator warned "Plugin 'auto_update' is enabled but not found" on
every display start with auto-update or a dim schedule on; the reserved
plugin-id check missed auto_update, sync, location and the rest; and
ConfigManager's (uncalled) orphan cleanup would have deleted display,
schedule and auto_update. All three now read src/core_config_keys.py, which
also gains CORE_SECRETS_KEYS for the github/youtube secrets sections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): partial JSON saves to /config/main change only what they send

A JSON body with one field reset every checkbox in the sections it touched:
the MQTT bridge's brightness slider turned off disable_hardware_pulsing,
inverse_colors, show_refresh_rate and use_short_date_format, and a
timezone-only save turned off web-UI autostart and weekly auto-updates.
Missing-means-unchecked now applies only to form posts: form-encoded bodies
and the v3 forms, which mark themselves with a hidden __form_section input.

Also on the config routes:
- vegas_min/max_cycle_duration no longer match the generic *_duration rule,
  so they stop landing in display_durations and a blank one no longer
  rejects the whole Display save;
- saving from the Raw JSON editor calls start_setup_if_needed like the
  General form, so enabling auto-update there finishes its setup;
- the schedule and dim-schedule POSTs accept the per-day days.<day> shape
  their GETs return, as well as the flat form keys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): install plugin dependencies from the configured plugins directory

install_plugin_dependencies.sh scanned only plugins/, but the Plugin
Store installs into plugin_system.plugins_directory (default
plugin-repos), so the documented "Recommended" fix found 0 plugins on
every store install. It now reads plugins_directory from
config/config.json (relative to the project root or absolute, default
plugin-repos) and also scans plugins/ for dev symlinks, installing a
plugin reached through both only once.

With set -e alone, `pip ... | tee` took tee's exit status, so a failed
pip install was reported as success; set -o pipefail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: replace stale API names, line numbers and the api_v3.py path

- ADVANCED_FEATURES: StreamManager methods that exist
  (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the
  real on-demand status envelope ({status, data: {state, service}})
- app.py:199 / :144 / :607-619 line citations and
  web_interface/blueprints/api_v3.py (now a package) replaced with file and
  function names in ADVANCED_FEATURES, CONFIG_DEBUGGING,
  PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE,
  PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README
- CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use
  /config/raw/main to replace the file; describe where validation runs
- TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints
  usage)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): verify the web interface that actually ships, on port 5000

verify_installation.sh failed every healthy install: it required the
long-removed web_interface_v2.py and looked for a listener on port 5001,
while the web interface binds 5000 (web_interface/start.py). It now
checks the files ledmatrix-web.service runs (start_web_conditionally.py,
web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the
same 5001 port in its listen check, HTTP probe and printed URLs.

Port matches are anchored so :50001 no longer counts as :5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): one display-size contract: display_manager.width/height

CLAUDE.md (#580) says to read display_manager.width/height because
matrix is None when hardware init fails; the development guide, the
safety-harness doc and two DisplayManager docstrings still recommended
matrix.width/height. The bundled starlark-apps plugin read matrix.width
unguarded, so its magnify recommendation and frame scaling raised in
fallback mode (e.g. after the Pi 5 hardware refusal).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): make install_service.sh --help print usage instead of installing

install_service.sh parsed no arguments, so `sudo ./scripts/install/
install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md)
rewrote ledmatrix.service, ledmatrix-web.service and both update-verify
units and enabled/started them. It now handles -h/--help (usage, exit 0,
no changes) and rejects any other argument with exit 2 before doing
anything. Running it with no arguments, as first_time_install.sh does,
is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scroll): describe the fixed-step model and document frame_hold

Since #545 a crisp speed from scroll_config.configure() makes the helper
advance a fixed whole-pixel step per presented frame with no clock, and the
display manager's frame hold is part of the speed. The docs still described
the removed wall-clock model:

- scroll_config's module and configure() docstrings said speed is applied
  in time-based mode and that omitting the hold "falls back to fractional
  pixels"; omitting it actually runs the scroll frame_hold times too fast.
- SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both
  modes, and read a 20 ms stats median as missed refreshes although that
  is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step,
  the hold-dependent healthy median, that target_fps plays no part, and
  that a hand-added scroll_pixels_per_second loses to a schema-default pair.
- PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling)
  without frame_hold; it now documents the parameter (core 3.4.0) with a
  configure() + set_scrolling_state example.
- update_scroll_position/set_scroll_speed and set_scrolling_state
  docstrings say the same.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark target_fps legacy; describe what Vegas scroll_delay does

- General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after
  the sports_scroll fix nothing in core scrolling reads it. The field and
  its API validation stay so saved configs and plugins that read
  global_config['target_fps'] keep working. CONFIG_REFERENCE says the same.
- Vegas frame_based_scrolling/scroll_delay were described as frame-count
  stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based
  mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and
  still advances by elapsed time, so the applied speed is
  clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The
  config comments, render_pipeline comment and CONFIG_REFERENCE rows now
  say so. No behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(deps): describe how plugin dependencies are really installed

The guides said the web service runs as root, that installs pick --user
from os.geteuid(), and quoted a warning and a
PluginManager._install_plugin_dependencies() method that don't exist. The
web unit runs as the installing user; store installs go through
install_requirements_file() and sudo safe_pip_install.sh (root), with a
user-level fallback that says so, and load-time installs run in the
display service's own (root) interpreter.

Manual paths now use the configured plugins directory (plugin-repos/ by
default) instead of plugins/, which store installs no longer use, and
install_plugin_dependencies.sh is described as scanning that directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): count local changes one way for the preflight and the pull

The automatic update's preflight ignored mode-only changes and anything
whose status line contained plugins/ or plugin-repos/, then promised
"Automatic updates will not stash your changes". perform_core_update
used plain git status (modes count) and ignored only 'plugins/', then
ran 'git stash push -- :!plugins', which nothing ever pops. So an edit
to a bundled plugin under plugin-repos/, or the installer's chmods on
tracked scripts, passed the preflight and was stashed away for good.

- auto_update.local_changes() is the one predicate both use:
  core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/
  excluded by leading folder rather than substring (a core file under
  web_interface/static/v3/js/plugins/ now counts).
- Update Code's explicit stash leaves out both plugin folders; the
  pull's --autostash carries their edits and mode changes across and
  reapplies them.
- The automatic updater calls perform_core_update(stash_local_changes=
  False), which refuses instead of stashing edits that appeared after
  the preflight; update_core reports that as 'blocked'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): diagnostics follow the web autostart default and api_v3 package

#556 made a missing web_display_autostart mean "start" (only an explicit
false/off keeps the web interface down), but the diagnostics still said
otherwise: diagnose_web_ui.sh reported a missing key as "defaults to
false", diagnose_web_interface.sh said the web interface "will not start
unless this is set to true" and recommended enabling it, and
debug_web_manual.py printed False. Troubleshooting a down web UI pointed
users at a non-cause.

Both shell scripts now evaluate the setting with the launcher's own
autostart_enabled() (inline fallback if it cannot be imported) and report
on / off / not set (on) / unparseable config; debug_web_manual.py uses
the same function. They also check web_interface/blueprints/api_v3/
__init__.py: api_v3.py became a package in #553, so every healthy
checkout was reported as missing a file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(install): what install_service.sh installs; verify script port; no sudo for --help

install_service.sh installs and starts ledmatrix, ledmatrix-web and the
update-verify units, not only ledmatrix.service (systemd/README.md,
README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help'
as a harmless check; it now shows --help without sudo and warns what a
real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh
checks the web interface on port 5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note update-all, plugin system settings and script fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): size the preview after orientation and pixel mappers

display_geometry.physical_size claimed to give DisplayManager's answer but
only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured
after the library's pixel mappers, so a Rotate:90 / orientation 90 chain
previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for
128x64.

Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc
does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown
names leave it alone), and move the orientation composition here so
DisplayManager and the preview share it. The module docstring no longer
claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH.

Audit finding F18.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): refuse settings the rgbmatrix library aborts on, on every board

The library answers several settings with a NULL matrix or abort() rather
than an error, so the display service crash-looped (Restart=on-failure)
instead of reaching fallback mode: rows above 64, chain_length above 255
(uint8_t binding setter, documented as "no upper limit"), a misspelled
hardware_mapping, and parallel 2-3 on a single-output mapping, reachable
from the Display form on the default adafruit-hat(-pwm) mapping. #586 only
guarded the Pi 5 subset.

- src/matrix_support.py holds the rules for every board (Options::Validate
  ranges, binding integer types, mapping names and outputs from
  lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the
  API's numeric ranges.
- DisplayManager checks them before building options and raises
  MatrixSettingsRefused, so a hand-edited config falls back with a logged,
  reported reason. Emulator mode only warns.
- The config API refuses them with a 400 naming the setting; combinations
  are checked against stored values but reported only when the request
  sets a field involved.
- The hardware status file gains "cause" (settings/library/forced). The
  fallback log and Display banner give the Pi 5 rebuild hint only for a
  library failure instead of rebuild + gpio_slowdown advice for every
  failure; one Pi 5 slowdown recommendation (1-3, start at 1).
- The Display form offers classic/classic-pi1 and orientation 90/270 and
  renders any other stored mapping selected with a warning, so an
  unrelated save no longer rewrites them; the API accepts 90/270.

Audit findings F03, F16, F19, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(display): library limits, template defaults and Pi 5 slowdown

- rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs,
  classic/classic-pi1 mappings and orientation 90/270 documented.
- Defaults are the config.template.json values: config migration adds
  missing keys from the template, so the listed "code defaults" never
  applied.
- One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode,
  starting at 1.
- Troubleshooting describes the refused-settings fallback, and CHANGELOG
  corrects the Unreleased "no upper limit" entry.

Audit findings F19, F20, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py opens the panel with the service's options

--measure and --demo built RGBMatrixOptions from a private copy of the
display service's builder that had drifted: gpio_slowdown came from
display.hardware (default 2) instead of display.runtime (default 3), and
rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors,
pixel_mapper_config and orientation were skipped, with different defaults
(hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown
was measured -- or garbled -- in a setup the service never drives.

The option filling in DisplayManager._setup_matrix moves, unchanged, into
DisplayManager.apply_matrix_options(options, config), which _setup_matrix
calls and the script reuses (overriding only limit_refresh_rate_hz for
--measure). The script now loads the whole config rather than the hardware
block. Tests pin the script's options to the service's attribute for
attribute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py recommends keys the resolver honours

The ladder ended by telling users to set
display_options.scroll_pixels_per_second. scroll_config ranks that key
below the scroll_speed + scroll_delay pair, deliberately, and several
plugin schemas default the pair into config, so the advised key was
silently ignored (a schema-default 1/0.02 pair plus an advised 66 still
resolved to 50 px/s).

The advice is now the pair that selects the crisp speed exactly
(pixels_per_frame every frame_hold/refresh seconds), explains that the
pair outranks scroll_pixels_per_second, and gives the scoreboards'
per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the
printed pair over a schema-default pair and check it lands on the
advertised speed and hold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula

- SPORTS_UNIFICATION.md still presented honouring global target_fps as
  sports_scroll's added behaviour and its one user-visible gain; note that
  it was withdrawn because it had become a speed multiplier.
- ADVANCED_FEATURES.md gave Vegas scrolling as
  (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s
  by elapsed time, through a 0.1-5 px per scroll_delay clamp when
  frame_based_scrolling is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): scroll model fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(dev): link-github links plugins from the ledmatrix-plugins monorepo

link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git,
and those per-plugin repositories no longer exist: official plugins are
directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the
monorepo once into the dev directory, finds plugins/<name>,
plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and
links it under its manifest id. With an explicit repo URL it still links a
single-repository plugin as before.

dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a
fork), plus plugins_repo and plugins_branch; github_pattern, which was
documented but never read, is dropped and warned about. Ships
dev_plugins.json.example and git-ignores dev_plugins.json, both of which
the guide promised. Reading JSON falls back to python3 when jq is missing
(get_plugin_id silently returned nothing without jq).

update/status/list find the git checkout above a monorepo plugin
directory (its .git is not in the plugin dir), and update pulls a shared
checkout once. status no longer exits 1 when nothing is broken.

Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration,
workflow, store integration, hello-world link, submission), and the
nonexistent scripts/git-hooks/pre-push-plugin-version and
scripts/bump_plugin_version.py replaced with the real rule: bump the
manifest version and run update_registry.py. scripts/dev/README.md and
CLAUDE.md updated to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scripts): monorepo workspace layout; fix_perms and install READMEs

MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin;
setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into
plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the
workspace file opens LEDMatrix plus ../ledmatrix-plugins.

scripts/fix_perms/README.md listed cache directories
fix_cache_permissions.sh never touches and a 'ledmatrix' service user
that doesn't exist (also in scripts/install/README.md); adds
safe_pip_install.sh. install/README: install_service.sh installs the web
and update-verify units too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): keep the rollback's pip retries inside the unit time limit

The health check reinstalled the previous requirements by trying the
next bash path after any failure, including a 600 s pip timeout. Two
files, two paths: up to 40 minutes of pip alone, while systemd stops
ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the
rollback half-way and leaving the update 'verifying' until the web UI
calls it lost.

- Like permission_utils.install_requirements_file, only a sudo refusal
  moves on to the next bash; a pip that ran and failed or timed out is
  not repeated. The refusal wording is one list
  (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only
  verifier and pinned equal by a test.
- All reinstalls in one rollback share a 600 s budget.
- WORST_CASE_SECONDS adds up every timeout on the longest path (27.5
  min); a test holds it under the unit's TimeoutStartSec and that under
  the web UI's VERIFY_LOST_SECONDS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools

Plugin config was prepared differently depending on how it arrived:

- JSON POST /plugins/config built a partial body on schema defaults, so
  {"enabled": true} reset every other setting of the plugin. It now merges
  onto the stored section first, as the form path already did.
- Legacy-boolean normalization (#588) ran only at load: GET /plugins/config
  returned the raw boolean, posting it back failed validation, and hot
  reload handed plugins the raw section (a legacy dynamic_duration: true
  came back as a boolean). schema_manager.prepare_plugin_config (normalize,
  then defaults) is now used by PluginManager.load_plugin, both save paths,
  GET, the save notifications and DisplayController's hot-reload callback.
- The JSON save's filter kept only enabled/display_duration/live_priority
  and dropped a submitted skin, skin_options or vegas_* tuning key. There
  is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES,
  used by validation and by the save filter; PluginManager's
  CORE_OWNED_CONFIG_KEYS is its vegas subset.
- Plugin sections posted to /config/main were stored verbatim, including
  values /plugins/config rejects. They now go through the same preparation
  (_prepare_plugin_config_for_save, extracted from save_plugin_config), and
  a failing section rejects the whole save before anything is written.
- dev_server read only top-level defaults and let a schema enabled:false
  win; build_full_config shallow-merged overrides, dropping sibling
  defaults; the harness extracted defaults differently from the device.
  loading.build_config now uses the device's extraction and preparation,
  and dev_server, check_plugin, render_plugin and the harness all use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt-bridge): brightness changes apply live and touch nothing else

The display service's hot reload applies a saved brightness within a few
seconds, and /config/main no longer resets other display settings on a
brightness-only JSON body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): automatic update hardening

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI

It described web_interface_v2.py and index_v2.html (both gone), client-side
form generation, one POST per field with {key, value}, and 'no nested
objects'. The v3 UI renders plugin forms server-side from the schema
(pages_v3 partial + plugin_config.html macros, nested sections and
x-widgets), posts the whole form once, and save_plugin_config() merges onto
the stored section, validates, splits x-secret fields and notifies the
plugin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt): brightness saves apply via hot reload and leave other settings alone

The bridge README said brightness is applied on the display's next
restart; the display controller's config hot reload applies it within
seconds. It also now states that the bridge's partial JSON save changes
only brightness (the /config/main merge fix in this change set).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): don't log pip's output from the health check's reinstall

pip can echo a private index URL with embedded credentials;
permission_utils redacts it, the stdlib-only verifier cannot, so it
logs the exit code only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark the plugin_system toggles as unused legacy keys

auto_discover, auto_load_enabled and development_mode are read by
nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE
said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and
the REST reference listed them as live settings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): docs and developer tools group

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): legacy plugin-system toggles no longer count as a General save

auto_discover, auto_load_enabled and development_mode have left the General
form, so a post carrying only one of them is not a general-settings save and
must not treat web_display_autostart and auto_update as unchecked. The
plugin_system block itself is left as on main for the branch that reworks it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): config-save and plugin-config preparation fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(claude): re-check matrix_support.py rules when the library submodule is bumped

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address Codacy findings on the core audit PR

- plugin_manager.prepare_plugin_config: when the fallback legacy-boolean
  pass also fails, log a warning instead of a bare except/pass.
- api_client.js: request() refuses any endpoint that is not a plain path
  under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace,
  control characters) with INVALID_ENDPOINT before calling fetch(), and
  plugin ids are URL-encoded wherever they are put into a URL (also in the
  app-shell batch load).
- test_update_all.js: pins both against the shipped client.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): check endpoint control characters without a control-character regex

Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags
the \x00-\x1f range in checkEndpoint's regex. Test the char codes
instead; the endpoints refused are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(auto-update): make the seed script executable on disk, not only in the index

On Linux Repo.publish() commits with -a, which recorded scripts/run.sh
as 100644 upstream because the seed file was never chmod +x. The pull
then brought in the same mode the installer chmod had made locally, so
installer_chmod saw no mode change left to check. The updater was fine:
with the upstream commit at 100755 the --autostash carries the device's
chmod across. Verified under Linux (WSL, git 2.43): the old helper fails
exactly as CI did, the fixed one passes all 63 tests in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 16:37:29 -04:00
ChuckandClaude Opus 5 f475038895 fix(cache): web UI can read what the display service caches again (#593)
* fix(cache): web UI can read what the display service caches again

ledmatrix-web.service carried CacheDirectory=ledmatrix. With User= set to
the installing user, systemd re-owns /var/cache/ledmatrix and everything
in it to that user and its primary group whenever the directory's owner
differs -- for a directory root created, on the first start. That erased
the root:ledmatrix setgid layout the installers set up, so every file the
display service (root) wrote afterwards was root:root 0660 and unreadable
by the web interface:

  WARNING - Permission denied loading cache for display_current_state ...

Since #547 install_service.sh renders the web unit from the template, so
every fresh install hit this. Measured on one rig: 392 unreadable files,
and the web UI's display status, on-demand state and plugin health empty.

Existing installs only receive `git pull`, never a reinstalled unit, so
the fix for them is in the code the root display service runs:

- DiskCache.set gives each file the directory's group (when the directory
  is group-writable) and 0660 on the open descriptor before the rename,
  independent of setgid. This also closes a window where a fresh file was
  visible as mkstemp's 0600.
- DiskCache.share_existing_files repairs files an older version left
  behind, once per process from the cleanup thread. It works through
  O_NOFOLLOW descriptors and skips hard links and other users' files: the
  directory is writable by the web user, and root must not be steered
  into changing a file outside it.

For new installs, the web unit drops CacheDirectory=/CacheDirectoryMode=,
and install_web_service.sh stops replacing an existing directory's
ledmatrix group with the user's group.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): on-demand and current-display status read the display's latest state

Found testing the cache-permission fix on a rig: once the web interface
could read display_on_demand_state at all, /display/on-demand/status kept
answering "active" for over 100 seconds while the file on disk said
"idle". Both status routes read the display service's keys through the
web process's memory tier, which serves the first copy it read for the
full max_age (120s). Read them with memory_ttl=0, as every other
cross-process reader (plugin health/metrics, the on-demand mailbox)
already does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): re-group the cache dir whenever the web user is outside its group

install_web_service.sh replaced an existing cache directory's group only
when it was root's. A directory in any other group the web user is not a
member of -- root:ledmatrix, for a user who is not in ledmatrix -- was left
alone, and every file root wrote there stayed unreadable to the web
interface. Replace the group whenever the installing user is not in it.

A directory whose group the user is already in (ledmatrix, or the user's
own group where CacheDirectory= left it) is still left as it is: re-grouping
a working directory strands the files already in it on the old group.

When the group does change and root-owned JSON files carrying the old group
are present, try-restart ledmatrix.service so DiskCache.share_existing_files
re-groups them through its symlink- and hard-link-safe path, rather than a
recursive chgrp.

Verified under WSL's systemd for seven directory states (user group,
ledmatrix member, ledmatrix non-member with and without root files,
root:root, missing, unnamed gid); the previous version left the non-member
case unchanged.

Addresses CodeRabbit review on #593.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 12:38:39 -04:00
ChuckandClaude Opus 5 869e36fb2f feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback

A General-tab toggle (off by default) checks for and installs LEDMatrix and
plugin updates once a week, overnight in the configured timezone.

- Pre-update checks skip (and report) instead of forcing: local edits or
  commits, merge/live rebase, no upstream, low disk, missing health check, or
  a version that was already rolled back. An abandoned rebase (HEAD back on a
  branch) is cleared, since it would otherwise block every pull.
- The pull reuses the Update Code path (now perform_core_update(), which
  reports dependency install failures as data).
- ledmatrix-update-verify.service, started via a .path unit from a request
  file, restarts the services from its own cgroup, requires them to come up
  and stay up, and otherwise resets to the previous commit and reinstalls the
  previous requirements. It runs a copy of the checker taken before the pull.
- No SSH needed: switching the toggle on restarts the display service, which
  (as root) installs the two units from the repo templates for the web user.
  first_time_install.sh installs them too and takes --enable-auto-update /
  LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh).
- Plugins update after the code passes its check; failures, blocks and
  rollbacks raise an Overview banner and show under the toggle.

Tested end to end on a Pi: web-UI setup, a good update, a broken web service
and a broken display (both rolled back), a blocked local edit, and an
abandoned rebase found on the device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(auto-update): address static-analysis findings

- Replace the subprocess.CompletedProcess the verifier fabricated for a
  command that could not start with a plain namedtuple; nothing is executed
  there, but the scanner flags any CompletedProcess built from variables.
- Mark the subprocess imports with the repo's standard B404 annotation (all
  calls are list-form argv, no shell).
- Mark the rollback-failed message as not SQL (B608 matched its wording).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): CI failures on Linux

- Keep the setup result when chown fails. CI runs as a non-root user, where
  chown to the web user raises; that discarded the result file, so the
  General tab would never learn whether setup worked. Regression test added.
- Register the two new /api/v3/system/auto-update routes in the URL map
  snapshot.
- Use utility classes app.css defines (space-y-1, hover:text-red-600).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): address review feedback

- Health check: a failed restart command no longer lets the check run
  against the still-running old process; it counts as a failure (and after a
  rollback, as a failed rollback). An unreadable restart count is never
  treated as stable, since a crash loop looks healthy between attempts.
- Installer writes the auto_update setting to a temp file and swaps it in,
  keeping mode and owner, so a running config watcher never reads a
  truncated config.json.
- Verify unit quotes its command-line paths (install folders with spaces);
  setup refuses folder names systemd would reinterpret (%, quotes,
  backslashes, control characters) and says so on the General tab.
- The auto-update status route no longer returns exception text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): keep error detail in the status route's 500

test_web_error_detail requires every 5xx handler to log the traceback and
return describe_exception(e), which redacts credentials, so failures are
diagnosable from the web UI. Dropping it for CodeQL broke that policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): dismiss route rejects non-object JSON with 400

A JSON array or scalar body made `.get('alert_id')` raise, returning 500.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): let the app-wide handler answer status-route errors

CodeQL (py/stack-trace-exposure, #709) flagged the route's own except,
which returned describe_exception(e). web_interface/app.py's error handler
already logs the traceback and returns the same redacted detail for any
unhandled exception, so the local copy is removed: same response, no new
exception-to-response flow, and test_web_error_detail's policy still holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:58:57 -04:00
ChuckandClaude Opus 5 814c21de1c chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0 (#580)
* chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0

Skins: no current scoreboard plugin builds on src.base_classes, so the only
skin hook (SportsCore._render_game) never runs. The plugin schema endpoint no
longer injects the Visual Skin dropdown, the store hides and refuses
"type": "skin" registry entries, and GET /api/v3/skins reports
supported: false with a message. Stored skin config still loads and saves.
src/skin_system/ and its tests are unchanged apart from the support flag.

Docs: check_plugin.py/render_plugin.py examples use --plugin; document
BasePlugin.get_update_interval() and its interaction with the manifest
update_interval; CLAUDE.md drops the stale template line number and
recommends display_manager.width/height.

Preview size: new src/display_geometry.py holds the size computation and
defaults DisplayManager uses (double-sided applied, chain_length default 2).
The web preview, /display/current, Starlark magnify default, sync handshake
and two dev scripts use it.

Release: __version__ 3.4.0, CHANGELOG 3.4.0 section plus a 3.3.0 tag note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address CodeRabbit review on #580

- Preview fallbacks (SSE stream and /display/current) use logical_size({})
  (128x32, the shared default) instead of a hard-coded 128x64.
- display_geometry treats a non-mapping display/hardware block as missing,
  so a malformed config.json falls back to defaults instead of raising
  AttributeError (which turned the Starlark render into an HTTP 500).
- Docs: the static update interval falls back manifest -> plugin config
  -> 60s, in both the API reference and the architecture spec.

Not taken: validating double_sided copies against chain_length/parallel.
An orientation Rotate: or U-mapper pixel mapper decides which axis panels
lie on, so the counts would reject working setups (the existing
vertical-split test is one).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display_geometry): a non-finite hardware size raises ValueError, not OverflowError

CodeRabbit flagged the Starlark magnify default in
_standalone_render_starlark_app for truthy non-mapping display values. That
case was already handled by a9e1bd0b (_display/_hardware treat a non-mapping
block as missing, covered by test_non_mapping_display_config_uses_the_defaults),
and the magnify it produces from the 128x32 defaults is the same as from 64x32.

Checking the same path found one input that still escaped: Python's JSON
parser accepts Infinity, and int(inf) raises OverflowError, which neither the
Starlark path (TypeError, ValueError) nor the preview stream in app.py caught,
so a hand-edited "rows": Infinity returned HTTP 500. physical_size now raises
ValueError for it, matching its documented contract, so every caller's
existing fallback applies. DisplayManager already caught Exception.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:36:58 -04:00
ChuckandClaude Opus 5 d51f7ada14 chore(scroll): drop the dead sub-pixel path, and two dev-tooling papercuts (#570)
Three independent changes, none of which alter runtime behaviour.

1. Remove ScrollHelper._get_visible_portion_subpixel and
   _interpolate_subpixel (162 lines). get_visible_portion dispatches only to
   _blend_visible_portion, so _get_visible_portion_subpixel had no caller, and
   _interpolate_subpixel was reachable only from inside it -- a closed island.
   _blend_visible_portion's own docstring already records that the scipy path
   it replaced was dead; the replacement landed but the corpse stayed.

2. scripts/check_plugin.py: also search ../ledmatrix-plugins/plugins. The
   scoreboards live in the sibling checkout, so --all silently skipped every
   one of them and only --plugin-dir reached them.

3. .gitignore: ignore config/.config_secrets.json.tmp.*, which the suite
   leaves behind several of per run.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:42:50 -04:00
ChuckandClaude Opus 5 92ac231138 fix(fonts): load 4x6 on its pixel grid, from any working directory (#565)
* fix(fonts): load 4x6 on its pixel grid, from any working directory

`extra_small_font` loaded 4x6-font.ttf at 6, off the face's 7px grid.
Under `draw.fontmode = "1"` the mono rasteriser thresholds each glyph at
50% coverage, so every glyph lost its fourth column and deformed:
christmas-countdown rendered "UNTIL" as "VM1JL". The advance is 5px at
both sizes, so snapping to 7 reflows nothing.

- Sizes in DisplayManager._load_fonts go through crisp_size() instead of
  literals. crisp_size / FONT_PIXEL_GRID / FONT_NAME_ALIASES move to
  src/common/font_layout.py; sports_card re-exports them.
- Mirror the fix in VisualTestDisplayManager, the harness's fork of
  _load_fonts. Without it every golden is blessed at the old size.
- Resolve bundled font paths against the install root, not the cwd.
  FontManager._resolve_asset_path now delegates to
  font_layout.resolve_asset_path (kept by name; plugins probe for it).
- The startup banner's middle rung snaps to 7; the 5 rung stays off-grid
  on purpose (the only size that fits a dotted quad on 64px).
- loading.py reads all plugin JSON as UTF-8 (cp1252 on Windows aborted
  check_plugin.py on a 0x9d byte).
- check_plugin.py reports in ASCII and never dies on an unencodable char.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(fonts): resolve relative asset paths from the install root, not the cwd

resolve_asset_path checked os.path.exists(relative_path) unconditionally,
so a relative asset path was still resolved against the process cwd first
-- exactly the dependency this module exists to remove. An unrelated
working directory that happens to contain assets/fonts/4x6-font.ttf (a
stale checkout, a copied assets folder, another project) would shadow the
real bundled font instead of the install root ever being consulted.

Only an absolute path is now returned as-is; a relative path always
resolves against _INSTALL_ROOT first, matching the docstring's stated
contract. FontManager._resolve_asset_path delegates to this function, so
it's covered by the same fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 10:47:01 -04:00
ChuckandClaude Opus 5 59997594ac test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths

Two shared-state problems, both of which show up as a permanently dirty
checkout or an unreproducible test failure.

test_display_dirty_tracking.py builds a real DisplayManager, whose
_snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web
UI reads. Every pytest process on the machine shares that one file, so two
concurrent runs -- CI shards, a second worktree, an agent running the suite
alongside -- overwrite each other's snapshot and the mtime assertions stop
meaning anything. The module fixture now points it at a session-unique temp
path; the individual tests that care still override it further.

To be clear about what this does and does not fix: this is a real shared-path
hazard, but it is NOT the cause of the intermittent 15-test failure in that
module. That turned out to be the emulator's fixed TCP port, fixed in the
follow-up commit. This change stands on its own merits.

web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json
and data/operation_history.json as the web interface runs, into a directory
that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig
that ever opened the web UI -- and every test run that constructs the app --
left three untracked files behind and a permanently dirty `git status`. Only
data/.gitkeep is tracked under data/, so the negation keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop the emulator binding a fixed port, so concurrent runs can't collide

This is the cause of the intermittent full-suite failures we have been chasing:
runs of identical code landing anywhere between 100 and 130 failures, while
every implicated test passed in isolation.

Six test modules set EMULATOR=true and build a real DisplayManager. The repo's
emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to
serve the dev preview. That port is a machine-wide singleton, so a second pytest
process -- a CI shard, another worktree, an agent running the suite alongside --
loses the bind. RGBMatrix construction then raises, DisplayManager catches it and
falls back to `self.matrix = None`, and every test that subsequently touches the
matrix dies with

    AttributeError: 'NoneType' object has no attribute 'SwapOnVSync'

which names neither a port nor a socket, and points at the wrong file entirely.
Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its
matrix-touching tests fail together or not at all -- the 15-test swing that made
the totals look random.

Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process
and running test_display_dirty_tracking.py:

    without this change    15 failed, 6 passed
    with this change       21 passed

The "raw" adapter renders in memory and binds nothing. Only display_adapter is
overridden, in a throwaway config written per pytest process; the repo's
emulator_config.json is untouched and `run.py -e` still opens the browser
preview on 8888. Nothing in the suite referenced the adapter, and the tests
wrap SwapOnVSync on the matrix object itself, so they are indifferent to what
sits underneath. allow_adapter_fallback is forced off -- falling back would
land us on the browser adapter and its fixed port, which is the whole problem.

CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set
to an absolute path: the previous behaviour depended on where pytest was invoked
from, and silently wrote a default config into whatever directory that was.

Verified no regressions: full suite on this branch and with origin/main's
versions of the touched files, same machine, back to back -- 115 failed /
4347 passed on both sides, zero failures unique to either. That 115 is the
pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell
scripts, Linux-only binaries); CI on Linux remains authoritative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: mark the shell entry points executable

Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh`
fails with "Permission denied" and only works if you know to prefix `bash`.
That one matters most: the web UI's own error hint, added in #560, tells users
to run exactly that path when a system action fails for want of passwordless
sudo, and following that instruction verbatim did not work.

All eleven carry a shebang and are invoked directly, never sourced. The two
sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately
left non-executable, which is what distinguishes a library from an entry point.

Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied
with `git update-index --chmod=+x` because this checkout is on Windows, where
core.fileMode is off and the working-tree bit is not tracked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:45:13 -04:00
ChuckandClaude Opus 5 f6367d63ae security: triage the CodeQL backlog — 129 alerts, three of them live (#561)
* fix(web): escape quotes in every HTML escaper, not just & < >

The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:

    value="${escapeHtml(v)}"   title="${escapeHtml(v)}"

so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.

It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.

app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.

cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `&#39;`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.

url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).

test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): stop request-supplied names from reaching paths outside their base

Three of the py/path-injection alerts were live, not lint:

* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
  whose resolved path *string-prefixed* the plugin directory. Flask's
  default converter forbids a slash but not dots, and
  get_plugin_directory('..') returned the parent of the plugins directory
  because it exists -- so every file under the project root then prefixed
  that directory, config/config_secrets.json included. The prefix check
  was also wrong on its own terms: with plugin dir "plugin-repos/foo",
  "../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
  start with "plugin-repos/foo".

* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
  body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
  file_id of "../../../../etc/something" deleted that file. This is the
  one finding in the batch that destroyed data rather than exposing it.

* POST /api/v3/cache/delete passed the body's key through
  CacheManager.clear_cache to DiskCache, which joined it as a filename and
  called os.remove. Same shape, same result. The guard goes in
  DiskCache.get_cache_path, the single choke point get/set/clear share, so
  every caller is covered rather than just this route. Real keys are the
  stems of files already flat in the cache directory -- that is how
  list_cache_files derives them -- so nothing legitimate is turned away.

The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.

Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.

test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): refuse a plugin id that is not a plain name, don't truncate it

pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.

Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(web): say what the plugin web_ui iframe actually is

The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.

Also drops the now-unused os/os.path imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): inline url-input's scheme guard at the previewLink.href sink

CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.

Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.

Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): address CodeRabbit findings on the CodeQL triage PR

- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
  credential-saving or connect flow runs. NetworkManager only accepts
  printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
  was previously saved/attempted before nmcli itself rejected it.

- web_interface/blueprints/api_v3/config.py: fail closed when the
  plugin config schema path can't be resolved under the plugins
  directory (e.g. a symlinked plugin dir). Previously this fell
  through with secret_fields left empty, so submitted credentials for
  that plugin were saved as ordinary, unencrypted configuration.

- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
  splicing the JSON day/column key into an inline oninput="..." handler
  string. escHtml() escapes quotes for a normal HTML attribute, but the
  browser HTML-decodes the attribute before running it as script, which
  undoes that escaping and lets a crafted column name (e.g. from an
  uploaded JSON file) break out of the JS string and execute. Cell
  edits now travel through data-day/data-col attributes read by one
  delegated 'input' listener instead.

  While in this file: fixed 6 pre-existing missing-')' typos on
  multi-line safeSetHTML(...) calls (already flagged by Biome in this
  PR's own CodeRabbit run as syntax errors blocking its lint pass).
  These predate this PR (present on main too) but made the whole file
  fail to parse in any JS engine, which is a bigger problem than the
  XSS finding itself and directly touches the same lines.

Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:32:03 -04:00
ChuckandClaude Opus 5 5137e86d16 feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab

PR #544's change, ported onto the api_v3 package split (#553). Identical
behaviour; only the placement of the new code differs.

The original added 508 lines to web_interface/blueprints/api_v3.py, which #553
deletes, so every hunk of it would conflict irreconcilably. Ported by AST:
26 new top-level items sorted to where the split puts each kind --

  __init__.py   2 imports, 11 constants, 7 helpers
  starlark.py   4 routes  (/starlark/editor/{apps,status,start,stop})
  misc.py       2 routes  (/integrations/mqtt-bridge{,/config})

Everything outside api_v3.py -- the Tools partial, the installer scripts, the
JS tests -- applied unchanged.

Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is
regenerated to match, which is exactly what test_api_v3_url_map.py is designed
to make you do when routes are added.

Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR
ships could not be run here -- node is not installed on this machine -- so
test/js/dom/test_tools_sections.js is unverified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST

The AST-based port of #544 onto the api_v3 package split dropped `time`
from starlark.py's import list. start_pixlet_editor() and
stop_pixlet_editor() both call time.time()/time.sleep() directly, so
every start (NameError building `state['started_at']`) and every stop
that has to wait out the EXIT trap crashed with a 500. No test caught
it because the route's own tests mock subprocess.Popen but never
actually invoked it before now.

Also carries over #544's later fix that this port branched before:
env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an
operator who had already pinned PIXLET_EDITOR_HOST to loopback,
forcing the unauthenticated `pixlet serve` process onto the LAN
regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as
api_v3.starlark.py's siblings already do for _pkg-owned names.

Both fixes route the shared _pkg.time reference the rest of the
package's route modules already use for anything a test might need to
patch, rather than a bare `import time` local to this file.

Ported the existing regression test from #544
(TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's
module layout (web_interface.blueprints.api_v3.starlark instead of the
old monolithic api_v3 module), which is what caught the NameError.

Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this
branch and on origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to
this change).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): clear the six lint errors this rebase introduced

All six were introduced by rebasing this branch onto the merged blueprint
split, not by the split itself. Confirmed by diffing pyflakes output against
main with line numbers normalised -- everything else it reports is present on
main too and is the package's deliberate re-export pattern.

starlark.py used _STARLARK_APPS_DIR three times without importing it (F821).
The rebase resolved an import-list conflict as a union of both sides, and that
symbol was on neither side of the conflict hunk, so it was silently lost. It is
defined in __init__.py and is now imported like its neighbours. This was the
only one of the six that would fail at runtime rather than merely lint.

__init__.py imported contextlib twice (F811): the cherry-pick added one next to
the existing import. Removed the duplicate; the original at line 19 is used.

__init__.py imported signal purely to re-export it to starlark.py, so pyflakes
saw it as unused (F401). signal is stdlib and does not need routing through the
blueprint package, so starlark.py imports it directly and __init__.py no longer
does. contextlib stays re-exported because this module genuinely uses it.

_read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule
this module imports at the bottom for its route side effects (F811). Renamed to
`settings`, with a comment saying why, since the name is otherwise the obvious
one to reach for.

Verified: pyflakes now reports nothing on this branch that main does not, the
package imports, all nine route modules load, and 117 routes register, matching
the pinned URL-map snapshot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply

Two CodeRabbit findings on the bridge settings endpoint, both of which returned
200 while doing something other than what the caller asked.

`request.get_json(silent=True) or {}` turned a missing or unparseable body --
and the JSON literals null, [] and false -- into an empty dict, which then
satisfied the isinstance(data, dict) guard on the very next line. The guard was
there to reject exactly those bodies. Dropping the `or {}` lets None fail it.

The same `or {}` on /errors/clear is left alone: its docstring documents the
body as optional, so an absent body legitimately means "use the defaults". The
difference is that saving settings has nothing sensible to do with no body.

`if data.get('clear_password'):` accepted any truthy value, and the string
"false" is truthy in Python -- so a client echoing the field back as a string
wiped a password it meant to keep. Now coerced through the package's existing
_coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else.

test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus
a missing body, and clear_password across truthy and falsy spellings. Verified
against the unfixed code -- reverting the body guard fails 5, reverting the
coercion fails 3.

Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials
when TLS is off (CWE-319). That is a policy decision about the feature rather
than a defect -- unencrypted MQTT on a trusted LAN is common and often
deliberate -- so it is raised on the PR for a maintainer call instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: work through the remaining review findings on the editor and bridge

allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses
the network in cleartext. Refused now rather than merely warned about -- but
refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal
deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults
false, and is coerced like the other booleans so the string "false" cannot
switch the guard off.

starlark.py:796 -- the supported service runs Flask threaded, so two start
requests could each see running=False, each launch an editor, and the second
state write replace the first PID, orphaning a process that holds the display
down with nothing recording it. The check-launch-write sequence now takes a
module-level lock.

starlark.py:848 -- if the state write failed the route returned success with an
editor running and no PID recorded: status and stop both reported no session
while the display stayed down until the timeout expired. It now terminates the
process group and returns an error.

starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so
nothing hands the display back, yet the response said "the display is
restarting". After an escalation the display is now restarted explicitly, and a
failure to do so returns an error naming the manual step instead of a success.

pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no
timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the
failure lands before the display is stopped rather than after.

pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface
address are three cases, not two. Collapsing the last two printed a URL saying
"localhost" whenever PIXLET_EDITOR_HOST named a LAN address.

tools.html:1254 -- escHtml does not encode single quotes, and the app id was
interpolated into an inline onclick="startPixletEditor('...')", so a directory
containing an apostrophe could break out of the JS string and run script. The
handler binds with addEventListener and reads the id from dataset, where it is
only ever parsed as an HTML attribute.

Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in
in both directions. The tools DOM suite gains three guards asserting the edit
buttons carry no inline onclick and pass the id via dataset -- those need jsdom
and did not run here, so CI verifies them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): log the traceback on the editor state-write failure

The 848 fix answers 500 when the session state cannot be written, and logged
that at error level -- but without exc_info, so the traceback never reached the
log. test_web_error_detail.py guards exactly this: a handler returning 5xx must
write an error-level record *with* the traceback and return the sanitized
detail, because checking that merely something was logged is too weak.

Caught by Core unit tests on the previous commit, not locally: the guard parses
every module under web_interface/blueprints/api_v3 as one source, so it only
fires once the whole package is read together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 11:32:12 -04:00
ChuckandClaude Opus 5 12f3790994 fix(install): render the systemd units from their templates, not from heredocs (#547)
* fix(install): render the systemd units from their templates, not from heredocs

The installers carried their own inline copies of units that also exist as
templates under systemd/, and the copies drifted.

install_service.sh renders ledmatrix.service from the template correctly, then
wrote ledmatrix-web.service from a heredoc that predated it -- missing
Wants=network-online.target, RestartSec=10, SyslogIdentifier, CacheDirectory,
CacheDirectoryMode and Environment=USE_THREADING=1. install_web_service.sh had
a third copy, and install_wifi_monitor.sh a fourth, that one already differing
from its template (syslog where the template says journal).

startup_validator.py compares the installed unit against the template, so a
rig installed this way warned on every boot -- and the remedy the warning
names, "re-run scripts/install/install_service.sh", reinstalled the same stale
copy. The warning could never clear. Reproduced on a live rig running exactly
that unit.

All three installers now render systemd/*.service through the same placeholder
substitution. The template gains a __USER__ placeholder rather than hardcoding
User=root, because the web interface runs as whoever installed it.

That last point was a second, independent cause of a permanent warning: the
validator substituted a fixed "root", so any non-root install reported drift
forever. It now reads User= from the installed unit -- an install-time
decision, not something the template dictates -- and compares everything else
strictly. first_time_install.sh already reads the installed User= the same way.

Tests cover a non-root web unit not warning, a genuinely changed directive in
that unit still warning, the User= fallback, and a grep-based guard that no
installer under scripts/install/ contains an inline unit body. That guard is
what found the install_wifi_monitor.sh copy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(install): escape sed replacements, use mktemp, and make render failures fatal

Address CodeRabbit findings on install_service.sh, install_web_service.sh and
install_wifi_monitor.sh:

- Values interpolated into each script's sed expression (project root path,
  username) were not escaped, so a value containing &, \ or the | delimiter
  would corrupt the rendered systemd unit. Add a shared
  sed_escape_replacement() helper in the new scripts/install/lib_systemd_render.sh
  (sourced by all three scripts) and apply it to every sed replacement.
- install_service.sh rendered the main and web units to the predictable path
  /tmp/ledmatrix.service.tmp before installing them -- a symlink/TOCTOU race
  (CWE-377). Use mktemp for both, with a trap to clean up on exit.
- install_service.sh treated a missing template as a mere warning and then
  checked only whether a unit already existed at the destination before
  enabling/starting it, so a render failure could silently fall back to
  enabling a stale, previously-installed unit. Both unit blocks now exit
  non-zero on a missing template or a failed render.

Also rename the ambiguous loop variable `l` to `line` in
test/test_systemd_unit_drift.py (Ruff E741); ruff isn't wired into any CI
workflow in this repo today, so this isn't currently CI-blocking, but the
rename is trivial and correct regardless.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5

* test(install): cover sed_escape_replacement against sed-special characters

CodeRabbit asked for regression coverage using a project path containing an
ampersand; the earlier commits on this branch already fixed the escaping,
mktemp usage, and enable/start-on-fatal-render-failure findings, and the
l->line rename was already applied -- this closes the one remaining gap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:19 -04:00
ChuckandClaude Opus 5 2df273ecfc fix(web): default web_display_autostart to true, as the installer already does (#556)
start_web_conditionally.py read the flag with
`config_data.get("web_display_autostart", False)`, so a config that simply
lacked the key got no web interface. Both config/config.template.json and
first_time_install.sh ship the key as true, so the code default contradicted
the shipped default in two places: absence means an older or hand-edited
config, not a request to stay down.

The failure mode was silent in the worst way. The "not starting" path exits 0,
so `systemctl status ledmatrix-web` reported the unit as successfully started
while nothing was listening on the port, and the only trace was one journal
line saying the flag was "false or not set" -- which reads as a deliberate
setting rather than a missing key.

Also start the web interface when config.json is missing or unparseable,
instead of exiting. The web interface is how a config gets created and
repaired, so a broken config is exactly when the user needs it most; leaving
it down means there is no way back in. Only an explicit false/off disables
autostart now, and the disabled message says "explicitly disabled" so the
journal distinguishes a real setting from a default.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 17:42:00 -04:00
ChuckandClaude Opus 5 f8e2e89edc refactor(sports): put the scoreboards on the shared scroll resolver (#542)
* refactor(sports): put the scoreboards on the shared scroll resolver

Eight sports scoreboards -- afl, baseball, basketball, football, hockey,
lacrosse, nrl, soccer -- scrolled through this module's own pacing while the
other eleven scrolling plugins went through src/common/scroll_config. Two
implementations of the same job, and this one was on the losing side of every
difference.

It never called set_scrolling_state. Two consequences, both of which this
release's work was about:

- The frame hold is applied through that call, so a speed the crisp ladder
  could render in whole pixels still presented a new frame every refresh.
- Core only runs deferred updates while nothing is scrolling. Believing
  nothing was, it ran blocking work in the middle of these scrolls.

The default is non-crisp today: scroll_speed 50.0 with scroll_delay 0.01 is
50 px/s, which on a 100Hz panel is half a pixel per refresh. That cannot
render as motion -- it alternates 0px and 1px steps and judders at a 50Hz
beat, on every scoreboard, out of the box. Resolved through the ladder it
stays 50 px/s and holds each frame for two refreshes: same speed, whole-pixel
motion.

The stepping disagreement that used to justify a separate module is gone.
scroll_config avoided frame-based mode because it stepped on a wall clock at
1/scroll_delay with scroll_delay set to the frame period, so the decision sat
on its own threshold and flipped on sub-millisecond jitter. That branch now
accumulates elapsed time, identical arithmetic to the time-based one, so the
two differ only in the units the speed arrives in.

What is NOT shared, and must not be: the two modules read identically-named
keys with different meanings. Here scroll_speed is px/SECOND and scroll_delay
only converts to px/frame; in scroll_config scroll_speed is px per STEP, so
px/s is speed/delay. Passing this module's settings dict to the resolver turns
50 px/s into 5000, clamped to 500 -- a tenfold speed-up everywhere. So
_get_scroll_settings keeps sole ownership of reading sports config, including
the league merging, and hands the resolver a plain px/s. A test pins that
specific number, because it is the mistake the refactor invites.

MIN/MAX_PIXELS_PER_FRAME are gone; the resolver bounds speed and the helper
clamps FPS. _resolve_target_fps stays, re-purposed: under the old model that
key was the rate frames were presented at, so it is the faithful translation
into the refresh the ladder is computed against, used when no hardware
refresh is configured.

Speed changes for panels that are not 100Hz: 50 px/s becomes 60 at 60Hz
(+20%) and 48 at 120Hz (-4%). At 100Hz it is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): drop the frame hold when a scroll times out, not just when it says so

set_scrolling_state(False) clears the hold. The other way a scroll ends is
is_currently_scrolling() deciding, after scroll_inactivity_threshold of
silence, that it is over -- which is what happens when the rotation moves on
mid-scroll or a plugin is torn down. That path cleared the flag and kept the
hold, so every later plugin, scrolling or static, was presented at refresh/N
by whoever scrolled last, until something called the explicit stop.

The method's own docstring already states the rule this breaks: the hold "must
not outlive the scroll that asked for it". The timeout was the exception it
did not cover.

Pre-existing, but reachable by three plugins before and eleven after the
sports scoreboards moved onto the shared resolver, so it belongs with that
change. The test ages the activity timestamp past the threshold rather than
sleeping.

Also adds scripts/sports_scroll_check.py. The sports scroll path is per-league
opt-in, so a rig showing static game cards never constructs a
SportsScrollDisplay and none of its pacing can be observed from a normal run
-- which is exactly what happened when this change was first put on hardware:
26 minutes, zero sports scroll lines. The script drives the path directly with
synthetic games and asserts the three things the resolver is meant to buy: the
speed lands on whole pixels, the hold is published, and it is released after.
It never starts or stops the display service, matching scroll_speeds.py, so a
crash here cannot leave the panel dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): refuse to grab the panel while the display service has it

The module docstring already said to stop ledmatrix first. Nothing enforced
it, and running the script against a live service is not a harmless mistake:
rpi-rgb-led-matrix configures GPIO directions and the hardware PWM inside
RGBMatrix(), and when the root check fails it calls exit() from C with no
cleanup. The service keeps rendering and swapping onto pins that have been
reconfigured underneath it, so the panel goes black while every diagnostic
says the display is healthy -- fresh framebuffer, every pixel lit, "RGB Matrix
initialized successfully", nothing in the log. A restart fixes it, once you
work out that is what happened.

Found the hard way: this is what took the panel down on the test rig, not the
change the script was written to verify.

--fallback skips the check, since it never opens the matrix. --force is there
for anyone who means it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(scripts): annotate the subprocess call the way this repo already does

Codacy fails a PR on one new issue, and bandit B404 fires on any subprocess
import. scripts/run_plugin_tests.py carries the same suppression with the same
justification -- list-form argv, no shell -- so this follows it rather than
inventing a second convention.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:13:04 -04:00
ChuckandClaude Opus 5 968b953a51 fix(display): pin one text layout engine, and give the 5x7 BDF face a size (#539)
* fix(display): pin one text layout engine, and give the 5x7 face a size

Two ways a font could render differently on two machines running the same
code, both found while diagnosing four plugins whose golden images passed on
the machine that generated them and failed everywhere else.

**Layout engine.** `ImageFont.truetype` picks its engine at load time: Raqm
where the host Pillow was built with libraqm, Basic otherwise. The two round
fractional glyph advances differently. `PressStart2P-Regular.ttf` at 8px has
whole-pixel advances, so they agree — which is why most of the fleet matched
everywhere and hid this. `4x6-font.ttf` at 6px does not: glyph positions drift
cumulatively along a run, and the four plugins that draw body text in it
(geochron, of-the-day, christmas-countdown, ledmatrix-weather's almanac) are
exactly the four whose goldens travelled badly.

Every core font load now goes through `src/common/font_layout.load_truetype`,
which pins the Basic engine, so a render depends on the font file and the size
and nothing else. Basic gives up complex-script shaping and kerning pairs;
neither applies to bitmap-grid faces on an LED panel. Output is unchanged on a
host without libraqm.

**Zero font height.** `DisplayManager` built the 5x7 BDF face with
`freetype.Face(path)` and never called `set_char_size`, so `face.size.height`
stayed 0 and `get_font_height()` returned 0 for it — callers stacking rows by
`prev_y + prev_height + gap` drew two lines on top of each other. The
start-up line `Calendar font size: 0 pixels` has been printing the symptom all
along. `font_manager._load_bdf_font` already called `set_char_size`, so
whether measurement worked depended on which path loaded the face.

`DisplayManager` now sets it too, and `get_font_height()` falls back to the
strike the file declares rather than returning a zero line height.

Fixes ChuckBuilds/ledmatrix-plugins#397
Refs ChuckBuilds/ledmatrix-plugins#371, #375, #378, #391

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): give the startup banner a rung that fits a full address at 64px

CI caught what pinning the layout engine exposed rather than caused.
`_fitting_font` walks PressStart2P then 4x6 at 6px, and "255.255.255.255" --
the widest thing the startup banner ever shows -- measures 66px at 4x6/6px
against the 62 a 64x32 panel has to give. It used to squeak in only because
the measurement depended on which layout engine the host Pillow happened to
have; with the engine pinned it does not, so the rung the worst case actually
needs is now in the ladder instead of implied: 4x6 at 5px, which measures 51.

The fallback was wrong in the same place. When nothing in the ladder fit, it
returned `self.font` -- the *widest* option, and precisely how "Initializing"
came to run off the side of a 64px panel to begin with. It returns the
narrowest face that loaded now.

test/test_initializing_screen.py: 34 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): name the exceptions the BDF strike read can raise

Codacy flagged the try/except/pass. It was already narrow in intent -- a
malformed strike table on the measurement path must degrade to "size unknown"
rather than take the display down -- but a bare `except Exception: pass` says
neither of those things and hides a genuinely broken font behind a silent 8px
fallback. It now catches what reading `available_sizes` can actually raise and
logs which face failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: drop logo PNGs the render harness downloaded into the worktree

These are fetched at runtime by the logo cache; they are not source, and they
rode in on a `git add -A` while I was running check_plugin.py against this
branch. Nothing in the change needs them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 09:40:00 -04:00
ChuckandClaude Opus 5 e23f1f45d3 feat(starlark,on-demand): the third-party fixes worth taking, plus a Home Assistant MQTT bridge (#538)
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge

Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of
patches and services built while running this project on Starlark apps
under MQTT control. Its patches are whole-file copies taken against an
older tree, so applying them as written would revert #523's frame
pacing, #534's display() bool returns and the GitHub token masking in
plugins_manager.js. Three of its claimed fixes are already in main, and
its api_v3 Starlark routes are #535's. What follows is the rest --
verified against current code, and reimplemented where the patch's
approach did not hold up.

**On-demand display.** `pinned` reached the controller from the API, was
stored on it and republished in the status payload, but never narrowed
the rotation -- a pinned request still cycled every mode its plugin
owns. Right for a sports plugin, whose modes are views of one subject;
wrong for a plugin whose modes are unrelated, which is every Starlark
app. Now honoured, and it survives a restart.

Restarting while on-demand was active loaded *only* the on-demand
plugin, so normal rotation had nothing to return to for the life of the
process -- and a restart mid-session is routine, since that is how an
update is applied. The panel came back cycling one plugin's modes with
no way out but clearing the cache by hand. Every enabled plugin loads
now; on-demand still resumes on its saved mode.

Stop requests are exempt from the duplicate guards on purpose, so that a
second click stops a mode a race left running -- which means consuming
the mailbox is the only thing that ends one. It was never consumed, so
the same stop was re-read and re-processed on every poll, forever. Both
paths now share one compare-before-delete helper.

**Starlark rendering.** `extract_schema` parsed the source with a regex,
which can only see option lists written out literally: an app whose
dropdown is filled from a live API call inside `get_schema()` came back
empty, and the config form offered nothing to pick. Now runs `pixlet
schema`, which executes the app, and falls back to the parser when
Pixlet is absent, too old for the subcommand, or the app fails to run.
The third-party patch replaced the parser outright and hardcoded
/usr/local/bin/pixlet; this keeps the fallback and the binary search.

A `|` in a config value was dropped by a shell-metacharacter filter,
though the command is a list with no shell involved -- and apps do use
it as a separator inside one value. The key went missing silently and
the app rendered its own "not configured" screen with nothing to say
why. And a 0-byte render was reported as success: Pixlet exits 0 and
writes nothing when an app has no content, which read downstream as a
working app drawing a black panel.

**Starlark display.** `display()` ignored the mode it was called with,
so a specific app could not be addressed. It now accepts `display_mode`
-- which is the whole mechanism, since the controller inspects the
signature before passing it. Found while there: `_select_next_app` ran
only while `current_app` was unset, so with several apps installed the
first was picked once and shown forever while the rest were rendered on
schedule and never displayed. And `enable_scrolling` was missing, so
multi-frame apps were called once per rotation slot and never advanced
past frame one.

**GET /api/v3/display/modes.** Every mode that can be requested
on-demand, with the plugin that owns it. Nothing exposed this, so
anything driving the display from outside the web UI read each plugin's
manifest.json off disk and reimplemented PluginManager's fallbacks. It
also triggers discovery, which is otherwise lazy and normally happens
because a person opened the dashboard.

**integrations/mqtt_bridge.** Home Assistant control over MQTT
Discovery: a mode select, a stop button, power, brightness. Rewritten
against the API rather than the filesystem, so it needs no read access
to config.json and cannot drift from the web UI. paho-mqtt 2.x
VERSION2, TLS, an availability topic that is also the last will, and
secrets from the environment.

**Two opt-in extras.** A DNS single-request unit, for glibc's parallel
A/AAAA lookup stalling ~5s per name on routers that answer only the A
query -- which makes any plugin calling an external API slow and
Starlark apps, which have a render timeout, fail outright. And a Pixlet
config editor: a script you run and Ctrl+C rather than the third-party
version's always-on unauthenticated Flask service, since it stops the
display for the length of a session. Neither is installed by default.

Long Starlark app names now wrap instead of overflowing their card.

115 new tests across 5 files. Also unblocked
test_starlark_display_contract.py, which was silently skipping wherever
fcntl is absent. Whole suite: no new failures against main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(mqtt_bridge): the five issues Codacy flagged on this branch

All in the new bridge, all real:

  * requests floor was 2.31.0, which carries CVE-2024-35195,
    CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which
    is what the project's own requirements.txt already pins.
  * `import time` was never used.
  * `"mqtt_password": None` in DEFAULTS read as a hardcoded credential.
    It is the "no password configured" default; marked nosec B105, the
    convention used elsewhere in the repo.

Also dropped an unused `build_app` from the display-modes test imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: the review findings on this PR

Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and
is answered below.

**One bad config section blanked the whole mode list.**
`/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`,
so a non-dict under a plugin id -- a shape DisplayController already
guards, so it happens -- raised AttributeError mid-loop and answered 500
with no modes at all. Every MQTT bridge entity is built from that list.
Now skipped with a warning.

**The DNS scripts reported success they had not earned.** Three separate
paths: `resolvconf -u` failing was swallowed by `|| true`; the
systemd-resolved branch exited 0 without applying anything, so the
oneshot unit recorded success while the workaround was inactive; and the
installer's `|| echo` turned a failed start into "installation
complete." with exit 0. All three now fail loudly. `single-request` is a
glibc resolv.conf option with no resolved.conf equivalent, so on those
hosts the honest answer is that it cannot be applied.

A NetworkManager-generated resolv.conf is regenerated on connection
changes, not only at boot, and the unit is oneshot with RemainAfterExit
-- so the option can vanish mid-boot with nothing to put it back. Now
detected and stated plainly rather than implied to be permanent.

**`Before=` does not order a manual restart.** It only orders units
already in the same transaction, so `systemctl restart ledmatrix` could
bypass the fix. install_dns_fix.sh now writes a ledmatrix.service
drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround
failing should not stop the display.

**The Pixlet editor's `--lan` is gone.** `pixlet serve` has no
authentication, and a printed warning is not access control. Loopback
only, with the SSH port-forward in the header where the flag used to be
documented -- SSH does the authenticating and nothing is left listening.

**The MQTT example config now defaults to TLS** on 8883. The installer
copies it verbatim, and without TLS the broker password and every
command cross the network in cleartext. A plaintext broker is still
supported and documented, and the bridge warns once at startup when a
password is configured without TLS.

**Not taken: "the upstream Pixlet CLI has no `schema` subcommand."**
Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh`
installs `tronbyt/pixlet`, whose `cmd/schema.go` is
`schema [PATH]` -> JSON on stdout, built on
`runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is
exactly what extract_schema_via_pixlet calls. A binary without the
subcommand exits non-zero and falls back to the source parser, which is
already covered by a test.

**CodeQL stack-trace exposure: not taken either.** I removed `details`
first and that broke
test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception,
which enforces `describe_exception` across all ~75 handlers -- written
because a device with failing storage answered "see logs for details"
from the log viewer itself. describe_exception redacts credentials; the
trade-off is the project's and is already made. Restored, with the
reasoning in a comment.

11 new tests. Whole suite: no new failures against main, 4127 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 08:34:52 -04:00
ChuckandClaude Opus 5 c9289e3a1d fix(store): update_plugin silently did nothing for four installed plugins (#536)
install_plugin() deliberately renames a plugin's directory to the MANIFEST id
when it differs from the REGISTRY id, so registry `stocks` lands in
`ledmatrix-stocks/`. Every lookup in _find_plugin_path() is by directory name,
so update_plugin("stocks") found nothing, logged "Plugin not installed", and
returned False.

Nothing surfaced that to the user. Clicking update in the web UI was a no-op
with no error, and the plugin stayed on a stale version indefinitely. Four
installed plugins hit this on a real device -- leaderboard, music, stocks and
weather -- found because a scripted update of eleven plugins failed on exactly
those four.

Adds a manifest-id scan as the LAST step of the resolution chain, so the two
documented lookups above it (configured dir, then the sibling plugins/
fallback) keep their exact meaning and ordering. That ordering is pinned by
test_discovery_path_contract.py, which characterises the divergence between
the three resolvers on purpose; this extends the chain rather than reordering
it. Directories renamed aside with '.standalone-backup-' during an install or
rollback are skipped, since matching one would report a half-finished install
as a live plugin.

Also adds scripts/audit_render_path.py, which walks the call graph from
display() and reports blocking calls reachable from it. display() runs on the
render thread, so anything slow there stalls the panel; on a vsync-paced loop
a single 15ms call drops a frame and a network round trip freezes the marquee.
Two instances were already found the slow way, by reading frame-time
histograms -- odds-ticker reading the scoreboard cache per frame, and
soccer-scoreboard timing out inside update(). The audit finds that shape in
the source instead. It is a heuristic and says so: a hit behind an interval
check may be fine.

It currently flags 23 calls across six plugins. The clearest is
ledmatrix-music, whose display() falls back to an inline
requests.get(timeout=5) when album art has not been prefetched -- a deliberate
"show the art rather than go blank" tradeoff by its author, but up to five
seconds of frozen panel. Reported, not changed; that is its owner's call.

185 store tests pass. Three of the six new tests fail without the fix.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 15:55:46 -04:00
ChuckandClaude Opus 5 d12323e7f1 perf(scroll): pace frames to the panel — 44→100 fps, stalls 14% → 0.02% (#523)
* perf(scroll): pace frames to the panel, not to a fixed sleep

Scrolling ran at 44-46 fps on a 2x128x64 chain and 14-17% of frames took
41-53ms, which reads as judder. Four independent causes, each measured on
the hardware; details and the diagnostic recipe are in
docs/SCROLL_PERFORMANCE.md.

The high-FPS loop slept a flat 8ms after every render. display() has
already blocked on the panel's vsync by then, so that sleep was added to a
wait that had happened: ~4ms of render plus 8ms put each iteration at ~12ms
against a 10ms refresh grid, so every swap missed a refresh and the loop
settled at 50fps while asking for 125 -- with no headroom, so a further
14% of frames slipped again. It now sleeps only the remainder, with a 1ms
floor so plugin threads still get the GIL.

ScrollHelper stepped position on a wall clock at 1/scroll_delay steps per
second. Plugins set scroll_delay to the frame period, so that comparison
sat exactly on its own threshold: a frame arriving a hair early moved zero
pixels and rendered an identical frame, dirty-tracking skipped the swap, it
returned in ~2ms, and the beat repeated. No scroll_delay value tunes that
out -- a shorter delay trades stalled frames for periodic double-steps.
Both modes now accumulate elapsed time at the same configured speed, so
position stays proportional to real time.

Sub-pixel blending goes back to off by default. It renders a half-step by
mixing two adjacent columns, which on a coarse panel showing pixel-font
text alternates crisp and smeared frames and reads as shimmer -- visibly
worse than integer stepping on the hardware. Vegas mode still opts in.

disk_cache uses orjson when importable, falling back to the stdlib. Encoding
a ~1MB record drops from 14.8ms to 5.4ms end-to-end, and that work holds the
GIL while a marquee is on screen. display_manager also checksummed the whole
framebuffer twice per frame (dirty tracking, then the preview snapshot); the
snapshot now takes the checksum the caller already computed.

New src/common/scroll_config.py resolves scroll settings in one place. Five
ticker plugins each hand-rolled this and disagreed: odds-ticker ranked the
deprecated scroll_pixels_per_second above the documented scroll_speed/delay
pair, and because that key carries a schema default the documented settings
were dead for every user (ChuckBuilds/ledmatrix-plugins#408), while
ledmatrix-leaderboard read the same key only as a fallback. The resolver also
warns when a speed will not advance a whole number of pixels per refresh,
which is the property that actually determines whether a scroll looks smooth.

scripts/build_rgbmatrix_nogil.sh rebuilds the rgbmatrix binding so it
releases the GIL. Upstream declares SwapOnVSync without nogil, unlike
SetPixel/Clear/Fill beside it, so the render thread held the GIL for the
whole vsync wait and starved background threads into long uninterruptible
bursts. The script patches, builds and self-verifies into a scratch tree;
--install backs up the original and rolls back if the service does not come
back healthy.

Measured after: 100 fps locked, no stalls observed, render thread down from
51% to 19% of one core.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): keep the panel swap locked to vsync while scrolling

Dirty tracking skipped SwapOnVSync for byte-identical frames. That is the
right call for static content, but SwapOnVSync is also what paces the render
loop, so skipping it skips the wait for the panel: a duplicate frame returns
in ~8ms instead of ~10ms on a 100Hz panel, advances the strip only 0.8px
instead of 1.0px, and so makes the next frame more likely to repeat as well.
The effect sustains itself once it starts.

Measured over 20 minutes on a 2x128x64 chain, both scrollers configured
identically at 100 px/s:

    leaderboard   10ms x35, 11ms x3            (clean)
    odds-ticker   10ms x26, 8ms x7, 15ms x5    (~20% duplicates mid-scroll)

The duplicates were not end-of-cycle idling -- 38% of fast frames fell within
90s of a scroll completion against 35% of normal frames, a null result. The
trigger is per-frame work: odds does more of it, and more variably, so it is
first to land a frame that advances less than a whole pixel.

Pushing an identical frame costs one canvas copy. Falling out of vsync lock
costs smooth motion. Static content is untouched, because
is_currently_scrolling() expires on its own inactivity threshold -- covered
by test_stale_scrolling_state_stops_forcing_pushes so a plugin that stops
scrolling without saying so cannot pin the panel into always-push.

Also de-flakes test_snapshot_still_written_on_skip, which asserted a strict
mtime increase between two writes that can land in the same filesystem tick;
it failed about two runs in three on Windows regardless of the code under
test. The file is now backdated before the check.

156 tests pass on the Pi. Not yet confirmed by eye on the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): report the frame-time tail, and stop the row-major blit

Two problems, both found by looking at the panel rather than the metric.

The frame-stats line reported ONE instantaneous frame every 5 seconds --
about 1 frame in 500 -- printed beside a 100-frame average. Both hide exactly
the fault they are used to chase: a 2ms duplicate and a 21ms double-wait
average to precisely 10ms, so a ticker stalling on half its frames still
reports a healthy "Avg FPS: 100.0". That reading cost several rounds of
chasing the wrong layer. The line now aggregates every frame since the last
log and reports median, p95, max, min, and explicit stall and skip rates
(past 1.5x the median missed a refresh; under half never reached the panel,
because dirty tracking skipped the swap so the frame never waited on vsync).

On the hardware this now reads:

    leaderboard  100.0 fps over 501 frames | median 10.00ms p95 10.05ms
                 max 10.34ms | stalls 0 (0.0%) skips 0 (0.0%)

The binding rebuild's blit patch becomes opt-in (RGB_PATCH_BLIT=1, default
off). Reordering that loop to row-major changes what a torn frame looks like:
column-major tearing shows as a vertical seam, row-major as a horizontal split
between the panel's upper and lower halves. On a 1/32 scan panel that reads as
a one-pixel fold across the middle of every panel, which is what was reported
on hardware and what went away when the blit was reverted. All of the measured
gain comes from the SwapOnVSync change, so the risky half is simply not worth
taking; the header says so.

Also fixes --install resolving its paths against $HOME, which is /root under
sudo, so it looked in /root/rgbmatrix-nogil-build and died with "no built
module found" on a machine where the build had just succeeded. It now resolves
SUDO_USER's home. Both build paths are verified on the Pi: default yields one
GIL-release site, RGB_PATCH_BLIT=1 yields two.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(scroll): let users pick a crisp speed for their own panel

Whole-pixel motion was previously only available at multiples of the refresh
rate -- 100, 200, 300 px/s on a 100Hz panel. 100 px/s crosses a 256px panel in
2.6s, which is brisk for reading, and everything slower had to blend (blur) or
repeat frames unevenly (judder). There was no way to ask for 50 px/s and get
clean motion.

SwapOnVSync takes a framerate_fraction the display manager never passed. It
holds each frame for N panel refreshes; the panel keeps refreshing at its full
rate throughout, so holding costs nothing in flicker and only changes how often
a NEW image is presented. That turns 50 px/s into one whole pixel every second
refresh instead of half a pixel every refresh.

The crisp speeds are therefore refresh_hz / hold * pixels_per_frame, and that
ladder depends on the panel: a Pi Zero on a long chain has a different set of
good speeds from a Pi 4 on a short one. crisp_ladder() enumerates them and
solve_crisp() picks the best match for a requested speed.

solve_crisp weights motion quality rather than picking the numerically nearest
entry, which matters more than it sounds. Asked for 30 px/s, nearest-by-value
answers 28.6 -- 2px jumps at 14fps -- over 33.3, which is single-pixel motion
at 33fps and obviously better on the panel. The target is also clamped into the
ladder's range first, because relative error saturates near 1.0 for a target
far outside it and the quality penalty would otherwise answer "10000 px/s" with
the slowest entry.

configure() snaps to the ladder and applies the hold when given a display
manager. Without one the hold silently cannot happen and motion falls back to
fractional pixels, so it warns rather than failing quietly. set_frame_hold()
resets to 1 when scrolling stops, so one plugin's pacing cannot leak into
whatever is on screen next.

scripts/scroll_speeds.py is the user-facing part: it prints the ladder for the
configured rate, measures what the panel ACTUALLY manages (--measure, for
hardware that cannot reach its configured limit), highlights the nearest option
to a wanted speed, and demos one live. It never starts or stops the display
service itself -- doing that inside a script stranded the panel twice today.

Speeds below ~20 px/s remain stepped regardless. That is the pixel pitch, not a
software limit.

Also fixes the dirty-tracking test spy, which stubbed SwapOnVSync with a
single-argument function and would have masked the new call as a failed push,
and rewrites a configure() test that had started passing for the wrong reason:
it asserted a judder warning, which snapping now prevents, and was matching the
unrelated "hold could not be applied" warning instead.

183 tests pass on the Pi.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): tie the frame hold to the scroll, not the plugin

The hold applied in configure() never reached the panel. Plugins share one
display manager, and set_scrolling_state(False) -- fired whenever ANY other
plugin finishes its scroll -- reset the hold to 1. A hold set once at plugin
construction was therefore always gone by the time that plugin rendered.

The symptom was a log line that lied. ledmatrix-stocks reported

    Scroll configured: 50.0 px/s (1px every 2 refreshes = 50.0 fps, smooth)

while the panel measured 100.0 fps, median 10.00ms. Config, resolution and
snapping were all correct; only the pacing silently was not applied.

set_scrolling_state(is_scrolling, frame_hold=1) now carries it, so the hold
lives exactly as long as the scroll that asked for it. configure() reports the
value as ScrollSettings.frame_hold instead of applying it -- applying it behind
the caller's back could never have been right on a shared display manager.
Existing callers are unaffected; the default keeps one frame per refresh.

Verified on hardware: stocks at 50 px/s now measures

    50.0 fps over 251 frames | median 20.00ms p95 20.09ms | stalls 0 skips 0

20.00ms being exactly two refreshes, with the panel still refreshing at 100Hz
underneath so flicker is unchanged.

test_another_plugin_stopping_does_not_strand_a_hold pins the interaction that
broke this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll,cache): resolve CodeRabbit review on #523

Eight findings, all reproduced before fixing.

scroll_config.configure() read the refresh rate *after* resolve() had
already used it. resolve() fills in target_fps, pixels_per_frame and the
judder warning from that rate, so on a 60Hz panel every one of them
described 100Hz -- and with snap_to_crisp=False nothing downstream
corrected it, so set_target_fps() paced the helper to 100 FPS. The rate
is now settled first, and falls back to the global config rather than
straight to the default.

refresh_hz_from_config() used `(cfg.get("display") or {}).get(...)`,
which raises AttributeError when either level is truthy but not a
mapping -- out of a function whose whole contract is a rate or a default.

The frame-stats line reported the upper-middle sample as the median and
the 96th sorted sample as p95 of 100. Both are also thresholds (stalls
at 1.5x the median, skips at 0.5x), so the counts were biased too. The
arithmetic is now in frame_stats()/format_frame_stats(), testable
without a clock.

configure()'s docstring and docs/SCROLL_PERFORMANCE.md still said it
applies the frame hold and warns when it cannot. It deliberately does
neither since "tie the frame hold to the scroll, not the plugin"; a
caller following the old text would omit set_scrolling_state() and slow
snapped speeds would still present every refresh.

disk_cache had no policy for non-finite floats: orjson writes null,
the stdlib writes NaN/Infinity, and orjson then rejects those legacy
files so DiskCache.get deleted them as corrupt. One behaviour on both
paths now -- write null, keep legacy records readable. allow_nan=False
detects the values; the replacement walk runs only when there is one,
so the ordinary write path is byte-identical and pays nothing.

build_rgbmatrix_nogil.sh picked the build artifact with a glob piped to
`head -1`, which sorts cpython-311 ahead of cpython-313, so a stale .so
staged in from the source tree was installed as core.so while the GIL
check -- which reads the generated core.cpp, not the .so -- still passed.
It now requires the current interpreter's exact ABI name and fails
closed. Its systemctl calls were also unchecked under `set -uo pipefail`:
a failed stop left the old service running, the following start
succeeded as a no-op, and the health check reported SUCCESS for a
binding that was never loaded.

orjson floor raised to 3.11.6 for CVE-2025-67221 (unbounded recursion
in dumps); it covers the project's Python 3.10-3.13 range.

Adds test/test_cache_nonfinite_floats.py (14) plus regression tests in
test_scroll_config.py and test_scroll_helper.py. 9 of the cache tests
and 9 of the scroll_config tests fail against the pre-fix code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(harness): keep the visual double's signature tied to production

Moves set_scrolling_state's frame_hold into the test double here, where
DisplayManager gains it, rather than in #534 where it arrived a PR early.
CodeRabbit flagged the #534 version correctly: a double that accepts an
argument production does not lets the call pass every harness run and
raise TypeError on the panel, which is the one failure a safety harness
exists to prevent.

The drift has now gone both ways across two branches -- double behind
production on this branch, double ahead of it on #534 -- so it is pinned
instead of remembered. test_display_double_parity.py compares the two
signatures and fails with the direction of the drift named. It reads the
files with ast rather than importing them, because display_manager
imports rgbmatrix at module scope and this check should hold on a laptop
and in CI as well as on a Pi.

Plugins begin passing frame_hold in ledmatrix-plugins#462, which is why
production and the double both need it before that lands.

Full suite: 3889 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 13:37:54 -04:00
ChuckandClaude Opus 5 a0d3e64099 fix: ten defects found validating the whole plugin fleet on hardware (#534)
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers

Three independent fixes found while validating every plugin on a 256x64 rig.

FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf
ships with the core, so every plugin offering it logged "Font family
'tom_thumb' not found" (16 warnings per countdown render) and had to carry a
private loader to use a bundled font. Closes #524.

VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter
that DisplayManager gained, so any plugin passing it died with TypeError at
render time and failed every size. Nine plugins now make that call;
ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other
seven only passed because their scroll path was unreachable without data.
Closes #525.

api_v3 declared module-level config_manager/plugin_manager = None that nothing
ever assigned -- app.py sets the blueprint attributes, which the other 150+
call sites use. Three sites read the decoys, so /health reported the config
unreadable and the plugin system uninitialised (making "degraded" permanent and
unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on
every rig. The decoys are removed rather than assigned, so a bare name is now a
NameError at test time instead of a silent None. The same function's first-call
uptime was computed from two separate clock reads and came out negative.
Closes #529.

Verified on the rig: both previously-failing plugins render, the tom_thumb
warnings are gone, /health reports "healthy" with all three checks passing, and
/display/current reports the real 256x64.

* fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark

The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is
world-writable and sticky, and the display service runs as a different user
from the tooling, so a leftover temp owned by anyone else became unopenable
even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a
sticky directory. The preview and the health check's liveness proxy then froze
until someone deleted the file by hand; on the test rig that meant 23 hours of
a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with
cleanup on failure, matching the hardware-status write a few hundred lines
above. Closes #528.

_poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults
the in-memory TTL to max_age -- so the first request was pinned in memory for an
hour and every later poll returned that stale copy. No second on-demand request
was honoured until the service restarted, while the API kept returning 200.
get() already documents memory_ttl=0 for exactly this cross-process case.
The consumed request is also now deleted: leaving it on disk meant a restart
replayed the previous request, activated it, and ignored the one the caller had
just made. Closes #530.

starlark-apps returned None from display() when it has no app to show, which is
the state of every install without Pixlet and of a fresh one before any app is
added. The controller only skips on a boolean False, so that held a black panel
for the full display_duration instead of rotating on. Closes #456 (core side).

Verified on the rig: two consecutive on-demand requests with no restart between
them are both activated, where the second was previously dropped in silence.

* perf(harness): share one cache across a plugin's renders

_instantiate built a fresh MockCacheManager for every (size, mode), and that
mock is a per-instance in-memory dict, so each render was a cold start. A plugin
that fetches per game or per player re-fetched everything N times over --
baseball-scoreboard at one size took 840s for nine renders where the arithmetic
said ~72s, and at eight sizes it exceeded a 900s timeout.

The second and later renders also never exercised the cache-hit path, which is
what a running rig executes almost all of the time, so a caching regression
could not be caught here.

The cache is now built once per render_plugin_matrix call and threaded down.
The display manager stays per-render -- the bounds checking depends on that --
so only fetched data is shared.

Measured on the rig, same render counts and same goldens:
  tide-display         2s -> 1s   (32 renders)
  cricket-scoreboard  10s -> 3s   (24 renders)
No pass/fail change across tide-display, cricket-scoreboard, clock-simple,
geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages.

Closes #533.

* fix(scripts): run standalone plugin tests instead of collecting nothing

run_plugin_tests.py discovered every plugin test file and handed the lot to
pytest. Most plugin tests are standalone scripts -- module-level main() plus an
`if __name__ == "__main__"` guard, signalling through an exit code -- and pytest
collects zero items from those. The run printed how many files it had *found*,
then "no tests ran", and exited without executing any of them. On a rig with all
44 first-party plugins that is 151 of 248 files.

Files are now classified and each kind runs under the right runner: pytest for
real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1
fail convention ledmatrix-plugins' own runner established (a script that wants a
tty or an LED matrix is a skip, not a regression).

Before:
    $ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos
    Found 1 test file(s)
    collected 0 items
    no tests ran in 0.31s                      rc=0

After:
    Found 1 test file(s) -- 0 collectable, 1 standalone script(s)
    1 passed, 0 skipped, 0 failed (scripts)    rc=0

Verified across three shapes: countdown (1 script), jellyfin-now-playing and
pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files
split 4 collectable / 7 scripts, all seven of which had never run).

Closes #532.

Running the flights scripts for the first time also surfaced four genuinely
failing tests there, hidden by the mirror-image bug in the plugins repo's own
runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465.

* fix(harness): give an empty-looking mode a few frames before warning about it

check_plugin's "drew nothing but display() returned X" warning fired on a single
frame, rendered with force_clear=True, under a frozen clock. All three defeat a
scrolling plugin, whose first frame is legitimately its blank scroll-in buffer.
Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which
people stop reading a warning, which matters because the true positives are
real: a mode that draws nothing and does not return False holds a blank panel
for its whole display duration.

An apparently-empty frame is now re-driven for up to 48 more frames with
force_clear=False (force_clear means "reset the scroll", so repeating it would
redraw frame 1 for ever) and with the clock advancing -- freezegun's factory
where time is frozen, a real sleep where it is not, since scroll position is
usually a function of elapsed time. The first frame that draws content replaces
the result.

The clock is moved back afterwards. It is shared by every render in the matrix,
so time borrowed by the probe leaked into later modes and drifted their goldens
-- f1_upcoming picked up 5 spurious drifts before this was restored.

Measured on the rig:

                        empty warns          check
                        before  after
  f1-scoreboard            42      0    48 PASS / 0 FAIL, goldens intact
  ledmatrix-elections      16      0    16 PASS / 0 FAIL
  on-air                    8      8    true positive, kept
  nfl-draft                 8      8    true positive, kept
  clock-simple/geochron/    0      0    unchanged
  christmas-countdown

58 false positives gone, both true positives kept, no golden regressions. Cost
is confined to modes that really are blank: plugins that draw immediately are
unchanged (clock-simple and tide-display still 2s), while on-air -- eight
deliberately blank modes -- goes to 21s.

Closes #527.

* fix(harness): load nested schema defaults, and merge caller config at leaf level

load_config_defaults read only top-level properties. An object property carries
its defaults on its children, not on itself, so everything nested was dropped --
2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of
565. render_plugin_matrix's comment says the plugin then "behaves like a real
install", which for most of the fleet it did not.

_defaults_from_properties now recurses. merge_config deep-merges the caller's
config onto the result so an override lands at the leaf: a shallow merge would
let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard
every other nhl default, which is the same class of bug being fixed here.

Measured before/after across all 49 installed plugins on the rig: **no render
changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact.
Plugins already fall back to the same values internally via config.get(key,
default), so supplying them explicitly agrees with what they were doing. The
defaults really are arriving now:

  ufc-scoreboard        9 -> 87 defaults
  ledmatrix-flights    51 -> 95
  masters-tournament   10 -> 51
  cricket-scoreboard   22 -> 50
  tide-display         12 -> 18

and hockey-scoreboard, which used to load nhl.enabled=None, now gets
nhl.enabled=True with its full display_modes block.

Caveat worth carrying: the eight plugins with the most nested config
(soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of
the 2,386 dropped defaults, 68%) could not be measured. They import
src.common.sports_shared, which the test rig's core branch predates, so they
fail to load there identically before and after. Re-run this comparison against
a core that has that module before trusting the "nothing changed" result for
them; those are exactly the plugins whose renders should change most.

Closes #531.

* refactor: narrow the exception handlers this branch introduced

Codacy flagged the new code; it passes on other recent PRs, so the finding is
mine. Four of the five broad `except Exception` clauses I added were catching
far more than they needed to, which is the same shape as several bugs this
branch fixes -- hello-world's TypeError sat invisible for exactly this reason.

  freezer() / move_to() / tick()   -> (AttributeError, TypeError, ValueError)
  cache_manager.delete()           -> (OSError, AttributeError, KeyError)

The fifth stays broad and now says why: it wraps a call into a plugin's own
display(), which can raise anything, and the first frame has already rendered --
so a failure there must not turn a good result into an error.

Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty
warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows
7 golden drifts both before and after this branch, so it is not from these
changes -- its committed goldens predate #521's 1-bit text rendering.

* fix: resolve CodeRabbit review and Codacy findings on #534

CodeRabbit raised six; all six were real.

The test double had drifted ahead of production. VisualTestDisplayManager
accepted set_scrolling_state(frame_hold=...) while DisplayManager did not,
so such a call passed every harness run and would raise TypeError on the
panel -- the one failure a safety harness exists to prevent. frame_hold
belongs to the change that adds it to DisplayManager (#523), so it moves
there and the double matches main again.

The harness swallowed exceptions from re-rendered frames. _settle_loop
re-renders a mode that came back blank, to give a scroll time to draw;
returning silently on a crash meant a mode that renders one good frame
and then explodes was reported as passing. Recorded on result.error now,
keeping the captured frame so the failure stays inspectable.

starlark-apps display() returned True after _display_frame() failed, so
the controller held a dead frame for the whole display_duration instead
of rotating on. _display_frame now returns bool on all three paths.

run_plugin_tests.py used env.setdefault for PYTHONPATH and
LEDMATRIX_CORE, so an inherited value won and the subprocess imported a
different core than the one under test -- ledmatrix-plugins#467 exactly.
Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally.

The on-demand mailbox is polled after every frame, ~125x/second on a
scrolling mode, and the read is deliberately uncached, so it was that
many disk reads per second to find nothing. Floored at 250ms, which is
imperceptible for a web-UI click. Consuming it also deleted whatever was
present rather than what had just been processed, so a request posted
while the previous one was in flight was thrown away and never ran; the
delete is now keyed by request_id. That narrows the window rather than
closing it -- a true atomic claim needs a primitive the cache layer does
not offer, and the code says so rather than implying otherwise.

Codacy's 2 criticals were bandit B404/B603 on the subprocess call added
to run_plugin_tests.py. Fixed interpreter, argument list, no shell;
annotated with the repo's existing nosec convention. Bandit is clean on
the file.

Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py
(4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of
those fail against the pre-fix code. Full suite: 3961 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: satisfy Codacy's subprocess checks on the new test runner

Codacy runs Bandit and Opengrep (its Semgrep fork). The new
subprocess.run in scripts/run_plugin_tests.py trips three patterns, on
two different lines:

  Bandit   B404 on the import, B603 on the call
  Opengrep dangerous-subprocess-use-audit          on the run( line
           dangerous-subprocess-use-tainted-env-args on the argv line

A nosemgrep applies only to its own line, so the call line and the argv
line each need one; a single comment on the call covered neither rule
fully. Suppression is the right answer here rather than a rewrite: the
interpreter is sys.executable, the arguments are a list, and no shell is
involved, so there is nothing to word-split or expand.

Matches the pair the rest of the repo already uses for this shape --
permission_utils.py, plugin_loader.py, install_dependencies_apt.py.

Codacy: 0 new issues, up to standards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: leave visual_display_manager untouched so #523 can merge

The only change this branch made to that file was a docstring, and it
collided with #523's rewrite of the same method -- so #534 and #523 each
merged cleanly against main but conflicted with each other. Reverted to
main's text; #523 owns this method and adds frame_hold to it.

The note the docstring carried ('frame_hold arrives in #523') would have
been stale the moment #523 landed anyway. The parity test in #523 is
what actually keeps the two signatures honest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: add the Ruff suppression nosec/nosemgrep do not cover

Ruff reports S603 on the same call Bandit and Opengrep do, and none of
the three suppressions covers the others. Confirmed the precondition
first: path comes from discover_plugin_tests(), which globs test files
inside the repo, and the call is a fixed interpreter with a list argv
and no shell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 13:37:38 -04:00
ChuckClaude Opus 5coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
696acdbc7b feat(render_plugin): add --display-mode so multi-mode plugins can be rendered (#522)
* feat(render_plugin): add --display-mode so multi-mode plugins can be rendered

render_plugin.py always called plugin.display(force_clear=True) with no mode.
A plugin that declares one display mode is fine, but the sports scoreboards
declare three or more and keep their per-mode state on sub-managers; their
no-argument path selects nothing and returns False, so the render came out
blank with nothing to say why. Measured on nrl-scoreboard with identical
seeded state:

  live.display() directly                 True,  1892 lit pixels
  plugin.display(display_mode="nrl_live") True,  1892 lit pixels
  plugin.display()                        False,    0 lit pixels

--display-mode passes the requested mode through. It is only passed when
asked for, so the many plugins whose display() takes no display_mode keep
working untouched, and a plugin that declares modes but does not accept the
argument degrades to its default screen with a warning rather than a
TypeError.

This is what lets the plugin READMEs show a scoreboard at all, and it also
unblocks screens like birdnet_stats and the weather plugin's hourly, daily and
almanac modes, which could previously only be described in prose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Only fall back when plugin display rejects display_mode

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-09-06 17:28:22 -04:00
ChuckandClaude Opus 5 9c0c0dc851 consolidate(web): credential exposure, secret loss, and the update path (#485)
* fix(web): stop /config/main handing out every credential it holds

The endpoint returned the raw config to anyone who could reach the port, and
this web interface has no authentication of any kind. An unauthenticated
request against a live rig returned:

    github.api_token                40 chars
    incoming-packages.ha_token     183 chars
    jellyfin-now-playing.api_key    32 chars
    ledmatrix-weather.api_key       32 chars
    on-air.mqtt_password             8 chars
    youtube.api_key                 20 chars
    youtube-stats.api_key           39 chars

A GitHub token and a Home Assistant long-lived token among them. Anything on
that LAN could read them.

The x-secret masking the plugin config endpoints use does not reach here: this
route never consults a schema, and core keys such as github.api_token have no
schema to carry the marker. Several of the fields above *are* tagged x-secret
in their plugin's schema and were still returned in full, which is what rules
out the schema route as the fix for this endpoint.

Credential-named fields are now blanked. Matching on the name is blunt, and
for a whole-config dump that is the right default: anything named like a
credential should not leave the process, and a new plugin adding a
differently-shaped secret is covered without anyone remembering to tag it.

Blanked rather than removed, and safe to blank: POST /config/main merges into
the freshly loaded config and writes only the keys it was given, so a client
that round-trips this response cannot erase a secret it never saw. The web API
suites confirm it -- 81 passing, unchanged.

On the test that matters: the first version of this suite exercised the two
helpers and nothing else, and reverting the single line that wires the
redactor into the route passed all thirty of them. A property asserted on a
helper is not a property asserted on the endpoint, and it is the endpoint that
is exposed to the network. The added test goes through the view function, and
it does fail on that revert.

This also corrects an earlier claim of mine. I reported that GET /api/v3/config
did not expose these values; that path 404s, so the check proved nothing. The
real route is /config/main and it exposed all of them.

* fix(web): stop an unrelated config edit from erasing a plugin's secret

Saving any field on a plugin's config form destroyed that plugin's stored
credential. On a rig with a weather API key, changing the city silently
emptied the key, and the plugin stopped working at the next fetch with no
indication why.

The path had no guard at any step. The config partial masks secrets before
rendering (pages_v3.py:740), so the browser posts them back blank; _parse_value
deliberately preserves "" for optional string fields; separate_secrets routes
that "" into secrets_config, which is a truthy dict; deep_merge writes it over
the stored value; save_raw_file_content persists it.

The blank does not even need the round-trip. merge_with_defaults injects the
schema's api_key default ("") into every save, so a client that never sends
the field at all still erases it. test_secret_count_message_counts_top_level_keys
was counting exactly that injected blank as a saved secret field -- the visible
edge of the bug, pinned as expected behaviour.

remove_empty_secrets() already existed for this, with seven unit tests and a
docstring describing this precise scenario ("clients will send those empty
strings back ... so that existing stored secrets are not overwritten with
blanks"). It was never wired into a call site. This wires it into both save
paths that merge into the secrets file.

A blank now means "unchanged" rather than "delete", which is the same contract
the helper's tests already describe. The cost is that a secret can no longer be
cleared by emptying the field; clearing needs its own affordance, since a
control that erases credentials as a side effect of ordinary edits is not one.

Verified by reverting the guard: the new round-trip test then fails with the
stored key read back as ''. 262 web tests pass with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop dumping the config and request headers to the journal

save_main_config logged its entire POST body and the full request headers at
ERROR on every save. The body is the configuration itself, and the headers
carry the session cookie, so a routine settings change wrote both to the
journal -- at a level that guarantees they survive any sane log filter.

The lines are leftover debug output: they say "DEBUG:" in the message while
calling logging.error, and they went through the root logger rather than the
module logger, bypassing the level configured for this blueprint.

Replaced with a debug-level line recording the shape of the request, which is
the part with diagnostic value. The local `import logging` went with them; it
shadowed a module-level import that was already there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop /config/secrets handing out every credential it holds

GET /api/v3/config/secrets returned config_secrets.json in full to anyone who
could reach the port, and this interface has no authentication. Probed against
a real rig it produced six populated credential fields: a 40-character GitHub
token, a 183-character Home Assistant token, and Jellyfin and weather API keys.
This is the second door onto the same credentials; #477 closes the first.

Masking the response alone would have been worse than the leak. The only
client fetches every secret, edits one field and posts all of them back, and
save_raw_file_content replaces the file wholesale -- so a masked GET followed
by the client's own save would write the mask over every credential the user
had not touched. That is why this was left open when the leak was found; it
needs both halves.

Read side: mask_all_secret_values(), which already existed for exactly this
endpoint -- its docstring names it -- and had never been wired to a call site.
It leaves empty values and YOUR_* placeholders alone, so a client can still
tell "set" from "not set" without being told the secret.

Write side: strip the echoed mask and blanks from the submission, then merge
onto what is stored, so "unchanged" means unchanged. The cost is that a secret
can no longer be cleared by blanking it; that wants its own affordance, since
a control that erases credentials as a side effect of saving an unrelated one
is not one.

Browser side: the token field is now left empty rather than filled from the
response. Filling it with the mask would have stored eight bullet characters
as the token the next time the user pressed Save, and filling it with the real
value is the thing being fixed. It reports whether a token is saved instead.

Verified end to end through the Flask endpoints, not the helpers. Reverting
the masking fails the leak tests; reverting the merge fails the preservation
tests; both halves are independently guarded. 278 web tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop reporting "no update" when the update check could not run

check-update returned update_available=False whenever git failed. The banner
is the only route to the update button, so a checkout git refuses to touch
looked exactly like a current one -- permanently, with nothing on screen to
act on and only a log line recording why.

The common cause is an install performed as root. scripts/install/one-shot-install.sh
clones into ${HOME}/LEDMatrix, never consults SUDO_USER, and contains no chown
at all, while its own error text suggests running the whole thing under sudo.
The result is a root-owned checkout, and on a rig this is what every git
command in it does:

    fatal: detected dubious ownership in repository at '...'

including the fetch this endpoint runs. Verified on real hardware rather than
assumed.

A failed check now reports check_failed with a message the user can act on --
for dubious ownership, the chown that fixes it. The banner shows that message
instead of hiding itself, with the update button suppressed since updating
cannot work until the cause is fixed. The success path is untouched.

This does not fix the installer, which is the real cause; it stops the symptom
being invisible. The installer needs SUDO_USER handling and a chown, and its
suggestion to run as root should go.

Reverting the endpoint change fails four of the five new tests; the fifth
guards the success path and correctly does not move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop the installer chmod stripping exec bits on every update

git tracks five scripts as mode 644 that first_time_install.sh then chmods to
755 (start_display.sh, stop_display.sh, the two install_*_service.sh, and
one-shot-install.sh does the same to first_time_install.sh). With
core.fileMode true, the default on Linux, git reports all five as modified
from then on, in files the user never touched.

The update button stashes local changes before pulling, so it is not blocked
by this. But it never pops that stash -- stash pop and stash apply appear
nowhere in the update flow -- so the mode change is stashed away and left
there, and the files revert:

    === file modes after the update button's stash ===
      664  first_time_install.sh      <- installer had made these 755
      664  start_display.sh
      664  stop_display.sh
      664  scripts/install/install_service.sh

So every web-UI update silently strips the executable bit from the installer's
own scripts, and leaves a stash entry holding the difference. start_display.sh
and stop_display.sh stop working from the shell afterwards.

A manual `git pull --rebase` over SSH fails outright, since nothing stashes for
it: "cannot pull with rebase: You have unstaged changes". That is the likely
source of the reports, since plenty of people update that way.

Tracking the five as 755 -- what they should always have been, as the
installer chmodding them attests -- removes the spurious mode change
entirely: nothing to stash, nothing stripped, no stash entry, and manual
pulls work.

The pull also passes --autostash, for the case the code explicitly tolerates:
when the stash fails it logs a warning and pulls anyway, and that pull is what
then fails. Autostash also pops what it stashes, which the manual stash does
not.

Note that `git add -A` after `git update-index --chmod=+x` silently reverts
the index to the on-disk mode, so the modes here were set by chmodding the
files themselves.

Regression test asserts the five stay tracked executable; reverting any one
of them fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): ask for the restart that makes an update take effect

The update button pulls new code and restarts nothing. There is no systemctl,
restart, reload or reboot anywhere in the 172-line git_pull handler -- it
stashes, pulls, installs changed requirements, re-removes plugins the user had
uninstalled, and returns "Code updated successfully."

Meanwhile both services go on running the code they loaded at boot. So the
display keeps rendering the old build, the web interface keeps serving the old
build, and the user is told the update worked. Nothing on screen suggests
otherwise, and the next reboot is what actually applies it -- whenever that is.

The affordance for this already exists: the restart-pending banner, raised
after main-config saves, with a Restart Now button wired to the display
service. A code update is a stronger reason to show it than a config save is.

The response now reports restart_required, and applyUpdate raises the banner
with wording for a code update rather than a config save. The banner's message
became a parameter and is persisted next to the flag, since it outlives the
page that raised it.

restart_required is only true when the pull actually moved HEAD. "Already up
to date" is a success too, and prompting after a no-op would train users to
dismiss the prompt unread.

This covers the display service, which is what the Restart Now button drives
and what users notice. The web interface still picks up its own new code on
its next restart; restarting it from inside a request it is serving is a
larger change than this one.

Reverting the flag fails the test that a pull which moved HEAD asks for a
restart. 290 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* Mask list-shaped secrets element-wise, close two vacuous tests

mask_all_secret_values treated any non-empty list as a scalar, so a
secrets file holding

  "accounts": [{"name": "a", "token": "tok-a"}, {...}]

came back as a single "••••••••". The caller could not see how many
entries existed, and the raw editor was handed a string where the file
holds an array. Recurse into lists in both _mask_value and _contains_mask.

Lists merge by replacement, not key-wise, so strip_masked_values now
drops a list outright if any element still carries the mask -- storing a
half-masked list would discard the untouched entries.

Two tests could pass without exercising what they claim to check:

- test_git_pull_resolution asserted modes only for paths git ls-files
  returned. A renamed or deleted installer target is simply absent from
  that output, so its mode was never checked. Assert every CHMODDED path
  is tracked first.
- test_config_secrets_masking never checked the POST status. A 500
  leaves the old file in place, which satisfies every assertion that
  follows. Assert 200 before reading the file back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 14:30:29 -04:00
ChuckandClaude Opus 5 71739d85d1 harden: cap malloc arenas, warn on unit drift, and grant the portal's sudo commands (#476)
* perf(systemd): cap glibc malloc arenas on the display service

Measured on a live rig 2.5 hours after start:

    RSS                          1030 MB
    Private_Dirty                 988 MB
    anonymous mappings > 10 MB       23
    largest        104, 79, 66, 63, 63 MB, on 64 MB-aligned addresses
    threads                           9
    cores                             3   -> glibc ceiling = 8 x 3 = 24 arenas

23 against a ceiling of 24, all 64 MB-aligned: these are glibc's per-thread
malloc arenas, not live objects. The data the process was actually holding
accounts for perhaps 15 MB -- the widest scroll strip observed was 35,746 x 64,
about 7 MB as RGB and the same again for its numpy mirror.

It is bloat rather than a leak: sampled four times over 135 seconds, RSS sat
between 990 and 1030 MB rather than climbing. glibc gives each allocating
thread its own arena, grows them to hold peak demand, and never gives them
back. A process that builds and drops large images across several threads is
exactly the shape that produces this.

The device had 59 MB free at the time, on 1845 MB total.

MALLOC_ARENA_MAX=2 trades a little allocator concurrency for that resident
memory. It is a tuning knob rather than a fix for a defect, so the rationale
and the measurements sit next to it in the unit file, and a test asserts they
stay there -- a bare environment variable invites removal by whoever meets it
next.

Two things this is NOT, both checked rather than assumed:

- Not an OOM problem today. A grep for "oom" in the service journal returned
  24 matches, all of which were the radar logging zoom=9 and zoom=7. The kernel
  OOM killer has not fired: dmesg has zero matches.
- Not currently capped by the unit's MemoryMax=85% either. That directive is in
  this file but absent from the unit actually installed on the rig, which
  reports MemoryMax=infinity, so nothing is enforcing a ceiling there.

The saving is unmeasured on hardware: applying it needs a service restart,
which blanks the panel, so that is the user's call rather than something to do
mid-audit. If p99 frame time regresses -- it sits at 18.4 ms against a 16.7 ms
budget for 60 FPS, so there is not much headroom -- raise the value rather than
remove it.

(cherry picked from commit 446207ffbc)

* test(systemd): pin the arena value instead of accepting a range

Review follow-up. The range check accepted 1, 3 and 4, so a change to 4 --
which hands most of the resident saving back -- passed a test whose whole
purpose is to notice that.

Pinned to the value the unit ships, in one named constant. Raising it is still
a legitimate response to a frame-time regression, but it should be a visible
edit here rather than silent drift, and the failure message says so.

Mutation-checked: changing the unit to 4 now fails.
(cherry picked from commit 73fff8d2d5)

* fix(startup): warn when an installed systemd unit has drifted from the repo's

Nothing re-applies systemd units after the first install. `git pull` -- which
is what the web UI's update button runs -- brings a new template into the
checkout, but no code in web_interface/ or src/ copies it to
/etc/systemd/system, and nothing anywhere runs `systemctl daemon-reload`. The
unit that actually runs is whatever first_time_install.sh wrote on day one.

So every hardening added to a unit is inert on existing installs, silently.
Measured on a live rig:

    installed  /etc/systemd/system/ledmatrix.service   2026-08-06
    template   systemd/ledmatrix.service               2026-08-19
    contents                                           differ

with the practical result that the MemoryMax=85% the repo's template specifies
was not being enforced at all -- `systemctl show` reported
MemoryMax=infinity. Anyone reading the template would reasonably believe the
service was capped.

Startup now compares each installed unit against its substituted template and
warns when they differ, naming install_service.sh as the remedy.

A warning, not an error, and deliberately not a silent rewrite: editing files
under /etc and restarting services is the installer's job, not something a
display process should do to a machine while it is booting. Making it fatal
would also brick every development checkout whose unit is legitimately absent
or hand-edited.

Comparison ignores comments, blank lines and ordering. The template carries
explanatory comments the installed copy will not have, and systemd does not
care about order within a section, so a literal comparison would warn on every
boot and be ignored within a week.

Mutation-checked three ways: never reporting drift fails, making it fatal
fails, and -- after the first attempt missed it -- comparing raw text now fails
too. That last gap is worth noting: the comment-insensitivity tests originally
exercised the helper directly, so a comparison that stopped calling the helper
passed them all. The test that catches it goes through _validate_systemd_units.

29 startup-validator tests pass.

(cherry picked from commit cf521bdfd8)

* fix(install): grant the sudo commands the captive portal actually runs

The installers write two allow-lists, /etc/sudoers.d/ledmatrix_web and
ledmatrix_wifi. Anything the code runs under sudo that is not in one of them
needs a password, which a service cannot supply, so the call fails.

Five commands were being run and none of them granted:

    sysctl -w net.ipv4.ip_forward=0|1     wifi_manager.py:788, 883
    nft add|delete table ip ledmatrix     wifi_manager.py:835, 895
    rfkill unblock wifi                   wifi_manager.py:1811
    iptables ...                          wifi_manager.py:796, 813, 818, 871
    mkdir -p .../dnsmasq-shared.d         wifi_manager.py:922

Together these are the captive portal: unblock the radio, bring up the AP,
add the redirect, turn on forwarding, and undo all of it afterwards. Without
the grants a hardened install would associate clients to the access point and
then fail to route them.

Why it has gone unnoticed: a stock Raspberry Pi image ships
/etc/sudoers.d/010_pi-nopasswd granting the default user

    <user> ALL=(ALL) NOPASSWD: ALL

which satisfies every one of these regardless of what the allow-lists say.
Confirmed on a live rig -- `sudo -n -l` permits sysctl there, and the blanket
rule is why. The allow-lists are effectively decorative on a default image and
only start mattering once that rule is removed or the service runs as another
user.

test_sudo_allowlist_covers_calls.py extracts every argv-style sudo call in
src/ and web_interface/ and asserts an installer grants it, so the next command
added without a rule fails here rather than on someone's hardened box.

Getting that test honest took three passes, each worth recording:

- Matching the literal "systemctl" against rules written as
  `$SYSTEMCTL_PATH enable ...` reported six gaps that did not exist. Binary
  path variables are now normalised before comparing.
- Scanning the whole installer let `NFT_PATH=$(command -v nft)` -- a variable
  definition, not a grant -- satisfy the check on its own, so deleting the
  actual nft rules still passed. Only NOPASSWD lines are considered now.
- `sudo -n <tool>` reported "-n" as the binary. sudo's own flags are skipped.

Each of the five grants is individually mutation-checked: removing any one
fails the suite.

(cherry picked from commit a372b43cd1)

* fix(install): drop the iptables wildcard, and pin each grant properly

Review follow-up. Two findings, both right, and the first is a hole I opened
myself.

`NOPASSWD: iptables *` is a root shell for the web user by another name.
`iptables --modprobe=/path/to/anything` runs that path as root, so a wildcard
grant on iptables escalates rather than restricts. I added that rule while
fixing a permissions gap, which is a worse outcome than the gap. It is gone,
and a test now fails on any trailing-wildcard grant to a tool that can execute
another program -- iptables, nft, tcpdump, find, awk, sed, perl, python, env.

The other finding: checking only the binary made the coverage test far weaker
than it looked. With `sysctl` present anywhere in the allow-list, deleting the
`net.ipv4.ip_forward=0` grant still passed -- and the portal would then be
unable to restore forwarding on teardown. Each required command is now matched
in full, and each is mutation-checked individually, including that exact
single-line case.

Scope pulled in deliberately. The first version of this test tried to assert
that *every* sudo call in the codebase is granted. Run honestly, it showed the
portal also runs iptables, nft, `ip addr`, `ip link` and `cp` with arguments
built at runtime -- an interface name, a port. Those cannot be granted safely
in a sudoers file: the rule needs a trailing wildcard, and that is the
escalation above. Closing that half needs a privileged helper that builds the
rules itself and takes only an interface and a port, granted the way
safe_plugin_rm.sh already is. That is a design decision, not a one-line grant,
so the test now pins the four commands this change actually grants and the
docstring says plainly what it does not cover.

Better a narrow test that is true than a broad one that is not.

(cherry picked from commit 500cfbc9f4)

* fix(install): pin PATH, keep unit order, and tighten the sudoers assertions

Three review findings, all correct.

The installer resolved binaries through an inherited PATH and wrote whatever
it found into sudoers as NOPASSWD grants. first_time_install.sh re-execs
itself with `sudo -E`, which preserves the caller's environment, so a writable
directory early in PATH turned a compromise of the low-privilege web user into
permanent root -- via a file the installer itself wrote. PATH is now pinned to
the system directories before anything is resolved, and every resolved binary
must be root-owned and unwritable by anyone else before it reaches the
sudoers file.

_unit_body() sorted a unit's lines before comparing. Order is not noise in a
systemd unit: repeated ExecStartPre=/ExecStartPost= run in the order they
appear, and a directive that moves between [Unit], [Service] and [Install]
means something different where it lands. The drift check reported no drift
for units that had genuinely changed. Order is preserved now.

Two of that check's own tests asserted the wrong thing --
test_reordered_directives_are_not_drift said so in its name -- and are
inverted, with a second covering a directive moved between sections. The
cosmetic-difference test now varies comments, blank lines and indentation,
which is what the installer actually drops, rather than reversing the file.

The sudoers assertions matched command prefixes, so
`sysctl -w net.ipv4.ip_forward=0 *` satisfied the requirement while granting
the caller arbitrary trailing arguments as root. They are exact now. The
wildcard check also normalises ${NFT_PATH} the same way as $NFT_PATH; the
brace is not a word boundary, so that spelling was skipped entirely.

Verified by reintroducing each: a widened required grant fails the exact
match, `${NFT_PATH} *` fails the wildcard check, and require_trusted_binary
refuses a non-root-owned, world-writable, or missing binary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 14:29:45 -04:00
ChuckandClaude Opus 5 cc258aaffd fix(install): tag the secondary installer's journalctl grants too (#491)
Review was right on all three counts, and the first is the one that matters:
scripts/install/configure_web_sudo.sh writes the same three wildcard
journalctl rules as first_time_install.sh and none of them carried NOEXEC. So
this PR closed the pager escape on one installer path and left it open on the
other, which is close to no fix at all -- a rig configured through that script
still hands out a root shell via less's "!command".

The test could not have caught it, for two independent reasons. INSTALLERS
did not list the file. And even listed, _grant_lines() kept the raw source
line: that installer echoes its rules, so each one ends in a quote rather
than the wildcard, and the trailing-* check skipped every one of them. Either
alone would have hidden it.

Both fixed: the file is covered, and an echoed rule is unwrapped to the
sudoers line it actually emits.

The selector test now covers -t ledmatrix as well. It asserted only the two
-u forms, so deleting the -t rule would have passed.

Verified by removing NOEXEC again from the secondary installer: four of the
six tests fail, where before the suite passed with the vulnerability present.


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 14:28:23 -04:00
ChuckandClaude Opus 5 9083df9f5c fix(pixlet): resolve the release tag correctly when downloading (#461)
* fix(pixlet): resolve the release tag correctly when downloading

Starlark apps render through the pixlet binary, and the installer that
fetches it silently produced nothing, so every app failed with "Pixlet
not available - Starlark apps will not work".

Two compounding defects:

The version lookup parsed the wrong token. GitHub returns the release
JSON on a single line, so `grep '"tag_name"'` matches the whole document
and the greedy `sed 's/.*"([^"]+)".*/\1/'` captures the LAST quoted
string in it. That resolved to "mentions_count", giving a download URL
for a release that does not exist. The `[ -z "$PIXLET_VERSION" ]`
fallback never fired, because the value was not empty -- just wrong.

And `curl -L -o` without `-f` writes a 404 body to the file and exits 0,
so the download was reported as successful and the first sign of trouble
was tar complaining "not in gzip format" about a page of HTML:

    → Downloading linux-arm64...
      Extracting...
    gzip: stdin: not in gzip format
    ✗ Failed to extract archive: .../pixlet_mentions_count_linux-arm64.tar.gz
    Download complete: 0/1 succeeded

Now the tag field is matched directly and the value taken from it, and
the result is checked for a version shape rather than merely being
non-empty -- a wrong-but-non-empty value is exactly what made this
silent. curl gets -f so an HTTP error is a failure, and the archive is
gzip-tested before extraction, since a proxy can return 200 with an
error page.

Verified on an arm64 rig: v0.53.1 resolved, 1/1 downloaded, the binary
runs, and the plugin's own detection finds it at
bin/pixlet/pixlet-linux-arm64.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(pixlet): anchor the version check, and don't echo response bytes raw

Both CodeRabbit findings were valid.

The shape check accepted partial matches, so "v0.53garbage", "0.53" and
"v0.5" passed it and built a download URL for a release that cannot exist
-- the failure the check was added to stop, just one step later. Anchored
at both ends now. Every tronbyt/pixlet release to date is vX.Y.Z (all 38
verified against the API), with an optional suffix left for a future -rc.1
or +build tag.

The invalid-response diagnostic printed bytes straight from whatever
answered the request. NUL and newline were filtered but escape, carriage
return and backspace were not, so an error page could rewrite the output
or bury it in a CI log. Non-printable bytes are stripped and it goes
through printf. CodeRabbit suggested hex-encoding the lot; printable
characters are kept instead, because "<!DOCTYPE html>" is the diagnostic
-- hex would make the line safe and useless.

Also corrected the comment above the parse. It asserted GitHub returns
this JSON on a single line; the API is pretty-printed by default, and I
could not get a single-line response from two machines across five header
variants. The single-line case is real (it is what produces
"mentions_count", and the failing device's error named
pixlet_mentions_count_linux-arm64.tar.gz), but it is a shape to be robust
against, not a constant. As written the comment invites the next reader to
check by hand, see pretty JSON, and conclude the fix was unnecessary.

Tests drive the real script with a stubbed curl: the tag resolves from
both response shapes, non-release values fall back, an HTTP error is
reported as a download failure rather than surfacing later as a tar error,
a non-archive body is rejected before extraction, and the diagnostic
cannot carry control bytes. The stub honours -f the way real curl does --
without that, the HTTP-error test passed against the old script too, since
both end at 0/1 and only the reporting layer differs.

Mutation-checked: 10 of the 16 fail against the pre-fix script, and the 5
covering these two findings fail against this branch's previous state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:35:02 -04:00
ChuckandClaude Opus 5 f887063434 test(harness): flag a mode that draws nothing without reporting it (#447)
The controller skips a mode whose display() returns False and treats
anything else -- including None -- as "content was shown". A mode that
draws nothing and does not return False is therefore never skipped, and
because a mode switch clears the panel first, it sits on a blank screen
for its whole display duration. Two sports plugins shipped exactly that.

The harness rendered those modes and passed them, because it called
display() and discarded the result. Capture it, and warn when a render
produced no lit pixels while claiming content.

Warn-only by default, and deliberately so: a scroll mode's first frame
is legitimately its blank scroll-in buffer, which is 42 of these on the
F1 scoreboard alone. Plugins whose modes are known to draw on their
fixture data can opt into failing via harness.json {"empty_check":
"strict"}, matching how the fill check is staged.

Worth being clear about the limit: this only sees what the fixtures
render. It would not have caught the sports bug, whose fixture seeds
games so the empty path never renders -- that needs the source-level
gate in the plugins repo. What it does catch is the same mistake in any
plugin whose empty state the harness does happen to reach, which is
coverage there was none of before.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 08:17:57 -04:00
ChuckandClaude Fable 5 d9683e28be Codebase audit: fix shipping bugs, remove verified-dead code, repair doc drift, add regression guards (#438)
* fix(web): implement delete_cached so the font catalog cache actually invalidates

api_v3.py's font upload/delete handlers import delete_cached from
web_interface.cache, but the function was never defined. The surrounding
except ImportError silently swallowed the failure, so the fonts_catalog
cache entry survived uploads/deletes and newly uploaded fonts did not
appear until the TTL expired or the service restarted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(web): remove dead weather/stocks partial routes that returned 500

The partial dispatcher still routed 'weather' and 'stocks' to loaders
rendering v3/partials/weather.html and stocks.html — templates that no
longer exist since weather and stocks became store plugins. Requesting
either partial raised TemplateNotFound, which the catch-all turned into
a 500. No template or JS references these partials (the only 'weather'
hit in the front end is a plugin-store category filter option), so the
branches and both loader functions are removed; unknown partials now
fall through to the existing 404 handler.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(deps): align contradictory psutil/Flask-Limiter/freetype-py pins

requirements.txt's optional-install comment recommended psutil>=5.9,<6.0
while web_interface/requirements.txt hard-requires >=6.0,<7.0 — anyone
following the comment ends up with an unsatisfiable pair. The comment now
recommends the same range the web interface requires (all psutil APIs
used — Process, boot_time, cpu_percent, disk_usage, virtual_memory — are
stable in 6.x). Flask-Limiter gains the same <4.0 cap in both files and
freetype-py the same >=2.5.1 floor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(config): add template keys the code already reads

display.hardware gains pixel_mapper_config, row_address_type,
multiplexing and panel_type (read at display_manager.py with these exact
fallbacks — users on non-standard panels previously had no way to
discover them from the template). vegas_scroll gains
frame_based_scrolling and scroll_delay, the only two of its 27 keys the
template omitted (read in src/vegas_mode/config.py). plugin_system gains
development_mode, which the web UI reads and writes but the template
never declared.

Every added value is byte-identical to the code-side .get() fallback, so
ConfigManager._migrate_config() merging these keys into existing user
configs cannot change behavior on any installed device.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(scripts): repair broken sys.path setup in utility scripts

clear_cache.py and download_nba_logos.py pointed sys.path at a 'src'
directory relative to the script's own folder (scripts/utils/src and
scripts/src — neither exists), so both crashed on import; they now insert
the project root and import via the src package like the other scripts.
debug_web_manual.py resolved 'project root' to scripts/debug/ instead of
two levels up. fix_nhl_cache.sh is removed: it used Python docstring
syntax in a bash script and invoked clear_nhl_cache.py, which does not
exist anywhere in the repo — it cannot ever have worked in its current
location.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: correct stale file:line references and the loader-fallback contradiction

CLAUDE.md and .cursorrules disagreed about plugin-directory fallback
behavior; the code (SchemaManager.get_schema_path) probes plugins/
BEFORE plugin-repos/, and the main discovery path has no fallback at
all — both files now describe the real behavior, preferring symbol names
over line numbers so the references rot slower. REST_API_REFERENCE.md
pointed at app.py:144/:607 for mounts that live at :199/:799 and counted
92 routes where there are 94. PLUGIN_ARCHITECTURE_SPEC.md's historical
banner gains a note that its example imports
(src/plugin_system/base_classes/*_plugin.py) never shipped — the real
base classes are src.base_classes.sports.SportsCore and
src.base_classes.hockey.Hockey.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: fix broken links, phantom script references, and stale CI description

Repairs every broken relative link in active docs (targets renamed or
archived long ago: PLUGIN_DEVELOPMENT.md -> PLUGIN_DEVELOPMENT_GUIDE.md,
API_REFERENCE.md -> REST_API_REFERENCE.md, PLUGIN_STORE_USER_GUIDE.md ->
PLUGIN_STORE_GUIDE.md, plugin_docs/ dir, TROUBLESHOOTING_QUICK_START.md,
and MIGRATION_GUIDE's README link that silently resolved to the docs
index instead of the project README). Replaces commands invoking scripts
that do not exist (scripts/update_stats.py, validate_registry.py,
check_updates.py, fix_permissions.sh) with the real tooling, and
rewrites HOW_TO_RUN_TESTS.md's CI section, which described a
security-audit workflow that was never committed and a pytest workflow
'queued to land' that landed long ago as test.yml.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: complete the docs index and refresh the web interface file tree

docs/README.md's own policy says every page must be linked from the
index, yet five weren't — including the entire skin system
(SKIN_SYSTEM.md, CREATING_SKINS.md), ADAPTIVE_LAYOUT.md,
plugin-safety-harness.md and SPORTS_UNIFICATION.md. Each is now listed
in the section it belongs to, and PLUGIN_ARCHITECTURE_SPEC.md is marked
historical in the index (the doc itself already carries the banner).
web_interface/README.md's static/v3 tree showed only app.css/app.js;
it now reflects the actual contents.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove dead modules confirmed unused in-repo and across all store plugins

- src/common/cli.py: imports a 'ledmatrix_common' package that exists
  nowhere (not in this repo, any requirements file, or the plugin
  monorepo), so it cannot ever have run; its README section claimed
  scripts/dev/* used it, which was also untrue.
- src/web_interface/logging_config.py: zero callers — the web app uses
  web_interface/logging_config.py (a different module), and nothing
  imports the src copy.
- handle_errors decorator in src/web_interface/error_handler.py: zero
  call sites (the module's response helpers stay — they are used).
- ConfigManager.get_clock_config(): reads a 'clock' config key that no
  longer exists anywhere; only caller was its own unit test.

Deliberately kept despite zero in-repo callers: DisplayError,
src/common/config_helper.py and display_helper.py — all documented as
plugin-facing API (docs/PLUGIN_ERROR_HANDLING.md, src/common/README.md),
and third-party plugins outside the official monorepo cannot be
enumerated. Verified against a fresh clone of ledmatrix-plugins (43
plugins): zero references to any removed symbol.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove manager-era NBA test files and one-off debug scripts

The four test_nba_*.py files imported nba_managers, leaderboard_manager
and odds_manager — top-level modules deleted when sports displays became
plugins — inside try/except blocks that swallowed the ImportError, so
they passed while exercising nothing. test_nba_data_structure.py and
debug_nba_api.py (a diagnostic script living in test/) made live ESPN
API calls rather than testing repo code. None were enrolled in CI.

scripts/debug/direct_fix_imports.py and check_imports.py were one-shot
artifacts that edited/inspected a hardcoded ~/LEDMatrix/web_interface/
app.py to fix an import problem solved long ago; nothing references
them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove generate_report.py, which aggregates artifacts of CI jobs that do not exist

The script's only function is to merge JSON artifacts
(bandit/semgrep/pip-audit/safety/gitleaks results) produced by a
security-audit workflow that was never committed —
.github/workflows/ has no such jobs, so there is nothing for it to
aggregate and no way to run it usefully. Its siblings stay:
prove_security.py and audit_plugins.py both run standalone (verified),
and .codacy.yml stays because the Codacy service (README badge) reads it
server-side without a workflow file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(web): load the three widget scripts store plugins already declare

time-picker.js, file-upload-single.js and plugin-file-manager.js
register widgets that installed store plugins reference in their config
schemas (countdown uses x-widget: time-picker and file-upload-single;
of-the-day uses plugin-file-manager), but base.html never included the
scripts. plugin_config.html renders such fields as an empty container
that polls LEDMatrixWidgets.get(...) on a 50ms loop forever, so those
plugin config fields appeared permanently blank. The audit initially
flagged these files as dead code; the monorepo cross-check proved the
opposite — they were unreachable, not unused.

example-color-picker.js (the documented custom-widget example) gains an
explicit warning that including it in base.html would shadow the
built-in color-picker widget.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: drop the legacy youtube block from the secrets template

No code reads a top-level youtube secrets key: the youtube-stats plugin
receives its API key namespaced under its own plugin id (declared via
x-secret in its config schema), like every other store plugin. The key
survives only in state_reconciliation.py's non-plugin-key exclusion set,
which stays — existing installs still carry the key in their generated
config_secrets.json, and the exclusion prevents it from being
misclassified as a plugin config. New installs simply stop being asked
for a YouTube API key they have nowhere to use.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore(deps): remove packages nothing imports, declare direct imports, move mypy to test deps

Removed from requirements.txt: python-socketio, python-engineio,
websockets, websocket-client — zero imports anywhere in this repo, and
the one store plugin that needs Socket.IO (ledmatrix-music) declares it
in its own requirements.txt, which the plugin store installs. Removed
the same quartet plus timezonefinder, geopy, google-auth-oauthlib,
google-auth-httplib2, google-api-python-client, unidecode, icalevents,
python-dateutil, flask-wtf and the werkzeug pin from
web_interface/requirements.txt — all leftovers from the deleted built-in
weather/calendar/music displays (flask-wtf was doubly dead: app.py
explicitly disables CSRF and sets csrf=None). scripts/
install_dependencies_apt.py, which mirrors these lists for the
first-time installer, drops the same packages.

Added: urllib3 (imported directly in four core modules), jinja2 and
markupsafe (imported directly in pages_v3.py) — previously reachable
only as transitives. mypy moves from runtime requirements to
requirements-test.txt.

Verified in a fresh venv: all four requirements files co-install, pip
check is clean, the full CI-enrolled suite (907 tests) and a Flask boot
smoke pass with the trimmed dependency set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* refactor: single canonical DateTimeEncoder

src/cache_manager.py and src/cache/disk_cache.py each defined an
identical DateTimeEncoder (datetime -> ISO-8601). The disk_cache copy is
the only one actually used for serialization; cache_manager now
re-exports it instead of defining a twin, so the two can never silently
diverge. Import compatibility is preserved — from src.cache_manager
import DateTimeEncoder still works and is the same class object.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs(code): document deliberate duplicates instead of merging them

The audit surfaced several near-duplicate implementations that turned
out to be either deliberate forks or behaviorally different — merging
any of them would risk changing behavior on installed devices, so each
now carries an explicit comment stating the relationship:

- VisualDisplayManager: headless fork of DisplayManager; header now
  lists the ~15 mirrored methods and warns that DisplayManager changes
  must be mirrored.
- normalize_abbreviation: LogoDownloader's version (called directly by
  nine scoreboard plugins) replaces filesystem-unsafe characters;
  LogoHelper's strips spaces. Logo filenames on existing installs
  depend on both behaviors staying put.
- The two PluginTestBase classes: the shipped one is plugin-author
  API, the repo's own richer harness lives in test/plugins/ — now
  cross-referenced.

Also verified (no change needed): ConfigManager's backup/rollback
methods genuinely delegate to AtomicConfigManager, and SportsCore
already delegates _read_bdf_native_size to FontManager.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: add a unified configuration reference

There was no single place documenting what lives in config.json —
display.* keys were scattered across README sections, vegas_scroll lived
in ADVANCED_FEATURES.md, and dim_schedule, display.double_sided,
sync.follower_position, plugin_system.development_mode and the four
newly-templated hardware keys were documented nowhere. CONFIG_REFERENCE.md
now lists every template key plus the code-read-only keys, each with
type, default, and the code location that reads it, and explains the
secrets file's plugin-id namespacing. Linked from the docs index.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: bring the README's feature tour into the plugin era

The Core Features section still presented clock/weather/sports/stocks/
music displays as built into the project, when all of them are store
plugins installed from the ledmatrix-plugins monorepo — only
starlark-apps and web-ui-info ship in this repo. The intro now says so
(the showcase itself is unchanged; those are real displays available in
the store). The display_durations reference drops its built-in-calendar
example in favor of plugin-id keys, and the Configuration section links
the new CONFIG_REFERENCE.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: archive the custom-icons status report, cross-link config docs, document assets/

PLUGIN_CUSTOM_ICONS_FEATURE.md was a 'What Was Implemented' status
report duplicating the actual guide (PLUGIN_CUSTOM_ICONS.md) — moved to
docs/archive/ per the docs index's own policy. The overlapping
plugin-config docs keep their content but PLUGIN_CONFIG_ARCHITECTURE.md
now states up front which doc is canonical for which purpose.

assets/README.md is new and load-bearing: assets/stocks, weather,
news_logos and broadcast_logos have zero references in this repo's code,
which makes them look deletable — but store plugins (ledmatrix-stocks,
ledmatrix-weather, news, odds-ticker) resolve those exact paths at
runtime against the install directory. The README records that evidence
so a future cleanup doesn't break installed plugins.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* test: add regression guards for the bug classes fixed in this PR

Three lightweight static checks, all enrolled in CI's unit-test
allowlist along with the new web-cache test:

- test_template_targets.py: every literal render_template() target must
  exist (would have caught the weather/stocks partial 500s at commit
  time).
- test_widget_scripts.py: every widget JS file must be script-included
  in base.html or explicitly allowlisted with a reason (would have
  caught the unloaded time-picker/file-upload-single/plugin-file-manager
  widgets), and allowlisted files must NOT be included (prevents the
  example widget from shadowing the real color-picker).
- test_doc_links.py: relative markdown links in active docs must
  resolve (docs/archive/ exempt).

Each guard was verified to fail against the pre-PR tree and pass now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(deps): restore the werkzeug version floor

Commit 1ec22db removed the werkzeug>=3.1.6,<4.0.0 pin along with the
genuinely-unused packages, but this one was a version floor on Flask's
transitive dependency, not a phantom: Flask 3.1.3 itself only requires
werkzeug>=3.1.0, so dropping the pin let fresh installs resolve
3.1.0-3.1.5. Restored with a comment explaining why it exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(web): route pixel_mapper_config into display.hardware; guard time-picker registration

pixel_mapper_config was the only display.hardware key absent from both
the display_fields detection allowlist and the hardware write loop in
the settings save path. No form posts it today, but if one ever did the
key would fall through to the generic handler and land at the TOP level
of config.json — where state_reconciliation would mistake it for a
missing plugin id and loop auto-repair attempts (the failure class the
'github'/'youtube' exclusion comment documents). It now round-trips
into display.hardware like its siblings.

time-picker.js gains the same LEDMatrixWidgets-undefined guard its two
sibling widgets already have; correct today only via defer ordering.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: align installer leftovers with the dependency cleanup

first_time_install.sh's fallback secrets heredoc (used only when the
template is missing) still wrote the legacy youtube block — now matches
the template (github only). install_dependencies_apt.py drops the
IMPORT_NAME_MAP entries for packages no longer in its install lists and
a stale google-api reference in a docstring.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix: address CodeRabbit review findings

Verified each finding against the code; fixes for the valid ones:

- install_dependencies_apt.py: the installer listed 'freetype', but the
  declared dependency is freetype-py — an apt miss would pip-install the
  wrong PyPI package. Now installs freetype-py with an import-name
  mapping (pre-existing bug, surfaced by the review).
- api_v3.py: pixel_mapper_config is validated as a string before being
  saved to display.hardware (JSON callers could previously store an
  object/list the matrix library can't use).
- .cursorrules: the Plugin Loading Process and File Organization
  sections still said discovery scans plugins/ — now consistent with the
  corrected overview (configured directory, default plugin-repos/).
- README.md: removed the stale '(except the core calendar)' claim — no
  core calendar exists in src/ — and qualified the plugin inventory
  (official plugins in the monorepo; third-party from their own repos).
- CONFIG_REFERENCE.md: hardware_mapping now shows the code fallback
  (adafruit-hat-pwm) alongside the template value.
- PLUGIN_REGISTRY_SETUP_GUIDE.md: check_plugin.py takes --plugin, not a
  positional id.
- scripts/fix_perms/fix_*.sh: exec bits set so the documented
  'sudo ./...' invocations work.
- Guard tests hardened: template guard now catches multi-line
  render_template() calls; widget guard parses actual <script> src
  values and fails if the widgets dir goes missing; type hints and
  docstrings added per repo coding guidelines.

Skipped with reasons (noted on the PR): limit_refresh_rate_hz 100-vs-90
is documented as intentional in CONFIG_REFERENCE.md; the psutil comment
already names the enforcing manifest; docs/archive/ findings are out of
scope per the docs policy (archive may rot).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove Cursor IDE tooling, consolidate its guidance into CLAUDE.md

The maintainer no longer uses Cursor. .cursorrules, .cursorignore and the
.cursor/ tree (rules, plugin templates, a parallel 751-line plugins
guide) are removed; measurement showed near-zero literal overlap risk —
the canonical content already lives in docs/. Unique guidance worth
keeping moved before deletion:

- CLAUDE.md gains the dev workflow (dev_plugin_setup.sh, dev_server.py,
  run.py -e, check_plugin.py), the plugin-secrets namespacing contract,
  and the no-draw_image()/paste-onto-PIL pitfall.
- PLUGIN_DEVELOPMENT_GUIDE.md absorbs the plugin version-management
  rules (pre-push hook install, SKIP_TAG, version resolution order) that
  its own text previously linked out to .cursorrules for.
- The one completed plan doc (.cursor/plans/) is archived to
  docs/archive/ per the docs policy rather than deleted.

One of the deleted rule files (sports-managers.mdc) targeted
src/*_managers.py globs that have matched nothing since the plugin
migration.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: second-pass cleanup — dead installer branch, broken script, orphaned JS, misfiled test deps

- first_time_install.sh: removed the pip fallback branch that installed
  from requirements_web_v2.txt — a file that has not existed since the
  v2 web interface was removed (the branch always printed its own
  'not found; skipping' warning).
- scripts/remove_plugin_backups.sh deleted: its PROJECT_ROOT resolved to
  the repo's PARENT directory, and its verify_submodules() checks for
  plugin submodules from an era before plugins moved to the store — it
  could never have worked from its current location.
- plugins_manager.js: removed three functions with zero call sites
  anywhere (addKeyValuePair, formatCommit, togglePasswordVisibility) —
  verified against all templates, all JS, and the dynamic window[name]
  dispatch sites, which resolve widget-registry keys only. Also replaced
  base.html's misleading 'Legacy ... during migration' label: the file
  is deliberately loaded last and provides the LIVE implementations of
  seven window.* plugin actions that shadow same-named definitions in
  app.js/app-shell.js.
- pytest/pytest-cov/pytest-mock moved from runtime requirements.txt to
  requirements-test.txt (CI already installs both files; the installer's
  line-by-line loop simply installs three fewer packages on devices; no
  store plugin declares pytest). HOW_TO_RUN_TESTS.md updated.
- scripts/add_defaults_to_schemas.py and analyze_plugin_schemas.py
  scanned the empty legacy plugins/ dir — now scan plugin-repos/.

Verified: fresh venv installs all four requirements files with pip check
clean and pytest available; bash -n on the installer; node --check on
the JS; widget/cache guard tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: correct semantically stale content across the user and developer guides

A second-pass content audit checked the guides' substantive claims
against the code (the first pass only fixed mechanical drift). Fixes:

- GETTING_STARTED: described booting a prebuilt SD image and seeing
  default clock/weather plugins — neither exists. Now documents the real
  install (Pi OS Lite + one-shot installer / first_time_install.sh) and
  that displays come from the Plugin Store. Duration and ordering
  instructions moved to the Rotation tab where the controls actually
  live.
- WEB_INTERFACE_GUIDE: three whole tabs were undocumented (Rotation,
  Backup & Restore, Tools) and the Display tab's Vegas Scroll section
  was unmentioned. Fonts overrides are per display element (not per
  plugin); Logs has an Auto-scroll checkbox (not a Pause button); the
  aspirational keyboard-shortcut list and no-JS claim removed.
- TROUBLESHOOTING: the hand-written service-file template (wrong user,
  wrong ExecStart, dropped the autostart gate) replaced with the real
  systemd/ units + install scripts; recovery steps no longer copy
  placeholder units verbatim; WiFi curl endpoint corrected to /api/v3/;
  cache-clearing advice now targets the real cache locations.
- ADVANCED_FEATURES: removed a false claim that CacheManager has no
  delete(); fixed two example snippets that raise TypeError
  (BackgroundDataService and get_config_file_mode signatures); fixed
  cache paths, a 5-minute TTL that is actually 1 hour, and the vegas
  table now links the complete 26-key reference.
- EMULATOR_SETUP_GUIDE: documented run.py flags that don't exist
  (--plugin/--test-plugins) removed in favor of dev_server.py and
  check_plugin.py; shipped emulator config values corrected (browser
  adapter default on :8888, not pygame).
- PLUGIN_QUICK_REFERENCE: drag-and-drop reordering is shipped, not
  'not yet supported'; discovery-fallback and registry-repo claims
  corrected. PLUGIN_API_REFERENCE: get_vegas_segment_width returns
  panels, not pixels. CONTRIBUTING: the repo uses flake8/mypy/bandit
  pre-commit hooks, not black/ruff, and tests need requirements-test.txt.
- SKIN_SYSTEM/DEVELOPER_QUICK_REFERENCE: stale module paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix: raise the two new dependency floors past their CVEs, silence a deliberate re-export

All three of these were introduced by this PR, which is what makes them
worth fixing here rather than deferring.

`urllib3` and `jinja2` were added to the requirements so that direct
imports stop relying on transitives — right call, but both floors were
set to the version that introduced the API rather than a version that is
safe to install. `urllib3>=1.26.0` sits below roughly ten CVEs including
a decompression-bomb safeguard bypass, and `jinja2>=3.1.0` below five
including two sandbox breakouts. Raised to 2.7.0 and 3.1.6, which is what
a working device already runs, so no install is disturbed. The comments
now say the floor is a security floor, since the next person to read
"imported directly" would otherwise reasonably lower it again.

This is the same reasoning the PR already applied to werkzeug; these two
just missed it.

The `DateTimeEncoder` import in cache_manager is unused on purpose — the
canonical class moved to src.cache.disk_cache and this re-export keeps
the documented import path working. flake8 cannot see intent, so it gets
an explicit `# noqa: F401` rather than being removed and quietly breaking
anything importing it from here. Verified the re-export still resolves to
the same object and still serialises datetimes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix: address CodeRabbit re-review — installer robustness and doc lint

- install_dependencies_apt.py: an import-only check let Debian
  Bookworm's python3-freetype 2.3.0 satisfy the freetype-py>=2.5.1 pin.
  check_package_installed() now verifies the installed freetype-py
  version, and an apt install that lands below the minimum falls through
  to pip instead of counting as success.
- first_time_install.sh: the .web_deps_installed marker was created even
  when the smart installer failed, so re-runs skipped installation with
  dependencies missing. The marker is now created only on success.
- CONTRIBUTING.md: document installing the pre-commit CLI before
  'pre-commit install' (the requirements files don't provide it).
- Doc lint: fence language on the on-demand cache example (MD040),
  blockquote continuation in GETTING_STARTED (MD028), and the
  suppress_adapter_load_errors key removed from the emulator debug
  example to match the options table.

Skipped one finding with reason (noted on the PR): the per-plugin
display_duration field in PLUGIN_QUICK_REFERENCE's example is not
obsolete — BasePlugin.get_display_duration() reads it and
PLUGIN_CONFIG_CORE_PROPERTIES.md documents it as a core property;
display.display_durations is a per-mode override, not a replacement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 14:04:23 -04:00
ChuckandClaude Opus 5 f2b246ef03 fix(version): make the core version have exactly one answer (#428)
* fix(version): make the core version have exactly one answer

Plugin compatibility floors compare against src.__version__, so that string
has to be trustworthy. It has not been. v3.1.0 was tagged 2026-05-31 while
src/__init__.py still said "1.0.0"; the bump did not land until 2026-07-12.
Every device installed from that release reports 1.0.0, which is below the
(2, 0, 0) floor in PluginLoader._warn_if_incompatible -- so those users are
silently exempt from every plugin compatibility warning.

web_interface carried a third answer, a hardcoded "3.0.0" that nothing read
and that had drifted two majors from the core. It now re-exports the
canonical value, so it cannot disagree again.

Adds:

- test/test_version_consistency.py (enrolled in the core unit CI job):
  src.__version__ is parseable semver, matches the newest CHANGELOG heading,
  the CHANGELOG's headings are unique and descending, and web_interface
  tracks the core. src.plugin_system.__version__ is deliberately excluded --
  it versions the plugin API and moves independently.

- scripts/check_release_version.py + a release-version-check workflow that
  asserts the tag, the CHANGELOG and src.__version__ agree. Runs on pushed
  v* tags and published releases, and via workflow_dispatch so a tag can be
  checked *before* it is created:

      python scripts/check_release_version.py v3.2.0

Verified: 757 core unit tests pass including the four new ones; the script
exits 0 for v3.2.0 and non-zero for both a mismatched tag (v3.1.0) and a
non-semver one (v2.5); web_interface and web_interface.app still import.

Prerequisite for cutting v3.2.0 -- phase B4 in docs/SPORTS_UNIFICATION.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(version): address review — regex strictness, OSError, stale doc claim

From CodeRabbit on #428, all three valid:

- The module docstring claimed the tag check "runs at release time in
  .github/workflows/release-version-check.yml". That workflow is held back to
  a follow-up PR (the pushing token lacks the `workflow` scope), so the claim
  was false as written. Both files now describe the script as a manual
  pre-flight and say the CI wiring is still to come.

- `\d` also matches non-ASCII decimal digits, which int() happily parses, and
  `\s` matches newlines -- so "##\n3.2.0" read as a version heading. Patterns
  now use [0-9] and [ \t], kept in step across the test and the script, with a
  regression test pinning both behaviours.

- A missing or unreadable CHANGELOG.md raised OSError out of read_text() and
  printed a traceback. In a release gate that reads as "the tooling is
  broken"; it now reports the path and a recovery action and exits 1.

Verified: v3.2.0 passes, a mismatched tag exits 1, and a missing CHANGELOG
exits 1 with the new message instead of a traceback. 5 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:21:47 -04:00
ChuckandClaude Opus 5 16fbb7ebeb fix(install): survive the rgbmatrix build on low-memory Pis (#430)
* fix(install): survive the rgbmatrix build on low-memory Pis

The one-shot installer failed at Step 6 on a 1GB Pi with "Failed building
wheel for rgbmatrix", and told the user to install build tools they already
had. The real cause was the kernel OOM killer.

Upstream's pyproject.toml declares no [tool.scikit-build] options, so
scikit-build-core drives Ninja at its default of nproc+2 jobs -- six
concurrent compiles on a 4-core Pi. CMakeLists.txt compiles the same 14
sources three times (~45 translation units), two of them Cython-generated
C++ where a single cc1plus peaks near 800MB. That does not fit in 512MB-1GB
of RAM.

Add scripts/install/lib_lowmem.sh and wire it into the installer:

- Cap build parallelism at max(1, min(cores, RAM/768)) via
  CMAKE_BUILD_PARALLEL_LEVEL, which is what cmake --build actually reads.
  MAKEFLAGS is ignored by Ninja and is set only as a Makefile-generator
  fallback. A 4GB Pi 4 still gets 4 jobs; 512MB and 1GB boards get 1.
- Add a temporary swapfile sized to bring RAM+swap to 3GB (capped at 2GB),
  removed once the build finishes. An EXIT trap is the backstop for the
  error path. Nothing is written to /etc/fstab or /etc/dphys-swapfile.
  Existing swap is measured excluding zram, which is compressed RAM and so
  does not help a build OOM.
- Keep pip's build tree off tmpfs. Debian 13 mounts /tmp as tmpfs, so the
  default held the whole C++ build tree in RAM alongside the compiler.
- Diagnose OOM failures from the build log and the kernel ring buffer,
  instead of always blaming missing build tools. The OOM killer writes
  nothing to pip's output, which is why this was misreported.
- Report RAM and the chosen job count in the Step 1 preflight, and emit a
  heartbeat during the compile so a deliberately serial 15-25 minute build
  does not look like a hang.

New flags --skip-swap and --build-jobs N, with LEDMATRIX_SKIP_SWAP and
LEDMATRIX_BUILD_JOBS equivalents.

Also skip the duplicate apt-get update that the one-shot installer and
first_time_install.sh each ran a minute apart, and complete the
dphys-swapfile advice in diagnose_dependencies.sh with the CONF_MAXSWAP
line, without which raising CONF_SWAPSIZE above 2048 is silently clamped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VdsJs65WnUo8BtKHMAGA1q

* fix(install): address review findings on the low-memory build path

Three fixes from PR review:

- Validate --build-jobs / LEDMATRIX_BUILD_JOBS before check_memory's
  fallback return. When lib_lowmem.sh is absent that return also honoured
  the override, so a non-numeric value skipped validation and instead blew
  up later in an arithmetic test in Step 6 with a generic error.
- Fall back to the default TMPDIR when the disk-backed build directory
  cannot be created, rather than pointing the build at a path that does
  not exist. A nearly-full disk is the likely cause on exactly the devices
  this targets.
- Pass LEDMATRIX_APT_UPDATED explicitly to the sudo child instead of
  relying on -E, which a sudoers env_reset/env_keep policy can strip.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VdsJs65WnUo8BtKHMAGA1q

* fix(install): don't add 30s to every rgbmatrix build

The build progress heartbeat slept for the full 30s report interval
before re-checking whether the compile had finished, so every build paid
up to 30 seconds of dead wall time -- including fast ones on a Pi 4/5 and
every --force-rebuild run.

Poll every 2s and report every 30s instead. Measured: 30s of overhead on
an instant build drops to 2s, with heartbeats still emitted on the same
schedule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VdsJs65WnUo8BtKHMAGA1q

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-03 15:21:18 -04:00
ChuckandClaude Sonnet 5 5b45f35888 Vegas mode: reclaim dead space and pace the rotation (#423)
* Vegas mode: reclaim dead space and pace the rotation

On a wide panel Vegas mode spent much of its time showing black. At 50px/s
on a 512px display, one display width of blank is 10.2 seconds, which makes
several long-standing behaviours expensive:

- ScrollHelper prepended a full display width of black as an "initial gap",
  charged once per cycle — 10.2s of black at the start of every rotation.
- Plugins without get_vegas_content() are captured off a full-display canvas,
  so their blank margins entered the ticker too. Measured: of-the-day drew
  35px of "No Data" on a 512px canvas (92% blank), youtube-stats 142px of
  content with 185px of black either side. Only the scroll_helper path had
  any trimming.
- Cycle transitions deliberately pushed a blank frame and then recomposed
  synchronously: 84ms at best, 4.8s at worst, every millisecond of it black.
- buffer_ahead doubled as the cycle size, so a 21-plugin install showed 3
  plugins per cycle and took ~7 cycles to come around.
- separator_width was applied between every image rather than at plugin
  boundaries, so a per-row ticker like the F1 scoreboard (116 images, which
  it renders 4px apart internally) got a 32px chasm between each row — and
  the width budget didn't count those gaps, so the plugin quietly occupied
  far more of the panel than intended.

Changes:

- src/vegas_mode/geometry.py: numpy column-ink primitives shared by the
  trimmer and the audit tool, so the number reported is the number acted on.
  A Python per-column loop over a 17,000px strip is far too slow for the
  render path.
- PluginAdapter trims every content path, not just scroll_helper. Only outer
  edges are cropped: interior blank columns are the plugin's own layout
  (logo left, score right) and closing them would corrupt the design. A
  plugin on a non-black background is inherently unaffected.
- ScrollHelper.create_scrolling_image takes an explicit lead_gap, still
  defaulting to display_width so the many standalone-ticker callers are
  unchanged. Vegas passes lead_in_width (default 0).
- Cycle end holds the last rendered frame instead of blanking, turning the
  recompose into a brief freeze rather than the panel switching off.
- plugins_per_cycle (default 6) is split from buffer_ahead, which goes back
  to being only a prefetch low-water mark.
- max_plugin_width_ratio (default 3x display width) caps one plugin's share
  of a cycle. Overflow is deferred, not discarded: a rotation offset advances
  each fetch so later rows appear on subsequent cycles. Single oversized
  images are cropped at a blank column so the cut misses glyphs.
- Composition groups images by plugin: rows are joined by intra_plugin_gap
  (default 8) and separator_width applies only between plugins. The width
  budget now counts those gaps.
- Plugin data updates no longer run on the Vegas render path.

All new settings are user-configurable in Display -> Vegas Scroll, including
min/max cycle duration and dynamic duration, which previously existed in code
but were reachable only by hand-editing config.json.

Measured with scripts/dev/vegas_audit.py on a 512x64 panel:

  mean ink coverage    42.7% -> 69.4%
  fully blank           5.9% -> 0%
  reads as empty        13.6% -> 0%
  worst blank stretch    4.8s -> 0s
  full rotation          414s -> 123s
  plugins per cycle         3 -> 6

Note the metric choice: a "fully blank" scan (>=95% black viewport) reported
only 0.4% and badly understated the problem, because two full-width segments
with mid-canvas content never fully blank the viewport — they hold it at ~28%.
window_coverage_stats grades every viewport position by how much ink it
carries, which is what tracks perceived dead time.

Known remaining: cycle transitions still freeze ~3.5s while the next cycle is
fetched. Fixing that needs background prefetch, which is deferred because the
fallback-capture path mutates the shared display_manager.image and racing it
against the render loop risks torn frames.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Drop unused Optional import from the vegas audit script

Flagged by Codacy (F401). Any, Dict and List are all still used.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Align Vegas API bounds with validate(), fix audit config plumbing

Both from review feedback on #423.

The web API's accepted ranges disagreed with VegasModeConfig.validate(),
which is what actually gates Vegas starting:

  scroll_speed      1-100  -> 1-200   (a slider value of 150 returned 400)
  separator_width   0-500  -> 0-128
  target_fps        1-200  -> 30-200
  buffer_ahead      1-20   -> 1-5

The three loose ones were the dangerous direction: the value saved with a
200, then VegasModeCoordinator.start() failed validation with only a log
line, so the ticker silently never ran. The UI already matched validate() in
all four cases, so the API was the odd one out.

test_vegas_api_bounds_match_validate parses the numeric_fields map out of
api_v3 and asserts every bound against validate(), plus that validate()
accepts both endpoints and rejects just outside them, so these cannot drift
apart again. That test immediately caught a missing upper bound on
min_plugin_width, now added — unbounded it would drop every segment and
leave a blank ticker.

Separately, vegas_audit.py constructed PluginAdapter without the config, so
it fell back to VegasModeConfig() defaults and would report trimming and
width-budget behaviour that differed from the user's config.json. It now
passes the loaded config exactly as the coordinator does. This is the same
class of drift the explicit lead_gap and grouping arguments already guard
against. Output is unchanged on a rig whose config matches the defaults.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Vegas mode: render plugins narrower, space rows by measured separation

Trimming reclaims blank margins but cannot compact a layout that genuinely
spans the display — a five-column forecast, a progress bar drawn at 100%
width, a stat block with the panel's whole width between its elements. Those
need the plugin to make different layout decisions, which means telling it the
screen is narrower while it renders.

DisplayManager.render_size() presents a smaller logical canvas for the
duration of a Vegas content fetch, reusing the same _LogicalMatrix
indirection double-sided mode already relies on so plugins see a consistent
size from every accessor. Plugins that size themselves from matrix.width need
no changes at all; one that wants to be explicit can read the new
BasePlugin.get_vegas_render_width().

Width is a percentage so a single setting travels across panel sizes:
vegas_scroll.render_width_pct globally, or vegas_width_pct in an individual
plugin's config. Measured on a 512x64 panel with real data:

  ledmatrix-weather   1536px -> 576px   (forecast becomes narrow cards)
  youtube-stats        353px -> 199px   (2% blank left, so genuinely compact)
  geochron             453px -> 153px   (ink density rises to 100%)
  ledmatrix-flights    950px -> 740px

The youtube-stats figure is the clearest evidence the layout itself changed
rather than being cropped: at full width the content had to be trimmed from
512px to 353px, whereas at 40% it arrives with almost no blank to reclaim.

Row spacing is now measured rather than added. A flat gap gets it wrong in
both directions at once — content drawn flush to its own edges ends up nearly
touching (reported for recent sports scores, which sat 8px apart), while
content already carrying wide margins gets pushed even further out.
separation_gap() measures the blank each pair already has and adds only the
shortfall, up to min_content_separation (default 24). intra_plugin_gap stays
as a floor applied regardless.

Two tests shipped in the previous commit encoded the old flat-gap arithmetic
and are updated to the measured semantics, including one renamed to reflect
that zero intra_plugin_gap alone no longer butts rows together.

Also fixes a real bug found while testing: the harness display manager had no
render_size(), and because the adapter catches broadly that surfaced as "no
content" rather than an error, silently dropping five plugins. Added the
context to VisualTestDisplayManager for parity, and _render_at() now degrades
to a no-op on any display manager lacking it, so a third-party or older
harness loses the narrowing rather than the content.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Vegas mode: end cycles before the wrap, keep the width budget honest

Three fixes, the first a regression from lead_in_width defaulting to 0.

get_visible_portion wraps: once scroll_position + display_width passes the end
of the strip it fills the right of the frame from the *head* of the same strip.
So the final display_width of travel showed the cycle's first plugin re-entering
on the right while its last plugin exited on the left, and the recompose that
followed replaced both at once. On a 512px panel at 50px/s that was 10.2s of
two plugins on screen at once, ending in a hard cut — reported as the ticker
"switching mid-scroll" from F1 to news.

That used to be invisible because the strip began with a full display_width of
blank, so the wrapped-in region was black. Removing that blank (it was 10s of
dead panel per cycle) exposed the wrap. Cycles now end one display width
earlier, before any wrapped content appears, clamped for strips no wider than
the display so they don't complete instantly and spin the recompose loop.

Verified on hardware: a 3936px strip now completes at 68.5s, exactly
(3936 - 512) / 50.

Second, auto_trim=False also skipped the width budget, which is an unrelated
concern — turning off margin cropping should not let one plugin hold the panel
for minutes. Seen in the field: the F1 scoreboard contributed 116 images and
14,848px untouched, giving a 33,821px cycle (11 minutes of content). The budget
now applies regardless of trimming; with it restored that cycle is 6,362px.

Third, the budget accounted for row gaps using the flat intra_plugin_gap while
the compositor had moved to measured separation, so it under-counted by up to
(min_content_separation - intra_plugin_gap) per row and a many-row plugin
overran its cap. Both now use the same separation_gap() rule, and a test
asserts the composed block fits the budget end to end rather than trusting the
two paths to agree.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Fix IndexError in find_blank_cut when the cut lands on the image edge

A cut position after the last column is legitimate — _crop_to_budget asks for
min(start + budget, img.width), which equals the width whenever the remaining
strip is shorter than the budget. find_blank_cut clamped target to width but
then walked leftwards starting at target itself, so ink[width] raised
IndexError.

Caught on hardware: it killed the ledmatrix-stocks fetch, and because
_fetch_plugin_content catches broadly that surfaced as the plugin silently
contributing nothing for the cycle.

Only reachable on the second or later pass of the rotating window over a single
oversized image, which is why the existing tests missed it — they all exercised
the first pass, where start is 0 and start + budget is comfortably inside the
image. Added TestRotationAcrossMultipleCycles, which walks the window round
several times and asserts content is never lost, plus direct coverage of
find_blank_cut at and beyond the image edge.

Both bounds now stop at width - 1 so neither direction can index past the end.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Only cut oversized segments at real gaps between items

The width-budget crop snapped to the nearest blank column, and in rendered text
the gap between two characters is a single column. So a cut routinely landed
inside a word: the cycle showed "Wednesda" and the orphaned "y" turned up as a
lone floating letter in the next cycle, positioned after whatever plugin
happened to precede it.

Measured on the clock-simple segment to confirm: its blank runs are
[1, 1, 1, 1, 1, 8, 8] — five single-column letter gaps, every one of which
find_blank_cut would happily have chosen.

Cuts now only land in a run of at least min_cut_gap blank columns (default 6),
which excludes letter spacing while still finding the gaps plugins put between
items (the stocks ticker uses 32px, baseball 48px). Where no boundary falls
inside the budget the cut waits for the next one and overruns, because
splitting an item is worse than a slightly long segment.

Continuous content is treated differently on purpose: an image with no internal
gaps is a map or a chart, where any column is as good as another, so it is still
cut to the budget exactly. The gap rule protects discrete items; letting a solid
image escape the cap in its name would be wrong.

blank_runs() is vectorised — 48ms for a 17,000px strip, against seconds for a
per-column Python loop.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Hold capture_mode for every plugin render, not just narrowed ones

The native content path only entered capture_mode when it was also narrowing
the canvas, so at full width — which is every plugin without a vegas_width_pct
override, i.e. most of them — a plugin calling update_display() while building
its Vegas content wrote straight to the hardware. That is a visible flash
mid-scroll, and it lines up with the flash reported at cycle transitions, when
several plugins are fetched back to back.

Suppression is now unconditional; the narrowing context stays separate because
it is already a no-op at full width.

Both contexts are reached through helpers that degrade to nullcontext when the
display manager lacks them. That matters more than it looks: the adapter's
handlers are deliberately broad, so an AttributeError from a missing context
does not surface as an error — it surfaces as the plugin contributing nothing.
Making the call unconditional without this turned 44 tests red for exactly that
reason, all of them reporting lost content rather than the real cause.

The test double now provides capture_mode and render_size too, so tests
exercise the real contexts instead of silently taking the degraded path.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Vegas mode: one continuous strip instead of swapping cycles

A cycle used to be a discrete strip that got replaced: motion stopped, every
pixel was substituted at once, and the next group started with the viewport
already full. That is the freeze, the flash and the jump.

The strip is now extended rather than replaced. ScrollHelper gains
append_content(), which adds items on the right without touching
scroll_position or total_distance_scrolled, so motion continues and the next
group simply arrives from the right. Because completion is measured against
total_scroll_width, extending also defers completion — there is no longer a
cycle boundary to see.

drop_scrolled_prefix() reclaims what has gone past, keeping the strip bounded
however long Vegas runs (observed 5,000-11,000px against an unbounded strip
otherwise). It shifts total_distance_scrolled and total_scroll_width together so
the completion arithmetic is unchanged, and refuses to run while the viewport is
wrapping: wrapping reads the head of the strip into the right of the frame, so
trimming the head there would visibly change the picture. A test caught that.

Groups are prepared off the render thread. The constraint is that the canvas and
the matrix proxy are process-wide mutable state, so narrowing or capturing
through them from another thread would corrupt the frame the render loop is
pushing. get_content() therefore takes offscreen_only: the background thread uses
only paths that avoid the canvas, and anything needing it is marked and picked up
on the render thread. That puts the expensive work (native renders of leaderboard
and baseball cards, seconds each) in the background and leaves the cheap work
(display capture, 40-600ms) in the foreground.

DisplayManager's capture flag is now thread-local. As a shared flag, a background
capture would have suppressed the render loop's own frame pushes for its
duration, freezing the panel precisely when the point was to avoid a freeze.

Canvas-bound plugins are drained one at a time rather than as a batch: six at
once held the render thread for 1.75s. Drains are also spaced by two seconds
while the lookahead is healthy, since taking them back to back turns one long
stall into a run of short ones. When the strip is genuinely running short the
throttle is ignored, because content matters more than smoothness there.

Measured on hardware: zero cycle-complete swaps, drains landing 2-4s apart,
lookahead holding at 1,200-3,500px, no errors.

Set continuous_scroll false to restore the swap behaviour; the old path is intact.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Pace the Vegas frame loop adaptively: 31.5 -> 78.7 fps

The loop slept a fixed frame_interval on top of however long the frame took, so
at a measured 31.6ms per frame a flat 8ms of that was pure idle — a quarter of
the budget spent not rendering. It now sleeps only the remainder of the budget.

Measured on hardware: 31.5 fps to 78.7 fps sustained, with CPU going *down* from
150% to 127%. Scroll speed is unchanged at 49.9px/s against a configured 50,
because motion is derived from elapsed time rather than frame count — this buys
smoothness, not speed.

Worth recording what the bottleneck was not: the per-frame render path measures
0.34ms in total (0.18ms for the numpy slice, 0.17ms for the dirty-tracking
digest), which is a theoretical 2900 fps. Optimising any of that would have been
wasted effort. The frame was idle, not busy.

Also nices the prefetch thread. Its work is PIL and numpy that releases the GIL,
so the scheduler can act on the priority, and without it the prefetch competes
for the same cores as the render loop.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Sub-pixel scrolling: motion at the frame rate, not the pixel rate

With integer positioning the number of distinct frames per second equals the
scroll speed in px/s, however fast the loop renders. Measured at 50px/s and
78.7fps, 36% of frames were byte-identical: the extra frames cost work and
bought no motion, and what was left was 50 discrete 1px steps a second.

Two things were wrong with the pre-existing sub-pixel support. get_visible_portion
never consulted sub_pixel_scrolling — it always took the integer path, so the flag
and _get_visible_portion_subpixel were dead code. And that implementation needed
scipy.ndimage.shift, which is not installed on the target devices (HAS_SCIPY is
False there), so it would not have interpolated even if reached. Verified both:
positions 1000.0 and 1000.5 produced identical frames either way.

Blending is now wired up and implemented with numpy. Two details make it
affordable: slice cached_array directly instead of building two PIL images only
to convert them straight back (the naive version measured 15x the integer path),
and use fixed-point uint16 multiply-add rather than float32, which suits the Pi's
cores and gives finer weighting than the panel can resolve. Result 0.939ms
against 0.237ms — 0.70ms added per frame, a 1065fps ceiling.

Measured on hardware: 81.2 fps with blending on, against 78.7 with it off, so no
cost within noise — and every frame is now a distinct position rather than one in
three being a repeat.

The trade is a slight horizontal softening of text, since each frame blends two
positions. Set smooth_scroll false for maximum crispness.

Also benchmarked and cleared as non-issues: extending the strip costs 9.4ms on an
11,000px strip and trimming 2.5ms, both under one frame at this rate.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Add overflow handling: keep ordered content whole instead of rotating a window

The width budget split any oversized plugin by advancing a window each cycle.
That is right for interchangeable items — news headlines, odds, stock prices —
but wrong for ordered content: a league table showed ranks 1-6, then resumed at
7 two rotations later, which reads as out of order and out of context. Nobody
needs rank 23 in a ticker; they need the top of the table, every time.

overflow_mode chooses between them:

  rotate   — advance a window each cycle so everything is seen eventually
             (unchanged default)
  truncate — always show the start and drop the rest, keeping ordered content
             coherent. Records no window state, so every pass starts at the top.

Per-plugin vegas_overflow overrides the global setting, since one install has
both kinds of plugin. Also adds per-plugin vegas_max_width_screens, so content
that must stay whole can be given more room — or uncapped with 0 — without
lifting the cap on every ticker.

Applied on the test rig: f1-scoreboard and ledmatrix-leaderboard set to
truncate, and baseball given 4.5 screens because it was showing 8 of 9 games
when the whole slate needed only a little more room. Verified: F1 now reports
"the first 10 of 116 ... the rest are not shown", baseball has dropped out of
the budget log entirely, and stocks, odds-ticker and stock-news still rotate.

Also corrects the crop log, which claimed "window advances next cycle"
unconditionally and so misreported truncated crops. A test now pins the
behaviour behind the message: truncate must leave no offset recorded.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Stop Vegas mode showing last night's games as if they were live

A game that was live in the evening was still being drawn as live the next
morning. Two faults combined to freeze plugin visuals indefinitely.

PR #291 added a call to plugin_adapter.invalidate_plugin_scroll_cache() so
a plugin's own cached scroll image would be rebuilt from fresh data. That
method was never implemented. hot_swap_content() wraps the call in a broad
except, so every hot swap has raised AttributeError and been swallowed
silently ever since — which is why the visuals it was meant to keep fresh
never were.

Continuous scrolling then removed the only path that reached it at all:
should_recompose() and hot_swap_content() are called from the
non-continuous branch of run_frame(), and continuous_scroll defaults to
True. So on a default install the pending-update flags were set by the
update tick, never consumed, and grew without bound.

Together these froze content completely, because refetching is not enough
on its own: the sports plugins' get_vegas_content() regenerates only "if
the cache is empty", so take_next_group() kept receiving the same picture
however often it asked.

Fixed by:

- Implementing invalidate_plugin_scroll_cache(). It covers both layouts —
  a helper directly on the plugin (stocks, news, odds-ticker) and one
  owned by a scroll-display manager (the sports scoreboards, which is the
  shape that produced this bug) — and clears cached_image and
  cached_array together, since the array is the image's numpy mirror.

- Adding StreamManager.invalidate_pending_updates() and calling it from
  the continuous branch. It only drops the caches; the plugin recomposes
  when it next comes round in the rotation. process_updates() is wrong
  here: it refetches synchronously and merges into the active buffer that
  continuous mode bypasses, and hot_swap_content() rebuilds and
  repositions the whole strip, which is the freeze-and-jump this mode
  exists to avoid.

Tests assert the fix rather than the implementation: 14 of the 17 new
tests fail without it. Includes the wiring itself, since the regression
was a call that was simply absent, and a check that the scroll position is
untouched so this cannot regress into the swap's visible jump.

All Vegas suites pass (355 tests). test_display_controller_vegas_tick.py
still cannot be collected off-device for want of rgbmatrix, identically
with and without this change.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Fix two CodeRabbit-flagged test assertions in vegas density tests

test_prepared_group_is_used_without_refetching had a tautological final
assertion; now checks stream.calls directly. test_no_partial_letter_at_either_edge
required both crop edges to be blank, but the left edge here is always the
crop's start position with no lead-in gap in word_strip, so it legitimately
carries ink — only the right edge is an actual cut and needs the check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 09:40:38 -04:00
ChuckandClaude Fable 5 cdf03fb107 Add skin system: user-installable visual overlays for sports scoreboards (#419)
* Add skin system: user-installable visual overlays for sports scoreboards

Skins restyle a scoreboard's live/recent/upcoming rendering while the
host plugin keeps doing data fetching, scheduling, caching, live
priority, and vegas mode — the anti-fork alternative for users who only
want a different layout.

- src/skin_system/: ScoreboardSkin API, SkinContext (canvas + adaptive
  layout + logo/font helpers), discovery/loading runtime with API major
  version gating and per-skin module namespacing
- src/base_classes/sports.py: _render_game() seam at the three
  _draw_scorebug_layout call sites; skin-first with built-in fallback,
  3-strikes session disable, slow-render warning; per-mode skin config
- scripts/validate_skin.py: headless multi-mode/multi-size validator
  with bundled per-sport fixtures (no hardware or network needed)
- skins/example-classic-baseball/: working reference skin
- Web UI: served-schema Visual Skin dropdown (validation never
  enum-restricted, so uninstalled skins can't invalidate configs) and
  GET /api/v3/skins
- Store: registry entries with type "skin" install to skins/
- docs/SKIN_SYSTEM.md (architecture), docs/CREATING_SKINS.md (author
  guide incl. Claude Code prompt), view-model contract locked by tests

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LrCusPasy1qeUN5anK3aA1

* Address review feedback on skin system

- skin_runtime: cache the entry module so the 2nd/3rd load of the same
  skin (live/recent/upcoming hosts) doesn't re-execute it with unbound
  sibling aliases; rebind cached sibling modules to their bare names
  around entry import and restore prior bindings after; include per-
  manifest mtimes in the discovery cache fingerprint so in-place skin
  updates are picked up
- sports.py: count render_skin_card exceptions toward the 3-strike
  session disable
- store_manager: validate skin ids (pattern + resolved-path containment
  in skins/), reject registry/manifest id mismatches, and stage+validate
  downloads in a temp sibling before replacing an existing skin
- schema_manager: leave the schema untouched when the configured skin
  value is a per-mode mapping (a string dropdown could overwrite it)
- validate_skin.py: reject non-positive sizes and non-object --options
  at parse time; support --output-dir outside the repo; type annotations
- example skin: validate accent_color once at load with logged fallback
- fixtures: pregame 0-0 scores in football/hockey upcoming fixtures
- api /skins: rely on the self-invalidating discovery cache instead of
  force_refresh
- docs: valid JSON manifest example, load_logo caching semantics spelled
  out, language ids on fenced blocks
- tests: view-model contract test now exercises the real extractor;
  regression test for repeated same-skin loads with sibling modules

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LrCusPasy1qeUN5anK3aA1

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-18 11:07:08 -04:00
ChuckandClaude Sonnet 5 66f9950a30 chore(ci): add security-audit tooling scripts (workflow files pushed separately) (#414)
* chore(ci): add security-audit workflow and plugin security-proof scripts

- scripts/prove_security.py, audit_plugins.py, generate_report.py --
  automated checks (dangerous eval()/exec() calls, dependency scanning,
  report generation) for plugins.
- .github/workflows/security-audit.yml + bandit.yaml -- CI wiring for
  the above plus gitleaks secret scanning and bandit static analysis.
- .github/workflows/tests.yml -- pytest matrix across Python 3.10-3.12.

Also fixes two Codacy findings while these files are freshly landing:
- prove_security.py: dropped a pointless f-string prefix with no
  placeholders.
- security-audit.yml: pinned gitleaks/gitleaks-action to a full commit
  SHA (matching this repo's existing pinning convention in test.yml)
  instead of the floating v2 tag.

Split out of the original chore/dead-code-removal commit, which had
accidentally bundled this in alongside unrelated dead-code deletions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* chore: drop workflow files -- pushed separately (needs workflow OAuth scope)

* fix(security-tooling): address PR review findings across bandit.yaml, audit_plugins.py, generate_report.py, prove_security.py

bandit.yaml:
- Removed scripts/prove_security.py's file-level exclusion. Ran bandit
  directly to get ground truth: the real false positive is B105 (dict key
  "PASS" misread as password-like), not the eval/exec pattern the old
  comment claimed. Added a targeted # nosec B105 there, and found+fixed
  the identical pattern already present in generate_report.py.
- Left the repo-wide B607 skip as-is: confirmed via AST scan that properly
  narrowing it touches 100+ bare-name subprocess call sites across
  wifi_manager.py, store_manager.py, permission_utils.py, app.py, and
  start.py -- none of which are part of this PR. Out of proportion to fix
  here; flagged as a dedicated follow-up.

scripts/audit_plugins.py:
- SyntaxError/OSError while scanning a plugin file now report CRITICAL
  (blocking) instead of WARNING/INFO -- a file that couldn't be parsed or
  read was never actually checked for danger, so it must not silently
  pass the audit.
- --plugin <name> now tracks whether the requested plugin was found across
  all PLUGIN_BASE_DIRS and exits 1 with a clear error if not, instead of
  silently scanning zero plugins and reporting success.
- The AST visitor now tracks import aliases (import subprocess as sp;
  from builtins import eval as e) and resolves them before checking
  against dangerous APIs, closing a straightforward evasion of every
  PLUGIN-001 through PLUGIN-005 check. Verified against both aliased and
  unaliased evasion patterns.

scripts/generate_report.py:
- _md_table_row now escapes pipe characters and normalizes newlines in
  every cell, so scanner-controlled content (a matched secret, a bandit
  issue_text) can't corrupt the Markdown table structure.
- _load now distinguishes "artifact missing/malformed" from "valid empty
  result": each summarizer returns an availability flag, and main() now
  reports INCOMPLETE (not PASSED) with exit code 1 when any artifact is
  unavailable, instead of silently folding it in as 0 findings.
- Gitleaks suppression now uses exact-match placeholder values (pulled
  from the actual config_secrets.template.json) plus a template-path
  allowlist, replacing broad substring checks that could hide a real
  secret containing something like "example.com" as part of its value.

scripts/prove_security.py:
- T1b (dangerous plugin calls): a file that fails to parse/read now
  reports CRITICAL with the exception details instead of being silently
  swallowed by `except (SyntaxError, OSError): pass`.
- T6 (Docker hardening): base images must now be pinned to an @sha256
  digest; a specific tag like python:3.12 is mutable and is now correctly
  flagged as unpinned, not just missing tags or :latest.
- T2a (API surface): no config mechanism for enforcing local-only access
  exists in this codebase today (app.py hardcodes host='0.0.0.0'), so the
  "environment-aware" check as described isn't buildable without adding
  new config infrastructure -- out of scope here. Applied the achievable
  part: upgraded from INFO to WARNING, since enforcement can never
  currently be confirmed.
- T1a (zip-slip): replaced the whole-file substring check with an AST
  walk that finds every extract()/extractall() call and confirms an
  is_relative_to() guard + "Zip-slip detected" log precede it in the same
  function. Verified it still passes on the real store_manager.py (both
  the per-member and validate-then-bulk-extract call sites) and correctly
  flags a synthetic unguarded extractall().
- T3a (hardcoded secrets): violation details no longer include the
  matched credential text -- only file, line, pattern type, and a
  redacted SHA-256 fingerprint, so a real finding doesn't get published
  into CI logs/artifacts/PR comments with wider exposure than the
  original leak. Verified with a synthetic secret that no raw content
  reaches the output.

Validated: all four files compile; bandit scans all three scripts clean
(2 legitimate targeted suppressions, 0 unaddressed findings); each new/
changed code path exercised directly (alias evasion, unmatched --plugin,
missing/malformed/valid-empty artifacts, digest-pinning, zip-slip
guard/no-guard, secret redaction); full audit_plugins.py -> prove_security.py
-> generate_report.py pipeline run end-to-end producing a correct report.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(security-tooling): address follow-up review findings on PR #414

scripts/generate_report.py:
- Added an explicit `object` type annotation to _md_sanitize_cell's value
  parameter -- it deliberately accepts any stringifiable value (calls
  str(value) unconditionally), so `object` reflects its actual contract
  more accurately than leaving it untyped.

scripts/prove_security.py:
- Dockerfile FROM-line parsing: renamed the comprehension variable `l` to
  `line` (ambiguous single-letter name). More importantly, fixed a real
  false-positive: `FROM --platform=<platform> <image>` was reading the
  --platform= flag itself as the image token, so a properly digest-pinned
  image behind a platform flag was incorrectly reported as unpinned.
  Verified against platform+digest, platform+tag-only, and digest+AS-alias
  Dockerfiles.

scripts/audit_plugins.py:
- Consolidated visit_Call's dangerous-API detection: previously, alias
  resolution only covered ast.Name calls for eval/exec/compile and
  ast.Attribute calls for subprocess/os.system, missing from-imported
  subprocess/os functions called as bare names (from subprocess import
  run as prun; prun(cmd, shell=True) or from os import system as s;
  s(cmd)). Added _resolve_call_target() to resolve both call shapes to a
  single fully-qualified target, then run all five PLUGIN-00x checks
  against that one resolved value. Verified against 10 evasion
  combinations (from-import aliases, direct/attribute calls, aliased
  module imports) and confirmed zero false positives on benign os/
  subprocess usage without shell=True.

Validated: all three files compile, bandit scans clean (same 2 legitimate
suppressions as before, 0 new findings), audit_plugins.py/prove_security.py
re-run against the real repo with no regressions from the prior fix pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-15 14:33:17 -04:00
2a1c47fa76 chore: remove dead modules and unused dependencies (~1,180 LOC) (#412)
* chore: remove dead modules and unused dependencies (~1,180 LOC)

Deletions, each re-verified with a fresh repo-wide grep (core, web,
scripts, docs, plugin monorepo) immediately before removal:

Modules with zero live importers:
- src/background_cache_mixin.py + src/generic_cache_mixin.py (134+150
  LOC — referenced only by each other)
- src/font_test_manager.py (134 LOC)
- src/image_utils.py (22 LOC, self-documented deprecated)
- src/layout_manager.py (408 LOC — only its own test imported it) +
  test/test_layout_manager.py
- src/common/basketball_plugin_example.py (328 LOC sample)

requirements.txt entries with zero importers in core (pre-plugin-era
manager deps): icalevents, geopy, timezonefinder, unidecode. Plus the
google-auth trio (google-auth-oauthlib, google-auth-httplib2,
google-api-python-client): their only importer is the calendar PLUGIN,
which declares all three in its own requirements.txt (verified in the
monorepo and on an installed copy) — the plugin dependency installer
owns them. Existing venvs are unaffected (removal doesn't uninstall);
fresh installs get them when calendar is installed.

Two stale references cleaned (a comment in test_pillow_compat.py, a
directory listing in HOW_TO_RUN_TESTS.md). Full suite green except the
two documented pre-existing failures (circuit_breaker mock drift, fixed
in #400; clock-simple 64x32 overflow, pre-dates this series); all core
entry modules verified importing cleanly under the emulator.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix: remove content accidentally bundled into the dead-code-removal commit

The previous commit's git add/commit swept in a lot of unrelated,
unreviewed work alongside the intended dead-code deletions: a new Plugin
Composer web UI, new security-audit/CI tooling, a new march-madness
plugin, and 23 local-development-only symlinks under plugin-repos/ (per
scripts/setup_plugin_repos.py's own docstring, these are meant to be
generated locally, never committed -- .gitignore has no entry for them,
which is how they slipped in).

Removed here, split into their own PRs instead (except plugin-repos/*
symlinks and march-madness/ncaa_logos, which are dropped rather than
carried forward -- see PR discussion):
- web_interface/blueprints/composer.py + composer-app.js +
  composer-canvas.js + composer.html + manager.py.j2
- scripts/prove_security.py, audit_plugins.py, generate_report.py
- .github/workflows/security-audit.yml, .github/workflows/tests.yml,
  bandit.yaml
- All plugin-repos/* symlinks (local dev artifacts, not meant to be
  committed at all)
- plugin-repos/march-madness/* and the 4 new assets/sports/ncaa_logos/*
  PNGs it needed (left out of every split PR pending a decision on
  whether march-madness belongs in the core repo or the plugin monorepo)

This PR now contains only what its title describes: the dead-code
removal from the previous commit, untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 09:58:31 -04:00
14a59c863c perf(wifi): one status fetch per monitor tick instead of three (#411)
* perf(wifi): one status fetch per monitor tick instead of three

The wifi monitor daemon fetched WiFi status + ethernet state before its
AP-mode check, the check internally fetched the same state again (with
retry), and the daemon fetched a third time afterwards — each fetch is
several nmcli subprocess forks, every 30s, forever, even on a perfectly
healthy link.

check_and_manage_ap_mode's decision logic is extracted to
_manage_ap_mode(status, ethernet, ap_active); the new
check_and_manage_ap_mode_with_state() runs the single (retrying) fetch
battery and returns (changed, status, ethernet, ap_active_after) — the
post-state is derivable because state only ever flips via one enable or
one disable. The original bool-returning method delegates, so existing
callers are untouched. The daemon's dead pre-fetch is removed and its
post-check reads use the returned state; the AP-enable retry semantics
(_get_wifi_status_with_retry) are preserved exactly. ~10-15 forks/30s
drops to ~4-6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(wifi): add missing type hints on AP-mode state helpers, fix unused-var lint in tests

- check_and_manage_ap_mode_with_state now declares its
  Tuple[bool, WiFiStatus, bool, bool] return type instead of being untyped.
- _manage_ap_mode's status/ethernet_connected/ap_active parameters are now
  typed (WiFiStatus, bool, bool), matching its existing -> bool return hint.
- test_wifi_check_state.py: rename unused unpacked variables (status/
  ethernet/ap_after) to underscore-prefixed names at the three call sites
  that don't use them, leaving the used ones untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 11:22:37 -04:00
Chuck 3d347a368a fix(testing): stop plugin enabled:false schema defaults from silently disabling harness tests (#408)
check_plugin.py, render_plugin.py, and the pytest plugin matrix each built
config as {"enabled": True} then merged in config_schema.json's defaults on
top, letting a plugin's own enabled:false default (a reasonable choice for
a seasonal/opt-in plugin -- 15 of 23 real plugins ship one) silently win.
Every harness/CI render of those plugins was testing "disabled, do
nothing" rather than real behavior.

Extract build_full_config() into testing/loading.py (already the shared
home for plugin-discovery/config-default logic) and use it from all three
call sites: schema defaults, then a forced enabled=True, then harness.json's
config, then the caller's explicit config -- so a test can still
deliberately disable a plugin on purpose, it just can't happen by accident
via the plugin's own shipped schema default anymore.
2026-07-14 08:21:35 -04:00
7f7f0d6464 feat: adaptive layout system — size-aware regions, crisp font ladders, image fitting (#393)
* feat(layout): adaptive layout & font scaling system for plugins

Add src/adaptive_layout.py — opt-in core helpers so plugins render
legibly on any panel size without hand-tuned per-display layouts:

- Region: integer rect algebra (bands/columns/weighted splits/centering)
  that partitions space so text bands can't overlap by construction
- Font ladders: ordered (family, size) steps known to render crisply
  (LADDER_GRID: X11 BDFs at native sizes; LADDER_ARCADE: PressStart2P at
  8px multiples) — fitting walks the ladder instead of scaling pixel
  fonts fractionally
- LayoutContext: breakpoint tiers, geometry scale vs. a declared design
  size, and cached fit_text/fit_lines/font_for_rows queries

Generalizes the three patterns proven in the field: f1-scoreboard's
scale factor, masters-tournament's tiers, baseball-scoreboard's font
fallback ladder.

Wiring: BasePlugin gains a lazy .layout property and draw_fit();
FontManager gains get_native_bdf_size() and a cache_generation counter;
manifest schema gains display.design_size and requires.display_size
max_width/max_height; 96x48 joins DEFAULT_TEST_SIZES; the bounds-check
harness records negative-coordinate draws; TextHelper's broken
measurement helpers are fixed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): adaptive image fitting + composite region helpers

Add src/adaptive_images.py — the image counterpart to fit_text:
- fit_image(img, box, mode=contain|cover|fill_height|stretch,
  crop_to_ink, anchor, resample, upscale) promoting the proven plugin
  patterns (football's crop-to-ink fill-height logos, masters' cover
  crop + NEAREST flags, static-image's letterbox). Upscales by default —
  thumbnail()'s downscale-only behavior is why imagery stays tiny on
  big panels.
- draw_fitted_image() pastes aligned within a Region with alpha mask.
- One central Pillow>=9.1 RESAMPLE shim replacing ~15 plugin copies.

LayoutContext.fit_image() caches results per (identity, box size,
options) with a 64-entry LRU; id()-keyed entries pin the source image.
BasePlugin.draw_image() is the one-liner adoption path beside draw_fit.

Composites in adaptive_layout.py: Region.offset() (user x/y-offset
passthrough), scoreboard_regions() (the two-logos-plus-score card math
duplicated across six sports plugins, logo_slot = min(H, W//2)), and
media_row() (art-left/text-right).

Fix LogoHelper's size-blind cache key (stale sizes on panel change);
deprecation note on dead image_utils.py.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(harness): scale-up fill check, config variants, multi-size dev gallery

Quality gates for adaptive layout:

- fill_metrics()/check_scale_up() in the safety harness: overflow catches
  content too big for a panel, but nothing caught content that stays tiny
  on panels >= 2x the plugin's declared design size. The check measures
  lit-content extents and warns (or fails, when a plugin opts into
  "fill_check": "strict" in test/harness.json) below 50% coverage on the
  doubled axis. Warn-only by default so no existing plugin breaks.

- harness.json "variants": extra runs with config overlays and their own
  golden dirs, so an opt-in mode (e.g. layout_mode: adaptive) is golden-
  tested beside the classic default. check_plugin.py loops base + variants
  and labels variant results mode@name.

- Dev preview server: GET /api/sizes (harness size sample), POST
  /api/render-matrix (render at up to 12 sizes in one call), size-preset
  dropdown, and an "All Sizes" side-by-side gallery in the preview UI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(plugins): adaptive-lib discoverability + advisory version compat warning

Discoverability: re-export the adaptive layout/image API from src.common
(the blessed-helpers package plugin authors already know) — canonical
paths stay src.adaptive_layout / src.adaptive_images so nothing breaks.
Document it in src/common/README.md and cross-link ADAPTIVE_LAYOUT.md
from the developer docs authors actually read (quick reference, API
reference, advanced dev, font manager, dev preview, plugin dev guide);
ADAPTIVE_LAYOUT.md gains adaptive-images, composite-layouts and
preserving-user-customization sections.

Compat: PluginLoader now logs one advisory warning (never raises) when a
plugin's manifest declares a min LEDMatrix version newer than the running
core, checking the min_ledmatrix_version / requires.* / versions[]
spellings found in the wild. Guarded against stale core version numbers.

src/__init__.py __version__ bumped 1.0.0 -> 3.1.0 to match the latest
release tag (v3.1.0) — it had never been updated and the compat check
needs a truthful number. NOTE: verify this matches the intended release
numbering before the next tag.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): add measure_font_crispness — verify a ladder rung isn't blurry

PIL antialiases TTF outlines by default; a 'pixel-style' font only
rasterizes without antialiasing at specific sizes (for PressStart2P:
exact multiples of its 8px design grid). A ladder rung at an unverified
size silently renders blurry on an LED panel — this exact bug shipped in
both text-display's and football-scoreboard's custom TTF ladders
(non-8-multiple PressStart2P sizes, and '5by7.regular'/'4x6-font' at
sizes that were never actually crisp).

measure_font_crispness(font, sample_text) renders the sample and reports
the fraction of ink-bbox pixels that are neither pure black nor pure
white. BDF fonts (real bitmaps) always score 0.0; TTF ladders should be
verified against this before shipping — see the new
TestFontFitting::test_ladder_arcade_is_crisp pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): add fit_text_proportional — proportional sizing vs. always-maximize

fit_text always picks the largest ladder rung that fits its box. That's
right when an element owns dedicated space, but wrong when several
independently-fitted elements need to stay visually harmonious as the
panel grows: a score's box might have generous room while a neighboring
logo scales by a fixed geometry factor via px() — fit_text lets the score
balloon out of proportion (even overlapping the logo) even though its
individual pick is technically correct.

fit_text_proportional(text, box, base_size_px, ladder) instead targets
base_size_px * self.scale (the same scale factor px() already uses),
picking the nearest ladder rung at or below that target, still capped to
what fits the box, floored at the smallest rung when the target is below
every rung. Refactored the shared largest-that-fits/ellipsize walk into
_walk_ladder() so fit_text and fit_text_proportional don't duplicate it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): fit_text_proportional gains an axis-specific scale override

self.scale (min(width_ratio, height_ratio)) is the right conservative
default for anything whose aspect ratio matters, but a caller whose
surrounding composition already scales along a single axis — e.g.
football-scoreboard's logo_slot = min(height, width // 2), which tracks
height alone — needs text sized the same way, or it reads as
under-scaled next to logos that grew on a panel that only got taller
(128x32 -> 128x64: self.scale stays 1.0 since width didn't grow, but
logos still double).

fit_text_proportional(..., scale=None) now accepts an explicit override;
None keeps the existing self.scale default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(layout): scoreboard_regions reserves real center space at 2:1 aspect ratios

logo_slot = min(height, width // 2) has a blind spot: at exactly 2:1
aspect ratio (width == 2 * height -- a very common shape: two, four, or
more square modules stacked into a taller panel) width // 2 and height
are equal, so the two logo slots claim the ENTIRE width and leave zero
pixels for a center column, no matter how large the panel gets. Not a
'small panel' problem -- 96x48, 128x64, and 256x128 (all exactly 2:1) hit
it identically, while the 128x32 design baseline and panels like 192x48
or 256x32 never do, because height is already the tighter constraint
there.

Two new parameters fix it in the one shared helper every scoreboard-style
plugin composes through:

- min_center_fraction / min_center_design_px reserve at least
  max(width * fraction, design_px * ctx.scale) for the center column,
  capping logo_slot further when needed. The scaled design-px term
  matters on small panels where a flat fraction alone reserves too little
  absolute space.
- score_bleed_fraction extends the score's own fit box (not the logo
  slots themselves) a controlled amount into each side -- the same way
  real broadcast scoreboards let a big score number's edges cross into
  the team marks flanking it. Without this the reserve alone can still be
  too narrow for a short score to render without truncating.

score_area is now genuinely narrower than the full card width (previously
identical to status_band/detail_band, which still span the full width and
overlay the logos -- short text there was never the problem).

Verified against the full harness size spread: a real game score like
'17-21' never needs ellipsis at any tested 2:1-or-tighter aspect ratio
(test_score_never_needs_ellipsis_for_a_short_score), and wide panels
(128x32/192x48/256x32-style) are provably unaffected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: document scoreboard_regions' center-reserve and score-bleed params

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: address CodeRabbit review on PR #393

- docs: scope the self.layout note to BasePlugin subclasses (others build
  a LayoutContext directly) and make explicit that adaptive layout is
  opt-in — classic rendering stays unless a plugin adopts the APIs.
- dev_server: broaden the render-request catch (a bad manifest.json now
  returns a clean 400 instead of an unhandled 500) and stop echoing raw
  exception text in the loader-failure responses — full tracebacks go to
  the dev server's console log instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): allowlist plugin_id before any path lookup

CodeQL (py/path-injection): plugin_id arrives in request input and flows
into filesystem paths via find_plugin_dir. Gate it with the same
^[a-zA-Z0-9_-]{1,64}$ allowlist the web UI's pages_v3 uses, at the
single choke point every route resolves through.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): lexical containment check on resolved plugin dirs

CodeQL doesn't recognize the interprocedural allowlist as a
path-injection barrier; add the canonical one — normalize (without
following symlinks, since dev plugins are commonly symlinked into
plugins/) and require the result to stay inside the search dir.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): inline normpath containment barrier before render

CodeQL doesn't credit the sanitization inside find_plugin_dir along
this flow; apply its documented barrier (normpath + startswith against
the allowed roots) inline in _parse_render_request, on the exact path
that reaches the render/load sinks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): derive plugin dir from trusted directory listings

CodeQL's barrier-guard recognition doesn't see a startswith check
inside an any() comprehension, so the normalize-and-prefix approach
still flagged. Break the taint outright instead: after lookup, re-derive
the directory by enumerating the search dirs (iterdir) and matching by
path equality — the Path used for all downstream file access is built
solely from trusted listings, never from request input.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): use os.scandir for path-injection barrier, redact stack traces from render responses

CodeQL doesn't model Path.iterdir() as a taint-clearing enumeration the
way it does os.scandir() -- _trusted_plugin_dir's iterdir-based rebuild
still traced plugin_id through to the manifest.json open(). Switched to
scandir, matching the pattern already verified clean on PR #396.

Also stops surfacing raw exception text (update()/display() failures)
in the JSON render response -- logs full detail server-side via
exc_info instead, returning only the exception class name to the
client. And drops path values from three plugin_loader debug/error
logs that CodeQL flags as clear-text-logging of externally-influenced
data, keeping plugin_id (not flagged) for context.

* fix(dev-server): remove conditional-reassignment ambiguity in plugin_dir resolution

CodeQL's path-injection flow still traced through _parse_render_request
after the scandir fix -- the tainted find_plugin_dir() result and the
scandir-derived _trusted_plugin_dir() result shared the same variable
name (plugin_dir), reassigned only on the truthy branch. That merge
point apparently isn't treated as a barrier by the flow analysis, so it
kept tracing the pre-reassignment value through to the manifest open().

Split into two distinct names -- candidate_dir (tainted, used only to
call _trusted_plugin_dir) and trusted_dir (the only name used for any
downstream file access) -- so there's no reassigned variable for the
flow to walk through.

* fix: remove unused imports flagged by Codacy

Union in adaptive_images.py and field in adaptive_layout.py are both
imported but never used -- the last two Codacy findings on this PR,
matching the same fix already applied on PR #396.

* fix(layout): bound the fit cache; never alias the source image in fits

Two latent issues found in a self-review pass:

- LayoutContext._fit_cache was an unbounded dict (the image cache got an
  LRU cap, the text-fit cache didn't). Cache keys embed the fitted TEXT,
  so a plugin fitting changing strings — a live game clock, a ticker —
  on a 24/7 service grows it forever. Now LRU-bounded at 512 entries via
  the same pattern as the image cache.

- fit_image returned the caller's ORIGINAL image object when the source
  was already RGBA at target size (contain/fill_height, no ink crop).
  ImageFitResult is documented as an independent copy, and LayoutContext
  caches results — an aliased image lets later mutations of the source
  corrupt cached fits (or vice versa). Copy in that branch.

Both covered by new regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 10:38:52 -04:00
ChuckandClaude Sonnet 5 05e7c43b27 fix(plugins): replace dependency marker files with a real satisfaction check (#390)
* fix(plugins): replace dependency marker files with a real satisfaction check

The .dependencies_installed hash-marker system only tracked "was this exact
requirements.txt hashed before" — not whether the packages it names are
actually present. That made it fragile (a wiped venv, a manually removed
package, or a lost/corrupted marker forces a needless full pip reinstall or,
worse, a false skip) and produced dead weight for the ~10 plugins whose
requirements.txt is comment-only (they still paid a pip subprocess on first
boot before a marker existed).

Replace it with requirements_are_satisfied() in plugin_loader.py, which
checks each real requirement line against importlib.metadata directly, so
install_dependencies() only shells out to pip when something is actually
missing or version-mismatched. Drops the marker file entirely: removed all
marker read/write sites in plugin_loader.py and store_manager.py, the
now-pointless marker-cleanup step in the git-update path, the unused legacy
marker implementation in plugin_manager.py, and the already-stale
clear_dependency_markers.sh script.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(security): close path-injection gap in dependency-satisfaction checks

CodeQL flagged 2 new high-severity "uncontrolled data used in path
expression" alerts at the open() calls inside this PR's new
requirements_has_real_deps()/requirements_are_satisfied() -- both are
reachable from paths that were never run through the basename+trusted-base
sanitiser this codebase already uses elsewhere:

- PluginLoader.install_dependencies() only applied that sanitiser when its
  optional plugins_dir argument was actually passed; the "no plugins_dir"
  branch trusted plugin_dir_real directly. Made plugins_dir required (not
  Optional) so that branch can't exist, and added an explicit guard in
  load_plugin() so install_deps=True without a plugins_dir fails loudly
  instead of silently. Production's only real caller (PluginManager) always
  passes plugins_dir already; the harness/dev-server/render-plugin callers
  all use install_deps=False and are unaffected.

- StoreManager._install_dependencies() never sanitised plugin_path at all,
  and its call sites ultimately derive that path from a plugin's own
  manifest.json "id" field (install_plugin_from_url) -- a malicious plugin
  could otherwise point requirements_file outside plugins_dir. Applied the
  same os.path.basename()-based containment pattern PluginLoader already
  uses (and that CodeQL recognises as a real sanitiser).

Added test_install_dependencies_requires_plugins_dir and
test_install_dependencies_rejects_path_outside_plugins_dir to lock in the
actual security property, not just quiet the scanner. Verified: all 20
tests in test_plugin_loader.py pass, plus the PR's existing test plan
(test_plugin_system.py, test_store_manager_caches.py: 53 passed) and the
full CI plugin-safety suite (test_harness.py, test_visual_rendering.py,
test_plugin_matrix.py: 52 passed, 2 pre-existing skips) all still pass.

* fix(security): replace basename-only sanitiser with a trusted-enumeration check

The previous commit's os.path.basename() + os.path.join() pattern (which a
pre-existing code comment claimed CodeQL recognises as a sanitiser) did not
actually clear the alert -- the next CodeQL run still flagged the same 2
sink lines, plus a new one at the os.path.join() call itself. Taking a
substring of tainted data apparently isn't treated as a barrier by this
query, whatever the comment assumed.

Replaced it with find_trusted_subdir(): enumerate the trusted plugins_dir
via os.scandir() and only use a name that scandir itself produced, matched
by equality against the caller's requested name. The path is then built
from that enumerated entry, not from the caller's string -- a value
sourced from iterating a trusted, non-tainted directory carries no taint
regardless of what it happens to equal, which is a stronger and more
conventional allowlist-style barrier than string-stripping. Applied
identically in both PluginLoader.install_dependencies() and
StoreManager._install_dependencies(), sharing one implementation.

Re-verified: all 65 tests across test_plugin_loader.py (20, including the
2 new security regression tests), test_store_manager_caches.py (35),
test_plugin_system.py (10) pass, plus the full CI plugin-safety suite
(test_harness.py/test_visual_rendering.py/test_plugin_matrix.py: 52
passed, 2 pre-existing skips).

* fix(security): redact URL credentials from pip subprocess output before logging

CodeQL flagged 3 clear-text-logging-of-secrets alerts in
install_requirements_file() (src/common/permission_utils.py:353,360,371).
Pre-existing on main, unrelated to this PR's own diff, but now visible
since the path-injection alerts that previously took priority in the
annotation list are fixed.

The underlying risk is real: pip can echo a private index URL's embedded
basic-auth credentials (from a requirements.txt --index-url line or
PIP_INDEX_URL) back verbatim in its own stderr/stdout on failure, and this
function both logs that output directly and returns it to callers --
store_manager.py's _install_dependencies() logs result.stderr from this
same function too.

Added _redact_url_credentials(), applied immediately after each of the two
subprocess.run() calls (mutating result.stderr/stdout in place) rather
than patching each log call site individually. This closes the leak at
the source: every downstream use -- the three flagged log lines, the
"note" string embedded in the returned stdout, and store_manager.py's own
logging of the returned result -- gets the redacted text for free.

Verified the fixed-phrase "denied" check (`"a password is required" in
result.stderr`) is unaffected, since URL syntax and those phrases don't
overlap -- covered explicitly by
test_does_not_touch_denied_check_phrases. Added
test/test_permission_utils.py (6 tests) covering the redaction helper
directly and both subprocess.run() call sites (the sudo-wrapper branch,
which this repo's scripts/fix_perms/safe_pip_install.sh makes live, and
the no-wrapper fallback branch). All pass.

* fix(security): stop interpolating req_file/pip-output into log calls

The previous commit's redaction (mutating result.stderr/stdout right after
each subprocess.run()) didn't clear CodeQL's clear-text-logging alerts --
same lesson as the path-injection fix earlier in this PR: a static
analyzer can't tell "this value was already sanitised two lines up" from
"this is still the raw tainted value" just by looking at a single log
call in isolation, so it conservatively keeps flagging it regardless of
what the redaction function actually does.

Removed all dynamic interpolation (req_file, result.stderr) from the 3
flagged logger.warning() calls entirely, replacing them with fixed
messages plus (for the one that had it) result.returncode, which is a
plain int with no possible taint. The full redacted detail is still
available where it actually matters -- in the returned
CompletedProcess.stderr/stdout and the "note" text -- just not duplicated
into a log line a scanner has to reason about in isolation.

Re-verified: all 6 test_permission_utils.py tests still pass (they assert
on the returned result, not log call arguments), plus the full
test_plugin_loader.py/test_store_manager_caches.py/test_plugin_system.py
suite (71 passed, 1 pre-existing deselect, 4 subtests).

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-11 08:55:50 -04:00