* fix(web): on-demand no longer restarts a running display service
POST /display/on-demand/start treated start_service (default true, sent by
"Preview on display", the on-demand dialog and the MQTT bridge) as
"restart": with the service running it ran systemctl stop, slept 1.5s and
started it again. Every request cold-started the display process -- every
plugin reloaded, panel blank -- to deliver a request the running process
already reads from the cache mailbox every ON_DEMAND_POLL_INTERVAL (0.25s),
including mid-dwell, mid-screen and mid-Vegas. The restart bought nothing:
startup only restores a session the display saved itself
(display_on_demand_config), so the new request arrived through the same
mailbox either way.
start_service now means "start it if it is not running". The stop route
coerces stop_service to a boolean so "false" no longer stops the service.
test_api_v3_on_demand_restart.py pinned the old restart path; it now pins
the replacement. Docs updated.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): on-demand loads a disabled plugin live instead of failing
The display process only loads enabled plugins, so an on-demand request for
a disabled one -- "Preview on display" offers it on every config page, with a
note that the plugin will be enabled for the preview -- failed with
invalid-mode. Nothing enabled it short of a restart, and the on-demand route
no longer restarts the service.
_activate_on_demand now loads an installed-but-not-running plugin through
the live-enable path (load_plugin + _register_loaded_plugin), with a new
load_plugin(force_enabled=True) so the instance runs enabled while
config.json keeps saying disabled. The plugin is tracked in
_on_demand_loaded_plugins, and the main loop unloads it through
_unregister_plugin once on-demand moves off it (stop, expiry, another
request, or a failed request that ends the session) -- right after its own
poll, where no display() is on the stack. A failed load publishes status
error with load-failed. A plugin enabled during the session stays loaded.
A session restored after a restart uses the same tracking instead of
setting enabled in the config dict config_manager caches, so its plugin is
unloaded when the session ends rather than staying loaded until the next
restart. Ending a session no longer resumes the rotation onto a plugin that
is about to be unloaded, which a restored session did.
Also: a stop sent while on-demand is inactive clears a failed request's
error, instead of /display/on-demand/status reporting status: error until
the state aged out.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Any test that imported web_interface.app and sent a request fired the app's
startup reconciliation, which runs against the checkout's real config.json
and plugin-repos/ and reinstalls every configured-but-missing plugin from the
live store. A full Windows run left basketball-scoreboard, calendar,
football-scoreboard, leaderboard and ledmatrix-stocks untracked in
plugin-repos/ (not gitignored) from that daemon thread.
test/conftest.py now installs an import hook that sets the app's run-once
_reconciliation_started latch as the module finishes executing, so lazy
imports, module-level imports and reloads all start disarmed.
StateReconciliation's own tests are unaffected. A regression test pins that
a request to the imported app launches no reconciliation thread.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
POST /display/on-demand/start treated start_service (default true, sent by
"Preview on display", the on-demand dialog and the MQTT bridge) as
"restart": with the service running it ran systemctl stop, slept 1.5s and
started it again. Every request cold-started the display process -- every
plugin reloaded, panel blank -- to deliver a request the running process
already reads from the cache mailbox every ON_DEMAND_POLL_INTERVAL (0.25s),
including mid-dwell, mid-screen and mid-Vegas. The restart bought nothing:
startup only restores a session the display saved itself
(display_on_demand_config), so the new request arrived through the same
mailbox either way.
start_service now means "start it if it is not running". The stop route
coerces stop_service to a boolean so "false" no longer stops the service.
test_api_v3_on_demand_restart.py pinned the old restart path; it now pins
the replacement. Docs updated.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
DiskCache.set skipped a payload identical to the last one written for the
key, but CacheManager.set stamps every record with time.time(), so the
payload always differed and the skip never fired: unchanged API data was
rewritten to the SD card on every plugin update cycle.
Header-first records are now compared without their timestamp (the digest
also carries the content length, since a collision is now a missed write).
A skipped write moves the file's mtime to the skipped record's timestamp,
and a real write sets it to the embedded one, so only a skip moves it
forward. DiskCache.get, including the header fast path, treats such a
record as fresh from the later of the two and returns that time as the
record's timestamp. The skip also checks the file is still the one this
process wrote (inode and size), so a file another process replaced is
rewritten.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Bumps src.__version__ to 3.7.0 and turns Unreleased (#672: sports_celebration,
sports_fetch and sports_card_wrappers) into ## 3.7.0; src/common/README.md and
docs/SPORTS_UNIFICATION.md say 3.7.0 for the three modules.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Three new hardware-free modules holding code the scoreboard plugins carry
as identical copies (executable AST, docstrings stripped, checked across
every carrying plugin at ledmatrix-plugins 30455671). The bodies are the
plugins'; the changes are type annotations for the mypy ratchet, the
colour helpers losing their leading underscore as public free functions,
and two comments that described the plugins' files.
- src/common/sports_celebration.py: SportsCelebrationMixin, the score/win
takeover drawn by afl, football, hockey, nrl and soccer
(_draw_celebration_layout and the palette, backdrop, scenery, confetti,
crest and _fit_font steps, with their class constants), plus the colour
helpers (logo_palette, lift_color, cap_luminance, mix_color, ...). Only
the drawing: _start_celebration, _check_for_goal/_check_for_score,
_check_for_win and display() differ between the plugins and stay there.
- src/common/sports_fetch.py: SportsFetchMixin, the four SportsCore methods
identical in all nine scoreboards: _fetch_season_directly,
_background_fetches_espn_ranges, _needs_previous_day and
_wants_live_odds, with _LOOKBACK_CUTOFF_HOUR and _LIVE_ODDS_LOOKAHEAD.
_get_timezone, _extract_game_details and _fetch_data are as identical
and stay behind, for the reasons sports_shared gives (a per-plugin
import; the abstract contract); so does SportsUpcoming.__init__, since
no src/common mixin has a constructor.
- src/common/sports_card_wrappers.py: SportsCardWrappersMixin, the
seventeen sports_card delegations the eight game renderers carry (15 in
all eight, 2 in all but football, whose own versions override them).
_schema_font_size/_resolve_font_size look identical but read each
plugin's own _SCHEMA_PATH, so they stay.
Each mixin has no __init__ and creates no attributes (the host contract is
declared as annotations only), defines no name the mixins beside it
define, and documents the attributes it reads; a host-contract test
parses each and fails on an undocumented read. A method kept on a
plugin's class wins over the mixin's.
Tests: behaviour ported from the plugins' celebration, odds, lookback and
date-range tests against stub hosts carrying exactly the contract, with
crests drawn by the test (test_sports_celebration.py, test_sports_fetch.py,
test_sports_card_wrappers.py), and test_sports_stage3_parity.py, which with
LEDMATRIX_PLUGINS set compares every body with every plugin copy that is
left (58 pass against the plugins today; a copy that is gone counts as
adopted). All three modules are on the mypy ratchet, in
src/common/README.md, the CHANGELOG's Unreleased section and
SPORTS_UNIFICATION's module table. Nothing in core uses them yet.
Full suite: the same 67 failing test ids as main (Windows-only), 77 more
passing.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Bumps src.__version__ to 3.6.2 and turns Unreleased (#670, the favourite
check's false "season has finished" for list-calendar competitions between
rounds) into ## 3.6.2.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
On 2026-09-29 ESPN's uefa.europa scoreboard still showed the 17 September
matchday, so every event was past. Its calendar is a "list" of rounds
(League Phase to 30 Jan 2027, then the knockout rounds to the final), not
a match-day whitelist, and the league's season type is a soccer id rather
than 2/3, so neither 3.6.1 rule applied and the check said the season had
finished.
When every event is past, a round in a list calendar that has not started
yet now draws no conclusion. Only a round's start date counts: end dates
are padded past the last game (AFL's Grand Final round still had a day to
run three days after the Grand Final), and rounds in an offseason phase
(college football's All-Star week) are skipped. Season end dates are still
ignored, so PLL (season to 2027-01-01) stays "finished", as do the World
Cup and AFL. Of 28 live ESPN scoreboards only uefa.europa's message changes.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Bumps src.__version__ to 3.6.1 and records #667 (the favourite check's false
"season has finished") under ## 3.6.1; #667 had no CHANGELOG entry. Plugins
that drop their bundled favourite-check copy floor on 3.6.1.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(common): favourite check no longer calls a started postseason a finished season
The day after a regular season ends, ESPN's default scoreboard still
returns that last regular-season day, while leagues[0].season has moved
to Postseason. All events were in the past, so the check logged "the
season has finished" for MLB on 2026-09-29 while the upcoming manager in
the same process was showing TB's wild-card games.
When every event is past and the league is in a later in-season phase
(regular season or postseason) than all of the returned events, draw no
conclusion. The offseason is excluded, so a genuinely finished season is
still reported as finished.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(common): favourite check reports the next matchday between soccer rounds
Between matchdays ESPN's soccer scoreboard keeps showing the last one, so
every event is in the past and in the league's current phase, which the
postseason rule does not cover; the check said the Premier League season
had finished on 2026-09-29 (last games 20 September, next 10 October).
When the league calendar is a "day" whitelist, its entries are days with
games, so a future one is used as the next fixture. MLB's day calendar is
a blacklist and is not read that way; PLL's whitelist has no future days
and is still reported as finished.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Turns the CHANGELOG's Unreleased section into ## 3.6.0 and bumps
src.__version__, the value plugin ledmatrix_min_version floors compare
against. 3.6.0 ships the two modules from #665 (favorite_team_check,
sports_timezone); nothing else has changed since 3.5.0. src/common/README.md
and docs/SPORTS_UNIFICATION.md say 3.6.0 for them instead of Unreleased.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(common): favorite_team_check and sports_timezone, promoted from the scoreboards (sports consolidation stage 2)
Two new hardware-free modules, taken from files the scoreboard plugins carry
as copies:
- src/common/favorite_team_check.py: FavoriteTeamCheck(logger, leagues), the
seven byte-identical <sport>_favorite_check.py copies. Same code; the only
additions are two type annotations (for the mypy ratchet).
- src/common/sports_timezone.py: resolve_timezone_name(), resolve_timezone(),
system_timezone_name(), from the ten <sport>_timezone.py copies. They
differed only in the plugin label named in the nothing-resolved warning and
the write-back-bug values, which become keyword-only arguments
(plugin_label, writeback_fixed_in). Same resolution order and log text.
Tests are ported from the plugins' own (test_favorite_check.py,
test_schedule_note_uses_game_dates.py, test_timezone_resolution.py; the
timezone ones run once per plugin's values and pin the exact warning text).
Both modules are on the mypy ratchet, in src/common/README.md, the CHANGELOG's
Unreleased section and SPORTS_UNIFICATION's module table. Nothing in core uses
them yet.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(common): bdf_font and json_body shipped in 3.5.0
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(common): annotate the favourite check's deliberate except/pass for Bandit
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): fold Unreleased into 3.5.0 for the release
Every Unreleased entry (#605-#663) moves into the 3.5.0 section, grouped
with the existing 3.5.0 areas; new groups for Display and Vegas, Plugin
error reporting, Wi-Fi, Fonts and Removed. "## Unreleased" stays as an
empty heading.
Module list: add src/common/json_body.py (espn_dates imports it with a
fallback) and src/common/bdf_font.py; list the other modules new since
v3.4.0 as core-internal; add the new names in existing modules
(handles_espn_date_ranges, register_plugin_fonts(plugin_dir),
forget_manager_fonts). Record the src.common and plugin_system modules
#608 deleted.
Add entries for merged PRs that had none: #604, #605, #606, #607, #608,
#609, #613, #616, #618, #622, #625, #628, #630, #633. Note that three
scripts named in older 3.5.0 entries were later deleted by #607.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): list the hardware-free test under developer tools, not plugin modules
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Plugins import their own files by bare name (`from sports import ...`),
which resolves to the first directory on sys.path that has the file. The
loader added a plugin's directory only if it was missing, so on a reload --
a live re-enable from the web UI -- the plugin's directory stayed behind
every plugin loaded since, and its bare imports found their files first.
Seen on ledpi: re-enabling UFC with hockey running failed with "cannot
import name '_status_is_final' from 'sports'" (it got hockey's sports.py).
A loading plugin's directory is now always moved to the front. Every
scoreboard ships its own sports.py, so any of them was exposed on reload.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* ci: mypy ratchet -- keep type-clean modules clean
mypy-clean.txt lists the 71 modules under src/ that type-check clean;
scripts/check_types.py runs mypy (--follow-imports=silent) on exactly
those files and fails on any error or a missing/unsorted/duplicate entry.
A new "Type check (mypy ratchet)" CI job runs it with mypy 1.20.2 and
pinned stubs; the manual pre-commit mypy hook now runs the same script
(a local hook, so mypy sees the installed requirements like CI does).
35 modules were made clean with annotation-only fixes: hints, typing.cast,
TYPE_CHECKING imports, implicit-Optional defaults made explicit, and
annotations widened (never guards removed) where mypy called a defensive
isinstance check unreachable. No runtime behaviour change.
mypy.ini: numpy and orjson are treated as Any (follow_imports=skip, also
for stubs). numpy 2.3+ stubs use 3.12 `type` statements that mypy won't
parse at python_version 3.10, and orjson is optional, so seeing its stubs
made the result depend on whether it was installed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: annotate check_types.py's list-form mypy subprocess
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
A new "Web UI JS tests" job installs jsdom, starts the web interface in
emulator mode and runs test/js/run_all.js with REQUIRE_DOM=1, which makes a
DOM suite that can't run a failure rather than a silent skip. (The unit
suites were already covered through pytest.)
Two suites failed against main when run for real:
- test_tools_sections rendered the Tools partial without LEDEscape, which
base.html's app-early.js defines; it now installs it in beforeParse, and
supplies two sample Starlark apps (one id with a quote) when the server
has none, instead of assuming a device with apps and Pixlet.
- test_store_dom assumed the live registry had at most 48 plugins; it now
checks pagination whichever side of 48 it is.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): split PluginStoreManager into mixins
src/plugin_system/store_manager.py (2,977 lines) keeps the class, its
shared state, locks, the uninstall registry, directory lookup and
uninstall; its methods are split by area into:
- store_registry.py (_RegistryMixin): registry, GitHub metadata, search,
manifest validation
- store_install.py (_InstallMixin): install paths and dependencies
- store_update.py (_UpdateMixin): updates, rollback, local git state
Pure move: all 56 members are byte-identical (checked with ast) and the
assembled class has exactly the same attributes as before (checked at
runtime). PluginStoreManager is imported from store_manager.py as before.
Tests that patched shared modules (subprocess, requests, tempfile, shutil)
through store_manager now reach them through the module whose code they
exercise; a source-text contract test reads all store_*.py modules.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: annotate findings the split moved into new store modules
subprocess imports and a list-form git clone (no shell), and the config
template's placeholder token string -- existing code that Codacy reported
as new because it moved. Annotated with the repo's nosec/nosemgrep style.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: annotate the default-branch git clone the split moved
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
A game ESPN had no odds for is cached as {"no_odds": True}, so it isn't
re-requested on every update. On the next update get_odds() returned that
marker from the cache as if it were odds: a truthy dict that callers took
to mean the game had some. It's still a cache hit (its ttl decides when to
ask again), but get_odds() now returns None for it -- on the cache hit and
in the stale-cache fallback after a failed fetch -- as the plugins' bundled
copies already did.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): split api_v3/plugins.py by area
web_interface/blueprints/api_v3/plugins.py (3,285 lines) becomes:
- plugins.py: installed list, enable/disable, plugin actions
- plugin_store.py: install, update, uninstall, store, saved repositories
- plugin_config.py: config get/save, schema, reset
- plugin_assets.py: asset uploads and plugin static files
- plugin_health.py: health, metrics, limits
- plugin_operations.py: operation history, state reconciliation
- plugin_calendar.py: calendar credentials and auth
Pure move: all 44 functions and 38 route decorators are byte-identical
(checked with ast), URLs and endpoint names are unchanged (url-map test).
Each module imports only what it uses. Tests and config.py that reached
into plugins.py for moved names now import from the new module.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): keep exception text out of calendar responses; annotate moved code
The split made scanners report existing findings in the moved code as new:
- CodeQL: the calendar auth and calendar-list routes returned exception
text (redacted, but still derived from the exception). Both now log the
exception and return a fixed message pointing at the log.
- MD5 in the asset upload only makes a filename unique: usedforsecurity=False.
- pickle reads/writes the calendar plugin's own OAuth token (as before):
annotated. Token-status labels and a log line naming the secrets path are
false positives: annotated with the repo's nosec/nosemgrep convention.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): name uploaded assets with SHA-256 instead of MD5
The hash only makes an uploaded image's filename unique. Codacy flags MD5
even with usedforsecurity=False, and SHA-256 does the job as well; existing
files keep their names.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): keep the redacted exception detail in calendar errors
Reverts the calendar part of 5e695b7c. The project's policy
(test_no_api_v3_handler_discards_its_exception) is that an API error
carries the redacted exception detail -- describe_exception runs it
through the credential redactor -- so a failure is diagnosable from the web
UI. Dropping it for CodeQL broke that; CodeQL can't see the redaction, so
its two alerts here are false positives, like the existing ones on main.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
- BackgroundDataService: the session adapter retried connection errors 3x
inside each attempt of the service's own retry loop (up to 16 connection
attempts per request on a dead network). The adapter no longer retries;
ESPN date chunks, which bypass the loop and skip a failed chunk, get a
small connection retry of their own (_ConnectionRetryingSession).
- CI installs web_interface/requirements.txt. The brotli header test now
checks its intent (core never hand-sets br; requests may advertise it when
a decoder is installed) instead of failing whenever brotli is present.
- Every Discord link uses the LEDMatrix server's invite (RdrC37rEag).
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin dir resolver in routes, nmcli AP detection, daemon config reload, upload safety
- Route plugin lookups (installed list, update, recorded version, config
form, web UI pages) through the plugin manager's resolver so plugins in
ledmatrix-<id> directories work.
- Captive-portal detection also sees the nmcli fallback AP (cached).
- WiFi monitor daemon re-reads wifi_config.json when its mtime changes.
- Drop the AP check in disconnect_from_network that could never fire.
- LED status file per WiFiManager; config path falls back to this checkout.
- BDF font preview via src.common.bdf_font.
- Asset uploads validate every file before saving; metadata and calendar
credentials written atomically; no absolute path in the response;
asset delete answers 400 for a missing body.
- Coerce string booleans in plugin toggle, on-demand start and AP force.
- SSE broadcaster clears its thread handle before exiting.
- start.py log filter handles every exc_info form.
- Cleanups: unused plugins/fonts partial work, duplicate backup catch-alls,
raw-config error helper, update-route tidy, redundant imports.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): request BDF font previews now that the server renders them
The Fonts tab skipped the preview request for .bdf files because the server
used to refuse them; /fonts/preview now draws BDF with the shared loader.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): take the update route's plugin directory from a directory listing
CodeQL flagged the path built from the request's plugin_id (the id was
already validated with safe_path_component, which CodeQL doesn't model; the
same flow on main is alerts 738/739). The directory is now the entry of
plugins_dir matched by name, so nothing built from user input reaches the
filesystem; an id with nothing installed goes to the store manager, which
reports it not found as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): read the blueprint's plugin_manager defensively in _plugin_directory
_get_plugin_version now goes through _plugin_directory, which read
api_v3.plugin_manager directly; the attribute exists only once the app sets
it, so test_path_traversal_guards::test_a_real_manifest_is_read failed
when run on its own (order-dependent in the full suite).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugin-system): unload/update race, failed-load module cleanup, limits validation, schema lookup, install rollback, op-queue dedupe
- unload_plugin takes the per-plugin lock (5s bounded) before cleanup(),
and an update() that finishes after its plugin was unloaded no longer
sets the state back to ENABLED.
- A load that fails after import drops plugin_<id> and its submodules
and forgets its manager fonts, so a fixed plugin reloads new code.
- Resource limits are validated as non-negative numbers: 400 at
POST /plugins/limits, bad cached records ignored with one warning.
Route docstrings note health/metrics reset and limits only change the
web process's view.
- SchemaManager.get_schema_path resolves each search dir via
resolve_plugin_dir (manifest id, ledmatrix-<id>) before the literal
paths; plugins/ still before plugin-repos/. Misses cached 30s and
logged once at DEBUG.
- install_from_url sets an existing copy aside and restores it if the
move fails, under the per-plugin reinstall lock.
- Operation queue refuses a second pending op for a plugin and trims
_operations with history.
- get_vegas_render_width reads display_manager.width first.
- get_logger in store/schema/health/resource/saved_repositories;
UTF-8 reads in store_manager and state_manager.
- Docs: update_interval precedence (manifest over config) stated where
users are told to set it in config.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): build the limits 400 message from the field name, not an exception
CodeQL flagged str(e) flowing into the response. invalid_limit_field()
returns the offending field without raising, and limits_from_dict uses it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
- Plugin-supplied widgets load as /static/plugin-widgets/...js?v=<plugin
version>, so an update isn't hidden behind the year-long immutable cache.
- Fire-and-forget loadInstalledPlugins() calls catch the rejection it has
already reported, so the global handler no longer adds a second toast.
- Timezone picker renders again when the General partial is re-injected.
- Remove dead code: executePluginAction's six plugin-id fallbacks and
[DEBUG] logging, window.currentPluginConfig and every read of it, the
file-upload JSON delete branch, unused PluginAPI / PluginInstallManager /
PluginStateManager helpers, loadPluginWidgetsFromManifest, the stale
install_manager.js and LEDVisibility fallbacks, error_handler.js's global
escapeHtml, 13 unused CSS rules, and stale comments/no-op returns.
- pytz < 2027, psutil < 7 in requirements-test.txt, pytest-cov < 8.
- Pin anthropics/claude-code-action to the commit v1 resolves to.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Six Starlark fixes are on main via #535 and #537, but a follow-up commit
carrying tests for half of them was pushed to fix/starlark-pixlet-install
six minutes after #535 merged, so those tests never landed. This ports
them onto the api_v3 package split:
- a failed toggle write answers 500, and a loaded app is not flipped in
memory when the manifest write fails
- each manifest writer gets its own temp file; concurrent writes leave
readable JSON; no temp files are left behind
- a failed dynamic import of tronbyte_repository / pixlet_renderer does
not stay cached in sys.modules
- a failed save_config() leaves config and timing untouched and does not
re-render; a successful save still applies
It also logs when the timing update to the manifest is not persisted.
_update_manifest_safe answers False rather than raising, so the existing
except branch never saw that failure.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(core): font zip cache, monotonic timers, resolver back-off, and other core/common fixes
- font_manager: a .zip font URL is served as its extracted font after a
restart (the cached-file check returned the archive first); downloads
use requests with a 30s timeout into a temp file + os.replace.
- api_helper / sync_manager: rate-limit and heartbeat/leader timeouts use
time.monotonic(); last_request_time and the status file's ts stay
wall-clock. set_on_new_cycle docstring no longer claims core uses it.
- logo_helper: the placeholder uses the same scaled box as a real logo.
- permission_utils: one _sudo_bash_candidates() helper (with the sudoers
exact-argv rationale) shared by sudo_remove_directory, which now retries
the next bash path on a sudo refusal, and install_requirements_file.
- dynamic_team_resolver: failed/empty fetch backs off 5 min; duplicate
INFO log and contradictory docstring example fixed.
- element_style: scale default looked up through element aliases.
- background_data_service: cache-hit callback runs outside the lock.
- config_arrays: union-aware type check (["array","null"]); stale
dotToNested() reference removed.
- auto_update_setup: non-dict auto_update reads as off; temp result file
unlinked when the write fails.
- exceptions: constructors copy the caller's context dict.
- logging_config: StructuredFormatter json.dumps(default=str).
- error_aggregator: removed unused export_path/export_to_file/_auto_export.
- Docstrings: validate_file_upload max_size_mb, raise_on_errors.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(sync): retry the status-file rename like the other atomic writers
On Windows os.replace can fail with "Access is denied" while a scanner
briefly holds the target open; config_manager_atomic._replace already
retries that (and re-raises at once on other platforms). The sync status
writer called os.replace directly, which made
test_concurrent_writers_each_use_their_own_temp_file flaky on Windows.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
- DisplayManager.defer_update()/process_deferred_updates(): one lock around
every queue mutation (appends from the update thread were lost to the
render thread's filter/slice reassignments); callables run outside it.
- FontManager and element_style no longer cache BDF freetype.Face objects
process-wide (load_bdf_face caches them per thread); element_style's LRU
is locked against get/move_to_end vs eviction races.
- limit_refresh_rate_hz default is one constant, DEFAULT_REFRESH_LIMIT_HZ =
100 (the template's), for the library options, refresh_hz, the matrix
guard, Vegas and scroll_config. Previously a missing key capped the panel
at 90 while pacing assumed 100.
- Sync follower: the TCP thread queues the leader's scroll image; the render
thread swaps image/array/width in between frames.
- update_display() error log rate-limited (traceback first, then once a
minute with a count); swallowed DisplayController exceptions log at DEBUG.
- Root display_controller.py runs run.py via runpy.
- stream_manager: correct the RLock release comments; merge duplicate if.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The static trigger peeked at the front of StreamManager's segment buffer,
which continuous scrolling (the default) never advances -- it extends the
strip with take_next_group() -- so the same first segment was examined on
every frame. A STATIC plugin paused the scroll only if it was first, once,
at startup; otherwise it scrolled past as ordinary content. Swap mode had
the same problem for any STATIC plugin not first in its cycle.
The render pipeline now records a marker (strip column, plugin id) for
each STATIC plugin where the strip is built -- composition and every
extension -- shifts the markers when the scrolled prefix is trimmed, and
clears them on reset. The coordinator pauses when the scroll reaches the
next marker: a tuple comparison per frame instead of a lock, a plugin
lookup and a get_vegas_display_mode() call. take_next_group() no longer
renders STATIC plugins' content. The pause calls display() under the
plugin lock and is timed with the monotonic clock.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): stop the run loop spinning when no mode has anything to show
A mode whose display() reports nothing rotates to the next at once, with no
dwell. With every enabled mode empty (only a sports plugin in its
off-season, say) the loop went round with no sleep: on ledpi, 169% CPU and
~1,800 "No content" log lines every 10 seconds. After one full rotation of
empty passes it now pauses EMPTY_ROTATION_PAUSE (1s) per pass, servicing
plugin updates and returning early on on-demand or schedule changes; live
priority is still checked at the top of every pass, and the streak resets
as soon as any mode shows something.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): restart the loop if the empty-rotation pause starts on-demand; per-rotation streak
- an on-demand request serviced during the pause returned early into the
on-demand branch, which advanced past the mode just requested; the loop
now restarts when the pause changed the mode, on-demand state or schedule
- the streak is reset when the rotation changes (on-demand start/stop, a
plugin enabled or disabled), so a streak from one rotation can't make
another pause before its own modes are tried
- docstring: live content is picked up within the pause, not "at once"
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
- mypy.ini parses again (multi-line exclude and inline value comments made
mypy reject the file); the mypy pre-commit hook is manual-only until the
~500 existing errors in src/ are paid down, and CONTRIBUTING says so
- .gitignore: ignore all of config/ except the templates (ytm_auth.json and
others weren't ignored)
- .gitattributes: LF for .sh and .service
- claude-code-review: skip fork PRs, which have no secrets
- check_system_compatibility.sh: 3.13 supported, <3.10 an error
- docs/scripts drift: emulator guide, README API Metrics, route count,
docs index, scripts README; pyflakes nits in dev scripts
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
- Operation History: the plugin filter lists the installed plugin ids
instead of one option, "plugins" (Object.keys of {plugins: [...]}).
- Ctrl/Cmd+S submits the active tab's first visible form with
requestSubmit() (validation and onsubmit guards run) instead of a bare
Event on the first form in the document; skipped inside a modal dialog
and on tabs without a form.
- Overview "Check Updates" confirms like "Update Code", takes its button
explicitly (no implicit global event) and shows the server's message.
Both, and the Tools tab git pull, raise the restart-pending banner on
restart_required.
- Tools: toolsAction and diagnostics show the server's error message;
only a non-JSON body falls back to HTTP <status>.
- Installed list after uninstall: PluginAPI writes clear the throttler's
GET cache, a forced loadInstalledPlugins clears it too, and the
post-uninstall reload goes through refreshInstalledPlugins().
- Plugin widgets load from /static/plugin-widgets/ only (the other two
paths have no route).
- Raw JSON editor escapes the parse error; slider escapes value/min/max/step.
- Removed the unreferenced array-of-objects and key-value helpers from
plugins_manager.js, the textarea auto-resize and Ctrl+R handlers in
app.js, and a redundant ?v= on the plugins_manager.js script tag.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
- load_plugin: an on_enable() that raises unregisters the instance, so the
next load retries instead of returning True "already loaded".
- get_plugin_info: guard plugin.get_info(); one plugin raising no longer
breaks /api/v3/plugins/installed.
- plugin_state.json and the operation history are written with
atomic_write_text under their lock.
- plugin_loader: module-level lock serialises pip installs across the
parallel startup loaders.
- store_manager._install_via_download: extract dir cleanup moved to finally.
- Test doubles: draw_image() warns (DeprecationWarning; the real
DisplayManager has none), MockDisplayManager.draw_text accepts the real
signature's optional params, VisualTestDisplayManager logs draw errors at
WARNING.
- Docs/comments: compatibility.py method name, PluginState.LOADED meaning,
brittle schema count, why _report_skip_once uses setdefault.
- Remove unused PluginOperationQueue.get_active_operations().
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
- Vegas: a live-priority pause was only lifted from inside run_frame(),
which returns before that check while paused, so the ticker never came
back until a restart. run_iteration() now resumes it (the controller
only calls it when nothing preempts Vegas); start()/stop() clear the
pause state. Iteration length is timed with the monotonic clock.
- Dim schedule: a per-day disabled day now updates the minute-gate cache,
so brightness no longer flips back to dim within each minute.
- On-demand: a second request no longer overwrites the rotation resume
index with the first request's mode.
- Render pipeline: reset() drops the prepared group and deferred queue,
and a prefetch in flight across a reset discards its result.
- Sync: stop() removes the status file (and the controller's cleanup now
calls it), standalone removes a stale one at startup, and writes use a
unique mkstemp temp file.
- render_gate.swap_releases_gil() delegates to frame_timing.
- Stale docstrings/comments corrected.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(security): refuse unsafe plugin ids, keep secrets private, validate bodies
- install_from_url and the registry install's manifest rename refuse a
plugin id that is not a single safe name (no ../ out of plugins_dir).
- Uninstall and config reset refuse core config sections and ids with
path parts; uninstall of a plugin whose directory is gone still works.
- separate_secrets checks a field's own x-secret marker before recursing,
so object/array secrets no longer land in config.json.
- Backup restore creates missing secrets/wifi/ytm files with mode 640;
export skips non-object manifests and no longer collides on same-second
exports.
- SYSTEM_FONTS includes every bundled font from BUNDLED_FONTS.
- Raw config/secrets saves and validate_request_json require a JSON object.
- A blank max_dynamic_duration_seconds keeps the stored value; other values
are validated to 30-1800 instead of raising a 500.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(security): validate the id before install_plugin moves anything; claim backup names atomically
- install_plugin set aside plugins_dir / plugin_id before any id check, so
"../x" moved a directory outside the plugins dir (the rollback moved it
back, but only if the install path got that far)
- two exports finishing in the same second could both see a free name and
the later os.replace destroyed the first archive; the name is now
claimed with O_EXCL before the archive is swapped in
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix: six bugs found testing main on a real Pi (ledpi)
- Stopping the service now runs cleanup. systemd stops ledmatrix.service
with SIGTERM, whose default action ended Python before run()'s finally
block, so the update worker, Vegas and the panel were never torn down.
main() now turns SIGTERM into KeyboardInterrupt, the Ctrl-C path.
- "Now showing" no longer turns into "unknown". display_current_state was
only written on a mode change and the web UI reads it with max_age=120,
so a live game or a single plugin on screen for longer read as unknown.
It is republished every 30 s while unchanged.
- Switching Vegas on in the web UI works when it was off at startup. The
coordinator was only created at startup; the config watcher now flags it
and the render thread creates it.
- configure_web_sudo.sh finds reboot and poweroff in /usr/sbin. Run as the
web user it could not, silently dropped their rules and still said it
granted them, so the web UI's Reboot/Shutdown stopped working.
- check_system_compatibility.sh reports installed packages as installed.
`dpkg -l | grep -q` under pipefail failed when grep exited early.
- A network failure fetching GitHub repo info logs a WARNING, not ERROR.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): fixes found testing on a Pi
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test: run the Linux-only script tests correctly
The sbin-lookup test set PATH=/nonexistent and then could not find bash
itself; call it by absolute path. The dpkg-query stub read $4, but the
package name is the third (last) argument. Both now pass on a Pi.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): cover the follower, long-render and startup cases
Review follow-ups on the ledpi fixes:
- The pending Vegas start is applied before the sync-follower branch too
(_apply_pending_vegas_init), which skips _is_vegas_mode_active() while a
follower is connected but needs the coordinator for the leader's image.
- _service_pending_changes(), which runs inside Vegas iterations and long
screens, republishes a stale display_current_state as well; the main
loop alone could be away for a 240 s Vegas iteration.
- The SIGTERM handler is installed after DisplayController() is built, so a
stop during parallel plugin loading keeps the default immediate exit.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(config): normalize nullable arrays and objects instead of refusing them
`element_style._nullable` widens every `customization.modes.<mode>` override
with 'null' so a blank means "inherit the base", which turns a colour declared
"array" into ["array", "null"]. `normalize_config_values` only knew how to
convert null/integer/number/boolean out of a union, so a valid [0, 249, 0]
matched nothing and logged
Could not normalize field customization.modes.upcoming.odds_text.text_color:
value=[0, 249, 0], type=<class 'list'>, schema_type=['array', 'null']
The warning was the harmless half. It `continue`d past the single-type handling
below, where `prop_type == 'array'` coerces items, so a nullable array never had
its items normalized while a plain one did. A form posts numbers as strings, so
["0", "249", "0"] survived to the validator and was rejected with "Expected type
integer, got str" -- setting a per-mode colour in the web UI failed outright.
Every per-mode override of a structural or string type was exposed, across all
eight scoreboard plugins, not only colours.
Re-enter the single-type handling with the matched member rather than bailing,
accept a string that matches, and warn only on a genuine mismatch so the
diagnostic still reaches the validator.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ET8e5weDrb5Ju5QTKLU7zh
* fix(config): convert only integral numbers for integer array items
Review catch on the previous commit. Routing nullable arrays into the shared
item handling made its integer coercion reachable for them, and that coercion
called int(v) on any number: a client sending [2.5, 249, 0] for an RGB array
got 2 stored and a 200 back, so a wrong value was silently corrected into a
valid-looking one rather than refused.
Convert only genuinely integral values, at both the union-item and the plain
'array' item branch so the two cannot drift. A whole float -- 2.0, which is all
JSON can express for an integer -- still converts. This also settles an
inconsistency that predates the change: int('2.5') raises, so the string form
was always preserved and rejected while the numeric form was truncated.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ET8e5weDrb5Ju5QTKLU7zh
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
DisplayManager.offscreen() gives a thread its own canvas, so Vegas renders every plugin's ticker content on its prefetch thread instead of pausing the scroll for canvas-bound plugins on the render thread. A render gate (src/common/render_gate.py, vegas_scroll.prefetch_gate, on by default with the GIL-releasing binding) lets the prefetch thread run Python only while the render thread waits in SwapOnVSync: on hdpi, frames 2+ refreshes late fell eightfold and late frames overall from 0.90% to 0.60%. See docs/OFFSCREEN_RENDERING.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A 1:N-scan HUB75 panel lights the two rows either side of its middle at opposite ends of each refresh, so a scroll at one pixel per refresh shows a 1px step across the middle of every panel. While something scrolls at one frame per refresh, DisplayManager now shows the half whose seam row lights first one refresh behind the other (src/scan_order.py), which lines the two up again. Only for layouts whose row order is known; display.scan_order_compensation "off" disables it. Confirmed on hdpi before and after.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
src/common/frame_timing.py times every frame the display presents, whoever drew it, and writes cumulative counters to /dev/shm. scripts/frame_soak.py grades a running service (late frames, freezes, where the time goes) and scripts/render_bench.py the hardware and render path alone. A stall watchdog logs the stacks behind any scroll held up for 250 ms or more (LEDMATRIX_STALL_WATCHDOG_MS lowers that). See docs/SCROLL_PERFORMANCE.md, "Soaking a rig".
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Vegas scrolls a whole number of pixels per panel refresh, locked to SwapOnVSync, against the refresh the panel really holds (measured from swap gaps), instead of blending sub-pixel positions against the refresh cap. The web preview PNG is encoded off the render thread while scrolling, with writes ordered and retried. On hdpi, late frames fell from 6.3% to 0.7%. See docs/SCROLL_PERFORMANCE.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): tear down Vegas mode on controller cleanup
DisplayController.cleanup() never called VegasModeCoordinator.cleanup(),
so the Vegas teardown (stop, pipeline/stream reset, adapter cache drop)
was unreachable. Call it before the display manager is cleaned up, and
skip it when Vegas was never created.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(vegas): default max_cycle_duration to the documented 240s
The template, the web UI help, CONFIG_REFERENCE and the controller all
say 240, but the code defaulted to 600 in two places, so a config
without the key ran Vegas iterations 2.5x longer than documented.
from_config now falls back to the dataclass field defaults instead of
repeating each one, so the two copies can no longer drift, and the
controller's follower scroll-speed default reads VegasModeConfig's.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): let run.py -d show display_manager's DEBUG output
display_manager pinned its logger to INFO at import, overriding the root
level, so debug mode never showed its DEBUG lines. Use get_logger() from
src.logging_config like the rest of the core and leave the level to the
logging setup.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): run each startup validation check once
StartupValidator.validate_all() ran twice at boot, before and after the
plugin manager was created, so every config, cache, display and
systemd-unit warning was logged twice. The second pass now runs only the
plugin checks. Drop the commented-out raise_on_errors line.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(vegas): one INFO line per plugin-list refresh
StreamManager logged "=" * 60 banners and a line per plugin (INCLUDED,
SKIPPED, FETCHING CONTENT, SEGMENT CREATED) at INFO on every refresh and
fetch, i.e. at each cycle start and every 30s. Log one INFO summary of
the rotation per refresh and move the per-plugin detail, the weighting
breakdown and "no content this cycle" to DEBUG (the adapter still warns
when every content path fails).
Also drop the check/cross marks from the controller's log messages.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(vegas): drop the per-iteration static-mode plugin scan
run_iteration() rebuilt _static_mode_plugins on every iteration, asking
every plugin for its display mode and logging the set at INFO, but
nothing ever read it: static pauses are triggered by
_check_static_plugin_trigger() from the next segment. Delete it, the
coordinator's get_ordered_plugins() that only it used, and the
write-only _static_pause_plugin / _static_pause_start.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(vegas): remove the staging buffer that was never filled
StreamManager and RenderPipeline carried a double-buffer design that
nothing used: _staging_buffer was only ever cleared or swapped, so
swap_buffers() never did anything and should_recompose()'s
staging_count > 0 branch was dead, and _active_scroll_image,
_staging_scroll_image, _is_rendering, _last_frame_time and
_frame_interval were written but never read. Delete the machinery and
rewrite the docstrings around what actually carries updates:
_pending_updates, consumed by process_updates() in swap mode and
invalidate_pending_updates() in continuous mode.
should_recompose() no longer builds a buffer-status dict every frame.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): tidy the display controller without changing behaviour
- Import VegasModeCoordinator locally instead of through module globals
(there is no circular import to avoid).
- Drop hasattr() checks on attributes PluginManager.__init__ always sets
(plugin_executor, plugin_last_update, get_plugin_lock,
run_scheduled_updates*, stop_update_worker) and the dead "older
manager" fallbacks; keep the health_tracker None checks, now via
_health_tracker().
- Extract _display_once() for the per-frame display call both render
loops copied, _advance_on_demand() for the two on-demand rotations,
_reset_on_demand_fields() for the error and clear paths, and
_timezone() / _in_window() for the two schedule checks.
- Remove always-true conditions and the unreachable non-plugin else
branch in run(), and read _was_display_active / _last_published_mode /
vegas_coordinator directly now that __init__ declares them.
- Declare the follower render state in __init__, name its tuning
constants, add _follower_sign(), and share the 90/s sync send
interval with the render pipeline (SYNC_SEND_INTERVAL).
- Delete history narration and the "Opt #N" labels, fix the comment
that called _scroll_speed constant (hot reload updates it), and drop
a startup timing log that measured nothing.
- render_pipeline / plugin_adapter: read display_manager.width/height
as the properties they are, drop an empty TYPE_CHECKING block, an
aliased threading import and a duplicated `if result and
self.sync_manager:`.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): trim dead code from display_manager
- Add _new_canvas() for the image/draw/fontmode="1" setup that was
copied six times.
- Call resolve_double_sided() and compose_pixel_mapper_config() directly
instead of through a module alias and a passthrough method, and replace
the comment that said the passthrough read class attributes.
- Delete the unused _initialized flag and _ORIENTATION_ROTATE_DEGREES
alias (no core or monorepo reader; tests stop resetting the flag), the
test pattern's unreachable no-matrix branch (it only runs once the
matrix exists), `del old_image # help GC` (a no-op on a local), a
duplicated early return in process_deferred_updates, and stale
comments.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(vegas): remove unread fields and test-only helpers, fix docstrings
- ContentSegment: drop total_width, fetched_at, is_stale, image_count and
is_static, none of which is read.
- StreamManager: drop _current_index (never advanced) and the test-only
get_all_content_for_composition() and has_pending_updates();
VegasModeConfig: drop the test-only is_plugin_included().
- geometry.find_blank_cut() has had no production caller since the crop
moved to item boundaries; delete it and its tests.
- PluginAdapter: the _finalize docstring described separator_width
between every image, and _crop_to_budget's said cuts snap to the
nearest blank column; both now describe what the code does.
- Coordinator: the static-pause interrupt log no longer blames follower
mode for every interrupt, and set_update_callback names the callback
the controller actually wires.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(scroll): correct ScrollHelper comments and drop dead branches
- Four comments said the strip always starts with display_width of
blank; it does only when lead_gap is None (Vegas passes its own).
- Delete the "Width calculation mismatch" warning: the image is created
at the calculated width, so the two can never differ.
- Remove the two scroll_delay <= 0 fallbacks (which disagreed with each
other): set_scroll_delay clamps it to at least 0.001 and nothing in
core or the plugin monorepo assigns it directly.
- Trim the scipy history from the blend docstring.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(run): drop a redundant comment
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): log set_scrolling_state only when it changes
Vegas and scrolling plugins set the scrolling state every frame, so once
display_manager's DEBUG output became visible in debug mode it printed
"Scrolling state set to: True" about 120 times a second. Log only when
the value differs from the previous one; the state, activity timestamp
and frame hold still update on every call.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): display-vegas
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): keep Cache and Logs helpers out of each other's way
Both partials declared top-level showError and escapeHtml. Their scripts
run at global scope after every HTMX swap, so whichever tab was opened
last owned window.showError, and a Cache failure after visiting Logs
rendered into the Logs panel (and the other way round). Each script is
now an IIFE; Cache still exports deleteCacheFile for its row buttons.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): make the HTMX-failure fallbacks for tab panels actually run
- The "HTMX never loaded" fallback read appElement.__x.$data, which is
Alpine 2. The page ships Alpine 3, so the check was always false and
the Overview never loaded without HTMX. It now reads Alpine.$data().
- The Overview and WiFi panels used hx-on::htmx:response-error, which
htmx expands to "htmx:htmx:response-error", an event that never fires.
- loadTabContent sent requests with <body> as the source, so htmx fired
its events on <body> and no panel's hx-on handler ran at all. The
panel is now the source. htmx also resolves its promise on a 4xx/5xx,
and the panel was stamped data-loaded anyway, leaving a skeleton that
never retried; it is now stamped only when no responseError fired.
loadPluginsDirect, loadOverviewDirect and loadWifiDirect are merged into
one window.loadPartialDirect(id, url), which also runs the partial's
inline scripts before Alpine sees the markup (as htmx-config.js does on
htmx:afterSwap). The ~10 s "htmx never arrived" path in loadTabContent
uses it for every tab instead of four hard-coded ones.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): store and registry failures no longer wipe the Plugin Manager
showError replaced the whole #plugins-content with an error message, so
one failed store search, custom-registry install or saved-repository
call took the installed list, the store and every control with it, with
no way back short of reloading the tab. Those failures are now error
notifications. The full-panel message is kept only for a first load of
the installed list that failed (nothing to show yet); a failed refresh
of an already-rendered list is a notification too. showSuccess's
fallback branch, which wrote the message into innerHTML unescaped, is
gone: showNotification always exists.
The "Please try refreshing your browser" hint tested for the text
"Failed to Fetch", which no browser produces (Chrome says "Failed to
fetch", Firefox "NetworkError..."), so it never appeared. It now keys on
the failure itself: a TypeError from fetch(), or PluginAPI's
NETWORK_ERROR wrapper around one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): escape plugin action and install output on every path
executePluginAction escaped data.message and data.output when an action
failed but put data.message straight into innerHTML when it succeeded,
and set the OAuth step-2 button's innerHTML from the manifest's
step2_button_text. Plugin actions run plugin code, so that is plugin- or
server-controlled markup in the page. Both paths now escape, and the
button label is set with textContent.
The same pattern sat in the install-from-GitHub-URL status lines
(plugin_id, the server's message, and error.message, which can echo a
repository URL) and the custom-registry load error; those are escaped
too.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): file-upload widget owns the image list and schedule editor
plugins_manager.js loads after the widget bundle, so its older copies of
deleteUploadedFile, updateImageList, hideUploadProgress, formatDate,
openImageSchedule, toggleImageScheduleEnabled, updateImageSchedule{Mode,
Time,Day} and updateCheckboxGroupData replaced the widget's. They are
deleted; the widget files are the only definitions.
Before switching over, the two sets were diffed and fixed so nothing
regresses:
- The old copy labelled the schedule/delete buttons for screen readers
and lazy-loaded thumbnails; the widget now does both.
- The schedule button did nothing on a card rendered by
plugin_config.html whenever the image id is a UUID (every upload): the
template turns "-" into "_" in the editor's id, and neither JS copy
did. Both now use the template's rule.
- The widget's "keep the open editor open" copied the editor's innerHTML
into the new list. That dropped its event listeners and showed the old
values, so after the first change the editor looked live but ignored
input. A schedule edit now saves to the hidden input and updates the
card's summary in place without re-rendering the list; a list re-render
(upload, delete) rebuilds an open editor from the data. Editor controls
are routed by one delegated change listener, so there are no
per-element listeners to lose.
- The old deleteUploadedFile had a JSON branch that removed a
#file_<id> element and skipped the re-render. No template or script
renders such an element, and JSON uploads are listed through
updateImageList like images, so re-rendering (the widget's behaviour) is
the consistent one; the branch was not carried over.
- The template always renders the summary line (".image-schedule-summary",
"Always shown" when unscheduled) so an edit has a line to update.
The inline-handler test evaluated plugins_manager.js's updateImageList;
test_file_upload_widget.js now covers the widget's list and editor.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): delete the unused handleCredentialsUpload
Its last caller went when plugin_config.html switched credential uploads
to the file-upload widget's handleSingleFileSelect. Nothing in the web
UI, the tests or the plugin monorepo references it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): delete dead and shadowed front-end code
Nothing calls any of these (checked across web_interface/, test/ and the
ledmatrix-plugins monorepo, including hx-*/x-*/onclick attributes):
- app-shell.js: the Alpine methods refreshPlugins (it called a
nonexistent this.searchPluginStore), loadPluginConfig,
savePluginConfig, getSchemaPropertyType, escapeCssSelector,
formatCommitInfo and formatDateInfo, and the top-level copies of
savePluginConfig, getSchemaPropertyType, escapeCssSelector,
formatCommitInfo, formatDateInfo and togglePluginFromTab. Plugin config
forms save through hx-post in plugin_config.html.
- window.reconnectSSE (app-shell.js); window.updateArrayTableAddButtonState
(array-table.js).
- toggleNestedSection, defined twice (app-shell.js and
plugins_manager.js) and called from nowhere.
- plugins_manager.js: the window.initializePlugins wrapper around an
IIFE-local origInit that was always undefined, and __pluginDomReady,
which was written but never read.
- display.html's fixInvalidNumberInputs fallback: app-shell.js defines it
before any partial loads.
- base.html's window.loadCodeMirror and the two CodeMirror stylesheet
preloads, and the .CodeMirror rules in plugins.html. The raw JSON
editor is a plain textarea.
Also deleted: app-shell.js definitions that a later script always
replaced, so they never ran: executePluginAction (plugins_manager.js
assigns its own), uninstallPlugin and its pollUninstallOperation
(plugins_manager.js), and updateAllPlugins (install_manager.js).
vendor/codemirror stays: test/test_web_smoke.py still requests
codemirror.min.js as a sample static asset.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): call showNotification without checking it exists
app-shell.js defines window.showNotification (a stand-in that queues
until the notification widget loads) and base.html runs it, deferred,
before every other script that notifies: app.js, the utilities, the
widget bundle, plugins_manager.js, and all partials, which HTMX loads
after the page. The 81 `typeof showNotification === 'function'` /
`!== 'undefined'` checks, the `window.showNotification || console.log`
and `|| alert` fallbacks, and their else branches (alert(), console
output, and schedule.html's own hand-built toast) could never take the
fallback path. They are removed, as is fonts.html's second copy of the
queueing stand-in.
The stand-in in app-shell.js keeps its guard (it must not replace the
widget's implementation if load order ever changes), and BaseWidget's
public notify()/getNotificationFunction() keep their shape for widgets
that plugins ship.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one HTML escaper, window.LEDEscape
About 30 files each carried their own escapeHtml / escapeAttr / escHtml /
_esc / escapeJs. They disagreed: several (notification.js, display.html's
escapeHtml, operation_history.html, the app() stub) did not escape
quotes, google-calendar-picker.js and tools.html's escHtml left ' alone,
and some turned 0 into ''. Most were fine only because the quote-safe
widget copies were preferred at runtime.
window.LEDEscape now lives at the top of app-early.js, a blocking script
in <head>, so it exists before any other script runs:
html(v) & < > " ' as entities, null/undefined as ''
attr(v) the same, for call sites that want to say "attribute"
jsStringAttr(v) a JS string literal safe inside an inline handler
Every former copy is now a one-line name for it (kept so call sites do
not change), widgets included, with no fallback. plugins_manager.js
loses its four escapeJs wrappers (callers use jsStringAttr), the
duplicate escapeAttr and escapeHtml inside renderInstalledCards and
renderCustomRegistryPlugins, and the window.escapeHtml /
window.escapeAttribute exports, which nothing read.
addArrayObjectItem's fallback markup (with a sixth hand-written escape
chain) is gone too: window.renderArrayObjectItem is defined earlier in
the same file, so the fallback could not run. The unused escapeHtml
methods on the Alpine app (app-early.js stub and app-shell.js) are
deleted.
test_html_escaping.js now runs LEDEscape and every remaining name for it,
and fails if a hand-rolled escaper reappears anywhere in web_interface/.
Suites that evaluate slices of plugins_manager.js or widget files load
LEDEscape from app-early.js through test/js/led_escape.js.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): stop htmx re-running partial scripts after every tab load
htmx-config.js runs each swapped-in <script> itself on htmx:afterSwap
and meant to turn htmx's own script handling off with
htmx.config.allowScriptTags = false. It did that once, while setting up,
but base.html loads htmx with a dynamic <script>, so htmx was not defined
yet and the setting never applied. On every tab load htmx then tried to
run each script again in its settle phase, found it already replaced
(no parent node) and threw "Cannot read properties of null (reading
'insertBefore')" into the console, which also skipped the rest of that
swap's settle tasks.
The setting is now applied in the afterSwap handler, which always runs
after htmx exists and before htmx settles the same swap.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): show "--" for a system stat the server could not read
The stats stream and /system/status now send null for a metric they
cannot read (cpu_temp off a Pi, for one) instead of 0. updateSystemStats
built the header and Overview text as value + unit, so a null showed as
"null°C". CPU, memory and temperature, in the header and on the
Overview, now render "--" plus the unit for null or a missing field --
the same placeholder the page starts with, and what tools.html already
shows.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one Alpine accessor and one plugin-list signal
window.getApp() (app-early.js) returns the root <body x-data="app()">
component through Alpine's public Alpine.$data, or null before Alpine
has initialised it. It replaces the private el._x_dataStack[0] reads in
app.js, app-early.js, app-shell.js, settings-search.js, overview.html and
plugins_manager.js, the three local getAppComponent/appData/getAppData
copies, and the Alpine 2 el.__x.$data fallbacks, which Alpine 3 never
provides.
Publishing the installed-plugin list: one load set window.installedPlugins
and dispatched pluginsUpdated twice (loadInstalledPlugins, then
renderInstalledPlugins), then wrote into the Alpine component through
_x_dataStack[0] and called its updatePluginTabs() directly, and
app-early.js's global listener set window.installedPlugins a third time
and called updatePluginTabs() again. Now renderInstalledPlugins is the
one publisher: it sets window.installedPlugins and dispatches
pluginsUpdated once, and the full app()'s listener (app-shell.js) is the
receiver. The app-early.js listener only builds the tab row while the app
is not the full implementation yet. The "grid not loaded yet" case is a
normal state (Plugin Manager tab not opened), so it logs through
pluginLog instead of console.warn.
updatePluginTabs had a "Debounce" comment and clearTimeout over a timer
nothing ever set, and two identical branches; it now just calls
_doUpdatePluginTabs (app-early.js detects the full implementation by
that name in its source, which the new comment says).
app()'s unused baseComponent lookup is removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): reload the plugin list after installs and failed toggles
Several callers refreshed the installed list with
if (typeof loadInstalledPlugins === 'function') loadInstalledPlugins();
else if (typeof window.loadInstalledPlugins === 'function') ...
but loadInstalledPlugins is local to the plugin-manager IIFE and
window.loadInstalledPlugins is never defined, so from outside that IIFE
both tests were false and nothing reloaded:
- A failed plugin toggle left the switch drawn in the new state while
the data said the old one. It now re-renders from the reverted data.
The optimistic in-place edit also has to forget the grid's
last-rendered markup, or setGridHtmlIfChanged sees identical HTML and
skips the revert. A successful toggle still keeps the switch (and
focus) as drawn.
- Installing from a GitHub URL (the early handleGitHubPluginInstall),
installing or uploading a Starlark app, and toggling a Starlark app on
its config tab never refreshed the list, so the new app had no tab or
Installed badge until the page was reloaded. They now force a reload
through window.pluginManager.loadInstalledPlugins(true), and the
Starlark grid redraws when that finishes instead of after a fixed
500 ms.
- The Starlark uninstall inside the IIFE reloaded from the 3 s cache,
which could still hold the app; it now forces a reload.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): route debug output through debugLog
base.html defines window.debugLog, gated on localStorage.pluginDebug.
plugins_manager.js read the same key twice more into its own flags
(_PLUGIN_DEBUG_EARLY, and PLUGIN_DEBUG behind a pluginLog() wrapper), and
api_client.js's RequestThrottler had a separate `debug` property with a
setDebug() that nothing called. All of it now goes through debugLog. The
"functions defined" dumps with their ✓ lines, and two per-plugin
"enabled=" loops that ran on every render, are dropped; "[PLUGINS STUB]"
labels on code that has not been a stub for a long time read
"[PLUGINS]".
Ungated console.log calls that announced normal events on every page
load or action (settings search and tooltips registering, every toast
repeated to the console, the schedule pickers initialising, widget
registry unregister/clear) go through debugLog too. What remains on
console.log is the widget registry's on-demand LEDMatrixWidgets.debug()
dump and BaseWidget.notify's no-notifier fallback.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop waits and guards that could never fire
- handlePluginAction polled up to 10 x 50 ms for window.togglePlugin,
configurePlugin, updatePlugin and uninstallPlugin before calling them.
All four are defined when the scripts load, before any card can be
clicked, so the poll always succeeded at once; it now calls them.
The long thinking-aloud comment over the toggle state is replaced by
two lines on why the stored state, not the checkbox, decides.
- initializePlugins checked typeof on setupGitHubInstallHandlers and
applyStoreFiltersAndSort, function declarations in the same IIFE, and
wrapped window.checkGitHubAuthStatus(), which returns a promise with
its own .catch, in try/catch.
- searchPluginStore wrapped each "#store-count" update (a getElementById
and an innerHTML assignment) in try/catch four times; one
setStoreCount() helper does it. The store's post-render re-attach of
the GitHub token handler dropped its try/catch and existence checks
for the same reason.
- The load-time fallback outside the IIFE tested typeof
initializePluginPageWhenReady, which is IIFE-local and so always
undefined there; it calls window.initPluginsPage directly.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): delete two unused plugin-manager helpers
stopOnDemand (IIFE-local; the page's stop button calls window.stopOnDemand
from app-shell.js) and debounce had no callers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): document the plugin-config handlers templates call
validatePluginConfigForm, handleConfigSave, handleToggleResponse,
handlePluginUpdate and refreshPluginConfig each get a JSDoc naming the
attribute in partials/plugin_config.html that calls it and what the
return value means (only validatePluginConfigForm's matters: false
cancels the submit).
- The `if (!window.__pluginConfigHandlersInitialized)` wrapper is gone:
app-shell.js runs once per page, so it was never false. The block is
dedented one level; `git diff -w` shows the real change.
- The three handlers read xhr.responseJSON first. XMLHttpRequest has no
such property (it is jQuery's), so that branch never ran; one
xhrJson(xhr) helper parses responseText for all of them, with the same
fallbacks as before.
- runPluginOnDemand and stopOnDemand checked that plugins_manager.js's
openOnDemandModal/requestOnDemandStop exist; plugins_manager.js is on
every page, so they call them.
- fixInvalidNumberInputs had a stray "Notification helper function"
comment on top of its own; a leftover "section toggle ... duplicate
definition removed" note is gone.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): one toast per save, and a failed durations save says so
app.js's global htmx:afterRequest listener showed the server's message
for every htmx request, and every form and button that posts through
htmx (plugin config save/toggle/update, Display, Durations, General,
Schedule, Dim schedule, the Overview actions) also reports its own result
from hx-on after-request. Each save showed two toasts. The global
listener now stays quiet for a request whose element, or its form, has
its own after-request handler.
That exposed the Rotation & Durations form's handler, which read
xhr.responseJSON: XMLHttpRequest has no such property, so it always said
"Durations saved" in green, even when the save failed (the global toast
had been the only place the error showed). display.html already had a
correct version (2xx only counts as saved; the server's message wins;
its status may refine success but never overturn failure). That is now
window.showSaveResult(xhr, savedText, failedText) in app.js, used by the
Display, Durations and General forms; General's inline copy of the same
logic is gone.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(web): file headers and comments that say what the code does now
- plugins_manager.js, app-shell.js, app.js and app-early.js open with a
header: what the file owns, how base.html loads it and in what order
relative to the others, and the globals it defines. app-early.js's
app() stub also says why it exists and that, with app-shell.js now
loaded before Alpine, it does not run in practice.
- base.html's note on plugins_manager.js said it must load last to win
over same-named functions in app.js/app-shell.js; there are none left,
so it now gives the real reason (it uses everything loaded before it).
- Change-narration and "already defined at the top, no need to redefine"
notes are gone or rewritten as present-tense reasons; comments that
were wrong are fixed ("Toggle password visibility" over the function
that opens the token panel, "Insert before the closing </nav>" over an
appendChild, "(from v2)", the export note that still listed
escapeHtml). About forty comments that restated the line below them
are removed, and a second window.currentPluginConfig = null outside the
IIFE is dropped (the IIFE sets it).
- The file-upload, checkbox-group and custom-feeds widgets' render()
stubs say plainly that the widget is rendered server-side, instead of
"for now" / "placeholder for future client-side rendering".
test_plugin_action_delegation.js sliced the source up to one of the
removed notes; it now ends the slice at the next section header.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): keep the escapeHtml/escapeAttribute globals for plugin pages
6da77363 removed window.escapeHtml and window.escapeAttribute because nothing in core or the plugin monorepo read them. Plugin web UIs served through serve_plugin_web_ui and third-party plugin pages may still call them, so they come back as aliases of window.LEDEscape.html and .attr, defined in app-early.js before any other script runs. test_html_escaping.js checks the aliases exist.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): web-frontend
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): encode the image thumbnail path; match script tags case-insensitively
CodeQL flagged the upload widget building an <img> src from a stored path,
and the escaper test extracting inline scripts with a case-sensitive regex.
Each path segment is now URL-encoded (still a same-origin path, and correct
for names with spaces or
* fix(web): clear Codacy findings in the escaper, app shell and upload widget
- LEDEscape looks entities up in a Map instead of indexing an object.
- showNotification is declared as a global for app-shell.js.
- openImageSchedule checks the index is a non-negative integer and reads
the image with Array.prototype.at.
- The schedule editor calls escapeHtml directly and documents why its
innerHTML template is safe: every value is escaped or constrained.
The remaining rule hits are suppressed on that line with the reason.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): build the image schedule editor with DOM calls
Codacy does not honour inline suppressions, and the editor's innerHTML
template kept tripping its XSS rules even though every value was escaped.
The editor is now built with a small element helper (createElement and
setAttribute), so no value is ever parsed as HTML, and the file's own
escapeHtml goes away.
Also for Codacy:
- LEDEscape.attr is its own function rather than a second name for html.
- The tab loader records a failed load on the panel (data-load-failed)
from a named handler, instead of a closure over a local flag.
The fake DOM in test_file_upload_widget.js gains append/replaceChildren,
its hostile-id check now asserts the id arrives as attribute data with no
innerHTML anywhere in the editor, and test_html_escaping.js drops the
file-upload.js escaper it no longer has.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): schedule editor helpers as plain functions
Codacy's lint flags arrow functions held in local constants and a forEach
callback that returns a value. The editor's pieces are now named function
declarations (displayStyle, scheduleModeOption, scheduleRangeTime,
scheduleDayTime, scheduleDayRow) taking what they need as arguments, and
the element helper loops with for...of. htmx is declared as a global in
app-shell.js. Output is unchanged; test_file_upload_widget.js passes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): drop repeats from uniqueItems lists before validating a plugin save
dedup_unique_arrays lost its only caller in #330, so submitting a value a
uniqueItems list already holds (a stock symbol saved once and posted again)
failed the whole save with a validation error. _prepare_plugin_config_for_save
runs it again just before validation, which covers both POST /plugins/config
and plugin sections posted to /config/main.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): /health counts the discovered plugins and logs the checks it fails
The plugin check counted plugin_manager.get_available_plugins(), which
PluginManager does not have, behind a hasattr guard that made plugin_count 0
on every device. It now counts the discovered manifests, discovering first
when nothing has been scanned yet.
The config, plugin and hardware checks answered "see logs for details"
without logging anything. Each now logs a warning with the traceback.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): store refresh no longer claims a commit-metadata refresh
POST /plugins/store/refresh read fetch_commit_info (or fetch_latest_versions)
only to append "(with refreshed commit metadata from GitHub)" to its message.
It never fetched any: the route re-downloads the registry and nothing else.
search_plugins takes the flag, but it reads commit info through its cache,
so passing it on would not refresh anything either. The flag is ignored now
and the message says what happened.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): refuse a malformed Vegas plugin order instead of clearing it
A vegas_plugin_order or vegas_excluded_plugins value that was not JSON, or
not a list, was stored as [] and the save answered 200, so a bad value wiped
the saved order or exclusions. Both now answer 400 and save nothing, the way
plugin_rotation_order already did; the three share one parser. A list that
holds anything but plugin-id strings is refused as well.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): per-plugin health and metrics read the display service's latest
GET /plugins/health/<id> and /plugins/metrics/<id> called get_health_summary
and get_metrics_summary without force_reload, so they answered with whatever
the web process read first and kept in memory, while the display service kept
writing newer state. They now pass force_reload=True, as the list routes do.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin config reset saves through the shared atomic save
POST /plugins/config/reset called config_manager.save_config directly, so it
took no backup, and a failed write escaped as an unhandled exception. It then
handed on_config_change the raw stored section, not the prepared config a
loaded plugin runs with. It now saves through _save_config_atomic with a
backup, answers CONFIG_SAVE_FAILED when that fails, and notifies with
_prepared_plugin_config, as POST /plugins/config does.
POST /plugins/toggle carried its own copy of _save_config_atomic's
save_config_atomic-or-save_config fallback; it calls the shared helper now.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): one reading and one "unavailable" for each system metric
system_metrics.collect_system_metrics() promised None for a metric it could
not read, but returned cpu_temp as 0 off a Pi, and the whole no-psutil
fallback as zeros. GET /system/status measured the same numbers a second time
with its own code, and answered None there. Now both come from
collect_system_metrics(), and "unavailable" is None everywhere.
/system/status keeps its 0.1s CPU sample and its 10s cache, and gains
nothing it did not already send. Two differences: without psutil it answers
200 with null metrics instead of 503, and a disk it cannot stat is null
instead of a 500.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): /display/current sends the snapshot as-is and logs a failed read
GET /display/current PIL-decoded the preview snapshot and re-encoded it before
base64-ing it, spending CPU on the Pi to send the same picture, and dropped
any failure with `except Exception: pass`. The /stream/display SSE stream
already passed the PNG's bytes straight through.
Both now read through web_interface/display_preview.py and answer with the
same payload. A missing snapshot is still a null image; any other read
failure is logged as a warning. /health reads the snapshot path from the same
module.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one helper puts a submitted plugin config's lists back
The plugin-config save turned position-keyed dicts ({"0": ..., "1": ...})
back into lists in five copies: four in the form path's
fix_array_structures (whose prefix branches never ran, since no caller
passed one), and _fix_json_arrays on the JSON path. It then force-fixed
the news plugin's feeds.custom_feeds by name, in case the generic pass had
missed it. src/web_interface/config_arrays.coerce_array_shapes now does it
for both paths, custom_feeds included. ensure_array_defaults duplicated
_fix_none_arrays and is gone.
In the same function: the union-type re-checks that the null handling
above them made unreachable, the "(temporary)" random_seed debug log, and
a commented-out log line are removed. A failed validation is logged once
as a warning, not four ERROR lines and a WARNING.
Element types are left to normalize_config_values, which already converted
them for both paths. One difference: the form path no longer adds an empty
{} for a nested object the post left out that has no defaults.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): import at module top and log through the module logger
The web_interface.cache imports in config.py and fonts.py were wrapped in
`except ImportError` fallbacks. It is an in-repo module that imports nothing
from the project, so it cannot fail to import; it is imported once at module
top, as system.py now does. cache.py's docstring said blueprints import it
lazily "to avoid circular imports"; it now says why that is unnecessary.
Five logging.error calls in the dim-schedule GET and three logging.warning
calls in plugins.py went to the root logger; they use the module logger.
Function-local re-imports of json, os, shutil, logging and Path, all
already imported by the module, are gone. The `import os` inside two except
blocks of save_plugin_config also made os a local name for the whole function.
execute_plugin_action's step-1 handler gets a comment saying why it stays:
it looks like a copy of the blueprint handler, but without it a
TimeoutExpired from the plugin's script would reach the route's own
`except subprocess.TimeoutExpired` and be answered as a 408.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): app.py loses dead CSRF and reconciliation state, comments fixed
- csrf was always None, so `if csrf: csrf.exempt(...)` never ran, and its
note that the api_v3 blueprint "is exempted above" named an exemption that
does not exist. Both are gone; the reason there is no CSRF protection stays,
shortened.
- The SSE rate-limit comment called the default "tight" at 20 per minute. The
default is 1000 per minute and the streams' 200 is the tighter one; the
comment now says so. The limits are unchanged.
- _reconciliation_done was written and never read. The docstring that
explains why reconciliation runs once keeps its reason, in the present
tense.
- Removed: a dangling "import cache functions" comment with no import under
it, a "security check ... within project_root" label on an existence check,
the "(simplified version)" narration, and the note that no redirect route is
needed. The preview loop's sleep comment no longer mentions a PIL encode
that the loop does not do.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(web): api_v3 comments name the package __init__, not a _common module
Every route module's docstring said the shared blueprint comes "from
._common", a module the package split never created; they name the
package __init__. The PROJECT_ROOT comment described the path from
_common.py; it now describes this package and keeps the incident it
guards against. The "(corrected) in this commit" note in
resolve_pull_command and the /health comment the split's mechanical
time -> _pkg.time rewrite garbled ("Stamp the start _pkg.time") read
correctly again.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop hasattr checks for attributes PluginManager always has
PluginManager.__init__ sets health_tracker and resource_monitor (to None
until they are configured), so the seven
hasattr(api_v3.plugin_manager, ...) guards in the health, metrics and limits
routes were always true. The falsy checks that do the work stay.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): pages_v3 dispatches partials from a dict with one error handler
load_partial chose a loader through a fourteen-branch if/elif, and thirteen
of the loaders then wrapped themselves in the same try/except, logging
"Error loading partial" without saying which. The route now looks the name up
in _PARTIAL_LOADERS and has the one handler, which logs the partial's name.
The loaders just render. _load_tools_partial keeps its own messages. The
search index's _partial_html already catches a loader that raises.
serve_plugin_web_ui repeated _plugin_dir_for inline (containment plus the
ledmatrix- prefix fallback); it calls it now. Also removed: the unused
markupsafe.escape import, function-local json/Path re-imports, and unused
exception bindings.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): remove unused imports, locals and a try that cannot fail
- get_error_aggregator was imported by the api_v3 package and used by no
one; seven names config.py imported, and Path in misc.py and logging in
plugins.py, likewise.
- branch_info in install_plugin was built and never logged; test_config in
/health was bound and never read (the load_config call is the check).
- An f-string with no placeholders in the asset upload route.
- _installed_plugin_ids wrapped list(manifests.keys()) in try/except;
_discovered_plugin_manifests always returns a dict.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): start.py logs its startup lines and drops unreachable branches
The startup banner went to stdout with print(); it goes through a logger
now, which the app import has already configured, so it reaches the journal
with a level and timestamp like every other line. The "no addresses" branch
is gone: get_local_ips() always returns at least "localhost".
The except around app.run re-raised "only if it's not a client
disconnection error" from inside the branch that had just established it
was one, so that raise could not run. It is one check now, on a named
tuple of the errnos, which the werkzeug log filter uses too. The comment
on threaded=True counts three SSE endpoints, which is how many there are.
Trailing whitespace is stripped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): save_main_config names its General fields once
The General tab's field names were listed twice, once to detect a General
form post and again, with four more, to keep the remaining-keys merge from
storing them as top-level keys. GENERAL_FIELDS and _MAPPED_TOP_LEVEL_FIELDS
hold them now, and the four per-section skip checks are one set.
The comment on that merge said plugin configs are handled "here too", and
"(including plugin keys)". Plugin sections are handled and removed from the
body before it runs; the comment says so.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): plugin directories come from the plugin manager only
Six lookups fell back to PROJECT_ROOT/plugins/<id> when there was no plugin
manager: GET /plugins/config's of-the-day data, POST /plugins/action, the
plugin static-file route, the calendar credentials upload and the calendar
OAuth routes. The loader never scans plugins/ (PluginManager.discover_plugins
reads only the configured directory, plugin-repos by default), so what they
found there was a plugin that never runs. _plugin_directory() asks the
manager and answers None without one, which each route already reports as
"not found".
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): web-backend
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(errors): record the exception's own stack trace
record_error() called traceback.format_exc(), which only sees an
exception while its except block is running. plugin_executor records
exceptions caught on a worker thread after that block has ended, so
every trace on /errors read "NoneType: None". The trace is now built
from the exception's __traceback__. The executor's log call had the
same problem with exc_info=True and now passes the exception.
record_error() also merged LEDMatrixError context into the caller's
dict in place; it now works on a copy.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(wifi): point at configure_wifi_permissions.sh instead of a sudoers list
The module docstring told users to grant NOPASSWD sudo on iptables and
ip. configure_wifi_permissions.sh refuses those grants on purpose: a
wildcard rule for either runs an arbitrary program as root. Point at
the script and say why it leaves them out.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(wifi): disconnect finds the saved profile by SSID
disconnect_from_network() asked `nmcli -f NAME,802-11-wireless.ssid
connection show` for the profile to take down, but nmcli rejects that
column for `connection show`, so the lookup always failed and only the
device was disconnected. The per-profile lookup _connect_nmcli() already
used is now _find_profile_for_ssid(), and both callers share it. It
also splits terse output on the last colon and unescapes "\:", so a
profile name containing a colon is found.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(wifi): write wifi_config.json atomically and report a failed save
_save_config() opened the file for writing in place and swallowed any
error, so a wifi_config.json left owned by root made the web toggle for
auto-enabling AP mode report success while nothing was saved, and a
crash mid-write could truncate the file. It now uses atomic_write_json,
which also keeps the file's owner and shared group when root saves it,
and returns False on failure. POST /wifi/ap/auto-enable answers 500 in
that case.
The file is now written with indent=4, like the other config files.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(fonts): resolve plugin:// fonts in the plugin's own directory
FontManager looked for a plugin's bundled fonts under Path("plugins") /
plugin_id: relative to the process cwd, and not the default install
directory (plugin-repos/), so a manifest's plugin:// fonts never loaded.
register_plugin_fonts() takes an optional plugin_dir, and PluginManager
passes the directory it loaded the plugin from. Callers that omit it get
a lookup in the configured plugin_system.plugins_directory, then plugins/,
resolved against the install root.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(api-helper): cache responses for the requested cache_ttl
APIHelper.get(cache_ttl=...) and set_cache(ttl=...) dropped the ttl on
the claim that CacheManager does not support one, but CacheManager.set()
takes a ttl, stores it with the entry, and both cache tiers honour it
over a reader's max_age. Without it every response expired after the
300-second default read age, whatever the plugin asked for. The ttl is
now passed through, and the cache read passes cache_ttl as max_age for
entries written without one. The class docstring describes what the
helper actually does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(style): one scale range for the schema, element_scale and LogoHelper
The generated Scale field allowed 0.1 to 10, element_style's reader
capped at 10 with no floor, and LogoHelper accepted 0.05 to 8 and reset
anything else to 1.0. A logo scale of 9, which the form accepts, drew at
the shipped size.
MIN_ELEMENT_SCALE / MAX_ELEMENT_SCALE (0.1, 10.0) in src.element_style
are now the schema bounds and the clamp every reader applies through
coerce_scale(): a positive number outside the range is clamped, and
anything that is not a finite positive number means the default. That
also stops element_scale() passing NaN through, since min(nan, 10.0)
is nan.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(logos): placeholder lands at the requested path; empty logos list
download_missing_logo() wrote its fallback placeholder to
<normalize_abbreviation(abbr)>.png in the logo directory rather than to
the logo_path the caller passed, so it could return True while nothing
existed where the plugin looks (e.g. "TA&M.png" vs "TAANDM.png").
create_placeholder_logo() takes an optional filepath, and
download_missing_logo passes the requested one.
download_missing_logo_for_team() only caught KeyError, so a team whose
"logos" list is empty raised IndexError; it now treats KeyError,
IndexError and TypeError as "no logo URL".
The placeholder is drawn with PLACEHOLDER_SIZE / PLACEHOLDER_BG, the
constants is_placeholder_logo() recognises it by, instead of repeated
literals.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(fonts): resolve bundled font paths against the install root
TextHelper's default font_dir, the logo placeholder's font and
FontManager's font_overrides.json were all relative to the process cwd,
so a process started anywhere but the install root (the plugin safety
harness, a manual run, a unit without WorkingDirectory) drew with PIL's
default face and read no overrides. They now go through
font_layout.resolve_asset_path; the overrides file sits in the install
root's config/.
The resolver docstrings described an order the code does not follow:
resolve_asset_path never consults the cwd, and sports_shared's
_resolve_font_path tries the cwd first. Both docstrings now say what
the code does, and _resolve_font_path calls resolve_asset_path instead
of probing FontManager for it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(sync): the web UI reads the sync status file the display writes
sync_manager writes its status to tempfile.gettempdir(), but
GET /api/v3/sync/status read a hardcoded /tmp/led_matrix_sync_status.json
and defaulted the port to a literal 5765. Wherever TMPDIR is set (or on
any non-/tmp host) the page only ever showed "starting". The endpoint now
uses sync_manager.STATUS_FILE and SYNC_PORT.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(http): the rankings resolver sends the project's User-Agent
DynamicTeamResolver fetched ESPN rankings with a bare requests.get, so
it sent python-requests' default User-Agent, which ESPN rejects; the
AP_TOP_N favourites then resolved to nothing. It now sends
DEFAULT_HTTP_HEADERS. BaseOddsManager carried its own copy of the
User-Agent string and now uses the same shared headers (which also adds
Accept-Language).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(backup): record the core release and read the configured plugin dir
The manifest's ledmatrix_version came from a VERSION file that does not
exist, then from .git/HEAD: a 12-character sha, or "ref: refs/he" when
the branch's ref was packed. It is now src.__version__.
list_installed_plugins() scanned a hardcoded plugin-repos/, so on an
install whose plugin_system.plugins_directory points elsewhere, plugins
missing from plugin_state.json were left out of the backup. It now reads
the configured directory from config/config.json, defaulting to
plugin-repos.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(startup): report a missing display section once
A config without a display section produced three errors for the one
problem ("Missing required configuration key: display", "Display
configuration is missing or empty" and "Display configuration is
missing"), and an empty one produced two. _validate_config now reports
it once, as a missing key or an empty section, and
_validate_display_config leaves it to that.
The module docstring said the validator fails fast; nothing in the
display service calls raise_on_errors(), so it now says the errors are
reported and startup continues.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(wifi): share the copied blocks and name the AP constants
- _parse_nmcli_wifi_list() is the one parser behind _scan_nmcli and
_scan_nmcli_cached.
- _verify_connected(), _wait_for_device_idle(), _failsafe_ap() and
_mark_forced() replace blocks that were pasted two or three times in
the connect and enable-AP paths. The device-idle wait now checks
before its first one-second sleep instead of after it.
- _check_command() calls _find_command_path() instead of repeating it.
- AP_IP, PORTAL_PORT, AP_PROFILE_NAME and AP_PROFILE_NAMES name values
that were spelled out 14, 12, 8 and 2 times; the two deletion loops
now walk the same tuple. The iwconfig status path compares the AP
address exactly: startswith() also skipped 192.168.4.10-19.
- Dropped a second WIFI.SIGNAL query that repeated the first, a no-op
"if ssid: continue", the try/except around _connect_wpa_supplicant's
constant return, and a second save of a scan scan_networks already
saves.
- _ensure_wifi_radio_enabled's docstring says it returns True when the
radio state cannot be read at all.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(config): drop dead branches and history comments in ConfigManager
- The module docstring pointed plugin authors at update_plugin_config(),
which does not exist; it now names save_config_atomic() and
save_raw_file_content().
- load_config's FileNotFoundError handler tested the message for
"config_secrets.json", but a missing secrets file is handled where it
is read, so only config.json reaches it; the check is gone.
- save_raw_file_content's `file_type == "main" or "secrets"` guard was
always true (anything else raised earlier).
- get_raw_file_content('secrets') already returns {} for a missing file,
so the os.path.exists() in front of two calls to it is gone.
- Comments that narrated earlier behaviour are rewritten as what the
code does now.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(background-data): present-tense comments, drop unused API
- Comments that told the history of each fix (what "used to" happen,
"the old per-delivery release") now state the invariant the code keeps.
- get_statistics() no longer reports a constant 'queue_size': 0, and the
uncalled clear_completed_requests() is gone (_cleanup_completed_requests
does that job on every completion). Neither is referenced in core, the
web UI or the plugin monorepo.
shutdown_background_service() has no production caller either, but it
is the only way to tear down the get_background_service() singleton,
which the tests rely on, so it stays.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(odds): drop the unread cache_ttl and merge the odds_data branches
BaseOddsManager loaded base_odds_manager.cache_ttl from config and never
used it: cached odds live for the update interval (get_odds' ttl=interval).
No core or monorepo code reads the attribute, so it is gone along with
its log line. The two consecutive `if odds_data:` blocks are one.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(backup): one table for the single-file sections
config, secrets, wifi and ytm_auth were each spelled out in create,
preview, validate and restore. _SINGLE_FILE_SECTIONS lists them once,
with the RestoreOptions flag that restores each, and all four walk it.
Restore error messages keep their wording ("Failed to restore
<file name>").
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(fonts): drop FontManager's write-only state and duplicate logs
- fonts_config, font_metadata and font_dependencies were written and
never read; the performance_stats keys font_load_times, render_times,
total_renders and the per-call "resolve" timings
(_record_performance_metric) likewise. get_performance_stats() reads
only the counters that remain. Nothing in core or the plugin monorepo
references any of them.
- A failed BDF load was logged twice, by _load_bdf_font and again by
get_font; get_font's line is the one kept.
- Removed "NEW:" and commented-out cozette entries, the "Copy font to
assets/fonts" comment on code that copies nothing, and local imports
of names the module already imports. The deprecated add_font() now
resolves assets/fonts against the install root.
The @deprecated methods stay.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(text-helper): cache loaded fonts; drop the pre-textlength fallback
TextHelper declared _font_cache, cleared it and reported its size, but
never stored anything in it. load_fonts() now keeps each (file, size)
it loads there, so clear_font_cache() and get_font_cache_stats() mean
what they say and repeated load_fonts() calls reuse the fonts.
get_text_width() no longer catches AttributeError for Pillow releases
without ImageDraw.textlength; requirements.txt pins Pillow>=12.2.
The class docstring describes what the helper does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(common): fix wrong docstrings in api_helper, permission_utils, snapshot_policy
- permission_utils called 0o2775 "sticky bit"; the 2 is setgid, which is
what makes new files take the directory's group.
- snapshot_policy pointed at web_interface/blueprints/api_v3.py, which
is a package now; the health check is in api_v3/misc.py.
- APIHelper.clear_cache() lost a history note and a fallback to a
clear() method that neither CacheManager nor the testing
MockCacheManager has. The session headers are built from
DEFAULT_HTTP_HEADERS instead of a copy of them, and the module
docstring says what the module offers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(sports): present-tense comments in the shared scoreboard renderers
- sports_scroll and sports_game_renderer comments that referred to "this
PR", "the old flat 128px card" or what the renderer "previously" did
now describe the current behaviour and its reason.
- The block explaining why non-finite settings are rejected sat above
_score_reserve_width; it describes _center_gap_width and now lives in
it.
- unshare_element_fonts wrapped its import of font_layout.load_truetype
in an `except ImportError` that cannot fire inside core; the import
stays at call time so tests can spy on the pinned loader.
- sports_card docstrings that told the history of a fix say what the
code does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(sports-shared): drop dead code, name the ESPN limit
- _get_weeks_data asked for limit=1000, which fetch_espn_scoreboard
clamps to ESPN_MAX_LIMIT anyway; it now names that constant. Its
unused `immediate_events = []` is gone.
- _get_season_schedule_dates() returned ("", "") and has no caller in
core or the plugin monorepo.
- _should_log keeps its warning_type parameter (part of the inherited
signature, though nothing in core or the monorepo calls it) and its
docstring says the cooldown is shared across types.
- An unused ImageFont import is gone.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(sync): one follower-mode switch, shared panel defaults
- The class docstring said the leader sends PNG frames. Frames go over
UDP as raw RGB; PNG is only the Vegas scroll image sent over TCP. It
now describes both paths.
- _enter_follower_mode() replaces the two copies of "note the leader,
switch from standalone to follower, log, write status" in the frame
and scroll-position handlers.
- The rows/cols fallbacks use DEFAULT_ROWS / DEFAULT_COLS from
src.display_geometry, as chain_length already did.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(style): drop _layout_axis, name the layout group title
- ElementStyleResolver._layout_axis() had no caller in core or the
plugin monorepo.
- _element_block_from_spec checked spec['size'] was a dict again after
size_spec already had; it reads size_spec.
- The "Layout Offsets" title written into three generated schema blocks
is _LAYOUT_TITLE.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(logo-helper): say what the placeholder draws; name the 1.5 box factor
- _create_placeholder_logo's docstring said it draws the team
abbreviation; it draws an outlined grey box and nothing else. The
docstring says so, and the "in a real implementation you'd want text"
comments are gone.
- The 1.5 x panel default logo box, written out six times, is
DEFAULT_LOGO_BOX_FACTOR.
- ImageDraw is imported with Image at the top of the module.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(logos): drop dead code and a duplicate regex in logo_downloader
- _SAFE_LEAGUE_CODE_RE was the same pattern as _SAFE_LEAGUE_RE; both
checks use the one.
- get_logo_filename_variations reassigned the TA&M case to the list it
already had; the function returns the two names directly.
- _get_team_name_variations() had no caller in core or the plugin
monorepo.
- fetch_single_team's docstring was copied from fetch_teams_data; a log
message read "for{team_id}".
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor: drop the Pillow<9.1 resample shim and a catch-and-reraise
- adaptive_images fell back to Image.LANCZOS/NEAREST for Pillow < 9.1;
requirements.txt pins Pillow>=12.2. RESAMPLE_LANCZOS and
RESAMPLE_NEAREST keep their names (src.common re-exports them).
- CacheManager.save_cache caught CacheError only to re-raise it; the
disk write is now called directly, with the same result.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(api-helper): stop the real CacheManager's cleanup thread
The cache-lifetime tests built a CacheManager and left its cleanup
thread's class-wide claim on the directory in place, which broke
test_cache_cleanup_thread_ownership when it ran later in the session.
The fixture now stops the thread on teardown.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): core-common
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(store): don't read a ZIP-installed plugin's remote from the LEDMatrix repo
update_plugin looked up remote.origin.url with `git -C <plugin> config
--local` for plugins that are not git checkouts. Under plugin-repos/ git
walks up to the enclosing LEDMatrix repository, so the lookup returned
LEDMatrix's own URL and a plugin missing from the registry was
"reinstalled" from the LEDMatrix repo. Only ask git when the plugin
directory has its own .git, the test _get_local_git_info already uses.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(schema): report each missing required field once, by name
validate_config_against_schema ran its own required-fields loop after
Draft7Validator.iter_errors, which already yields one `required` error
per missing field, so every missing top-level field was listed twice.
The validator's copy also printed the schema's whole `required` list
("Missing required property '['api_key', 'city']'") instead of the field.
Drop the loop and take the field name from the error itself.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(store): stop mangling repository URLs that contain ".git"
install_from_url and fetch_registry_from_url cleaned URLs with
`rstrip('/').replace('.git', '')`, which removes ".git" anywhere:
https://github.com/user/my.github.io became .../myhub.io, so installing
or browsing that repository asked GitHub for one that does not exist.
Add src/plugin_system/repo_urls.py with one anchored normalize_repo_url(),
same_repo() for comparisons, github_owner_repo() and github_api_headers(),
and use them for the five copies of the owner/repo parsing and GitHub
headers in the store and for saved repositories. GitHub URLs are now
recognised by urlparse().hostname everywhere: _get_latest_commit_info
used a substring test, and _install_from_monorepo_api parsed any host.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(store): install a repository whose only branch is not main/master
_install_via_git returned None both when every clone failed and when the
last-resort clone of the repository's default branch succeeded.
_install_plugin_impl papered over it with `and not plugin_path.exists()`;
install_from_url did not, so a repository whose only branch is e.g.
`develop` was cloned, then treated as a failure, then "downloaded" from
main/master archives that do not exist.
After a default-branch clone, return the branch the clone checked out
(read from .git/HEAD), so None means failure and nothing else, and give
both callers the same `branch_used is None` fallback.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugins): judge the memory limit on each call's own growth
monitor_call stores `metrics.memory_mb = max(previous, growth)`, and
_check_limits compared that high-water mark with max_memory_mb. It never
decreases, so once one update() grew the process past the limit every
later call raised ResourceLimitExceeded and the circuit breaker kept
reopening. Pass the call's own RSS growth to _check_limits; keep the
high-water mark for reporting and document what it measures.
Remove ResourceMetrics.update_average_execution_time: nothing called it,
and it overwrote the running total with the average.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugins): reload_plugin re-reads the manifest from the discovered directory
reload_plugin read `plugins_dir / plugin_id / "manifest.json"`, ignoring
the discovery map and the plugin_dirs rules. For a plugin whose
directory name differs from its manifest id the path did not exist, the
re-read was skipped without a word, and the reload kept the stale
manifest. Resolve the directory with find_plugin_directory, as
load_plugin does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugins): drop the always-null last_display from plugin state info
PluginStateManager reported `last_display` from `_last_display`, which
nothing ever wrote, so it was null for every plugin. Recording it in
PluginExecutor.execute_display would not help: get_state_info's only
reader is the web process, whose PluginManager never calls display().
Remove the field, its dict and get_last_display() (no caller in core,
the web UI or the plugin monorepo).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(store): share the rollback and requirements helpers, drop dead code
- install_plugin and _reinstall_with_rollback set aside, discard and
restore the old copy through _set_aside/_discard_backup/_restore_backup
instead of two copies of the same blocks.
- The loader and the store run the same pre-pip checks through
contained_plugin_dir() and requirements_to_install() in plugin_loader.
They still invoke pip differently (sys.executable -m pip vs. the sudo
wrapper). `except (BrokenPipeError, OSError)` + `isinstance(e, OSError)`
becomes `except OSError` checking errno.EPIPE.
- load_module never returns None, so load_plugin's check is gone and the
docstring says what it raises.
- Remove the always-true JSONSCHEMA_AVAILABLE, the inline re-imports of
re and permission_utils, the fake status_result object nobody reads,
hasattr(git_error, 'cmd'), a redundant "merge conflict" test and
`import traceback` (exc_info=True does it).
- Correct comments: install_from_url names the directory for the
caller's id when given (not always the manifest id), _get_local_git_info
saves one git subprocess (not four), _enrich calls two helpers,
search_plugins documents all its arguments, _find_plugin_path states
its behaviour instead of a TODO, and history narration is gone.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): tidy base_plugin, correct plugin_manager/state comments
- base_plugin: drop the unused `import logging`; get_display_duration
runs the instance value and the config value through one
_positive_seconds() helper instead of two copies of the coercion; the
'static'/'none'/fallback branches of get_vegas_display_mode, which all
returned FIXED_SEGMENT, are one; fix the mis-indented validate_config
example; say that get_supported_vegas_modes/get_vegas_segment_width
are not consulted by core (kept, plugins override them).
- schema_manager: import expand_style_elements normally rather than
swallowing an ImportError of a core module.
- plugin_manager: the plugins directory is the configured one
(plugin-repos/ by default), not plugins/; get_config() returns the live
dict, not a copy, so the interval cache comments say what it saves.
- state_manager: config_version and the file version are not used to
detect corruption; say what they are.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): stop writing data/plugin_operations.json
PluginOperationQueue wrote its finished-operation history to
data/plugin_operations.json after every operation, and read it back only
into its own in-memory list, which only get_operation_history() exposes
-- and nothing calls that. The operation-history endpoint reads
OperationHistory (data/operation_history.json). No code in src/,
web_interface/, scripts/ or test/ reads the file.
Drop the history_file/lazy_load parameters and the load/save code; the
bounded in-memory history stays. web_interface/app.py and the
integration test stop passing the removed arguments. An existing
data/plugin_operations.json is left in place (data/* is gitignored).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): plugin-system
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs: add ARCHITECTURE and PERMISSIONS guides
ARCHITECTURE.md maps the processes, the state the display and web
services share through the cache, the display loop, the plugin system,
the web UI and the update path, with links into the code and a
where-to-start table.
PERMISSIONS.md lists who owns what after install, both sudoers files
(and why iptables is not granted), the polkit rule, and which
scripts/fix_perms script to run as which user.
Both are linked from the docs index, along with the MQTT bridge README
and src/common/README.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: correct stale setup, service and troubleshooting claims
- README: quick actions run systemctl on ledmatrix.service (run.py), not
display_controller.py; use_short_date_format has no effect; the
installer uses system pip with --break-system-packages, not a venv.
- CONFIG_DEBUGGING: LEDMATRIX_DEBUG must be "true"; logs are in journald.
- GETTING_STARTED, WEB_INTERFACE_GUIDE, TROUBLESHOOTING: enabling a
plugin, plugin settings, brightness and Vegas settings apply without a
restart; matrix hardware settings still need one.
- TROUBLESHOOTING: install dependencies with sudo so the root service
sees them; point permission problems at PERMISSIONS.md instead of a
project-wide chown.
- ADVANCED_FEATURES: real BackgroundDataService stats keys; Vegas hooks
return VegasDisplayMode and None falls back to capture; cache files
are 0660; fix_web_permissions.sh runs as the web user and does not
touch sudoers.
- STARLARK_APPS_GUIDE: only the linux-arm64 pixlet binary is downloaded.
- HOW_TO_RUN_TESTS: test class examples that exist.
- CLAUDE.md: PluginStoreManager, plugin_dirs.py, monorepo installs via
the Trees API with ZIP fallback, requirements.txt is optional.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: mark deprecated plugin APIs and state manifest fields once
Methods @deprecated("3.7.0") (the set pinned in test_deprecation.py)
were shown as current API in the quick reference, API reference,
advanced guide, development guide and FONT_MANAGER. Each is now marked
deprecated with its replacement. FONT_MANAGER is rewritten around the
current API; the override editor is gone and override methods are
deprecated.
Required manifest fields were stated three different ways. The API
reference now has one section: the 7 schema-required fields, the 4 the
store refuses without, class_name for the loader, and the 8 to set.
The other guides link to it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: document every src/common module and every widget
- src/common/README.md covered 7 of 17 modules. It now has a table of
all of them (purpose, whether plugins import it, release to floor
on), a short entry each, and logging advice that matches the code.
- SPORTS_UNIFICATION listed two shared modules and called
sports_helpers the first; it now lists all six.
- The widgets README lists all 28 registered widgets plus the support
files, and absorbs the parts that only docs/widget-guide.md had
(x-options.labels, x-advanced, x-display hidden, plugin-file-manager).
docs/widget-guide.md is now a pointer to it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(security): fix_web_permissions.sh re-hardens the root sudo helpers
The script chowns the whole project to the web user. That included
scripts/fix_perms/safe_plugin_rm.sh and safe_pip_install.sh -- the two
helpers /etc/sudoers.d/ledmatrix_web lets the web user run as root -- so
running it turned both into a root shell for whoever can edit them. It
also re-grouped config_secrets.json away from ledmatrix.
After the chown it now does what first_time_install.sh's Steps 11 and
11.1 do: helpers back to root:root 755, and config_secrets.json back to
the web unit's User=:ledmatrix 640. Each step is non-fatal and prints the
manual command if it fails.
Also fixes what the script and its docs claimed: it never configured
sudoers, its closing hint pointed at ./configure_web_sudo.sh (wrong
path), and the README and ADVANCED_FEATURES.md said to run it with sudo,
which it refuses.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(security): validate and harden every sudoers drop-in the scripts write
configure_wifi_permissions.sh copied its rules into
/etc/sudoers.d/ledmatrix_wifi without `visudo -c`. A malformed drop-in
makes sudo refuse every command for every user, which on a headless Pi
leaves no way back in. It now checks first and leaves the installed file
alone when the rules do not parse, as the other two writers do. (It
already used mktemp, so that part of the review did not apply.)
It also grants the two literal commands wifi_manager.py runs for
NetworkManager's shared-mode dnsmasq drop-in -- `cp
/tmp/ledmatrix-nm-dnsmasq.conf .../dnsmasq-shared.d/ledmatrix-captive.conf`
and `rm -f` of that file. The directory's mkdir was granted, the file was
not. Both are pinned in test_sudo_allowlist_covers_calls.py.
configure_web_sudo.sh wrote its rules to /tmp/ledmatrix_web_sudoers_$$,
a predictable name in a world-writable directory; it now uses mktemp with
an EXIT trap, as first_time_install.sh does. It sets mode 440 on the
installed file instead of leaving the temp file's mode, and finds visudo
in /usr/sbin when that is not on the user's PATH, which skipped the
check silently.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(install): escape the project path in the DNS-fix and MQTT unit renderers
install_dns_fix.sh and install_mqtt_bridge.sh substituted
__PROJECT_ROOT_DIR__ with the raw path, while the other three renderers
go through sed_escape_replacement from lib_systemd_render.sh. A checkout
under a path containing `&`, `\` or `|` rendered a corrupted unit from
these two only. Both now source the helper and use it, and a test checks
that every placeholder substitution in scripts/install uses an escaped
value.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(install): stop the installer scripts reporting things that are not true
- first_time_install.sh printed "Password: ledmatrix123" for the setup
access point. wifi_manager creates it as an open network ("No
password" on the panel), so it now says so.
- Step 10.1 printed "✓ WiFi management permissions configured" straight
after its own failure message; install_wifi_monitor.sh printed
"✓ Package installation completed" after a failed apt install. The
tick now only follows success.
- Step 7 printed "Web dependencies already installed ... in Step 5" in
the one branch that runs because Step 5 did not install them, then
created .web_deps_installed on that basis. It now warns and leaves the
marker off so the next run retries, as the comment below it intends.
- check_system_compatibility.sh called Debian 12 Bookworm "full
compatibility confirmed" while first_time_install.sh refuses anything
but Debian 13. Bookworm, older Debian and non-Debian systems are now
errors. Its counters used ((X++)), which under `set -e` exits the
script at the first warning or error (the expression is 0), so the
check never reached its summary on any system with one.
- configure_web_sudo.sh and configure_wifi_permissions.sh finished by
testing `sudo -n test -f ...` and `sudo -n nmcli device status`,
neither of which is granted, so they always reported a failure. They
now ask `sudo -n -l` about commands the new rules do grant, which
checks the rule without running anything.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(install): print the completion summary before rebooting
With -y -- and so for every one-shot `curl | bash` install, which always
passes -y -- first_time_install.sh ran `reboot` about 180 lines before
its "Installation Complete / Web UI Access" summary. reboot returns at
once, so the summary printed while the Pi was going down and the SSH
session usually dropped before the web UI address could be read.
The reboot block moves, unchanged, to the very end of the script. The
interactive prompt now also follows the summary. Because the summary now
runs before the -y reboot, its one command that could fail under
`set -Eeuo pipefail` (the SSID lookup, when nmcli reports a connected
device but no active network line) gets `|| true`; a missing SSID was
already handled as "SSID unknown".
one-shot-install.sh prints its "Next steps" after the installer returns,
by which time the reboot is under way, so it now says so, and README's
Quick Install mentions the automatic reboot.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(scripts): correct wrong comments and messages, drop dead code
No behaviour change except the output text noted below.
- 2775 is setgid, not the sticky bit (first_time_install.sh Step 3.1,
fix_plugin_permissions.sh), and root needs no "PWM hardware access"
to plugin files.
- The 777 comments in first_time_install.sh Step 3's fallback and
fix_assets_permissions.sh said root needs it to write. Root ignores
mode bits; the comments now say what 777 actually opens. The 777
itself is unchanged.
- apt_remove ends in `|| true`, so Step 12's "Some packages could not be
removed" branch could never run; it is gone and the helper stays
non-fatal.
- detect_web_service_user's comment named Step 8 for the web unit
(install_service.sh installs it in Step 7.5) and now says which
branch actually runs.
- Step 5 described an "already installed" check that does not exist;
the ACTUAL_USER comment described the re-exec backwards.
- on_error printed a literal "\n" before "Common fixes:".
- Dead code: one-shot-install.sh's uncalled fix_tmp_permissions,
LEDMATRIX_ELEVATED=1 (never read) on the sudo re-exec, and
configure_web_sudo.sh's unused PYTHON_PATH, which also made a missing
python3 fatal for rules that never mention it.
- start_display.sh / stop_display.sh said "for user: <you>"; the
service runs as root.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(fix_perms): fix_cache_permissions.sh uses setup_cache.sh's model
There were two models for /var/cache/ledmatrix. setup_cache.sh (the
installer's Step 2) and install_web_service.sh share it through the
ledmatrix group: root:ledmatrix, 2775, files 660, which is also what
DiskCache relies on to give files the directory's group.
fix_cache_permissions.sh instead made it 777 and re-grouped it to the
invoking user's group, undoing that.
It now runs setup_cache.sh for /var/cache/ledmatrix and keeps its own
handling of ~/.ledmatrix_cache. Dropped: /var/cache/ledmatrix/
placeholder_logos (nothing reads it) and the checks against the
`daemon` user (no service runs as daemon).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* ci: pin actions/checkout in the Claude workflows, drop template comments
claude.yml and claude-code-review.yml used actions/checkout@v4 while
test.yml and release-version-check.yml pin the v4.2.2 commit SHA; they
now pin the same SHA. The commented-out starter-template settings
(prompt, claude_args, paths, author filter) are removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(scripts): index every script and list removal candidates
New scripts/README.md gives one line per top-level script and scripts
directory, marked keep, dev-only or diagnostic, and lists the eight
scripts nothing in the repo refers to as candidates for removal (kept
for now). The install, utils and dev READMEs now list the files they
were missing.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test: tighten two checks that mutation testing showed were too loose
- The wifi sudoers check matched `visudo -c -f "$TEMP_SUDOERS"` in the
error report too, so replacing the check with `if false` still passed.
It now requires the command as the condition.
- The summary test never had the setup access point up, so reinstating
the bogus "Password: ledmatrix123" line went unnoticed. A case with
hostapd active now checks the AP is described as open.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(permissions): describe the repaired fix_perms scripts and new WiFi grants
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): docs-scripts
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): answer unhandled api_v3 errors from one blueprint handler
Fifty-three api_v3 routes ended in a copy of the same catch-all: log the
traceback, return {status, "An error occurred; see logs for details",
details: describe_exception(e)} with a 500. They are replaced by one
errorhandler on the api_v3 blueprint that returns exactly that body.
It lives on the blueprint rather than falling through to app.py's global
handler because the two answers differ: the global one adds
error_code: UNKNOWN_ERROR, and api_client.js sends a body with an
error_code to the error modal and one without to a plain toast. A
blueprint handler also gives tests that mount api_v3 on a bare Flask app
the same answer the real app gives.
Only handlers that were byte-for-byte that shape were removed (matched on
the AST, and each rewritten function re-parsed and compared). Handlers
with their own message, extra keys, operation-history records or cleanup
stay, as does execute_plugin_action's step-1 handler, which sits inside
an `except subprocess.TimeoutExpired` arm that would otherwise turn a
plugin's timeout into a 408.
HTTPExceptions raised inside a route go back as themselves in the global
handler's 4xx shape. Where a removed catch-all used to swallow one (only
delete_plugin_asset's non-silent get_json() is reachable), a malformed
request now gets its 415/400 instead of a 500.
Most of the diff is re-indentation from unwrapping the try blocks;
`git diff -w` shows the real change.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin action errors name the real failure, not UnboundLocalError
execute_plugin_action bound a local `logger` in its JSON-parsing arm,
which made `logger` local to the whole function. Every other
`logger.error` in it then raised UnboundLocalError, so a failing OAuth
step-1 script was reported as "UnboundLocalError: cannot access local
variable 'logger'" -- from the step-1 handler, and before the previous
commit from the route's outer catch-all too. Use the module logger.
Found by comparing every api_v3 route's forced-failure response before
and after the catch-all consolidation.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop the error category and exception-name code guessing
WebInterfaceError derived an ErrorCategory from every error code and put
it in each structured error body as `error_category`. Nothing reads it:
not the web UI (static/ and templates/), not the tests beyond the ones
pinning the mapping itself, and not any plugin in ledmatrix-plugins. The
enum, the inference table and the JSON key go.
from_exception() could also guess an error code from the exception's
class name ("Config" -> CONFIG_LOAD_FAILED, and so on). Every caller
passes a code, so the guess never ran; error_code is now required.
suggested_fixes stays: the error dialog in static/v3/js/utils/
error_handler.js lists them.
The REST reference loses error_category and says what an unanticipated
exception in an /api/v3 route answers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one call for the from_exception error responses
Nine plugin routes built a structured error by hand:
from src.web_interface.errors import WebInterfaceError
error = WebInterfaceError.from_exception(e, ErrorCode.X)
return error_response(error.error_code, error.message,
details=error.details, context=error.context,
status_code=500)
That is now exception_error_response(e, ErrorCode.X) in api_helpers, so
error_response() is the only structured-error entry point the routes
use. The three operation-history routes never passed the context, and
with_context=False keeps their bodies exactly as they were; a test
compares the helper against the hand-written pair for both forms.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): one api_v3 error-response path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): one resolver for plugin id -> directory
Five places mapped a plugin id to its directory, each with its own rules
and each re-reading manifests per lookup: PluginManager discovery and
get_plugin_directory, PluginLoader.find_plugin_directory,
PluginStoreManager._find_plugin_path / list_installed_plugins, and
state_reconciliation.disk_plugin_ids. They disagreed on backup dirs,
on whether the manifest id or the directory name is the id, on duplicate
ids and on path safety.
src/plugin_system/plugin_dirs.py now holds the rules once:
PluginDirectoryIndex scans one directory and reads each manifest once;
resolve_plugin_dir() searches directories in order. What legitimately
differs per caller is an explicit argument: search dirs (discovery and
the loader: configured dir only; the store: configured then sibling
plugins/), ledmatrix- prefix (not for the store), case folding (loader
only), manifest pass (not for get_plugin_directory, whose discovery map
already holds it).
Behaviour changes, all for layouts installs do not produce:
- a directory whose manifest declares the id beats one merely named for
it (discovery already worked this way; the loader and store now agree)
- the store searches the configured dir completely before plugins/
- backup and hidden dirs are skipped everywhere (the loader's case and
manifest scans and list_installed_plugins used to return them)
- duplicate ids resolve deterministically (exact name, then
ledmatrix-<id>, then by name) with a one-time warning; discovery no
longer lists the id twice
- disk_plugin_ids / list_installed_plugins report manifest ids, falling
back to the directory name; auto-update looks the directory up
- ids that are not one plain path segment resolve to nothing in every
caller (the loader used to truncate them, the store to join them)
The .standalone-backup- marker is one constant, BACKUP_MARKER, used by
store_manager's rename-aside names and every lookup.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): one plugin-directory resolver
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): use src.logging_config in the web process; routine requests to DEBUG
The web interface had its own logging setup (web_interface/logging_config.py)
that replaced the root handlers with a plain stdout formatter. The web
service's journal lines therefore never carried a syslog priority, so
`journalctl -p err -u ledmatrix-web` returned nothing while errors were
logged, and the line shape differed from the display's (the log viewer's
prefix stripping only matched the display format). It also ran after the
module-level managers were built, so their INFO lines at import (including
"Re-removed N uninstalled plugin(s)") were dropped.
app.py now calls src.logging_config.setup_logging() first thing, the same as
run.py: journald priorities under systemd, LEDMATRIX_DEBUG honoured,
LEDMATRIX_JSON_LOGGING still selects JSON.
Per-request logging moves to web_interface/request_logging.py. Every request
used to be logged at INFO, so the UI's polling filled the journal
("GET /api/v3/errors/summary - 200" every minute per tab). Now a successful
GET/HEAD/OPTIONS is DEBUG, a successful write is INFO, 4xx WARNING, 5xx
ERROR. Durations use perf_counter and print to 0.1ms.
The duplicate module is deleted; nothing else imported it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one thread-safe TTL cache for the web process
web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.
Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
counted and get_cached defaulted to 60s. An entry now expires after the TTL
it was stored with; a reader's ttl_seconds can only shorten that. Both
current callers pass the same value on both sides (fonts_catalog 300s,
system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
same expired key could raise KeyError (reproduced), which the endpoints
turned into a 500.
app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.
Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): web logging and TTL cache
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): only ask systemctl about known units
Codacy flagged the systemctl argv built from a variable. The unit now has
to be one of two literals, and anything else raises.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): response_time_ms reads the same clock request_logging stamps
request_logging now stamps request.start_time from perf_counter, but
success_response still subtracted it from time.time(), so metadata
reported ~1.8e12 ms. Found testing on ledpi.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(fonts): one BDF loader and one BDF rasterizer
BDF faces were loaded three ways (FontManager._load_bdf_font,
element_style._load_bdf, DisplayManager._load_fonts) and drawn by two
copies of the same per-pixel loop (DisplayManager._draw_bdf_text and the
plugin test harness's "replicated" copy), which golden images and
check_plugin/dev_server previews rely on matching the panel.
src/common/bdf_font.py now owns both:
- load_bdf_face(path, size) -> (face, realised_px): native-strike fallback
for sizes the file lacks, one bounded LRU cache keyed on path, size and
mtime. FontManager, element_style and DisplayManager delegate to it;
read_bdf_native_size moves here (the old names delegate).
- draw_bdf_text(draw, text, x, y, face, color, clip): builds each glyph as
a 1-bit mask and fills it with ImageDraw.bitmap instead of a draw.point
per pixel. A blending Draw (RGB image, "RGBA" mode) keeps the point path
so translucent colours still blend.
Pixel-identical: 220,032 renders (every bundled BDF at native and
off-strike sizes, 14 strings, 4 colours, clipped on every edge, through
each old loader x rasterizer) match origin/main byte for byte.
test/test_bdf_font.py keeps a lightweight version against a frozen copy of
the old loop. DisplayManager._draw_bdf_text goes from 1.4-23 ms to about
0.1 ms per string (the old loop re-read FreeType's buffer as a Python list
for every pixel).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(testing): harness calendar_font is sized like the panel's
VisualTestDisplayManager built its 5x7 calendar_font / bdf_5x7_font as a
bare freetype.Face. With no size set its ascender reads 0, so BDF text
drawn with it landed 6px above where DisplayManager draws it -- entirely
off the canvas at y=0 -- and get_font_height() returned 0. Golden images
and check_plugin / dev_server previews showed text the panel does not.
Load it through load_bdf_face at the panel's 7px, so it is the very face
DisplayManager uses. Across the differential run this changes only the
cases drawn with the harness's own calendar_font (968 of 220,032), which
now match the panel's output.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(fonts): one BDF face per thread
The shared face cache now hands every loader (FontManager, element_style,
DisplayManager, the harness) the same freetype.Face. FreeType does not allow
two threads to use one face at once, since load_char rewrites its glyph
slot, so key the cache by thread as well.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(sports): wrap the sports_card twins that behave identically
SportsCoreSharedMixin (switch mode, via each scoreboard's sports.py) and
sports_card (scroll/Vegas mode, via game_renderer.py) carried the same
helpers twice. test/test_sports_twins.py now calls every pair with the
same inputs -- the eight scoreboards' harness fixture games in flat,
flat+nested and nested-only shapes, plus edge cases (favourites by id and
abbreviation, NRL's colliding abbreviations, missing and non-numeric
scores, bad zones, out-of-range dates, shared font faces).
Identical pairs become thin wrappers over the sports_card function:
_card_option, _vs_text, _format_game_time, _coerce_rgb, _crisp_size (with
the class's own tables), _unshare_element_fonts (with the class's own
element map, via a new optional argument), and the colour/month/weekday/
font-grid tables (dicts copied, not aliased). _format_game_date shares the
card's formatting body but keeps its own setting, weekday zone and month
table; _schema_font_size shares the parser but keeps its per-class cache,
because a reloaded plugin gets new classes and a shared path cache would
stop it seeing an edited schema. _resolve_font_size agrees but keeps its
body so it still dispatches through the overridable hooks.
No behaviour change: old and new mixin/card agree on all 22,994
comparisons over the test corpus, and the pairs that do differ
(favourite-result colours on nested payloads and by favourites source,
the weekday's timezone, the element-name map, per-mode colours) are left
alone and pinned in TestPinnedDivergence for an owner decision.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(sports): pin that an ambiguous NRL abbreviation tints in both modes
NRL's resolver passes a shared abbreviation ("NEW") through with an error
and its _is_favorite_game matches ids only, but both favourite-colour
helpers match on abbreviation as well, so both display modes tint a
Knights or Warriors result for a user who typed "NEW". The twins agree;
neither consults the _favorite_key seam. Pinned so a fix is deliberate.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(install): generate the web sudoers rules in one place
/etc/sudoers.d/ledmatrix_web was written by two copies of the same
allow-list: a heredoc in first_time_install.sh Step 10 and a block of
echo lines in scripts/install/configure_web_sudo.sh. They drifted before
(safe_pip_install.sh was granted by one only), and a test existed just
to catch that.
Both now call web_sudoers_rules() from the new
scripts/install/lib_sudoers.sh and keep their own validate (visudo -c),
install and confirm flows.
- first_time_install.sh output is byte-for-byte unchanged, so a device
re-running the installer gets "already up to date". If the library is
missing, Step 10 keeps the installed file and carries on, the same way
it handles rules that fail visudo (an empty file would pass visudo).
- configure_web_sudo.sh now writes the installer's layout: same 18 rules,
different comments and order. It still leaves out reboot, poweroff and
journalctl when they are missing; the library does that for both.
The drift test now pins the generator's grants, checks that neither
installer writes rules of its own, and runs each installer's call line
to check the argument order. Tests that read the rule text now read the
library.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(install): detect the web service user in one function
first_time_install.sh pasted the same WEB_SERVICE_USER detection block
three times (Step 3.1's fallback, the plugin-repos setup and Step 11).
The copies were identical apart from comments; they now call
detect_web_service_user(), whose body is that block unchanged.
Behaviour is the same: the function sets the same global and always
returns 0, as the inline if-chain did. Checked on Linux against all
three original copies across 13 layouts (installed unit with and without
User=, the repo as shipped, each grep branch, template placeholders).
The comment notes that the install_web_service.sh / install_service.sh
greps no longer match anything, so until Step 8 installs the unit the
result is "root". That behaviour is left as it was.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The sports plugins cache whole season schedules: 53MB for MLB, 18MB for
NHL, 17MB for NCAA baseball. On a Pi 4, orjson.loads of the MLB file
takes ~1.8s with the GIL held, and every thread in the display service
waits -- the stall watchdog caught the render thread frozen 0.5-1.3s with
the interpreter itself blocked, right on these reads. When a season record
expired, DiskCache.get paid that whole parse only to find the timestamp
too old and throw the result away.
CacheManager.set now writes timestamp and ttl ahead of the data, and
DiskCache.get reads them from the first 256 bytes of the file, applying
the same rule as before (a per-entry ttl wins over max_age; no limit
means never stale). A record that is stale is refused without being
parsed. Files in the old layout, and records from other writers, don't
match the header and are parsed in full as before.
Also: ESPN responses in the background data service and espn_dates are
parsed with orjson when it is installed (src/common/json_body.py). The
stdlib parser behind response.json() takes 3.1s on the MLB season
against orjson's 1.8s, both with the GIL held. espn_dates imports it with
a fallback, since plugins bundle copies of that module for older cores.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs(scroll): explain the tear across the middle on fast scrolls
A 1:32-multiplexed 64-row panel lights row 31 almost a whole refresh after
row 32, so fast scrolls show a sideways offset at mid-height of about
speed x refresh period. Documents the cause, how to read the real refresh
rate (show_refresh_rate prints with a carriage return), what was measured on
a single-chain 2x128x64 Pi 4 (pwm_bits, gpio_slowdown and an uncapped
refresh barely help; gpio_slowdown 2 glitches), and the fix that does help:
fewer pixels per output via parallel chains.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(scroll): limit the 1:32 row-pair explanation to panels that scan that way
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(redaction): make URL-userinfo redaction linear, not quadratic
_REDACT_URL_USERINFO could start a match at every letter of a run of
scheme characters, and each attempt read to the end of the run looking
for `://`. On a long unbroken run of letters or digits (a hex digest, an
ID, part of a response body) that is quadratic: 1.6s for 20k characters.
The display service redacts every message, stack trace and context value
it publishes in the error snapshot, holding the aggregator lock, and
re.sub holds the GIL for the whole call, so one such exception stalled
every thread, render loop included (~0.5s measured for 20k chars of hex).
It also made test_snapshot_stays_small the slowest test in the suite by
far: 142s of a 383s run, 139s of it in this one regex.
A match may now only start where a run of scheme characters starts
(negative lookbehind). Leading digits and `+.-` are captured in group 1
so the substitution restores them, and the scheme still has to start
with a letter, so what gets redacted is unchanged: old and new output
were identical on 300k fuzzed inputs. 20k chars now take ~0.5ms, 200k
~6ms, and test_snapshot_stays_small takes 0.8s.
test/test_redaction.py pins the exact output for schemes that begin after
digits or `+.-`, and bounds 50k-character runs at 1s; against the old
pattern those timing tests fail at 3-11s each.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK
* fix(redaction): make Authorization-header redaction linear too
_REDACT_AUTH_HEADER matched the value's opening as `\s*["\']?\s*`: two
`\s*` separated only by an optional quote. With no quote, a whitespace
run could be split between them in every possible way, and when no
credential followed (end of text, or `,` `"` `<` ...) the engine tried
them all before giving up: quadratic, 8s for `authorization:` and 20k
spaces, 17s with `Proxy-Authorization:` (tried again at the inner
`authorization`). Same stall as the URL pattern: re.sub holds the GIL,
and the display service redacts everything it publishes.
The quote and the whitespace after it are now one optional unit,
`\s*(?:["\']\s*)?`, which matches the same strings with only one way to
split them. Output is identical to the old pattern on 300k fuzzed
inputs; 20k spaces now take ~1.6ms. A scan of all three redaction
patterns over prefix/run/suffix shapes finds none left that scales
superlinearly.
test/test_redaction.py pins exact output for quoted, tabbed, multi-line
and credential-less headers, and bounds header + 20k whitespace at 1s;
against the previous pattern those fail at 8-17s each.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK
---------
Co-authored-by: Claude <noreply@anthropic.com>
The Pixlet editor stop route restarts the display after a SIGKILL with
_run_systemctl_command, which starlark.py never imported (since #554). The
Starlark device-location resolver fell back to _ensure_cache_manager, which
#609 deleted; the resolver already accepts no cache manager. Both raised
NameError on the rare path that reaches them. pyflakes finds no other
undefined names in src/ or web_interface/.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* feat(web): show which plugins use each font, warn before deleting one
The Fonts tab lists font files from the web process's own scan, and the
plugins that register fonts run in the display process, so the tab had no
way to say whether a font was in use before deleting it.
The display service now publishes {catalog key: [plugin ids]} to the
shared cache (font_usage_snapshot, src/font_usage.py), built from the
loaded plugins' FontManager.register_manager_font() registrations. A
daemon thread checks every 10 s and writes only when the usage changed
(plus a daily refresh so cache cleanup cannot expire it); it never raises.
Families, aliases (press_start, four_by_six, ...) and paths are resolved
through FontManager's catalog to the file stem the Fonts tab keys rows by;
fonts outside assets/fonts are left out. Unloading a plugin drops its
registrations (new FontManager.forget_manager_fonts).
GET /api/v3/fonts/catalog merges used_by into each row per request (the
5-minute scan cache is copied, never edited): a list of plugin ids, or
null when the display service has not reported. The tab shows a Used by
column ("unknown" / "-" / ids, rendered as text) and deleting an in-use
font names the plugins in the confirmation, from a fresh read. The server
still refuses only system fonts. Catalog fetches bypass the browser's
5-second API cache, which otherwise served the pre-delete list.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix: call forget_manager_fonts through a hasattr check pylint can follow
getattr(..., None) then callable() is fine at runtime, but pylint's E1102
("not callable") can't see through it, and Codacy fails the check on it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(starlark): blank app locations use the device location, not San Francisco
A Starlark (Tidbyt) app whose Location field is blank rendered at its
author's hard-coded DEFAULT_LOCATION -- usually San Francisco -- even with
the device city set under General settings. A user in Charlotte, NC got San
Francisco weather and radar with nothing in config.json to explain it.
src/device_location.py fills unset location fields at render time (display
plugin and the web standalone render): the device city is geocoded once via
Open-Meteo, preferring a match in the configured state/country, and cached
permanently. A saved location always wins; if the lookup fails the field is
dropped so the app uses its own default, and the failure is not retried for
30 minutes.
Also fixes the config form: clearing a location omitted the key, and the
save merges, so the old value could never be removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(starlark): say what happens when the device location can't be used
A blank app Location only renders at the device's city when one is set and
the Open-Meteo lookup finds it. With no city, no match, or the geocoder
unreachable (retried after 30 minutes), the app gets no location and keeps
its author's default. The guide, the config page hint, CONFIG_REFERENCE and
the CHANGELOG entry now say so.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor: remove the skin system
Skins never rendered with the current scoreboard plugins: the only hook was
SportsCore._render_game in src/base_classes, which no plugin builds on, so
the UI and store already treated them as unsupported. The owner decided on
2026-09-23 to remove them outright.
Removed src/skin_system/ (runtime, base class, fixtures), skins/,
scripts/validate_skin.py and their tests; the store's "type": "skin"
installer, uninstaller and hide/refuse filters (the official registry lists
no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins.
Stored skin/skin_options config values are handled in the next commit.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(config): drop retired skin/skin_options keys instead of validating them
A config.json written while the skin system existed can carry skin and
skin_options in any plugin section, and most plugin schemas set
additionalProperties: false. They are no longer core plugin properties;
RETIRED_PLUGIN_KEYS in schema_manager lists them and
drop_retired_plugin_keys removes them (unless the plugin's own schema
declares the name) in prepare_plugin_config, which loading, hot reload,
GET /plugins/config and both web saves already share, and in
validate_config_against_schema for callers that validate a raw section.
POST /plugins/config and /config/main also drop them from the stored
section they merge into, so they leave config.json on the next save.
Tests cover the load path (real PluginManager.load_plugin: no schema
warning, not degraded), raw and prepared validation,
validate_all_plugin_configs, and the JSON, form and /config/main saves.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor: remove the unused src/base_classes package
No scoreboard plugin builds on src.base_classes: the nine monorepo
scoreboards ship their own sports.py and share code through src/common
(docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins
imports it. The one import anywhere, baseball-scoreboard's
rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing
instantiates.
Removed the package and the eight test files that only tested it
(test_api_extractors, test_data_sources, test_sports_base_characterization,
test_sports_capabilities, test_sports_core_promotions,
test_sports_logo_cache_bounded, test_sports_modes_promotions,
test_sports_odds_fanout). test_common_is_hardware_free no longer lists
src.base_classes as a forbidden import, and comments in sports_helpers.py
and base_odds_manager.py stop pointing at it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: drop the skin system and src/base_classes from the docs
Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to
them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins
section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the
term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was
removed and shared code lives in src/common, in the Layering section and
the view-model-contract rule. Other docs stop pointing at the removed
package. CHANGELOG records both removals under Unreleased.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(store): hide and refuse registry entries that aren't plugins
The skin filters went with the skin system, but a custom registry can still
list "type": "skin" entries, and installing one as a plugin would unpack it
into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing
type means plugin) now hides non-plugin entries from the store and
custom-registry listings, and install refuses them, in the route with a
clear 400 and in _install_plugin_impl for any other caller.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(starlark): stop the root display service locking the web UI out
Reported after a fresh install: installing an app from the Starlark tab
failed with "install failed: Failed to install from repository", and so did
uploading a .star file and installing from a GitHub directory. The reporter
found the cause only by reading service logs, and fixed it with
sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/starlark-apps
starlark-apps is gitignored, so it is never checked out -- it is created
lazily by whichever process reaches it first. Those processes run as
different users. systemd/ledmatrix.service is User=root and constructs this
plugin at startup, which is where _get_apps_directory() is called from;
systemd/ledmatrix-web.service runs as the login user and is what actually
installs apps.
The documented first step is to install pixlet and reboot, so on a fresh
machine the display service usually wins that race and mkdir() leaves the
directory root-owned. The web process then fails in _install_star_file() on
app_dir.mkdir(), which catches nothing, so PermissionError reaches the
route's outer `except Exception` and becomes the generic message the user
saw. All three install paths write to the same directory, which is why all
three failed.
The web user cannot repair this -- chown needs root. So root does it, on
every startup, which also heals machines already broken by this without the
owner having to find the chown themselves. It is a no-op when not root, when
the platform has no POSIX ownership, and when the checkout genuinely belongs
to root; a chown that fails warns rather than killing startup.
Also made the failure legible if the handover is ever prevented: a
PermissionError now names the directory, the automatic repair, and the
manual chown, instead of a message that names neither path nor cause.
Verified by mutation: dropping the handover call, chowning a genuinely
root-owned checkout, and letting a non-root process chown each fail their
own test. 121 starlark tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(starlark): address the review on the ownership repair
Findings from the automated review of #604.
Symlinks (CWE-59, the serious one). A root chown that follows links is a
privilege-escalation primitive: anyone able to write in starlark-apps could
point a link at a root-owned file and have the repair hand it over. Entries
are now read with os.lstat, symlinks are skipped outright, and the chown
passes follow_symlinks=False. Descendants are processed before the directory
itself, so the container does not change hands while its contents are still
being walked.
install_app() caught PermissionError in its broad handler and returned
False, which both routes report as a generic install failure -- the exact
shape of the bug this PR exists to fix, since the caller could not tell
"this app is broken" from "this process cannot write here". PermissionError
is now re-raised; every other failure still returns False.
The test fixtures skipped on bare Exception, which would have turned a
syntax error or NameError in the plugin into a green run. They now skip only
for a named absent dependency and re-raise anything else.
Also fixed the _Stat stub that failed in CI but passed locally: it carried
only st_uid/st_gid, and pathlib reads st_mode while walking. It now wraps
the real stat result and overrides ownership alone.
NOT taken: the CodeQL "information exposure through an exception" finding on
the hint response. Dropping `details` would contradict this package's
documented rule -- "if it returns 5xx, it says why" -- which
test_no_api_v3_handler_discards_its_exception enforces with an allowance
that may shrink and never grow. The Starlark routes are the ones that policy
was written for: they answered 500 with no detail for three releases.
describe_exception already redacts credentials and truncates. Keeping the
detail is the deliberate trade-off, so the finding is declined rather than
silently worked around.
Verified on hdpi with the updated code: a symlink to /etc/shadow planted in
starlark-apps was skipped while the directory was handed back, and
/etc/shadow stayed root:shadow.
Mutation-checked all three behaviours. The symlink test was vacuous on the
first attempt -- the link already had the target owner, so it was skipped
for the wrong reason and the mutation passed. It now forces the link to look
like it needs handing over, and fails when the check is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(errors): serve /api/v3/errors/* from the display service's aggregator
The error aggregator is a per-process singleton and only the display
service runs plugins, so only its aggregator records anything. The routes
read the web process's own, empty one and always reported no errors.
The display service now publishes a bounded snapshot of its aggregator to
the shared cache (plugin_error_snapshot) from a daemon thread: at most once
every 10 s and only when something changed, never raising into the caller.
The routes read it and keep their response shapes, adding
snapshot_available, generated_at and clear_pending; exception text has
credentials redacted.
POST /errors/clear writes a clear request (plugin_error_clear_request) that
the display applies on its next 5 s tick via the new clear_before(), which
keeps errors recorded after the cutoff and rebuilds the counts. Until the
snapshot acknowledges the request, reads hide everything before the cutoff,
so a snapshot written just before the click cannot bring errors back. Adds
"all": true; cleared_count is null when only the display can know it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(web): show plugin errors in the Logs tab
A compact panel under the log viewer: per-plugin error counts, repeating
errors (type, count, affected plugins, a sample message, last seen) and a
Clear button, with empty states for "no errors" and "display service
hasn't reported yet". Polls every 15 s while the tab is active; all text
goes through escapeHtml.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: describe where plugin error reports come from and how clear works
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(errors): redact the published snapshot before clipping it
Keeping only a traceback's tail (or clipping a message) could cut an
`api_key=` marker off while keeping the secret after it, and the web side's
redaction would then have nothing to match. The display now redacts every
free-text field of the snapshot first. The patterns move to a Flask-free
src/redaction.py so the display service can use them; redact_text in the web
error handler uses the same function, unchanged in behaviour.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): apply on-demand, brightness and schedule changes mid-screen
The main loop read the on-demand mailbox, the on/off schedule and the
brightness target once per pass -- once per screen. A dwell can be a minute
and a Vegas iteration runs for max_cycle_duration (240s), so on a Pi an
on-demand request posted at 10:54:27 was activated at 10:57:24, and two
brightness saves 12s apart inside one 30s screen never reached the panel.
During Vegas nothing read the mailbox at all: _check_vegas_interrupt only
checked on_demand_active, which only the main-loop read sets.
_service_pending_changes does the main loop's on-demand poll, expiry,
schedule and brightness steps, throttled to PENDING_CHANGES_INTERVAL (the
existing 0.25s mailbox floor), on the display thread. It runs from the Vegas
interrupt checker, the high-FPS and once-a-second render loops (replacing
their direct on-demand poll) and _sleep_with_plugin_updates; between passes
it costs one monotonic compare. A brightness change re-pushes the current
frame, since the panel only shows it from the next push.
Callers act on what it leaves behind: Vegas yields on an on-demand start or
the display being scheduled off (and the main loop then blanks instead of
rendering a screen), the render loops break on a schedule-off as they
already did on a mode change, and the dwell sleep returns early on an
on-demand start/stop or a schedule flip -- so the 60s scheduled-off sleep
now wakes for an on-demand request. The main loop no longer rotates after
a dwell that ended that way, which advanced a new on-demand session past
the mode that was asked for.
A brightness set_brightness() refuses is not retried until the target
changes, so the 4Hz pass doesn't log the same failure (fallback mode)
four times a second.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(display): a screen scheduled off midway stops rendering
Covers the schedule-off break added to the high-FPS and once-a-second
render loops: with the display scheduled off halfway through a 120s screen,
neither loop renders for more than one redraw plus one service interval
past the boundary.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(logos): harden the plugin logo download and share core HTTP headers
download_missing_logo / LogoDownloader.download_logo, the path the
scoreboard plugins use, read response.content with no size cap and wrote
straight to the final path, so a failed or corrupt download could be left
in place and cached as the logo. It now goes through fetch_logo: streamed
with a 10 MB cap, image/* only, decoded by Pillow, converted to RGBA once,
and moved into place atomically. A failure leaves no partial or temp file
and keeps any logo already on disk. LogoHelper._download_logo delegates to
the same code. Public signatures and return values are unchanged; saved
files are pixel-identical to before (RGBA, palette+tRNS, L+tRNS, LA, JPEG).
download_missing_logo reuses one downloader per thread instead of a new
Session per logo. Per thread rather than behind a lock: Session is not
documented thread-safe, and a lock would serialise every plugin's
downloads behind the slowest one.
Placeholders are written atomically, without the test_write.tmp probe.
The logo downloader and background data service now send the real
ChuckBuilds User-Agent from src.common.api_helper (USER_AGENT,
DEFAULT_HTTP_HEADERS) instead of a yourusername/contact@example.com
placeholder, and no longer hand-set Accept-Encoding: br (brotli is not
installed).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(http): drop APIHelper's hand-set brotli encoding; LogoHelper sends the real UA
APIHelper advertised `br` though brotli isn't installed, so a server that
honoured it would send a body requests can't decode. LogoHelper sent a bare
`LEDMatrix-Common/1.0`, the kind of User-Agent ESPN has been rejecting.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
project_root was only assigned in the relative-path branch, so an absolute
plugin_system.plugins_directory made web_interface/app.py raise NameError
at import (first use: the SchemaManager construction). Define it before the
if/else; plugins_dir resolution is unchanged.
Adds a regression test that imports the real module in a fresh interpreter
with an absolute and a relative plugins_directory.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(cache): collapse CacheStrategy's all-60 defaults table and twin soccer branch
get_sport_live_interval() without a config manager looked the sport up in
a table where every value was 60, with 60 as the fallback; it now returns
60. get_data_type_from_key() had an `if 'soccer'` branch returning the
same 'sports_live' as its else.
test_cache_strategy_intervals pins the returned strategy for every data
type x sport key x config-manager shape; it passes unchanged on the old
code. A 2,544-entry dump of every CacheStrategy method over a wider grid
is identical before and after.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(cache): drop CacheStrategy's `<sport>_scoreboard` config lookup
get_sport_live_interval() and get_cache_strategy() read live/recent/
upcoming intervals from config[f"{sport}_scoreboard"]. Those sections
belonged to the built-in scoreboards the plugin system replaced; plugin
config is keyed by plugin id ("football-scoreboard"), so on a current
config the lookup always fell through to the defaults (60 live, 1800
recent, 10800 upcoming), which are now returned directly.
The one input where this differs: a config.json upgraded from the
pre-plugin era that still carries e.g. an "nfl_scoreboard" section (no
code removes them), queried with an explicit sport key. No caller in core
or the plugin monorepo passes a sport key here -- get_with_auto_strategy
only derives one for keys classed sports_live/live_scores, and its callers
(odds managers, odds-ticker) use odds keys -- so the stale section was
unreachable in practice. A dump of every CacheStrategy method over 2,544
inputs differs from the previous commit only in those 45 legacy-config
entries; the test grid now includes that shape.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(cache): list cache files without holding the memory-tier lock
CacheManager.list_cache_files() held the in-memory cache's lock while it
listed and stat'd the whole cache directory -- 8,864 files on a real rig
-- so every get()/set() from the display loop and plugins waited out the
scan. The lock never protected the disk: DiskCache writes and deletes
under their own lock, and a file vanishing between listdir and stat was
already handled (logged and skipped). The body is unchanged apart from
the dedent.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(cache): delegate memory-tier cleanup and stats to MemoryCache
CacheManager._cleanup_memory_cache() was a line-for-line copy of
MemoryCache.cleanup(), and get_memory_cache_stats() a copy of
MemoryCache.get_stats(), both reaching into the component's private
_cache/_timestamps/_lock through "backward compatibility" aliases bound
in __init__. So the component's own cleanup and stats only ever ran in
tests, and the aliases went stale whenever the component was swapped
(test_cache_ttl_honoured does). Both now delegate, and the aliases are
gone: nothing in core, the tests, or the ledmatrix-plugins monorepo reads
them.
Behaviour is the same. Compared line by line, the two cleanups differ
only in the sort key's fallback (0 vs 0.0, which orders identically),
range+bounds check vs slice for the eviction, and the logger name on the
DEBUG summary line (src.cache_manager -> src.cache.memory_cache). A
differential run over 20,000 random memory states (str/None/garbage/
future timestamps, orphan keys, sizes 0-12, forced and throttled runs)
gives identical removed counts, resulting dicts and last-cleanup times;
the same harness catches each of three seeded mutations of
MemoryCache.cleanup. The throttle clock also moves with it:
CacheManager kept its own copy of last-cleanup, the component's is used
now, and they started equal.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(background): inline the sport cache key and drop the unused request queue
get_sport_cache_key() constructed a whole CacheManager -- ConfigManager,
config parse, cache-dir probing with test-file writes -- to return
f"{sport}_{date}". It now builds the key itself in the same format as
CacheManager.generate_sport_cache_key() (UTC date, %Y%m%d); tests check
the two agree for explicit dates and, with a frozen clock at 03:30 UTC,
for the default date. Median per call on Windows: ~0.6 ms -> ~2 us
(alternating runs); on a Pi the old path also wrote a probe file per call.
request_queue was a PriorityQueue nothing ever put into: requests go
straight to the executor, so `priority` never did anything. The queue is
gone; the `priority` parameter and FetchRequest field stay (every
monorepo scoreboard passes priority=) and are documented as ignored, and
get_statistics() keeps reporting queue_size, now a literal 0 as it
always was in practice.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
save_config() opened config.json with 'w' and streamed json.dump into it,
so a power cut or an unencodable value left the file truncated.
save_config_atomic() renamed a temp file into place but never fsynced it,
rewrote the unchanged secrets file on every save, and re-parsed every
backup to rotate them. save_raw_file_content() had its own third copy.
All of them, plus rollback and config creation from the template, now go
through atomic_write_text(): temp file in the same directory, fsync,
final mode set before the rename, rename (retried on Windows while a
reader holds the file), directory fsync. A root save copies the previous
owner onto the new file so a rename by the display service no longer
hands config.json to root; the shared-group fix-up is unchanged. The
mode is chosen from the file name, so a "secrets" directory in the
install path no longer makes config.json 0640.
The secrets file is rewritten only when its content changes, and backup
rotation works from filenames alone. Backups keep their names
(config/backups/config.json.backup.<version>, paired secrets backup) and
the five newest are kept, as before.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
35 methods on CacheManager, DisplayManager, FontManager and PluginManager
have no caller in core, the ledmatrix-plugins monorepo or the registry's
third-party plugins, but plugins live elsewhere, so they stay for one
release. src.deprecation.deprecated logs a warning (and emits a
DeprecationWarning) the first time each is called in a process, naming the
release that removes it. The list and replacements are in CHANGELOG and
PLUGIN_API_REFERENCE's new Deprecated APIs section; a test pins the set.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop validators nothing calls
escape_html, validate_image_url, validate_font_awesome_class,
validate_mime_type, validate_numeric_range, validate_string_length and
sanitize_plugin_config had no callers outside their own tests. Only
validate_file_upload (fonts upload) is imported by the web interface.
dedup_unique_arrays is kept: its one caller in save_plugin_config was
removed by the unrelated sync PR (#330), which looks accidental.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(api): remove the music-auth and of-the-day JSON routes
POST /plugins/authenticate/spotify and /plugins/authenticate/ytm had no
caller but their tests: the music plugin authenticates through its
web_ui_actions (authenticate_spotify.py / authenticate_ytm.py) via
/plugins/action.
POST /plugins/of-the-day/json/upload and /json/delete looked the plugin
up by the id ledmatrix-of-the-day (its manifest id is of-the-day), were
reachable only from a file_type "json" upload field that no schema
declares, and put the plugin directory on sys.path per request to
import scripts.update_config. of-the-day manages its files through
plugin-file-manager and its own web_ui_actions.
The of-the-day branch of GET /plugins/config stays: it matches the real
manifest id and still merges the on-disk category files into the form.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(api): read managers only from the blueprints
api_v3/__init__.py and pages_v3.py declared module globals
(plugin_store_manager, saved_repositories_manager, schema_manager,
operation_queue, plugin_state_manager, operation_history, sync_manager,
config_manager, plugin_manager) that nothing assigns: app.py sets the
managers as attributes on the Blueprint objects, and every route reads
them there. The one reader, backup restore's fallback to the module
plugin_store_manager, could only ever fall back to None.
_ensure_cache_manager() built a second CacheManager in the web process
instead of using the one app.py puts on api_v3. The display routes now
read api_v3.cache_manager, creating it on the blueprint only when
nothing set it (the same None handling as the /cache routes).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(web): drop run.sh and the unused log_config_change
web_interface/run.sh was referenced only by web_interface/README.md;
the service starts the UI through scripts/utils/start_web_conditionally.py
and the README already documents `python3 web_interface/start.py`.
log_config_change() in web_interface/logging_config.py was never called.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): delete unreferenced store_manager.js, diff_viewer.js, htmx-sse.js
- js/plugins/store_manager.js (window.PluginStoreManager) and
js/config/diff_viewer.js (window.ConfigDiffViewer) were loaded on every
page but nothing reads either global.
- htmx-sse.js (plus its CDN fallback) was loaded after HTMX, but no
template or plugin page uses sse-connect / hx-ext="sse": the live
streams run through LEDStreams in app-shell.js.
js/plugins/state_manager.js stays: install_manager.js's updateAll()
reads and refreshes window.PluginStateManager.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): remove app.js helpers nothing calls
- hexToRgb, rgbToHex, validateForm, uploadFont and switchTab (whose
'switch-tab' event had no listener) have no caller in the templates,
static JS or the plugin monorepo.
- installPlugin: plugins_manager.js (loaded last) assigns
window.installPlugin, and its own store cards are the only callers.
- The showNotification fallback could never install: app-shell.js is
deferred ahead of app.js and defines the same fallback at top level.
- performanceMonitor only logged with ?debug=perf and read an unset
this.measures; the marks it took on every load had no reader.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop app-shell.js refreshPlugin
A top-level function in app-shell.js, so a window global, but nothing
calls it (no inline handler, no window lookup, no string-built name).
The other plugin actions in that block stay. updatePlugin is the live
window.updatePlugin: plugins_manager.js only installs its own copy when
none exists. uninstallPlugin/pollUninstallOperation, updateAllPlugins,
executePluginAction and toggleNestedSection are replaced by later
deferred scripts, but a click that lands while those scripts are still
downloading reaches the app-shell copies, so removing them is not a
pure no-op.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): remove definitions plugins_manager.js always overrides
All of these are replaced before anything can call them, checked
against the load order in base.html and the live window.* values:
- openOnDemandModal/requestOnDemandStop stubs: the IIFE later in the
same script assigns the real functions synchronously.
- updatePlugin and uninstallPlugin stubs (`window.X || stub`): app-shell.js
already defined both, so the fallback never installed. Same for the
later updatePlugin override, gated on the live function containing
'[UPDATE]', which app-shell.js's never does.
- The first addArrayObjectItem/removeArrayObjectItem: reassigned by the
top-level copies after the IIFE.
- The first `function formatDate` in the IIFE: a later declaration of
the same name in the same scope wins.
- deleteUploadedImage, getCurrentImages, showUploadProgress,
formatFileSize and getScheduleSummary: character-for-character
copies of js/widgets/file-upload.js, which stays the owner.
- `typeof X === 'undefined'` fallbacks and `typeof X !== 'undefined'`
re-exports after the IIFE: always false, or a self-assignment.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): render the shell directly and delete index.html
index.html extended base.html with {% block content %}, but base.html
defines no blocks, so none of index.html ever rendered: rendering both
with jinja2 gives byte-identical output. index() still loaded the config,
read config.json and config_secrets.json raw and json.dumps'd them on
every page load for variables base.html never reads, and flashed errors
that base.html never shows. It now renders base.html with no context.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): stop htmx-config.js replacing console.error and console.warn
It swapped both globals for filters that dropped any error mentioning
insertBefore / "Cannot read properties of null" when "htmx" appeared in
the message or stack, and a list of Permissions-Policy warnings. That
hid real errors from every script on the page, and made every logged
error and warning report htmx-config.js as its source. The beforeSwap
target validation above it, which prevents the insertBefore errors in
the first place, stays.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(web): quiet the widget load announcements and debug logs
About 30 lines hit the console on every page load: one "... widget
registered" per widget file, one "[WidgetRegistry] Registered widget: X"
per registration, plus the registry, base widget and plugin loader
announcing themselves. The load-time announcements are removed; the
per-call ones (registry register, plugin widget loads, "Render called")
now go through the page's debugLog switch (localStorage.pluginDebug),
guarded because the widgets also load in node tests without it.
fonts.html and wifi.html debug logging goes through debugLog as well.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(api): drop the removed music-auth and of-the-day JSON routes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* chore(scripts): delete unreferenced helper scripts
None of these is referenced by an installer, systemd unit, CI workflow,
test, the web UI or src/:
- utils/cleanup_venv.sh removes venv_web_v2, which nothing creates
- utils/clear_python_cache.sh hardcodes ~/LEDMatrix and a .webassets-cache
nothing uses
- install/migrate_config.sh only copies the template, which the installer
and ConfigManager already do
- install/debug_install.sh, debug/debug_web_manual.py
- diagnose_web_ui.sh and verify_web_ui.sh overlap diagnose_web_interface.sh,
which the docs point to
- fix_internet_connectivity.sh is iptables-only (stale on nftables)
- diagnose_plugin_permissions.sh, dev/validate_python.py
- download_nba_logos.py + README_NBA_LOGOS.md: logo_downloader fetches
logos on demand
- setup_plugin_repos.py linked into the production plugin-repos/ dir; the
dev workflow is scripts/dev/dev_plugin_setup.sh, and
MULTI_ROOT_WORKSPACE_SETUP.md now uses it
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(config): drop unused plugin_system flags and a dead unit comment
- config.template.json: remove plugin_system.auto_discover,
auto_load_enabled and development_mode. Nothing reads them; the web UI
only stores them when a client sends them. ConfigManager's migration
only adds template keys, so existing configs keep theirs unchanged.
- config.template.json: re-indent vegas_scroll's live_* keys.
- systemd/ledmatrix.service: remove the comment documenting
LEDMATRIX_ON_DEMAND_PLUGIN / on_demand_env.conf; nothing reads either.
- CONFIG_REFERENCE.md: say the legacy keys are no longer in the template.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: delete docs/archive and PLUGIN_IMPLEMENTATION_SUMMARY.md
- docs/archive/: superseded guides; the repository history keeps them
and no live doc links into the directory. The one open document in it,
WEB_UI_AUDIT_2026-09.md, moves to docs/audits/ and is linked from the
docs index.
- PLUGIN_IMPLEMENTATION_SUMMARY.md invented usage statistics, called
v2.0.0 current, listed shipped auto-updates as future work and
documented a BasePlugin.get_config() that does not exist.
- docs/README.md: drop both, and stop telling contributors to archive
obsolete pages instead of deleting them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(plugin-api): fix extra_small_font size, cache metric key and scroll pacing example
- PLUGIN_API_REFERENCE: extra_small_font loads at 7, not 6 (crisp_size
snaps it, src/display_manager.py); get_cache_metrics() returns
cache_hit_rate, not hit_rate (src/cache/cache_metrics.py).
- ADVANCED_PLUGIN_DEVELOPMENT: the basic scrolling example slept in a loop
and never passed frame_hold; use ScrollHelper + scroll_config.configure()
and set_scrolling_state(True, frame_hold=...) as PLUGIN_API_REFERENCE does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(plugin-config): match the config tab, icon and web-action docs to the code
- PLUGIN_CONFIG_QUICK_START / PLUGIN_CONFIGURATION_TABS /
PLUGIN_CONFIGURATION_GUIDE: there is no "Reset to Defaults" button (the
tab has Refresh, Update, Uninstall, Save Configuration); plugin config
hot-reloads (ConfigService + on_config_change), so no restart; the
schema is found by the fixed name config_schema.json, not a manifest
config_schema field; the tab row is "Plugin Manager", not "Plugins";
forms are server-rendered from /v3/partials/plugin-config/<id>; the
duration hook is get_display_duration()/display_duration; a class_name
mismatch raises PluginError; the store requires id, name, class_name and
display_modes (not version); plugin_system.debug/log_level do not exist
(use run.py -d / LEDMATRIX_DEBUG). Drop "future" features that shipped.
- PLUGIN_CONFIG_CORE_PROPERTIES: list all of CORE_PLUGIN_PROPERTIES,
including skin, skin_options and the vegas_* tuning keys.
- PLUGIN_CUSTOM_ICONS: icon is only a Font Awesome class (fallback
fa-puzzle-piece); emoji/URL icons and getPluginIcon() never existed in
v3. Note that /api/v3/plugins/installed currently omits icon.
- PLUGIN_WEB_UI_ACTIONS (+ example JSON): success_message, error_message
and step1_message are never read.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(store): describe the monorepo registry and the store UI as they are
- PLUGIN_STORE_GUIDE: the Plugin Store is a section of the Plugin
Manager tab; URL installs are "Install from GitHub" -> "Install Single
Plugin"; bulk update exists (Check & Update All) plus opt-in weekly
auto-update; PluginStoreManager() defaults to plugins/, so the Python
examples pass plugin-repos; registry plugins are downloaded (GitHub API,
ZIP fallback), not cloned; updates compare version with latest_version.
- PLUGIN_REGISTRY_SETUP_GUIDE: replace the per-plugin-repo + tag
walkthrough with a short page on the monorepo registry (plugin_path,
latest_version, update_registry.py) that points at the monorepo's own
SUBMISSION.md. Drops the reference to the deleted
PLUGIN_IMPLEMENTATION_SUMMARY.md and setup_plugin_repos.py.
- plugin_registry_template.json: use the real entry shape.
- PLUGIN_QUICK_REFERENCE: automatic background updates exist (opt-in);
registry example and publishing steps use the monorepo, not tags.
- PLUGIN_DEVELOPMENT_GUIDE: tags/releases are not read by the store.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(readme): fix the Triple Bonnet mapping, install prerequisites and backup names
- README: the Adafruit Triple Bonnet uses `regular` (3 outputs), not
`regular-pi1` (1 output) -- src/matrix_support.py MAPPING_OUTPUTS, and
the README's own hardware_mapping section; the template default mapping
is adafruit-hat, the PWM mod switches it to adafruit-hat-pwm; manual
install only needs git up front (first_time_install.sh installs
python-dev-is-python3, cmake, ninja-build etc.; cython3/scons are not
used); the Pi Zero 2 W is a supported low-memory board, consistent with
PRODUCT.md, LOW_MEMORY_BOARDS.md and the installer's low-memory build;
fix the "First_time_install.sh" spelling, an orphan "2." list item and
the hello-world starter link (it lives in the plugins monorepo).
- CONFIG_DEBUGGING: automatic backups are
config/backups/config.json.backup.<YYYYMMDD_HHMMSS_ffffff> (five kept),
not config_YYYYMMDD_HHMMSS.json.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(dev): correct the test-running and rgbmatrix build instructions
- HOW_TO_RUN_TESTS: coverage is not collected by a plain pytest run and
pytest.ini has no threshold; the only one is --cov-fail-under=52 in the
core unit-test job of .github/workflows/test.yml, which runs the whole
test/ tree (not an allowlist). Almost no tests carry markers, so
-m integration / -m slow select nothing; drop them and -m unit as the
quick check. Replace the hardcoded /home/chuck path.
- DEVELOPMENT: the rgbmatrix package is built with pip install . from
the submodule root (scikit-build-core + CMake + Ninja), as
first_time_install.sh does; there is no make build-python /
bindings/python step, and the build deps are python-dev-is-python3,
cmake and ninja-build, not cython3/scons.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(wifi): the setup AP is open; auto-enable can be turned off without code changes
- WIFI_NETWORK_SETUP / SSH_UNAVAILABLE_AFTER_INSTALL: both AP paths in
src/wifi_manager.py create an open network and nothing reads
ap_password, so drop the "ledmatrix123" password and the ap_password
key/advice.
- SSH_UNAVAILABLE_AFTER_INSTALL: disabling automatic AP mode does not
need code changes -- auto_enable_ap_mode is a WiFi-tab toggle and
POST /api/v3/wifi/ap/auto-enable; note the monitor daemon reads
wifi_config.json at start, so restart it after changing the setting.
Use the ledpi username and a relative install path like the other docs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(reference): add auto_update, drop drifted line numbers, fix UI and service details
- CONFIG_REFERENCE: document the top-level auto_update.enabled key (read
by web_interface/auto_update.py and src/auto_update_setup.py); replace
drifted file:line references with function names; the template's
dim_schedule mode is "global".
- ADVANCED_FEATURES: core does not read a per-plugin background_service
block (the sports plugins read their own), and priority is "higher
number = higher priority" on FetchRequest but not used for ordering.
- WEB_INTERFACE_GUIDE: the General tab toggle is "Web Display Autostart"
(web interface service), brightness is 1-100, and config paths are
relative to the LEDMatrix folder, not /config.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: drop references to code removed in #608
get_installed_plugin_info, WiFiManager's saved_networks and the six
always-skipping plugin test files are deleted there. NetworkManager already
remembers joined networks; LEDMatrix no longer stores WiFi passwords.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: don't link SKIN_SYSTEM.md from the core-properties page
#615 deletes SKIN_SYSTEM.md; with this link, whichever of the two merged
second would break test_doc_links. The skin/skin_options entries go when
#615 removes the keys.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): remove the no-op PluginHealthMonitor
Its monitor loop did nothing (`if callbacks: pass`), register_health_check
had no callers and api_v3.health_monitor was never read by any route. The
live health data comes from PluginHealthTracker, which is untouched.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(store): drop the never-set uninstall tombstones
Nothing in production called mark_recently_uninstalled, so the
reconciler's was_recently_uninstalled check was always False. The
persistent uninstall registry is what actually stops resurrection; the
reconciler test now exercises that gate instead.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(common): delete unused config/display/game helpers, utils and error_handler
Nothing in core, the web UI, scripts or the plugin monorepo imports
config_helper, display_helper, game_helper, utils or error_handler; only
their own tests did. The error_handler re-exports leave src.common's
__all__; APIHelper, TextHelper, ScrollHelper, LogoHelper and the adaptive
layout exports are unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(config): drop ConfigService's unused versioning and save API
ConfigVersion, get_version/get_version_history/get_version_config,
rollback, save_config, reload, get_plugin_config and the backward-compat
load_config/get_config_path/get_secrets_path had no callers. The display
controller only uses get_config, subscribe, unsubscribe and shutdown,
plus the file watcher. Change detection now compares against the
current checksum instead of the last history entry.
The subscriber tests asserted `callback.called or True`; they now
reload the way the watcher does and assert the notification.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): drop unread plugin state history and callbacks
plugin_state.PluginStateManager kept a bounded per-plugin transition
history that only get_state_history (tests only) read; get_state_info
reports a separate lifetime count, which stays. set_error_info and
record_display had no callers, and set_state_with_error's `error`
argument only fed the history.
The web-side state_manager.PluginStateManager loses
subscribe_to_state_changes, _notify_callbacks, set_plugin_error and
get_state_version, none of which had callers; with no subscribers the
old-state copy in update_plugin_state went with them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): remove unused PluginManager methods and attribute guards
update_all_plugins was only called by a test (the display loop uses
run_scheduled_updates); get_plugin_health_metrics,
get_plugin_resource_metrics and get_plugin_state had no callers; and
plugin_modules was written but never read. plugin_directories is now
initialised in __init__, so the hasattr() guards around it go.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): remove unused executor, loader, store and package helpers
- PluginExecutor.execute_safe: no callers.
- PluginLoader._parse_semver: only its own tests; compatibility.parse_semver
is the live copy and test_compatibility.py already covers it.
- PluginStoreManager.get_installed_plugin_info: no callers.
- PluginResourceMonitor._local: never read.
- src.plugin_system.get_store_manager and __api_version__: no importers in
core, scripts or the plugin monorepo.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(wifi): stop storing Wi-Fi passwords in wifi_config.json
WiFiManager appended every joined network's SSID and password, in
plaintext, to saved_networks in config/wifi_config.json, and nothing
(web UI, backup restore, scripts) ever read them back: NetworkManager
keeps its own credentials. The writes are gone, and loading the config
now drops any saved_networks key and rewrites the file, so passwords
already on disk are scrubbed.
Also removes _check_dnsmasq_conflict (never called) and _detect_trixie,
whose result only reached one log line, along with the
NM_CONNECTIONS_PATHS constant only it used.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): remove unreachable and unused DisplayController code
- _follower_rebuild_scroll_image: never called.
- mode_duration (never read) and last_mode_change (write-only).
- The `chosen_cap <= 0` branch: chosen_cap is either the minimum of
caps already filtered to > 0 or DEFAULT_DYNAMIC_DURATION_CAP (180).
- The `max_duration < min_duration` branch directly after
`max_duration = max(min_duration, max_duration)`.
- The circuit-breaker branch's `display_result = False` and
`manager_to_display = None`: the first is overwritten a few lines
later, the second is already None there.
- The bool-to-bool conversion of execute_display's result, which is
always a bool.
- The `loaded_plugins` lookup in _update_modules: PluginManager has no
such attribute, so it always fell through to `plugins`.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(vegas): remove unused config update, boundary finder and refresh
VegasModeConfig.update had no callers outside its own tests (the
coordinator rebuilds the config with from_config on a change);
geometry.find_item_boundary and StreamManager._refresh_plugin_content
had no callers at all.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(run): drop the debug block that pretended to import the plugin system
In debug mode run.py put src/plugin_system itself on sys.path and printed
"Plugin system import successful" without importing anything. Nothing
imports plugin_system modules by bare name, so the path entry did
nothing either.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test: delete tests that test nothing
- test/plugins/test_{basketball_scoreboard,calendar,clock_simple,
odds_ticker,soccer_scoreboard,text_display}.py skip everywhere the named
plugins are not installed, including CI (LEDMATRIX_PLUGINS_DIR holds only
the fixture plugin); test_plugin_matrix.py already covers every
discovered plugin. Their PluginTestBase and the fixtures only it used
(plugins_dir, mock_display_manager, mock_cache_manager,
mock_plugin_manager, base_plugin_config in test/plugins/conftest.py) go
with them.
- test_plugin_system.py: test_discover_plugins (body was `pass`) and
test_dependency_check (a comment), plus the test_plugin_manager fixture
only the former requested.
- test_display_manager.py: test_draw_image asserted that an image it had
just assigned was not None.
- test_display_controller.py: the rotation and schedule-override tests
re-implemented the run-loop arithmetic inline and asserted on their own
result without calling the controller.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test: expect one plugin_last_update success stamp after update_all_plugins
EveryStampRecordsACompletion required at least two success-path stamps;
the second was update_all_plugins, removed as test-only. The worker and
synchronous paths share the remaining stamp in _execute_update_now, and
the check that every stamp calls _note_update_completed is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
retry() in one-shot-install.sh used `if ! "$@"; then status=$?`, where $? is
the status of the negation -- always 0. A failed command was never retried
and retry() reported success, so a failed `git clone` carried on until a
later check noticed the missing checkout. It now retries (3 attempts) and
returns the command's status. The two apt steps stay non-fatal: warning and
continuing is what they effectively did before, and making them fatal would
stop installs that work today. A clone that keeps failing stops the install,
as it already did, just sooner and with the one-shot's own error message.
Both installers granted the web user NOPASSWD root on display_controller.py,
start_display.sh and stop_display.sh. Those files are owned by the user after
Step 11's chown, so the grant let the web user rewrite them and run them as
root, and nothing ever ran them through sudo. Removed from both installers,
with a test that every project file granted as root is a root-owned
fix_perms helper.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(errors): stop affected_plugins growing without bound
Each repeat of an error pattern appended every plugin in the time window to
the pattern's list again, so a plugin failing in a loop grew the display
process's memory without limit: 3,000 errors from three plugins reached 2.5
million entries. Keep the list unique.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(fonts): load a BDF font at its native size instead of PIL's default
FreeType rejects any size but a BDF strike's own, and FontManager answered
that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf
requested at 8 or 10px rendered as PIL's default font. Retry at the native
strike, as element_style already does.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin toggle failures no longer claim "operation in progress"
Every exception in POST /plugins/toggle was mapped to
PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation
is already in progress". Report the failure as what it is, and record the
plugin id in the operation history for form posts too (it read a `data`
variable that only the JSON path set).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): route plugin card clicks through handlePluginAction
The document-level delegation checked `typeof handlePluginAction`, which is
scoped inside the plugin-manager IIFE and so never visible to it. Every card
click took a copied fallback that stopped propagation (the grid's own
listener never ran), confirmed an uninstall twice, and sent Starlark app
uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>.
Expose the handler on window and delegate to it.
Also run every test/js/unit suite under pytest: they need only node, but CI
ran one of the eight.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): apply Rotation durations, WiFi messages and Vegas settings
Three settings the web UI saves never reached the display:
- Rotation & Durations: display.display_durations was never read. Every
plugin inherits get_display_duration() and the plugin was asked first. A
saved value now wins. The page shows unsaved screens blank with the
plugin's own duration as a placeholder, and saving a blank removes the
override, so one save no longer pins every screen.
- WiFi status overlay: the controller looked for wifi_status.json one
directory above the repo. Both sides now use
wifi_manager.get_wifi_status_path(). The message is written by rename so
the display never reads it half-written, and the resumed plugin redraws the
whole panel afterwards.
- Vegas: nothing called coordinator.update_config(), so saved Vegas settings
never reached a running scroll. They are now queued when
display.vegas_scroll changes, and applied while Vegas is stopped too, so a
disable then re-enable works. The follower's scroll-speed default (75) now
matches VegasModeConfig's (50).
Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us
per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows
with each plugin.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix: keep affected_plugins order when serialized; guard non-Element targets
ErrorPattern.to_dict() ran the now-ordered list through set(), so
get_error_summary() listed plugins in an unstable order. The document-level
card-action listener called event.target.closest() without checking the
target is an Element.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): record #600 in the 3.5.0 section
#600 merged into main while the release PR was open, so the 3.5.0 section
went in without it. Nothing in that PR touched the CHANGELOG, and no check
covers "everything merged since the last tag is written down", so tagging
v3.5.0 as main stands would ship the standings-endpoint fix undocumented.
The entry goes under Sports data, next to the other ESPN fetch changes, and
is written from the commit: what the old order did, why a college league's
200 defeated the 404 fallback, and what is now treated as routine.
No version change: 3.5.0 is not tagged yet, so this belongs in that section
rather than a new one. `scripts/check_release_version.py v3.5.0` still passes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X
* docs(changelog): record #602 in the 3.5.0 section
#602 merged into main after #601, the same way #600 merged during it, and
also touched no CHANGELOG. So the section was still a commit short of what
v3.5.0 will actually ship.
It gets its own "Installers" subsection rather than a line under "Small
fixes": a malformed drop-in in /etc/sudoers.d makes sudo refuse every command
for every user, which on a headless Pi is unrecoverable over SSH. That is not
a small fix, and someone reading the release notes to decide whether to update
should see it.
Written from the commit: what both installers did, what `visudo -c` now gates,
and the fixed /tmp path that mktemp replaced.
`scripts/check_release_version.py v3.5.0` still passes, and this branch is
rebased onto 967f3a05 so the section now covers every commit since v3.4.0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X
---------
Co-authored-by: Claude <noreply@anthropic.com>
Both installers generated the ledmatrix_web rules and copied them straight
into /etc/sudoers.d without ever parsing them. Every rule is built from
`which` lookups, so an empty or surprising path produces a malformed
drop-in -- and a malformed file in /etc/sudoers.d makes sudo refuse every
command for every user. On a headless Pi that is unrecoverable over SSH.
first_time_install.sh now runs `visudo -c` on the generated file and, if it
does not parse, prints what visudo said and leaves the installed file
untouched rather than replacing it with a broken one. configure_web_sudo.sh
does the same before it offers the rules for confirmation.
first_time_install.sh also built the file at a fixed /tmp path as root;
mktemp now picks the name.
test/test_sudoers_is_validated.py renders the installer's own sudoers
heredoc and checks the result with visudo -- the check neither installer
had -- and asserts the install stays gated on it.
Claude-Session: https://claude.ai/code/session_01Dby94z9PV3zVM25fqGNXTt
Co-authored-by: Claude <noreply@anthropic.com>
* chore: prepare the 3.5.0 release
Turns the CHANGELOG's Unreleased section into `## 3.5.0` and bumps
`src.__version__`, the value plugin `ledmatrix_min_version` floors compare
against. No behaviour change; nothing outside the CHANGELOG, `src/__init__.py`
and one docs line is touched.
The staged entries are reshaped into the `### ` subsections every released
section already uses, and the "new modules a plugin may import via `src.*`"
block moves to the top as the plugin-facing summary, the same shape as 3.4.0.
Its floor, written as "the release that ships this" while it was staged, is now
3.5.0, and `docs/SPORTS_UNIFICATION.md` says 3.5.0 for `sports_helpers.py`
instead of "(unreleased)".
Four merged changes had never been written down. They are added under the
subsection each belongs to, from the commits and their measurements:
- the idle back-off clamped to the next kickoff (#599)
- concurrent ESPN date chunks (#596)
- the three web routes that consulted plugin manifests before anything had
discovered plugins, one of which wrote a plugin API key to config.json in
plain text (#594)
- the cache permission fix and its systemd unit changes (#593), which get
their own subsection
No tag and no release: `scripts/check_release_version.py v3.5.0` passes, so
tagging is a separate, deliberate step.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X
* ci: let Claude Code Review run on PRs the Claude app opens
The review action refuses a workflow whose actor is a GitHub App unless the
app is named in `allowed_bots`, which this workflow never set:
Actor is a GitHub App: claude[bot]
Actor type: Bot
Action failed with error: Workflow initiated by non-human actor: claude
(type: Bot). Add bot to allowed_bots list or use '*' to allow all bots.
It aborts about two seconds in, before the diff is read, so the check is red
on every such PR and re-running cannot help: the actor does not change. Until
now no PR here had a bot author, so nothing tripped it.
`'claude'` rather than `'*'`: the action lowercases each entry and strips a
trailing `[bot]` before comparing it to the actor
(`isAllowedBot` in `src/github/validation/actor.ts`), so this admits
`claude[bot]` and no other app. `'*'` would admit any app that can trigger a
workflow here, with a prompt it controls — the action's own docs warn about
that on public repositories, and this one is public.
The write-permission check already allowed the app; `checkHumanActor` was the
only gate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(sports): ask the endpoint the league actually publishes for standings
ESPNDataSource.fetch_standings tried /standings first regardless of league
and fell back to /rankings only on a 404. College leagues answer /standings
with a 200 that carries no poll, so the fallback never fired and the poll
came back empty every time. Nothing failed; the rank badge simply never
appeared, and anything keyed off rankings quietly did nothing.
Endpoints are now ordered by whether the league publishes a poll, a 200
that lacks the key counts as a miss so a league answering both still ends
up with whichever one carries the poll, and only a 404 is treated as
routine -- it is how a league says it has none. A connection error, a
timeout or an unparseable body is logged as an error again.
This is the implementation the football, baseball and hockey boards already
ship; core was the last copy still on the old one. Verified against live
ESPN: mens-college-basketball returns a populated rankings key where it
previously returned nothing, and nba still resolves from /standings alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(standings): stop the endpoint handler from swallowing its own bugs
Addresses both CodeRabbit findings on #600.
The handler caught `Exception`, so an AttributeError or TypeError raised
while *inspecting* the payload was indistinguishable from an endpoint that
failed. The loop would move on and, if the other endpoint had nothing
either, return {} -- silently dropping rankings for a league that has them.
That is the precise failure this function was written to fix, so the
handler was able to reintroduce it.
Only the request is guarded now. `requests.RequestException` covers the
transport failures and `ValueError` covers a body that will not parse;
payload inspection happens after the handler, where a bug surfaces instead
of being logged as a missing poll. A non-dict payload is treated as a miss
explicitly rather than by tripping over `.get`.
Tests: the fallback paths had no coverage -- the old single-endpoint code
would have passed the suite unchanged. Added order assertions for both
league kinds, a 200-without-a-poll fall-through, 404 and non-404 recovery,
a non-object payload, and a guard proving a bug is no longer swallowed.
`test_fetch_standings_returns_empty_on_error` faked a transport failure
with a bare `Exception`, which only passed because the handler caught
everything. It now raises ConnectionError, which is what actually happens.
Verified by mutation: restoring standings-first fails 5 tests, restoring
the catch-all fails the bug-not-swallowed guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(sports): stop the idle back-off sleeping through a kickoff
A league with no live games backs its poll off as empty checks mount,
capped by live_idle_max_interval. The escalation counts empty looks and
nothing else, so a league three hours before kickoff is indistinguishable
from one three months out of season. Both reach the ceiling -- and the
ceiling then *is* the blind spot.
Measured on two rigs on 2026-09-19: gaps of up to 928s between looks, ten
of them at or above 900s. Reproduced in the wild on 2026-09-20, where an
unpatched rig sat for fifteen minutes with eight NFL games in progress and
had not noticed any of them. That is the "it doesn't pick up new live
games until I restart it" report -- restarting being the one thing that
forces an immediate look.
The clamp costs no extra request: the live fetch already downloads the
whole day's scoreboard, upcoming games included, so the earliest start
still ahead of us falls out of the payload the manager already has.
Before a kickoff the wait is shortened so it cannot run past it; just
after one, the live cadence is held for _KICKOFF_GRACE_SECONDS, because a
provider that has not yet flipped the status would otherwise look like
another empty check and escalate the back-off again, right when the game
is starting.
The grace window needed a second pass. A soak caught it as dead code: the
just-passed kickoff was replaced by the next fixture on the card the
instant it passed, `now < start` went true again, and the back-off
returned to its ceiling. Observed live -- the rig polled at 13:00:45,
found nothing because ESPN had not flipped the status, then went quiet for
a quarter of an hour. A kickoff inside the grace window is now kept.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* test(sports): pin absolute tolerances and correct a wrong grace expectation
pytest.approx defaults to a relative tolerance. On a unix timestamp that is
roughly 1790 seconds, so every kickoff assertion here was effectively
vacuous -- it called a kickoff half an hour away "equal". All seven now
pin abs=1.
That hid a wrong expectation. test_an_earlier_kickoff_still_wins_during_the_grace
asserted a game ten minutes out should displace one that kicked off moments
ago. It should not, and the code does not: while the grace holds, the wait
is the live cadence (30s), which is strictly tighter than clamping to the
nearer kickoff would give (~600s). Letting the candidate win would set a
ten-minute wait at the exact moment games are starting -- the dead grace
window this branch exists to fix.
The test now pins the real behaviour plus the safety property that makes it
correct, and is renamed to say what it checks.
Reported by CodeRabbit on the PR. The finding was right that code and test
disagreed; the suggested fix was the wrong way to resolve it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The server-rendered plugin settings partial rendered straight from the
saved config, so an option added in a plugin update (geochron 1.2.0's
show_date / show_date_line, default true) drew as an unchecked box, and
the save route's missing-checkbox handling then stored it as false.
Enum dropdowns likewise showed their first option instead of the default.
- _load_plugin_config_partial runs the stored section through
prepare_plugin_config (as GET /plugins/config does) before masking
secrets, so a secret's schema default is masked too.
- render_field falls back to the field's own default, covering children
of objects that declare a default of their own (where the defaults
extraction stops).
- The legacy-boolean parity test now compares against the config the
plugin actually runs with (defaults included).
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* perf(sports): fetch ESPN date chunks concurrently
Since ESPN started rejecting `dates=YYYYMMDD-YYYYMMDD` on 2026-09-15, one
season request became a chunk per month -- and a month over the 500-event
cap becomes a request per day. A cold college-baseball season is about 130
requests, and they went out one at a time.
That is slower than the 20s budget `_update_plugins()` shares across every
plugin at startup, so scoreboards were logging `update() timed out` on
first run and being deferred to the scheduled tick with nothing on the
panel. Measured on a Pi 4 against live ESPN, March+April college baseball
(63 requests, 3101 events): 11.2s sequential, 1.6s concurrent. Over a whole
boot that moved football-scoreboard, ledmatrix-flights and birdnet-go
inside the budget -- 13 plugins deferred before, 10 after.
Chunks now go out six at a time, in two passes: months and edge days first,
then the days of any month that came back capped. Six keeps the shared
Session under requests' default pool_maxsize of 10, so no connection is
discarded. Merged events still follow `espn_date_chunks` order -- a capped
month's days are spliced back into its own slot -- so the payload does not
depend on which request won the race.
Request order is no longer significant, so the three tests that pinned it
compare the chunks as a set and keep asserting the merged event order,
which is the part callers actually see.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(sports): drop capped month payloads before fetching their days
Review of the concurrent chunk fetch found it raised the worst-case peak
memory more than the concurrency explains. The old loop discarded a month
that came back at the 500-event cap the moment it saw it; the rewrite kept
every capped month alive in `results`/`slots` until all of their day
requests had finished.
Measured on a Pi 4 fetching 20260201-20260531 college baseball (four capped
months, 5462 events), peak RSS growth over the call:
sequential (main) 83 MB
concurrent, months retained 121 MB (+43)
concurrent, one worker 108 MB -- the retention alone was +25
concurrent, months dropped 98-100 MB (+16)
docs/LOW_MEMORY_BOARDS.md puts a 1 GB Pi 3B+ at under 200 MB of headroom,
where running out makes the board unreachable until a power cycle, so the
difference matters. The remaining +16 MB is six responses parsing at once;
three workers saved about 6 MB more, within run-to-run noise, so the worker
count stays at six.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(sports): state what ESPN_CHUNK_WORKERS was measured to do, not more
The comment claimed the sequential fetch made scoreboards blow the 20s
startup update() timeout. A boot on this branch still deferred 12 plugins
and timed out baseball-scoreboard while its season fetches took 0.74s and
1.12s: the startup budget is spent on other per-plugin work. Say what was
measured -- 17.7s sequential, 2.6-3.3s concurrent -- and nothing else.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(sports): share the ESPN rejected-range memo with the background service
BackgroundDataService always sent a season range first and, on a 400,
fell back to chunks without recording the rejection, so every background
season fetch spent a doomed request and live scoreboards learned nothing
from it (or it from them). The worker now consults and sets the same
6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight
to month/day chunks, and if every chunk fails the range is asked once for
a real error without re-spending the chunks.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): keep plugin asset and action routes inside their directories
POST /plugins/assets/upload, GET /plugins/assets/list and POST
/plugins/assets/delete joined the request's plugin_id onto assets/plugins
unchecked, so '../../config' created, wrote, listed and deleted outside
it. #561 guarded only the route that serves the files. All three now go
through path_safety.resolve_under and answer 400 for anything but a
plain name, and delete only unlinks a metadata path that resolves into
that plugin's uploads directory.
PluginManager.get_plugin_directory refuses ids that are not one plain
path segment, so /plugins/action (which runs a manifest script from the
returned directory) and every other caller get the guard; the action
route also rejects such ids up front, covering its no-manager fallback.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): report a no-op plugin update as already up to date
update_plugin() returns True both for a real update and for "nothing to
do" (a ZIP-installed monorepo plugin already at the registry version, a
bundled plugin). With no git commit to compare, POST /plugins/update
called every such success "updated successfully", so Check & Update All
counted most official plugins as updated on every run.
The route now reads what changed off the plugin itself (commit, else
manifest version, else last_updated) and returns data.update_status
(updated / up_to_date / local_only). The update-all toast is summarised
by PluginInstallManager.summarizeUpdateResults from that status, falling
back to the message for older servers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(sports): scoreboard scroll speed no longer follows target_fps
sports_scroll computed the crisp speed ladder against the global
target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since
frame-locked presentation (#545) the helper steps a fixed number of whole
pixels per presented frame and the panel presents at its real refresh, so
the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a
50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s.
The ladder now uses the display manager's refresh_hz, then
display.hardware.limit_refresh_rate_hz, then the default. target_fps is
not consulted. Docstrings now say scroll_delay is ignored for pacing (no
behaviour change there) and describe the fixed-step model.
Tests: replace the tests that pinned target_fps as the ladder refresh and
described time-based stepping; assert speed independence from target_fps
(unit and end-to-end presented px/s against the real helper), that the
fixed per-frame step is applied, and that scroll_delay does not change
speed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): escape registry and upload values in plugin manager inline handlers
The store, saved-repository and custom-registry buttons built
onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone,
so a custom registry entry whose id contained ' closed the attribute and
added its own handler. One helper, jsStringAttr(), now HTML-escapes the
JSON literal for every one of those handlers, and the store View button
opens only http(s) repo links.
The live window.updateImageList (plugins_manager.js loads last, so its
copy wins over the file-upload widget's) wrote the uploaded file's
original name, path and ids into markup raw; they are escaped now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): note plugin asset, action and inline handler guards
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(update): let the root pip wrapper install web_interface/requirements.txt
Update Code, the automatic update's health check and Install Base
Requirements install web_interface/requirements.txt through
safe_pip_install.sh, which only allowed the root requirements.txt. The
first commit changing that file would fail its dependency install, and
the automatic updater rolls back any update whose dependencies did not
install -- on every device, for every newer commit.
The wrapper now lists both core requirement files. Only their folders
are resolved, so a requirements.txt symlinked out of the project is
compared by its target and refused (previously the root file's own
symlink target was what got allowed). The updater's file list is a
named constant, and a test runs the real wrapper (pip stubbed) on
every file Update Code and the rollback install.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): do not retry plugin requests that got an HTTP answer
PluginAPI.request wrapped everything that was not a structured error as
NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a
JSON error without error_code included. Check & Update All retries
NETWORK_ERROR, so those updates were re-sent five more times with
backoff, contrary to the #587 contract that an HTTP error response is
the server's answer.
NETWORK_ERROR now means only that fetch() rejected. Any HTTP response
without an error_code, or with a body that is not JSON, is API_ERROR
with the HTTP status attached. Tested against the shipped api_client.js.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scroll): restart the stats window when an idle gap is dropped by size
#582 dropped an idle gap from the frame stats two ways: the reset_scroll()
sentinel, which also restarts the 5s window timer, and a size guard for
scrollers that never call reset_scroll(), which did not. On that path the
first real frame after the gap found the boundary overdue and logged a
stats line for a one-frame window. Both paths now share one seeding helper.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(update): leave plugins alone when update_core's own rollback fails
update_core returns rollback_failed directly when a partial pull or an
update whose health check never started cannot be rolled back. run()
only held plugins back for 'verifying', so those devices still got new
plugin versions and a display restart on top of a core in an unknown
state -- the opposite of what the health-check path does, and of the
3.4.0 changelog (plugins are left alone if the rollback fails).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(api): make the REST reference match the api_v3 package
Every documented request body, query parameter and response shape was
re-checked against the handlers in web_interface/blueprints/api_v3/.
Fixes calls that failed as documented (repo_url, action_id/params,
files/image_id, font_file+font_family, ?font=, cache key,
auto_enable_ap_mode, plugin limit keys), removes the font-override
endpoints dropped in #566, corrects response shapes (plugins/config,
plugins/schema, health, metrics, operation history, github-status,
fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds
the 26 routes it omitted (backup, system auto-update/git, wifi radio,
starlark editor, MQTT bridge, status endpoints, skins).
Documents the merge semantics of partial JSON saves to /config/main and
/plugins/config and the dim-schedule POST accepting GET's days shape,
which land in the same change set. Replaces app.py line numbers and the
removed api_v3.py path with file and function names.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): remove the General-tab plugin system toggles that did nothing
plugin_system.auto_discover, auto_load_enabled and development_mode had
General-tab toggles whose help tips promised dormant plugins and verbose
logging, but nothing reads them: every enabled plugin is discovered and
loaded regardless. Remove the three toggles.
The keys stay tolerated in stored configs. The save handler now stores
a flag only when a client sends it; treating a missing key as an
unchecked box would otherwise rewrite all three to false on every
General-tab save, which still posts plugins_directory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(scroll): remove dead code left by #523/#570
- Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read
them since the numpy blend replaced the scipy path.
- Drop ScrollHelper._last_integer_position and frame_time_target, which
were written but never read.
- Keep target_fps and set_target_fps() but document them as
informational: nothing paces off them, yet ledmatrix-elections'
test_scroll_pacing.py reads helper.target_fps back and third-party
plugins may call the setter.
- Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to
throttle", set_sub_pixel_scrolling's "default: True", and
set_frame_based_scrolling's claim that it steps.
The plugins monorepo was grepped for every removed name; none is used.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI
FONT_MANAGER.md told plugins to read display_manager.font_manager, which
does not exist, so a plugin following it failed to load with
AttributeError. The shared FontManager lives on the PluginManager and
BasePlugin._get_font_manager() returns it (with a fallback for harnesses).
Also removes the Fonts-tab override workflow and element-override panels
that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls
/plugins/store/search does not exist (404) and the list endpoint reads
query, not q. The registry guide's curl examples omitted the JSON
Content-Type, so the handlers saw an empty body and answered 400.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(config): use the shared core-key list in the last three private copies
StartupValidator warned "Plugin 'auto_update' is enabled but not found" on
every display start with auto-update or a dim schedule on; the reserved
plugin-id check missed auto_update, sync, location and the rest; and
ConfigManager's (uncalled) orphan cleanup would have deleted display,
schedule and auto_update. All three now read src/core_config_keys.py, which
also gains CORE_SECRETS_KEYS for the github/youtube secrets sections.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): partial JSON saves to /config/main change only what they send
A JSON body with one field reset every checkbox in the sections it touched:
the MQTT bridge's brightness slider turned off disable_hardware_pulsing,
inverse_colors, show_refresh_rate and use_short_date_format, and a
timezone-only save turned off web-UI autostart and weekly auto-updates.
Missing-means-unchecked now applies only to form posts: form-encoded bodies
and the v3 forms, which mark themselves with a hidden __form_section input.
Also on the config routes:
- vegas_min/max_cycle_duration no longer match the generic *_duration rule,
so they stop landing in display_durations and a blank one no longer
rejects the whole Display save;
- saving from the Raw JSON editor calls start_setup_if_needed like the
General form, so enabling auto-update there finishes its setup;
- the schedule and dim-schedule POSTs accept the per-day days.<day> shape
their GETs return, as well as the flat form keys.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scripts): install plugin dependencies from the configured plugins directory
install_plugin_dependencies.sh scanned only plugins/, but the Plugin
Store installs into plugin_system.plugins_directory (default
plugin-repos), so the documented "Recommended" fix found 0 plugins on
every store install. It now reads plugins_directory from
config/config.json (relative to the project root or absolute, default
plugin-repos) and also scans plugins/ for dev symlinks, installing a
plugin reached through both only once.
With set -e alone, `pip ... | tee` took tee's exit status, so a failed
pip install was reported as success; set -o pipefail.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: replace stale API names, line numbers and the api_v3.py path
- ADVANCED_FEATURES: StreamManager methods that exist
(get_next_segment, take_next_group, refresh, advance_cycle, ...), and the
real on-demand status envelope ({status, data: {state, service}})
- app.py:199 / :144 / :607-619 line citations and
web_interface/blueprints/api_v3.py (now a package) replaced with file and
function names in ADVANCED_FEATURES, CONFIG_DEBUGGING,
PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE,
PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README
- CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use
/config/raw/main to replace the file; describe where validation runs
- TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints
usage)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scripts): verify the web interface that actually ships, on port 5000
verify_installation.sh failed every healthy install: it required the
long-removed web_interface_v2.py and looked for a listener on port 5001,
while the web interface binds 5000 (web_interface/start.py). It now
checks the files ledmatrix-web.service runs (start_web_conditionally.py,
web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the
same 5001 port in its listen check, HTTP probe and printed URLs.
Port matches are anchored so :50001 no longer counts as :5000.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(plugins): one display-size contract: display_manager.width/height
CLAUDE.md (#580) says to read display_manager.width/height because
matrix is None when hardware init fails; the development guide, the
safety-harness doc and two DisplayManager docstrings still recommended
matrix.width/height. The bundled starlark-apps plugin read matrix.width
unguarded, so its magnify recommendation and frame scaling raised in
fallback mode (e.g. after the Pi 5 hardware refusal).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(install): make install_service.sh --help print usage instead of installing
install_service.sh parsed no arguments, so `sudo ./scripts/install/
install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md)
rewrote ledmatrix.service, ledmatrix-web.service and both update-verify
units and enabled/started them. It now handles -h/--help (usage, exit 0,
no changes) and rejects any other argument with exit 2 before doing
anything. Running it with no arguments, as first_time_install.sh does,
is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(scroll): describe the fixed-step model and document frame_hold
Since #545 a crisp speed from scroll_config.configure() makes the helper
advance a fixed whole-pixel step per presented frame with no clock, and the
display manager's frame hold is part of the speed. The docs still described
the removed wall-clock model:
- scroll_config's module and configure() docstrings said speed is applied
in time-based mode and that omitting the hold "falls back to fractional
pixels"; omitting it actually runs the scroll frame_hold times too fast.
- SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both
modes, and read a 20 ms stats median as missed refreshes although that
is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step,
the hold-dependent healthy median, that target_fps plays no part, and
that a hand-added scroll_pixels_per_second loses to a schema-default pair.
- PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling)
without frame_hold; it now documents the parameter (core 3.4.0) with a
configure() + set_scrolling_state example.
- update_scroll_position/set_scroll_speed and set_scrolling_state
docstrings say the same.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(config): mark target_fps legacy; describe what Vegas scroll_delay does
- General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after
the sports_scroll fix nothing in core scrolling reads it. The field and
its API validation stay so saved configs and plugins that read
global_config['target_fps'] keep working. CONFIG_REFERENCE says the same.
- Vegas frame_based_scrolling/scroll_delay were described as frame-count
stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based
mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and
still advances by elapsed time, so the applied speed is
clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The
config comments, render_pipeline comment and CONFIG_REFERENCE rows now
say so. No behaviour change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(deps): describe how plugin dependencies are really installed
The guides said the web service runs as root, that installs pick --user
from os.geteuid(), and quoted a warning and a
PluginManager._install_plugin_dependencies() method that don't exist. The
web unit runs as the installing user; store installs go through
install_requirements_file() and sudo safe_pip_install.sh (root), with a
user-level fallback that says so, and load-time installs run in the
display service's own (root) interpreter.
Manual paths now use the configured plugins directory (plugin-repos/ by
default) instead of plugins/, which store installs no longer use, and
install_plugin_dependencies.sh is described as scanning that directory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(update): count local changes one way for the preflight and the pull
The automatic update's preflight ignored mode-only changes and anything
whose status line contained plugins/ or plugin-repos/, then promised
"Automatic updates will not stash your changes". perform_core_update
used plain git status (modes count) and ignored only 'plugins/', then
ran 'git stash push -- :!plugins', which nothing ever pops. So an edit
to a bundled plugin under plugin-repos/, or the installer's chmods on
tracked scripts, passed the preflight and was stashed away for good.
- auto_update.local_changes() is the one predicate both use:
core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/
excluded by leading folder rather than substring (a core file under
web_interface/static/v3/js/plugins/ now counts).
- Update Code's explicit stash leaves out both plugin folders; the
pull's --autostash carries their edits and mode changes across and
reapplies them.
- The automatic updater calls perform_core_update(stash_local_changes=
False), which refuses instead of stashing edits that appeared after
the preflight; update_core reports that as 'blocked'.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scripts): diagnostics follow the web autostart default and api_v3 package
#556 made a missing web_display_autostart mean "start" (only an explicit
false/off keeps the web interface down), but the diagnostics still said
otherwise: diagnose_web_ui.sh reported a missing key as "defaults to
false", diagnose_web_interface.sh said the web interface "will not start
unless this is set to true" and recommended enabling it, and
debug_web_manual.py printed False. Troubleshooting a down web UI pointed
users at a non-cause.
Both shell scripts now evaluate the setting with the launcher's own
autostart_enabled() (inline fallback if it cannot be imported) and report
on / off / not set (on) / unparseable config; debug_web_manual.py uses
the same function. They also check web_interface/blueprints/api_v3/
__init__.py: api_v3.py became a package in #553, so every healthy
checkout was reported as missing a file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(install): what install_service.sh installs; verify script port; no sudo for --help
install_service.sh installs and starts ledmatrix, ledmatrix-web and the
update-verify units, not only ledmatrix.service (systemd/README.md,
README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help'
as a harmless check; it now shows --help without sudo and warns what a
real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh
checks the web interface on port 5000.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): note update-all, plugin system settings and script fixes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display): size the preview after orientation and pixel mappers
display_geometry.physical_size claimed to give DisplayManager's answer but
only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured
after the library's pixel mappers, so a Rotate:90 / orientation 90 chain
previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for
128x64.
Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc
does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown
names leave it alone), and move the orientation composition here so
DisplayManager and the preview share it. The module docstring no longer
claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH.
Audit finding F18.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display): refuse settings the rgbmatrix library aborts on, on every board
The library answers several settings with a NULL matrix or abort() rather
than an error, so the display service crash-looped (Restart=on-failure)
instead of reaching fallback mode: rows above 64, chain_length above 255
(uint8_t binding setter, documented as "no upper limit"), a misspelled
hardware_mapping, and parallel 2-3 on a single-output mapping, reachable
from the Display form on the default adafruit-hat(-pwm) mapping. #586 only
guarded the Pi 5 subset.
- src/matrix_support.py holds the rules for every board (Options::Validate
ranges, binding integer types, mapping names and outputs from
lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the
API's numeric ranges.
- DisplayManager checks them before building options and raises
MatrixSettingsRefused, so a hand-edited config falls back with a logged,
reported reason. Emulator mode only warns.
- The config API refuses them with a 400 naming the setting; combinations
are checked against stored values but reported only when the request
sets a field involved.
- The hardware status file gains "cause" (settings/library/forced). The
fallback log and Display banner give the Pi 5 rebuild hint only for a
library failure instead of rebuild + gpio_slowdown advice for every
failure; one Pi 5 slowdown recommendation (1-3, start at 1).
- The Display form offers classic/classic-pi1 and orientation 90/270 and
renders any other stored mapping selected with a warning, so an
unrelated save no longer rewrites them; the API accepts 90/270.
Audit findings F03, F16, F19, F21.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(display): library limits, template defaults and Pi 5 slowdown
- rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs,
classic/classic-pi1 mappings and orientation 90/270 documented.
- Defaults are the config.template.json values: config migration adds
missing keys from the template, so the listed "code defaults" never
applied.
- One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode,
starting at 1.
- Troubleshooting describes the refused-settings fallback, and CHANGELOG
corrects the Unreleased "no upper limit" entry.
Audit findings F19, F20, F21.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scripts): scroll_speeds.py opens the panel with the service's options
--measure and --demo built RGBMatrixOptions from a private copy of the
display service's builder that had drifted: gpio_slowdown came from
display.hardware (default 2) instead of display.runtime (default 3), and
rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors,
pixel_mapper_config and orientation were skipped, with different defaults
(hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown
was measured -- or garbled -- in a setup the service never drives.
The option filling in DisplayManager._setup_matrix moves, unchanged, into
DisplayManager.apply_matrix_options(options, config), which _setup_matrix
calls and the script reuses (overriding only limit_refresh_rate_hz for
--measure). The script now loads the whole config rather than the hardware
block. Tests pin the script's options to the service's attribute for
attribute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scripts): scroll_speeds.py recommends keys the resolver honours
The ladder ended by telling users to set
display_options.scroll_pixels_per_second. scroll_config ranks that key
below the scroll_speed + scroll_delay pair, deliberately, and several
plugin schemas default the pair into config, so the advised key was
silently ignored (a schema-default 1/0.02 pair plus an advised 66 still
resolved to 50 px/s).
The advice is now the pair that selects the crisp speed exactly
(pixels_per_frame every frame_hold/refresh seconds), explains that the
pair outranks scroll_pixels_per_second, and gives the scoreboards'
per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the
printed pair over a schema-default pair and check it lands on the
advertised speed and hold.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula
- SPORTS_UNIFICATION.md still presented honouring global target_fps as
sports_scroll's added behaviour and its one user-visible gain; note that
it was withdrawn because it had become a speed multiplier.
- ADVANCED_FEATURES.md gave Vegas scrolling as
(scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s
by elapsed time, through a 0.1-5 px per scroll_delay clamp when
frame_based_scrolling is on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): scroll model fixes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dev): link-github links plugins from the ledmatrix-plugins monorepo
link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git,
and those per-plugin repositories no longer exist: official plugins are
directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the
monorepo once into the dev directory, finds plugins/<name>,
plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and
links it under its manifest id. With an explicit repo URL it still links a
single-repository plugin as before.
dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a
fork), plus plugins_repo and plugins_branch; github_pattern, which was
documented but never read, is dropped and warned about. Ships
dev_plugins.json.example and git-ignores dev_plugins.json, both of which
the guide promised. Reading JSON falls back to python3 when jq is missing
(get_plugin_id silently returned nothing without jq).
update/status/list find the git checkout above a monorepo plugin
directory (its .git is not in the plugin dir), and update pulls a shared
checkout once. status no longer exits 1 when nothing is broken.
Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration,
workflow, store integration, hello-world link, submission), and the
nonexistent scripts/git-hooks/pre-push-plugin-version and
scripts/bump_plugin_version.py replaced with the real rule: bump the
manifest version and run update_registry.py. scripts/dev/README.md and
CLAUDE.md updated to match.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(scripts): monorepo workspace layout; fix_perms and install READMEs
MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin;
setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into
plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the
workspace file opens LEDMatrix plus ../ledmatrix-plugins.
scripts/fix_perms/README.md listed cache directories
fix_cache_permissions.sh never touches and a 'ledmatrix' service user
that doesn't exist (also in scripts/install/README.md); adds
safe_pip_install.sh. install/README: install_service.sh installs the web
and update-verify units too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(update): keep the rollback's pip retries inside the unit time limit
The health check reinstalled the previous requirements by trying the
next bash path after any failure, including a 600 s pip timeout. Two
files, two paths: up to 40 minutes of pip alone, while systemd stops
ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the
rollback half-way and leaving the update 'verifying' until the web UI
calls it lost.
- Like permission_utils.install_requirements_file, only a sudo refusal
moves on to the next bash; a pip that ran and failed or timed out is
not repeated. The refusal wording is one list
(permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only
verifier and pinned equal by a test.
- All reinstalls in one rollback share a 600 s budget.
- WORST_CASE_SECONDS adds up every timeout on the longest path (27.5
min); a test holds it under the unit's TimeoutStartSec and that under
the web UI's VERIFY_LOST_SECONDS.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools
Plugin config was prepared differently depending on how it arrived:
- JSON POST /plugins/config built a partial body on schema defaults, so
{"enabled": true} reset every other setting of the plugin. It now merges
onto the stored section first, as the form path already did.
- Legacy-boolean normalization (#588) ran only at load: GET /plugins/config
returned the raw boolean, posting it back failed validation, and hot
reload handed plugins the raw section (a legacy dynamic_duration: true
came back as a boolean). schema_manager.prepare_plugin_config (normalize,
then defaults) is now used by PluginManager.load_plugin, both save paths,
GET, the save notifications and DisplayController's hot-reload callback.
- The JSON save's filter kept only enabled/display_duration/live_priority
and dropped a submitted skin, skin_options or vegas_* tuning key. There
is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES,
used by validation and by the save filter; PluginManager's
CORE_OWNED_CONFIG_KEYS is its vegas subset.
- Plugin sections posted to /config/main were stored verbatim, including
values /plugins/config rejects. They now go through the same preparation
(_prepare_plugin_config_for_save, extracted from save_plugin_config), and
a failing section rejects the whole save before anything is written.
- dev_server read only top-level defaults and let a schema enabled:false
win; build_full_config shallow-merged overrides, dropping sibling
defaults; the harness extracted defaults differently from the device.
loading.build_config now uses the device's extraction and preparation,
and dev_server, check_plugin, render_plugin and the harness all use it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(mqtt-bridge): brightness changes apply live and touch nothing else
The display service's hot reload applies a saved brightness within a few
seconds, and /config/main no longer resets other display settings on a
brightness-only JSON body.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): automatic update hardening
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI
It described web_interface_v2.py and index_v2.html (both gone), client-side
form generation, one POST per field with {key, value}, and 'no nested
objects'. The v3 UI renders plugin forms server-side from the schema
(pages_v3 partial + plugin_config.html macros, nested sections and
x-widgets), posts the whole form once, and save_plugin_config() merges onto
the stored section, validates, splits x-secret fields and notifies the
plugin.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(mqtt): brightness saves apply via hot reload and leave other settings alone
The bridge README said brightness is applied on the display's next
restart; the display controller's config hot reload applies it within
seconds. It also now states that the bridge's partial JSON save changes
only brightness (the /config/main merge fix in this change set).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(update): don't log pip's output from the health check's reinstall
pip can echo a private index URL with embedded credentials;
permission_utils redacts it, the stdlib-only verifier cannot, so it
logs the exit code only.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(config): mark the plugin_system toggles as unused legacy keys
auto_discover, auto_load_enabled and development_mode are read by
nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE
said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and
the REST reference listed them as live settings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): docs and developer tools group
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): legacy plugin-system toggles no longer count as a General save
auto_discover, auto_load_enabled and development_mode have left the General
form, so a post carrying only one of them is not a general-settings save and
must not treat web_display_autostart and auto_update as unchecked. The
plugin_system block itself is left as on main for the branch that reworks it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): config-save and plugin-config preparation fixes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(claude): re-check matrix_support.py rules when the library submodule is bumped
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: address Codacy findings on the core audit PR
- plugin_manager.prepare_plugin_config: when the fallback legacy-boolean
pass also fails, log a warning instead of a bare except/pass.
- api_client.js: request() refuses any endpoint that is not a plain path
under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace,
control characters) with INVALID_ENDPOINT before calling fetch(), and
plugin ids are URL-encoded wherever they are put into a URL (also in the
app-shell batch load).
- test_update_all.js: pins both against the shipped client.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): check endpoint control characters without a control-character regex
Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags
the \x00-\x1f range in checkEndpoint's regex. Test the char codes
instead; the endpoints refused are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(auto-update): make the seed script executable on disk, not only in the index
On Linux Repo.publish() commits with -a, which recorded scripts/run.sh
as 100644 upstream because the seed file was never chmod +x. The pull
then brought in the same mode the installer chmod had made locally, so
installer_chmod saw no mode change left to check. The updater was fine:
with the upstream commit at 100755 the --autostash carries the device's
chmod across. Verified under Linux (WSL, git 2.43): the old helper fails
exactly as CI did, the fixed one passes all 63 tests in the file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The web process discovers plugins lazily: plugin_manifests is empty until
some endpoint calls discover_plugins(). Three routes consulted it without
discovering, so they misbehaved for as long as nothing else had run --
which, after every ledmatrix-web restart, is until someone opens the
dashboard:
- POST /display/on-demand/start answered 404 "Plugin <id> not found"
(or "Mode <mode> not found"). Measured on a rig: 404 for over three
minutes after a web restart, until GET /plugins/installed ran. The
browser UI loads the plugin list first, so API-only callers (the Home
Assistant MQTT bridge, scripts) are the ones who hit it.
- POST /plugins/toggle answered 404 "Plugin not found".
- POST /config/main did not recognise a plugin section, so it skipped
secret separation and merged the section as-is: the plugin's API key
was written to config.json in plain text instead of config_secrets.json.
Add _discovered_plugin_manifests(), which discovers when nothing has been
yet, and rescans once when a specific plugin id (or, for on-demand by
mode, a mode) is not found, so a plugin installed since the last scan is
found too. _installed_plugin_ids() now uses it.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(cache): web UI can read what the display service caches again
ledmatrix-web.service carried CacheDirectory=ledmatrix. With User= set to
the installing user, systemd re-owns /var/cache/ledmatrix and everything
in it to that user and its primary group whenever the directory's owner
differs -- for a directory root created, on the first start. That erased
the root:ledmatrix setgid layout the installers set up, so every file the
display service (root) wrote afterwards was root:root 0660 and unreadable
by the web interface:
WARNING - Permission denied loading cache for display_current_state ...
Since #547 install_service.sh renders the web unit from the template, so
every fresh install hit this. Measured on one rig: 392 unreadable files,
and the web UI's display status, on-demand state and plugin health empty.
Existing installs only receive `git pull`, never a reinstalled unit, so
the fix for them is in the code the root display service runs:
- DiskCache.set gives each file the directory's group (when the directory
is group-writable) and 0660 on the open descriptor before the rename,
independent of setgid. This also closes a window where a fresh file was
visible as mkstemp's 0600.
- DiskCache.share_existing_files repairs files an older version left
behind, once per process from the cleanup thread. It works through
O_NOFOLLOW descriptors and skips hard links and other users' files: the
directory is writable by the web user, and root must not be steered
into changing a file outside it.
For new installs, the web unit drops CacheDirectory=/CacheDirectoryMode=,
and install_web_service.sh stops replacing an existing directory's
ledmatrix group with the user's group.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): on-demand and current-display status read the display's latest state
Found testing the cache-permission fix on a rig: once the web interface
could read display_on_demand_state at all, /display/on-demand/status kept
answering "active" for over 100 seconds while the file on disk said
"idle". Both status routes read the display service's keys through the
web process's memory tier, which serves the first copy it read for the
full max_age (120s). Read them with memory_ttl=0, as every other
cross-process reader (plugin health/metrics, the on-demand mailbox)
already does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(install): re-group the cache dir whenever the web user is outside its group
install_web_service.sh replaced an existing cache directory's group only
when it was root's. A directory in any other group the web user is not a
member of -- root:ledmatrix, for a user who is not in ledmatrix -- was left
alone, and every file root wrote there stayed unreadable to the web
interface. Replace the group whenever the installing user is not in it.
A directory whose group the user is already in (ledmatrix, or the user's
own group where CacheDirectory= left it) is still left as it is: re-grouping
a working directory strands the files already in it on the old group.
When the group does change and root-owned JSON files carrying the old group
are present, try-restart ledmatrix.service so DiskCache.share_existing_files
re-groups them through its symlink- and hard-link-safe path, rather than a
recursive chgrp.
Verified under WSL's systemd for seven directory states (user group,
ledmatrix member, ledmatrix non-member with and without root files,
root:root, missing, unnamed gid); the previous version left the non-member
case unchanged.
Addresses CodeRabbit review on #593.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
_copy_file() replaces each restored file and then carries the previous
owner across with os.chown. On Windows os.chown does not exist and
st_uid/st_gid are 0 rather than absent, so the ownership branch always
ran and raised AttributeError. That is not an OSError, so it escaped
every per-section handler in restore_backup(): a restore over any
existing config aborted at config.json and restored nothing.
Skip the ownership step where os.chown is missing, as
auto_update_setup.py already does. No change on POSIX.
test_restore_over_a_file_the_user_cannot_write simulates root-owned
files with chmod 0o444; on Windows that sets the read-only attribute,
which blocks any rename over the file, so it is skipped there. The
modes the app writes (0o644/0o640/0o600) replace fine on Windows.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(sports): recover from ESPN rejecting scoreboard date ranges
Since 2026-09-15 ESPN's site API answers `dates=YYYYMMDD-YYYYMMDD` with
400 "Failed to get events endpoint." for every sport. Single days, months
(`YYYYMM`) and season years still work. Every season and weeks-window fetch
in core failed, including the background service the scoreboards submit
their season schedules to.
src/common/espn_dates.py re-asks a rejected range as whole-month chunks
plus the leftover edge days, which tile the window exactly (a season is
8 requests, not 213). A month that comes back with exactly 500 events is
truncated (college baseball's March) and is re-asked day by day.
It also clamps `limit` to 500: above that ESPN truncates silently, e.g.
college football returns 25 of 68 games for one Saturday at limit=1000.
BackgroundDataService recovers rejected ranges on the worker thread and
advertises `handles_espn_date_ranges` so plugins can tell whether to hand
it a range. SportsCore, sports_shared, ESPNDataSource and APIHelper route
through the helper or the clamped limit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): ESPN date-range fallback and limit clamp
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(sports): stop re-sending ESPN date ranges once one is rejected
Live scoreboards refresh every 30 seconds, and each refresh sent the range
first, got the 400, then fetched the chunks: three requests where one used
to do. After a rejection, ranges now go straight to chunks for six hours,
then the range is tried again so the workaround retires itself if ESPN
reverts. A single-day 400 does not set the memo, and when every chunk fails
the range request supplies the error without the chunks being fetched a
second time. Per-fetch chunk logging drops to debug; the rejection itself
stays a warning.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(sports): clamp limit only on ESPN scoreboard submissions
The background service is generic, and limit above 500 only truncates
scoreboards. /teams needs limit=1000 (college football has 762 teams and
limit=500 returns 500), so a teams submission must keep its limit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
ensure_shared_group_ownership() - the chgrp self-heal ConfigManager runs
before reading config_secrets.json (#416) - looked up os.geteuid
unguarded. That name does not exist on Windows, and the AttributeError
is not an OSError, so it escaped the helper's best-effort handling and
every except clause in load_config(). Any Windows checkout with a
config/config_secrets.json got a ConfigError from every config load and
could not import web_interface.app.
That is what made test_update_all_plugins.py error at setup: its client
fixture imports web_interface.app. It was not state leaked between test
files - the trigger is whether the checkout has a secrets file.
Return early when os.geteuid or os.chown is missing. No change on POSIX.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(web): accept every panel size and row address type the rgbmatrix library does
The Display form capped columns at 128 and chain length at 24, and its
submit handler (fixInvalidNumberInputs) rewrote anything larger to the cap,
so wide panels and long chains silently saved as the wrong size. The config
API checked none of the hardware numbers, so values the library rejects (odd
rows, parallel 4, PWM dither bits 3) saved and the matrix then refused to
start.
- Form limits now match the pinned library: rows even 8-64, cols >= 16 and
chain_length >= 1 with no upper bound, parallel 1-3, PWM dither bits 0-2,
PWM LSB nanoseconds 50-3000.
- save_main_config rejects out-of-range rows, cols, chain_length, parallel,
brightness, scan_mode, pwm_bits, pwm_dither_bits, pwm_lsb_nanoseconds and
gpio_slowdown with a 400.
- A stored gpio_slowdown or pwm_dither_bits of 0 renders as 0 instead of the
default, so saving the tab no longer overwrites it.
- Row Address Type offers 5 (SM5368 / B707 row shift register). Verified on a
Waveshare 96x48 V2 (24S-A1) on a Pi 4 with the Adafruit Triple LED Matrix
Bonnet: rows 48, cols 96, row address type 5, BGR, GPIO slowdown 8.
- Help text and docs: FM6124-family panels use Panel Type Standard; on a Pi 5
the library supports only row address types 0 and 2.
No change to the rpi-rgb-led-matrix submodule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): drop the rows cap and document every display setting accurately
Rows: no upper limit in the form or the API. Still even and at least 8. The
current rgbmatrix library rejects more than 64 per panel, so a larger value
saves but the matrix won't start; the help tip, README, config reference and
troubleshooting section all say so, and nothing here needs changing if the
library lifts the limit.
limit_refresh_rate_hz: the form accepts 0 (the library's "no cap"), a stored
0 no longer renders and re-saves as 120, and the API rejects negatives.
pwm_dither_bits stays 0-2: the library rejects 3 and 4, so the old form's
0-4 only ever let users save a config the display couldn't start with.
Docs and help tips, checked against the pinned library and its README:
- panel_type and rp1_rio get README entries
- show_refresh_rate prints to stdout; it never drew on the panel
- dither bits raise the refresh rate; the tip said they lowered it
- scan_mode is about interlacing at low refresh, not wrong colours
- disable_hardware_pulsing: hardware pulsing needs OE on GPIO 18 and the
onboard sound driver off; software timing makes rows flash brighter
- gpio_slowdown guidance agrees between the README and the UI
- all 22 multiplexing values listed; every numeric setting states its range
- troubleshooting for a blank panel after a settings change, jumping rows
and brightness flashes
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): reject true and 5.5 for row_address_type and multiplexing
Both still went straight through int(), so a JSON true saved as 1 and 5.5
as 5. They now use the shared hardware range check like the other panel
fields. Review feedback on #586.
Also: the RP1 Backend tooltip said it is ignored on Pi 3/4 (it is ignored
on every model but the Pi 5), and the README gave the dynamic-duration
default cap as 90s; the code default is 180s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat: refuse matrix settings a Raspberry Pi 5 can't drive
On a Pi 5 the pinned rgbmatrix library drives the panel through the RP1
chip, and that path supports only row address types 0 and 2, parallel 1-3
and the regular / regular-pi1 / classic / adafruit-hat(-pwm) mappings
(Rp1PioConfigSupported in lib/rp1/rp1_pio_backend.cc). For anything else
CreateFromOptions returns NULL; the Python binding doesn't check, so the
display process crashed on its first call into the matrix and systemd
restarted it into the same crash every 10 seconds.
- src/pi5_matrix_support.py: the rule and Pi 5 detection, matching the
library's /proc/device-tree/model check
- DisplayManager raises before creating the matrix, so it is a logged init
failure (reported by /api/v3/hardware/status) and fallback mode
- the config API rejects those settings on a Pi 5 when a request sets
row_address_type, parallel or hardware_mapping
- the Display form offers only row address types 0 and 2 on a Pi 5, and
warns when a stored value can't be used
- CLAUDE.md: re-check the rule whenever the submodule is bumped
Review feedback on #586.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
v3.4.0 shows "Plugin Config Warning - In config but not installed:
auto_update. Reinstall via the Plugin Store, or remove these entries from
config.json." auto_update is the core weekly-update setting from #581.
Reconciliation treated every top-level dict not in its private
_SYSTEM_CONFIG_KEYS list as a plugin id, and #581 could not know to extend
that list.
- Move core top-level keys into src/core_config_keys.py (CORE_CONFIG_KEYS)
and use it in reconciliation. Tests fail if a config.template.json key or
a key written by the general-settings save is missing from it.
- A secrets-file key only counts as a non-plugin when no installed plugin
has that id. Plugin secrets are namespaced by id, so installed plugins
with secrets were reported as missing from config on every run.
- still_unresolved() drops "not on disk" findings whose id is no longer a
plugin entry in config, so a stored verdict clears without a restart.
- A plugin whose id is a core key is skipped with a warning, and the fix
never writes a plugin stub over or in place of a core setting.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The news plugin's schema turned global.dynamic_duration from a boolean
into an {enabled, min_duration_seconds, ...} object. Installs that have
not saved the news settings since still hold `true`, so every start
logged "Plugin news config does not match its schema (loading anyway):
Field 'global.dynamic_duration': Expected type object, got bool" and
flagged news degraded.
The settings form already reads such a boolean as {"enabled": <bool>}
(render_nested_section in plugin_config.html) and the next save writes
the object. The loader did not. It now applies the same rule before
merging schema defaults and validating, so the defaults fill in the rest
of the object and the plugin receives it in the new shape.
The rule lives in schema_manager.legacy_bool_as_object /
normalize_legacy_booleans. It applies at any depth of nested objects
but not inside arrays, matching the form, and only to a real bool under
an object-typed property with an `enabled` child. Every other mismatch
still warns. A parity test renders the template macro against the helper
so the two cannot drift.
Nothing is written to config.json at load: the normalization is in
memory, and the next save of the plugin's settings persists the object.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(web): update-all skips Starlark apps and no longer misses plugins
Check & Update All posted every entry from /plugins/installed to
POST /plugins/update, including the virtual starlark:<app_id> entries
that list installed Starlark apps. The store manager cannot find those,
so each answered 500 "plugin not found". Update-all now sends only
plugin ids (install_manager.js, and the older app-shell.js copy), and the
route answers a starlark: id with a 400 saying it is a Starlark app.
A request that got no HTTP answer was recorded as failed and never sent
again. On a device, a web-service restart mid-run killed the in-flight
request and refused the next one, stock-news, which was left on 2.6.2
with 2.8.0 available. Such requests are now re-sent with backoff
(about 30s) before being reported as failed. HTTP error answers are not
retried.
Tests: test/js/unit/test_update_all.js (run from pytest via
test/web_interface/test_update_all_plugins.py so CI covers it) and the
route contract for starlark: ids.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(web): walk update-all retry delays without indexed lookup
Codacy's ESLint security/detect-object-injection rule flagged
retryDelays[attempt] as a High issue. The index was a bounded loop
counter over a fixed array, but shifting a per-plugin copy of the
schedule gives the same backoff without the pattern. No behaviour
change: test/js/unit/test_update_all.js (21) and
test/web_interface/test_update_all_plugins.py (7) pass unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(web): honour x-display: hidden in plugin settings
Plugins keep deprecated and internal keys declared so stored configs keep
validating (weather api_key/radar_zoom, countdown's auto-generated row id),
but the settings form drew them as live controls.
A property marked "x-display": "hidden" -- or an object whose children are
all hidden -- now gets no control at any depth: top level, nested sections,
Advanced Settings (not counted either), array-table columns and the row
editor. A hidden top-level key is not reported in __rendered_section.
Saving never changes a hidden value. Plain and nested fields aren't posted,
so the save's deep merge keeps them; _set_missing_booleans_to_false skips
hidden booleans at every depth. A posted array row replaces the stored item,
so hidden row properties are carried as JSON-encoded hidden inputs and
decoded exactly on save (an id "1" stays a string). New rows get none.
JSON API saves are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(web): read hidden row keys without dynamic property access
Build the set of x-display: hidden item properties once and look values up
through Object.entries, instead of indexing objects by a variable key on
the lines this branch added (Codacy: object injection sink, 6 warnings).
Behaviour is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(common): sports_helpers, the helpers all nine scoreboards copy verbatim
Add src/common/sports_helpers.py: the helpers the scoreboard plugins'
sports.py carry byte-identical copies of (docstring-stripped AST, checked at
ledmatrix-plugins f09bff2), so a later plugins PR can delete its copies once
it floors on the core release that ships this.
- Free functions: clamp_window, clamp_seconds, logo_needs_refresh (lazy
src.logo_downloader import, as in the plugins), spread_weighted_order,
MIN_WINDOW_DAYS / MAX_WINDOW_DAYS. All nine plugins.
- SportsHelpersMixin (no __init__, stateless): _mode_customization,
_setting_int, _reset_dwell_on_reentry, _next_switch_index,
_spread_weighted_order (all nine), _odds_color and
_upcoming_date_and_time_text (all but ufc), plus the _favorite_key seam
from base_classes core.py for later phases.
A new module rather than more methods on sports_shared: a plugin that
deletes a copy and relies on an existing module having grown the method
fails at runtime with AttributeError on an older core, which neither the
loader nor check_min_core_version.py can see; a missing module fails at load.
Tests: behaviour for every helper, a derived host contract, and a parity
test that AST-compares every body against every plugin copy when
LEDMATRIX_PLUGINS points at a checkout (skipped otherwise).
test_common_is_hardware_free.py imports src.common and every sports_* module
with rgbmatrix blocked and scans src/common for module-level imports of
src.base_classes, src.display_manager and src.plugin_system (no existing
violations).
Nothing in core imports the new module; no behaviour change. CHANGELOG
Unreleased entry and a converging note in docs/SPORTS_UNIFICATION.md.
__version__ is not bumped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(common): address review on sports_helpers and the hardware-free test
- SportsHelpersMixin docstring and CHANGELOG: constructor-free, but it keeps
lazy state on its host (_reset_dwell_on_reentry, _next_switch_index).
- test_common_is_hardware_free: the runtime check now filters every
FORBIDDEN package, src.plugin_system included; the AST scan resolves
relative imports against src.common, so `from .. import plugin_system`
and `from ..plugin_system import x` are caught. Guard tests for both.
- Parity skip reason names the CI guard that runs the same comparison:
ledmatrix-plugins scripts/check_sports_helpers_parity.py (#495).
- _odds_color: line-level pylint disable for a not-callable false positive
(getter is None-checked); the AST is unchanged, parity still passes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): complete 3.4.0 with weekly auto-updates and scroll timing fix
#581 (weekly automatic updates) and #582 (scroll frame-stats idle gap)
merged after the 3.4.0 section was written in #580. v3.4.0 will be tagged
on main including both, so they belong in 3.4.0.
Also record src.common.font_layout (#539, #565), a src.* module plugins
may import that shipped in 3.4.0 but was never listed, and mark
display_geometry and auto_update_setup as core-internal.
Correct the 3.3.0 historical note: remote tags v3.3.0 (bc2dbf38) and
v3.3.1 (32d637a4) both report "3.3.0" and both ship sports_shared.py.
The "3.2.0" claim came from a stale local tag.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): cover the whole 3.4.0 release since v3.3.1
The 3.4.0 section listed plugin-facing API and per-element customization
but not the rest of what merged since v3.3.1. Group it under subheadings:
Install and updates, Scrolling, Plugins, Web interface, Tools and
security, Fixes, and put the existing customization block under its own
heading.
Omitted on purpose: #569 (fixes a regression and an editor race in the
unreleased per-element framework), #570 (no runtime change), and
test-only, refactor and dev-tooling PRs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(web): weekly automatic updates with health check and rollback
A General-tab toggle (off by default) checks for and installs LEDMatrix and
plugin updates once a week, overnight in the configured timezone.
- Pre-update checks skip (and report) instead of forcing: local edits or
commits, merge/live rebase, no upstream, low disk, missing health check, or
a version that was already rolled back. An abandoned rebase (HEAD back on a
branch) is cleared, since it would otherwise block every pull.
- The pull reuses the Update Code path (now perform_core_update(), which
reports dependency install failures as data).
- ledmatrix-update-verify.service, started via a .path unit from a request
file, restarts the services from its own cgroup, requires them to come up
and stay up, and otherwise resets to the previous commit and reinstalls the
previous requirements. It runs a copy of the checker taken before the pull.
- No SSH needed: switching the toggle on restarts the display service, which
(as root) installs the two units from the repo templates for the web user.
first_time_install.sh installs them too and takes --enable-auto-update /
LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh).
- Plugins update after the code passes its check; failures, blocks and
rollbacks raise an Overview banner and show under the toggle.
Tested end to end on a Pi: web-UI setup, a good update, a broken web service
and a broken display (both rolled back), a blocked local edit, and an
abandoned rebase found on the device.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(auto-update): address static-analysis findings
- Replace the subprocess.CompletedProcess the verifier fabricated for a
command that could not start with a plain namedtuple; nothing is executed
there, but the scanner flags any CompletedProcess built from variables.
- Mark the subprocess imports with the repo's standard B404 annotation (all
calls are list-form argv, no shell).
- Mark the rollback-failed message as not SQL (B608 matched its wording).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auto-update): CI failures on Linux
- Keep the setup result when chown fails. CI runs as a non-root user, where
chown to the web user raises; that discarded the result file, so the
General tab would never learn whether setup worked. Regression test added.
- Register the two new /api/v3/system/auto-update routes in the URL map
snapshot.
- Use utility classes app.css defines (space-y-1, hover:text-red-600).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auto-update): address review feedback
- Health check: a failed restart command no longer lets the check run
against the still-running old process; it counts as a failure (and after a
rollback, as a failed rollback). An unreadable restart count is never
treated as stable, since a crash loop looks healthy between attempts.
- Installer writes the auto_update setting to a temp file and swaps it in,
keeping mode and owner, so a running config watcher never reads a
truncated config.json.
- Verify unit quotes its command-line paths (install folders with spaces);
setup refuses folder names systemd would reinterpret (%, quotes,
backslashes, control characters) and says so on the General tab.
- The auto-update status route no longer returns exception text.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auto-update): keep error detail in the status route's 500
test_web_error_detail requires every 5xx handler to log the traceback and
return describe_exception(e), which redacts credentials, so failures are
diagnosable from the web UI. Dropping it for CodeQL broke that policy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auto-update): dismiss route rejects non-object JSON with 400
A JSON array or scalar body made `.get('alert_id')` raise, returning 500.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auto-update): let the app-wide handler answer status-route errors
CodeQL (py/stack-trace-exposure, #709) flagged the route's own except,
which returned describe_exception(e). web_interface/app.py's error handler
already logs the traceback and returns the same redacted detail for any
unhandled exception, so the local copy is removed: same response, no new
exception-to-response flow, and test_web_error_detail's policy still holds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
ScrollHelper.last_frame_time was set once in __init__ and thereafter only
at the end of log_frame_rate(). Nothing re-armed it when a scroll began, so
the first frame of every scroll was timed against the last frame of the
*previous* one and the whole idle period between them was recorded as a
single frame.
Measured over 3 hours on a 256x64 Pi 4, that produced 31 windows reading
Scroll frame stats - 0.0 fps over 1 frames | median 136776.02ms
p95 136776.02ms max 136776.02ms min 136776.02ms | stalls 0 (0.0%)
and -- worse, because it is not obviously wrong -- put the same gap in the
max field of otherwise healthy windows, where the worst values were 537s
and 604s. It also counted as one stall per scroll start: at ~500 frames to
a window that is ~0.2%, against measured stall rates of 0.07-0.16%. The
stall rate is the number used to judge whether a scroll change worked, and
it was the same order of magnitude as its own artefact.
The first frame of a scroll has no predecessor, so it has no frame time.
last_frame_time is now None until one is rendered, and reset_scroll() puts
it back -- the same treatment last_update_time already gets three lines
above, for the same reason. reset_scroll() alone is not enough, because the
scrollers actually emitting these lines never call it, so a sample at or
past the 5s log interval is dropped as well: nothing that renders a scroll
takes that long over one frame. Seeding also restarts the window timer, or
the boundary is already overdue when the second frame arrives and every
scroll opens by reporting a window of exactly one frame. A window whose
samples were all dropped now logs nothing rather than reporting the gap.
docs/SCROLL_PERFORMANCE.md documented the diagnostic in terms of a
"Frame time: N ms" line that 6031e705 replaced with the aggregate, so its
grep matched nothing on any rig. The section now describes the line that is
actually emitted, reads duplicate frames off skips and a below-median
result rather than a 2ms mode, and adds a command that ranks every scroller
by p95 -- verified against 3 hours of journal, where it reproduces
src.base_odds_manager at p95 44.08ms against 10.19ms for the two scrollers
already on src/common/scroll_config.py.
requirements.txt still offered scipy for the sub-pixel interpolation path
deleted in #570. Installing it has no effect; the entry says so.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0
Skins: no current scoreboard plugin builds on src.base_classes, so the only
skin hook (SportsCore._render_game) never runs. The plugin schema endpoint no
longer injects the Visual Skin dropdown, the store hides and refuses
"type": "skin" registry entries, and GET /api/v3/skins reports
supported: false with a message. Stored skin config still loads and saves.
src/skin_system/ and its tests are unchanged apart from the support flag.
Docs: check_plugin.py/render_plugin.py examples use --plugin; document
BasePlugin.get_update_interval() and its interaction with the manifest
update_interval; CLAUDE.md drops the stale template line number and
recommends display_manager.width/height.
Preview size: new src/display_geometry.py holds the size computation and
defaults DisplayManager uses (double-sided applied, chain_length default 2).
The web preview, /display/current, Starlark magnify default, sync handshake
and two dev scripts use it.
Release: __version__ 3.4.0, CHANGELOG 3.4.0 section plus a 3.3.0 tag note.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: address CodeRabbit review on #580
- Preview fallbacks (SSE stream and /display/current) use logical_size({})
(128x32, the shared default) instead of a hard-coded 128x64.
- display_geometry treats a non-mapping display/hardware block as missing,
so a malformed config.json falls back to defaults instead of raising
AttributeError (which turned the Starlark render into an HTTP 500).
- Docs: the static update interval falls back manifest -> plugin config
-> 60s, in both the API reference and the architecture spec.
Not taken: validating double_sided copies against chain_length/parallel.
An orientation Rotate: or U-mapper pixel mapper decides which axis panels
lie on, so the counts would reject working setups (the existing
vertical-split test is one).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display_geometry): a non-finite hardware size raises ValueError, not OverflowError
CodeRabbit flagged the Starlark magnify default in
_standalone_render_starlark_app for truthy non-mapping display values. That
case was already handled by a9e1bd0b (_display/_hardware treat a non-mapping
block as missing, covered by test_non_mapping_display_config_uses_the_defaults),
and the magnify it produces from the 128x32 defaults is the same as from 64x32.
Checking the same path found one input that still escaped: Python's JSON
parser accepts Infinity, and int(inf) raises OverflowError, which neither the
Starlark path (TypeError, ValueError) nor the preview stream in app.py caught,
so a hand-edited "rows": Infinity returned HTTP 500. physical_size now raises
ValueError for it, matching its documented contract, so every caller's
existing fallback applies. DisplayManager already caught Exception.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
first_time_install.sh granted the web user safe_plugin_rm.sh but not
safe_pip_install.sh, unlike scripts/install/configure_web_sudo.sh. On devices
set up only by the first-time installer, install_requirements_file could not
use the root wrapper and fell back to a user-level install that root-run
ledmatrix.service may not see.
Also harden both sudo-granted helpers to root:root 755. first_time_install.sh
never did this, and Step 11's project-wide chown to the user would undo it if
placed in Step 10, so it runs at the end of Step 11.1.
Add a test that parses the ledmatrix_web sudoers rules from both installers
and asserts they grant the same commands, and that every granted helper is
hardened (after the chown, in first_time_install.sh).
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The pinned rpi-rgb-led-matrix commit emits the ARMv7-only `dmb ishst`
instruction in lib/rp1/rp1_rio_backend.cc, guarded only by __arm__, so the
build fails on every ARMv6 board ("selected processor does not support
`dmb ishst' in ARM mode"). Bump the pin to upstream 1ee4f76, which merges
12d839f (guard on __ARM_ARCH >= 7) plus docs only.
The installer also needed two changes for that bump to reach anyone:
- git pull never moves an existing submodule checkout, so a device that
already failed would keep building the broken commit. The build step now
moves the checkout forward to the pin — never backward or sideways (a
`git submodule update --remote` checkout is left alone), and never fatal.
- The submodule git commands ran as root on the user's clone (git's SUDO_UID
exemption allows it), leaving .git/modules/rpi-rgb-led-matrix-master
root-owned and the user unable to run git in it. They now run as the
project directory's owner, and root-owned leftovers are handed back.
Root-owned installs keep running as root.
test/test_install_rgb_checkout.py covers the non-root sync scenarios under
the installer's strict mode, checks that every called _helper is defined
before use, and pins the one-shot-install.sh -> first_time_install.sh
contract. Verified the tests fail on five deliberate mutations. Root/owner
scenarios were exercised manually under WSL Ubuntu, and the library was
cross-compiled for arm1176jzf-s at both pins (old: rp1_rio_backend.cc fails
at line 120; new: 16/16 sources compile).
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Two stored shapes broke the config form:
* A scalar under a field that is now an object. News' dynamic_duration
was a boolean and is becoming an object; render_nested_section did
`key in true` and the whole page failed to render. Look into dicts
only, and carry a legacy boolean over as the object's `enabled`, so
the next save upgrades it without switching the feature off.
* A custom feed logo with a path but no id. The template always emitted
an empty `logo.id` input, which the save route parsed to null, failing
the id's string type on every save. Emit it only when there is an id,
as custom-feeds.js already does.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#569 fixed _normalize_color and #572 covered the resolver path. That test's own
docstring notes the resolver "normalizes colour separately from element_color",
and the other path had no test: the stateless element_color(), which
src.common.sports_card delegates to and which every one of the nine scoreboard
plugins takes for each per-element colour it draws.
That is the path that regressed. element_color() moved here with the per-element
customization framework, the coercion rejected out-of-range components where the
reader it replaced clamped them, and a rejection reads as "not configured" -- so
one component over 255 painted the element white while the user's colour sat in
their config. Every scoreboard's test_element_text_colors.py failed on it, and
it took two plugin PRs red on CI to surface.
Six cases: clamping, in-range untouched, hex, unparseable fallback, missing
element, and agreement with sports_card.coerce_rgb. The last is the point --
the two shared readers disagreed about the same value, so this asserts against
coerce_rgb directly rather than restating the arithmetic, and any future move
of element_color has to keep them consistent.
Verified by mutation: restoring the rejecting coercion fails two of the six,
alongside the resolver test from #572.
Tests only; no source change.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(web): let the plugin config form use the full page height
The form wrapper has carried `max-h-96 overflow-y-auto` since #145, but
the class was a no-op until #568 defined `.max-h-96` in app.css. That
silently capped the whole config form at 24rem with a nested scrollbar.
Drop the cap so the form flows naturally and the page scrolls.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): give installed plugin card descriptions the full card width
The enable/disable toggle was a flex sibling of the whole text column
(name, metadata, description), so it reserved its width for the full
height of the card body. Descriptions wrapped into a narrow strip,
leaving blank space under the toggle and making cards very tall.
Move the toggle into a header row with just the name and badges, and
render the metadata and description below at full width.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
A failing plugin action returns a 400 whose JSON body carries the
script's own message, but both file-manager widgets threw it away:
plugin-file-manager's toggle always said "Toggle failed", and
json-file-manager's request helper threw "Server error 400" before
reading the body. That hid of-the-day's "Category ... not found in
config", which is why its toggles looked broken for no reason.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(web): render widget-less arrays of objects as a table, not comma text
An array of objects with no x-widget (geochron's `cities`) fell through to
the comma-separated text input. Jinja joined each item as a Python dict
repr, the save route read them back as a list of strings, and the schema
rejected them -- so every save of the plugin returned 400 "Configuration
validation failed", whatever setting was changed.
Default such arrays to the existing array-table widget, which already
edits arrays of objects and posts `field.N.key` inputs the save route
rebuilds into a list.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): don't leave an empty object stub in array items on save
The unchecked-checkbox pass walked into every nested object of an array
item looking for booleans, creating it when absent. A news custom feed
with no logo came out with `logo: {}`, which fails the logo's
`required: [id, path]`, so every save of the news plugin returned 400.
Recurse into a scratch dict instead and attach it only if a boolean was
actually set in it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(colour): clamp out-of-range text_color components instead of dropping them
_normalize_color returned None for a triple with a component outside 0..255,
and None means "not configured" to element_color -- so configuring
[300, 0, 20] silently handed the element its *default* colour rather than red.
Every scoreboard reads its per-element colours through this path, so the bug
reached all eight.
It is also the odd one out: sports_card.coerce_rgb and
SportsShared._coerce_rgb both clamp, and core's own test is named
test_coerce_rgb_clamps_rather_than_rejecting. The rejecting normaliser arrived
with the shared readers in 82a65ad2 (#425) while the eight plugins' colour
tests kept asserting the clamping behaviour they had before, so the two sides
have disagreed ever since.
Clamped inline rather than delegating to coerce_rgb: sports_card already
imports element_style, so importing back would be circular.
Adds the core assertion whose absence let this drift -- element_color had no
test covering an out-of-range component.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(element-style): cover ElementStyleResolver's own colour clamp path
CodeRabbit flagged that the new sports_card clamp regression test only
exercises element_color(); ElementStyleResolver._resolve() normalizes
configured colours through a separate call to the same _normalize_color,
comparing against a schema/classic reference to decide user_forced_color.
Add a resolver-level case so a future regression in that path (e.g. going
back to rejecting out-of-range components instead of clamping) is caught
too.
Mutation-checked: fails if _normalize_color rejects instead of clamps.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KP6kWxjUtJi72c56GaMmC8
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(element-style): clamp out-of-range colour components instead of rejecting
A regression this framework shipped. The eight scoreboards used to read their
colours through sports_card.coerce_rgb, which clamps; routing them through the
shared element_color sent them through _normalize_color, which rejected any
component outside 0..255 and fell back to the default. So a configured
[999, -5, 20] -- a typo'd bright red -- rendered white instead of (255, 0, 20).
Their own test_element_text_colors.py caught it: one case of nineteen, in all
eight plugins, failing only once the core change reached main.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): the style editor takes over its own blocks -- and gets to at all
Two defects, both found by rendering the real partial in a browser rather than
by reading the code.
It was losing a race to its own fields. The hand-off guard asked "do any
fallback controls differ from their server-rendered defaults?" as a proxy for
"is someone editing this?". But the fallback holds this block's own font
fields, and the font-selector widget populates them on the same 50ms timer --
so a plain page load, with nobody touching anything, raced into "dirty" and the
editor removed itself, leaving the 701-line accordion form it exists to
replace. Measured: seven customization.*.font selects dirty ~60ms after
injection, clean again by 400ms. The question is whether a *person* typed, and
event.isTrusted answers exactly that; the listeners now go on synchronously,
because the edit worth protecting can happen before initWidget runs.
It took over too much. Taking over removed the whole fallback section, but a
customization block can hold more than styling -- football keeps
favorite_result_colors there -- so that removed the only UI those fields had,
and the editor also rendered them as an element, giving every row an "enabled"
and three colour columns. Core now marks the blocks it recognises as styling
(the compact declaration already did; hand-written adoption did not), the
widget renders only those, and the template drops only the children the widget
reports owning.
Verified on football's real schema: 28 rows across four mode tabs, columns
Element/Font/Size/Colour/X/Y, favorite_result_colors still editable with its
ten inputs, no duplicated field names, no console errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): style editor no longer drops layout-only fields it never rendered
CodeRabbit flagged elementKeys() in style-editor.js: render() claims the
whole customization.layout child as the widget's own (removing it from the
generic fallback renderer, since posting the same offset twice is worse),
but elementKeys() only listed keys that also have their own top-level style
block. A hand-written schema can put a key under layout that never got one
-- a logo, a timeout indicator, a possession arrow with a position but no
font or colour -- and that key's only control silently disappeared: no row
in the style editor's table (elementKeys never listed it) and no fallback
section either (layout was removed wholesale).
elementKeys() now appends any layout-declared key not already covered by a
style element, so table() renders a row for it (layout columns only, no
style columns) and the wholesale layout ownership claim stays truthful.
Verified against current code before fixing. New regression test
(test/js/unit/test_style_editor_element_keys.js, following this repo's
existing eval-extraction pattern for testing widget JS without a browser)
fails against the reverted function and passes with the fix; added to
run_all.js and the suite table in test/js/README.md.
Full pytest suite: 4887 passed, 62 skipped, 2 failed -- both the
pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias gap on this sandbox,
identical on origin/main, unrelated to this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dpg3HLWohdCUdzz2QNHanm
* fix(web): style editor no longer strands leaf-valued layout fields
A prior fix on this PR made elementKeys() append any layout-only key with
no style block of its own (a logo, a timeout indicator, a possession
arrow), so table() draws a row for it instead of losing it when the
wholesale `layout` claim removes the generic fallback. That covers a
layout-only key shaped like an object (x_offset/y_offset, ...), because
columnsFor() only ever produced columns from a key's *sub-fields*.
It missed the case where the layout-only key's own value is itself a
leaf -- a plain "show_logo" boolean directly under layout, no x/y object
underneath. elementKeys() still lists it (any row: no matching column),
so it renders as an uneditable blank row and its only control -- the
generic fallback checkbox -- is still gone. Confirmed by executing the
real widget's render() against a synthetic schema in Node (a DOM-stub
harness, not committed): the field's name never appeared as an <input>.
columnsFor() now gives such a leaf key a column keyed to itself
('layout-leaf'), and elementRow() binds it to the leaf's own path
(customization.layout.<key>, matching the name the fallback would have
used) instead of leaving every cell blank.
New regression test (test/js/unit/test_style_editor_layout_leaf_columns.js,
following this PR's existing eval-extraction pattern) checks the leaf
column is produced, is self-keyed, doesn't duplicate, and that a schema
with no leaf-valued layout key is unaffected; wired into run_all.js and
the suite table in test/js/README.md.
test/js/run_all.js: 84 + 6 + 6 = all suites passed (jsdom unavailable
here, DOM suites skip as before). Python suite untouched by this change;
test_style_editor_extra_fields.py, test_style_editor_save_roundtrip.py
and the one PIL-dependent style_editor_takeover.py case fail identically
before this commit -- missing flask/PIL in this sandbox, not this PR.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): keep layout-leaf style-editor columns distinct from name collisions
columnsFor() keyed a layout-only leaf field's column by its bare field
name. If an unrelated element's style block or another element's layout
axis block happened to declare a sub-field with that same name, the
`!seen.has(key)` guard skipped creating the leaf's column, silently
dropping its only control again -- the same failure the leaf-column fix
was meant to close, just reached through a name collision (CodeRabbit
review on 324a7ea).
Key layout-leaf columns under a namespaced id so they can never be
shadowed by an unrelated column sharing their name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): CSS.escape() the owned key before it becomes a selector
container.dataset.ownedKeys round-trips schema property keys through a
DOM dataset attribute, and the takeover handoff spliced each one
straight into '[data-child-key="' + k + '"]' with no escaping --
inconsistent with this codebase's own convention elsewhere
(plugin-file-manager.js, app-shell.js's escapeCssSelector) for building
a selector from a dynamic value. A key containing a quote or backslash
would break the selector or be steerable; Codacy's static analysis
flagged this pattern (1 high ErrorProne finding on PR #569, current
head at the time) as a new issue, though its dashboard is unreachable
from this sandbox (egress to app.codacy.com is blocked) and the
check-run API returned no detail text -- verified and fixed by reading
the diff directly rather than the tool's own description.
Added a source-assertion regression test alongside this file's
existing ones (this behavior lives in an inline script no Python test
executes).
Full suite: 4888 passed, 62 skipped, 2 failed -- both the pre-existing
Europe/Kiev/Asia/Calcutta tzdata-alias gap in this sandbox, identical
on origin/main, unrelated to this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(web): every advertised layout offset gets a control in the style editor
The editor took the whole layout section over but matched offsets to style
rows by exact key. A hand-written schema's two blocks were never named alike --
football styles score_text but positions score -- so of football's eleven
positionable things only status_text had a control. Score, odds, both logos,
timeouts, possession, down-and-distance, date, time and records were options
the schema advertised and the renderer reads, reachable nowhere in the UI.
Core now resolves each style element's layout key through alias_keys, the map
the resolver already reads offsets with, and records it as x-layout-key. The
widget reads that rather than carrying a second copy of the rules, and posts
under the key the schema declares: football's own offset reader looks up
layout.score, so a value saved as layout.score_text would be kept and never
drawn. Layout entries no style element claims get an "Other positions" table
with its own columns, in every mode panel as well as the base one, in the order
the plugin declared them.
Verified in a browser against football's real schema: 92 of 92 layout fields
(23 base, 23 per mode) rendered exactly once under their declared names, none
posted under a style key, no duplicated field names, favorite_result_colors
still editable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(wifi): make Connect work from the setup AP
Joining a network from LEDMatrix-Setup has to take the AP down first, which
drops the phone that sent the request. The connect endpoint answered only
after the attempt finished, so the browser never got a reply and the WiFi
tab's Connect button appeared to do nothing.
- /wifi/connect answers 202 immediately while the AP is active and connects
in a background thread; the result (never the password) is reported via
/wifi/status as last_connect_attempt. A second connect while one is
pending gets 409.
- connect_to_network holds a /tmp flag for the attempt; the monitor daemon
skips AP management while it is fresh. Previously the daemon's
disconnected counter, accumulated over the whole AP session, re-enabled
the AP on its next tick in the middle of the connect.
- The WiFi tab and captive setup page explain the handoff up front, and on
reopening show why the last attempt failed. The wrong-password message
now works: the route sets the error_type the captive page checks.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(wifi): serialize connect attempts on both paths
Addresses CodeRabbit review on #571:
- Check for a pending attempt before branching on AP state. A background
attempt takes the AP down long before it finishes, so a second click
used to bypass the 409 and start a competing synchronous connect.
- Record pending for the synchronous (non-AP) path too, so two requests
can't overlap and have the first clear the daemon's in-progress flag
while the second is still connecting.
- Clear the pending state if the background thread fails to start, rather
than refusing every later request until restart.
- Say the setup network returns "within a few minutes": a stale flag plus
the daemon's grace period can take longer than one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Three independent changes, none of which alter runtime behaviour.
1. Remove ScrollHelper._get_visible_portion_subpixel and
_interpolate_subpixel (162 lines). get_visible_portion dispatches only to
_blend_visible_portion, so _get_visible_portion_subpixel had no caller, and
_interpolate_subpixel was reachable only from inside it -- a closed island.
_blend_visible_portion's own docstring already records that the scipy path
it replaced was dead; the replacement landed but the corpse stayed.
2. scripts/check_plugin.py: also search ../ledmatrix-plugins/plugins. The
scoreboards live in the sibling checkout, so --all silently skipped every
one of them and only --plugin-dir reached them.
3. .gitignore: ignore config/.config_secrets.json.tmp.*, which the suite
leaves behind several of per run.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(web): harden, polish and optimize the web UI per the September 2026 audit
Works through docs/archive/WEB_UI_AUDIT_2026-09.md (health 8/20).
Implementation integrity (P0)
- app.css now defines every utility class the templates and JS use,
including .hidden, so the ~145 JS show/hide toggles work. Button reset,
and base component rules (.btn, .form-control) wrapped in :where() so
utility classes on the same element win. New static-audit test fails
when a used utility class has no rule.
Accessibility
- Focus rings render (the old ring rule referenced undefined variables);
one :focus-visible outline everywhere; skip link; labelled nav landmarks.
- Shared dialog helper (js/utils/dialog.js): role/aria-modal, focus trap,
Escape, focus return, applied to every modal.
- Named icon-only buttons and labelled ~70 form fields.
- Toasts announced once; errors persist >= 10s; one showNotification.
- Captive WiFi page: live region, timeouts, dark mode, 16px inputs.
Performance (Pi Zero 2 W)
- SSE streams and tab timers pause when hidden or off-tab; the display
stream only runs while a preview is visible. app-shell.js deferred.
- Widget scripts served as one versioned bundle (/assets/widgets.js):
52 -> 21 script tags, 66 -> 35 requests on first load.
- Stdlib gzip fallback when flask-compress is missing: first-load JS/CSS
1358 KB -> 291 KB on the wire. SSE untouched.
Theming and responsive
- File managers, form fields and Fonts upload on theme tokens; bare
inputs themed in dark mode; no more white surfaces.
- No horizontal overflow at 375px on any tab; 44px touch targets on
coarse pointers; reduced-motion respected; header title truncates.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): clear Codacy findings on #568
- json-file-manager: focus-trap releases kept in a Map (no dynamic
property access or delete; no value-returning forEach callback)
- notification / schedule-picker: style and day-label lookups via Map
- app.js: move the pending-queue assignment out of the expression
- diff_viewer / error_handler: named function declarations instead of
arrow consts
No behavior change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test: check the OAuth widget ships in the widget bundle
base.html no longer tags widget scripts one by one; they load through
/assets/widgets.js. Assert the page requests the bundle and the bundle
contains google-oauth.js, which is what the test was protecting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): address review feedback on #568
- widget bundle version fingerprints every file (name, mtime_ns, size)
- gzip fallback appends Accept-Encoding to an existing Vary header
- dialog helper: releasing a non-top dialog no longer moves focus out of
the dialog the user is in
- labels: file-upload targets its file input; fallback config fields get
label for/id pairs; native color input has a fallback name
- utility audit also reads class names inside bound :class expressions
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): give the native color-picker input an accessible name
CodeRabbit flagged this on PR #568 as an outside-diff finding (never
posted inline, so it was missed in the round of fixes that addressed
the other 6 review comments). The <input type="color"> only carried a
title attribute; screen readers don't reliably announce title, and
there's no other label naming the control when showHexInput is false.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): clear Codacy findings in app-shell.js
- drop the unused catch binding on the SSE JSON parse
- move the pending-notification queue assignment out of the expression
No behavior change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): contain plugin widgets/ dir and bound style-editor retries
From CodeRabbit review on #568 (code that arrived with the main merge):
- serve_plugin_widget resolves widgets/ with resolve_under before
resolving the manifest script under it, so a symlinked widgets
directory can't become the containment base (CWE-22). New test.
- style-editor init stops polling after ~10s when the widget never
registers and leaves the plain fallback fields in place.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(sports): rebuild un-shared faces through the pinned layout engine
unshare_element_fonts re-instantiates a duplicate font face so two
elements can be told apart by id(). It did so through bare
ImageFont.truetype, which takes PIL's default layout engine rather than
the one src/common/font_layout.py pins. Raqm and Basic disagree on
fractional advances -- that disagreement is the reason the pin exists,
having broken golden images across machines -- so a rebuilt face could
measure differently from the shared face it replaced, on any host where
Raqm is installed.
These were the only two call sites in src/ bypassing the pin.
The guard asserts that the rebuild goes through the pinned loader rather
than comparing engine values: where Raqm is absent, bare truetype returns
BASIC anyway, so an engine comparison passes whether or not the pin is
honoured. The first draft of this test did exactly that and passed with
the bug reintroduced.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(web): drop the two dead client-side config-form renderers
generateConfigForm and generateSimpleConfigForm (580 lines) were defined
on the Alpine component and never called: server-side Jinja replaced them,
as pages_v3.py:641 records. Nothing in any template invokes them -- there
is no x-html in the templates and no bracket access on the component.
They carried their own x-widget dispatch, which made them an active trap:
the next person adding a widget would reasonably think both renderers
needed updating.
plugins/config_manager.js (PluginConfigManager, 133 lines) goes for the
same reason -- loaded on every page from base.html, referenced only by
itself and by an archived doc.
Kept, having checked them: widgets/example-color-picker.js is the worked
example docs/widget-guide.md points plugin authors at, and
widgets/plugin-loader.js is the client half of a documented feature
(manifest-declared plugin widgets) whose server route is missing --
soccer-scoreboard already ships a widgets/custom-leagues.js that this
loader is meant to fetch. That is an unfinished feature to complete, not
dead code to delete.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(web): serve plugin-declared widgets, and actually ask for them
LEDMatrixWidgets.loadPluginWidget has always fetched
/static/plugin-widgets/<plugin>/<widget>.js, and docs/widget-guide.md has
always documented that path, but nothing served it. soccer-scoreboard has
shipped a 17KB widgets/custom-leagues.js since August that could never
load. Both halves were missing, not just the route:
- serve_plugin_widget serves the script from the plugin's widgets/
directory as text/javascript. The manifest is the allowlist -- only a
widget the plugin declares is reachable -- so installing a plugin does
not publish everything it ships. Path handling mirrors the sibling
serve_plugin_web_ui: allowlist regexes, os.path.basename, resolve() +
relative_to() containment, and the ledmatrix- prefix fallback. The
declared script name is guarded too, since it comes from the plugin
rather than the request.
- The config form never requested one. Its x-widget dispatch is a
hardcoded list of core widget names, so a plugin's own widget fell
through to a plain text input. An unrecognised x-widget on a string
field now asks ensureWidget() for it. The text input stays as the
fallback and is removed only once the widget has actually rendered, so
a missing or broken widget costs the user an editor rather than their
configured value on the next save.
- manifest_schema.json gains "widgets", so the declaration is validated
rather than merely tolerated by additionalProperties.
Verified in a browser against the real partial: a declared widget loads,
registers and renders, and its field posts exactly one value; a field
whose widget 404s keeps its text input and still posts its value.
Not addressed: loadPluginWidgetsFromManifest still has no caller. The
per-field ensureWidget path is lazier and is what the form now uses, so
that bulk helper is dead weight -- worth removing, but left alone here
rather than inventing a call site for it.
Known limitation, documented: only string-typed fields take this path.
object/array/boolean/number fields and enums are dispatched by the
template's own branches, which still only know core widgets.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(element-style): a wrong-size BDF now keeps its font, not its size
BDF fonts are fixed-size bitmap strikes: FreeType accepts only the pixel
size baked into the file and raises for anything else. 32 of the 35
shipped fonts are BDF, so a size picked in the web UI usually is not a
valid strike -- and load_font caught that failure with its generic
"unloadable font" handler, which substitutes PressStart2P. Asking for
5x7.bdf at size 10 therefore rendered a completely different typeface,
silently.
It now falls back to the file's own native size instead, which is what
SportsCore._load_custom_font_from_element_config has always done. The
native size is read via FontManager._read_bdf_native_size rather than a
fourth copy of that parser, matching how core.py already delegates.
Also here, because they are the same code path:
- native_bdf_size() is exposed for the web UI, which needs to know when a
size field can take effect at all. None means "free choice".
- ElementStyle.font_size now reports the size actually realised rather
than the one requested. Callers lay out from it, and reserving space
for a size nothing was drawn at is how this surfaces.
- The module font cache is a bounded LRU (256) instead of an unbounded
dict. The display process runs for weeks and every config save can add
a (font, size) pair; every other hot cache in the codebase is bounded
this way.
Untouched configs are unaffected: the shipped classic fonts are the three
TTFs, so nothing was hitting the substitution path by default.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(element-style): per-mode style and offset overrides
Lets one element be styled differently per situation -- a scoreboard's
live / upcoming / recent cards, weather's current / hourly / daily
screens -- under customization.modes.<mode>.
The mode is bound at construction rather than passed per call. That is
what makes this cheap to adopt: SportsUpcoming and SportsRecent are
already separate instances with distinct SKIN_MODE values, so binding
once makes every existing style()/offset_value() call site mode-aware
without editing any of them. A per-call mode argument exists for the rare
host that renders more than one mode.
The two layers answer different questions, deliberately:
- The base layer keeps the existing "differs from the schema default"
rule, because the save flow writes the full default object into
config.json whether or not the user touched it.
- A mode layer is pure override -- its fields default to None, so
presence is intent. Nothing writes into it unasked, so there is nothing
for the stricter rule to protect against.
None therefore means inherit, and has to stay distinct from 0: a mode
y_offset of 0 means "sit at the base position", not "no preference".
This is the same distinction scroll_card.switch_* draws with "inherit".
A malformed mode value falls back to the resolved base value rather than
to the caller's default -- caught by the degradation tests, which is what
they are for: resolving the mode first let one bad string in a mode block
silently discard a good base offset.
With no modes block, and for every existing caller, resolution is
unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(element-style): declare per-mode overrides in config_schema.json
A plugin adds "x-style-modes": ["live", "upcoming", "recent"] alongside
its x-style-elements declaration and gets a customization.modes.<mode>
group per mode, with every field of every declared element repeated as an
override.
Those override fields are typed nullable and default to null, which is
the whole trick. The save flow writes schema defaults into config.json
wholesale, so giving a mode field the base element's default would make
every mode a frozen copy of the base the first time a user pressed Save,
and the base would stop reaching them. Null means inherit. The mutation
test for this is explicit: with concrete defaults, a base font_size of 14
resolves as 10 with user_forced set.
min/max from the declaration carry into the mode blocks, so an
out-of-range override is rejected by validation rather than clamped
silently at render time.
Also: the emitted font field now carries "x-widget": "font-selector". The
widget already shipped and the config form already allowlisted it -- the
hint was simply never emitted, so the field rendered as a bare text box
that the user had to type a font filename into.
Verified through the real SchemaManager path -- load_schema, defaults
extraction, merge_with_defaults, validation, then resolution -- rather
than against a hand-built dict, since the thing at risk is what that
pipeline does to a null.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): render the config form from the schema the save route validates
The form read config_schema.json with a raw json.load while
api_v3.save_plugin_config went through SchemaManager. Those are not the
same schema: SchemaManager applies expand_style_elements, which turns a
compact customization.x-style-elements declaration into the per-element
blocks the form knows how to render.
Without it, that customization object has an x-style-elements key and no
"properties", so the template's object branch matched nothing and the
section rendered as empty space -- while saving still validated against
the expanded shape. of-the-day ships the compact form, so its
customization section has been invisible in the web UI.
pages_v3 gains a schema_manager the way it already has config_manager and
plugin_manager. use_cache=False matches the save route, so an edited
schema is not served stale during plugin development. The raw read stays
as a fallback for callers that register this blueprint without one.
Checked before making the change: load_schema does nothing here except
read, validate and expand -- inject_skin_selector is a separate method it
does not call -- so this is not a behaviour change for schemas without
the declaration.
The test pair renders the same compact schema with and without a
SchemaManager, so it documents exactly what was broken as well as what is
fixed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(web): style-editor widget -- a row per element instead of 65 accordions
Rendered element by element, a realistic scoreboard's customization block
is 65 nested sections, and reaching one per-mode font size takes five
levels of expanding. The widget collapses that to one compact row per
element -- font, size, colour, X, Y -- with a tab per declared mode.
It emits ordinary inputs under the same dotted names the generic renderer
would produce, so the save/validate/merge pipeline is untouched: no hidden
JSON blob and no new server-side parsing. It is driven entirely by the
schema block it is handed, so fields added to the schema later appear
without editing the widget. If it fails to load or throws, the generic
nested rendering it replaces is left in place.
Fixing two things the save path got wrong for nullable fields, found by
posting what the widget actually emits:
- The indexed-array recombiner (text_color.0/.1/.2 -> one list) compared
the declared type to the string 'array', so a per-mode colour, typed
["array", "null"], was never reassembled and failed validation on save.
_parse_form_value_with_schema had the same comparison.
- A blank nullable field became [] rather than None, which then failed the
minItems the colour array declares. Null is the inherit sentinel, so it
has to survive.
And two things the widget itself got wrong, found by looking at it:
- An unset base control fell back to the select's first option, so an
untouched scoreboard claimed every element used 10x20.bdf -- and the
size box then locked itself to that bitmap font's fixed size. Base
controls now show the schema default; mode controls stay blank, because
blank there means inherit.
- Elements arrived alphabetised (Detail and Odds above Score). Flask's
JSON provider sorts keys, so declaration order has to be stated
explicitly; expand_style_elements now emits x-propertyOrder, which the
generic renderer already honoured too.
Size is disabled and shown as fixed for a bitmap font, using the
scalable/native_size the font catalog now reports.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(element-style): visibility, alignment and scale per element
Completes the customization vocabulary: hide an element, align it, and
resize a logo, alongside the font/size/colour/offset that already existed.
All three per mode.
They resolve to "change nothing" until the user asks for something -- True,
None and 1.0 -- rather than to whatever the schema declares. That is the
same invariant the font fields keep: a caller that honours them still
renders an untouched config exactly as it did before they existed. A
schema default therefore does not count as a choice, which matters because
the save flow writes that default into config either way.
scale sits in the layout block with the offsets rather than in the element
block, because it is geometry: a logo has a scale and no font. The widget's
columns come from the schema, so a logo row shows visibility, offsets and
scale and no empty font cell.
Two bugs found by the tests rather than by reading:
- A nullable enum needs null in its enum list, not just in its type. The
mode copy of `align` defaulted to null and then failed its own schema, so
a plugin declaring any enum field with modes could not save at all. Six
tests failed on this before any of them reached what they were testing.
- defaults_from_schema only ever extracted font/font_size/text_color, so
the schema defaults for the new fields were invisible to the resolver and
a declared default read as a user choice.
Widget: the table scrolls horizontally and pins the element-name column.
Nine columns do not fit the config panel, and clipping them hid the offsets
entirely while scrolling them made every row anonymous.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(element-style): resolve elements under the names plugins actually use
Two naming conventions collided as the scoreboards grew. Counted across
the published schemas: the style block names elements with a _text suffix
(score_text, status_text, detail_text), while the layout block mostly uses
the bare noun (score, date, time, odds) -- except status_text, which kept
the suffix in seven plugins and lost it in two. records vs record splits
seven to two the same way.
A lookup now tries the exact name first and then the spellings that mean
the same thing. Exact-first is what makes this inert for any config that
already matches; the aliases only decide cases that resolved to nothing
before.
This is also what makes migrating to the compact declaration form safe.
That form uses one key for both blocks, so a scoreboard adopting it asks
for layout.score_text while its users have layout.score saved -- without
the aliases, every offset they had dialled in would silently become 0.
Applies to the style block, the layout block, the schema defaults and the
per-mode overrides, since the drift shows up in all four.
Not attempting to canonicalise on write: renaming keys in config.json
would break the plugins still reading the old spelling from their own
bundled code, and the drift costs a dict miss rather than correctness.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(plugins): BasePlugin.styles -- per-element styling every plugin inherits
Adopting the element-style system meant repeating three things in every
plugin: a guarded import, finding its own config_schema.json, and
rebuilding the resolver when on_config_change swapped the config dict.
This is those three things once, on the class all 45 plugins already
inherit from.
title = self.styles.style('title_text',
classic_font='PressStart2P-Regular.ttf',
classic_size=8, classic_color=(255, 255, 255))
The classic_* arguments are the adoption contract: with nothing configured
they come back verbatim, so a plugin that switches to this renders exactly
as before until a user changes something.
A plugin with one instance per display mode sets STYLE_MODE on the class
and every existing lookup becomes mode-aware without a call site changing
-- which is the point of binding the mode to the resolver rather than
passing it per call. styles_for() covers a plugin that renders several
modes from one instance.
Schema discovery reads the concrete class's own module rather than this
file, because this file lives in src/plugin_system where no plugin schema
exists -- the same trap SportsCore._config_schema_path documents. The
first mutation test for that passed anyway: an installed plugin's module
directory and its entry under plugins_dir are the same path, so the test
could not tell the two apart. The case where they diverge is a plugin
symlinked in for development, and the test now forces that shape.
Getting discovery wrong is silent rather than loud: with no schema the
resolver has no defaults to compare against, so every configured value
reads as a deliberate override and the plugin quietly stops honouring its
own shipped styling.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(element-style): adopt hand-written customization blocks, and widen the font list
Nineteen plugins spell their style elements out longhand instead of
declaring them -- football's block is 701 lines for seven elements -- and
predate this system entirely. Core now recognises that shape, so they pick
up the row-per-element editor and the real font picker on a core update
rather than on a plugin release. Checked against every published schema:
21 plugins adopt, and the defaults of each still validate against the
schema generated for it.
Detection requires *every* field in a block to be one this system
understands. A looser "has at least one style field" rule sweeps in
baseball's `count`, which carries a text_color beside geometry that means
nothing here. That distinction took three attempts to test: the first two
assertions passed under both rules, because an over-eager rule leaves a
fontless block looking untouched and only surfaces as an extra row in the
editor.
The hardcoded font enum is replaced rather than extended. Football lists
five of the thirty-five installed fonts, which is why a font a user
uploads can never appear in one. It is not a curated safe set -- it omits
some twenty other faces that fit the declared size cap just as well -- it
is the fonts that happened to exist when it was written.
Widening it does need a guard, though, and not the one the schema already
has: a bitmap font ignores font_size and renders at its size baked into
the file, so `maximum: 16` cannot stop a 27px face. The picker now filters
out fixed-size fonts taller than the element's own declared ceiling, which
drops exactly the four that would overflow a 32px panel and keeps the
other thirty.
Per-mode overrides stay opt-in: core cannot invent a plugin's display
modes, so `x-style-modes` remains the one line that unlocks them. Their
layout half covers every positionable element rather than only those with
a style block -- the two namespaces do not line up in a hand-written
schema, and football positions six things (logos, timeouts, possession)
that have no style block at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(web): remove the two Fonts-tab panels that reported invented data
"Element Font Overrides" let a user configure an override, showed a
success toast, and changed nothing. All three endpoints behind it were
stubs -- GET returned a hardcoded {}, POST and DELETE returned success
without calling anything -- each marked "This would integrate with the
actual font system".
Wiring them to FontManager would not have fixed it. The machinery there is
real (_load_overrides/_save_overrides persist config/font_overrides.json,
resolve_font applies them, and the countdown plugin genuinely consumes
it), but the panel's element dropdown offered eleven invented keys --
nfl.live.score, clock.time, weather.current -- that no plugin has ever
read. An override saved against one of those would have persisted
correctly and still done nothing.
"Detected Manager Fonts" goes for the same reason. It claimed to show
"fonts currently in use by managers (auto-detected)"; its own comment said
"we'll simulate this", and it listed every font in the catalog with a
hardcoded usage_count of 1 -- the panel beside it, with fabricated
numbers attached.
Per-element font choice now lives in each plugin's own config editor,
against the elements that plugin actually has, and covers size, colour,
offsets, visibility, alignment and scale rather than family and size.
Kept: the font library (upload, preview, delete), which works, and
/fonts/tokens, which is a stub but genuinely feeds the preview's size
dropdown. FontManager's override methods are untouched -- countdown uses
them.
Verified in a browser with the tab's JS running: no console errors, 35
fonts listed, upload and preview intact. Removing the panel meant unwiring
it from populateFontSelects too, which would otherwise have bailed out
early on the missing select and left the preview dropdown empty.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(sports): one reader for element colours and layout offsets
There were two copies of the per-element colour read and three of the
layout-offset read. They had already drifted -- the scroll-card renderer
carries a comment about having ignored offsets its own schema advertised
-- and each new capability had to be added to all of them or silently work
in some places and not others.
All of them now go through src.element_style, which is what carries the
alias handling and the per-mode lookup. That lands immediately for the
nine plugins importing these modules: a scoreboard asking for `score_text`
offsets finds the `layout.score` its users configured, and a Live instance
resolves its own colours through SKIN_MODE without any call site passing a
mode.
_normalize_color learned "#RRGGBB" in the process. The scoreboards' own
readers have always accepted it, so the shared one had to, or consolidating
would have quietly dropped a form users' configs may hold. _coerce_offset
picked up the non-finite guard the scroll-card reader had and the other two
did not.
_get_layout_offset is promoted onto SportsCoreSharedMixin. Each plugin
still carries its own copy in its bundled sports.py, which wins by MRO --
so adopting this is a deletion in the plugin, and until that deletion
nothing changes for it.
Note for whoever runs the suite next: test_display_dirty_tracking.py is
order-dependent. Fifteen of its tests failed in one full run and passed in
the next with no change in between, and pass in isolation. Pre-existing,
unrelated to this, but it makes a full-run diff untrustworthy until it is
fixed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(changelog): record the element-style work under Unreleased
This file's own preamble asks for it: a plugin may delete its bundled
fallback copy of a core module only when its manifest floors on the first
release that shipped that module, which requires the additions to be
recorded here against a version.
Names a plugin can now import and floor on -- the stateless layout_offset
and element_color readers, alias_keys, native_bdf_size, the resolver's mode
binding, BasePlugin.styles, and the promoted
SportsCoreSharedMixin._get_layout_offset -- plus the schema and web-UI
changes, the four fixes and the three removals.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(fonts): log the BDF native-size read failure instead of swallowing it
The bdf-native-size lookup in get_fonts_catalog() caught any exception
and silently discarded it. Every other guarded read added in this PR
(the manifest parse in _declared_widget_script, the SchemaManager
fallback in _load_plugin_config_partial) logs before falling through
to the same degraded behavior. This one didn't, which is the shape a
silent-exception-swallow lint rule flags. Behavior is unchanged --
native_size still comes back None -- but a corrupt or unreadable BDF
file now leaves a trace.
Verified: font-related tests (140) and the full suite still pass,
with only the 2 pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias
failures already present on origin/main.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: address CodeRabbit findings on the style-editor/font-selector PR
- Fix _load_font_sized double-wrapping the (font, size) tuple on the
missing-font path, which handed callers a tuple instead of a font.
- Fix _set_nested_value skipping an explicit None when the key already
existed, which silently kept stale overrides when a user cleared a
nullable per-mode field or blanked all channels of an indexed color.
- Preserve BDF scalable/native_size metadata through fetchFontCatalog's
catalog-format mapping so maxFixedSize filtering actually applies.
- Stop caching an empty array on a failed font-catalog fetch so a later
call can retry instead of being stuck with the failed result.
- Keep a saved font selected in the style editor even when it no longer
fits a newly declared maxFixedSize, instead of silently deselecting it.
- Don't drop in-progress user edits to fallback fields when a plugin
widget finishes loading asynchronously and takes over the form.
- Tighten the removed font-override endpoint test to assert 405, not
just != 200.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): a partial save no longer switches off checkboxes it never showed
An HTML checkbox posts nothing when unchecked, so the save route walked the
schema and forced every boolean missing from the form to False. That is right
for the rendered form and wrong for every other caller: a script, the MQTT
bridge or a curl against the documented endpoint never rendered a checkbox, and
reading its silence as "all off" turns a one-field save into a mass disable.
Found on hardware. Posting four customization.* keys to a live device switched
off nfl.enabled, ncaa_fb.enabled and every display-mode toggle in one request.
The form now reports the top-level sections it drew (__rendered_section), and
inside those an absent checkbox still means unchecked -- including a section
whose only fields are checkboxes that are all off, which no heuristic could
recover. A post with no marker only touches objects it actually posted a field
from. Meta fields are dropped before form keys are treated as config paths,
because unknown keys are otherwise written straight into config.json.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(sports): resolve element colour by name, and honour visible/align/scale
Two of the three gaps this framework shipped with.
Colour by name. A draw resolved its colour by comparing the *identity* of the
font object it was handed, which cannot tell two elements apart when they share
a face -- so those draws went out white. Every bitmap font is in that case,
because a freetype.Face cannot be re-instantiated to un-share it, which is how
an element rendered in any of the 32 shipped BDF fonts silently lost a colour
its picker had offered all along. _draw_text_with_outline now takes
element="score_text" and reads the colour by name; the identity path remains
for un-annotated callers, but narrows before giving up -- one configured colour
among the sharers is the only thing the user can have meant.
Visible, align and scale. The resolver has understood these since the
framework landed and nothing consumed them: an element could be marked hidden
in the web UI and still render. Adds the stateless readers, the mixin
accessors, and a scale parameter on the one shared logo-sizing seam (keyed into
the cache, so two elements scaled differently cannot be served each other's
image). Naming an element in a draw also honours its visibility.
Untouched configs are unaffected: every new parameter defaults to today's
behaviour, and all ten affected plugins render pixel-identically to main across
every harness size.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(plugins): how to declare styleable elements; harden the widget's lookups
The plugin-author guide for the compact x-style-elements declaration -- what
each key does, how to read values back without breaking the "user-forced only
when it differs from the default" rule, and why a hand-written block needs no
changes to be adopted.
Also clears the static-analysis findings on style-editor.js. Every lookup in
that file is keyed by something out of a schema or a saved config, so a key of
__proto__ or constructor would walk the prototype chain and hand back a
function instead of a schema; reads now go through an own-property helper. The
panel registry became a list, and the flagged vars moved to their function
roots.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): clear the remaining static-analysis findings
Five, all on lines this branch touched.
The Python one is not a new defect: _set_missing_booleans_to_false's first
parameter was always named `config`, which shadows the `config` submodule
imported for its side effects at the bottom of this module. Editing the
signature simply put the existing warning on a changed line. The parameter is
the plugin's config dict, so `plugin_config` is what it should have been called
anyway; callers pass it positionally and are unaffected.
The JavaScript ones are the object-injection rule firing on reads keyed by
data. own() now goes through a property descriptor, so the one unavoidable
data-keyed read is no longer a computed member access; at() consumes its path
instead of indexing it; and the column set is a Map, which has no prototype to
pollute and needs no guarded reads at all.
Verified the widget still renders identically against football's real schema:
29 element rows, all four mode tabs, values populated, no console errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): drop the hasOwnProperty alias the descriptor read made redundant
own() now reads through Object.getOwnPropertyDescriptor, so the alias it used to call has no remaining reference.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(fonts): load 4x6 on its pixel grid, from any working directory
`extra_small_font` loaded 4x6-font.ttf at 6, off the face's 7px grid.
Under `draw.fontmode = "1"` the mono rasteriser thresholds each glyph at
50% coverage, so every glyph lost its fourth column and deformed:
christmas-countdown rendered "UNTIL" as "VM1JL". The advance is 5px at
both sizes, so snapping to 7 reflows nothing.
- Sizes in DisplayManager._load_fonts go through crisp_size() instead of
literals. crisp_size / FONT_PIXEL_GRID / FONT_NAME_ALIASES move to
src/common/font_layout.py; sports_card re-exports them.
- Mirror the fix in VisualTestDisplayManager, the harness's fork of
_load_fonts. Without it every golden is blessed at the old size.
- Resolve bundled font paths against the install root, not the cwd.
FontManager._resolve_asset_path now delegates to
font_layout.resolve_asset_path (kept by name; plugins probe for it).
- The startup banner's middle rung snaps to 7; the 5 rung stays off-grid
on purpose (the only size that fits a dotted quad on 64px).
- loading.py reads all plugin JSON as UTF-8 (cp1252 on Windows aborted
check_plugin.py on a 0x9d byte).
- check_plugin.py reports in ASCII and never dies on an unencodable char.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(fonts): resolve relative asset paths from the install root, not the cwd
resolve_asset_path checked os.path.exists(relative_path) unconditionally,
so a relative asset path was still resolved against the process cwd first
-- exactly the dependency this module exists to remove. An unrelated
working directory that happens to contain assets/fonts/4x6-font.ttf (a
stale checkout, a copied assets folder, another project) would shadow the
real bundled font instead of the install root ever being consulted.
Only an absolute path is now returned as-is; a relative path always
resolves against _INSTALL_ROOT first, matching the docstring's stated
contract. FontManager._resolve_asset_path delegates to this function, so
it's covered by the same fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* docs: add PRODUCT.md product context for web UI design work
Captures durable product truth (users, positioning, operating context,
constraints, principles) so design passes on the web control panel share
one source. Open decisions (offline-only, CSS build step, WCAG target)
are recorded as undecided rather than adopted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: add PRODUCT.md and September 2026 web UI audit
PRODUCT.md captures durable product context (users, positioning,
operating context, constraints, principles) for web UI design work.
Open decisions (offline-only, CSS build step, WCAG target) are recorded
as undecided rather than adopted.
docs/archive/WEB_UI_AUDIT_2026-09.md records the technical audit of
web_interface/ (8/20): the hand-rolled Tailwind subset in app.css leaves
333 used utility classes undefined (including .hidden), focus rings never
render, modals lack dialog semantics, and SSE/polling never pause. Includes
a verified-and-rejected section so the cache-busting false positive is not
re-raised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(config): stop same-second backups overwriting each other
A backup's version is its identity. save_config_atomic() hands the path
back, rollback_config(backup_version=...) looks that version up, and the
paired secrets backup is found by reusing the same string.
The version was stamped at second granularity, so two saves inside the
same second produced the same filename and the second shutil.copy2()
silently overwrote the first backup. The path a caller was still holding
then pointed at different content, and rolling back to it restored the
wrong config. A user saving twice in quick succession lost a restore
point with no error.
list_backups() made it worse. It parsed the version off Path.stem, which
drops only the last dot-component, so for config.json.backup.20240101_120000
parts was ['config', 'json', 'backup'] and parts[-2] was 'json' -- never
'backup'. The filename branch was unreachable: every backup fell through
to the mtime fallback and reported a second-granularity restamp of its
mtime rather than the name on disk, so a unique filename alone would not
have been enough for rollback to find the right version.
Stamp microseconds, and never overwrite an existing backup -- on a
collision bump a -N suffix rather than lose a restore point. Parse the
version off the exact glob prefix so it round-trips with the filename,
still reading the legacy second-granularity format so restore points that
predate this keep working.
Two tests had encoded the bug:
- test_multiple_config_changes asserted a rollback produced plugin1=45
with plugin2=15, a state no single backup ever held -- 45 was only in
the second backup, 15 only in the first. It passed because the two
saves collided onto one file, so the first version resolved to the
second's content. Corrected to the state that backup actually holds.
- test_backup_rotation asserted against a hardcoded max of 3 while
setUp configured 5, and still passed: every save in its loop collapsed
onto a single filename, so there was only ever one backup to count and
rotation was never exercised. It now asks the manager for its limit
and overshoots it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(config): fold collision suffix into ordering, close backup-path race
_parse_backup_version() stripped any trailing "-segment" unconditionally,
so a collision-suffixed backup parsed to the exact same timestamp as its
sibling and list_backups() had no deterministic way to order them. Only
strip the suffix when it's numeric, and fold it back in as extra
microseconds so same-tick collisions sort newest-first reliably.
_create_backup() also checked backup_path.exists() before shutil.copy2(),
which two concurrent callers can both pass for the same path -- the second
copy2() then silently destroys the first call's restore point. Reserve
each path (config and, when configured, secrets) with exclusive file
creation instead of a check-then-copy, retrying on a real conflict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Follow-up to #562. That commit fixed the actual cause of the intermittent
15-test failure in test_display_dirty_tracking.py -- the emulator's fixed TCP
port 8888, a machine-wide singleton that a concurrent pytest process takes
away. This adds the two things that would have made it a five-minute
diagnosis instead of a long one, and closes the other door into the same
failure.
Confirmed the module is order-independent as it stands, on this checkout:
pytest test/ -q, three times 115 failed / 4464 passed / 63 skipped,
byte-identical failure sets, the
module 21/21 passed each time
module forced last (197 files first) identical failure set
module forced first identical failure set
module after each of test_display_manager, test_display_controller,
test_display_controller_vegas_tick, test_skin_system, test_sports_scroll,
test_initial_update_budget, test_display_double_parity,
test_initializing_screen all pass
four concurrent processes on the file 21/21 each
And reproduced the original, to be sure the diagnosis in #562 is the whole
story. Holding 0.0.0.0:8888 from a separate process:
HEAD's test/conftest.py 21 passed
pre-#562 test/conftest.py 15 failed, 6 passed
The 15/6 split is not arbitrary: the six survivors are the only tests in the
file that never touch dm.matrix.
conftest.py: DisplayManager is a process-wide singleton and the RGBMatrix /
RGBMatrixOptions names it constructs through are module globals, bound once at
import. All three are shared by every test module in the run, so a module that
leaves an instance in _instance -- or leaves patch('src.display_manager.
RGBMatrix') standing -- changes what the NEXT module builds, invisibly, and
only in a full run. A module-scoped autouse fixture now resets the singleton
and restores either binding if a patch outlived its module. Module-scoped
rather than per-test so that files sharing one manager across their own tests
keep doing so; only the leak across the module boundary is cut. Autouse
fixtures are set up ahead of requested ones, so this is finalised after a
module's own DisplayManager fixture. Verified with a throwaway pair of probe
modules -- one leaks a patch and a singleton, the next asserts both are clean
-- which passed and were then removed.
test_display_dirty_tracking.py: _setup_matrix() swallows every construction
failure and falls back to matrix=None, so a broken environment arrived as
fifteen identical "'NoneType' object has no attribute 'SwapOnVSync'" errors
naming neither the fixture nor the cause. The fixture now fails once, and
says where to look; under a held port it reads
DisplayManager fell back to matrix=None: RGBMatrix construction raised...
Known causes: the emulator adapter losing a fixed TCP port to another
process -- see pytest_configure in test/conftest.py -- or a
patch('src.display_manager.RGBMatrix') leaked from an earlier test module.
with WinError 10048 in the captured log directly above it.
No regressions: full suite with both changes is 115 failed / 4464 passed /
63 skipped, failure set identical to the pre-change baseline. The 115 is the
pre-existing Windows-environment baseline (os.geteuid, POSIX modes, fcntl);
CI on Linux remains authoritative.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* chore: stop tests and rigs writing to shared paths
Two shared-state problems, both of which show up as a permanently dirty
checkout or an unreproducible test failure.
test_display_dirty_tracking.py builds a real DisplayManager, whose
_snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web
UI reads. Every pytest process on the machine shares that one file, so two
concurrent runs -- CI shards, a second worktree, an agent running the suite
alongside -- overwrite each other's snapshot and the mtime assertions stop
meaning anything. The module fixture now points it at a session-unique temp
path; the individual tests that care still override it further.
To be clear about what this does and does not fix: this is a real shared-path
hazard, but it is NOT the cause of the intermittent 15-test failure in that
module. That turned out to be the emulator's fixed TCP port, fixed in the
follow-up commit. This change stands on its own merits.
web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json
and data/operation_history.json as the web interface runs, into a directory
that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig
that ever opened the web UI -- and every test run that constructs the app --
left three untracked files behind and a permanently dirty `git status`. Only
data/.gitkeep is tracked under data/, so the negation keeps it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test: stop the emulator binding a fixed port, so concurrent runs can't collide
This is the cause of the intermittent full-suite failures we have been chasing:
runs of identical code landing anywhere between 100 and 130 failures, while
every implicated test passed in isolation.
Six test modules set EMULATOR=true and build a real DisplayManager. The repo's
emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to
serve the dev preview. That port is a machine-wide singleton, so a second pytest
process -- a CI shard, another worktree, an agent running the suite alongside --
loses the bind. RGBMatrix construction then raises, DisplayManager catches it and
falls back to `self.matrix = None`, and every test that subsequently touches the
matrix dies with
AttributeError: 'NoneType' object has no attribute 'SwapOnVSync'
which names neither a port nor a socket, and points at the wrong file entirely.
Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its
matrix-touching tests fail together or not at all -- the 15-test swing that made
the totals look random.
Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process
and running test_display_dirty_tracking.py:
without this change 15 failed, 6 passed
with this change 21 passed
The "raw" adapter renders in memory and binds nothing. Only display_adapter is
overridden, in a throwaway config written per pytest process; the repo's
emulator_config.json is untouched and `run.py -e` still opens the browser
preview on 8888. Nothing in the suite referenced the adapter, and the tests
wrap SwapOnVSync on the matrix object itself, so they are indifferent to what
sits underneath. allow_adapter_fallback is forced off -- falling back would
land us on the browser adapter and its fixed port, which is the whole problem.
CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set
to an absolute path: the previous behaviour depended on where pytest was invoked
from, and silently wrote a default config into whatever directory that was.
Verified no regressions: full suite on this branch and with origin/main's
versions of the touched files, same machine, back to back -- 115 failed /
4347 passed on both sides, zero failures unique to either. That 115 is the
pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell
scripts, Linux-only binaries); CI on Linux remains authoritative.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore: mark the shell entry points executable
Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh`
fails with "Permission denied" and only works if you know to prefix `bash`.
That one matters most: the web UI's own error hint, added in #560, tells users
to run exactly that path when a system action fails for want of passwordless
sudo, and following that instruction verbatim did not work.
All eleven carry a shebang and are invoked directly, never sourced. The two
sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately
left non-executable, which is what distinguishes a library from an entry point.
Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied
with `git update-index --chmod=+x` because this checkout is on Windows, where
core.fileMode is off and the working-tree bit is not tracked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(web): escape quotes in every HTML escaper, not just & < >
The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:
value="${escapeHtml(v)}" title="${escapeHtml(v)}"
so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.
It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.
app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.
cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `'`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.
url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).
test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): stop request-supplied names from reaching paths outside their base
Three of the py/path-injection alerts were live, not lint:
* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
whose resolved path *string-prefixed* the plugin directory. Flask's
default converter forbids a slash but not dots, and
get_plugin_directory('..') returned the parent of the plugins directory
because it exists -- so every file under the project root then prefixed
that directory, config/config_secrets.json included. The prefix check
was also wrong on its own terms: with plugin dir "plugin-repos/foo",
"../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
start with "plugin-repos/foo".
* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
file_id of "../../../../etc/something" deleted that file. This is the
one finding in the batch that destroyed data rather than exposing it.
* POST /api/v3/cache/delete passed the body's key through
CacheManager.clear_cache to DiskCache, which joined it as a filename and
called os.remove. Same shape, same result. The guard goes in
DiskCache.get_cache_path, the single choke point get/set/clear share, so
every caller is covered rather than just this route. Real keys are the
stems of files already flat in the cache directory -- that is how
list_cache_files derives them -- so nothing legitimate is turned away.
The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.
Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.
test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): refuse a plugin id that is not a plain name, don't truncate it
pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.
Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(web): say what the plugin web_ui iframe actually is
The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.
Also drops the now-unused os/os.path imports.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): inline url-input's scheme guard at the previewLink.href sink
CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.
Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.
Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): address CodeRabbit findings on the CodeQL triage PR
- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
credential-saving or connect flow runs. NetworkManager only accepts
printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
was previously saved/attempted before nmcli itself rejected it.
- web_interface/blueprints/api_v3/config.py: fail closed when the
plugin config schema path can't be resolved under the plugins
directory (e.g. a symlinked plugin dir). Previously this fell
through with secret_fields left empty, so submitted credentials for
that plugin were saved as ordinary, unencrypted configuration.
- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
splicing the JSON day/column key into an inline oninput="..." handler
string. escHtml() escapes quotes for a normal HTML attribute, but the
browser HTML-decodes the attribute before running it as script, which
undoes that escaping and lets a crafted column name (e.g. from an
uploaded JSON file) break out of the JS string and execute. Cell
edits now travel through data-day/data-col attributes read by one
delegated 'input' listener instead.
While in this file: fixed 6 pre-existing missing-')' typos on
multi-line safeSetHTML(...) calls (already flagged by Biome in this
PR's own CodeRabbit run as syntax errors blocking its lint pass).
These predate this PR (present on main too) but made the whole file
fail to parse in any JS engine, which is a bigger problem than the
XSS finding itself and directly touches the same lines.
Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab
PR #544's change, ported onto the api_v3 package split (#553). Identical
behaviour; only the placement of the new code differs.
The original added 508 lines to web_interface/blueprints/api_v3.py, which #553
deletes, so every hunk of it would conflict irreconcilably. Ported by AST:
26 new top-level items sorted to where the split puts each kind --
__init__.py 2 imports, 11 constants, 7 helpers
starlark.py 4 routes (/starlark/editor/{apps,status,start,stop})
misc.py 2 routes (/integrations/mqtt-bridge{,/config})
Everything outside api_v3.py -- the Tools partial, the installer scripts, the
JS tests -- applied unchanged.
Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is
regenerated to match, which is exactly what test_api_v3_url_map.py is designed
to make you do when routes are added.
Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR
ships could not be run here -- node is not installed on this machine -- so
test/js/dom/test_tools_sections.js is unverified.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST
The AST-based port of #544 onto the api_v3 package split dropped `time`
from starlark.py's import list. start_pixlet_editor() and
stop_pixlet_editor() both call time.time()/time.sleep() directly, so
every start (NameError building `state['started_at']`) and every stop
that has to wait out the EXIT trap crashed with a 500. No test caught
it because the route's own tests mock subprocess.Popen but never
actually invoked it before now.
Also carries over #544's later fix that this port branched before:
env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an
operator who had already pinned PIXLET_EDITOR_HOST to loopback,
forcing the unauthenticated `pixlet serve` process onto the LAN
regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as
api_v3.starlark.py's siblings already do for _pkg-owned names.
Both fixes route the shared _pkg.time reference the rest of the
package's route modules already use for anything a test might need to
patch, rather than a bare `import time` local to this file.
Ported the existing regression test from #544
(TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's
module layout (web_interface.blueprints.api_v3.starlark instead of the
old monolithic api_v3 module), which is what caught the NameError.
Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this
branch and on origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to
this change).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(api-v3): clear the six lint errors this rebase introduced
All six were introduced by rebasing this branch onto the merged blueprint
split, not by the split itself. Confirmed by diffing pyflakes output against
main with line numbers normalised -- everything else it reports is present on
main too and is the package's deliberate re-export pattern.
starlark.py used _STARLARK_APPS_DIR three times without importing it (F821).
The rebase resolved an import-list conflict as a union of both sides, and that
symbol was on neither side of the conflict hunk, so it was silently lost. It is
defined in __init__.py and is now imported like its neighbours. This was the
only one of the six that would fail at runtime rather than merely lint.
__init__.py imported contextlib twice (F811): the cherry-pick added one next to
the existing import. Removed the duplicate; the original at line 19 is used.
__init__.py imported signal purely to re-export it to starlark.py, so pyflakes
saw it as unused (F401). signal is stdlib and does not need routing through the
blueprint package, so starlark.py imports it directly and __init__.py no longer
does. contextlib stays re-exported because this module genuinely uses it.
_read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule
this module imports at the bottom for its route side effects (F811). Renamed to
`settings`, with a comment saying why, since the name is otherwise the obvious
one to reach for.
Verified: pyflakes now reports nothing on this branch that main does not, the
package imports, all nine route modules load, and 117 routes register, matching
the pinned URL-map snapshot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply
Two CodeRabbit findings on the bridge settings endpoint, both of which returned
200 while doing something other than what the caller asked.
`request.get_json(silent=True) or {}` turned a missing or unparseable body --
and the JSON literals null, [] and false -- into an empty dict, which then
satisfied the isinstance(data, dict) guard on the very next line. The guard was
there to reject exactly those bodies. Dropping the `or {}` lets None fail it.
The same `or {}` on /errors/clear is left alone: its docstring documents the
body as optional, so an absent body legitimately means "use the defaults". The
difference is that saving settings has nothing sensible to do with no body.
`if data.get('clear_password'):` accepted any truthy value, and the string
"false" is truthy in Python -- so a client echoing the field back as a string
wiped a password it meant to keep. Now coerced through the package's existing
_coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else.
test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus
a missing body, and clear_password across truthy and falsy spellings. Verified
against the unfixed code -- reverting the body guard fails 5, reverting the
coercion fails 3.
Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials
when TLS is off (CWE-319). That is a policy decision about the feature rather
than a defect -- unencrypted MQTT on a trusted LAN is common and often
deliberate -- so it is raised on the PR for a maintainer call instead.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: work through the remaining review findings on the editor and bridge
allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses
the network in cleartext. Refused now rather than merely warned about -- but
refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal
deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults
false, and is coerced like the other booleans so the string "false" cannot
switch the guard off.
starlark.py:796 -- the supported service runs Flask threaded, so two start
requests could each see running=False, each launch an editor, and the second
state write replace the first PID, orphaning a process that holds the display
down with nothing recording it. The check-launch-write sequence now takes a
module-level lock.
starlark.py:848 -- if the state write failed the route returned success with an
editor running and no PID recorded: status and stop both reported no session
while the display stayed down until the timeout expired. It now terminates the
process group and returns an error.
starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so
nothing hands the display back, yet the response said "the display is
restarting". After an escalation the display is now restarted explicitly, and a
failure to do so returns an error naming the manual step instead of a success.
pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no
timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the
failure lands before the display is stopped rather than after.
pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface
address are three cases, not two. Collapsing the last two printed a URL saying
"localhost" whenever PIXLET_EDITOR_HOST named a LAN address.
tools.html:1254 -- escHtml does not encode single quotes, and the app id was
interpolated into an inline onclick="startPixletEditor('...')", so a directory
containing an apostrophe could break out of the JS string and run script. The
handler binds with addEventListener and reads the id from dataset, where it is
only ever parsed as an HTML attribute.
Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in
in both directions. The tools DOM suite gains three guards asserting the edit
buttons carry no inline onclick and pass the id via dataset -- those need jsdom
and did not run here, so CI verifies them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(api-v3): log the traceback on the editor state-write failure
The 848 fix answers 500 when the session state cannot be written, and logged
that at error level -- but without exc_info, so the traceback never reached the
log. test_web_error_detail.py guards exactly this: a handler returning 5xx must
write an error-level record *with* the traceback and return the sanitized
detail, because checking that merely something was logged is too weak.
Caught by Core unit tests on the previous commit, not locally: the guard parses
every module under web_interface/blueprints/api_v3 as one source, so it only
fires once the whole package is read together.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(api-v3): split the 10,469-line blueprint into a package
web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181
functions in one module -- 9% of the core by line count and three times the
next largest file. It becomes a package of nine route modules grouped by path
segment, plus __init__.py for the shared imports, constants, Blueprint and
helpers.
Every route module decorates the SAME api_v3 Blueprint object, so endpoint
names stay api_v3.<function>, the URL map is unchanged and app.py is untouched.
Verified: 111 routes before, 111 after, byte-identical rules, endpoints and
methods, and every endpoint still on the one blueprint.
plugins 3,867 config 1,178 starlark 692 system 619
fonts 452 misc 398 wifi 361 display 326 backup 212
__init__ 1,787 (imports, constants, Blueprint, 56 helpers)
Two things the URL-map check could not catch, both found by running the suite:
1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one
directory deeper made that resolve to web_interface/ instead of the project
root. Nothing failed at import; it surfaced as ~110 tests failing with 404s
and "installation script not found", because every path built from it was
one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts
PROJECT_ROOT/run.py exists so the next move cannot repeat it.
2. Module-attribute patching. Tests do
monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route
module that binds such a name by value never sees the patch. The shared code
therefore stays in __init__.py rather than moving to a _common submodule --
it has to live on the module the tests patch -- and the eleven names tests
patch are read back through the package (_pkg.X) instead of bound by value.
Those eleven were found by AST-scanning every setattr in the test tree, not
by guessing; "time" is among them, used to drive a fake clock through the
second-resolution credential-backup filenames.
Test changes are confined to what genuinely moved: patch targets that now name
the owning route module, imports of helpers, and six tests that scan the api_v3
source as a file and now read the package directory.
Full suite: 4,278 passed, 68 skipped, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(api-v3): address CodeRabbit findings from the blueprint-split review
Fixes to the api_v3 package split (PR #553), one per finding verified
against the actual code:
- __init__.py: _redact_credentials only blanked scalar values under a
credential-named key; a bare list of secrets under such a key (e.g.
tokens: ["a", "b"]) passed through untouched, since the list branch
recursed with no memory that its key looked like a credential. Nested
dicts still walk normally (a documented, tested behaviour -- a container
like secrets: {api_key: ..., note: ...} is a section name, not a value to
blank outright), but any value reached under a credential-shaped key is
now actually blanked.
- __init__.py: the OAuth helper script's raw stderr/stdout went to
logger.error unredacted (CWE-532) right next to a comment claiming this
was deliberate; the HTTP response already used the existing redact_text
helper. Routed the log line through the same helper.
- __init__.py / starlark.py: the standalone Starlark manifest fallback
(used when the plugin instance isn't loaded) read-modified-wrote
manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe
(plugin-repos/starlark-apps/manager.py), which already holds an flock for
the same file when the plugin is loaded. Added _starlark_manifest_lock,
mirroring that pattern, and wrapped every standalone read-modify-write
call site in it. The app-config update route also wrote config.json and
the manifest as two separate, non-transactional writes (a second,
distinct finding at the same call site); config.json is now rolled back
if the manifest write that follows it fails.
- backup.py: restore options used bare bool() on values from the request,
so {"restore_secrets": "false"} restored secrets anyway (bool("false") is
True). Switched to the existing _coerce_to_bool helper already used for
this exact purpose elsewhere in the package.
- config.py: an automated import-rewrite mangled four user-facing
validation strings and their neighbouring comments -- "Invalid start
time" had become "Invalid start _pkg.time" (and likewise for "end time")
in both the schedule and dim-schedule per-day validation paths.
- display.py: `import _pkg.time as time_module` -- _pkg is a local alias
for the package, not a real importable module, so this raised
ModuleNotFoundError whenever a caller restarted an already-running
display service via /display/on-demand/start, after the on-demand
request was already written to cache. Fixed to `import time`. Audited
the rest of the package for the same `_pkg.<module>` import mistake;
every other `_pkg.` reference is a legitimate attribute read-through
(`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken
import statement.
- fonts.py: validate_file_upload's max_size_mb parameter is silently
unused by that helper (it only checks filename/extension) -- the font
upload route saved arbitrarily large files as a result. Added the same
seek-and-check pattern already used for the sibling .star upload.
- wifi.py: two ad hoc, inconsistent bool coercions. POST
/wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled
it. POST /wifi/radio's enabled/force parsing recognized real bool and
some strings but not int 1/0 (1 is True is False in Python). Factored one
small _parse_bool_ish helper local to this file and used it at all three
sites.
Not changed: the "unknown/misspelled restore option keys default to True"
half of the backup.py finding -- the file's own comment documents that a
missing key deliberately means "restore everything," matching the
already-existing JSON-parse-failure guard a few lines above it; only the
bool-coercion defect was a real bug.
Added or extended regression tests for every fix, following each area's
existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed
on both this branch and origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to this
change) -- no new failures.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5
* fix(api-v3): reject unknown restore option keys
CodeRabbit's review of the blueprint split (#553) asked that
POST /backup/restore reject option keys outside RestoreOptions'
known set. The follow-up commit fixed the bool("false")-is-True
bug with _coerce_to_bool but never added the key check: a typo'd
or renamed key (e.g. "restoreSecrets") is silently ignored by
opts_dict.get(key, True), so the flag stays at its True default
and secrets get restored despite the caller's request saying
otherwise -- with no indication anything was wrong.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb
* fix(api-v3): address CodeRabbit findings on the blueprint split
- _redact_credentials: blank scalar descendants of objects reached
through a credential-owned list (e.g. tokens: [{"value": "secret"}])
regardless of field name -- the existing name-based walk only
protected direct dict values under a credential key, not list items.
- wifi.py: reject enabled/force/auto_enable_ap_mode values
_parse_bool_ish can't recognize (400) instead of silently treating
them as False, which could disable Wi-Fi or the radio itself.
- Starlark manifest locking: lock a stable manifest.json.lock sidecar
instead of manifest.json itself, in both the standalone route path
(_starlark_manifest_lock) and the plugin path
(StarlarkAppsPlugin._save_manifest / _update_manifest_safe).
manifest.json is replaced by an atomic rename on every write, which
swaps in a fresh inode; a lock held on the old inode does not
exclude a second locker that opens the path afresh right after the
rename and gets the new inode, so two writers could race despite
each holding "a lock". A sidecar that no write ever touches always
resolves to the same inode for every locker.
Skipped as stale: the "serialize the complete manifest
read-modify-write" finding at api_v3/__init__.py -- every standalone
handler that calls _write_starlark_manifest is already wrapped in
_starlark_manifest_lock() on this branch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(api-v3): re-check reconciliation findings by the reconciler's own rules
Both CodeRabbit findings on the merge commit, verified against the code first.
Major, plugins.py: the stale-findings filter derived its own notion of "in
config" and "on disk", and both were looser than the reconciliation module's.
set(load_config()) also contains system keys, the secrets-file keys load_config()
merges in, and non-dict values; and any directory holding a manifest.json
counted as installed even when that manifest does not parse. Either looseness
clears a finding that is still true -- and a secrets key read as a plugin is the
precise bug the filter exists to stop reporting, so reintroducing that asymmetry
while re-checking was the wrong way round.
The two extractions now live in state_reconciliation.py as config_plugin_ids()
and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys()
alongside. _get_config_state() and _get_disk_state() use them too, so there is
one definition rather than two that can drift. _get_disk_state() re-reads each
manifest for version/name after taking membership from the shared extractor;
that costs one extra small read per plugin on a path that runs once per boot.
Minor, the new test: the fixture assigned api_v3.config_manager and
api_v3.plugin_manager directly. Those live on a module-level blueprint
singleton, so the mocks leaked into every later test that imports api_v3 --
pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr,
which restores them. This is the same pollution class that made an earlier test
in this session break seven unrelated ones, so it is worth getting right.
Five cases added for the parity itself: a secrets key, a system key and a
non-dict value must not clear an "installed but missing from config" finding,
and neither an unparseable manifest nor a .standalone-backup- directory may
count as installed. All five fail against the looser version.
Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and
CodeRabbit all pass. Codacy reads action_required on every commit of this
branch including the first, so it is pre-existing and not from this work.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(web): say when a system action failed for want of passwordless sudo
POSTing reboot_system to a Pi returns, in full:
{"message": "Action failed; see logs for details", "status": "error"}
The cause is that the web interface runs unprivileged, and its
systemctl/reboot/journalctl calls only work once
scripts/install/configure_web_sudo.sh has granted NOPASSWD. first_time_install.sh
never invokes that script and no user-facing doc mentions it, so on a fresh
device every privileged action fails -- start_display, stop_display, the
autostart toggles, reboot, and the log viewer.
That last one closes the loop: "see logs for details" is unreachable advice
when journalctl is refused for the same reason. This is exactly the failure
src/web_interface/error_handler.py's describe_exception() was written to break,
and /system/action's exception handler was still discarding the cause instead
of using the helper the module already imports.
Two changes, no behaviour change when things work:
- The exception path now returns 'details': describe_exception(e), matching how
the other handlers in this blueprint already report.
- A failure whose stderr or exception text is sudo refusing to prompt ("a
password is required", "no tty present", "a terminal is required") reports
what to do about it, naming configure_web_sudo.sh. Unrelated failures keep
the generic message and their stderr, so a missing unit is not blamed on
sudo.
Granting the sudo rights is left alone deliberately: auto-running a script that
hands out NOPASSWD is a security decision for the maintainer, not something to
slip into an installer. Making the refusal legible is the part that is
unambiguously an improvement.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): apply the sudo hint on the on-demand start_display path too
start_display with a mode builds its own response and returns before the shared
nonzero-result path, so a recognized sudo refusal there reported only "Failed to
start display" and said nothing about the passwordless sudo that refused it --
the exact gap the rest of this PR closes everywhere else.
Raised by CodeRabbit on #560 and verified against the code before fixing: the
branch at api_v3.py:2058 does return early past the shared handler.
Three regression cases: the on-demand branch reports the sudo cause, keeps its
"Display started" message on success, and does not blame an unrelated failure on
sudo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
On a device running four installed, configured, working plugins, the overview
banner read:
Stale plugin config entries found: football-scoreboard, odds-ticker, data,
ledmatrix-weather, starlark-apps. Remove them from config.json or reinstall
via the Plugin Store.
Every claim in that sentence was wrong, and following its advice would have
deleted 4.9KB of working league settings. Four separate defects combined.
1. Secrets keys became phantom plugins. load_config() merges
config_secrets.json into the config it returns, and the ignore list named
only 'github' and 'youtube'. A 'data' key in that file therefore read as a
plugin id and was reported as "in config but not on disk" forever. Read the
secrets file's own top-level keys instead of hardcoding two of them.
2. The auto-fix clobbered real config. The handler for "on disk but not in
config" assigned `config[plugin_id] = {'enabled': False}` unconditionally,
so whenever detection was wrong it replaced a plugin's entire configuration
with a stub. On the reported device it only failed to do so because the
write hit EACCES. Now it refuses to overwrite an entry that already exists.
3. The banner gave backwards advice. plugin_missing_in_config ("on disk, not in
config") and plugin_missing_on_disk ("in config, not on disk") are opposite
problems, and both were rendered as "stale config entries ... remove them
from config.json" -- which is correct for the second and destructive for the
first. They are now reported separately, each with the advice that fits.
4. A stale verdict was served indefinitely. The result is a snapshot written
once per run to a status file, and a run that fails to apply a fix also
declares it will not retry. A condition that had since resolved kept being
reported for hours. The status endpoint now re-checks stored findings
against current state, dropping only what it can prove stale and keeping
any kind it cannot re-verify.
The secrets-key lookup is deliberately fail-safe: an unreadable, absent,
malformed or non-path secrets location narrows the ignore set rather than
raising. An earlier revision let TypeError escape, which the broad handler in
_get_config_state() swallowed as "Error reading config state" -- emptying the
config state and making every downstream detection wrong. The existing
reconciliation tests caught it; there is now a regression test for it too.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
src/base_classes/baseball.py _get_baseball_display_text 45 lines
src/web_interface/api_helpers.py validate_request_params 22
web_interface/blueprints/api_v3.py _validate_time_range 14
Each has exactly one occurrence across both repositories -- its own
definition. No decorator, no __all__, no getattr dispatch, nothing in
templates or JavaScript.
A fourth candidate was dropped after checking: _unshare_element_fonts in
src/common/sports_shared.py looked unreferenced, but eight scoreboard plugins
call SportsCore._unshare_element_fonts directly from their
test_element_text_colors.py, plus their own copies at runtime. It is live API.
The earlier reading came from a plugins checkout 84 commits behind main, which
is a good argument for re-verifying this kind of claim against a fresh tree
rather than trusting an earlier scan.
Full suite: 4,265 passed, 68 skipped.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
display_controller resolves once, and caches, whether a plugin's display()
takes a display_mode keyword -- self._plugin_accepts_display_mode, populated
right before the dispatch. It then handed the executor a
types.SimpleNamespace wrapping a closure, and execute_display() ran
inspect.signature() on that to work out the same thing.
Because the SimpleNamespace is rebuilt per call, the callable was new every
time, so nothing inside the executor could ever cache it either. Measured at
~39us per dispatch on a Pi 4, for a value the caller had a line earlier.
execute_display() now takes accepts_display_mode, falling back to inspecting
only when a caller does not pass it, so existing callers are unaffected.
Also documents two things that read as bugs and are not:
- execute_with_timeout()'s timeout is advisory. Nothing cancels the thread --
Python cannot -- so on expiry the operation runs to completion in the
background and only the caller gives up. A permanently hung plugin leaks a
daemon thread per attempt. This is why callers holding a lock across the
call must release it from inside the wrapped callable, as run()'s
_release_display_lock already does.
- Only the first display() of each mode goes through the executor; the
per-frame loops call display() directly. That is deliberate: a thread per
frame would cost more than an advisory timeout buys. Both loops now say so,
so the asymmetry does not read as an oversight.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The SSE status stream sent 'disk_used_percent': 0 as a literal, so every
consumer of the live view showed 0% disk no matter how full the card was.
/api/v3/system/status computed it correctly; the stream that the dashboard
actually watches did not. On a Pi with a modest SD card that is the warning a
user most needs, and it was guaranteed to never appear.
The stream also omitted memory_available_mb. /api/v3/system/status carries it
with a comment spelling out why it matters: MemAvailable accounts for
reclaimable page cache, so it is what separates a board reading 70% "used" that
is fine from one reading 70% that is about to fail fork(). A 1GB Pi 3B+ can sit
at either. The number that predicts the failure was missing from the live view.
An unreadable disk now reports None rather than 0. The UI already renders null
as '--'; a confident 0 reads as "plenty of room", which is worse than a blank.
Metric collection moves to web_interface/system_metrics.py, with no Flask or app
imports. That is not cosmetic: importing web_interface.app constructs the Flask
application and a CacheManager, and the latter claims the cache directory with a
cleanup thread. The first version of these tests imported the generator directly
and broke test_cache_cleanup_thread_ownership ("one thread per directory") plus
four starlark route tests through that side effect. Reading a CPU percentage
should not boot a web application, and testing it should not either.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
SportsCore._logo_cache was a plain dict keyed by team abbreviation with no
eviction. Its entries are not file bytes but decoded RGBA thumbnails sized to
display*1.5 -- roughly 36KB on a 256x64 panel, more for wide wordmarks -- and
assets/sports/ncaa_logos ships 307 of them. A plugin that walked a full league
held the whole league resident: about 11-18MB per manager instance, and a league
runs three (live/recent/upcoming) that each keep their own cache, so the same
logos were duplicated across them.
On the 1GB Pi 3B+ this was measured on, one board was sitting at 439MB resident
with ~290MB available, so tens of megabytes of duplicated league logos is real
money. Bounded to 64 entries, which holds a full "other games" cycle (on the
order of 20 games, 40 teams) without thrashing while capping the cache well
below a 307-team league.
Eviction is LRU rather than clear-when-full, using the OrderedDict/popitem
pattern the neighbouring caches in this codebase already use (_IMAGE_CACHE_MAX,
_FIT_CACHE_MAX, _TEXT_WIDTH_CACHE_MAX). That ordering matters: the logos on
screen right now are precisely the ones that must not be discarded, so a cache
hit moves the entry to the end.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(plugins): let a plugin ask to be polled faster while it has live content
Reported: "the football plugin with live games only updates the live game in
progress if I restart the display."
The data path was never the problem. NFLLiveManager fetches ESPN with no cache,
SportsLive.update() refreshes current_game in place when the game IDs are
unchanged, and the scorebug redraws from the game dict every frame -- which is
why the reporter's logs look healthy.
The problem is cadence. _get_plugin_update_interval() read only the manifest's
static update_interval, football's manifest pins that to 60, and the plugin's
own live_update_interval (15s) was invisible to the scheduler. Measured on a rig
during the fourth quarter of the game in the report:
23:21:49 23:22:50 23:23:50 23:24:50 23:25:50 <- exactly 60s apart
A clock and score up to a minute stale during a two-minute drill reads as a
frozen panel, and a restart is the one moment it is ever current.
A single static number cannot say "every 15 seconds while a game is on, every 15
minutes in July", and only the plugin knows which is true. get_update_interval()
lets it say so per tick; returning None means "no opinion" and the existing
manifest/config resolution applies, so every plugin that predates this is
unaffected.
Requests are clamped to MIN_DYNAMIC_UPDATE_INTERVAL (5s): a plugin returning 0
would otherwise be re-entered on every tick of the render loop, busy-waiting
against its own API. A hook that raises or returns a non-number is ignored
rather than propagated -- a scheduler that fails on one plugin's bug stops
updating all the others.
Deliberately NOT changed: the manifest still beats config in the static path.
That looked like the obvious fix -- user config being silently ignored -- until
checking a real rig, where football and baseball both carry update_interval 3600
in config against a manifest 60, and weather 1800 against 60. Those values are
stale precisely because nothing has been honouring them; making config win would
have slowed three plugins by 60x, turning a one-minute lag into an hour. The
dynamic hook makes the flip unnecessary. There is a test pinning the current
precedence with that reasoning attached.
Full suite: 4,283 passed, 68 skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* test(plugins): drive the real scheduler, not just the interval resolver
test_plugin_dynamic_update_interval.py asserts that
_get_plugin_update_interval() returns the number the plugin asked for. That is
not the same claim as "the plugin gets updated more often", and the gap between
those two is exactly where the original bug lived: the plugin knew it wanted
15s, said so in live_update_interval, and nothing downstream acted on it.
So this ticks the real run_scheduled_updates() through a simulated hour and
counts dispatches. Against pre-fix core it reports "10 updates in 10 minutes of
a live game" -- the 60s manifest cadence, matching what was measured on a rig
during the reported game. Against the fix it reports ~40.
Also pins the regression that would be worse than the bug: an idle hour must
still be ~60 updates, not 240. Asking for the live interval year-round would
poll ESPN four times a minute all summer.
Scope note, since it is easy to over-read this fix: the *switch* display path
already refreshed the manager immediately before drawing, via
_try_manager_display() -> _ensure_manager_updated(), which honours the manager's
own 15s interval. So a switch-mode card was already <=15s stale at draw time
before this change. What this fixes is the background cadence, which is what
live-priority detection, Vegas content and scroll preparation all read.
Full suite: 4,288 passed, 68 skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(plugins): reject bool and -inf hook results in dynamic interval
get_update_interval() ran bool through float() (bool is an int subclass,
so True/False became 1.0/0.0) and only checked for +inf, not -inf. Both
cases landed on the MIN_DYNAMIC_UPDATE_INTERVAL floor by coincidence
instead of falling back to the static/manifest interval as invalid
input should.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Tst9cied2ri9bH4QRWa6H
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
logo_downloader.py and LogoHelper write into assets/sports/<league>_logos/
whenever a plugin meets a team whose logo is not on disk. Those directories are
also tracked -- 209 NCAA logos and 153 soccer ones ship with the repo -- so
every rig accumulates untracked files nobody intended to commit. This checkout
had 62; hdpi shows the same.
The cost is not the files, it is that a permanently dirty `git status` trains
everyone to ignore the one signal that says a checkout is not what you think it
is. That is how a stale tree sat unnoticed on a rig for hours until a restart
surfaced four sports plugins that could no longer import.
Ignoring a directory does not untrack what is already in it, so the logos that
ship keep shipping -- verified: 209 and 153 still tracked, no deletions in the
diff. Only new downloads are hidden.
Adding a logo on purpose stays possible and is what the escape hatch in the
comment documents. It is also rare: the last deliberate addition was #415, four
named NCAA logos a plugin needed, and `git log` finds no other in a year. So the
common case is noise and the rare case is explicit, which is the right way round.
Untracked files: 62 -> 0.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(logo): remember a missing logo instead of re-warning every rotation
load_logo() stat'd the path and logged a WARNING on every call, and the
positive cache never covered it because a miss returns None and caches
nothing. A file that is simply not there therefore produced one warning per
rotation for as long as the process ran -- measured on a live rig at 114 lines
in 24 hours for a single missing ticker icon, for a file nobody was going to
add.
Misses are now remembered for 10 minutes: warn once, then return None without
touching the disk. Bounded rather than permanent because logo_downloader
writes logos at runtime, so a file that appears later must still be picked up
without a restart. Downloads through load_logo_with_download() clear the entry
outright -- load_logo() consults the miss record before it stats the disk, so
without that a freshly downloaded logo would stay invisible for the whole
window.
This is in the core rather than in ledmatrix-stocks, where it was found, so
every plugin that goes through LogoHelper gets it.
_cache_order stays a list. Swapping the pair for an OrderedDict would shave an
O(n) scan per cache hit, but n is capped at cache_size (100 by default) and
test_logo_helper.py pins the current structure; not worth the churn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(logo): make the miss TTL longer than the rotation it is meant to outlast
Deployed the previous commit to a live rig and measured it: no change at all.
"Logo not found for VOO" stayed at ~6 lines an hour, exactly the baseline.
The TTL was 600s and the display rotation is ~618s, so every recheck expired
just as the plugin came round again and the negative cache never once got to
suppress a warning. The fix was correct in shape and useless in practice,
which only measuring on the rig would show.
An hour instead. That is safe because the TTL is not the main way an entry
clears: load_logo_with_download() drops it the moment a download succeeds and
clear_cache() drops all of them. The TTL only covers a file that appeared some
other way -- someone copying one in by hand -- and waiting up to an hour for
that, or restarting, is a fair trade for not re-warning about a file nobody is
going to add.
The general lesson is in the comment: a TTL has to be long relative to the loop
that does the asking, not merely "a while".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(install): render the systemd units from their templates, not from heredocs
The installers carried their own inline copies of units that also exist as
templates under systemd/, and the copies drifted.
install_service.sh renders ledmatrix.service from the template correctly, then
wrote ledmatrix-web.service from a heredoc that predated it -- missing
Wants=network-online.target, RestartSec=10, SyslogIdentifier, CacheDirectory,
CacheDirectoryMode and Environment=USE_THREADING=1. install_web_service.sh had
a third copy, and install_wifi_monitor.sh a fourth, that one already differing
from its template (syslog where the template says journal).
startup_validator.py compares the installed unit against the template, so a
rig installed this way warned on every boot -- and the remedy the warning
names, "re-run scripts/install/install_service.sh", reinstalled the same stale
copy. The warning could never clear. Reproduced on a live rig running exactly
that unit.
All three installers now render systemd/*.service through the same placeholder
substitution. The template gains a __USER__ placeholder rather than hardcoding
User=root, because the web interface runs as whoever installed it.
That last point was a second, independent cause of a permanent warning: the
validator substituted a fixed "root", so any non-root install reported drift
forever. It now reads User= from the installed unit -- an install-time
decision, not something the template dictates -- and compares everything else
strictly. first_time_install.sh already reads the installed User= the same way.
Tests cover a non-root web unit not warning, a genuinely changed directive in
that unit still warning, the User= fallback, and a grep-based guard that no
installer under scripts/install/ contains an inline unit body. That guard is
what found the install_wifi_monitor.sh copy.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(install): escape sed replacements, use mktemp, and make render failures fatal
Address CodeRabbit findings on install_service.sh, install_web_service.sh and
install_wifi_monitor.sh:
- Values interpolated into each script's sed expression (project root path,
username) were not escaped, so a value containing &, \ or the | delimiter
would corrupt the rendered systemd unit. Add a shared
sed_escape_replacement() helper in the new scripts/install/lib_systemd_render.sh
(sourced by all three scripts) and apply it to every sed replacement.
- install_service.sh rendered the main and web units to the predictable path
/tmp/ledmatrix.service.tmp before installing them -- a symlink/TOCTOU race
(CWE-377). Use mktemp for both, with a trap to clean up on exit.
- install_service.sh treated a missing template as a mere warning and then
checked only whether a unit already existed at the destination before
enabling/starting it, so a render failure could silently fall back to
enabling a stale, previously-installed unit. Both unit blocks now exit
non-zero on a missing template or a failed render.
Also rename the ambiguous loop variable `l` to `line` in
test/test_systemd_unit_drift.py (Ruff E741); ruff isn't wired into any CI
workflow in this repo today, so this isn't currently CI-blocking, but the
rename is trivial and correct regardless.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5
* test(install): cover sed_escape_replacement against sed-special characters
CodeRabbit asked for regression coverage using a project path containing an
ampersand; the earlier commits on this branch already fixed the escaping,
mktemp usage, and enable/start-on-fatal-render-failure findings, and the
l->line rename was already applied -- this closes the one remaining gap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
start_web_conditionally.py read the flag with
`config_data.get("web_display_autostart", False)`, so a config that simply
lacked the key got no web interface. Both config/config.template.json and
first_time_install.sh ship the key as true, so the code default contradicted
the shipped default in two places: absence means an older or hand-edited
config, not a request to stay down.
The failure mode was silent in the worst way. The "not starting" path exits 0,
so `systemctl status ledmatrix-web` reported the unit as successfully started
while nothing was listening on the port, and the only trace was one journal
line saying the flag was "false or not set" -- which reads as a deliberate
setting rather than a missing key.
Also start the web interface when config.json is missing or unparseable,
instead of exiting. The web interface is how a config gets created and
repaired, so a broken config is exactly when the user needs it most; leaving
it down means there is no way back in. Only an explicit false/off disables
autostart now, and the disabled message says "explicitly disabled" so the
journal distinguishes a real setting from a default.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(scroll): advance whole pixels per frame, not per wall-clock second
Smooth motion is not a frame-rate property, and measuring it as one is why
this survived three rounds of fixes. odds-ticker's frame timing is excellent
-- 100.0 fps, 10.00ms median, 0% stalls, worst in-scroll frame 19.95ms -- and
it still visibly stuttered.
What the eye judges is whether the strip advances the same number of whole
pixels on every presented frame. update_scroll_position derived position from
scroll_speed * delta_time and get_visible_portion truncated it with int(), so
jitter in delta_time decided which side of a pixel boundary the position
landed on. The live windows show why that matters: a rock-steady 100.0 fps
whose individual frames still range 5.6ms to 15.2ms, which at 100 px/s is
0.57px to 1.44px of movement.
Run the measured frame times through the real helper and 5.8% of frames
advance 0 or 2 pixels instead of 1 -- about six hitches a second. A frame that
moves nothing followed by one that jumps two is exactly what micro-stutter
looks like.
It is worst at a crisp speed, which is the part that stings: at 100 px/s on a
100Hz panel the accumulator sits exactly on integer boundaries, so
sub-millisecond jitter flips it either way and the motion beats at around
50Hz. Snapping to the crisp ladder fixes the average and the wall clock then
throws away the per-frame uniformity the ladder was bought for.
So when scroll_config snaps to a crisp speed it now also puts the helper in
fixed-step mode: each presented frame advances exactly pixels_per_frame and no
clock is consulted. 100% of frames move by the same amount, whatever the
jitter.
This is only correct because SwapOnVSync blocks until the panel has taken the
frame, which makes the frame count a truer clock than time.time(). Before the
swap was locked to vsync it would have run at whatever speed the loop spun at.
Related: frame-based mode used to step discretely and was converted to
elapsed-time accumulation earlier in this series, because its threshold
comparison flipped on jitter. That was right for the code as it stood -- but
it treated the symptom, replacing a broken discrete step with a smooth-looking
accumulator instead of asking why a wall clock was involved at all.
Non-crisp speeds keep pacing off time, and set_scroll_speed() clears the fixed
step so a legacy caller changing speed is not silently ignored.
Trade-off worth naming: speed is now tied to the presentation rate rather than
to real time. If the loop cannot keep up with the panel the scroll runs slow
rather than jumping to catch up. That is the better failure -- uniform motion
at a slightly wrong speed beats correct average speed with a hitch six times a
second -- and a loop that cannot hit the resolved rate is a measurement
problem for the crisp ladder, not something to paper over with uneven steps.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(scroll): make the time-based pin actually pin something
Review caught that test_time_based_stepping_is_what_it_replaces could pass
against perfectly uniform motion, and it was right.
update_scroll_position sets last_update_time on its way through, so the very
first call sees a delta_time of zero and moves nothing in time-based mode.
_advances counted that synthetic frame, which put a guaranteed zero in every
histogram -- enough on its own to satisfy "uneven > 0". The test asserting the
defect exists would have passed after the defect was gone.
The first call is now primed and discarded, and the assertion is a proportion
rather than "more than zero": against these frame times the old path misses
roughly one frame in twenty, so 1% is well below the real rate and far above
anything a stray frame could produce.
Re-measured with the artefact removed, the numbers in the PR description are
unchanged: 5.85% of frames uneven before (114 zero-advance and 120 double
frames in 4000), 0.00% after.
Also fills in the docstrings the review flagged: everything in the new test
file, plus three pre-existing one-liners in scroll_config that the diff
touched.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Tests has been red on main since #535. All 13 failures in
test/web_interface/test_starlark_pixlet_routes.py are the same
ModuleNotFoundError: No module named 'yaml'.
The test loads plugin-repos/starlark-apps/tronbyte_repository.py by path --
deliberately, "the way the blueprint does", since the core web blueprint
really does exec that plugin module -- and the plugin imports yaml.
Nothing is undeclared. The plugin's own requirements.txt already pins
PyYAML>=6.0.2, and on a real rig the plugin store installs it. CI installs
only requirements.txt and requirements-test.txt, so a core test that reaches
into a plugin gets none of the plugin's dependencies.
PyYAML goes in the test requirements rather than the core ones because it is
not a core dependency: nothing in src/ or web_interface/ imports yaml. This is
the same shape as the psutil entry directly above it -- a package the core does
not require, installed so a test can exercise a real path instead of a stub.
Verified locally: with yaml available the file goes from 13 failures to 64
passing. (One unrelated failure remains on Windows only, where os.geteuid does
not exist.)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* refactor(sports): put the scoreboards on the shared scroll resolver
Eight sports scoreboards -- afl, baseball, basketball, football, hockey,
lacrosse, nrl, soccer -- scrolled through this module's own pacing while the
other eleven scrolling plugins went through src/common/scroll_config. Two
implementations of the same job, and this one was on the losing side of every
difference.
It never called set_scrolling_state. Two consequences, both of which this
release's work was about:
- The frame hold is applied through that call, so a speed the crisp ladder
could render in whole pixels still presented a new frame every refresh.
- Core only runs deferred updates while nothing is scrolling. Believing
nothing was, it ran blocking work in the middle of these scrolls.
The default is non-crisp today: scroll_speed 50.0 with scroll_delay 0.01 is
50 px/s, which on a 100Hz panel is half a pixel per refresh. That cannot
render as motion -- it alternates 0px and 1px steps and judders at a 50Hz
beat, on every scoreboard, out of the box. Resolved through the ladder it
stays 50 px/s and holds each frame for two refreshes: same speed, whole-pixel
motion.
The stepping disagreement that used to justify a separate module is gone.
scroll_config avoided frame-based mode because it stepped on a wall clock at
1/scroll_delay with scroll_delay set to the frame period, so the decision sat
on its own threshold and flipped on sub-millisecond jitter. That branch now
accumulates elapsed time, identical arithmetic to the time-based one, so the
two differ only in the units the speed arrives in.
What is NOT shared, and must not be: the two modules read identically-named
keys with different meanings. Here scroll_speed is px/SECOND and scroll_delay
only converts to px/frame; in scroll_config scroll_speed is px per STEP, so
px/s is speed/delay. Passing this module's settings dict to the resolver turns
50 px/s into 5000, clamped to 500 -- a tenfold speed-up everywhere. So
_get_scroll_settings keeps sole ownership of reading sports config, including
the league merging, and hands the resolver a plain px/s. A test pins that
specific number, because it is the mistake the refactor invites.
MIN/MAX_PIXELS_PER_FRAME are gone; the resolver bounds speed and the helper
clamps FPS. _resolve_target_fps stays, re-purposed: under the old model that
key was the rate frames were presented at, so it is the faithful translation
into the refresh the ladder is computed against, used when no hardware
refresh is configured.
Speed changes for panels that are not 100Hz: 50 px/s becomes 60 at 60Hz
(+20%) and 48 at 120Hz (-4%). At 100Hz it is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display): drop the frame hold when a scroll times out, not just when it says so
set_scrolling_state(False) clears the hold. The other way a scroll ends is
is_currently_scrolling() deciding, after scroll_inactivity_threshold of
silence, that it is over -- which is what happens when the rotation moves on
mid-scroll or a plugin is torn down. That path cleared the flag and kept the
hold, so every later plugin, scrolling or static, was presented at refresh/N
by whoever scrolled last, until something called the explicit stop.
The method's own docstring already states the rule this breaks: the hold "must
not outlive the scroll that asked for it". The timeout was the exception it
did not cover.
Pre-existing, but reachable by three plugins before and eleven after the
sports scoreboards moved onto the shared resolver, so it belongs with that
change. The test ages the activity timestamp past the threshold rather than
sleeping.
Also adds scripts/sports_scroll_check.py. The sports scroll path is per-league
opt-in, so a rig showing static game cards never constructs a
SportsScrollDisplay and none of its pacing can be observed from a normal run
-- which is exactly what happened when this change was first put on hardware:
26 minutes, zero sports scroll lines. The script drives the path directly with
synthetic games and asserts the three things the resolver is meant to buy: the
speed lands on whole pixels, the hold is published, and it is released after.
It never starts or stops the display service, matching scroll_speeds.py, so a
crash here cannot leave the panel dark.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scripts): refuse to grab the panel while the display service has it
The module docstring already said to stop ledmatrix first. Nothing enforced
it, and running the script against a live service is not a harmless mistake:
rpi-rgb-led-matrix configures GPIO directions and the hardware PWM inside
RGBMatrix(), and when the root check fails it calls exit() from C with no
cleanup. The service keeps rendering and swapping onto pins that have been
reconfigured underneath it, so the panel goes black while every diagnostic
says the display is healthy -- fresh framebuffer, every pixel lit, "RGB Matrix
initialized successfully", nothing in the log. A restart fixes it, once you
work out that is what happened.
Found the hard way: this is what took the panel down on the test rig, not the
change the script was written to verify.
--fallback skips the check, since it never opens the matrix. --force is there
for anyone who means it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(scripts): annotate the subprocess call the way this repo already does
Codacy fails a PR on one new issue, and bandit B404 fires on any subprocess
import. scripts/run_plugin_tests.py carries the same suppression with the same
justification -- list-form argv, no shell -- so this follows it rather than
inventing a second convention.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(plugins): add search, filter and sort to Installed Plugins, on a shared helper
The Installed Plugins grid had no way to narrow it down: no search, no way to
see only what's enabled, disabled, or out of date. On a rig with a couple dozen
plugins that means scrolling the whole grid to find one.
The two sections below it already solved this, twice, independently — the
Plugin Store and Starlark Apps carried a copy-paste fork of the same ~600 lines
(filter state, apply-filters-and-sort, page renderer, pagination strip,
active-filter badge, listener wiring). Rather than add a third copy, this
extracts the shared machinery and builds the new toolbar on it.
New: web_interface/static/v3/js/plugins/list_filter.js — ListFilter.create()
owns debounced search, filter axes, sort, the active-filter count, Clear, and
optional pagination/persistence. Callers keep their own card markup via a
`render` callback. Three control types cover every axis the page uses: pills
(new), select (store category, starlark author) and cycle (the tri-state
All -> Installed -> Not Installed button).
Installed Plugins gets a compact toolbar: search box, one-click All / Enabled /
Disabled / Updates pills, and a sort dropdown (A-Z, Z-A, updates first,
recently updated, category). Filters reset on load, so you never come back to a
mysteriously short list. No new CSS — this is the first consumer of the
.filter-pill rules already sitting unused in app.css.
renderInstalledPlugins() is split so it still publishes canonical state while
renderInstalledCards() draws only the visible subset; the filtered list is
never assigned to window.installedPlugins, which the toggle handler,
isStorePluginInstalled(), runUpdateAllPlugins() and the Alpine config tabs all
read as their source of truth. Toggling a plugin while filtered pins its card
so it doesn't vanish from under the cursor.
The Store and Starlark migrations are behaviour-preserving: same element ids,
same localStorage keys (storeSort/storePerPage, starlarkSort/starlarkPerPage),
same tri-state button markup, same pagination. Verified by differential tests
that run the old and new implementations side by side against identical
fixtures and compare every observable after each interaction. The only visible
change is the pagination attribute (data-store-page/data-starlark-page ->
data-list-page), which nothing outside its own click handler referenced.
Net -156 lines in plugins_manager.js while adding a feature.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(plugins): keep raw search text, and stop the store search refetching
Two review findings from CodeRabbit on #540.
Do not write the trimmed search value back into the input. setSearch() trimmed
before storing, and syncControls() then copied that trimmed value back over what
the user had typed. Pausing longer than the debounce after typing a space
deleted the space (and reset the caret), making multi-word terms effectively
untypable. The raw text is now kept alongside the trimmed one: filtering and
activeCount() still use the trimmed value, while the input keeps exactly what
was typed.
Remove the legacy #plugin-search / #plugin-category listeners in
initializePlugins(). They bound searchPluginStore as the event handler, so the
DOM event arrived as its `fetchCommitInfo` argument — always truthy, which
skipped the cached-filter fast path and refetched /api/v3/plugins/store/list
with commit info on every keystroke burst and category change. The store's
ListFilter controller already filters the cached list, which is what those two
controls should do. This double-binding predates this PR (the old code guarded
with _listenerSetup and _storeFilterInit, two different flags, so both sets
stayed live); it is fixed here because the refactor owns that wiring now.
Both fixes are covered by tests that fail without them: the trailing-space
regressions in the installed-plugins DOM suite, and a new whole-file jsdom test
that counts fetches while typing (1 request at init, 0 thereafter; previously
1 -> 2 -> 3 -> 5).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(plugins): build pagination via DOM APIs, drop computed member access
Addresses the five Codacy security findings, all in list_filter.js.
Pagination no longer assembles an HTML string (3 findings: 2 critical + 1 high,
"unsafe assignment to innerHTML"). The interpolated values were only page
integers and local class constants, so there was no injection path, but
concatenating markup into innerHTML is the pattern the scanners flag and
createElement is no less clear. Each button now also owns its click listener
directly instead of the container being re-queried afterwards, and the strip is
cleared with textContent = '' rather than by assigning empty markup. No
innerHTML assignment remains in the file.
haystack() now walks Object.entries(item) and keeps the configured fields,
instead of reading item[field] per field ("generic object injection sink").
Field order no longer drives the haystack order, which is irrelevant to the
substring test. matches() iterates controls with for...of instead of an index
("variable assigned to object injection sink").
The rendered pagination is unchanged: same buttons, labels, page numbers,
disabled states and classes. The old-vs-new differential tests now compare
pagination structurally (tag, text, page, disabled, sorted class list) rather
than as an HTML string, since building nodes legitimately serialises
differently — «/» as characters rather than «/», disabled="" rather
than a bare attribute. That comparison is stronger than the string one it
replaces, and the real-DOM suite still drives the actual page buttons.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(plugins): keep configured field order when building the search haystack
The previous commit swapped item[field] for Object.entries(item) to clear a
static-analysis object-injection warning, and in doing so changed the order of
the haystack: entries follow the object's own key insertion order, not the
configured `fields` order. Since the values are concatenated, that order decides
which values end up adjacent, so a multi-word query spanning a field boundary
matched differently. For store fields [name, description, author, id, ...] and
API objects keyed {id, name, description, author, ...}, "bob plugin-01" matched
before and stopped matching after.
That contradicted the behaviour-preservation claim for the store and starlark
migrations, and the differential tests missed it because every fixture query was
a single word.
Values now come out of a Map built from Object.entries, iterated in `fields`
order: the original haystack is restored, and there is still no computed member
access for the analyser to flag.
Regression coverage for the ordering itself, at both levels:
- unit: phrases spanning name->id and category->tags, plus the reverse
(object-key) order asserted NOT to match
- differential: the same class of query compared old-vs-new, with a guard that
the phrase actually matches something so a mutual zero-result cannot pass
vacuously
Verified both fail without this fix (3 unit, 2 differential) and pass with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* test(web): add JS suites for ListFilter and the plugin-manager grids
No JS toolchain exists in this repo, so these are plain node scripts with no
framework: each prints ok/FAIL lines and exits non-zero. `node test/js/run_all.js`
runs everything, skipping the DOM suites (rather than failing) when jsdom is
absent or nothing is listening, so it stays useful in a bare checkout.
unit/test_list_filter.js ListFilter search/filter/sort/count/sticky, driven
through the installed-plugins config eval'd
verbatim out of plugins_manager.js so the test
cannot drift from the real configuration
unit/test_render_cards.js renderInstalledCards markup, both empty states,
and escaping of hostile plugin metadata
dom/test_installed_dom.js the toolbar in a real DOM, including the HTMX
partial re-swap and a getComputedStyle check that
.filter-pill[data-active] matches what we emit
dom/test_store_dom.js store pagination, per-page, category, tri-state
Installed button, persistence across a re-boot
dom/test_no_double_fetch.js loads the whole plugins_manager.js and counts
requests, so a keystroke cannot refetch the store
The DOM suites deliberately fetch the partial and the plugin data from a running
web interface instead of using fixtures, so a renamed element id or a changed
payload shape fails them loudly. Point them at a rig with a full plugin set when
it matters (BASE=http://host:5000); a dev box with two plugins installed passes
while exercising very little.
Several assertions exist to stop specific bugs recurring: trailing spaces
surviving the search debounce, a query spanning two adjacent search fields
(haystack field order is load-bearing), and window.installedPlugins staying at
full length while the grid is filtered. Others guard against passing vacuously —
counting only non-skeleton cards, and checking a search phrase matches something
before comparing two result sets.
The old-vs-new differential suites that verified the store and starlark
migrations are not included: they compared against the pre-refactor code, which
now exists only in git history.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
#535 restored the thirteen routes, so the store stopped answering 404 --
and still would not load. Confirmed against a running device before
anything was changed: /repository/browse answers 200 with 1000 apps in
27s, so the routes are fine. Two things underneath them are not.
**The store never used the token the user configured.** The three
repository routes read `github_token` off config.json. Nothing writes
that key -- it is not in config.template.json, no setting offers it, and
it appears nowhere else in the codebase. The configured token goes to
config_secrets.json as `github.api_token`, which PluginStoreManager
loads and every other GitHub caller uses. So the store could never be
authenticated: 60 requests/hour, on the same per-IP budget 48 installed
plugins spend on update checks, while the 5000 the user had already
configured sat unused. On the device, /plugins/store/github-status
reported authenticated with a limit of 5000 at the same moment
/starlark/repository/browse reported 60, with 18 left. The store going
blank was that 60 running out.
**Every failure looked identical.** list_all_apps_cached turned any
listing failure -- rate limit, DNS, timeout, non-200 -- into an empty
app list, and the route sent that out as `status: success`, so a rate
limit and an empty repository drew the same blank grid with no error
anywhere. It now returns the reason, the route answers 502 with it, and
a failure is no longer cached as an empty repository for two hours.
The guard for a bad response was itself a crash: _make_request catches
`(json.JSONDecodeError, ValueError)` but `json` was never imported, so
evaluating the tuple raises NameError and the guard written for exactly
this case never ran. Reachable whenever something on the path answers
with HTML -- a captive portal, a proxy page, a DNS-hijacking router.
Seventeen handlers answered 5xx with no detail at all.
test_no_api_v3_handler_discards_its_exception is meant to prevent that
across api_v3, but it matched one exact message string, and all thirteen
Starlark routes wrote their own wording. The guard now keys on the shape
that matters: if it returns 5xx, it says why. The 15 pre-existing
non-Starlark functions are listed as a set that may shrink, never grow.
**The listing was capped at 1000 and did not say so.** The contents API
truncates a directory silently; tronbyt/apps has 1075 app directories,
so the store showed a truncated repository and looked complete doing it.
Now listed via the git trees API, which reports `truncated`, with the
contents API kept as a fallback.
Not addressed: the 27-second cold load -- 1075 manifests fetched five at
a time behind skeleton placeholders -- which is probably the largest part
of what "does not load" feels like, and wants its own change.
25 new tests.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(display): pin one text layout engine, and give the 5x7 face a size
Two ways a font could render differently on two machines running the same
code, both found while diagnosing four plugins whose golden images passed on
the machine that generated them and failed everywhere else.
**Layout engine.** `ImageFont.truetype` picks its engine at load time: Raqm
where the host Pillow was built with libraqm, Basic otherwise. The two round
fractional glyph advances differently. `PressStart2P-Regular.ttf` at 8px has
whole-pixel advances, so they agree — which is why most of the fleet matched
everywhere and hid this. `4x6-font.ttf` at 6px does not: glyph positions drift
cumulatively along a run, and the four plugins that draw body text in it
(geochron, of-the-day, christmas-countdown, ledmatrix-weather's almanac) are
exactly the four whose goldens travelled badly.
Every core font load now goes through `src/common/font_layout.load_truetype`,
which pins the Basic engine, so a render depends on the font file and the size
and nothing else. Basic gives up complex-script shaping and kerning pairs;
neither applies to bitmap-grid faces on an LED panel. Output is unchanged on a
host without libraqm.
**Zero font height.** `DisplayManager` built the 5x7 BDF face with
`freetype.Face(path)` and never called `set_char_size`, so `face.size.height`
stayed 0 and `get_font_height()` returned 0 for it — callers stacking rows by
`prev_y + prev_height + gap` drew two lines on top of each other. The
start-up line `Calendar font size: 0 pixels` has been printing the symptom all
along. `font_manager._load_bdf_font` already called `set_char_size`, so
whether measurement worked depended on which path loaded the face.
`DisplayManager` now sets it too, and `get_font_height()` falls back to the
strike the file declares rather than returning a zero line height.
Fixes ChuckBuilds/ledmatrix-plugins#397
Refs ChuckBuilds/ledmatrix-plugins#371, #375, #378, #391
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display): give the startup banner a rung that fits a full address at 64px
CI caught what pinning the layout engine exposed rather than caused.
`_fitting_font` walks PressStart2P then 4x6 at 6px, and "255.255.255.255" --
the widest thing the startup banner ever shows -- measures 66px at 4x6/6px
against the 62 a 64x32 panel has to give. It used to squeak in only because
the measurement depended on which layout engine the host Pillow happened to
have; with the engine pinned it does not, so the rung the worst case actually
needs is now in the ladder instead of implied: 4x6 at 5px, which measures 51.
The fallback was wrong in the same place. When nothing in the ladder fit, it
returned `self.font` -- the *widest* option, and precisely how "Initializing"
came to run off the side of a 64px panel to begin with. It returns the
narrowest face that loaded now.
test/test_initializing_screen.py: 34 passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display): name the exceptions the BDF strike read can raise
Codacy flagged the try/except/pass. It was already narrow in intent -- a
malformed strike table on the measurement path must degrade to "size unknown"
rather than take the display down -- but a bare `except Exception: pass` says
neither of those things and hides a genuinely broken font behind a silent 8px
fallback. It now catches what reading `available_sizes` can actually raise and
logs which face failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore: drop logo PNGs the render harness downloaded into the worktree
These are fetched at runtime by the logo cache; they are not source, and they
rode in on a `git add -A` while I was running check_plugin.py against this
branch. Nothing in the change needs them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge
Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of
patches and services built while running this project on Starlark apps
under MQTT control. Its patches are whole-file copies taken against an
older tree, so applying them as written would revert #523's frame
pacing, #534's display() bool returns and the GitHub token masking in
plugins_manager.js. Three of its claimed fixes are already in main, and
its api_v3 Starlark routes are #535's. What follows is the rest --
verified against current code, and reimplemented where the patch's
approach did not hold up.
**On-demand display.** `pinned` reached the controller from the API, was
stored on it and republished in the status payload, but never narrowed
the rotation -- a pinned request still cycled every mode its plugin
owns. Right for a sports plugin, whose modes are views of one subject;
wrong for a plugin whose modes are unrelated, which is every Starlark
app. Now honoured, and it survives a restart.
Restarting while on-demand was active loaded *only* the on-demand
plugin, so normal rotation had nothing to return to for the life of the
process -- and a restart mid-session is routine, since that is how an
update is applied. The panel came back cycling one plugin's modes with
no way out but clearing the cache by hand. Every enabled plugin loads
now; on-demand still resumes on its saved mode.
Stop requests are exempt from the duplicate guards on purpose, so that a
second click stops a mode a race left running -- which means consuming
the mailbox is the only thing that ends one. It was never consumed, so
the same stop was re-read and re-processed on every poll, forever. Both
paths now share one compare-before-delete helper.
**Starlark rendering.** `extract_schema` parsed the source with a regex,
which can only see option lists written out literally: an app whose
dropdown is filled from a live API call inside `get_schema()` came back
empty, and the config form offered nothing to pick. Now runs `pixlet
schema`, which executes the app, and falls back to the parser when
Pixlet is absent, too old for the subcommand, or the app fails to run.
The third-party patch replaced the parser outright and hardcoded
/usr/local/bin/pixlet; this keeps the fallback and the binary search.
A `|` in a config value was dropped by a shell-metacharacter filter,
though the command is a list with no shell involved -- and apps do use
it as a separator inside one value. The key went missing silently and
the app rendered its own "not configured" screen with nothing to say
why. And a 0-byte render was reported as success: Pixlet exits 0 and
writes nothing when an app has no content, which read downstream as a
working app drawing a black panel.
**Starlark display.** `display()` ignored the mode it was called with,
so a specific app could not be addressed. It now accepts `display_mode`
-- which is the whole mechanism, since the controller inspects the
signature before passing it. Found while there: `_select_next_app` ran
only while `current_app` was unset, so with several apps installed the
first was picked once and shown forever while the rest were rendered on
schedule and never displayed. And `enable_scrolling` was missing, so
multi-frame apps were called once per rotation slot and never advanced
past frame one.
**GET /api/v3/display/modes.** Every mode that can be requested
on-demand, with the plugin that owns it. Nothing exposed this, so
anything driving the display from outside the web UI read each plugin's
manifest.json off disk and reimplemented PluginManager's fallbacks. It
also triggers discovery, which is otherwise lazy and normally happens
because a person opened the dashboard.
**integrations/mqtt_bridge.** Home Assistant control over MQTT
Discovery: a mode select, a stop button, power, brightness. Rewritten
against the API rather than the filesystem, so it needs no read access
to config.json and cannot drift from the web UI. paho-mqtt 2.x
VERSION2, TLS, an availability topic that is also the last will, and
secrets from the environment.
**Two opt-in extras.** A DNS single-request unit, for glibc's parallel
A/AAAA lookup stalling ~5s per name on routers that answer only the A
query -- which makes any plugin calling an external API slow and
Starlark apps, which have a render timeout, fail outright. And a Pixlet
config editor: a script you run and Ctrl+C rather than the third-party
version's always-on unauthenticated Flask service, since it stops the
display for the length of a session. Neither is installed by default.
Long Starlark app names now wrap instead of overflowing their card.
115 new tests across 5 files. Also unblocked
test_starlark_display_contract.py, which was silently skipping wherever
fcntl is absent. Whole suite: no new failures against main.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(mqtt_bridge): the five issues Codacy flagged on this branch
All in the new bridge, all real:
* requests floor was 2.31.0, which carries CVE-2024-35195,
CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which
is what the project's own requirements.txt already pins.
* `import time` was never used.
* `"mqtt_password": None` in DEFAULTS read as a hardcoded credential.
It is the "no password configured" default; marked nosec B105, the
convention used elsewhere in the repo.
Also dropped an unused `build_app` from the display-modes test imports.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: the review findings on this PR
Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and
is answered below.
**One bad config section blanked the whole mode list.**
`/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`,
so a non-dict under a plugin id -- a shape DisplayController already
guards, so it happens -- raised AttributeError mid-loop and answered 500
with no modes at all. Every MQTT bridge entity is built from that list.
Now skipped with a warning.
**The DNS scripts reported success they had not earned.** Three separate
paths: `resolvconf -u` failing was swallowed by `|| true`; the
systemd-resolved branch exited 0 without applying anything, so the
oneshot unit recorded success while the workaround was inactive; and the
installer's `|| echo` turned a failed start into "installation
complete." with exit 0. All three now fail loudly. `single-request` is a
glibc resolv.conf option with no resolved.conf equivalent, so on those
hosts the honest answer is that it cannot be applied.
A NetworkManager-generated resolv.conf is regenerated on connection
changes, not only at boot, and the unit is oneshot with RemainAfterExit
-- so the option can vanish mid-boot with nothing to put it back. Now
detected and stated plainly rather than implied to be permanent.
**`Before=` does not order a manual restart.** It only orders units
already in the same transaction, so `systemctl restart ledmatrix` could
bypass the fix. install_dns_fix.sh now writes a ledmatrix.service
drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround
failing should not stop the display.
**The Pixlet editor's `--lan` is gone.** `pixlet serve` has no
authentication, and a printed warning is not access control. Loopback
only, with the SSH port-forward in the header where the flag used to be
documented -- SSH does the authenticating and nothing is left listening.
**The MQTT example config now defaults to TLS** on 8883. The installer
copies it verbatim, and without TLS the broker password and every
command cross the network in cleartext. A plaintext broker is still
supported and documented, and the bridge warns once at startup when a
password is configured without TLS.
**Not taken: "the upstream Pixlet CLI has no `schema` subcommand."**
Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh`
installs `tronbyt/pixlet`, whose `cmd/schema.go` is
`schema [PATH]` -> JSON on stdout, built on
`runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is
exactly what extract_schema_via_pixlet calls. A binary without the
subcommand exits non-zero and falls back to the source parser, which is
already covered by a test.
**CodeQL stack-trace exposure: not taken either.** I removed `details`
first and that broke
test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception,
which enforces `describe_exception` across all ~75 handlers -- written
because a device with failing storage answered "see logs for details"
from the log viewer itself. describe_exception redacts credentials; the
trade-off is the project's and is already made. Restored, with the
reasoning in a comment.
11 new tests. Whole suite: no new failures against main, 4127 passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The two review nitpicks left over from #535. Both are still on main after
that merge; the five findings alongside them landed with it.
**The toggle could not find what the list had just shown.**
`_starlark_virtual_plugins` publishes the raw manifest key as
`starlark:<key>`, and `_toggle_starlark_app` passed it back through
`_validate_and_sanitize_app_id`, which lowercases and rewrites every
character outside `[a-z0-9_]`. An app stored as `My-App` was listed as
`starlark:My-App` and looked up as `my_app`, so toggling an app the page
had drawn a moment earlier answered 404. Keys written by
`_install_star_file` are already sanitised, so this only shows up for
manifests written by the starlark-apps plugin itself or edited by hand.
`_validate_starlark_app_path` rejects traversal without rewriting, so it
is the check to use here -- listing and toggling now agree on one key.
The updater also uses `setdefault` rather than indexing: the app is
loaded but its on-disk entry need not exist, and `_update_manifest_safe`
does not catch `KeyError`, so that escaped as a 500 rather than writing
the entry.
**`star_file` was stored absolute.** Readers join it to the app's own
directory -- `_standalone_render_starlark_app` does `app_dir /
app_data.get('star_file', f'{app_id}.star')` -- so the key's default is a
bare filename and an absolute value gave it a second meaning. Since
`Path.__truediv__` discards the left side when the right is absolute,
the manifest was pinned to whatever PROJECT_ROOT installed it, and a
moved or redeployed install could not find its own file. Storing
`dest.name` matches the default and stays relocatable. Read paths are
unchanged, so manifests already holding an absolute path keep working.
7 new tests. Whole suite: no new failures against main, 4013 passed
against 4007.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The Pixlet install button reported "Pixlet install failed: Resource not found" -- Flask's 404 handler, because the route did not exist. #253 added thirteen Starlark routes; #330 rewrote api_v3.py and dropped all of them, along with the `starlark:<app_id>` entries that surface installed apps in the plugins list and the toggle branch that enables them.
Restores all thirteen routes, the plugin-list entries and the toggle path, so Pixlet installs, the app store browses and installs, and an installed app can be managed like any other plugin.
Not a straight revert. Three error paths stopped returning exception text to the caller; the manifest write moved off a shared temp filename that two concurrent writers could interleave; both dynamic importers stopped leaving half-initialised modules in sys.modules; the config update rolls back when the save fails; the toggle checks that persistence succeeded; and the path check returns the validated path instead of a boolean so callers stop re-joining the raw value. New tests no longer reach GitHub.
Verified on a 256x64 Pi: Pixlet installs and runs (v0.53.1), the store lists 1000 apps, install/toggle/uninstall round-trip, and traversal and command-injection probes are rejected at every entry.
25 CodeQL alerts dismissed as verified false positives -- path-injection where traversal is blocked, and one list-form subprocess with no shell. Both classes already present on main.
Full core suite: 3981 passed.
install_plugin() deliberately renames a plugin's directory to the MANIFEST id
when it differs from the REGISTRY id, so registry `stocks` lands in
`ledmatrix-stocks/`. Every lookup in _find_plugin_path() is by directory name,
so update_plugin("stocks") found nothing, logged "Plugin not installed", and
returned False.
Nothing surfaced that to the user. Clicking update in the web UI was a no-op
with no error, and the plugin stayed on a stale version indefinitely. Four
installed plugins hit this on a real device -- leaderboard, music, stocks and
weather -- found because a scripted update of eleven plugins failed on exactly
those four.
Adds a manifest-id scan as the LAST step of the resolution chain, so the two
documented lookups above it (configured dir, then the sibling plugins/
fallback) keep their exact meaning and ordering. That ordering is pinned by
test_discovery_path_contract.py, which characterises the divergence between
the three resolvers on purpose; this extends the chain rather than reordering
it. Directories renamed aside with '.standalone-backup-' during an install or
rollback are skipped, since matching one would report a half-finished install
as a live plugin.
Also adds scripts/audit_render_path.py, which walks the call graph from
display() and reports blocking calls reachable from it. display() runs on the
render thread, so anything slow there stalls the panel; on a vsync-paced loop
a single 15ms call drops a frame and a network round trip freezes the marquee.
Two instances were already found the slow way, by reading frame-time
histograms -- odds-ticker reading the scoreboard cache per frame, and
soccer-scoreboard timing out inside update(). The audit finds that shape in
the source instead. It is a heuristic and says so: a hit behind an interval
check may be fine.
It currently flags 23 calls across six plugins. The clearest is
ledmatrix-music, whose display() falls back to an inline
requests.get(timeout=5) when album art has not been prefetched -- a deliberate
"show the art rather than go blank" tradeoff by its author, but up to five
seconds of frozen panel. Reported, not changed; that is its owner's call.
185 store tests pass. Three of the six new tests fail without the fix.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* perf(scroll): pace frames to the panel, not to a fixed sleep
Scrolling ran at 44-46 fps on a 2x128x64 chain and 14-17% of frames took
41-53ms, which reads as judder. Four independent causes, each measured on
the hardware; details and the diagnostic recipe are in
docs/SCROLL_PERFORMANCE.md.
The high-FPS loop slept a flat 8ms after every render. display() has
already blocked on the panel's vsync by then, so that sleep was added to a
wait that had happened: ~4ms of render plus 8ms put each iteration at ~12ms
against a 10ms refresh grid, so every swap missed a refresh and the loop
settled at 50fps while asking for 125 -- with no headroom, so a further
14% of frames slipped again. It now sleeps only the remainder, with a 1ms
floor so plugin threads still get the GIL.
ScrollHelper stepped position on a wall clock at 1/scroll_delay steps per
second. Plugins set scroll_delay to the frame period, so that comparison
sat exactly on its own threshold: a frame arriving a hair early moved zero
pixels and rendered an identical frame, dirty-tracking skipped the swap, it
returned in ~2ms, and the beat repeated. No scroll_delay value tunes that
out -- a shorter delay trades stalled frames for periodic double-steps.
Both modes now accumulate elapsed time at the same configured speed, so
position stays proportional to real time.
Sub-pixel blending goes back to off by default. It renders a half-step by
mixing two adjacent columns, which on a coarse panel showing pixel-font
text alternates crisp and smeared frames and reads as shimmer -- visibly
worse than integer stepping on the hardware. Vegas mode still opts in.
disk_cache uses orjson when importable, falling back to the stdlib. Encoding
a ~1MB record drops from 14.8ms to 5.4ms end-to-end, and that work holds the
GIL while a marquee is on screen. display_manager also checksummed the whole
framebuffer twice per frame (dirty tracking, then the preview snapshot); the
snapshot now takes the checksum the caller already computed.
New src/common/scroll_config.py resolves scroll settings in one place. Five
ticker plugins each hand-rolled this and disagreed: odds-ticker ranked the
deprecated scroll_pixels_per_second above the documented scroll_speed/delay
pair, and because that key carries a schema default the documented settings
were dead for every user (ChuckBuilds/ledmatrix-plugins#408), while
ledmatrix-leaderboard read the same key only as a fallback. The resolver also
warns when a speed will not advance a whole number of pixels per refresh,
which is the property that actually determines whether a scroll looks smooth.
scripts/build_rgbmatrix_nogil.sh rebuilds the rgbmatrix binding so it
releases the GIL. Upstream declares SwapOnVSync without nogil, unlike
SetPixel/Clear/Fill beside it, so the render thread held the GIL for the
whole vsync wait and starved background threads into long uninterruptible
bursts. The script patches, builds and self-verifies into a scratch tree;
--install backs up the original and rolls back if the service does not come
back healthy.
Measured after: 100 fps locked, no stalls observed, render thread down from
51% to 19% of one core.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display): keep the panel swap locked to vsync while scrolling
Dirty tracking skipped SwapOnVSync for byte-identical frames. That is the
right call for static content, but SwapOnVSync is also what paces the render
loop, so skipping it skips the wait for the panel: a duplicate frame returns
in ~8ms instead of ~10ms on a 100Hz panel, advances the strip only 0.8px
instead of 1.0px, and so makes the next frame more likely to repeat as well.
The effect sustains itself once it starts.
Measured over 20 minutes on a 2x128x64 chain, both scrollers configured
identically at 100 px/s:
leaderboard 10ms x35, 11ms x3 (clean)
odds-ticker 10ms x26, 8ms x7, 15ms x5 (~20% duplicates mid-scroll)
The duplicates were not end-of-cycle idling -- 38% of fast frames fell within
90s of a scroll completion against 35% of normal frames, a null result. The
trigger is per-frame work: odds does more of it, and more variably, so it is
first to land a frame that advances less than a whole pixel.
Pushing an identical frame costs one canvas copy. Falling out of vsync lock
costs smooth motion. Static content is untouched, because
is_currently_scrolling() expires on its own inactivity threshold -- covered
by test_stale_scrolling_state_stops_forcing_pushes so a plugin that stops
scrolling without saying so cannot pin the panel into always-push.
Also de-flakes test_snapshot_still_written_on_skip, which asserted a strict
mtime increase between two writes that can land in the same filesystem tick;
it failed about two runs in three on Windows regardless of the code under
test. The file is now backdated before the check.
156 tests pass on the Pi. Not yet confirmed by eye on the panel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scroll): report the frame-time tail, and stop the row-major blit
Two problems, both found by looking at the panel rather than the metric.
The frame-stats line reported ONE instantaneous frame every 5 seconds --
about 1 frame in 500 -- printed beside a 100-frame average. Both hide exactly
the fault they are used to chase: a 2ms duplicate and a 21ms double-wait
average to precisely 10ms, so a ticker stalling on half its frames still
reports a healthy "Avg FPS: 100.0". That reading cost several rounds of
chasing the wrong layer. The line now aggregates every frame since the last
log and reports median, p95, max, min, and explicit stall and skip rates
(past 1.5x the median missed a refresh; under half never reached the panel,
because dirty tracking skipped the swap so the frame never waited on vsync).
On the hardware this now reads:
leaderboard 100.0 fps over 501 frames | median 10.00ms p95 10.05ms
max 10.34ms | stalls 0 (0.0%) skips 0 (0.0%)
The binding rebuild's blit patch becomes opt-in (RGB_PATCH_BLIT=1, default
off). Reordering that loop to row-major changes what a torn frame looks like:
column-major tearing shows as a vertical seam, row-major as a horizontal split
between the panel's upper and lower halves. On a 1/32 scan panel that reads as
a one-pixel fold across the middle of every panel, which is what was reported
on hardware and what went away when the blit was reverted. All of the measured
gain comes from the SwapOnVSync change, so the risky half is simply not worth
taking; the header says so.
Also fixes --install resolving its paths against $HOME, which is /root under
sudo, so it looked in /root/rgbmatrix-nogil-build and died with "no built
module found" on a machine where the build had just succeeded. It now resolves
SUDO_USER's home. Both build paths are verified on the Pi: default yields one
GIL-release site, RGB_PATCH_BLIT=1 yields two.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(scroll): let users pick a crisp speed for their own panel
Whole-pixel motion was previously only available at multiples of the refresh
rate -- 100, 200, 300 px/s on a 100Hz panel. 100 px/s crosses a 256px panel in
2.6s, which is brisk for reading, and everything slower had to blend (blur) or
repeat frames unevenly (judder). There was no way to ask for 50 px/s and get
clean motion.
SwapOnVSync takes a framerate_fraction the display manager never passed. It
holds each frame for N panel refreshes; the panel keeps refreshing at its full
rate throughout, so holding costs nothing in flicker and only changes how often
a NEW image is presented. That turns 50 px/s into one whole pixel every second
refresh instead of half a pixel every refresh.
The crisp speeds are therefore refresh_hz / hold * pixels_per_frame, and that
ladder depends on the panel: a Pi Zero on a long chain has a different set of
good speeds from a Pi 4 on a short one. crisp_ladder() enumerates them and
solve_crisp() picks the best match for a requested speed.
solve_crisp weights motion quality rather than picking the numerically nearest
entry, which matters more than it sounds. Asked for 30 px/s, nearest-by-value
answers 28.6 -- 2px jumps at 14fps -- over 33.3, which is single-pixel motion
at 33fps and obviously better on the panel. The target is also clamped into the
ladder's range first, because relative error saturates near 1.0 for a target
far outside it and the quality penalty would otherwise answer "10000 px/s" with
the slowest entry.
configure() snaps to the ladder and applies the hold when given a display
manager. Without one the hold silently cannot happen and motion falls back to
fractional pixels, so it warns rather than failing quietly. set_frame_hold()
resets to 1 when scrolling stops, so one plugin's pacing cannot leak into
whatever is on screen next.
scripts/scroll_speeds.py is the user-facing part: it prints the ladder for the
configured rate, measures what the panel ACTUALLY manages (--measure, for
hardware that cannot reach its configured limit), highlights the nearest option
to a wanted speed, and demos one live. It never starts or stops the display
service itself -- doing that inside a script stranded the panel twice today.
Speeds below ~20 px/s remain stepped regardless. That is the pixel pitch, not a
software limit.
Also fixes the dirty-tracking test spy, which stubbed SwapOnVSync with a
single-argument function and would have masked the new call as a failed push,
and rewrites a configure() test that had started passing for the wrong reason:
it asserted a judder warning, which snapping now prevents, and was matching the
unrelated "hold could not be applied" warning instead.
183 tests pass on the Pi.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scroll): tie the frame hold to the scroll, not the plugin
The hold applied in configure() never reached the panel. Plugins share one
display manager, and set_scrolling_state(False) -- fired whenever ANY other
plugin finishes its scroll -- reset the hold to 1. A hold set once at plugin
construction was therefore always gone by the time that plugin rendered.
The symptom was a log line that lied. ledmatrix-stocks reported
Scroll configured: 50.0 px/s (1px every 2 refreshes = 50.0 fps, smooth)
while the panel measured 100.0 fps, median 10.00ms. Config, resolution and
snapping were all correct; only the pacing silently was not applied.
set_scrolling_state(is_scrolling, frame_hold=1) now carries it, so the hold
lives exactly as long as the scroll that asked for it. configure() reports the
value as ScrollSettings.frame_hold instead of applying it -- applying it behind
the caller's back could never have been right on a shared display manager.
Existing callers are unaffected; the default keeps one frame per refresh.
Verified on hardware: stocks at 50 px/s now measures
50.0 fps over 251 frames | median 20.00ms p95 20.09ms | stalls 0 skips 0
20.00ms being exactly two refreshes, with the panel still refreshing at 100Hz
underneath so flicker is unchanged.
test_another_plugin_stopping_does_not_strand_a_hold pins the interaction that
broke this.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(scroll,cache): resolve CodeRabbit review on #523
Eight findings, all reproduced before fixing.
scroll_config.configure() read the refresh rate *after* resolve() had
already used it. resolve() fills in target_fps, pixels_per_frame and the
judder warning from that rate, so on a 60Hz panel every one of them
described 100Hz -- and with snap_to_crisp=False nothing downstream
corrected it, so set_target_fps() paced the helper to 100 FPS. The rate
is now settled first, and falls back to the global config rather than
straight to the default.
refresh_hz_from_config() used `(cfg.get("display") or {}).get(...)`,
which raises AttributeError when either level is truthy but not a
mapping -- out of a function whose whole contract is a rate or a default.
The frame-stats line reported the upper-middle sample as the median and
the 96th sorted sample as p95 of 100. Both are also thresholds (stalls
at 1.5x the median, skips at 0.5x), so the counts were biased too. The
arithmetic is now in frame_stats()/format_frame_stats(), testable
without a clock.
configure()'s docstring and docs/SCROLL_PERFORMANCE.md still said it
applies the frame hold and warns when it cannot. It deliberately does
neither since "tie the frame hold to the scroll, not the plugin"; a
caller following the old text would omit set_scrolling_state() and slow
snapped speeds would still present every refresh.
disk_cache had no policy for non-finite floats: orjson writes null,
the stdlib writes NaN/Infinity, and orjson then rejects those legacy
files so DiskCache.get deleted them as corrupt. One behaviour on both
paths now -- write null, keep legacy records readable. allow_nan=False
detects the values; the replacement walk runs only when there is one,
so the ordinary write path is byte-identical and pays nothing.
build_rgbmatrix_nogil.sh picked the build artifact with a glob piped to
`head -1`, which sorts cpython-311 ahead of cpython-313, so a stale .so
staged in from the source tree was installed as core.so while the GIL
check -- which reads the generated core.cpp, not the .so -- still passed.
It now requires the current interpreter's exact ABI name and fails
closed. Its systemctl calls were also unchecked under `set -uo pipefail`:
a failed stop left the old service running, the following start
succeeded as a no-op, and the health check reported SUCCESS for a
binding that was never loaded.
orjson floor raised to 3.11.6 for CVE-2025-67221 (unbounded recursion
in dumps); it covers the project's Python 3.10-3.13 range.
Adds test/test_cache_nonfinite_floats.py (14) plus regression tests in
test_scroll_config.py and test_scroll_helper.py. 9 of the cache tests
and 9 of the scroll_config tests fail against the pre-fix code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* test(harness): keep the visual double's signature tied to production
Moves set_scrolling_state's frame_hold into the test double here, where
DisplayManager gains it, rather than in #534 where it arrived a PR early.
CodeRabbit flagged the #534 version correctly: a double that accepts an
argument production does not lets the call pass every harness run and
raise TypeError on the panel, which is the one failure a safety harness
exists to prevent.
The drift has now gone both ways across two branches -- double behind
production on this branch, double ahead of it on #534 -- so it is pinned
instead of remembered. test_display_double_parity.py compares the two
signatures and fails with the direction of the drift named. It reads the
files with ast rather than importing them, because display_manager
imports rgbmatrix at module scope and this check should hold on a laptop
and in CI as well as on a Pi.
Plugins begin passing frame_hold in ledmatrix-plugins#462, which is why
production and the double both need it before that lands.
Full suite: 3889 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers
Three independent fixes found while validating every plugin on a 256x64 rig.
FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf
ships with the core, so every plugin offering it logged "Font family
'tom_thumb' not found" (16 warnings per countdown render) and had to carry a
private loader to use a bundled font. Closes#524.
VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter
that DisplayManager gained, so any plugin passing it died with TypeError at
render time and failed every size. Nine plugins now make that call;
ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other
seven only passed because their scroll path was unreachable without data.
Closes#525.
api_v3 declared module-level config_manager/plugin_manager = None that nothing
ever assigned -- app.py sets the blueprint attributes, which the other 150+
call sites use. Three sites read the decoys, so /health reported the config
unreadable and the plugin system uninitialised (making "degraded" permanent and
unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on
every rig. The decoys are removed rather than assigned, so a bare name is now a
NameError at test time instead of a silent None. The same function's first-call
uptime was computed from two separate clock reads and came out negative.
Closes#529.
Verified on the rig: both previously-failing plugins render, the tom_thumb
warnings are gone, /health reports "healthy" with all three checks passing, and
/display/current reports the real 256x64.
* fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark
The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is
world-writable and sticky, and the display service runs as a different user
from the tooling, so a leftover temp owned by anyone else became unopenable
even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a
sticky directory. The preview and the health check's liveness proxy then froze
until someone deleted the file by hand; on the test rig that meant 23 hours of
a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with
cleanup on failure, matching the hardware-status write a few hundred lines
above. Closes#528.
_poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults
the in-memory TTL to max_age -- so the first request was pinned in memory for an
hour and every later poll returned that stale copy. No second on-demand request
was honoured until the service restarted, while the API kept returning 200.
get() already documents memory_ttl=0 for exactly this cross-process case.
The consumed request is also now deleted: leaving it on disk meant a restart
replayed the previous request, activated it, and ignored the one the caller had
just made. Closes#530.
starlark-apps returned None from display() when it has no app to show, which is
the state of every install without Pixlet and of a fresh one before any app is
added. The controller only skips on a boolean False, so that held a black panel
for the full display_duration instead of rotating on. Closes#456 (core side).
Verified on the rig: two consecutive on-demand requests with no restart between
them are both activated, where the second was previously dropped in silence.
* perf(harness): share one cache across a plugin's renders
_instantiate built a fresh MockCacheManager for every (size, mode), and that
mock is a per-instance in-memory dict, so each render was a cold start. A plugin
that fetches per game or per player re-fetched everything N times over --
baseball-scoreboard at one size took 840s for nine renders where the arithmetic
said ~72s, and at eight sizes it exceeded a 900s timeout.
The second and later renders also never exercised the cache-hit path, which is
what a running rig executes almost all of the time, so a caching regression
could not be caught here.
The cache is now built once per render_plugin_matrix call and threaded down.
The display manager stays per-render -- the bounds checking depends on that --
so only fetched data is shared.
Measured on the rig, same render counts and same goldens:
tide-display 2s -> 1s (32 renders)
cricket-scoreboard 10s -> 3s (24 renders)
No pass/fail change across tide-display, cricket-scoreboard, clock-simple,
geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages.
Closes#533.
* fix(scripts): run standalone plugin tests instead of collecting nothing
run_plugin_tests.py discovered every plugin test file and handed the lot to
pytest. Most plugin tests are standalone scripts -- module-level main() plus an
`if __name__ == "__main__"` guard, signalling through an exit code -- and pytest
collects zero items from those. The run printed how many files it had *found*,
then "no tests ran", and exited without executing any of them. On a rig with all
44 first-party plugins that is 151 of 248 files.
Files are now classified and each kind runs under the right runner: pytest for
real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1
fail convention ledmatrix-plugins' own runner established (a script that wants a
tty or an LED matrix is a skip, not a regression).
Before:
$ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos
Found 1 test file(s)
collected 0 items
no tests ran in 0.31s rc=0
After:
Found 1 test file(s) -- 0 collectable, 1 standalone script(s)
1 passed, 0 skipped, 0 failed (scripts) rc=0
Verified across three shapes: countdown (1 script), jellyfin-now-playing and
pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files
split 4 collectable / 7 scripts, all seven of which had never run).
Closes#532.
Running the flights scripts for the first time also surfaced four genuinely
failing tests there, hidden by the mirror-image bug in the plugins repo's own
runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465.
* fix(harness): give an empty-looking mode a few frames before warning about it
check_plugin's "drew nothing but display() returned X" warning fired on a single
frame, rendered with force_clear=True, under a frozen clock. All three defeat a
scrolling plugin, whose first frame is legitimately its blank scroll-in buffer.
Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which
people stop reading a warning, which matters because the true positives are
real: a mode that draws nothing and does not return False holds a blank panel
for its whole display duration.
An apparently-empty frame is now re-driven for up to 48 more frames with
force_clear=False (force_clear means "reset the scroll", so repeating it would
redraw frame 1 for ever) and with the clock advancing -- freezegun's factory
where time is frozen, a real sleep where it is not, since scroll position is
usually a function of elapsed time. The first frame that draws content replaces
the result.
The clock is moved back afterwards. It is shared by every render in the matrix,
so time borrowed by the probe leaked into later modes and drifted their goldens
-- f1_upcoming picked up 5 spurious drifts before this was restored.
Measured on the rig:
empty warns check
before after
f1-scoreboard 42 0 48 PASS / 0 FAIL, goldens intact
ledmatrix-elections 16 0 16 PASS / 0 FAIL
on-air 8 8 true positive, kept
nfl-draft 8 8 true positive, kept
clock-simple/geochron/ 0 0 unchanged
christmas-countdown
58 false positives gone, both true positives kept, no golden regressions. Cost
is confined to modes that really are blank: plugins that draw immediately are
unchanged (clock-simple and tide-display still 2s), while on-air -- eight
deliberately blank modes -- goes to 21s.
Closes#527.
* fix(harness): load nested schema defaults, and merge caller config at leaf level
load_config_defaults read only top-level properties. An object property carries
its defaults on its children, not on itself, so everything nested was dropped --
2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of
565. render_plugin_matrix's comment says the plugin then "behaves like a real
install", which for most of the fleet it did not.
_defaults_from_properties now recurses. merge_config deep-merges the caller's
config onto the result so an override lands at the leaf: a shallow merge would
let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard
every other nhl default, which is the same class of bug being fixed here.
Measured before/after across all 49 installed plugins on the rig: **no render
changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact.
Plugins already fall back to the same values internally via config.get(key,
default), so supplying them explicitly agrees with what they were doing. The
defaults really are arriving now:
ufc-scoreboard 9 -> 87 defaults
ledmatrix-flights 51 -> 95
masters-tournament 10 -> 51
cricket-scoreboard 22 -> 50
tide-display 12 -> 18
and hockey-scoreboard, which used to load nhl.enabled=None, now gets
nhl.enabled=True with its full display_modes block.
Caveat worth carrying: the eight plugins with the most nested config
(soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of
the 2,386 dropped defaults, 68%) could not be measured. They import
src.common.sports_shared, which the test rig's core branch predates, so they
fail to load there identically before and after. Re-run this comparison against
a core that has that module before trusting the "nothing changed" result for
them; those are exactly the plugins whose renders should change most.
Closes#531.
* refactor: narrow the exception handlers this branch introduced
Codacy flagged the new code; it passes on other recent PRs, so the finding is
mine. Four of the five broad `except Exception` clauses I added were catching
far more than they needed to, which is the same shape as several bugs this
branch fixes -- hello-world's TypeError sat invisible for exactly this reason.
freezer() / move_to() / tick() -> (AttributeError, TypeError, ValueError)
cache_manager.delete() -> (OSError, AttributeError, KeyError)
The fifth stays broad and now says why: it wraps a call into a plugin's own
display(), which can raise anything, and the first frame has already rendered --
so a failure there must not turn a good result into an error.
Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty
warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows
7 golden drifts both before and after this branch, so it is not from these
changes -- its committed goldens predate #521's 1-bit text rendering.
* fix: resolve CodeRabbit review and Codacy findings on #534
CodeRabbit raised six; all six were real.
The test double had drifted ahead of production. VisualTestDisplayManager
accepted set_scrolling_state(frame_hold=...) while DisplayManager did not,
so such a call passed every harness run and would raise TypeError on the
panel -- the one failure a safety harness exists to prevent. frame_hold
belongs to the change that adds it to DisplayManager (#523), so it moves
there and the double matches main again.
The harness swallowed exceptions from re-rendered frames. _settle_loop
re-renders a mode that came back blank, to give a scroll time to draw;
returning silently on a crash meant a mode that renders one good frame
and then explodes was reported as passing. Recorded on result.error now,
keeping the captured frame so the failure stays inspectable.
starlark-apps display() returned True after _display_frame() failed, so
the controller held a dead frame for the whole display_duration instead
of rotating on. _display_frame now returns bool on all three paths.
run_plugin_tests.py used env.setdefault for PYTHONPATH and
LEDMATRIX_CORE, so an inherited value won and the subprocess imported a
different core than the one under test -- ledmatrix-plugins#467 exactly.
Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally.
The on-demand mailbox is polled after every frame, ~125x/second on a
scrolling mode, and the read is deliberately uncached, so it was that
many disk reads per second to find nothing. Floored at 250ms, which is
imperceptible for a web-UI click. Consuming it also deleted whatever was
present rather than what had just been processed, so a request posted
while the previous one was in flight was thrown away and never ran; the
delete is now keyed by request_id. That narrows the window rather than
closing it -- a true atomic claim needs a primitive the cache layer does
not offer, and the code says so rather than implying otherwise.
Codacy's 2 criticals were bandit B404/B603 on the subprocess call added
to run_plugin_tests.py. Fixed interpreter, argument list, no shell;
annotated with the repo's existing nosec convention. Bandit is clean on
the file.
Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py
(4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of
those fail against the pre-fix code. Full suite: 3961 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* chore: satisfy Codacy's subprocess checks on the new test runner
Codacy runs Bandit and Opengrep (its Semgrep fork). The new
subprocess.run in scripts/run_plugin_tests.py trips three patterns, on
two different lines:
Bandit B404 on the import, B603 on the call
Opengrep dangerous-subprocess-use-audit on the run( line
dangerous-subprocess-use-tainted-env-args on the argv line
A nosemgrep applies only to its own line, so the call line and the argv
line each need one; a single comment on the call covered neither rule
fully. Suppression is the right answer here rather than a rewrite: the
interpreter is sys.executable, the arguments are a list, and no shell is
involved, so there is nothing to word-split or expand.
Matches the pair the rest of the repo already uses for this shape --
permission_utils.py, plugin_loader.py, install_dependencies_apt.py.
Codacy: 0 new issues, up to standards.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* chore: leave visual_display_manager untouched so #523 can merge
The only change this branch made to that file was a docstring, and it
collided with #523's rewrite of the same method -- so #534 and #523 each
merged cleanly against main but conflicted with each other. Reverted to
main's text; #523 owns this method and adds frame_hold to it.
The note the docstring carried ('frame_hold arrives in #523') would have
been stale the moment #523 landed anyway. The parity test in #523 is
what actually keeps the two signatures honest.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* chore: add the Ruff suppression nosec/nosemgrep do not cover
Ruff reports S603 on the same call Bandit and Opengrep do, and none of
the three suppressions covers the others. Confirmed the precondition
first: path comes from discover_plugin_tests(), which globs test files
inside the repo, and the call is a fixed interpreter with a list argv
and no shell.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 13:37:38 -04:00
ChuckClaude Opus 5coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
* feat(render_plugin): add --display-mode so multi-mode plugins can be rendered
render_plugin.py always called plugin.display(force_clear=True) with no mode.
A plugin that declares one display mode is fine, but the sports scoreboards
declare three or more and keep their per-mode state on sub-managers; their
no-argument path selects nothing and returns False, so the render came out
blank with nothing to say why. Measured on nrl-scoreboard with identical
seeded state:
live.display() directly True, 1892 lit pixels
plugin.display(display_mode="nrl_live") True, 1892 lit pixels
plugin.display() False, 0 lit pixels
--display-mode passes the requested mode through. It is only passed when
asked for, so the many plugins whose display() takes no display_mode keep
working untouched, and a plugin that declares modes but does not accept the
argument degrades to its default screen with a warning rather than a
TypeError.
This is what lets the plugin READMEs show a scoreboard at all, and it also
unblocks screens like birdnet_stats and the weather plugin's hourly, daily and
almanac modes, which could previously only be described in prose.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Only fall back when plugin display rejects display_mode
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
An LED panel has no partial brightness. PIL defaults ImageDraw's fontmode to
"L", which anti-aliases TrueType glyphs into a grey fringe the panel can only
round off -- a 4px glyph arrives smeared into 3px.
DisplayManager creates its shared `draw` in six places and set fontmode at
none of them, while _load_fonts loads extra_small_font as 4x6-font.ttf at
size 6. Measured at draw time, that face at that size puts 74% of its lit
pixels at partial coverage. Every plugin drawing small text through the
shared draw inherited the blur; geochron was the case that surfaced it.
The harness's VisualDisplayManager had the same gap, which mattered more than
it looks: goldens were recording anti-aliased text that production would not
produce, so the harness could not have caught this. Fixing only production
left geochron still blurry under the harness -- that is how the second site
was found.
Both are set to "1" so the harness renders what the panel renders.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
_schema_font_size swallowed every exception and cached an empty dict. That is
not cosmetic. With no schema, a configured font size can no longer be compared
against the schema default, so every size is treated as a deliberate user
choice and skips the snap to the font's pixel grid -- which renders
4x6-font.ttf at 6 instead of 7: a 3px-wide glyph instead of 4px.
That shipped. On a 256x64 panel it made the odds, the team records and the date
row hard to read, and it was found by a user counting pixels on a photo of the
panel rather than by anything here. The cause (_plugin_dir returning None under
the real plugin loader) is fixed in #519; this makes the same class of failure
audible next time:
Orphan: could not read config_schema.json (FileNotFoundError: ...); every
font size will be treated as user-chosen and will skip its pixel grid
snap. Font sizes may render a pixel narrow.
The message names the consequence, not just the error, because the error alone
does not suggest "your fonts are a pixel narrow".
Logged rather than raised: an unreadable schema must not stop a plugin
rendering. The cache is built once per class (per schema path in sports_card),
so this cannot repeat per frame.
Scope deliberately small. An audit of the three shared modules found 23 handlers
that swallow and return a default, but all 23 catch specific types -- TypeError,
ValueError, ImportError -- turning bad config values into defaults, which is
what they are for. Of 77 broad handlers across the font and odds paths, 74
already log. Only these two were both broad and silent.
_plugin_dir() returned None on every device. The consequence was silent and
reached the panel:
_plugin_dir() -> None
_schema_font_size() -> None for every element
-> a configured size equal to the schema default stops looking like a
default and is treated as a deliberate user choice
-> the snap to the font's pixel grid is skipped
-> 4x6-font.ttf renders at 6 instead of 7: 3px-wide glyphs, not 4px
On a 256x64 panel that made the odds, the team records and the date row hard to
read. Both `odds` and `detail` were affected -- anything resolving a
grid-snapped schema default was a pixel narrow.
Why it was invisible here. PluginLoader._namespace_plugin_modules renames every
bare module a plugin brought in (sports, game_renderer, ...) to
"_plg_<plugin_id>_<module>" and REMOVES the bare sys.modules entry, so two
plugins owning a module of the same name cannot collide. A class defined in
sports.py still reports __module__ == "sports", but sys.modules["sports"] is
gone, so walking the MRO for a module with a __file__ finds nothing.
Every test here imported plugins directly, which leaves the bare entry in
place, so the walk succeeded. The safety harness loads plugins its own way and
never reproduced it either. It was found by a user counting pixels on the
panel.
The directory is now declared by the plugin (_PLUGIN_DIR) and only deduced as
a fallback, for hosts that declare nothing -- the plugins' own probe harnesses
build classes with type().
Verified on hardware, which is the only place the original failure appeared:
before, the live service logged plugin_dir=None and 4x6-font.ttf@6 for all six
football managers; after, plugin_dir resolves and both odds and detail are @7.
Five regression tests, including the production shape: a class whose __module__
is absent from sys.modules still resolves via its declared directory, and the
precondition that the MRO walk alone returns None is pinned so the test keeps
meaning something if the fallback changes.
* fix(store): read the core version from disk, not from a stale import
Updating the core to 3.3.0 and then updating plugins refused all eight sports
scoreboards:
Refusing to install nrl-scoreboard: NRL Scoreboard supports LEDMatrix
>=3.3.0, but this system is running 3.2.0.
while src/__init__.py on that machine read 3.3.0. Observed on hardware, not
theorised.
The gate ran `from src import __version__ as core_version`, which binds
whatever the process loaded at start. The plugin store's gate lives in the web
UI, a long-lived service of its own, and the update route deliberately restarts
nothing -- it replaces files on disk and asks the user to restart. Its prompt
named only the *display* service, so a user who followed it left the web
process holding the previous number.
Stale by exactly one release is the case that bites: every plugin flooring on
the release you just installed is refused, blaming a core version that is
already correct on disk. It reads as a broken plugin store. 3.3.0 is the first
release where this hits a whole family at once, since all eight scoreboards
floor there.
compatibility.current_core_version() reads the version from the file instead,
falling back to the imported value on any failure -- so it can only ever be as
correct as before, never worse. All four gate call sites use it: three in
store_manager (install, the git-pull update path, install_from_url) and one in
plugin_loader's advisory warning.
The restart prompt now names both services.
Twelve tests, including the hardware failure itself: a process holding 3.2.0
while disk says 3.3.0 refuses hockey, and reading fresh allows it. The inverse
is asserted too -- a genuinely old core still refuses, so the gate has not
become permissive. One test greps both modules for the old import-bound read;
reintroducing that line fails it, which is what stops this coming back.
Not changed: web_interface/__init__.py also imports __version__, but for
display rather than gating, and the API endpoint already reports a fresh
git describe.
* fix: drop the unused os import
Left over from a first draft that joined paths by hand before this used
pathlib. Flagged by CodeRabbit on #518; confirmed dead -- no os. reference
remains in the module.
* feat(sports): share the sports.py surface that is identical in all eight
Nine scoreboards ship their own sports.py -- 41,326 lines. Comparing executable
ASTs across the eight that share a lineage, 48 method bodies are byte-identical
in every one: 1,007 lines carried eight times, so 8,056 lines that must be
edited eight times to fix once.
They are the parts with no sport in them: the selection and rotation engine
(_round_robin_favorites, _favorites_first, _compose_selection,
_check_ranking_coverage, _game_divisions, _normalise_quality), the font/colour/
date subsystem (_scale_headline_fonts, _scorebug_font, _resolve_font_size,
_format_game_date, _font_color), and the switch-mode upcoming card
(_draw_upcoming_center_switch). Nothing here knows what an inning is.
Mixins rather than free functions: every one of these reads host state, so
rewriting 48 bodies into free functions would be a rewrite rather than a move,
and it is the move that keeps the renders identical.
Three of the 48 are deliberately left in the plugins, because a byte-identical
body is not automatically safe to move:
- _get_timezone calls resolve_timezone, imported from a per-plugin module
(hockey_timezone, soccer_timezone, ...). All eight of those differ -- each
carries its own _WRITEBACK_FIXED_IN -- so hoisting the caller would silently
bind every scoreboard to one plugin's copy.
- _extract_game_details and _fetch_data are @abstractmethod stubs. They are the
sport contract; satisfying them from a mixin would let a plugin instantiate
without implementing its own sport.
_resolve_font_path went the other way: a module-level function, identical in all
eight, that _scale_headline_fonts needs -- so it is inlined here.
_schema_font_size needed a real change rather than a move. It located the
plugin's config_schema.json with __file__, which here is src/common/, so the
load failed silently, the cache stayed empty and every element fell back to an
unsnapped size -- measured at 81% anti-aliased edges on a panel that should be
pixel-crisp. It now recovers the plugin directory from the instance. Note that
type(self).__module__ alone is not enough: SportsCore is an ABC, so a subclass
built with type(name, bases, ns) -- which the plugins' own tests do -- reports
its module as "abc". _plugin_dir walks the MRO past those synthetic classes to
the first module sitting beside a config_schema.json.
Worth recording: the 176 harness renders did NOT catch that regression. The
plugin's own test_fonts_are_crisp.py did. Renders alone were not a sufficient
gate here.
Not merged with src/common/sports_card.py despite fourteen same-named twins.
Only five are provably equivalent by source comparison; the other nine differ in
ways inspection cannot settle, and a wrong guess silently changes what every
scoreboard draws. That merge needs differential testing and is its own change.
* fix(sports): declare the constants the mixins read, and test the contract
CodeRabbit found _QUALITY_CHOICES and _RANKING_COVERAGE_SECONDS read by
_normalise_quality and _check_ranking_coverage but never defined on a mixin.
Confirmed: both are declared by all eight scoreboards, so nothing fails today --
it would only have bitten the ninth plugin to adopt this, at runtime, mid-render.
Both are identical everywhere, so they get defaults here; each plugin's own copy
still shadows them.
Auditing for others showed those two were the only ones, but also that the
host-contract docstring was substantially incomplete: it listed 21 attributes
where the mixins actually read about 40, and omitted five hooks
(_is_favorite_game, _is_game_really_over, _is_ranked_game,
_passes_other_filters, _get_timezone). The section is now derived from that
audit rather than remembered.
test_sports_shared.py covers what is genuinely new, not the moved bodies:
- The contract itself. It parses the module for every ALL-CAPS `self.X` the
mixins read and asserts each is defined, so the next omission fails here
rather than in the field.
- _plugin_dir, the only new logic in the move. Including the case that made it
necessary: SportsCore is an ABC, so a subclass built with type(name, bases,
ns) -- which the plugins' own tests build -- reports __module__ as "abc". The
test asserts that precondition before asserting the walk steps past it.
- The three SportsLive bodies. Hockey and lacrosse disable live mode in their
harness fixtures, so the 176 renders never reach this path; testing the mixin
directly means coverage no longer depends on which plugin happens to have a
unit test.
Two of those tests pin things that would otherwise be silently undone.
SportsRecentSharedMixin does carry an __init__ -- SportsRecent.__init__ was one
of the 48 byte-identical bodies. Its bare super() binds to where it is defined,
now the mixin, so it only reaches the host because the mixin is listed first in
the bases. One test proves the chain runs; the next proves that reversing the
order silently skips the host constructor.
* fix(sports): drop three unused imports and let the matcher narrow
Codacy flagged five issues on this file.
Three are unused imports: math, abc.abstractmethod and zoneinfo.ZoneInfo.
Nothing in the module references any of them -- the timezone work goes through
pytz, and the @abstractmethod mention in the module docstring describes the
two stubs that deliberately stayed behind in each plugin, not anything
declared here. pyflakes agrees; all three are removed.
The other two are "team_in is not callable" on the round-robin favourite
matcher. That call is already guarded by callable(), so it cannot raise at
runtime, but callable() is not a narrowing construct a static analyser
follows: the name still carries the None from getattr's default. Normalising a
non-callable to None and branching on `is None` gives the analyser a test it
does understand, and keeps the guard.
Behaviour is unchanged. _round_robin_favorites has no test coverage, so I
exercised it directly on both paths -- a host with _team_in (id matching, the
NRL case) and one without (abbreviation matching) -- across limits 1 to 4, and
the selections are identical before and after. A host whose _team_in is
present but not callable still falls back to abbreviation matching rather than
raising.
test/test_sports_shared.py: 27 passed. The 9 collection errors under
`pytest test/ -k sport` reproduce identically on the unmodified branch and are
not from this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
v3.3.0 is tagged, but src/__init__.py still says "3.2.0" -- and that string,
not the git tag, is what the compatibility gate compares
(store_manager.py: `from src import __version__ as core_version`).
The effect is that every plugin flooring at 3.3.0 is refused on a device
running 3.3.0. Checked against the real gate and the real manifest:
core __version__ reported to the gate : 3.2.0
hockey floor : 3.3.0
verdict : REFUSE
"supports LEDMatrix >=3.3.0, but this system is running 3.2.0"
That is all eight sports scoreboards plus calendar 1.2.3, and it would read as
a broken plugin store rather than a stale constant.
The TRUSTWORTHY_FLOOR escape hatch does not cover this: it exempts cores
reporting below 2.0.0 as "unknown rather than old", and 3.2.0 is above it, so
the number is trusted and compared.
This is the same slip as v3.1.0, which was tagged six weeks before its version
string was bumped and shipped __version__ = "1.0.0" -- the reason that escape
hatch exists at all.
With the bump, the same gate call returns ALLOW for hockey at both the
sports_card and sports_shared floors, and for calendar 1.2.3.
CHANGELOG.md gains the 3.3.0 section. That file is what plugin authors read to
decide which release to floor on, so it records the three new modules --
src/common/sports_card.py, sports_game_renderer.py and sports_shared.py --
against this version, along with the three traps in adopting the mixins.
No test pinned the old literal; the ones that care monkeypatch __version__.
135 compatibility tests pass.
* feat(sports): share the scroll-card geometry the scoreboards all duplicate
The card helpers moved to src/common/sports_card.py, which shared the eight
scoreboards' settings lookups. Their *geometry* stayed duplicated: nine
methods deciding how wide the centre strip is, how much room each logo gets,
and where an upcoming card's date and time land. Five were byte-identical in
all eight plugins; the other four were identical in seven, each with a
different single outlier.
That shape is why this is a mixin and not free functions. Comparing executable
ASTs against the eight plugins, 67 of the 70 method bodies are inherited
unchanged and 3 become ordinary overrides -- baseball keeps its own
_logo_slot_width and _draw_upcoming_game_status, hockey its own
_upcoming_date_and_time. No per-sport branching goes inside the base.
It deliberately has no __init__ and no state. The plugins' constructors differ
six ways and none of it is worth unifying, so adoption is one line on the
class statement plus deleting what now comes from here.
Placed in src/common/ rather than src/base_classes/sports/ on purpose:
importing that package pulls core.py -> DisplayManager -> rgbmatrix, and this
is pure geometry that must not drag a hardware import into every plugin that
uses it. It sits next to sports_card.py, which the same plugins already use.
Only _SCORE_PROBE varies between plugins, so leagues that reach three digits a
side override that one ClassVar; the four gap constants are identical
everywhere.
The tests drive the mixin through a host that provides exactly the surface the
module docstring names and nothing else, so the mixin growing a new self.*
dependency the plugins do not have fails the contract test rather than
shipping.
* fix(sports): reject non-finite card settings before they abort the render
A center_gap of inf passes `isinstance(x, (int, float)) and x >= 0` unharmed
and then raises OverflowError out of int(). The surrounding guards caught only
(TypeError, ValueError), so it escaped and took the whole card render with it.
The same holds for center_gap_ratio, the two clamp bounds, and layout offsets,
where "inf" arrives as a string and float() is happy to produce it.
Four of the five paths crashed; only a NaN ratio happened to survive, by
accident of min/max rather than by design.
This is pre-existing behaviour -- the bodies moved here verbatim from the eight
plugins and every one of them has it today. Fixing it in the mixin fixes it in
all eight at once, which is the argument for the mixin.
Guarded with math.isfinite() before any int()/round(), falling back to the same
defaults the finite paths already use, plus OverflowError added to the except
clauses as a backstop. Ordinary settings are untouched: all 192 scroll-card
renders (8 plugins x 8 panel sizes x 3 game types) stay byte-identical to
pristine main.
Found by CodeRabbit on #514 and confirmed by running it before fixing.
Twenty methods were byte-identical in all eight scoreboards' game_renderer.py:
the colour pickers, the scroll_card settings lookup, the date and time
formatting, the favourite-team rules and the font-size grid snapping. Every fix
to any of it had to be made eight times, and a new scoreboard began by copying
them a ninth.
src/common/sports_card.py holds them once. 245 lines leave each plugin.
**Free functions, not a base class.** Every helper takes config/logger/fonts as
arguments rather than reading them off an instance, so a plugin keeps its
method and delegates the body -- call sites, signatures and override points are
all untouched. Adoption is therefore per-function and reversible, which is what
let all eight move with byte-identical renders.
The bodies are the plugins' code moved, not rewritten. Two deliberate
differences, both verified:
- crisp_size takes the seven-plugin guard (`not desired`) rather than
football's. They agree on every real input; the extra guard only stops a
None size raising TypeError, so adopting it is a no-op for seven plugins and
removes a crash path for the eighth.
- schema_font_size caches per schema PATH. The plugins cached on their own
class, which is the same distinction expressed without a class to hang it
on; two plugins never share an entry. The path has to be passed in because
the plugins derived it from __file__, and __file__ here is the core's.
Verified before any plugin was touched: 534 differential comparisons of the
helpers against afl's originals and 704 more of the font-sizing chain against
all eight plugins' originals -- 1,238 comparisons, zero differences. Writing
the constant tables by hand introduced two errors that check caught: the tie
colour was (255,255,0) instead of the plugins' (255,200,0), and a
"five_by_seven" alias that does not exist. Both are now taken from the plugins
verbatim.
43 tests pin the contract, including the cases the plugins' own comments record
as having bitten: a three-character string must not iterate into a colour, a
shared font face must give up rather than guess an element, a font_size equal
to the schema default carries no intent, and a bad timezone falls back to UTC
rather than blanking the card.
Full suite 3768 passed, 6 skipped.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(logos): stop a failed download pinning a team to a grey box forever
When a logo download fails, create_placeholder_logo writes a 64x64 grey PNG
under the *real* logo's filename. Every later call then hits
`if filepath.exists(): return True` and reports success, so the real logo is
never attempted again. One transient failure -- no network at boot, ESPN
blipping -- permanently costs that team its logo.
This is not hypothetical. Five of the eleven cached AFL logos in my checkout
were 384-byte stubs written in a single bad minute, and they had stayed that
way ever since; the scoreboard rendered COLL, FRE, NMFC, PORT and SYD as grey
text boxes on every card.
Placeholders are now stamped with a `ledmatrix_placeholder` PNG text chunk
carrying their creation time, and `is_placeholder_logo` recognises them. It
also matches on the placeholder's exact geometry and background colour, so the
stubs already sitting on users' disks are picked up too -- without that, this
fix would only help teams whose logos break in future. Verified against the
real stubs: all five detected, all six real logos untouched.
`download_missing_logo` now treats an existing placeholder as the failed
download it is and retries, rather than as a satisfied request. The retry is
rate-limited to PLACEHOLDER_RETRY_SECONDS (6h) so this does not trade a
permanent grey box for an ESPN request every frame; a failed retry rewrites the
placeholder, restarting the clock. The age comes from the stamp rather than
mtime, so a backup restore, an rsync, or a permissions script cannot silently
reset it.
`download_missing_logos_for_league` gets the same treatment -- a bulk pass is
exactly where a previously failed logo should get another chance -- and
`LogoHelper.load_logo_with_download` no longer accepts a stale placeholder as a
cache hit. That import is lazy and guarded so the module still works against a
core build predating the marker.
`LogoHelper._create_placeholder_logo` needs no change: it returns an in-memory
image and never writes it to disk, which is the behaviour this bug argues for.
Tests cover marked and legacy-unmarked detection, the two false-positive cases
(a real 500x500 logo, and a 64x64 image that is merely the same size), the
retry, the rate limit, and that the age survives an mtime touch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(logos): address review — unify eligibility, invalidate cache, restart back-off
Three findings from the review on #512, all confirmed against the code:
1. The three download sites each had their own idea of "already have it".
download_missing_logos_for_league() retried *any* placeholder, ignoring the
back-off entirely, while download_all_ncaa_football_logos() was never
updated and still skipped placeholders forever. They now share one
should_attempt_download(), which also covers force_download, so the sites
cannot drift apart again. download_missing_logo() reads through the same
helper.
2. LogoHelper.load_logo_with_download() answered from the in-memory cache
before touching the disk, so after a stale placeholder was successfully
replaced the *cached placeholder image* was still returned -- the real logo
would not have appeared until the process restarted. The cache entry for
that file (every size of it) is now dropped after a successful download.
3. A failed retry left the stale placeholder on disk with its old timestamp,
so the next call saw it as stale again and retried immediately: a download
attempt per call, which is precisely what the back-off exists to prevent.
refresh_placeholder_timestamp() restamps it, and the helper calls that on
the failure path. It refuses to touch anything that is not a placeholder.
Tests cover both bulk loops in both directions (fresh placeholder skipped,
stale one retried), the eligibility rule including force_download, the
timestamp refresh, and the two LogoHelper paths -- including that a
freshly-downloaded logo is actually what comes back rather than the cached
placeholder.
Two of the new bulk-loop tests initially passed for the wrong reason: the
fetch_teams_data stub returned {}, which is falsy, so the loops bailed before
reaching the eligibility check at all. Fixed to return a truthy payload.
Re-verified end to end: with both halves in place, rendering the AFL scoreboard
took FRE.png from a 362-byte stub to a 12,928-byte logo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* docs(sports): record where B6 stands, and why it is waiting
The phase table had B4 as "next" and B5 as "after B4" while both had shipped,
and described B6 as blocked on B4's gate — which is now merged and released. A
plan that misreports which phase it is in is worse than no plan: the next
person reads it and repeats finished work.
Corrected, and three things that were only ever decided in conversation are now
written down:
* **B6 is deliberately held.** 3.2.0 published 2026-08-03; 3.1.0 ran nine
months before it. B6's premise is that cores without the module are gone,
and there is no release-asset count or install telemetry to show that.
Running it now strands users on their current plugin versions. The gate
that makes it safe is already built and tested — it is the calendar that is
missing, and no amount of further code changes that.
* **Stop adopting further shared modules** (data_sources, game_renderer,
base_odds_manager) until B6 closes. Each adoption adds a copy to keep in
step against a payoff contingent on B6.
* **A B5 retrospective**, because "the adoption went fine" is not what
happened: four of eight shipped with scroll mode broken on a 3.2.0 core.
The bundled fallback did not protect against it — the break was on the
modern path — which is an argument for the sunset, not against it. Records
the ledger too: net negative on disk until B6 runs.
Also replaces the "what's next" list, whose first five items were all done,
with what actually remains.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix: apply CodeRabbit auto-fixes
Fixed 1 file(s) based on 2 unresolved review comments.
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
* docs(sports): stop a wrapped PR reference reading as a heading
A line wrapped onto "#433), the newest manifest entry ...", which
markdownlint reads as a malformed ATX heading (MD018). Reflowed so the
line starts with "(#431, #433)" instead.
Not the suggested fix: adding a space after the hash would have turned
the PR reference into "# 433". The B5 safety claim raised alongside this
was already corrected in ac44b5a.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* docs(sports): scope the B5 safety claim to the fallback
The heading read "B5 — adoption is safe by construction", which this same
document disproves two sections later: four of the eight adopted plugins
shipped with scroll mode broken on a 3.2.0 core and were repaired in
plugins #251.
The body was already careful -- it says fallback compatibility is what is
guaranteed, and that correctness on a core which *does* ship the module
needs object-level and scroll-mode validation. The heading was not, and a
heading is what a reader scanning the plan actually takes away.
Retitled to name both halves, with a sentence up front saying why the
unqualified claim is false and pointing at the retrospective that shows
it. The phase intro said "one of them is safe by construction and the
other is not"; that now says what it actually means -- one cannot break a
user on an old core, the other can.
The second review point, MD018 on the ATX heading at line 409, does not
reproduce: that line now begins "(#431, #433)" rather than "#433)", so
there is no bare-hash heading. `grep -cE '^#+[^ #]'` returns 0 for the
whole file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* docs(sports): re-check the hold, and close two items that are already done
The remaining-work list had two entries that finished without the doc noticing,
which is the failure mode this file exists to prevent.
- The stale plugin-test tranche is gone. run_plugin_tests.py --all now reports
174 passed, 2 skipped, 0 failed across the whole fleet. Recorded how to
re-check it too: these are standalone scripts, not a pytest suite, and one
calls sys.exit(1) at import, so pointing pytest at a plugin directory
collapses into an INTERNALERROR that looks nothing like the real state.
- CLAUDE.md already says eight panel sizes.
That leaves the hardware soaks as the only open item needing work rather than
calendar time.
B6's prerequisite is now built -- core test/test_sports_sunset_matrix.py
(#505) -- so the phase table and the regression-test section say so, and the
two modelling traps it had to work through are recorded for whoever touches it
next: the copy-removed shape must be an unguarded import or the failure names
scroll_display_legacy instead of the core module, and only the leaf module may
be hidden because a pre-3.2.0 core still ships src/common/.
The hold itself is re-checked and unchanged: v3.2.0 is still latest,
__version__ is still 3.2.0, no 3.3.0, 23 days rather than the few months the
gate asks for. Also worth stating plainly -- the core updates by git pull, not
by downloading a release, so release-asset counts would not measure uptake even
if we had them. Whatever unblocks this has to come from the store side.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
#510 shows as merged, but into fix/gate-git-pull-updates -- #508's branch --
rather than main. #508 reached main first, so the sideload gate was left behind
on a branch. Same failure as plugins #350/#351, which merged into each other's
bases; worth knowing the pattern, because GitHub reports these as MERGED and
`gh pr list` shows nothing outstanding.
main today has two of the three routes gated: install_plugin (#431/#433) and
update_plugin's git branch (#508). install_from_url validates required manifest
fields and then installs whatever it found, never comparing the core version.
Cherry-picked unchanged from the orphaned branch -- it applies to main with no
conflict. TestSideloadGate pins the three cases the other routes pin: refuses a
floor above this core leaving nothing behind, still allows a compatible plugin
(the guard against a gate that refuses everything), and does not block a 2.0.0
floor on a core reporting an untrustworthy version.
Full suite 3725 passed, 6 skipped.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Callers that join an in-flight fetch share one FetchResult. #499 released the
payload inside the delivery loop, so the first callback got the data and every
joiner got `result.data is None`.
That is not a quiet degradation. Consumers read `result.data.get('events')`, so
they raise AttributeError -- which the delivery loop catches and logs. The
entire failure surfaced as one line:
ERROR - src.background_data_service - Error in callback for request
nhl_2026_...: 'NoneType' object has no attribute 'get'
and a manager that silently never received its schedule. Seen on hardware:
NHLRecentManager logs "Background fetch completed for 2026: 1000 events" and
the very next line is the error, from NHLUpcomingManager's callback on the same
request -- which had already logged "No events found in shared data."
Deduplication is the normal case, not a corner. A sport's recent, upcoming and
live managers all want the same season schedule, so the second and third are
joiners on almost every cycle. _release_payload's own docstring said "once A
callback has been handed the data", singular, which is the assumption that
broke: the loop above it was written for many, and says so.
Moved after the loop, and guarded on `callbacks` being non-empty. The guard
matters: a request submitted without a callback must keep its payload, because
polling get_result() is then the only way to collect it. The per-delivery
release got that right by accident -- an empty list never entered the loop body
-- and the existing test for it caught the omission.
test_background_payload_release.py gains TestJoinersAllGetTheData: two
submitters on one in-flight cache_key, asserting both are handed a populated
payload, plus that the memory fix still happens once they have all had it.
test_background_fetch_dedupe.py already proved the joiner's callback FIRES; it
never checked what the callback received, which is the gap that let this
through.
Verified the new test bites: restoring the release inside the loop fails it
with "'second' was handed a released payload".
Full suite 3716 passed, 6 skipped.
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(store): gate the git-pull update path
`install_plugin` gates every route that re-downloads, `_reinstall_with_rollback`
included. `update_plugin` has one branch that re-downloads nothing: a git
checkout pulls in place, installs dependencies, and returns True. A pull could
therefore deliver a manifest flooring above this core and nothing would notice
until the plugin failed to load — which surfaces as one line in the journal and
a display that silently stopped appearing.
Checked after the pull rather than before it, for the same reason
`_install_plugin_impl` checks after the download: the registry carries no
compatibility field, so the incoming floor is only knowable once the new commit
is on disk.
Undone with `git reset --hard` to the pre-pull commit rather than by removing
the directory. This is a live checkout, the old commit is still in the object
store, and the reset leaves the user on the exact version they were already
running — the same promise `_reinstall_with_rollback` makes, reached by the
means this path actually has, with no window where the plugin directory does
not exist. An unreadable manifest allows: it is not evidence of a floor.
Scope, stated plainly: monorepo plugins install as archives and update through
`_reinstall_with_rollback`, so they were already gated. Only registry entries
with no `plugin_path` reach this branch. It is closed anyway because the sunset
rule in the plugins repo's `08-shared-sports-code.md` names, as condition 3,
that the core enforces the floor "at install/update time" — and B6 rests on
that being true rather than merely written down. `install_from_url` is still
ungated; the tests say so rather than letting the next reader assume otherwise.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(store): do not pull what the gate cannot un-pull
Review of the gate found a data-loss path it had introduced, plus two smaller
scope errors. All three from CodeRabbit on #508.
**The stash failure was load-bearing and was not treated as one.** update_plugin
stashes local changes before pulling; when that stash failed or timed out it
logged a warning and pulled anyway. That was harmless while nothing ever undid
a pull. It is not harmless now: the gate's rollback is `git reset --hard`, which
discards uncommitted tracked edits -- exactly the edits the stash existed to
protect. A pull does not refuse on a dirty tree as long as the incoming commit
touches other files, so the sequence completed silently: pull succeeds, gate
refuses, reset takes the user's work with it.
update_plugin now returns before pulling unless the tree was already clean or
was successfully stashed. Refusing costs an update in a case that had already
gone wrong; the alternative costs data. That also makes `--hard` safe by
construction in _gate_pulled_commit, and its comment now says so rather than
observing it in passing.
Pinned by test_a_failed_stash_stops_the_update_before_pulling, which writes a
local edit, forces the stash to fail, and asserts both that HEAD did not move
and that the edit is still on disk. Verified it bites: with the new guard
removed the file comes back as `class P: pass`, the edit gone.
**_HAS_GIT could take the module down instead of skipping it.** With no git on
PATH, subprocess.run raises FileNotFoundError, and this runs at import time --
before skipif can act, so the whole file errors rather than skipping. Now
catches OSError.
**The doc overclaimed the gate's reach.** It said the floor is enforced on
"every route that installs or updates" while the same passage notes
install_from_url is ungated. Both spots now scope the claim to registry-managed
installs and the two supported update paths, and name the sideload exception.
Full suite 3720 passed, 6 skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four league fallback logos were each tracked at two paths differing
only in case:
assets/sports/mlb_logos/MLB.png + mlb.png
assets/sports/nba_logos/NBA.png + nba.png
assets/sports/nfl_logos/NFL.png + nfl.png
assets/sports/nhl_logos/NHL.png + nhl.png
On Windows and default macOS the filesystem is case-insensitive, so
both index entries map to one physical file. Whichever git writes last
wins and the other entry reports as permanently modified, so `git
status` can never be clean and any `git pull` flips which one is dirty.
For MLB/NBA/NFL the two entries pointed at the same blob, so the
collision was only cosmetic. NHL was not: NHL.png is 1439x1621
(672455 B) and nhl.png is 768x768 (107184 B), so which resolution the
league fallback logo loaded depended on checkout order rather than on
the code.
Keep the uppercase path in each pair. Every logo lookup uppercases the
abbreviation before building a filename -- LogoDownloader
.normalize_abbreviation and .get_logo_filename_variations
(src/logo_downloader.py) and LogoHelper.normalize_abbreviation
(src/common/logo_helper.py) all do -- and nothing in the tree requests
a lowercase league logo, so the uppercase name is what the code
actually asks for. For NHL that is also the higher-resolution asset.
Removed with `git update-index --force-remove` so the literal index
entry is dropped without the case-insensitive working tree deleting the
survivor.
B5 and B6 make different promises about an adopted scoreboard, and only the
second is obvious. The table:
| bundled copy present | bundled copy removed
pinned old core | loads (the fallback) | ERROR, naming the module
current core | loads, using core code | loads, using core code
The top-left cell is the one worth having. Nothing in the suite proved that an
adopted plugin still runs on a core predating src.common.sports_scroll, and
that claim is the entire basis for having shipped B5 ahead of B6's gate.
The bottom-left cell is asserted through PluginManager.load_plugin rather than
a bare import, deliberately. The manager catches the ModuleNotFoundError, so a
test written around pytest.raises would pass against a core where the module is
merely broken rather than absent, and would say nothing about what the user
meets: a plugin parked in ERROR and one log line. The assertion is the ERROR
state plus an error naming the module, which is what makes the log actionable.
Modelling notes, both of which were wrong first time and matter:
- The copy-removed shape is an UNGUARDED import, not the guarded one with the
legacy file deleted. Keeping the guard while removing its fallback only
mislabels the failure -- the plugin reports a missing scroll_display_legacy
and never mentions the core module that is actually absent.
- Only the leaf module is hidden. A pre-3.2.0 core still ships src/common/;
hiding the package would be a harsher core than any that shipped, and would
make the failure name the package instead of the module.
Ablation: disabling the old-core simulation fails three cells, and dropping the
error object from load_plugin's set_state fails the fourth, so none of it
passes vacuously. The install gate -- the other half of the guarantee -- stays
in test_plugin_compatibility_gate.py rather than being duplicated here.
Refs docs/SPORTS_UNIFICATION.md, "B6 -- why the sunset needs more than a
version floor".
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(memory): join an in-flight fetch instead of starting a duplicate
submit_fetch_request() had no notion of "already fetching this". request_id
embeds a millisecond timestamp, so every submit looked new, and
active_requests is keyed by that id rather than by what is being fetched.
Two submits for the same cache_key therefore started two identical fetches.
It is not a rare race. _fetch_data in the sports managers branches: the Live
manager fetches only today's games, but Recent and Upcoming both pull the
full season schedule under the SAME cache_key. On a cache miss both miss,
both submit, and nothing stops the second. On a running 512x64 board:
138 background fetches in 24 hours, arriving in pairs at identical
millisecond timestamps, roughly hourly:
2 2026-08-25 11:47:26.612
2 2026-08-25 10:46:28.064
2 2026-08-25 08:01:55.962
Half of them redundant. Each duplicate costs a second download, a second
JSON parse -- the expensive part on a Pi -- and a second parsed copy
resident at the same time. Schedules on that board run 256KB to 20MB, 106MB
across all sports. Because the pairs land in the same millisecond they also
occupy two of the three executor slots with identical work, which is what
makes two large parses peak simultaneously.
A submit for a cache_key already in flight now joins that request: its
callback is added to the existing one and the existing request_id is
returned, so get_result() works for both. Different keys are untouched, and
dedupe applies only while a fetch is in flight -- a submit after completion
fetches again, because this is not a second cache layer.
Three details:
- The in-flight entry is dropped and the callback list snapshotted in the
SAME critical section as filing the result. Otherwise a submitter could
join a fetch whose callbacks had already run and never be called back.
- Cancellation is the other way a request leaves active_requests, so it
releases the key too. And the join path looks the request up rather than
trusting the id, so an entry stranded any other way cannot wedge a key
permanently -- it is dropped and a fresh fetch starts.
- One callback raising no longer prevents the others being delivered.
Previously there was only ever one.
Interaction with #499, whichever merges second: that PR releases the payload
after the callback runs. With several callbacks the release must happen
after ALL of them, and must not happen at all if a joined submitter passed
no callback, since polling get_result() would then be its only delivery
path. The callback list built here is the hook for that.
test_background_fetch_dedupe.py -- 8 tests, covering the join, callback
delivery to both submitters, one callback raising, distinct keys not being
coalesced, a post-completion submit fetching again, cancellation releasing
the key, a stranded entry not wedging one, and the reported count. Verified
non-vacuous by removing only the join branch: 3 fail. The callback tests
assert the ids coalesced, without which they would pass trivially on two
independent requests.
Full suite: 3698 passed, 60 skipped, 1 failure that reproduces identically
on unmodified main (test_install_lowmem, environment-dependent: /var/tmp is
disk-backed on this machine).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(memory): discard a cancelled fetch instead of letting it commit
Review follow-up on the dedupe.
Cancelling releases the cache_key, so a replacement fetch for that key can
start immediately. But _fetch_data_worker() had no cancellation check: the
cancelled worker still wrote its response to the cache, flipped its own
status from CANCELLED to COMPLETED, and ran its callbacks. The stale
response could therefore land on top of the replacement's fresher data.
The worker cannot abort an HTTP call in flight, so the response is discarded
on return instead: no cache write, no callbacks, status left CANCELLED. The
check sits immediately before the cache write, which is the first
side effect.
Also fixed, found by the new test rather than by reading:
request_id was f"{sport}_{year}_{milliseconds}", which is not unique.
Two submits inside the same millisecond produced the SAME id -- the
test's two sequential fetches collided on a fast mocked response, and
one request silently replaced the other in active_requests and
completed_requests. Rare before this PR; load-bearing now, because
dedupe hands that id back to every joiner as their handle for
get_result(). A per-service counter is appended.
Two test problems of my own, both fixed here rather than left to flake:
- The cancellation test synchronised with time.sleep(0.4). A slow worker
would have made it pass for the wrong reason. It now waits for the
request to be filed in completed_requests.
- The id-uniqueness test patched session.get, but submits are async: the
50 workers outlived the patch and made real DNS calls to the dummy host.
It stubs the executor instead, which is what a submit-time test should
exercise.
20 consecutive runs of the dedupe file: 0 failures. Full suite: 3700 passed,
60 skipped, 1 failure that reproduces identically on unmodified main
(test_install_lowmem, environment-dependent).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(memory): make cancellation terminal, not advisory
Three paths wrote request.status without checking whether the request had
already been cancelled, so a cancel could be silently undone and the work it
was meant to stop went ahead anyway.
- A request cancelled while queued had CANCELLED overwritten with IN_PROGRESS
the moment its worker started, defeating the discard check entirely: it
downloaded, cached and called back for work the caller had withdrawn. It
now skips the fetch outright, which is also the cheapest possible cancel.
- The cancelled-check and the cache write were separate critical sections, so
a cancel landing between them left the payload in the cache with the
callbacks suppressed -- every submitter that joined the fetch waited for a
call that never came. The worker now claims the commit in the same critical
section that reads the status, and cancel_request refuses once claimed. The
write stays outside the lock: it serialises a multi-megabyte payload to the
SD card, and holding the service lock across that would stall every submit,
status query and cancel behind it.
- A cancelled request that then failed was relabelled FAILED, which slipped
past the CANCELLED-only callback gate and delivered a spurious error
callback. The except path now leaves CANCELLED alone.
get_request_status() also reported a cancelled request as FAILED, since it
inferred status from result.success; the final status is now recorded on the
result. Both early returns assign to `result` so completed_requests files the
outcome that was reported rather than the untouched placeholder.
Tests cover cancellation before worker start, during the commit, and during an
HTTP failure; all three fail against the unfixed code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* test(memory): reach the exception path as a cancelled request
test_a_failure_after_cancelling_stays_cancelled cancelled the request while
its worker was still queued, so once the pre-start branch landed the worker
returned there and never reached the exception handler the test is named for.
It passed against the unfixed code only because that branch did not exist yet;
with it, the test passed for the wrong reason and reverting the except-path
guard did not fail it.
Cancel while the worker is parked inside the HTTP call instead, and assert the
fetch actually started so the test cannot silently degrade into the pre-start
case again. Reverting each of the three guards now fails exactly one test.
Also read the payload inside the callback rather than off the FetchResult
afterwards: #499 releases result.data once delivery is done, so the later read
saw the released object and not what the caller was handed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(plugins): retain state history by age, with the count as a ceiling
Follow-up to the cap in this PR. A flat entry count answers the wrong
question: what a reader wants from this history is "the last couple of
hours", and how many transitions that is depends entirely on the plugin's
update interval. On a real board those span 2s to 3600s, so 200 entries is
interval 200 entries covers
2s 3.3 minutes (flights, live)
10s 16.7 minutes (jellyfin)
60s 1.7 hours (default)
300s 8.3 hours (news)
3600s 4.2 days
-- the plugin churning hardest, the one worth looking at, keeps the least.
So transitions are now trimmed by AGE first
(STATE_HISTORY_MAX_AGE_SECONDS, two hours), which makes the retained window
comparable whatever the cadence, and the count cap becomes purely a memory
ceiling for pollers fast enough to exceed it inside that window. The
ceiling rises 200 -> 2000: at ~230 bytes an entry that is ~0.5MB per plugin
worst case, and only plugins updating faster than roughly every 4s can
reach it. Steady-state memory is unchanged for everything slower, since the
age trim binds first.
Two details worth stating:
- The trim reads time.monotonic(), stored alongside each transition,
rather than the datetime already inside it. A DST shift or an NTP step
would otherwise make every entry look ancient and flush the history in
one go. The human-readable timestamp is untouched and still what
get_state_history() returns.
- Trimming happens on append, so a plugin that goes quiet keeps its last
window until it writes again. That is deliberate: it is bounded either
way, and a lazy trim costs nothing on the hot scheduling path. The
guarantee is therefore about the SPAN of retained history, not its age
against the current clock, and the test asserts it that way.
The public shape is unchanged: get_state_history() still returns the same
list of transition dicts, and state_history_count is still the lifetime
total.
test_plugin_state_history_retention.py adds 7 tests. Verified against this
branch with only the age trim removed: 4 fail, 3 pass -- the three that
survive are testing the count ceiling and the monotonic clock, which this
commit does not change. Full suite 3753 passed, 60 skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(plugins): build get_state_info() as one locked snapshot
Every field was read under its own lock, so an unload running concurrently
could be observed half-done: 'state' read before clear_state() removed it
and 'state_history_count' read after, handing PluginManager.get_plugin_info()
a plugin that is ENABLED with zero transitions.
The whole payload is now built in one critical section. _lock is an RLock,
so the helpers called inside it can still take it.
The regression test runs a reader against a thread that repeatedly fills and
clears the same plugin, and fails on the first torn snapshot. Verified by
removing only the lock: fails on 3 of 3 runs, passes on 3 of 3 with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(plugins): cap the per-plugin state transition history
PluginStateManager recorded every state transition in a per-plugin list
and never trimmed it. The only code that removed entries was
clear_state(), called solely from PluginManager.unload_plugin(), so a
plugin that stays loaded -- normal operation -- never released one.
The list is written on the hot scheduling path. Every update cycle
appends twice: _reserve_for_update() sets RUNNING and _finish() sets
ENABLED back again. At the default 60s update interval that is 2,880
entries per plugin per day, and nothing reads them -- get_state_info()
only takes their len(). Pure dead weight.
Measured against the unpatched class, ten plugins on a 60s interval:
sim uptime history entries heap growth
1 day 28,810 7.7 MB
7 days 201,610 53.9 MB
30 days 864,010 230.9 MB (still climbing)
With the cap it is flat at 2,000 entries / 0.5 MB from day one.
On a 1 GB board 231 MB of garbage is fatal on its own, and the failure
is not a clean OOM: once MemAvailable falls far enough fork() starts
returning ENOMEM, so sshd accepts connections and closes them before its
banner while the kernel still answers pings. The board looks like a
hardware fault and needs a power cycle. Same family as the ceilings
added in #464.
Retain the most recent 200 transitions per plugin in a deque and let the
rest age out. state_history_count is surfaced through the web API, so
the lifetime total is tracked separately rather than plateauing at the
cap. get_state_history() now returns a copy under the lock; it was
handing out the manager's own list, which a caller could mutate.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(plugins): copy history entries out, lock clear_state
Review follow-ups on the transition history.
get_state_history() copied only the outer list, so a caller holding a
returned transition could rewrite the manager's record of what happened
-- which contradicted the defensive-copy guarantee in its own docstring.
Copy each entry too. Every value in a transition is immutable, so a
shallow copy per entry is enough. test_get_state_history_entries_are_copies
pins it; without the change it fails with 'tampered' == 'enabled'.
clear_state() mutated five shared dicts without holding _lock, while
every other mutator takes it. A concurrent set_state() could interleave
and leave a plugin with history but no state. Drop the five as one unit.
This does not close the wider unload-vs-worker race, which lives in
PluginManager.unload_plugin() and predates this change: an update worker
still in flight can call set_state() after clear_state() returns and
recreate the entry. Serialising that needs the per-plugin lock held
across worker join in unload_plugin(), which is a separate change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(web): show available memory in Tools diagnostics
System Diagnostics reported memory as used-percent plus used/total GB.
Neither distinguishes a healthy board from one about to fail, because
page cache counts as used and is reclaimable on demand -- a Pi can read
70% used and be fine, or read the same and be minutes from trouble.
MemAvailable is the kernel's own estimate of what a new allocation can
actually obtain, and it is the number that tracked the failure on a 1GB
Pi 3B+: healthy running sat above 500MB, and the crash came at 73MB. By
that point fork() was failing, so sshd could not spawn a session and
systemd could not respawn the display, while the kernel carried on
answering pings at 0% loss. Used-percent gave no warning at any point on
the way there; available memory fell steadily for hours.
/api/v3/system/status now returns memory_available_mb from
psutil.virtual_memory().available, and Tools renders it as its own tile,
coloured against the thresholds that failure implies: red under 150MB,
amber under 300MB, green above.
The existing memory tile is left alone -- used/total is still what you
want when sizing a workload; this answers the different question of how
much room is left right now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): round available memory once, so the tile agrees with itself
The colour was classified from the raw value while the label was rounded,
and the API sends one decimal place. At the boundaries the two disagreed:
149.6 rendered as "150 MB" in red, and 299.6 as "300 MB" in amber -- each
contradicting the threshold its own colour claims to apply ("red under
150MB"). A reader checking the tile against the documented thresholds would
conclude the readout was broken.
Rounding once and using that number for both restores agreement. It moves
those two boundary cases up a band, which does not matter: the thresholds
come from a measured failure at 73MB, so which side of the line a spare
0.4MB falls on carries no information. The tile agreeing with itself does.
Null handling and the thresholds themselves are unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(memory): release fetched payloads once they have been delivered
BackgroundDataService kept the fetched body on the FetchResult it filed
in completed_requests, which is swept hourly and capped at 500 entries
by count. For a status record that costs nothing; for a season schedule
it costs a tenth of the board.
Measured on a 1GB Pi 3B+ with a 1-second RSS profile: the display
process sat at 404MB after plugin load, then stepped +21MB when NFL
fetched its season and +90MB when NCAA football fetched 946 games for
2026 -- and stayed at 494MB. Not a leak; a staircase that never came
down. When a later fetch landed while headroom was low, available memory
reached ~70MB, fork() began failing, and the board stopped being able to
start a process at all: sshd accepted connections and closed them before
its banner, systemd could not respawn the display, and the panel went
dark while the kernel carried on answering pings.
The cache-hit path was the worse of the two. It runs once per update
interval per sport, mints a fresh request_id each time, and files
whatever the cache returned. The memory tier is capped at 150 entries on
a 1GB board, so a miss re-parses the payload from disk into a genuinely
new object -- separate copies accumulating toward the 500-entry cap,
not shared references.
Releasing is safe: the payload is written to the cache under the
request's cache_key before the result is built, the callback is handed
the object directly, and consumers read it back from the cache
afterwards (the plugins' callbacks use it only in passing, to log a
count, before reading the cache). Nothing is lost -- it moves from RAM
to the disk cache that was already holding it.
Requests submitted without a callback keep their payload, since polling
get_result() is then the only way to collect it. That keeps the existing
contract, and the existing tests covering it, intact.
Not addressed here: max_workers=3 allows three concurrent fetches, so
three large parses can peak at once, and there is no in-flight dedupe by
cache_key -- a second submit for a key already being fetched starts a
second fetch. Both bound the transient peak rather than what stays
resident, and both are behaviour changes worth their own review.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(memory): file the cache-hit result before running its callback
Restores the original ordering. Releasing the payload after the callback
meant filing the result after it too, so a callback that queried
get_result() or is_request_complete() for its own request would not have
found it -- a behaviour change unrelated to the memory fix.
The dict holds a reference to the same object, so releasing after filing
still clears the payload.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test: wait for the payload release, not just the filing
The callback test waited on is_request_complete(), which goes true as soon
as the worker files the result in completed_requests. The worker then runs
the cleanup pass, then the callback, then releases the payload. Both of the
test's assertions therefore raced the worker: `seen` is populated by the
callback, and `data is None` only after the release that follows it.
It passes today because a one-line callback usually finishes inside the
20ms poll interval. Confirmed by making the callback sleep 0.4s: _wait()
returns with seen == {} and the payload still resident.
_wait_for_release() polls for the released payload instead. Release happens
strictly after the callback returns, so a released payload also means the
callback has finished and one wait covers both assertions. Verified against
the same 0.4s callback.
_wait() stays for the other three fetch-path tests, which assert only what
is already true when the result is filed -- the success flag, the error,
and the cache write that happened during the fetch itself. Its docstring
now says so, so the next reader picks the right one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(web): reject non-finite JSON numbers instead of raising
POST /api/v3/config/dim-schedule with {"dim_brightness": Infinity} answered
500. So did /api/v3/errors/clear with max_age_hours, and /api/v3/config/main
with multiplexing or row_address_type.
json.loads accepts Infinity/-Infinity/NaN by default -- they are not valid
JSON, but Python's parser emits them -- and Flask's get_json passes them
straight through. int(float('inf')) raises OverflowError, which is neither
ValueError nor TypeError, so validation blocks that carefully caught those let
it past and Flask turned it into a 500.
The status code was not the real damage. dim-schedule answered with
CONFIG_SAVE_FAILED and suggested "Check file permissions on config directory"
and "Check available disk space" for what was an invalid number. Every one of
these sites already had a correct 400 response written; they just never
reached it.
NaN already returned 400, because int(nan) raises ValueError. That is why this
only ever showed up for the infinities, and why it survived: the obvious test
case passes.
OverflowError is now caught alongside ValueError/TypeError at the 27 sites in
this file whose try block performs a numeric coercion. An AST sweep confirms
no int()/float() of request-derived data is left outside a block that catches
it.
Verified end to end through Flask's test client rather than by reasoning about
the parser: all four routes returned 500 before and 400 after.
Tests: five Infinity cases (which fail against the previous except tuples),
two NaN cases pinned so narrowing the tuple cannot quietly break them, and a
check that ordinary input is not rejected by the widened guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* test(web): assert 400 exactly, and prove valid input is accepted
Both review points were right, and the first is the failure mode this file
exists to catch.
Accepting any 4xx meant a 404 would have passed. Renaming one of these routes
would have left the test green while it tested nothing -- the same "looks like
coverage, points somewhere safe" shape that hid the composer injections. Now
asserts exactly 400.
Both infinity signs are exercised for every route. int() raises OverflowError
either way, but only +Infinity was in the original report, and a guard that
special-cased the sign would have passed a one-sided test.
The valid-input test previously asserted "not a 400", which did not show what
it claimed: the mocked save path fails for any input, so that assertion held
whether or not validation had accepted the value. It now gives load_config a
real dict and stubs _save_config_atomic, so the endpoint reaches its success
response and the test can assert 200 -- which only happens if the value passed
validation.
8 of the 11 checks fail with OverflowError removed from the except tuples.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
test_location_toggle asserted that ":42" -- a bare colon plus the record's
hardcoded lineno -- is absent from a line formatted with include_location=False.
But every formatted line starts with an HH:MM:SS.mmm timestamp, so ":42" also
matches the clock whenever the minute or the second is 42. The test fails for
roughly 3% of runs with nothing wrong:
2026-08-22 08:05:42.274 - INFO - test.logger - hello
^^^ matches ":42"
Assert on the whole "module.funcName:lineno" token the format string actually
emits ('%(module)s.%(funcName)s:%(lineno)d') instead of a fragment of it. That
cannot collide with a timestamp, and it checks the thing the test is named for.
Confirmed by formatting a record stamped 08:42:42 -- both minute and second
colliding: the old assertion fails, the new one passes.
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(display): retry a plugin that is enabled but failed to load
A plugin whose validate_config() returns False is treated as a hard load
failure. The API then reports enabled=true, loaded=false, error=null: the
plugin is simply absent, with nothing saying why. hockey-scoreboard sat in
that state on a live rig for four days.
The recovery path existed but could not be reached. _reconcile_enabled_plugins
computes to_add = desired - current, and a plugin that failed to load is never
in current, so it stays in to_add and would be retried. But the reconcile is
queued by _enabled_set_changed(), which compares only top-level `enabled`
flags -- and the edit that actually fixes such a plugin (enabling a league,
filling in an API key) is nested inside the plugin's own config section. No
top-level flag changes, so no reconcile is queued, and the save that should
have fixed it does nothing. Only toggling some unrelated plugin -- which does
change a top-level flag -- queues the global reconcile that recovers it.
Add a second gate: queue a reconcile when a discovered plugin is enabled in
config but absent from the running set.
It is deliberately narrow rather than "reconcile on any config change".
Reconcile calls discover_plugins(), a ~39-manifest filesystem scan, and it
runs on the render thread; doing that on every config save would trade this
bug for a frame hitch. Gating on plugin_manifests also keeps non-plugin
sections that carry their own `enabled` flag (schedule, display) from
queueing a reconcile they can never satisfy. In the steady state -- every
enabled plugin loaded -- the new check is False and costs nothing.
The same valid-but-unconfigured => hard-fail shape still exists in
text-display, youtube-stats, birdnet-go, ledmatrix-flights and
mqtt-notifications; this makes all of them recoverable without a restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(display): snapshot the plugin mappings under their locks
Addresses the review finding on the cross-thread reads.
_enabled_plugin_not_running runs on the config-watcher thread and read two
mappings the render thread mutates. Catching RuntimeError was not a fix: it
turned a torn read into a coin flip between an unnecessary discovery scan and
a missed retry, which is the bug this PR exists to remove.
Both reads are now snapshots taken under the lock that guards their writes:
- plugin_manifests via a new PluginManager.discovered_plugin_ids(), which
copies the ids while holding the existing _discovery_lock. Discovery
rebuilds that mapping entry by entry, so an unsynchronised reader can see
it half-populated.
- plugin_display_modes under a new controller lock, taken at the only two
sites that mutate it (_register_loaded_plugin / _unregister_plugin).
The locks are never nested -- each snapshot is taken and released before the
next -- so this cannot deadlock against discovery, which holds _discovery_lock
while it rebuilds.
No cost on the per-frame path. Both mutation sites run during reconcile, which
is rare, and every hot-path read of plugin_display_modes is on the render
thread itself, same thread as the writes, so those stay lock-free.
Tests: the accessor returns a snapshot rather than a live view, and actually
takes the discovery lock (proved from a second thread, since an RLock is
reentrant on the owning one) so a later refactor cannot quietly drop it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(display): consume the reconcile request before serving it
Addresses the second review finding: a lost update on
_pending_plugin_reconcile.
The flag was cleared after a successful reconcile. Reconcile has already read
its config by that point, so a config change arriving mid-flight set a flag
that the trailing clear then erased -- a request that was never served, and
the newest config never reconciled. That is the same "my save did nothing"
symptom this PR exists to remove, so leaving it would have undercut the fix.
Consume the request before running it instead, and re-arm only on a retryable
failure. A change that lands during reconcile now stays set and is picked up
on the next pass.
The per-frame read stays lock-free. It is a fast path that can only produce a
false negative -- the watcher setting the flag just after it is read is seen
on the next iteration -- never a false positive that loses a request. The lock
is taken only when a reconcile is actually pending or a config change arrives.
Extracted _service_pending_reconcile() so the sequence is testable rather than
buried in run()'s loop; the review asked for a regression test that invokes
the subscriber during reconciliation, which is not reachable otherwise.
Tests: 4 new, covering a request racing in mid-reconcile, the quiet success,
the retryable-failure re-arm, and not reconciling when nothing is pending.
Two of them fail against the previous clear-after-success semantics.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three review findings from #485 that I missed when addressing that PR;
it has since merged, so they land here.
1. Array-item secrets destroyed by any unrelated save (data loss).
remove_empty_secrets recursed into dicts but let a list fall through to
the scalar branch and kept it verbatim. Lists merge by *replacement*, so
the blanks the masked form posts back went straight over the stored
array:
stored [{"name":"a","token":"REAL-A"}, {"name":"b","token":"REAL-B"}]
posted [{"name":"a","token":""}, {"name":"b","token":""}]
merged [{"name":"a","token":""}, {"name":"b","token":""}]
-> both credentials gone
Same failure as the scalar api_key case fixed earlier, one container
deeper. Lists now prune element-wise, and a list with nothing real in it
is dropped so the stored one is left alone. Where one entry does change,
the new merge_secrets merges by index instead of replacing.
Two details the first attempt got wrong, both caught by existing tests:
- An emptied dict item must stay {}, not None. ConfigManager's
_strip_secrets_recursive treats a secrets list as *parallel* to the
regular one ({} = "item i has no secrets"); a None makes it stop
looking parallel, and it then drops the whole key from the main config
-- silently deleting the items' non-secret fields too.
- The incoming list's length wins. The regular config's list is
authoritative about how many items exist, so preserving surplus stored
entries would let the two fall out of step and make deleting an entry
impossible.
2. Submitted credentials written to the journal (security).
save_plugin_config logged `Full config: {plugin_config}` at INFO and
`Config that failed: {plugin_config}` at ERROR. Both run before
separate_secrets, so plugin_config still held the values just typed into
the form. Now keys only. Swept the rest of web_interface/ and src/ for
the same shape -- these were the only two.
3. Restart banner kept stale wording.
showRestartPending() cleared the stored custom text but left the DOM
element alone, so a config save could show the previous update's
message. The default is read back from the server-rendered copy rather
than duplicated in JS, so the template stays the one owner of the string.
Verified: 556 passed, 1 skipped across the web suite. Mutation-checked --
reverting api_v3 fails the logging guard and the array-merge test;
reverting either half of the secret_helpers change fails the unit tests.
New end-to-end coverage drives the real endpoint, not just the helpers.
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
test_returns_nothing_when_tmpdir_is_already_disk_backed asserted that
lm_disk_backed_tmpdir prints nothing when TMPDIR is already disk-backed,
and used pytest's tmp_path as the "disk-backed" directory:
# tmp_path is on the regular filesystem, so the default must be kept.
assert call("lm_disk_backed_tmpdir", env={"TMPDIR": str(tmp_path)}) == ""
That premise is false on the platform the helper was written for. Debian
13 mounts /tmp as tmpfs -- which is the entire reason lm_disk_backed_tmpdir
exists -- and pytest puts tmp_path under /tmp. So on the target platform
TMPDIR is memory-backed, the helper correctly answers /var/tmp, and the
test fails:
E AssertionError: assert '/var/tmp' == ''
The helper is right; the test was wrong. Reproduced on a box where
/tmp is tmpfs and / is ext4.
The test now looks for a directory whose backing store is actually disk
-- tmp_path, else a scratch dir under /var/tmp, else beside the library
-- using the same findmnt lookup the helper itself uses, and skips only
if no disk-backed directory exists anywhere. An earlier version of this
fix skipped whenever tmp_path was tmpfs, which made it skip on every
machine with a tmpfs /tmp; that is barely better than asserting the
wrong thing, so it now searches instead of giving up.
Verified: 31 passed, 0 skipped. Mutation-checked -- deleting the
"is the current TMPDIR memory-backed?" guard from lm_disk_backed_tmpdir
fails this test, so it still catches the regression it is there for.
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Vegas logged an FPS line at INFO every five seconds for the whole of
every run. Measured over two hours on a rig: 1410 samples, 98.5% of them
within 10% of target. The 1.5% that were not included a reading of
8.6fps against a target of 60 -- a real stall, completely invisible
inside 1389 lines reading "59.6". INFO is now reserved for a shortfall,
the recovery from one, and a slow heartbeat so a healthy marquee still
shows a pulse. Scroll-progress tracing drops to DEBUG for the same
reason: it runs for the whole of every scroll and is what you turn debug
on to watch.
Three review findings, all fixed here.
1. Per-frame timing used the wall clock (critical). The loop sleeps the
remainder of each frame budget:
frame_elapsed = <now> - frame_started
time.sleep(max(0.0, frame_interval - frame_elapsed))
These devices have no RTC, so the clock jumps by however wrong boot
time was when NTP first syncs. A backward step makes frame_elapsed
negative, `frame_interval - frame_elapsed` then exceeds the whole
budget, and the render loop stalls for the size of the correction. A
forward step instead inflates the p99 and worst-frame figures this
telemetry exists to report. Both per-frame timestamps are monotonic
now. start_time stays wall-clock: it is only used for the iteration
duration report, where a human-readable clock is the point.
2. FPS health state reset every iteration. last_fps_health_log and
was_degraded were locals of run_iteration(), which is called once per
cycle. Starting at 0.0 against a monotonic clock, `due` was true on
the first sample of every iteration, so the 300s heartbeat degenerated
into one report per cycle -- reintroducing the noise this change is
about. A recovery that crossed an iteration boundary was never
reported either, since was_degraded had already gone back to False.
Both now live on the coordinator and reset in start().
3. The degraded threshold read as an off-by-one. 90% of target is
deliberate -- a marquee jitters constantly, so "anything below target"
would report forever and mean nothing -- but nothing said so, leaving
55fps-against-60 looking like a missed case. The constant now states
the band and gives that exact example.
Also drops two soccer logo PNGs that a `git add -A` had swept into the
first commit. They are unreferenced, unrelated to frame-rate telemetry,
and 210KB.
Verified: each fix mutation-checked -- restoring the wall clock on either
per-frame timestamp, or making the health state local again, fails the
new tests. 566 passed across the vegas, coordinator and scroll suites.
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(sports): fetch odds for the games shown, not the whole window
SportsUpcoming.update() walked every upcoming game in the schedule
window and called _fetch_odds() on each one inside that collection
loop, narrowing to upcoming_games_to_show only afterwards. Each call is
a separate sequential ESPN request.
The comment sitting above it said odds were fetched "only for games that
will be displayed". The only narrowing it actually applied was
show_favorite_teams_only, which is not the default, so in the usual
configuration nothing narrowed it at all.
Measured on devpi, where the football plugin has the same shape:
467 odds requests in one 35s burst, 467 distinct events
315 NFL + 152 college-football -- roughly a whole season
plugin football-scoreboard operation timed out after 30.0s
The burst repeats each time the 1h odds TTL expires: 67 -> 327 -> 957 ->
1261 requests/hour across four consecutive hours. Between expiries the
cache works and the rate is zero, so this is a thundering herd on
expiry, not a caching failure.
The fetch now runs after selection, over team_games -- the list already
cut to upcoming_games_to_show. This mirrors the fix the football plugin
already carries; the shared base class never got it.
SportsLive is deliberately left as it is: it walks the raw event list
because it has to find which games are live, but only fetches odds for a
game that has already passed the is_live/is_halftime test, so its
fan-out is bounded by how many games are actually in progress. The test
pins that distinction rather than assuming it.
The test reads the AST rather than the source text, and asserts the full
set of call sites, so a new one has to be classified deliberately
instead of inheriting whichever behaviour it happens to land in. Writing
it that way is what turned up the SportsLive site, which I had missed.
Verified: reverting the fix fails the test with the offending iterable
named ("iterates over 'events'"). 525 passed, 9 skipped across the sports
and odds suites.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* test(sports): check the odds guard structurally, not by its text
Review caught that _guards_above() collected an `if` test even when the
call sat in that if's `else`, so moving _fetch_odds() into the else of
the is_live/is_halftime test would still pass -- while fetching odds for
exactly the non-live games the guard exists to exclude.
Verifying that turned up a wider hole in the same assertion. It matched
substrings of the *unparsed source*, so a negated condition satisfied it
too:
if not (details["is_live"] or details["is_halftime"]):
self._fetch_odds(details) # every non-live game
Both names still appear in that text, so `"is_live" in guards` held and
the test passed on code doing the opposite of what it claims to check.
The guard test is now structural. It walks the AST for an enclosing `if`
whose *body* (never its `else`) contains the call, and whose test
references both names without either sitting under a `not`.
Verified by mutation: fetching odds for non-live games now fails with
"does not sit in the true branch of a test requiring the game to be in
progress". Moving the call into the else of the *favourites* test still
passes, which is correct -- the game there is still live, so the
in-progress contract holds and the fan-out stays bounded by how many
games are actually in play.
525 passed, 9 skipped across the sports and odds suites.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(config): make the device location the default for plugin location fields
A user in Kansas City reported their radar centred on Dallas, TX with
nothing in config.json to explain it.
The radar is the `ledmatrix-weather` plugin's `radar` mode, and it centres
on the same coordinates as every other weather mode: `forecast_data`
lat/lon, geocoded from the plugin's own `location_city` /
`location_state` / `location_country`. Those ship with schema defaults of
Dallas / Texas / US. A user who never opened the weather plugin's config
form therefore has no `location_city` on disk, and `PluginManager` merges
the schema default in at load time — so the whole plugin (not just the
radar) silently runs on Dallas. Radar is just the only mode that draws a
recognisable map and gives the mismatch away.
Meanwhile the device-wide `location` block that General settings writes
was read by nothing at all, despite its own help text promising it was
"used for weather, sunrise/sunset, and other location-based content".
`SchemaManager.generate_default_config()` now substitutes the device
`location` into the three fully-namespaced `location_*` keys before
handing defaults back, so the promise holds:
- Only `location_city` / `location_state` / `location_country` are
substituted. A bare `state` key is left alone — `ledmatrix-elections`
uses it for a two-letter code, and rewriting it would break that plugin.
- A value the user saved on the plugin still wins: this replaces the
schema default, and `merge_with_defaults` puts user config on top.
- The substitution is applied on the way out of the defaults cache rather
than into it, so changing the device location takes effect immediately.
- No config manager, no `location` block, or an unreadable config all
fall back to the plugin's own schema defaults.
Every caller benefits: the plugin loader, the config form (which now
pre-fills the user's real city), config save, and reset-to-defaults.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GNLrSZ32FNKpHRaduKEJsg
* docs(web): name the exact plugin keys the device location seeds
Review follow-up. The General settings help text said the device location
was "the default for every plugin that asks for a city", which overstates
what the code does: only the fully-namespaced `location_city` /
`location_state` / `location_country` keys are substituted. A plugin with
a bare `city` key gets nothing — deliberately, since `ledmatrix-elections`
uses `state` for a two-letter code. The tips now name the exact keys.
Worth noting for anyone editing these: `ui.help_tip(...)` takes a
single-quoted Jinja string, so an apostrophe in the tip text has to be
escaped or written around. The wording here avoids them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GNLrSZ32FNKpHRaduKEJsg
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(plugins): say when discovery skips a directory
A plugin can be enabled in config, enabled in plugin state, present on disk
with a valid manifest and an importable entry point -- and simply absent from
the running process, with nothing anywhere to say why.
That is not hypothetical. hockey-scoreboard on a live rig is enabled in both
places, imports cleanly when loaded by hand, and is listed in the Vegas plugin
order, but is not among the 22 plugins the process actually holds. Establishing
even that much meant comparing cache-file mtimes to find it had last run three
days earlier. The journal had nothing, because discovery does not report what
it declines to load.
Two paths were silent. A directory with no manifest.json was skipped without
comment, which is defensible until it is the thing you are trying to explain.
Quieter still, a manifest that parsed but carried no "id" was read
successfully and then dropped on the floor -- no warning, no trace, and the
plugin simply does not exist as far as the rest of the system is concerned.
Both now log a warning naming the directory and the reason.
This does not explain the rig above; its manifest has an id. It makes the next
occurrence diagnosable from the journal instead of from file timestamps.
Reverting the change fails both tests. 65 plugin-system tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(plugins): warn once per directory, not once per scan
Self-review catch. Discovery runs on every web UI page load and every config
reconcile, so warning unconditionally about an unloadable directory would put
a line in the journal each time someone opened a page -- the same log-volume
problem this change exists to help diagnose.
The skip is now reported once per directory per process. The diagnostic value
is unchanged: the reason a plugin is missing still appears in the journal,
once, where before it appeared nowhere.
Test added covering five consecutive scans producing one warning.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(plugins): one unusable manifest no longer aborts the whole scan
json.load accepts any JSON value, so a manifest.json holding null, [],
"text" or 42 parses without complaint and then raises AttributeError on
manifest.get('id'). Nothing catches that: the outer handler around the
scan takes OSError and PermissionError only.
So a single malformed manifest did not skip that one directory -- it
aborted _scan_directory_for_plugins outright, and every other plugin on
disk, however healthy, silently failed to register. Reproduced with
three directories, the middle one holding `null`:
SCAN ABORTED -> AttributeError: 'NoneType' object has no attribute 'get'
the two valid plugins never registered
That is the same failure this PR set out to fix, in its most severe
form: a plugin enabled in config, enabled in plugin state, present on
disk, and absent from the running process with nothing to say why --
except here it takes every other plugin with it.
A manifest that is not a JSON object is now skipped like any other
unusable directory, named once, with what it actually was:
Skipping bad-null: its manifest.json is NoneType, not a JSON object
Skipping bad-list: its manifest.json is list, not a JSON object
scan returned: ['aaa-good', 'zzz-good']
Verified: removing the guard fails 6 of the 10 tests. Covers null, list,
string, int and bool, and asserts the healthy plugins either side of the
bad one still register.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(plugins): one bad metrics cache entry should not stop every plugin
Caught live on a rig: every plugin failing, once each, continuously.
ERROR - src.plugin_system.plugin_manager - plugin geochron operation failed:
ResourceMetrics.__init__() got an unexpected keyword argument
'consecutive_failures'
ERROR - ... plugin text-display operation failed: ...
ERROR - ... plugin news operation failed: ...
ERROR - ... plugin odds-ticker operation failed: ...
with /api/v3/health reporting plugin_system: not_initialized while the display
process itself kept running and updating the panel.
`consecutive_failures` is a plugin_health field, not a metrics one.
get_metrics() does ResourceMetrics(**cached), which raises TypeError on a
single unrecognised key, and that exception escapes into plugin_manager and is
reported per plugin. One malformed cache entry takes the whole plugin system
down.
How a health-shaped record came to sit under a plugin_metrics key on that
machine is not established, and I could not finish the diagnosis: the rig went
back into its EIO failure mode partway through -- SSH resetting pre-banner,
systemctl unexecutable -- while the web API kept answering from RAM. Checked
before that: the cache files on disk are correctly shaped and separate, and
CacheManager.get() returns the right record for each key, so it is not a live
key collision. A restored backup mixing two machines' caches is the likeliest
explanation, and that rig had one restored onto it.
Either way the loader should not be brittle enough for the answer to matter.
plugin_health already repairs its records field by field rather than trusting
what is on disk; this does the same. Known fields are kept, unknown ones are
dropped and named once in the log so a genuine schema change stays visible
rather than being silently discarded, and a non-mapping entry no longer raises.
Keeping the known fields matters: discarding the record wholesale would throw
away real call counts and timings because of an unrelated stray key.
Mutation-checked: restoring ResourceMetrics(**cached) fails 6 checks, dropping
the whole record fails the field-preservation check, and dropping unknown
fields silently fails the logging check. 28 tests pass across the resource
monitor and plugin health suites.
* perf(health): stop rewriting a health record on every healthy cycle
Every successful plugin update called record_success(), which persisted the
record unconditionally. In steady state the only fields that had changed were
total_successes and last_success_time -- a counter and a timestamp that
health_monitor surfaces for display and that nothing reads back after a
restart. Nothing alerts on the age of last_successful_update; it is carried in
the metrics dataclass and shown.
Measured on a rig running 24 plugins, all steady-state (0 consecutive
failures, circuit closed): a five-minute sample caught 22 health-file
rewrites, about 4.4 a minute or 6,300 a day. Each write is ~400 bytes through
cache_manager.set(), which writes a file per call, so each one costs a
filesystem block plus an ext4 journal write.
That lands on an SD card, where the unit of cost is an erase-block cycle
rather than the bytes involved, and where wear is what eventually kills the
card. Two cards have already failed on the other rig with the same
signature -- unreadable block device, EIO on exec, sshd unable to read its
host keys.
The circuit breaker still has to survive a restart, so the write is kept for
exactly the fields it is rebuilt from: consecutive_failures, circuit_state,
circuit_opened_time, half_open_start_time. A failure, a circuit opening and a
recovery are all still written the moment they happen. In-memory state is
updated every time either way, so the health API and web UI show what they
always did.
Tested: 100 healthy cycles now perform zero writes after the first, the
counters remain accurate in memory, and a failure, a recovery and a
half-open-to-closed transition each still reach disk. One test kills and
rebuilds the tracker from the cache to prove the breaker's state genuinely
survives what is no longer written.
Mutation-checked both ways: persisting unconditionally again fails the
steady-state test, and widening _DURABLE_FIELDS to include last_success_time
fails it too. The 46 existing health tests pass.
(cherry picked from commit 14abea2d24)
(cherry picked from commit 0f77bd2345)
* perf(vegas): trace the content path at DEBUG instead of INFO
plugin_adapter narrates every step of acquiring content from every plugin --
"Has get_vegas_content", "Native: calling get_vegas_content()", "Native
content returned None", "Has scroll_helper", per-item sizes -- once per plugin
per cycle, all at INFO.
Measured on a live rig: 13,408 log lines an hour, of which 13,366 were INFO
and 35 were WARNING. Roughly 223 lines a minute of string formatting on a Pi
that is also driving the panel, written through journald to the SD card, with
the 35 lines that actually indicate a problem buried among them.
Top repeated messages in that hour:
717 Scroll progress: elapsed=... total_scrolled=.../... px
399 [plugin] --> INCLUDED in Vegas scroll
323 [plugin] content_type=static, display_mode=fixed
195 [plugin] Has get_vegas_content: True
195 [plugin] Native: calling get_vegas_content()
168 [plugin] Native: get_vegas_content() returned None
168 [plugin] Native content returned None <- the same fact, twice
54 logger.info calls in plugin_adapter become logger.debug, along with the
per-frame scroll-progress line in scroll_helper. Together those are 3,174 of
the 13,408 lines an hour, a 23% cut, and the ~3,600 odds-manager lines are
addressed separately by ledmatrix-plugins#300.
Nothing is lost: the 19 warning/error/exception calls in the module are
untouched, so real failures still surface at their own level. This is a
logging-level change only -- no control flow, no behaviour.
One INFO call is deliberate and stays. The padding-strip message picks its
level at runtime (`logger.warning if (left and right) else logger.info`) and
test_vegas_plugin_adapter.py pins that choice; it survives because it is not a
direct logger.info call site. That test still passes.
Mutation-checked both ways: reintroducing a single INFO trace fails the guard,
and demoting the warning/error calls along with the trace fails a second guard
written for exactly that mistake. 537 vegas and scroll tests pass.
(cherry picked from commit e496d95dfe)
(cherry picked from commit 8d1e43c15a)
* fix(logging): give the journal the real severity of each line
Everything this process writes to stdout reaches the journal as PRIORITY=6,
whatever the Python level was, because journald has nothing else to go on.
Measured on a live rig over 24 hours:
lines containing " - ERROR - " 55
lines containing " - WARNING - " 13
journald PRIORITY recorded 6, for every one of them
So `journalctl -p err -u ledmatrix` returns nothing while errors are being
logged, and `-p warning` likewise. Triage falls back to grepping message text,
which is slower and unreliable: during this audit a search for "oom" matched
the radar logging "zoom=9" twenty-four times and briefly looked like the OOM
killer had been firing.
systemd reads a leading "<N>" on each stdout line and takes it as the priority
(sd-daemon(3)), so a formatter that prefixes one costs no dependency. Every
line of a multi-line record is tagged, not just the first -- the journal splits
them, and an untagged continuation reverts to the default, which would leave
the body of a traceback filed as informational while its first line was an
error.
Applied only when JOURNAL_STREAM is set, which systemd sets for services whose
output it captures. Run from a terminal, in the emulator or under pytest the
prefixes would be literal noise, and the file handler keeps the plain
formatter for the same reason.
Mutation-checked three ways: prefixing unconditionally fails the
outside-systemd test, prefixing only the first line fails the multi-line test,
and mapping ERROR to 6 fails the level mapping. 39 tests pass across the
logging suites.
(cherry picked from commit 780fca6365)
* fix(logging): let callers see through the journald formatter wrapper
CI caught what local testing could not: two existing tests in
test_logging_config.py assert that setup_logging() selected a
StructuredFormatter or a ContextualFormatter, by checking the console
handler's formatter directly. Wrapping that formatter to tag each line with
its syslog priority makes those assertions false.
They passed locally and failed on the runner because the wrapper is applied
only when JOURNAL_STREAM is set -- absent in a terminal, present in CI. An
environment-dependent break, which is the kind that gets shipped.
The wrapper now exposes the formatter it delegates to, and those two tests
look through it. They are about which formatter format_type selects, and that
behaviour is unchanged; only the object they have to reach for moved.
Verified both ways this time: 39 tests pass with JOURNAL_STREAM set and with
it unset.
* perf(plugins): stop rewriting a plugin's metrics file on every call
Plugin metrics were persisted to the cache inside monitor_call, so every
call by every plugin rewrote a small JSON file. Measured on a running rig:
one plugin's plugin_metrics file changed nine times a minute, with fourteen
such files active. Each is around 350 bytes, which on ext4 costs a 4KB block
plus a journal entry, so the cost is dominated by the write itself rather
than the payload. Cache writes accounted for essentially all of that device's
2.4 MB/min of SD traffic, on a card that wears out and has already failed
twice on the other rig.
Metrics cannot be de-duplicated the way health state can, because call_count
changes on every call and the timings usually do too. So they are rate-limited
instead: at most one write per plugin per 30 seconds.
The in-memory copy stays authoritative and exact -- a plugin's call_count is
still precise the instant after it runs. Only the cross-process snapshot the
web UI reads is delayed, and telemetry up to half a minute old is still a fair
description of a long-running plugin.
reset_metrics clears the throttle timestamp, so a reset is not left showing a
deleted key for the rest of the interval.
Extrapolating the sampled rate, this takes metric writes from roughly 126 a
minute to 28. Health persistence, the other half of the churn, is handled
separately in #475.
Verified by reverting the throttle: the churn test then reports 50 writes for
50 calls. 88 tests pass across resource monitor, plugin system and web API.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix: use a monotonic clock and only mark metrics persisted once written
Two review findings on the throttle, both right.
The interval compared wall-clock timestamps. These devices have no RTC, so
the clock jumps by however far off boot-time was the moment NTP first syncs
-- a forward jump would allow an early write, a backward one would stall the
snapshot well past the interval. time.monotonic() is not subject to either.
The timestamp was also recorded before cache_manager.set(). A set() that
raised would buy the next interval's silence without leaving a snapshot
behind, which is the one case where skipping the write is least affordable.
Recorded after the write lands instead, so a failure is retried on the next
call.
Verified by restoring the original ordering: the new test then reports one
write where two are expected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* Address all six review findings on the perf consolidation
CodeRabbit reported six; this is all six, checked against its own
"Actionable comments posted: 6" rather than against what I happened to
scroll past.
Three are real defects in the code:
1. _under_systemd() trusted the presence of JOURNAL_STREAM.
systemd publishes JOURNAL_STREAM as "dev:ino", and every child process
inherits it -- including one whose stdout has been redirected to a pipe
or a file. The variable outlives the descriptor it describes, so a
subprocess would decide it was talking to the journal and emit the "<N>"
priority prefixes as literal noise into that captured output. That is
exactly the noise the function exists to prevent. It now parses the pair
and fstats stdout, per systemd's own guidance, and returns False for
missing, malformed, mismatched, or unusable descriptors.
2. Cached metrics were not type-checked.
A dataclass does not enforce its annotations, so
ResourceMetrics(call_count="not a number") builds happily and only
fails later, deep inside monitor_call:
TypeError: can only concatenate str (not "int") to str
Values are now coerced to their declared type at load, where there is
still a cache key to name in the warning, and a value that cannot be
coerced starts the plugin fresh instead of arming a delayed failure.
A numeric string is accepted rather than discarded -- a JSON round-trip
can widen an int, and that is recoverable.
3. The first metrics snapshot was skipped for the first 30s of uptime.
_persist_metrics used 0.0 as the "never written" default. monotonic() is
time since boot on Linux and systemd starts this service at boot, so
`now - 0.0 < 30` was true for the first half-minute of every run: the
throttle swallowed the very first write, the one that matters most after
a restart. The sentinel is now None and the interval is only applied when
a previous write exists.
Three are tests that could pass without testing anything:
4. test_health_write_churn's fake cache stored by reference, so the
tracker kept mutating the object already in the store -- a record
could look persisted when no write had happened, which is precisely
what test_durable_state_survives_a_restart exists to detect. Both
directions now deep-copy, like a cache that serialises to a file.
Verified: disabling the one real cache write now fails three tests.
5. test_values_of_the_wrong_type_do_not_raise asserted only that a
dataclass had been constructed, which was true with the bad value
still in it. It now asserts the loaded metrics are usable -- the
field is numeric, and arithmetic on it does not raise -- across four
kinds of bad value.
6. test_vegas_log_volume counted "logger.error(" in the source text,
which also matches comments, docstrings and string literals --
including that module's own docstring, which names those levels. A
real error call could be demoted with the tally unmoved. It now walks
the AST, reusing the helper already in the file. Verified: demoting
all 18 warning/error/exception calls now fails the test.
Verified: every fix mutation-checked by reverting it and confirming the
matching test fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(web): stop /config/main handing out every credential it holds
The endpoint returned the raw config to anyone who could reach the port, and
this web interface has no authentication of any kind. An unauthenticated
request against a live rig returned:
github.api_token 40 chars
incoming-packages.ha_token 183 chars
jellyfin-now-playing.api_key 32 chars
ledmatrix-weather.api_key 32 chars
on-air.mqtt_password 8 chars
youtube.api_key 20 chars
youtube-stats.api_key 39 chars
A GitHub token and a Home Assistant long-lived token among them. Anything on
that LAN could read them.
The x-secret masking the plugin config endpoints use does not reach here: this
route never consults a schema, and core keys such as github.api_token have no
schema to carry the marker. Several of the fields above *are* tagged x-secret
in their plugin's schema and were still returned in full, which is what rules
out the schema route as the fix for this endpoint.
Credential-named fields are now blanked. Matching on the name is blunt, and
for a whole-config dump that is the right default: anything named like a
credential should not leave the process, and a new plugin adding a
differently-shaped secret is covered without anyone remembering to tag it.
Blanked rather than removed, and safe to blank: POST /config/main merges into
the freshly loaded config and writes only the keys it was given, so a client
that round-trips this response cannot erase a secret it never saw. The web API
suites confirm it -- 81 passing, unchanged.
On the test that matters: the first version of this suite exercised the two
helpers and nothing else, and reverting the single line that wires the
redactor into the route passed all thirty of them. A property asserted on a
helper is not a property asserted on the endpoint, and it is the endpoint that
is exposed to the network. The added test goes through the view function, and
it does fail on that revert.
This also corrects an earlier claim of mine. I reported that GET /api/v3/config
did not expose these values; that path 404s, so the check proved nothing. The
real route is /config/main and it exposed all of them.
* fix(web): stop an unrelated config edit from erasing a plugin's secret
Saving any field on a plugin's config form destroyed that plugin's stored
credential. On a rig with a weather API key, changing the city silently
emptied the key, and the plugin stopped working at the next fetch with no
indication why.
The path had no guard at any step. The config partial masks secrets before
rendering (pages_v3.py:740), so the browser posts them back blank; _parse_value
deliberately preserves "" for optional string fields; separate_secrets routes
that "" into secrets_config, which is a truthy dict; deep_merge writes it over
the stored value; save_raw_file_content persists it.
The blank does not even need the round-trip. merge_with_defaults injects the
schema's api_key default ("") into every save, so a client that never sends
the field at all still erases it. test_secret_count_message_counts_top_level_keys
was counting exactly that injected blank as a saved secret field -- the visible
edge of the bug, pinned as expected behaviour.
remove_empty_secrets() already existed for this, with seven unit tests and a
docstring describing this precise scenario ("clients will send those empty
strings back ... so that existing stored secrets are not overwritten with
blanks"). It was never wired into a call site. This wires it into both save
paths that merge into the secrets file.
A blank now means "unchanged" rather than "delete", which is the same contract
the helper's tests already describe. The cost is that a secret can no longer be
cleared by emptying the field; clearing needs its own affordance, since a
control that erases credentials as a side effect of ordinary edits is not one.
Verified by reverting the guard: the new round-trip test then fails with the
stored key read back as ''. 262 web tests pass with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop dumping the config and request headers to the journal
save_main_config logged its entire POST body and the full request headers at
ERROR on every save. The body is the configuration itself, and the headers
carry the session cookie, so a routine settings change wrote both to the
journal -- at a level that guarantees they survive any sane log filter.
The lines are leftover debug output: they say "DEBUG:" in the message while
calling logging.error, and they went through the root logger rather than the
module logger, bypassing the level configured for this blueprint.
Replaced with a debug-level line recording the shape of the request, which is
the part with diagnostic value. The local `import logging` went with them; it
shadowed a module-level import that was already there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop /config/secrets handing out every credential it holds
GET /api/v3/config/secrets returned config_secrets.json in full to anyone who
could reach the port, and this interface has no authentication. Probed against
a real rig it produced six populated credential fields: a 40-character GitHub
token, a 183-character Home Assistant token, and Jellyfin and weather API keys.
This is the second door onto the same credentials; #477 closes the first.
Masking the response alone would have been worse than the leak. The only
client fetches every secret, edits one field and posts all of them back, and
save_raw_file_content replaces the file wholesale -- so a masked GET followed
by the client's own save would write the mask over every credential the user
had not touched. That is why this was left open when the leak was found; it
needs both halves.
Read side: mask_all_secret_values(), which already existed for exactly this
endpoint -- its docstring names it -- and had never been wired to a call site.
It leaves empty values and YOUR_* placeholders alone, so a client can still
tell "set" from "not set" without being told the secret.
Write side: strip the echoed mask and blanks from the submission, then merge
onto what is stored, so "unchanged" means unchanged. The cost is that a secret
can no longer be cleared by blanking it; that wants its own affordance, since
a control that erases credentials as a side effect of saving an unrelated one
is not one.
Browser side: the token field is now left empty rather than filled from the
response. Filling it with the mask would have stored eight bullet characters
as the token the next time the user pressed Save, and filling it with the real
value is the thing being fixed. It reports whether a token is saved instead.
Verified end to end through the Flask endpoints, not the helpers. Reverting
the masking fails the leak tests; reverting the merge fails the preservation
tests; both halves are independently guarded. 278 web tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop reporting "no update" when the update check could not run
check-update returned update_available=False whenever git failed. The banner
is the only route to the update button, so a checkout git refuses to touch
looked exactly like a current one -- permanently, with nothing on screen to
act on and only a log line recording why.
The common cause is an install performed as root. scripts/install/one-shot-install.sh
clones into ${HOME}/LEDMatrix, never consults SUDO_USER, and contains no chown
at all, while its own error text suggests running the whole thing under sudo.
The result is a root-owned checkout, and on a rig this is what every git
command in it does:
fatal: detected dubious ownership in repository at '...'
including the fetch this endpoint runs. Verified on real hardware rather than
assumed.
A failed check now reports check_failed with a message the user can act on --
for dubious ownership, the chown that fixes it. The banner shows that message
instead of hiding itself, with the update button suppressed since updating
cannot work until the cause is fixed. The success path is untouched.
This does not fix the installer, which is the real cause; it stops the symptom
being invisible. The installer needs SUDO_USER handling and a chown, and its
suggestion to run as root should go.
Reverting the endpoint change fails four of the five new tests; the fifth
guards the success path and correctly does not move.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop the installer chmod stripping exec bits on every update
git tracks five scripts as mode 644 that first_time_install.sh then chmods to
755 (start_display.sh, stop_display.sh, the two install_*_service.sh, and
one-shot-install.sh does the same to first_time_install.sh). With
core.fileMode true, the default on Linux, git reports all five as modified
from then on, in files the user never touched.
The update button stashes local changes before pulling, so it is not blocked
by this. But it never pops that stash -- stash pop and stash apply appear
nowhere in the update flow -- so the mode change is stashed away and left
there, and the files revert:
=== file modes after the update button's stash ===
664 first_time_install.sh <- installer had made these 755
664 start_display.sh
664 stop_display.sh
664 scripts/install/install_service.sh
So every web-UI update silently strips the executable bit from the installer's
own scripts, and leaves a stash entry holding the difference. start_display.sh
and stop_display.sh stop working from the shell afterwards.
A manual `git pull --rebase` over SSH fails outright, since nothing stashes for
it: "cannot pull with rebase: You have unstaged changes". That is the likely
source of the reports, since plenty of people update that way.
Tracking the five as 755 -- what they should always have been, as the
installer chmodding them attests -- removes the spurious mode change
entirely: nothing to stash, nothing stripped, no stash entry, and manual
pulls work.
The pull also passes --autostash, for the case the code explicitly tolerates:
when the stash fails it logs a warning and pulls anyway, and that pull is what
then fails. Autostash also pops what it stashes, which the manual stash does
not.
Note that `git add -A` after `git update-index --chmod=+x` silently reverts
the index to the on-disk mode, so the modes here were set by chmodding the
files themselves.
Regression test asserts the five stay tracked executable; reverting any one
of them fails it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): ask for the restart that makes an update take effect
The update button pulls new code and restarts nothing. There is no systemctl,
restart, reload or reboot anywhere in the 172-line git_pull handler -- it
stashes, pulls, installs changed requirements, re-removes plugins the user had
uninstalled, and returns "Code updated successfully."
Meanwhile both services go on running the code they loaded at boot. So the
display keeps rendering the old build, the web interface keeps serving the old
build, and the user is told the update worked. Nothing on screen suggests
otherwise, and the next reboot is what actually applies it -- whenever that is.
The affordance for this already exists: the restart-pending banner, raised
after main-config saves, with a Restart Now button wired to the display
service. A code update is a stronger reason to show it than a config save is.
The response now reports restart_required, and applyUpdate raises the banner
with wording for a code update rather than a config save. The banner's message
became a parameter and is persisted next to the flag, since it outlives the
page that raised it.
restart_required is only true when the pull actually moved HEAD. "Already up
to date" is a success too, and prompting after a no-op would train users to
dismiss the prompt unread.
This covers the display service, which is what the Restart Now button drives
and what users notice. The web interface still picks up its own new code on
its next restart; restarting it from inside a request it is serving is a
larger change than this one.
Reverting the flag fails the test that a pull which moved HEAD asks for a
restart. 290 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* Mask list-shaped secrets element-wise, close two vacuous tests
mask_all_secret_values treated any non-empty list as a scalar, so a
secrets file holding
"accounts": [{"name": "a", "token": "tok-a"}, {...}]
came back as a single "••••••••". The caller could not see how many
entries existed, and the raw editor was handed a string where the file
holds an array. Recurse into lists in both _mask_value and _contains_mask.
Lists merge by replacement, not key-wise, so strip_masked_values now
drops a list outright if any element still carries the mask -- storing a
half-masked list would discard the untouched entries.
Two tests could pass without exercising what they claim to check:
- test_git_pull_resolution asserted modes only for paths git ls-files
returned. A renamed or deleted installer target is simply absent from
that output, so its mode was never checked. Assert every CHMODDED path
is tracked first.
- test_config_secrets_masking never checked the POST status. A 500
leaves the old file in place, which satisfies every assertion that
follows. Assert 200 before reading the file back.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* perf(systemd): cap glibc malloc arenas on the display service
Measured on a live rig 2.5 hours after start:
RSS 1030 MB
Private_Dirty 988 MB
anonymous mappings > 10 MB 23
largest 104, 79, 66, 63, 63 MB, on 64 MB-aligned addresses
threads 9
cores 3 -> glibc ceiling = 8 x 3 = 24 arenas
23 against a ceiling of 24, all 64 MB-aligned: these are glibc's per-thread
malloc arenas, not live objects. The data the process was actually holding
accounts for perhaps 15 MB -- the widest scroll strip observed was 35,746 x 64,
about 7 MB as RGB and the same again for its numpy mirror.
It is bloat rather than a leak: sampled four times over 135 seconds, RSS sat
between 990 and 1030 MB rather than climbing. glibc gives each allocating
thread its own arena, grows them to hold peak demand, and never gives them
back. A process that builds and drops large images across several threads is
exactly the shape that produces this.
The device had 59 MB free at the time, on 1845 MB total.
MALLOC_ARENA_MAX=2 trades a little allocator concurrency for that resident
memory. It is a tuning knob rather than a fix for a defect, so the rationale
and the measurements sit next to it in the unit file, and a test asserts they
stay there -- a bare environment variable invites removal by whoever meets it
next.
Two things this is NOT, both checked rather than assumed:
- Not an OOM problem today. A grep for "oom" in the service journal returned
24 matches, all of which were the radar logging zoom=9 and zoom=7. The kernel
OOM killer has not fired: dmesg has zero matches.
- Not currently capped by the unit's MemoryMax=85% either. That directive is in
this file but absent from the unit actually installed on the rig, which
reports MemoryMax=infinity, so nothing is enforcing a ceiling there.
The saving is unmeasured on hardware: applying it needs a service restart,
which blanks the panel, so that is the user's call rather than something to do
mid-audit. If p99 frame time regresses -- it sits at 18.4 ms against a 16.7 ms
budget for 60 FPS, so there is not much headroom -- raise the value rather than
remove it.
(cherry picked from commit 446207ffbc)
* test(systemd): pin the arena value instead of accepting a range
Review follow-up. The range check accepted 1, 3 and 4, so a change to 4 --
which hands most of the resident saving back -- passed a test whose whole
purpose is to notice that.
Pinned to the value the unit ships, in one named constant. Raising it is still
a legitimate response to a frame-time regression, but it should be a visible
edit here rather than silent drift, and the failure message says so.
Mutation-checked: changing the unit to 4 now fails.
(cherry picked from commit 73fff8d2d5)
* fix(startup): warn when an installed systemd unit has drifted from the repo's
Nothing re-applies systemd units after the first install. `git pull` -- which
is what the web UI's update button runs -- brings a new template into the
checkout, but no code in web_interface/ or src/ copies it to
/etc/systemd/system, and nothing anywhere runs `systemctl daemon-reload`. The
unit that actually runs is whatever first_time_install.sh wrote on day one.
So every hardening added to a unit is inert on existing installs, silently.
Measured on a live rig:
installed /etc/systemd/system/ledmatrix.service 2026-08-06
template systemd/ledmatrix.service 2026-08-19
contents differ
with the practical result that the MemoryMax=85% the repo's template specifies
was not being enforced at all -- `systemctl show` reported
MemoryMax=infinity. Anyone reading the template would reasonably believe the
service was capped.
Startup now compares each installed unit against its substituted template and
warns when they differ, naming install_service.sh as the remedy.
A warning, not an error, and deliberately not a silent rewrite: editing files
under /etc and restarting services is the installer's job, not something a
display process should do to a machine while it is booting. Making it fatal
would also brick every development checkout whose unit is legitimately absent
or hand-edited.
Comparison ignores comments, blank lines and ordering. The template carries
explanatory comments the installed copy will not have, and systemd does not
care about order within a section, so a literal comparison would warn on every
boot and be ignored within a week.
Mutation-checked three ways: never reporting drift fails, making it fatal
fails, and -- after the first attempt missed it -- comparing raw text now fails
too. That last gap is worth noting: the comment-insensitivity tests originally
exercised the helper directly, so a comparison that stopped calling the helper
passed them all. The test that catches it goes through _validate_systemd_units.
29 startup-validator tests pass.
(cherry picked from commit cf521bdfd8)
* fix(install): grant the sudo commands the captive portal actually runs
The installers write two allow-lists, /etc/sudoers.d/ledmatrix_web and
ledmatrix_wifi. Anything the code runs under sudo that is not in one of them
needs a password, which a service cannot supply, so the call fails.
Five commands were being run and none of them granted:
sysctl -w net.ipv4.ip_forward=0|1 wifi_manager.py:788, 883
nft add|delete table ip ledmatrix wifi_manager.py:835, 895
rfkill unblock wifi wifi_manager.py:1811
iptables ... wifi_manager.py:796, 813, 818, 871
mkdir -p .../dnsmasq-shared.d wifi_manager.py:922
Together these are the captive portal: unblock the radio, bring up the AP,
add the redirect, turn on forwarding, and undo all of it afterwards. Without
the grants a hardened install would associate clients to the access point and
then fail to route them.
Why it has gone unnoticed: a stock Raspberry Pi image ships
/etc/sudoers.d/010_pi-nopasswd granting the default user
<user> ALL=(ALL) NOPASSWD: ALL
which satisfies every one of these regardless of what the allow-lists say.
Confirmed on a live rig -- `sudo -n -l` permits sysctl there, and the blanket
rule is why. The allow-lists are effectively decorative on a default image and
only start mattering once that rule is removed or the service runs as another
user.
test_sudo_allowlist_covers_calls.py extracts every argv-style sudo call in
src/ and web_interface/ and asserts an installer grants it, so the next command
added without a rule fails here rather than on someone's hardened box.
Getting that test honest took three passes, each worth recording:
- Matching the literal "systemctl" against rules written as
`$SYSTEMCTL_PATH enable ...` reported six gaps that did not exist. Binary
path variables are now normalised before comparing.
- Scanning the whole installer let `NFT_PATH=$(command -v nft)` -- a variable
definition, not a grant -- satisfy the check on its own, so deleting the
actual nft rules still passed. Only NOPASSWD lines are considered now.
- `sudo -n <tool>` reported "-n" as the binary. sudo's own flags are skipped.
Each of the five grants is individually mutation-checked: removing any one
fails the suite.
(cherry picked from commit a372b43cd1)
* fix(install): drop the iptables wildcard, and pin each grant properly
Review follow-up. Two findings, both right, and the first is a hole I opened
myself.
`NOPASSWD: iptables *` is a root shell for the web user by another name.
`iptables --modprobe=/path/to/anything` runs that path as root, so a wildcard
grant on iptables escalates rather than restricts. I added that rule while
fixing a permissions gap, which is a worse outcome than the gap. It is gone,
and a test now fails on any trailing-wildcard grant to a tool that can execute
another program -- iptables, nft, tcpdump, find, awk, sed, perl, python, env.
The other finding: checking only the binary made the coverage test far weaker
than it looked. With `sysctl` present anywhere in the allow-list, deleting the
`net.ipv4.ip_forward=0` grant still passed -- and the portal would then be
unable to restore forwarding on teardown. Each required command is now matched
in full, and each is mutation-checked individually, including that exact
single-line case.
Scope pulled in deliberately. The first version of this test tried to assert
that *every* sudo call in the codebase is granted. Run honestly, it showed the
portal also runs iptables, nft, `ip addr`, `ip link` and `cp` with arguments
built at runtime -- an interface name, a port. Those cannot be granted safely
in a sudoers file: the rule needs a trailing wildcard, and that is the
escalation above. Closing that half needs a privileged helper that builds the
rules itself and takes only an interface and a port, granted the way
safe_plugin_rm.sh already is. That is a design decision, not a one-line grant,
so the test now pins the four commands this change actually grants and the
docstring says plainly what it does not cover.
Better a narrow test that is true than a broad one that is not.
(cherry picked from commit 500cfbc9f4)
* fix(install): pin PATH, keep unit order, and tighten the sudoers assertions
Three review findings, all correct.
The installer resolved binaries through an inherited PATH and wrote whatever
it found into sudoers as NOPASSWD grants. first_time_install.sh re-execs
itself with `sudo -E`, which preserves the caller's environment, so a writable
directory early in PATH turned a compromise of the low-privilege web user into
permanent root -- via a file the installer itself wrote. PATH is now pinned to
the system directories before anything is resolved, and every resolved binary
must be root-owned and unwritable by anyone else before it reaches the
sudoers file.
_unit_body() sorted a unit's lines before comparing. Order is not noise in a
systemd unit: repeated ExecStartPre=/ExecStartPost= run in the order they
appear, and a directive that moves between [Unit], [Service] and [Install]
means something different where it lands. The drift check reported no drift
for units that had genuinely changed. Order is preserved now.
Two of that check's own tests asserted the wrong thing --
test_reordered_directives_are_not_drift said so in its name -- and are
inverted, with a second covering a directive moved between sections. The
cosmetic-difference test now varies comments, blank lines and indentation,
which is what the installer actually drops, rather than reversing the file.
The sudoers assertions matched command prefixes, so
`sysctl -w net.ipv4.ip_forward=0 *` satisfied the requirement while granting
the caller arbitrary trailing arguments as root. They are exact now. The
wildcard check also normalises ${NFT_PATH} the same way as $NFT_PATH; the
brace is not a word boundary, so that spelling was skipped entirely.
Verified by reintroducing each: a widened required grant fails the exact
match, `${NFT_PATH} *` fails the wildcard check, and require_trusted_binary
refuses a non-root-owned, world-writable, or missing binary.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review was right on all three counts, and the first is the one that matters:
scripts/install/configure_web_sudo.sh writes the same three wildcard
journalctl rules as first_time_install.sh and none of them carried NOEXEC. So
this PR closed the pager escape on one installer path and left it open on the
other, which is close to no fix at all -- a rig configured through that script
still hands out a root shell via less's "!command".
The test could not have caught it, for two independent reasons. INSTALLERS
did not list the file. And even listed, _grant_lines() kept the raw source
line: that installer echoes its rules, so each one ends in a quote rather
than the wildcard, and the trailing-* check skipped every one of them. Either
alone would have hidden it.
Both fixed: the file is covered, and an echoed rule is unwrapped to the
sudoers line it actually emits.
The selector test now covers -t ledmatrix as well. It asserted only the two
-u forms, so deleting the -t rule would have passed.
Verified by removing NOEXEC again from the secondary installer: four of the
six tests fail, where before the suite passed with the vulnerability present.
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
journalctl starts a pager when its output is a terminal, and from less a "!sh"
is a shell with whatever privileges journalctl was given. That is the standard
journalctl escalation, and these rules end in a wildcard:
<user> ALL=(ALL) NOPASSWD: /usr/bin/journalctl -u ledmatrix *
Nothing this project runs needs the pager -- both call sites pass --no-pager,
in web_interface/app.py and api_v3.py. But a sudoers rule cannot require a flag
that sits in the middle of a command line, and reasoning about what a trailing
wildcard does and does not admit is exactly the kind of subtlety that produces
a hole. sudo's NOEXEC tag stops the command executing another program at all,
which closes it without depending on that reasoning.
NOEXEC works by LD_PRELOAD, so it applies to dynamically linked binaries.
Checked on the target hardware: journalctl there is dynamically linked. The
generated rules were run through `visudo -c` -- parsed OK.
Found while auditing the pre-existing wildcard grants, prompted by review
catching a far worse one I had added myself in the same area: `iptables *`,
where --modprobe runs an arbitrary path as root.
Reachability, stated plainly: on a stock Raspberry Pi image none of this
matters, because 010_pi-nopasswd already grants the default user
`ALL=(ALL) NOPASSWD: ALL`. It matters on a hardened install, or where the
service runs as a user without that blanket rule.
Two mutation checks: dropping NOEXEC from a rule fails, and deleting the rules
rather than tagging them fails too -- that second one matters, since "make the
test pass" and "remove the feature" would otherwise look the same.
* test(sync): cover the display sync protocol, and fix what that surfaced
DisplaySyncManager had no tests at all — it appeared in the suite only as
a MagicMock() stand-in, so none of its framing, handshake, or socket
handling was ever exercised. Writing that coverage surfaced three bugs.
Both receive loops caught the generic Exception and immediately retried.
A socket left in a bad state raises on every call, so the thread spun at
100% CPU logging the same line; the reverted-code run of the new
regression test takes 24 seconds where the fixed one takes 0.2. Both now
back off briefly before retrying.
The follower dispatched on `data[:8] == _RAW_MAGIC or len(data) > 512`.
That size threshold is not part of either wire format: a control message
over 512 bytes — a hello_ack carrying a long incompatibility error, for
instance — went to the image decoder and was dropped, and a raw frame
under 512 bytes went to the JSON parser. Both formats are already
self-describing, so dispatch on the magic prefix and treat a JSON parse
failure as the legacy unmarked PNG, with the shared frame bookkeeping
factored into _handle_received_frame().
_oversized_frame_warned was created on first use through
getattr(self, ..., False) rather than in __init__, alone among the
instance attributes.
75 tests: role parsing, the hello compatibility matrix, watchdog
timeouts, both receive loops, the TCP image server's length and
dimension caps and decompression-bomb guard, status shape per role, and
one end-to-end loopback handshake so the wire format is exercised for
real and not only against mocks.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(logos): cover LogoHelper, and stop bad downloads poisoning the cache
Nothing in test/ referenced logo_helper.py, so its caching, resizing and
download-fallback logic was entirely unexercised. Two bugs surfaced.
_download_logo wrote response.content to disk with no size limit and no
check that the bytes were an image. A logo URL is remote input, so the
response chose how much went into the assets directory; worse, an
undecodable one stayed there, and because load_logo() only reports the
decode failure and returns None, every later call re-read the same
corrupt file. The download path never retried, so a single bad response
made a logo permanently blank rather than falling back to the
placeholder. Cap the response, verify it decodes, and delete it if not,
which lets the existing fallback in load_logo_with_download do its job.
get_cache_stats() divided by self.cache_size with no guard, so a helper
built with cache_size=0 raised ZeroDivisionError from what is only a
stats call.
37 tests: size-qualified cache keys, LRU eviction and refresh, the four
load_logo_with_download paths, download permissions and timeout,
placeholder generation, and the abbreviation normalizer — including a
test pinning its deliberate divergence from
LogoDownloader.normalize_abbreviation, since logo filenames on existing
installs depend on both behaviors staying put.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(web): cover the error and response builders, and stop dropping empty values
errors.py and error_handler.py's response builders had no direct tests,
though every API response passes through them. Two bugs surfaced.
WebInterfaceError set suggested_fixes with `or`, so a caller passing []
to mean "I have no suggestions for this one" got the default list
instead. Only None should fall back.
create_success_response gated `data` on `is not None` but `message` and
`metadata` on truthiness, so an explicitly-passed "" or {} vanished from
the response while 0 and False survived — the response shape depended on
the value. api_helpers.success_response() then re-gated metadata the same
way, which is the path every api_v3 endpoint actually calls, so fixing
only the inner function would have changed nothing observable. Both now
use `is not None`.
That wrapper also merged request timing into the caller's own metadata
dict in place. A caller reusing a dict across requests would accumulate
previous responses' timings; it now copies before adding.
79 tests: category inference for every error code, mapped vs fallback
suggestions, the JSON shape including which keys are omitted when empty,
exception-to-code inference, and the success/error builders end to end.
Two behaviours are pinned as deliberate rather than fixed: an empty
context stays out of the response body, and from_exception's `message`
is the fixed per-code string, never the raw exception text.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(web): cover the input validators, and close three holes in them
validators.py had tests for dedup_unique_arrays only; the other eight
functions were untested. Three bugs surfaced.
validate_image_url checked for '..' only inside its relative-path
branch, so http://host/../secret passed validation while /../secret was
rejected — the traversal check now runs before the branch split, which
is where a safety check on the whole URL belongs.
validate_file_upload lowercased the uploaded filename's extension but
compared it against the caller's list verbatim, so allowed_extensions of
['.TTF'] rejected every valid .ttf file. Both sides are lowercased now.
The one in-tree caller passes lowercase already, so this only widens what
future callers can hand it.
validate_numeric_range accepted True and False, because bool subclasses
int; a boolean then compared as 1 or 0 against the range and validated
cleanly. Excluded explicitly, matching how base_plugin.py already handles
the same trap for display_duration.
84 tests. Two behaviours are pinned rather than changed:
sanitize_plugin_config deliberately does not HTML-escape strings, since
escaping at this layer would store the escaped form in config.json — the
docstring said "prevent injection", which read as a promise it does not
keep, and now says what it actually does. validate_font_awesome_class's
second 'fa-' check is unreachable behind its own regex; harmless, so
characterized rather than removed.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(api): cover wifi and registry endpoints, and fix bodyless POSTs
The /wifi/* routes drive the host's real networking and the registry
routes reach GitHub, and neither had endpoint-level tests. Covering them
surfaced a bug affecting six endpoints.
Six handlers read their body as `request.get_json() or {}`. The `or {}`
says every field is optional and a missing body should fall back to
defaults — but get_json() without silent=True raises UnsupportedMediaType
when there is no JSON Content-Type, and it raises before `or {}` is ever
evaluated. Each handler's catch-all then reported that as a 500. So
POSTing with no body — what curl sends by default, and what a fetch()
without options sends — failed on /plugins/store/refresh,
/display/on-demand/start, /plugins/config/reset,
/plugins/of-the-day/json/delete, /plugins/{id}/limits and
/plugins/authenticate/spotify. The shipped UI always sends a JSON object,
which is why this stayed hidden.
All six now use silent=True. test_api_v3_optional_body.py covers the
affected endpoints and adds a source check, since the combination of
`or <default>` with a non-silent read is self-contradictory wherever it
appears and is easier to catch by inspection than by exercising each
endpoint by hand.
Also adds test/_api_v3_test_helpers.py: the blueprint holds its managers
on a module-level singleton rather than in Flask app state, so a test
that mocks them leaks into every later test unless the originals are
restored. The existing _make_client() does this for unittest classes;
this is the pytest-fixture equivalent, for the five suites still to come.
69 endpoint tests: connect/disconnect/AP/radio including the string-aware
boolean coercion these endpoints deliberately use, the radio's
lockout-refusal path, registry refresh and fetch-from-URL, and a guard
that WiFiManager is never constructed for real.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(api): cover the music auth endpoints, and always clean up the wrapper
The Spotify step-2 handler writes a Python wrapper script to a temp file
with the user's redirect URL embedded in its source, then executes it.
That is the most dangerous shape in the blueprint and had no tests.
The wrapper was deleted in the success/failure branch and again in the
TimeoutExpired handler. Any other failure from subprocess.run — no
interpreter, a fork failure, an interrupted call — reached neither, and
left a world-readable temp file containing the user's redirect URL on
disk. Cleanup moves to a finally block, which is what "delete this
whatever happens" should have been from the start.
The injection tests are the point of this file. Eight adversarial
redirect URLs (embedded quotes, backslashes, newlines, triple quotes, a
full `"; import os; os.system("id"); "`) are each pushed through the
endpoint and the generated wrapper is parsed with ast: it must still be
valid Python, the URL must still be a single string literal bound to
redirect_url, and no os.system call may appear anywhere in the tree.
json.dumps holds up, but nothing was checking that it does.
40 tests. Also pins that the two endpoints are not symmetrical despite
the matching names — only Spotify has a two-step flow and a wrapper; YTM
runs its script directly — so a later change does not "restore" a parity
that was never there.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(api): cover the credentials upload, and stop it hoarding secrets
The endpoint that receives the user's Google OAuth credentials file had
no tests. Two bugs surfaced.
The OAuth-shape check ran inside `except Exception: pass`. A JSON
document that parses but is not an object — a bare 42, true, null, a
list — makes `'installed' not in creds_data` raise TypeError, which the
bare except swallowed, and the file was then written out as
credentials.json regardless. The check now decides the outcome instead
of being advisory, so anything not credentials-shaped is refused up
front rather than failing later inside the calendar plugin.
Every overwrite copies the old file to credentials.json.backup.<ts> and
nothing removed them, so a user who re-uploaded ten times had ten
complete sets of OAuth client credentials sitting in the plugin
directory, indefinitely. Keep the newest five. Pruning is housekeeping,
so a backup that cannot be removed logs and leaves the upload alone.
27 tests: size and extension limits, malformed JSON, the shape check,
0600 permissions on the written file, backup-on-overwrite, and pruning
including the repeated-upload case that stays bounded.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(api): cover the install endpoints, and make 14 dead guards reachable
/plugins/install and /plugins/install-from-url were tested only at the
PluginStoreManager layer, so the route logic — the queue-versus-direct
branch, schema invalidation, discovery, state and history recording — was
unexercised.
Covering them surfaced the wider form of the body-parsing bug fixed for
the `or {}` handlers in the previous commit. Fourteen handlers read
`data = request.get_json()` and immediately guard with `if not data:
return 400, 'No data provided'`. That guard cannot run: get_json()
without silent=True raises UnsupportedMediaType for a request with no
JSON body, so the catch-all answered 500 "an error occurred; see logs
for details" where the handler plainly meant to answer 400 and say
which field was missing. Every one of these endpoints told a caller who
simply forgot the body to go read the server logs.
All fourteen now use silent=True, so the guard each author already wrote
is the one that runs. This covers /config/raw/main and /config/raw/secrets
among them, whose own bodyless case had the same shape.
The two remaining bare reads are left alone: neither declares what a
missing body should do, so there is no stated intent to honour.
31 install tests plus 17 body tests. The install pair is checked against
each other rather than only individually — the same install logic is
written twice, once in the queue callback and once in the fallback, so
the tests assert both produce identical schema, discovery, state and
history effects. They agree today; the one difference is the success
message wording, which is characterized rather than changed.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(api): cover the raw config write endpoints
/config/raw/main and /config/raw/secrets write whatever JSON they are
given straight to config.json and config_secrets.json, bypassing the
secret-separation path the rest of the config surface goes through. Given
how carefully that surface keeps secrets out of config.json, the pair
that skips it was worth pinning precisely. Backed by a real
ConfigManager over tmp_path, so the assertions are against files on disk.
20 tests covering both routes: what lands in which file, that a raw
secrets write never touches config.json and vice versa, the GitHub token
reload, the uninitialized-manager and empty-body branches, and the
ConfigError path that carries config_path through to the response.
The bypass itself is pinned as intentional rather than changed — these
back the raw JSON editor, so writing the body verbatim is the feature.
The test says so explicitly, because the failure mode is someone later
routing plugin config through here as a convenience and silently losing
secret separation.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(api): cover backup restore and path containment, and fix restore scope
Restore is the most destructive thing the web interface can do — it
overwrites config, secrets, WiFi settings and fonts, then reinstalls
plugins — and neither it nor the file routes beside it had tests.
A malformed `options` field fell back to {}. Every RestoreOptions flag
defaults to True, so a caller who asked for a narrow restore and
mis-serialized the request got a full one instead, secrets included, and
was told it succeeded. Valid JSON that is not an object was worse:
`"null"` or `"[1,2]"` reached .get() on a non-dict and raised, so the
request died as a generic 500. Both are now refused with a 400 that says
what was wrong, and restore_backup is never reached.
The other file routes take a filename straight out of the URL and turn it
into a path — one to read, one to unlink. _safe_backup_path is the only
thing keeping those inside the export directory, and it was untested. No
bypass was found; the thirteen traversal shapes are pinned so a later
loosening of that pattern has to argue with something. The delete route's
by-name enumeration is covered too, including that a directory sharing a
backup's name is not removed.
84 tests. Two behaviours are pinned as intentional: a failed plugin
reinstall turns the whole restore into an error even though file
restoration succeeded, and omitting `options` entirely still means
restore everything — that is the documented default, and it is only the
mis-serialized case that was wrong.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* ci: raise coverage floor to 52%
Measured 54.45% after the Tier 1 and Tier 2 suites, up from 50%. Keeping
the same two points of headroom the 45 -> 48 ratchet used.
The modules this branch set out to cover: sync_manager 0 -> 97%,
logo_helper 0 -> 98%, errors and error_handler 0 -> 100%, validators
0 -> 97%. api_v3 moved less in percentage terms because it is 4,341
statements, but the endpoints covered are the destructive and
credential-handling ones.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(sync): probe for a free port on loopback, not every interface
CodeQL flagged the ephemeral-port probe in the handshake test for
binding to all interfaces. The probe only needs a free port number, so
loopback is both sufficient and correct — a test should not open a port
to the network to discover one.
The manager under test still binds to all interfaces, which is
deliberate and already marked nosec: a follower has to receive the
leader's UDP broadcast.
Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: bound the logo download, and stop malformed input reading as a fault
Review findings on the coverage branch.
The download size cap I added checked len(response.content), which has
already buffered the whole body -- it stopped the bytes reaching disk but
not memory, which was the point. A server that omits Content-Length and
never stops sending would still exhaust the process. Stream it instead,
counting as it arrives, into a sibling .part file that is replaced over
the target only once it decodes. A transfer that dies midway now leaves
nothing behind rather than a truncated logo for load_logo() to cache.
The follower's control-message handler caught three exception types, but
two reachable UDP payloads raise others: a bare JSON scalar makes
msg.get() raise AttributeError, and an "sx" carrying a non-numeric x
raises ValueError or TypeError from float(). Those escaped to the outer
handler, skipping the legacy-PNG fallback and -- since this branch added
a backoff there -- charging one malformed packet a 0.1s stall on the
receive path. The legacy-PNG path also decoded without the dimension cap
its TCP counterpart applies, so a crafted 65KB frame could force a large
allocation on the render thread; both paths now share one constant.
Three repo_url handlers called .strip() on client input without checking
it was a string, so {"repo_url": 12345} answered 500. The credentials
upload parsed the same file twice, the second time inside a bare except
that a preceding parse had already made unreachable. And both raw-config
handlers kept a json.JSONDecodeError arm that get_json(silent=True) had
turned into dead code, collapsing "sent something unparseable" into "sent
nothing" -- they now say which.
Two of the new tests were not testing what they claimed. The pruning
round-trip wrote ten backups inside one second, so all ten landed on the
same int(time.time()) filename and overwrote each other; it never reached
the limit it asserted. And the sync clock helper patched attributes on the
stdlib time module, freezing time process-wide for every daemon thread
earlier tests had left running.
Full suite: 3352 passed, coverage 54%.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(sync): probe broadcast by sending, not by listening
The broadcast check added in the previous commit bound INADDR_ANY to
receive its own probe datagram, and the free-port probe did the same to
pick a port. CodeQL flagged both, correctly: a test suite has no reason
to open a socket the whole network can reach.
Sending is enough for what the probe is actually for. An environment
that refuses broadcast raises on sendto, which is the case that occurs
in sandboxes and is the one worth skipping over; confirming delivery
would have required the listening socket. A network that accepts the
send and silently drops it still reaches the assertion, exactly as it
did before either commit. The port probe binds loopback -- it only needs
a number, and the manager's own bind is the one that has to succeed, with
the retry loop already covering a port taken elsewhere.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: keep callback faults out of the frame-decode fallback
Review follow-up on the previous two commits.
Widening the control-message except tuple put the callback dispatch
inside it, so an _on_new_cycle() that raised ValueError, TypeError or
AttributeError sent a perfectly good control packet to the legacy PNG
decoder -- which reported it as an image decode error and buried the
real fault. Split the two: whether the payload parses as JSON decides
frame vs control message, a second guard covers reading the fields of an
attacker-shaped body, and the callback fires outside both. It still
cannot kill the receive thread; the loop's own handler catches it, and
now says what actually went wrong.
The logo download's temp file was a fixed "<name>.part". Two plugins
asking for the same logo at once would interleave writes into it,
publish the mixture, or delete each other's partial. mkstemp gives each
download its own name in the same directory, so os.replace stays atomic.
Its descriptor is adopted by fdopen before the request runs, since a
request that raises before the write would otherwise leak the fd --
quietly, because load_logo_with_download swallows that.
Two test fixes: the oversized-frame test replaced PIL.Image.open
process-wide, the same hazard the clock helper documents, and Ruff B007
on an unused loop variable.
Full suite: 3355 passed, coverage 54%.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test(sync): cover the announce loop, and reject non-finite scroll positions
Three review findings from the follower receive path.
Non-finite scroll x reached follower rendering. json.loads accepts the
bare NaN/Infinity literals and float() accepts them as strings, so
"x": NaN arrived as a real float and was stored verbatim. NaN loses
every comparison the scroll code makes, so a follower given one sits on
a position it can never advance past. It now raises through the existing
malformed-control-message guard, which logs and drops the packet and
leaves the last good position in place.
_broadcast_available() only proves the host accepts sendto() for a
broadcast; a network that accepts the send and drops the packet would
let TestRealSocketHandshake run to its five-second deadline and fail on
assertions the code did not break. The deadline now distinguishes the
two: if not one packet crossed in either direction, that is the
environment, and the test skips rather than reporting a protocol
failure.
That skip could hide a real regression in the announcing side, so
TestFollowerAnnounceLoop covers it on mock sockets, where no network is
involved and nothing can skip: hello carries this display's hardware
config and goes to the broadcast address, heartbeats follow, an empty
hardware config falls back to 32x64x1, hello is not resent before its
interval, and a send failure is swallowed rather than killing the loop.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
---------
Co-authored-by: Claude <noreply@anthropic.com>
The enum that lets a checkbox group draw its options is also what validates
the saved value. When an option goes away -- a league retires a team code, a
schema drops a choice -- a config still holding the old value has no checkbox
to render for it, but the value stayed in the hidden _data input anyway:
that input is seeded from the stored array and only rebuilt by
updateCheckboxGroupData() on change.
So the stale value was posted back on every save the user did not happen to
touch that widget for. The schema rejected it and the save endpoint returned
400 CONFIG_VALIDATION_FAILED, which blocks editing *any* field on that
plugin until the user works out which invisible entry is at fault -- with
nothing on screen naming it, because the offending value is precisely the one
with no checkbox.
Runtime was never affected: load_plugin() treats schema violations as
warn/degrade, and a retired code already matched nothing. Only the web UI
blocked.
Values not in the enum are now dropped before the hidden input is seeded, and
listed above the group so the selection is not lost silently. Only when the
widget has options -- an empty enum means there is nothing to check against,
and filtering on it would wipe the field.
This is not hypothetical. ledmatrix-plugins #212 ("correct team abbreviations
so config save no longer 400s") and #234 (removed the retired NHL code UTA
from a picker across four plugins) are both this failure mode, fixed one
league at a time. Nine shipped plugins use checkbox-group today; all of them
get the fix.
Tested by rendering the checkbox-group block lifted out of the shipped
template, following test_enum_option_labels.py, so the tests exercise the
production expression rather than a copy. Mutation-checked: removing the
filter fails 2 tests, filtering unconditionally fails the empty-enum test,
and dropping the notice fails the one asserting the value is named.
* fix(web): verify the onboarding timezone step, don't compare it to the default
The Getting Started card's timezone step ticked when the saved timezone
differed from the value config.template.json ships (America/New_York),
OR-ed with the saved city differing from Tampa. Both halves were wrong.
"Differs from the default" answers "did somebody edit this?", but what the
checklist needs to know is whether the value is right. A user genuinely in
America/New_York could never satisfy it, so the card nagged forever with
four of five steps done -- the case that prompted this, on a panel whose
timezone was correct all along.
The city half was worse than useless: the saved city says nothing about
whether the timezone is set, and because the two were OR-ed, saving a city
ticked the step off with the timezone still wrong. That is the direction
that actually breaks displays, since event times then render in the wrong
zone.
The browser already knows its own zone, so compare against that. No new
persisted state, no network, and it catches the reverse case the old test
got backwards: a panel still set to the old zone after a move now stays
unticked, where before it ticked the moment the value stopped being the
default. Zones are compared by the wall-clock time they produce for one
instant rather than by identifier, so aliases (Asia/Calcutta vs
Asia/Kolkata, Europe/Kiev vs Europe/Kyiv) don't read as a mismatch. When
they genuinely differ the step names the browser's zone, so an unticked box
says why. Configs with no timezone, an unparseable zone, or a browser
without Intl leave the step open for the existing manual tick.
The step still deep-links to the General tab, and the location value stays
visible in its label -- it just no longer votes on whether the timezone is
configured.
Tests render the partial across configured zones and both cities: the step
never pre-ticks server-side, carries the configured zone for the client to
check, is unmoved by the city, and the panel-size step still resolves
server-side. Reverting the template fails 9 of the 11.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): compare zones on fields Intl has always had
dateStyle/timeStyle are late additions to Intl -- Firefox shipped them in
91 -- and an implementation that does not know them ignores them and
formats the date alone. The comparison would then read New York, Chicago
and Madrid as the same zone and tick the step for a timezone that is
plainly wrong, which is the failure the check exists to catch. Silent, and
only on older browsers.
Explicit numeric fields (year/month/day/hour/minute) have been in Intl
since ECMA-402 v1, so there is nothing left to degrade to.
The options look like a stylistic choice, so a test pins them: it reads the
comparison with comments stripped -- the comment names dateStyle to explain
why it is not used -- and fails if either style option comes back or a
time field is dropped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): sample both sides of DST when comparing zones
CodeRabbit caught this and it is right: comparing the wall clock at one
instant treats zones that merely coincide right now as the same one.
America/New_York and America/Lima hold the same offset all winter, so a
panel set to the wrong one of the two ticked the step in January and then
ran an hour off from March -- a silent false pass, which is the failure the
whole check exists to prevent. Same shape as the dateStyle problem in the
previous commit: a comparison coarser than it looks.
Three instants now, all of which must agree: now, and mid-January and
mid-July of the current year. Those sit either side of DST in both
hemispheres, so only zones that agree year-round match. Toronto still
matches New York, which is correct -- either renders the same times.
Two tests. A static one asserts the comparison samples more than the
current instant, since reverting to `[now]` looks like a simplification.
And a table pinning which pairs must count as the same zone: aliases and
same-rule zones equal, seasonal coincidences (New York/Lima,
Phoenix/Los_Angeles, Sydney/Guadalcanal) not. That table mirrors the
algorithm rather than executing the shipped JS -- there is no JS runtime
here and the repo has no JS test infra -- so it records the verdicts the
browser code has to reach, and the static guard keeps the two aligned.
Mutation-checked: reverting to a single instant fails the static guard.
Also documented what the city test compares. CodeRabbit read it as always
failing, on the grounds that the label differs between Tampa and Seattle.
It does, but timezone_step() returns the opening tag only, so the
comparison is over data-done and data-tz and the label is not in it. The
assertion is left as an equality over the whole tag, which is stronger than
checking the two attributes by name; the docstring now says so.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(service): survive corrupt health cache and clean exits
Three independent failure modes that each end with a dark panel and no
automatic recovery.
1. PluginHealthTracker._load_health_state returned the cached value
verbatim. If that value is not a dict, every caller raises
AttributeError: 'list' object has no attribute 'get' — during
DisplayController.__init__, so the process dies before the display
loop starts. systemd restarts it, the same bad entry is read back
from disk, and it dies again: an unattended restart loop that
survives reboots because the cause is persisted. Observed in the
field with plugin_health:<id> holding an unrelated plugin's list
payload. Now non-dict entries are discarded with a warning and the
defaults are rebuilt.
2. ledmatrix.service used Restart=on-failure, so any exit with status 0
left the unit stopped and the panel dark indefinitely — systemd
treats it as success and never brings it back. Restart=always.
3. ledmatrix-wifi-monitor.service used StandardOutput=syslog, which
systemd has marked obsolete; it warns and rewrites it to journal on
every load.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* perf(memory): size the cache to the board and stop reinstalling deps
On a 1GB Pi 3B+ the display process settles around 600MB RSS of 905MB
total. When the remaining headroom runs out the failure is not a clean
crash: fork() starts returning ENOMEM, so sshd accepts connections and
closes them before its banner, timer jobs stop running, and the panel
goes dark, while already-resident processes keep serving normally. The
board looks healthy from outside and cannot be logged into. Only a power
cycle clears it.
Three contributing causes:
- MemoryCache had a fixed 1000-entry ceiling. Entries are parsed API
payloads of tens of KB, so one ceiling cannot serve both a 512MB Zero
2 W and an 8GB Pi 5. Now scaled from MemTotal (150 entries at <=1GB,
1500 at >=8GB), overridable with LEDMATRIX_CACHE_MAX_ENTRIES.
- requirements_are_satisfied() returned False for any requirement with
extras, so a plugin depending on python-socketio[client] re-ran pip on
every single start: ~8s, a network dependency, and a 100-200MB spike
at the least convenient moment. During a restart loop it repeats for
each restart. Extras are now resolved one level deep against installed
metadata, keeping the conservative "anything unverifiable falls
through to pip" contract.
- ledmatrix.service had no memory ceiling. MemoryMax=85% expressed as a
percentage so one unit file suits every board. Note this needs the
memory cgroup controller, which Pi firmware disables by default;
first_time_install.sh now adds cgroup_enable=memory to cmdline.txt,
and the unit file documents how to verify it took effect.
first_time_install.sh also enables persistent journald storage (capped
at 64M). Default storage is volatile, so every reboot destroys the logs
that would explain why the board rebooted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: guidance for 512MB and 1GB boards
Documents the memory ceiling on small boards and, more usefully, what
running into it actually looks like: sshd accepting connections and
closing them before the banner, the web UI still responding normally,
clean ping, a dark panel, and a wrong clock after the next boot. None of
those read as "out of memory", which makes the failure hard to identify
from the symptoms.
Cross-referenced from SSH_UNAVAILABLE_AFTER_INSTALL.md, since "I can't
SSH in any more" is how most people will first meet this.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: address review findings on the low-memory work
Nine CodeRabbit findings, five in code.
**Health state (the one that matters).** The non-dict guard did not cover a
dict missing fields the callers index directly, which is the shape actually
seen in the wild: a record carrying only circuit_state produced
`plugin clock-simple operation failed: 'circuit_state'` about fifty times a
minute with the panel frozen. The record is now completed against the
defaults per field rather than trusted or discarded wholesale. Per field
matters: a first pass rejected any incomplete record outright, which reset a
tripped breaker and real failure counts to healthy because one optional
field was absent -- an existing test caught it. Values of the wrong type
(a counter persisted as a string, an unknown circuit_state) fall back
individually, valid neighbours survive, and newer fields the schema has
grown since (degraded, degraded_reason) are carried through untouched.
**Cache ceiling.** MemoryCache.set() accepted entries without bound between
cleanup sweeps, which run every 300s by default, so a burst could take the
cache far past max_size -- the unbounded growth the limit exists to stop.
Eviction now runs under the same lock on every write, sharing one helper
with the periodic sweep so the two cannot drift.
**Installer, cgroups.** Only cgroup_enable=memory was checked, so a board
carrying that without cgroup_memory=1 reported success and got no change,
leaving MemoryMax= inert. Each parameter is now checked and appended
independently; verified against all four combinations, single line preserved.
**Installer, journald.** Persistence was inferred from /var/log/journal being
non-empty, which proves neither Storage=persistent nor a size cap -- the
directory survives a switch back to volatile. The effective configuration is
read instead (systemd-analyze cat-config, falling back to the conf files),
and an explicitly configured SystemMaxUse is preserved rather than
overwritten. Verified across volatile, persistent-without-cap,
persistent-with-user-cap, cap-without-storage, and commented-only configs.
**Dependency extras.** _extras_are_satisfied stopped at one level, so a
gated dependency that itself requests an extra (requests[socks]) passed on
the base distribution's version while the extra's own dependency was
missing, and pip was skipped. It now recurses, with a visited
(distribution, extras) set so a cycle terminates.
Docs: both kernel command-line paths documented (the installer falls back to
/boot/cmdline.txt), daemon-reload and restart added after the systemd
override example, memory exhaustion added to the SSH summary with its
power-cycle-only recovery, and a language on the fenced block for MD040.
Tests: five for the health-state repair including the exact wild shape and
that record_failure/record_success no longer raise against it, and one for
the cache ceiling. Both mutation-checked. Full suite 2927 passed, with the
one pre-existing tmpfs failure that also fails on main.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix: harden the health-state repair and confirm journald took effect
Second review round; all three findings were valid and two were bugs in the
repair added last commit.
The repair could raise out of itself. An unhashable circuit_state (a list or
dict on disk) hit `value in {...}` and raised TypeError -- from the code
whose whole job is to stop a malformed record crashing the caller. It now
requires a str before the membership test.
bool is a subclass of int, so True passed the timestamp check and then
compared as 1.0: enough to expire a cooldown the instant the breaker opened,
while False would stop the elapsed check firing at all. Timestamps now
exclude bool explicitly.
The regression test for the original crash was seeded with a record that
*contained* circuit_state, so it passed against the old raw-return behaviour
too -- the counters are read with .get(), so circuit_state is the only field
whose absence used to raise. Reseeded to omit it, and it now fails against
raw-return as intended.
journald: drop-ins apply in lexical order, so a local file sorting after
ledmatrix-persistent.conf still wins and writing ours proves nothing. The
effective Storage is re-read afterwards and a warning naming the diagnostic
command is printed if persistence is still not active, rather than reporting
a success that was not verified.
Full suite 2934 passed, same single pre-existing tmpfs failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(pixlet): resolve the release tag correctly when downloading
Starlark apps render through the pixlet binary, and the installer that
fetches it silently produced nothing, so every app failed with "Pixlet
not available - Starlark apps will not work".
Two compounding defects:
The version lookup parsed the wrong token. GitHub returns the release
JSON on a single line, so `grep '"tag_name"'` matches the whole document
and the greedy `sed 's/.*"([^"]+)".*/\1/'` captures the LAST quoted
string in it. That resolved to "mentions_count", giving a download URL
for a release that does not exist. The `[ -z "$PIXLET_VERSION" ]`
fallback never fired, because the value was not empty -- just wrong.
And `curl -L -o` without `-f` writes a 404 body to the file and exits 0,
so the download was reported as successful and the first sign of trouble
was tar complaining "not in gzip format" about a page of HTML:
→ Downloading linux-arm64...
Extracting...
gzip: stdin: not in gzip format
✗ Failed to extract archive: .../pixlet_mentions_count_linux-arm64.tar.gz
Download complete: 0/1 succeeded
Now the tag field is matched directly and the value taken from it, and
the result is checked for a version shape rather than merely being
non-empty -- a wrong-but-non-empty value is exactly what made this
silent. curl gets -f so an HTTP error is a failure, and the archive is
gzip-tested before extraction, since a proxy can return 200 with an
error page.
Verified on an arm64 rig: v0.53.1 resolved, 1/1 downloaded, the binary
runs, and the plugin's own detection finds it at
bin/pixlet/pixlet-linux-arm64.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(pixlet): anchor the version check, and don't echo response bytes raw
Both CodeRabbit findings were valid.
The shape check accepted partial matches, so "v0.53garbage", "0.53" and
"v0.5" passed it and built a download URL for a release that cannot exist
-- the failure the check was added to stop, just one step later. Anchored
at both ends now. Every tronbyt/pixlet release to date is vX.Y.Z (all 38
verified against the API), with an optional suffix left for a future -rc.1
or +build tag.
The invalid-response diagnostic printed bytes straight from whatever
answered the request. NUL and newline were filtered but escape, carriage
return and backspace were not, so an error page could rewrite the output
or bury it in a CI log. Non-printable bytes are stripped and it goes
through printf. CodeRabbit suggested hex-encoding the lot; printable
characters are kept instead, because "<!DOCTYPE html>" is the diagnostic
-- hex would make the line safe and useless.
Also corrected the comment above the parse. It asserted GitHub returns
this JSON on a single line; the API is pretty-printed by default, and I
could not get a single-line response from two machines across five header
variants. The single-line case is real (it is what produces
"mentions_count", and the failing device's error named
pixlet_mentions_count_linux-arm64.tar.gz), but it is a shape to be robust
against, not a constant. As written the comment invites the next reader to
check by hand, see pretty JSON, and conclude the fix was unnecessary.
Tests drive the real script with a stubbed curl: the tag resolves from
both response shapes, non-release values fall back, an HTTP error is
reported as a download failure rather than surfacing later as a tar error,
a non-archive body is rejected before extraction, and the diagnostic
cannot carry control bytes. The stub honours -f the way real curl does --
without that, the HTTP-error test passed against the old script too, since
both end at 0/1 and only the reporting layer differs.
Mutation-checked: 10 of the 16 fail against the pre-fix script, and the 5
covering these two findings fail against this branch's previous state.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
run_scheduled_updates_with_changes() snapshotted plugin_last_update,
called run_scheduled_updates(), and diffed the two to answer "whose data
just changed".
But run_scheduled_updates() only enqueues. The work runs on the update
worker and stamps plugin_last_update there, after this method has already
returned, so the two snapshots were always identical and the result was
always an empty list. The only path that ever worked was the synchronous
kill-switch, where update() runs inline.
Vegas is the caller. That empty list is what feeds mark_plugin_updated(),
which drops the cached content for a plugin whose data moved -- so a
segment kept scrolling whatever it was first built from. It is the
failure the coordinator's own comments describe: last night's live game
still drawn as live the next morning. On a live rig: zero update ticks in
twenty minutes, with weather, stocks and news all updating on schedule.
The worker now records each completed update in a ledger and the call
drains it, reporting what has finished since the previous poll rather
than what this call enqueued. That costs one tick of latency -- Vegas
polls every ~4s -- and is correct whichever side of the queue the work
lands on. Failure paths are excluded: they stamp the timestamp too, to
space out retries, but no fresh data exists.
Verified on the rig it was found on: 0 update ticks before, 208 in
twenty-five minutes after, naming real plugins.
The behavioural tests here would pass with both production call sites
deleted, which mutation testing caught -- they drive the ledger directly.
So there is also a structural test asserting the invariant at the source:
wherever a successful update stamps plugin_last_update, it must record
the completion. Writing it immediately caught that _record_update_failure
stamps the same field and must not be included.
Mutation-checked: removing either call site, removing both, and dropping
the drain's clear are all caught.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(vegas): let live content keep its place in the ticker
Live content used to preempt Vegas outright: while any plugin reported
live priority the display controller refused to run the ticker at all and
showed a full-screen scoreboard instead. Keeping the marquee meant not
seeing live scores; seeing live scores meant losing the marquee.
Two changes, both off by default.
vegas_scroll.live_in_ticker keeps the ticker running through a live game.
Three places assumed the takeover and all three now honour it: the
controller's gate, the coordinator's per-frame pause, and the rotation
switch that would otherwise move current_mode_index underneath a ticker
that never yields.
And the rotation is no longer a strict round robin. It was one slot per
plugin per cycle, so with a dozen plugins enabled a live score came round
once a lap and could be minutes old on screen. A plugin can now hold
several slots, placed by Smooth Weighted Round-Robin -- the same
scheduler the sports plugins already use to rotate their own games. The
property that matters is that repeats are spread through the cycle
rather than clumped: three in a row and then silence would be worse than
no boost at all.
Weight comes from the plugin first, via a new optional
get_vegas_priority_weight(), then from the core: live content earns
live_weight, everything else 1. So existing plugins gain the behaviour
without changes, and the hook exists for the one thing the core cannot
work out -- the core can see that a game is live but not whose, so only
the plugin can say a favorite is playing.
Documented in ADVANCED_FEATURES (worked example, why weights are per
plugin not per game, and that frequency is not freshness),
CONFIG_REFERENCE, PLUGIN_API_REFERENCE, and the config template.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(vegas): carry the new keys through config, and correct two docs
Three findings from CodeRabbit, all valid.
to_dict() and update() enumerate keys explicitly and had not learned the
three new ones, so get_status() never reported them and a live config
change never applied -- turning live_in_ticker on in the web UI would
have done nothing until a restart. update() clamps the weights exactly
as from_config does.
The vegas_scroll key count in ADVANCED_FEATURES said 29; the template
has 30. My arithmetic, not the reviewer's.
The third was a documentation error rather than a code one, and I have
fixed it the other way round. The docs claimed a raising
get_vegas_priority_weight() is treated as weight 1. The code instead
falls through to the core's own live-content check, and that is the
better behaviour: the hook is only how a plugin asks for *more* than
live_weight, and has_live_priority/has_live_content are separate methods
guarded separately, so a plugin with a broken weight calculation should
lose the favorite distinction and keep the live boost. Said so in the
code, the base-plugin docstring and the API reference.
The test fake now fails in each place independently, because the two
failures mean different things: a broken hook still earns live_weight, a
plugin that cannot say whether it is live has nothing to fall back on
and weighs 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ui/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(vegas): stop the heaviest plugin doubling across the cycle seam
Smooth Weighted Round-Robin spaces repeats well within a pass, but it
schedules the heaviest item first and usually last as well. The strip
loops, so those two are neighbours: the marquee showed the same plugin
twice running at exactly the one join a within-cycle check cannot see.
Observed on a live rig at 28 slots -- gaps of 6, 7, 7, 7 and then 1.
Rotating the list does not fix it. Rotation preserves the cyclic order
exactly, so it moves where the seam is drawn rather than the adjacency
itself; the trailing entry has to be swapped with one from the middle.
The first version swapped with the first slot that merely fitted, which
undid the spacing this exists to protect -- it moved a repeat from a gap
of 7 into a gap of 2, more clumped than the seam had ever been. It now
picks the candidate furthest from any other appearance, so the repeat
lands in the widest gap.
Left alone when no candidate exists. A plugin holding most of the slots
has to neighbour itself, and scheduling it is better than refusing to.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(vegas): stop the seam repair creating the duplicate it removes
Swapping the trailing repeat with a middle slot moves two elements, and
the candidate filter only guarded one of them. It checked the neighbours
`repeated` would acquire at j, but not what the displaced element would
sit beside at the end -- so ['a','b','c','d','x','y','x','a'] came back
as [...,'x','x'], the seam duplicate traded for a fresh one. Reported by
CodeRabbit with that exact case.
Adding the missing condition fixed it and immediately broke something
else: schedule[j] is schedule[-2] when j is the second-to-last slot, so
that candidate was always excluded, and ['a','b','c','a'] lost the only
repair it has. The same class of mistake twice, from reasoning about
which neighbours two moved elements end up with.
So it no longer reasons. It performs each candidate swap, counts the
cyclic duplicates in the result, and keeps the best one that has none --
preferring whichever leaves the boosted plugin most evenly spread. When
no such swap exists the schedule is returned untouched, which is the
unavoidable case: a plugin holding most of the slots has to neighbour
itself.
Fuzzed across 6,956 seam schedules: none made worse, none lost an entry.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(calendar): implement the OAuth and calendar-listing endpoints
The plugin's config advertised a three-step setup and only step 1
existed. Step 3's picker fetched
/api/v3/plugins/calendar/list-calendars, which was never registered, so
Flask fell through to the global 404 handler and the user saw "Resource
not found" -- a message that names nothing and points nowhere. Step 2
had no endpoint at all, so even a working picker would have found no
token to list with.
Two routes, following the pattern the spotify and ytm plugins already
use for their own auth scripts:
POST /plugins/calendar/authenticate two-step Google OAuth
GET /plugins/calendar/list-calendars calendars for the picker
The authenticate route drives calendar_registration.py, which the plugin
already ships and which was written expressly for this -- it reads a
redirect URL on stdin and prints one JSON object. It takes two calls
because a human has to visit Google in between; the script persists the
PKCE verifier from the first call for the second, without which the
exchange fails with "Missing code verifier".
The listing route reads the token directly rather than shelling out
again: the picker is interactive and a subprocess per click is slower
than the API call it would wrap. It refreshes an expired token in place,
sorts the primary calendar first, and drops entries with no id, which
could not be selected anyway.
Both name the plugin when it is not installed, rather than reproducing
the anonymous 404 that started this.
Verified against the live Google API on the dev rig: HTTP 200 with the
account's real calendars.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(calendar): one input for the auth code, and a louder warning
Two things reported after testing the flow.
There were two boxes and no way to tell which to use. The config
template's string branch dispatches widgets from an allow-list of names,
and anything missing from it falls through to a plain input type=text --
so the field rendered both the widget's own box and a stray one for the
same key. google-oauth is now on that list, which is all the widget ever
needed to render in place of the fallback rather than beside it.
And the warning that the redirect page fails to load was small grey text
under a link, which is where it is least likely to be read. It is now an
amber callout that leads with "The next page will fail to load. That is
expected." The failure lands at exactly the moment the user has to act
on it, and it looks precisely like the flow breaking rather than
working. The paste box is labelled too, rather than relying on a
placeholder that vanishes on focus.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(calendar): redact script diagnostics, and page the calendar list
Three findings from CodeRabbit, all valid.
Raw subprocess output was being returned to the client -- the script's
stderr on one path, and its own error payload on another. CodeQL flagged
the same line. That script handles OAuth client secrets and interpolates
exceptions into its messages, so either could carry a secret or a path.
Both now go to the log unredacted, where they are worth having in full,
and reach the client through a redactor.
That redactor already existed inside describe_exception, which only
takes exceptions. Split out as redact_text: an exception is not the only
thing worth returning, and a subprocess's stderr is just as capable of
quoting a token.
calendarList.list returns 100 entries per page by default, caps at 250,
and hands back a nextPageToken when there are more. Reading one page
would have hidden calendars from the picker with nothing to say the list
was cut short. It now pages, asking for 250 at a time, bounded at ten
pages so a malformed token cannot spin.
And a test helper was a lambda where ruff wants a def.
The five new tests fail against the previous commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(calendar): redact the last two raw exception interpolations
CodeQL flagged four exposure paths. Two were mine and genuinely raw: the
OSError from failing to spawn calendar_registration.py, which carries the
interpreter path and whatever the OS chose to say, and the ImportError
for the Google libraries, whose message named the missing module by
interpolating the exception directly. Both now go through
describe_exception, and the unredacted text goes to the log.
The other two are the repo-wide pattern from PR #448 -- 67 handlers on
main already return details=describe_exception(e), and these two new
handlers follow it. That function is the sanitizer: it strips URL
userinfo, auth headers and credential-shaped key=value pairs, collapses
to one line and caps the length. CodeQL's taint tracking cannot see a
sanitizer it has no model for, so it reports the flow regardless.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(calendar): announce status changes, name the paste box, drop a no-op
Three more from the review, all valid.
The status line is written after every async call -- the consent link is
ready, the exchange failed -- and was a plain paragraph, so a screen
reader was told none of it. It is a live region now.
The paste box had a visible label that was never associated with it, so
its only accessible name was the placeholder, which disappears on focus:
precisely when the value is being pasted. The label now points at the
input by id.
And a conditional in the test helper returned the same value from both
branches, which Ruff flags as RUF034. It was left over from making the
fake page; one page is all those cases need, and TestPagination builds
its own sequences.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* test(calendar): assert the accessibility relationships, not their parts
The previous assertions searched for role="status", aria-live, a label
`for` and an input `id` independently, so they passed whether or not
those belonged together. Two attributes on different elements announce
nothing, and a `for` that names something other than the input leaves it
just as anonymous.
Both attributes are now asserted on the status element itself, and the
label and input are checked to go through the same identifier rather
than merely both existing. Verified by mutation: a mismatched pair and a
displaced aria-live are both caught.
Reported by CodeRabbit, against tests I had written two commits earlier.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(startup): bound the initial plugin update so the panel lights sooner
DisplayController.__init__ calls _update_modules() once to populate
plugin data before the first frame. It walks every loaded plugin in
turn, and each update blocks the calling thread for up to the executor's
30s timeout, so the uncapped total is the sum of every slow plugin on
the system. The rig's own log:
Initial plugin update completed in 82.255 seconds
Initial plugin update completed in 55.123 seconds
Initial plugin update completed in 25.975 seconds
The panel shows nothing for all of it.
Nothing is lost by stopping early. A plugin that has never updated is
immediately due, so run_scheduled_updates() collects it seconds later --
with the display already running rather than blank.
A deadline alone was not enough: it is checked before each plugin, so
the last one to start could still block for the full 30s, and a 20s
budget produced a 31.8s pass on the rig. The remaining budget is now
passed down as that update's timeout too, with a floor so a plugin
starting on the last sliver is not handed ~0s and recorded as having
timed out for a slot it never had. Measured after: 20.006s.
Found while profiling a scroll freeze with py-spy, which caught the main
thread 9.34s inside execute_with_timeout's join. Worth being clear that
this is startup latency, not the recurring stutter -- _update_modules
has exactly one caller and runtime updates already run off the display
thread.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* feat(display): show the device address on the startup screen
That screen is what the panel holds for the whole initial plugin update,
and on a headless Pi it is the only place the address appears without
going looking for it -- so it now carries the address under
"Initializing".
The lookup connects a UDP socket, which sends no packets: it only asks
the kernel which source address it would route from. That costs 0.03ms
and works with the network down so long as a route exists. Deliberately
not `hostname -I` plus a systemctl probe for AP mode, which is how the
web launcher does it -- two subprocesses with multi-second timeouts, on
the startup path this branch exists to shorten.
Two things had to change for the address to be worth putting there.
The text is now sized to fit rather than fixed at 8px: "Initializing" is
96px in PressStart2P, drawn at x=10, so it already ran off the side of a
64px panel before an address was added. It falls back to 4x6 where that
does not fit, and both lines are centred.
And the test pattern is punched out from behind the block, with the text
drawn white rather than blue. The diagonal runs through the middle of
the panel, which is exactly where this sits, and blue on black reads
fine on a monitor but is marginal on a dim panel. An address that cannot
be read off the wall is not worth showing.
The rendering tests assert against pixels -- no green left behind the
text at any supported size, enough lit pixels to be visible -- rather
than against the geometry that produced them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(display): keep the startup text blue -- it is a channel reference
The test pattern lights one pure channel per element: red border, green
diagonal, blue text. That is how a glance at the panel tells you whether
led_rgb_sequence is right -- wire it BGR and the border comes up blue
and the text red. Drawing the text white, as the previous commit did for
contrast, lights all three channels and destroys the only blue reference
on the screen.
Reverted to blue, with the reason written down so it is not treated as a
style preference again, and with tests that pin it: the text must be
pure blue, nothing on the screen may be white, and all three primaries
must be present.
The punched-out backdrop stays. It only removes the diagonal from behind
the glyphs, which costs nothing diagnostically -- the diagonal is still
plainly visible across the rest of the panel -- and it is what makes the
address readable at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(startup): defer a plugin with too little budget, rather than clamp it
The per-plugin timeout was clamped up to a floor, so a plugin that began
with a sliver of budget left was granted the full floor and ran on past
the deadline: a 20s budget could take 22. The floor existed to stop a
plugin being handed a slot too short to use and then recorded as having
timed out, which is a real concern, but clamping solved it by breaking
the bound.
Deferring solves both. Below the floor the plugin is left to the update
tick, which was already the fate of everything after the deadline, so
nothing new is lost -- a plugin that has never updated is immediately
due. Above it, the timeout is the exact remainder, and the pass cannot
outlast its deadline.
Measured on the rig after the change: 20.002s, 5 plugins deferred.
Also names an unused binding in the initializing-screen test.
Both reported by CodeRabbit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(vegas): make scroll stutter visible, and catch it in the act
The loop reported only a mean FPS over a five-second window. At 120fps
that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee --
moves the average from 120.0 to 115.4 and reads as healthy. Stutter was
literally unmeasurable.
The FPS line now carries p99, the worst frame, and a hitch count. On the
dev rig that immediately turned "it sometimes stutters" into a number:
two freezes of 3.2s and 0.7s in twenty minutes, with every other frame
under 81ms.
Statistics say a stall happened but not what caused it, and by the time
they are logged the stack is gone. So there is also a watchdog that dumps
every thread's stack while the loop is still wedged. It is off unless
LEDMATRIX_STALL_WATCHDOG is set to a threshold in seconds, since it
prints a lot. Pointed at the 3.2s freeze it named the culprit on the
first try: a plugin generating a 17,000px scroll image, logo PNG decode
and all, synchronously on the render thread.
The hitch threshold is relative to what frames actually cost, not to the
configured target. The target is routinely set above what the panel can
hold so vsync does the pacing; measured against that budget every
ordinary frame counts as a hitch, and the first version of this counter
duly reported 250 per window on a display running perfectly smoothly.
The watchdog is owned by the coordinator, not created per iteration --
run_iteration is called repeatedly, so building one there would leak a
thread each time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(vegas): let the watchdog see stalls that hold the GIL
The watchdog only noticed a late heartbeat, which a whole class of
freeze can never produce: if the loop is inside one long C call that
holds the GIL, this thread cannot run during the stall, and by the time
it does the loop has already checked in. On the dev rig that hid a
recurring 3.2s freeze completely -- twenty minutes of watching produced
one dump, for an unrelated 0.4s stall.
What it can still observe is that its own sleep ran long. A badly
overshot wait is now reported as a stall in its own right. The stacks
are stale by then and the message says so, but knowing the freeze is
GIL-holding is most of the diagnosis: it rules out lock contention and
scheduling, and points at a single long C call.
This also explains why lowering sys.setswitchinterval changed nothing --
the switch interval cannot preempt a C call that never releases the GIL.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* feat(vegas): report the worst frame, not just the mean
The loop logged only a mean FPS over a five-second window. At 120fps
that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee
-- moves the average from 120.0 to 115.4 and reads as perfectly healthy.
Stutter was unmeasurable, which is why "it sometimes freezes" went
unpinned for so long.
Adding p99 and the worst frame turned that into a number immediately: on
the dev rig, two freezes of 3.2s and 0.7s in twenty minutes with every
other frame under 81ms. Not general slowness -- two rare, total stalls,
which is a different problem with a different fix.
Costs 0.96us per frame, about 0.012% of an 8.3ms frame.
This replaces an earlier version that also shipped a stall watchdog and
a hitch counter. The watchdog never found anything -- one dump in
forty-five minutes, for an unrelated stall -- because it can only notice
a late heartbeat, and the freeze happens in coordinator.start() before
the frame loop begins beating. py-spy found the cause in one recording
by sampling the process externally, which needs no code here. The hitch
counter went with it: it needed a rolling median every frame, which was
most of the cost, to produce a number the worst frame already tells you.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(vegas): use the nearest-rank index for p99
int(n * 0.99) is off by one, and at exactly 100 samples it selects the
maximum -- which is the number logged immediately beside it as the worst
frame. The two columns exist to say different things, p99 the
bad-but-ordinary frame and worst the outlier, so they agreed precisely
when the sample was smallest and least informative.
Nearest rank is ceil(n * fraction) - 1. Extracted so it can be tested
directly rather than only through a five-second logging interval.
Reported by CodeRabbit on the PR.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a display.hardware.orientation config field ("normal" / "180")
so panels mounted upside down (e.g. to put the Pi/wiring on a more
convenient side) render correctly without custom pixel_mapper_config
edits. Composes onto the existing pixel_mapper_config as a trailing
"Rotate:180" mapper, so it stays independent of any custom mapper
string (e.g. U-mapper chain layouts) already in use.
Exposed as a "Panel Orientation" dropdown in the web UI's Display
settings, validated server-side, and documented in README and
CONFIG_REFERENCE.
Claude-Session: https://claude.ai/code/session_01FakipqMDHQLpsFjTuBdSFQ
Co-authored-by: Claude <noreply@anthropic.com>
The display process ran three cleanup threads over one directory:
14:22:59.954 display_controller (the real manager)
14:22:59.973 startup validation, run 1 (discarded)
14:23:01.055 startup validation, run 2 (discarded)
Two of those managers existed only to read a directory path.
StartupValidator._validate_cache_directory built a whole CacheManager to
call get_cache_dir(), and validation runs twice -- once before the
plugin manager exists and again after. Each construction also probes
writability by writing and deleting .writetest on the card.
The discarded ones never went away. cleanup_loop closes over `self`, so
the thread keeps its manager alive: two objects that could never be
collected, waking every 24 hours to re-scan the same 9,000-file
directory. Nothing stopped them either -- stop_cleanup_thread had no
callers anywhere in the tree.
Two changes. The validator now takes the CacheManager the application
actually uses, which is also the more correct thing to validate; when
no caller supplies one it still builds its own, but stops the thread
afterwards. And CacheManager now tracks which directory it is sweeping,
so the second manager over a directory skips starting a thread at all.
That is the right granularity regardless of call sites: the sweep lists
a directory and deletes from it, so a second thread only duplicates the
scan. Ownership is released on stop, so a survivor can take over rather
than leaving the directory permanently unclaimed by a dead owner.
Measured directly, three managers over one directory: 3 threads before,
1 after.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(cache): collect the temp files abandoned writes leave behind
DiskCache.set() writes through mkstemp then os.replace, and removes its
own temp file in a finally. That covers a write that fails, but not a
process that dies between the two -- a SIGKILL, a lost restart race, a
power cut, all ordinary on a Pi.
Nothing ever collected what was left. The temp names are
".<key>.json.<random>", and cleanup_expired_files listed only names
ending in .json, so every one of them was invisible to the sweep for as
long as the card had been in service. On the dev rig: 76 files,
1,050 MB, 81% of the whole cache directory, the oldest six months old.
The startup sweep reported "18/8864 files deleted, 0.01 MB freed" while
sitting on top of a gigabyte it could not see.
They are removed after an hour. A real write holds its temp file for
milliseconds, so that is far outside any in-flight write while still
clearing the same day's debris, and it is deliberately not tied to the
retention policies: those say how long data stays useful, and a
half-written file never was.
The predicate is tested harder than the sweep, because a false positive
deletes real data. It matches the shape set() creates rather than just a
leading dot, so a completed ".json", a stray .gitignore, and a
"weather.json.bak" are all left alone -- and one test drives set()
itself and asserts the names it produces are matched, so the writer and
the predicate cannot drift apart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(cache): count swept temp files as scanned
files_scanned only counted completed .json files, so a sweep that
removed orphans reported more deleted than it had looked at -- the
summary line renders "<deleted>/<scanned>", which came out as "76/1".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The odds fetch used a bare requests.get, so it went out as
python-requests/x.y -- the one agent ESPN is known to reject. Around
2026-08-04 it began 403ing browser strings and bare custom tokens alike;
what it accepts is a token carrying a URL that says who is calling.
Every other ESPN caller in the tree already sends that header
(src/common/api_helper.py, src/base_classes/data_sources.py); this path
was simply missed.
It is the worst one to miss. Odds are fetched per live game from inside
the live update loop, so its failures are the ones that cost the caller
its whole update budget -- the same path the 5s timeout and the cooldown
were added to protect.
Sent via a session rather than per-call, which also reuses the
connection across a slate. Deliberately no retry adapter, unlike
api_helper: retries multiply request_timeout, which is 5s precisely to
stay inside the 30s operation budget.
The existing tests patched the module's requests.get, which this change
bypasses -- test_base_odds_manager was consequently reaching the real
ESPN and taking 404s. Both files now patch the session, and the new
tests pin the agent against api_helper's live value so the two cannot
drift apart the next time ESPN moves the goalposts.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
CacheManager.set(key, data, ttl=...) stored the number and no read path
ever consulted it. Expiry came from a max_age inferred from substrings
in the key -- "live", "odds", "stock" -- so all 52 callers passing a ttl
were writing a value that did nothing. The docstring said so outright:
"stored for compatibility but expiration is still controlled via max_age
when reading". It is easier to read that as a note than as a defect,
which is presumably how it survived.
Both cache layers already hold the record when they decide, so each now
prefers an explicit ttl and falls back to max_age when there is none.
The caller that wrote the record knows what its data is; a substring
guess is a reasonable default for records that never said, and a poor
override for records that did.
Measured against a device's real cache of 8,875 entries carrying a ttl,
the inferred and intended values disagreed nearly everywhere:
stocks max_age 600 vs ttl 1800 4903 entries
news max_age 3600 vs ttl 600 1770 entries
odds max_age 1800 vs ttl 3600 1301 entries
images max_age 300 vs ttl 2592000 20 entries
In every case the ttl matches what the plugin plainly intended: stock
quotes cached for half an hour rather than ten minutes, headlines
refreshed every ten minutes rather than hourly, bird photographs that
never change kept for a month rather than five minutes.
Two things make this safe to land now. No sports_live entry carries a
ttl at all -- the live-score path does not use set(ttl=) -- so live
freshness is untouched, which matters with a season two weeks out. And
replaying the change against that real cache, 997 currently-expired
entries become live while not one live entry becomes expired, so there
is no invalidation spike on deploy.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Odds are fetched per live game from inside SportsLive.update(), with
show_odds defaulting on, and the plugin executor kills an operation at
30s. The odds request timeout was also 30s, so a single stalled request
consumed the entire budget and the update carrying every game's score
was killed.
Out of season that is invisible: preseason week 1 returns one game. A
Sunday slate is around sixteen, so the odds of at least one slow request
rise sharply just as the cost of losing the update does.
Shorten the request timeout to 5s, and after a network failure skip the
network for 60s. The timeout alone is not enough -- sixteen consecutive
5s timeouts still blow through -- and when ESPN is unreachable it is
unreachable for the whole slate, so the first failure already answers
the question for the rest of the pass.
before: one stalled request = 30s = the entire budget
after : 5s, the rest of the slate skipped, retry after 60s
The stale-cache fallback is unchanged: the cache is consulted before any
of this, and the failing request still falls back to it.
An earlier version of this branch also jittered the cache TTL to stagger
expiry across a slate. That has been dropped: CacheManager.set() stores
ttl for compatibility but the read path expires entries by a per-type
max_age (1800s for odds), so the jitter was inert. Making the read path
honour a per-entry ttl is a real fix but changes a contract 48 plugin
call sites already rely on, which is not a change to make two weeks
before the season.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(web): say what actually went wrong instead of "unknown"
Every failing endpoint returned "An error occurred; see logs for
details" and nothing else. That is survivable until the logs are the
thing you cannot reach: a device whose SD card was failing answered the
restart action, /system/status and /logs with that same sentence -- the
log viewer included, because journalctl could not be executed -- while
the exception underneath said
[Errno 5] Input/output error: 'systemctl'
which names the fault outright. The only endpoint that helped was
/health, and only because it happens to pass a subprocess's stderr
through. Diagnosis came down to guessing which endpoint leaked something.
Add describe_exception(), returning "TypeName: message" on one line, and
populate the `details` field that the response schema has always had and
nothing ever filled. The type alone carries information -- a bare
PermissionError says more than any generic sentence.
Exception text is not automatically safe to echo: a requests error
quotes the URL it failed on, and plugins that authenticate by query
string put their key there. Credential values are redacted while the
parameter name is kept, since knowing which credential was involved is
part of the diagnosis. Length is capped and newlines collapsed so a
parser's context cannot flood a JSON field.
Nine handlers in api_v3 bound the exception and never used it, so the
promised log entry was never written either -- "see logs for details"
was false, not merely unhelpful. Those now log with a traceback and
carry the detail. The other 60 already logged and are unchanged; they
can adopt the helper as they are touched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): redact auth headers and URL userinfo, and cover every handler
Three review findings.
The sanitizer missed two credential shapes that requests puts in its
exception text verbatim: `Authorization: Bearer <token>` and
`https://user:password@host`. Both would have gone straight into a
response. The auth-scheme name and the username are kept -- they say
which credential and whose without being the secret.
The AST test only asked whether *something* had been logged, so a
`logger.info("failed")` satisfied it while discarding the exception just
as completely. It now requires an error-level record carrying exc_info
and `describe_exception()` called on the handler's own bound exception.
Enforcing that revealed the first cut had scoped itself wrongly. I had
converted the nine handlers that logged nothing and left the sixty that
logged, reasoning their detail was at least in the journal. But
/system/status is one of the sixty, and on the failing device it told me
nothing -- the journal was exactly what could not be read. Splitting
them left most of the diagnostic surface unhelpful for the case this
change exists for, so all sixty-nine now carry the detail.
Two handlers had no bound exception name, and three passed the message
through a variable rather than a literal; both shapes needed doing by
hand. Full suite: 2383 passed, one pre-existing unrelated failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): stop reporting client errors as server faults
Werkzeug's HTTPExceptions subclass Exception, so the catch-all handler
saw them too and turned every 405, 400, 413 and 415 into a 500
UNKNOWN_ERROR. A GET on a POST-only route answered "an error occurred;
see logs for details", which tells the caller nothing and blames the
wrong side -- found while probing a device whose POST-only config
endpoints did exactly that.
Hand HTTPExceptions back as themselves, with their own status and
description. A genuine server fault still reports as one, with the
detail this branch adds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): redact any auth scheme, and require the detail in the response
Two review findings.
The auth-header pattern listed Bearer, Basic, Digest and Token, so
`Authorization: ApiKey SECRET` or `Negotiate SECRET` went to the client
intact. A fixed list silently leaks whatever it does not name, and
plugin APIs invent their own schemes, so match any scheme name and keep
it while redacting the credential.
The AST test accepted a describe_exception(e) call anywhere in the
handler, which a handler could satisfy by computing the detail and
dropping it before returning the generic message. It now requires the
call inside every return expression, which is where it has to be to
reach the caller.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(vegas): stop the width cap emitting fragments and stale windows
Two defects in the rotation that narrows an oversized plugin to its
width budget. Both were found while investigating "cut off early /
starts in the middle" reports and are the reason the cap is no longer
on by default; they still bite anyone who sets one.
A rotation's last window was whatever happened to be left over. Windows
are placed by walking forward from the previous one, with nothing
looking at the remainder, so a 1,840px stocks ticker against a 1,536px
budget split 1,492 + 348 -- every other appearance showed seven seconds
and cut. Absorb a remainder below half a budget into the window before
it. That overruns the budget by at most half, which is the better trade:
the budget guards against one plugin holding the panel for minutes, not
against a 20% overshoot. The floor is measured against the budget rather
than the panel because snapping to item boundaries already lands an
ordinary window short of it -- a 512px budget over 182px-pitch items
yields 348px windows, so an absolute floor merges windows that were
never fragments.
The stored offset also outlived the content it was recorded against. It
was a pixel column, reused verbatim after the plugin re-rendered, so
once anything ahead of it changed width the window pointed at unrelated
items -- observed as news refreshing 9,793px -> 9,505px mid-rotation.
Track the rotation as an index into the strip's item boundaries instead,
since the Nth boundary survives a digit appearing in a price, and record
alongside it what the offset indexes into: a row list, a boundary list,
or a column in a gapless image. A mismatch restarts the rotation rather
than reinterpreting the number, which also closes the case where one
plugin's row index was read back as a pixel column after its content
changed from several rows to one wide strip.
Replaying the four plugins that actually hit the cap on a live 512px
panel: no window is now a fragment, none exceeds 1.5 budgets, and every
rotation still covers the whole strip.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(vegas): apply the runt floor to the multi-row rotation too
The floor only guarded the single-image path. I had reasoned the row
path could not produce a runt because it wraps, which is wrong: wrapping
only helps when the row wrapped to actually fits. Rows of 450, 450 and
100 against a 512px budget give the 100 a pass of its own -- two seconds
against nine, which is the symptom this branch exists to remove.
Reproduced before changing anything:
pass 1: 450px pass 2: 450px pass 3: 100px
A window may now overrun the budget while it is still shorter than the
floor, bounded at the same 1.5 budgets the single-image path allows, so
the short row is carried with its neighbour instead of standing alone.
pass 1: 450px pass 2: 450px pass 3: 550px
A next row too wide to absorb within that cap still leaves a short
window standing -- rows of 900 and 100 keep alternating. Merging them
would mean a window of nearly two budgets, and the rule that always
shows an oversized first row already makes the same trade.
Three regression tests: the reported shape, that the overrun stays
bounded when a row cannot be absorbed, and that absorbing never drops a
row from the rotation. The single-image path is untouched -- the four
plugins that actually hit the cap on a live panel replay identically.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The controller skips a mode whose display() returns False and treats
anything else -- including None -- as "content was shown". A mode that
draws nothing and does not return False is therefore never skipped, and
because a mode switch clears the panel first, it sits on a blank screen
for its whole display duration. Two sports plugins shipped exactly that.
The harness rendered those modes and passed them, because it called
display() and discarded the result. Capture it, and warn when a render
produced no lit pixels while claiming content.
Warn-only by default, and deliberately so: a scroll mode's first frame
is legitimately its blank scroll-in buffer, which is 42 of these on the
F1 scoreboard alone. Plugins whose modes are known to draw on their
fixture data can opt into failing via harness.json {"empty_check":
"strict"}, matching how the fill check is staged.
Worth being clear about the limit: this only sees what the fixtures
render. It would not have caught the sports bug, whose fixture seeds
games so the empty path never renders -- that needs the source-level
gate in the plugins repo. What it does catch is the same mistake in any
plugin whose empty state the harness does happen to reach, which is
coverage there was none of before.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(vegas): stop capping plugin width by default
Vegas plugins read as "cut off early" or "starting in the middle". That
was the per-plugin width budget, not the scroll engine:
overflow_mode=rotate is designed to resume mid-content on each
appearance, so the symptom was the feature working as specified.
Measured over a 17-plugin fleet on a 512px panel, the cap was a bad
trade. Only four plugins were ever wide enough to hit the 3.0 default --
leaderboard 11,518px, news 10,021px, odds-ticker 4,643px, hockey
1,508px. Weather is 650px and flights 512px; the cap never touched them
or the other eleven. So it bought nothing on thirteen plugins while
costing two visible faults on four: content entering mid-item (a news
ticker started at column 6027 of its own strip), and a final rotation
window of whatever happened to be left -- 348px of an 1,840px stocks
ticker, seven seconds of panel time.
Default max_plugin_width_ratio to 0 (uncapped), so every plugin
contributes all of its content and is always entered at its beginning.
The cap remains available, and vegas_max_width_screens still caps an
individual plugin -- which is where the knob belongs, since a genuinely
long ticker is a property of that plugin rather than of the fleet.
Verified on a live 512px device: 327 budget crops in the preceding six
hours, none after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* test(vegas): cover the width cap still working when asked for
Defaulting the cap off must not quietly remove it. Asserts that an
explicit ratio is honoured and validates, alongside the existing check
that omitting it means uncapped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(web): use canonical secret helpers in api_v3; make ConfigManager secret strip/merge array-aware
api_v3.py carried three inline nested copies of find_secret_fields/
separate_secrets (main-config save, plugin-config save, plugin-config
reset). They drifted from each other (one lacked isinstance guards) and
none supported the canonical module's array-item secrets
(accounts[].token). All three endpoints now import from
src/web_interface/secret_helpers.
Adopting the canonical behavior makes array-item secrets reachable, and
their parallel-placeholder shape ([{'token': ...}, {}] alongside the
regular list) was not survivable by ConfigManager's round-trip:
_strip_secrets_recursive dropped the whole key (losing the regular
fields from config.json) and _deep_merge replaced the regular list
wholesale on load. Both are now array-aware:
- strip removes the secret fields from each item and ALWAYS keeps the
list so indices survive for merge-on-load; whole-key secrets (scalar
lists, shape mismatches) still drop the key entirely — never leak.
- merge folds each secrets item into the config item at the same index,
skipping {} placeholders. The regular list's length is authoritative
in both directions: a user deleting an array item never has it
resurrected from a stale secrets entry (extras warn and are ignored).
api_v3's own deep_merge intentionally still replaces lists wholesale —
form posts carry complete arrays and index-merging would resurrect
deleted items; a comment now documents that.
Tests: the parity guard flips from 'exactly 3 inline copies' to 'zero,
and the canonical import must exist'; TestArraySecretStripAndMerge
covers the new strip/merge semantics incl. length-mismatch contracts;
new test_api_v3_secret_roundtrip.py drives all three endpoints through
a Flask client with a REAL ConfigManager+SchemaManager over tmp_path,
proving secrets land in config_secrets.json, config.json stays clean,
and a fresh load merges them back into the right array items.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: repair broken helper paths across display, cache, odds, logging, resolver, repos, config, validator
Nine fixes for bugs surfaced while writing coverage for previously
untested modules (plus the bool-duration quirk pinned in PR #441):
- base_plugin.get_display_duration: exclude bools from both numeric
branches — display_duration=True no longer reads as a 1-second slot;
it falls through to config, then the 15.0 default.
- display_helper: draw_error_message/draw_no_data_message called
_draw_centered_text with the wrong arguments and crashed with
AttributeError — both now delegate to draw_centered_text.
draw_scorebug_layout drew status and clock at the same y, overprinting
each other — they now share one combined top line.
draw_ticker_layout drew its text starting at x=display_width (fully
off-canvas), returning a blank frame every time — now draws at x=0;
scroll_speed stays accepted-but-unused and is documented as such.
- api_helper.clear_cache guarded on a nonexistent CacheManager.clear()
method, silently never clearing anything; it now uses the real surface
(clear_cache/delete/list_cache_files) and no-ops safely otherwise.
- base_odds_manager._extract_espn_data raised AttributeError when ESPN
sent explicit JSON nulls ("homeTeamOdds": null) — every level now
null-safes with 'or {}'. format_odds_summary gated on
is_odds_available, which deliberately ignores money lines, so
ML-only odds formatted as "No odds available" — it now gates only on
empty/no_odds data and formats money lines.
- logging_config.ContextualFormatter mutated record.msg in place, so a
second handler prepended the context prefix twice; it now formats a
copy. log_error hardcoded exc_info=True and raised TypeError when the
caller passed exc_info — now kwargs.setdefault.
- dynamic_team_resolver wrote its "shared" class cache through self,
creating instance shadows — the cache was per-instance and every
scoreboard refetched rankings. Writes now go through the class.
- saved_repositories cleaned URLs with an unanchored .replace('.git','')
that mangled URLs merely containing '.git' (my.github.io -> myhub.io);
now strips only a trailing suffix. add/remove also roll back the
in-memory list when the save fails, so memory always matches disk.
- config_helper.merge_configs shallow-copied the base, aliasing every
un-overridden nested dict into the result — now deep-copies.
- startup_validator.validate_all accumulated errors/warnings across
calls — now resets both lists per run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: cover the previously untested modules
Nine new suites plus an extension, asserting the Phase-1b fixed behavior
and pinning the quirks deliberately left alone:
- test_logging_config.py: formatters (JSON shape, no record mutation,
single prefix through two handlers), PluginLoggerAdapter precedence,
setup_logging handler hygiene and LEDMATRIX_DEBUG, log_error exc_info.
- test_startup_validator.py: exact messages, error-vs-warning split,
accessor split (load_config vs get_config), cache-dir branches with
os.access monkeypatched (root can write anything in CI), idempotence,
raise_on_errors classification precedence.
- test_config_helper.py (full): load/save round trips, dot-notation
get/set incl. silent-failure contract, post-fix no-aliasing merge,
schema validation branches, the '{id}_config' key pin, default-enabled
pin.
- test_saved_repositories.py: three load shapes, bare-list rewrite pin,
trailing-only .git strip (my.github.io regression), save-failure
rollback, type-classification case-sensitivity pin.
- test_api_helper.py: rate-limit math, cache-hit short circuit, ESPN
URL/key formats, exact User-Agent guard, retry adapter, post-fix
clear_cache against the real CacheManager surface, ttl-dropped pin.
- test_base_odds_manager.py: cache-key/URL construction, no_odds
sentinel round trip, stale-cache fallback, null-safe extraction,
ML-only formatting, is_odds_available truth table (ML-blind by
contract), config key/attr mismatch pin.
- test_dynamic_team_resolver.py: expansion/dedup/slicing, dropped
unknown-dynamic names (TOP_ substring hazard pinned), genuinely
shared class cache (second instance: zero HTTP), TTL expiry,
failure degradation without raising.
- test_display_helper.py (full): the fixed error/no-data renders,
combined scorebug top line, non-blank ticker with scroll_speed
no-op pin, composite upconversion, logo bleed positions, square
orientation pin.
- test_skin_runtime_cache.py: discovery-cache hit/invalidation
semantics (manifest mtime, .py edits pinned as non-invalidating),
sys.modules namespacing contract incl. bare-name restore and stdlib
shadowing, entry-module execute-once, API minor-version tolerance,
skin_matches_target table.
- test_sports_capabilities.py (extended): _draw_celebration_layout
executed for real (flash window, matrix-dims fallback, highlight
alternation, logo-failure isolation), _should_celebrate_for direct,
strict duration boundary, score_to_int edges, both-teams-score
precedence, expired-coalesce refire, disabled-win baseline
preservation, id-less prune.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: real schedule/dim coverage for DisplayController; fix two vacuous schedule tests
New test_display_controller_schedule.py drives _check_schedule and
_check_dim_schedule on a bare controller stub: same-day and
midnight-crossing windows with inclusive boundaries, global vs per-day vs
legacy-inferred modes (and dim's global-only default — no legacy
inference), per-day disabled days, invalid %H:%M fallbacks, unknown
timezone -> UTC, dim_brightness default 30, inactive-display short
circuit, and the _was_display_active/_was_dimmed transition flags.
test_display_controller.py's test_schedule_disabled and
test_active_hours patched config_service.get_config — which
_check_schedule never reads — so both asserted the init-default value
and could not fail. Rewritten on the test_inactive_hours pattern
(inject controller.config['schedule'], reset the minute gate, flip the
flag to the opposite state first so the assertion has teeth).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* ci: raise coverage floor to 48%
Measured 50% with the new suites in place (was 47% baseline when the
gate was introduced at 45); floor stays two points under measured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: address CodeQL alert and review findings
- config_manager: the "secrets list longer than config list" warning now
interpolates only config-side data (no key name or secrets-derived
values), resolving the CodeQL clear-text-logging alert.
- base_plugin: validate_config rejects bool display_duration, matching
get_display_duration (bool is an int subclass and would otherwise pass
as a positive number).
- config_helper: merge_configs deep-copies override values in the
non-recursive branch so mutating the merged result cannot reach back
into override_config.
- saved_repositories: saves are atomic (temp file + fsync + os.replace),
so a failed write can no longer truncate saved_repositories.json.
- tests: regression cases for each fix, plus a pin that whole-item
array secrets (key[] + key[].field both marked) strip to empty {}
skeletons — no secret values can reach config.json.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
---------
Co-authored-by: Claude <noreply@anthropic.com>
Reported from a pi whose checkout sat on a local branch:
git pull failed (returncode=1): There is no tracking information for
the current branch. Please specify which branch you want to rebase
against.
The Tools tab reported that as "Update failed; check logs for details",
which tells the user nothing they can act on, and the underlying git
message never reached the UI at all.
A branch with no upstream is easy to end up on — checking one out by
name, restoring a backup, or following a guide that names a branch — and
until now it left the update button permanently broken with no way out
except SSH.
resolve_pull_command() now decides how to pull:
- upstream set -> git pull --rebase, as before
- no upstream, origin/<branch> -> git pull --rebase origin <branch>,
then attach tracking so the next update is a plain pull
- no upstream, no remote branch -> an error naming the branch and
pointing at Switch branch
- detached HEAD -> says so, rather than failing obscurely
That resolution happens BEFORE the stash. Previously the handler stashed
local changes and then discovered it could not pull, putting the user's
work away for an update that was never going to run.
Failures now surface git's own message instead of "check logs".
Adds a branch picker to the Tools tab, backed by GET
/system/git-branches (local + remote-only) and a checkout_branch action.
Switching attaches tracking, so Pull Latest works afterwards. Branch
names are validated against a strict pattern before reaching a subprocess
argument list.
Local edits block a checkout, as they should. Rather than a truncated
one-line error, the response carries git's full list of blocking files
and a can_retry_with_stash flag; the UI then offers "Stash and switch" as
an explicit choice. Stashing is never done unasked — putting someone's
edits away without consent is worse than refusing the switch.
Verified on the pi that produced the report: on its untracked 'audit'
branch the update now returns the actionable message, git-info reports
upstream='' and can_pull=false, and an injected branch name is rejected.
27 tests build real git repositories and cover each path, including the
stash route that could not be exercised safely on the device.
* feat(web): let schemas label enum dropdown options
An enum property renders as a dropdown whose option text is derived from
the value — underscores replaced, title case applied. That works when the
value reads as its own label and fails when it does not: "vs" renders as
"Vs", and "abbrev" tells the user nothing about the "Sep 19" it produces.
Schemas had no way to say otherwise, so the label was whatever the config
key happened to look like.
Enum dropdowns now take their option text from x-options.labels when the
schema supplies it. This is not a new convention: the checkbox-group
widget has read x-options.labels since it was written, with the same
humanised fallback. This extends it to plain enums and to array-table
columns.
Display only — the option value, and so the saved config, is unchanged.
The map may be partial; unlabelled values keep the humanised fallback, so
every existing schema renders exactly as before. Older cores ignore
x-options entirely, which means a plugin can ship labels without
requiring users to upgrade first.
Array-table columns get the same lookup but keep the raw value as their
fallback rather than the humanised one. Those columns hold values such as
ticker symbols, where "aapl" -> "Aapl" would be wrong, and they were not
being humanised before this change.
Verified against the running web service: with labels the hockey plugin's
date dropdown reads "Sep 19 / 9/19 / 19 Sep / 19/9 / Fri Sep 19"; with
the pre-change template and the same schema it falls back to
"Abbrev / Numeric / Day First / ...", confirming the degradation path.
* fix(web): label enum options in dynamically added table rows, and test the
shipped template
Both points from the CodeRabbit review on #442.
array-table.js built enum <option> elements with o.textContent = opt, so a
row added with "Add row" showed the raw value while the server-rendered
rows above it showed the schema's label — the same column reading two
different ways until the page was reloaded. Both option-building sites now
go through a shared enumOptionLabel(), which mirrors the template exactly:
x-options or x_options, labels map, raw value as the fallback.
The tests rendered a copy of the template expression, so they could pass
while production drifted. They now extract the live enum <select> block
out of plugin_config.html and render that, and assert on the full
value -> label map rather than substring presence.
Mutation-checked, since a guard that cannot fail is not a guard:
- remove the labels lookup -> 5 of 9 fail
- change only the fallback to -> 2 of 9 fail
option|upper (keeping the
enum_labels.get call intact)
- revert the JS to raw values -> 1 of 9 fails
The middle case is the one the review called out as able to slip through.
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(backup): make a restored device match the one that was backed up
Found by wiping a working device and reinstalling from scratch. Every
problem here is invisible until you actually do that, which is why a
green test suite and eleven hours of uptime had not surfaced any of them.
**Four enabled plugins vanished on restore.** Weather, stocks, music and
leaderboard have a registry `id` that differs from the `id` in their own
manifest: the registry calls them `weather`, everything else calls them
`ledmatrix-weather`. Installation already prefers the manifest id for the
directory name and warns when the two disagree, so on disk, in
config.json and in a backup they are `ledmatrix-weather` -- but nothing
resolved that in reverse. Restore asked the store for `ledmatrix-weather`
and got "Plugin not found in registry", four times, and the device came
back missing four plugins the user had enabled.
Registry lookup now falls back to matching `plugin_path`, which already
records `plugins/ledmatrix-weather`. Renaming the published ids would
have orphaned `plugin_state.json` entries keyed on the old ones. Exact id
still wins, so a path that collides with another entry's id cannot
shadow it. Against the live registry and a real 28-plugin install this
takes unresolvable directories from five to one -- the one being
starlark-apps, which is genuinely not in the registry.
**Secrets could not be restored at all.** A fresh install left
config_secrets.json group-readable but not group-writable, and the web
interface -- which is what performs a restore -- does not necessarily run
as the owner. Every other file in the backup restored; secrets failed
with EACCES. Now group-writable, so the account running the web UI can
put them back.
**A partial restore reported "Restore had errors" and nothing else.**
That is the same message whether the whole thing failed or it quietly
dropped your API keys. It now names what was restored, what failed, and
which plugins were not reinstalled.
**ytm_auth.json was never in the backup.** It sits in config/ beside the
three files that are, and is pure device-local auth: losing it silently
signs the user out of YouTube Music. Backed up and restored with the
wifi config, which it resembles.
**Backups were written inside the directory a reinstall deletes.**
config/backups/exports is destroyed by the reinstall the user was told to
make it before. Exports now go beside the install, falling back to the
old path when that is not writable.
**The installer reboots without asking in non-interactive mode**, which
the README did not mention -- easy to hit when piping the install, and
alarming when a device you are installing onto disappears. Documented,
with --no-reboot-prompt. Its log also claimed root:ledmatrix while
printing a hardcoded group name rather than the one it used.
Tests: registry resolution gets its own suite, including the collision
case and third-party entries with an empty plugin_path. The existing
round-trip test passed throughout this because its fixture plugin has a
directory name equal to its id -- the one shape that cannot fail -- so it
now carries ytm_auth too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* ci: run the backup suites
test_backup_manager.py existed but was never enrolled, so the tests that
should have guarded backup and restore have not run on a pull request.
That is part of why the restore bugs in the previous commit reached a
device: the suite was there, it just was not watching. Adds it alongside
the new registry-resolution tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(backup): restore must not depend on owning the file it replaces
Follow-up from testing the previous commit on real hardware, where the
secrets fix turned out to be both too narrow and slightly wrong.
Too narrow: config.json, wifi_config.json and ytm_auth.json are installed
root-owned and group-readable exactly like the secrets file, so all four
were unrestorable by the web service, not just one. `shutil.copy2` opens
the destination for writing, which needs permission on the *existing
file*; the web user could create files in that directory all day and
still not replace them.
Slightly wrong: the previous commit loosened the secrets file to
group-writable. That was treating the symptom. The real error was
deciding ownership from `ledmatrix.service` -- the display service, which
runs as root and only ever *reads* secrets -- when the account that
*writes* them is the web interface, which deliberately does not run as
root. Ownership now follows the web service's user and the mode stays
640.
`_copy_file` writes a temporary file alongside the target and renames
over it. That needs only directory permission, so a restore no longer
cares who owns the destination, and it is atomic: a crash mid-restore can
no longer leave a half-written config. The destination's mode is carried
across so restoring secrets does not widen them to the umask, and its
owner is carried across too when the OS allows it -- only root can hand a
file to another user, so a restore run by the web service keeps its own
ownership rather than pretending to preserve root's.
Verified on a device with all four config files set root-owned 640 and
unwritable by the web user: before, every one failed with EACCES; after,
the restore reports success with no errors and all four sections
restored, mode still 640, root still able to read them and the web
service still able to write them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(backup): address CodeRabbit and CodeQL findings on PR #439
- first_time_install.sh: verify chown/chmod succeed and the final
owner/group/mode on config_secrets.json before reporting success;
exit with a clear error otherwise instead of swallowing failures.
- api_v3.py: replace the predictable .writetest probe with an
exclusive NamedTemporaryFile to avoid a race with concurrent
resolvers; log the preferred/fallback export path and OSError when
falling back to the reinstall-deleted directory.
- api_v3.py: mark a restore as failed when plugin reinstalls fail,
even if file restoration itself succeeded, so the endpoint no longer
reports HTTP 200 success on a partial restore.
- api_v3.py: stringify plugin IDs before joining them into the error
message so a malformed backup's non-string plugin_id can't raise a
TypeError and mask the detailed response.
- backup_manager.py / api_v3.py: stop putting raw exception text (originating
from a user-controlled backup file) into restore results returned to
the client; log full details server-side instead. Addresses the
CodeQL "stack trace information exposure" alert.
- test coverage: add a test for get_plugin_info() resolving a
manifest id, and assert the disabled restore_wifi path also skips
and omits ytm_auth.json.
Co-authored-by: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: run the whole test tree and make the plugin-safety job assert something real
The unit-tests CI job ran an explicit 24-file allowlist that had rotted:
63 of 90 test files (display, vegas, store manager, web API, web_interface)
never ran on a PR. The job now runs all of test/ (minus test/plugins, which
the plugin-safety job owns) so new test files are enrolled by default and
any exclusion needs a visible, commented --ignore.
The plugin-safety job was a green no-op: plugins/ is empty in CI, so every
test skipped with 'Manifest not found'. It now renders a bundled
deterministic fixture plugin (test/fixtures/plugins/ci-fixture-plugin,
golden images included for all 8 default sizes) via LEDMATRIX_PLUGINS_DIR,
and sets LEDMATRIX_REQUIRE_PLUGINS=1 so discovering zero plugins fails
loudly instead of skipping green. The per-plugin suites document that they
target dev machines with real plugins installed.
Coverage is now measured and enforced in exactly one place — the CI
unit-tests step (--cov=src --cov=web_interface --cov-fail-under=45, from a
measured 47% baseline). pytest.ini previously declared --cov-fail-under=30
but CI always passed --no-cov, so the gate had never run anywhere; local
pytest is now coverage-free and fast.
Enabling the 63 unenrolled files surfaced three cases of test rot, fixed
here: test_display_controller_vegas_tick.py could not collect without the
hardware rgbmatrix module (now uses the emulator convention), the
state-reconciliation unrecoverable-cache tests broke when production added
the is_plugin_uninstalled tombstone check (bare Mock returned truthy),
and test_get_system_status assumed the optional psutil dependency
(now installed via requirements-test.txt and guarded by importorskip).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: replace can't-fail tests with real assertions
test_font_manager.py was 5 of 6 tests shaped as 'try: call(); assert True /
except: assert True' — running in CI while unable to fail on any
regression. Rewritten against the real FontManager API and the bundled
assets/fonts: returned font types, cache-hit identity, distinct entries per
size, default-font fallback for unknown families and corrupt files
(recorded in failed_loads), BDF native-size reading, text measurement, and
cache lifecycle.
test_display_manager.py's test_draw_text ended in 'assert True'; it now
renders onto a known-black canvas and asserts pixels were actually lit —
which required un-breaking the fixture's freetype MagicMock so draw_text's
isinstance check doesn't silently swallow the draw.
test_display_controller.py carried a permanently-skipped test whose skip
reason already declared it redundant; deleted.
Both display test files now set EMULATOR=true before importing
display_manager (the same convention as test_display_dirty_tracking.py) so
they collect standalone instead of depending on which test module imports
display_manager first.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: cover the untested fragile logic (compatibility gate, secrets, config merges, durations, skin cards)
New unit tests for pure or filesystem-only logic that previously had zero
direct coverage:
- test_compatibility.py: the semver install gate (parse_semver suffix
handling, every range operator, TRUSTWORTHY_FLOOR behavior for cores
reporting untrustworthy versions, 'more restrictive wins', and the
malformed-manifest shapes that used to raise).
- test/web_interface/test_secret_helpers.py: the canonical x-secret
helpers — find/separate/mask/remove, array-item secrets, no input
mutation, and a separate->recombine round-trip.
- test/web_interface/test_api_v3_helpers.py: the module-level helpers
behind the plugin config save endpoint (_is_plugin_update_available,
_coerce_to_bool including the int==1 quirk, deep_merge including its
shared-subtree shallowness, _parse_form_value, dotted-key-aware
_get_schema_property/_set_nested_value).
- test_base_plugin_duration.py: get_display_duration's full coercion
ladder (instance attr -> config -> 15.0), including the bool-is-int
quirk where display_duration=True means one second.
- test_config_manager_secrets.py: the secrets round-trip — deep-merge on
load, strip on save, group pruning, the load fast path — and two
characterized sharp edges marked SUSPECTED BUG: an unreadable secrets
file at save time writes secrets into config.json in plaintext, and a
same-mtime-same-size content swap is served stale.
- test_schema_manager_merge.py: merge_with_defaults branch behavior (None
replacement vs falsey preservation, dict-vs-scalar mismatches, arrays
replaced wholesale, defaults never mutated).
- test_skin_system.py (extended): render_skin_card shares _render_game's
3-strike counter but never resets it on success — the asymmetry is
pinned in both directions, along with card fallthrough and the disable
interaction between the two paths.
Suspected bugs are characterized, not fixed — each carries a comment so a
future behavior change is deliberate rather than accidental.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: add drift guards for cross-file contracts
Three guard suites that pin contracts spanning multiple files, where one
side changing unilaterally breaks the other silently:
- test_version_comparison_consistency.py: the repo's four version
comparators (compatibility.parse_semver, api_v3's packaging-based
_is_plugin_update_available, store_manager update_plugin's raw string
equality, skin_runtime._major) answer differently on the same inputs.
A table pins each one's verdict; update_plugin is driven through its
real code path to show the SUSPECTED BUGs: 'v1.2.0' vs '1.2.0'
triggers a full reinstall the UI calls unnecessary, and a locally-ahead
plugin gets downgraded. A pairwise-ordering check keeps parse_semver
agreeing with packaging on plain X.Y.Z.
- test/web_interface/test_secret_separation_parity.py: api_v3.py carries
three inline copies of find_secret_fields/separate_secrets that lack
the canonical module's array-item support. The copy count is asserted
exact (it may only go down; new copies must import
src/web_interface/secret_helpers), the missing-array-support gap is
asserted so it can't grow silently, and the canonical behavior that
migration will adopt is documented executably.
- test_discovery_path_contract.py: the three 'where is plugin X'
resolvers (PluginManager discovery, StoreManager._find_plugin_path,
SchemaManager.get_schema_path) agree on the configured directory, and
their divergent fallback chains are characterized. Also pins the
.standalone-backup- naming contract shared by store rollback and
discovery, and _resolve_skin_target's path-traversal rejection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: address review feedback — fixture lifecycle, test names, ClassVar
- ci-fixture-plugin: call display_manager.clear() before rendering (per
plugin guidelines — the fixture should model a well-behaved plugin),
add a class docstring, and document why Pillow is deliberately not
pinned in its requirements.txt (core dependency; harness installs
nothing).
- Rename two tests whose names contradicted their assertions:
test_unparseable_core_version_is_compatible ->
test_unparseable_core_with_high_floor_is_blocked, and
test_unreadable_secrets_file... -> test_corrupt_secrets_file...
- Annotate TestGetSchemaProperty.SCHEMA as ClassVar (RUF012).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* ci: allow manual test.yml runs via workflow_dispatch
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: unify version comparison, refuse secret-leaking saves, reset skin strikes on card success
Fixes the three suspected bugs this PR's characterization tests pinned,
flipping those tests to assert the corrected behavior:
- plugins/store: ONE shared update comparator. New
compatibility.is_update_available() (PEP 440 via packaging) is now used
by both the web UI's update badge (api_v3._is_plugin_update_available
is a thin alias) and store_manager.update_plugin's reinstall decision.
Previously update_plugin used raw string equality: 'v1.2.0' vs '1.2.0'
triggered a full reinstall the UI called unnecessary, and a locally-
ahead plugin (2.0.0 installed, registry 1.9.0) was silently DOWNGRADED.
Now equivalent spellings skip the reinstall and locally-ahead versions
are never downgraded; unparseable versions still reconcile by
reinstalling from the registry.
- config: save_config and save_config_atomic now refuse (ConfigError)
when config_secrets.json exists but cannot be loaded. Both previously
proceeded without stripping, writing the merged secrets into
config.json in plaintext. The shared _load_secrets_for_save() helper
raises with an actionable message instead; a missing secrets file is
still fine (nothing to strip), and _migrate_config's catch-all keeps
boot resilient.
- skins: render_skin_card resets _skin_failures on both success paths
(vegas card returned, or mode renderer handled), mirroring
_render_game. Transient card failures no longer accumulate across a
session until they permanently disable a working skin.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: harden shared comparator edges from review
- is_update_available: reject truthy non-string versions (a malformed
manifest can carry a number; packaging raises TypeError on those) by
surfacing the mismatch instead of raising.
- store_manager.update_plugin: drop the truthiness gate around the
comparator so a missing version on either side follows the shared
'no update' verdict, keeping the store consistent with the UI badge;
a missing manifest still uses the reinstall recovery path.
- config_manager._load_secrets_for_save: catch only expected read/parse
failures (OSError/ValueError/RecursionError) so implementation bugs
propagate as themselves, and log with traceback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(web): implement delete_cached so the font catalog cache actually invalidates
api_v3.py's font upload/delete handlers import delete_cached from
web_interface.cache, but the function was never defined. The surrounding
except ImportError silently swallowed the failure, so the fonts_catalog
cache entry survived uploads/deletes and newly uploaded fonts did not
appear until the TTL expired or the service restarted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(web): remove dead weather/stocks partial routes that returned 500
The partial dispatcher still routed 'weather' and 'stocks' to loaders
rendering v3/partials/weather.html and stocks.html — templates that no
longer exist since weather and stocks became store plugins. Requesting
either partial raised TemplateNotFound, which the catch-all turned into
a 500. No template or JS references these partials (the only 'weather'
hit in the front end is a plugin-store category filter option), so the
branches and both loader functions are removed; unknown partials now
fall through to the existing 404 handler.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(deps): align contradictory psutil/Flask-Limiter/freetype-py pins
requirements.txt's optional-install comment recommended psutil>=5.9,<6.0
while web_interface/requirements.txt hard-requires >=6.0,<7.0 — anyone
following the comment ends up with an unsatisfiable pair. The comment now
recommends the same range the web interface requires (all psutil APIs
used — Process, boot_time, cpu_percent, disk_usage, virtual_memory — are
stable in 6.x). Flask-Limiter gains the same <4.0 cap in both files and
freetype-py the same >=2.5.1 floor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(config): add template keys the code already reads
display.hardware gains pixel_mapper_config, row_address_type,
multiplexing and panel_type (read at display_manager.py with these exact
fallbacks — users on non-standard panels previously had no way to
discover them from the template). vegas_scroll gains
frame_based_scrolling and scroll_delay, the only two of its 27 keys the
template omitted (read in src/vegas_mode/config.py). plugin_system gains
development_mode, which the web UI reads and writes but the template
never declared.
Every added value is byte-identical to the code-side .get() fallback, so
ConfigManager._migrate_config() merging these keys into existing user
configs cannot change behavior on any installed device.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(scripts): repair broken sys.path setup in utility scripts
clear_cache.py and download_nba_logos.py pointed sys.path at a 'src'
directory relative to the script's own folder (scripts/utils/src and
scripts/src — neither exists), so both crashed on import; they now insert
the project root and import via the src package like the other scripts.
debug_web_manual.py resolved 'project root' to scripts/debug/ instead of
two levels up. fix_nhl_cache.sh is removed: it used Python docstring
syntax in a bash script and invoked clear_nhl_cache.py, which does not
exist anywhere in the repo — it cannot ever have worked in its current
location.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: correct stale file:line references and the loader-fallback contradiction
CLAUDE.md and .cursorrules disagreed about plugin-directory fallback
behavior; the code (SchemaManager.get_schema_path) probes plugins/
BEFORE plugin-repos/, and the main discovery path has no fallback at
all — both files now describe the real behavior, preferring symbol names
over line numbers so the references rot slower. REST_API_REFERENCE.md
pointed at app.py:144/:607 for mounts that live at :199/:799 and counted
92 routes where there are 94. PLUGIN_ARCHITECTURE_SPEC.md's historical
banner gains a note that its example imports
(src/plugin_system/base_classes/*_plugin.py) never shipped — the real
base classes are src.base_classes.sports.SportsCore and
src.base_classes.hockey.Hockey.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: fix broken links, phantom script references, and stale CI description
Repairs every broken relative link in active docs (targets renamed or
archived long ago: PLUGIN_DEVELOPMENT.md -> PLUGIN_DEVELOPMENT_GUIDE.md,
API_REFERENCE.md -> REST_API_REFERENCE.md, PLUGIN_STORE_USER_GUIDE.md ->
PLUGIN_STORE_GUIDE.md, plugin_docs/ dir, TROUBLESHOOTING_QUICK_START.md,
and MIGRATION_GUIDE's README link that silently resolved to the docs
index instead of the project README). Replaces commands invoking scripts
that do not exist (scripts/update_stats.py, validate_registry.py,
check_updates.py, fix_permissions.sh) with the real tooling, and
rewrites HOW_TO_RUN_TESTS.md's CI section, which described a
security-audit workflow that was never committed and a pytest workflow
'queued to land' that landed long ago as test.yml.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: complete the docs index and refresh the web interface file tree
docs/README.md's own policy says every page must be linked from the
index, yet five weren't — including the entire skin system
(SKIN_SYSTEM.md, CREATING_SKINS.md), ADAPTIVE_LAYOUT.md,
plugin-safety-harness.md and SPORTS_UNIFICATION.md. Each is now listed
in the section it belongs to, and PLUGIN_ARCHITECTURE_SPEC.md is marked
historical in the index (the doc itself already carries the banner).
web_interface/README.md's static/v3 tree showed only app.css/app.js;
it now reflects the actual contents.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: remove dead modules confirmed unused in-repo and across all store plugins
- src/common/cli.py: imports a 'ledmatrix_common' package that exists
nowhere (not in this repo, any requirements file, or the plugin
monorepo), so it cannot ever have run; its README section claimed
scripts/dev/* used it, which was also untrue.
- src/web_interface/logging_config.py: zero callers — the web app uses
web_interface/logging_config.py (a different module), and nothing
imports the src copy.
- handle_errors decorator in src/web_interface/error_handler.py: zero
call sites (the module's response helpers stay — they are used).
- ConfigManager.get_clock_config(): reads a 'clock' config key that no
longer exists anywhere; only caller was its own unit test.
Deliberately kept despite zero in-repo callers: DisplayError,
src/common/config_helper.py and display_helper.py — all documented as
plugin-facing API (docs/PLUGIN_ERROR_HANDLING.md, src/common/README.md),
and third-party plugins outside the official monorepo cannot be
enumerated. Verified against a fresh clone of ledmatrix-plugins (43
plugins): zero references to any removed symbol.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: remove manager-era NBA test files and one-off debug scripts
The four test_nba_*.py files imported nba_managers, leaderboard_manager
and odds_manager — top-level modules deleted when sports displays became
plugins — inside try/except blocks that swallowed the ImportError, so
they passed while exercising nothing. test_nba_data_structure.py and
debug_nba_api.py (a diagnostic script living in test/) made live ESPN
API calls rather than testing repo code. None were enrolled in CI.
scripts/debug/direct_fix_imports.py and check_imports.py were one-shot
artifacts that edited/inspected a hardcoded ~/LEDMatrix/web_interface/
app.py to fix an import problem solved long ago; nothing references
them.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: remove generate_report.py, which aggregates artifacts of CI jobs that do not exist
The script's only function is to merge JSON artifacts
(bandit/semgrep/pip-audit/safety/gitleaks results) produced by a
security-audit workflow that was never committed —
.github/workflows/ has no such jobs, so there is nothing for it to
aggregate and no way to run it usefully. Its siblings stay:
prove_security.py and audit_plugins.py both run standalone (verified),
and .codacy.yml stays because the Codacy service (README badge) reads it
server-side without a workflow file.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(web): load the three widget scripts store plugins already declare
time-picker.js, file-upload-single.js and plugin-file-manager.js
register widgets that installed store plugins reference in their config
schemas (countdown uses x-widget: time-picker and file-upload-single;
of-the-day uses plugin-file-manager), but base.html never included the
scripts. plugin_config.html renders such fields as an empty container
that polls LEDMatrixWidgets.get(...) on a 50ms loop forever, so those
plugin config fields appeared permanently blank. The audit initially
flagged these files as dead code; the monorepo cross-check proved the
opposite — they were unreachable, not unused.
example-color-picker.js (the documented custom-widget example) gains an
explicit warning that including it in base.html would shadow the
built-in color-picker widget.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: drop the legacy youtube block from the secrets template
No code reads a top-level youtube secrets key: the youtube-stats plugin
receives its API key namespaced under its own plugin id (declared via
x-secret in its config schema), like every other store plugin. The key
survives only in state_reconciliation.py's non-plugin-key exclusion set,
which stays — existing installs still carry the key in their generated
config_secrets.json, and the exclusion prevents it from being
misclassified as a plugin config. New installs simply stop being asked
for a YouTube API key they have nowhere to use.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore(deps): remove packages nothing imports, declare direct imports, move mypy to test deps
Removed from requirements.txt: python-socketio, python-engineio,
websockets, websocket-client — zero imports anywhere in this repo, and
the one store plugin that needs Socket.IO (ledmatrix-music) declares it
in its own requirements.txt, which the plugin store installs. Removed
the same quartet plus timezonefinder, geopy, google-auth-oauthlib,
google-auth-httplib2, google-api-python-client, unidecode, icalevents,
python-dateutil, flask-wtf and the werkzeug pin from
web_interface/requirements.txt — all leftovers from the deleted built-in
weather/calendar/music displays (flask-wtf was doubly dead: app.py
explicitly disables CSRF and sets csrf=None). scripts/
install_dependencies_apt.py, which mirrors these lists for the
first-time installer, drops the same packages.
Added: urllib3 (imported directly in four core modules), jinja2 and
markupsafe (imported directly in pages_v3.py) — previously reachable
only as transitives. mypy moves from runtime requirements to
requirements-test.txt.
Verified in a fresh venv: all four requirements files co-install, pip
check is clean, the full CI-enrolled suite (907 tests) and a Flask boot
smoke pass with the trimmed dependency set.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* refactor: single canonical DateTimeEncoder
src/cache_manager.py and src/cache/disk_cache.py each defined an
identical DateTimeEncoder (datetime -> ISO-8601). The disk_cache copy is
the only one actually used for serialization; cache_manager now
re-exports it instead of defining a twin, so the two can never silently
diverge. Import compatibility is preserved — from src.cache_manager
import DateTimeEncoder still works and is the same class object.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs(code): document deliberate duplicates instead of merging them
The audit surfaced several near-duplicate implementations that turned
out to be either deliberate forks or behaviorally different — merging
any of them would risk changing behavior on installed devices, so each
now carries an explicit comment stating the relationship:
- VisualDisplayManager: headless fork of DisplayManager; header now
lists the ~15 mirrored methods and warns that DisplayManager changes
must be mirrored.
- normalize_abbreviation: LogoDownloader's version (called directly by
nine scoreboard plugins) replaces filesystem-unsafe characters;
LogoHelper's strips spaces. Logo filenames on existing installs
depend on both behaviors staying put.
- The two PluginTestBase classes: the shipped one is plugin-author
API, the repo's own richer harness lives in test/plugins/ — now
cross-referenced.
Also verified (no change needed): ConfigManager's backup/rollback
methods genuinely delegate to AtomicConfigManager, and SportsCore
already delegates _read_bdf_native_size to FontManager.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: add a unified configuration reference
There was no single place documenting what lives in config.json —
display.* keys were scattered across README sections, vegas_scroll lived
in ADVANCED_FEATURES.md, and dim_schedule, display.double_sided,
sync.follower_position, plugin_system.development_mode and the four
newly-templated hardware keys were documented nowhere. CONFIG_REFERENCE.md
now lists every template key plus the code-read-only keys, each with
type, default, and the code location that reads it, and explains the
secrets file's plugin-id namespacing. Linked from the docs index.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: bring the README's feature tour into the plugin era
The Core Features section still presented clock/weather/sports/stocks/
music displays as built into the project, when all of them are store
plugins installed from the ledmatrix-plugins monorepo — only
starlark-apps and web-ui-info ship in this repo. The intro now says so
(the showcase itself is unchanged; those are real displays available in
the store). The display_durations reference drops its built-in-calendar
example in favor of plugin-id keys, and the Configuration section links
the new CONFIG_REFERENCE.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: archive the custom-icons status report, cross-link config docs, document assets/
PLUGIN_CUSTOM_ICONS_FEATURE.md was a 'What Was Implemented' status
report duplicating the actual guide (PLUGIN_CUSTOM_ICONS.md) — moved to
docs/archive/ per the docs index's own policy. The overlapping
plugin-config docs keep their content but PLUGIN_CONFIG_ARCHITECTURE.md
now states up front which doc is canonical for which purpose.
assets/README.md is new and load-bearing: assets/stocks, weather,
news_logos and broadcast_logos have zero references in this repo's code,
which makes them look deletable — but store plugins (ledmatrix-stocks,
ledmatrix-weather, news, odds-ticker) resolve those exact paths at
runtime against the install directory. The README records that evidence
so a future cleanup doesn't break installed plugins.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* test: add regression guards for the bug classes fixed in this PR
Three lightweight static checks, all enrolled in CI's unit-test
allowlist along with the new web-cache test:
- test_template_targets.py: every literal render_template() target must
exist (would have caught the weather/stocks partial 500s at commit
time).
- test_widget_scripts.py: every widget JS file must be script-included
in base.html or explicitly allowlisted with a reason (would have
caught the unloaded time-picker/file-upload-single/plugin-file-manager
widgets), and allowlisted files must NOT be included (prevents the
example widget from shadowing the real color-picker).
- test_doc_links.py: relative markdown links in active docs must
resolve (docs/archive/ exempt).
Each guard was verified to fail against the pre-PR tree and pass now.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(deps): restore the werkzeug version floor
Commit 1ec22db removed the werkzeug>=3.1.6,<4.0.0 pin along with the
genuinely-unused packages, but this one was a version floor on Flask's
transitive dependency, not a phantom: Flask 3.1.3 itself only requires
werkzeug>=3.1.0, so dropping the pin let fresh installs resolve
3.1.0-3.1.5. Restored with a comment explaining why it exists.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(web): route pixel_mapper_config into display.hardware; guard time-picker registration
pixel_mapper_config was the only display.hardware key absent from both
the display_fields detection allowlist and the hardware write loop in
the settings save path. No form posts it today, but if one ever did the
key would fall through to the generic handler and land at the TOP level
of config.json — where state_reconciliation would mistake it for a
missing plugin id and loop auto-repair attempts (the failure class the
'github'/'youtube' exclusion comment documents). It now round-trips
into display.hardware like its siblings.
time-picker.js gains the same LEDMatrixWidgets-undefined guard its two
sibling widgets already have; correct today only via defer ordering.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: align installer leftovers with the dependency cleanup
first_time_install.sh's fallback secrets heredoc (used only when the
template is missing) still wrote the legacy youtube block — now matches
the template (github only). install_dependencies_apt.py drops the
IMPORT_NAME_MAP entries for packages no longer in its install lists and
a stale google-api reference in a docstring.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix: address CodeRabbit review findings
Verified each finding against the code; fixes for the valid ones:
- install_dependencies_apt.py: the installer listed 'freetype', but the
declared dependency is freetype-py — an apt miss would pip-install the
wrong PyPI package. Now installs freetype-py with an import-name
mapping (pre-existing bug, surfaced by the review).
- api_v3.py: pixel_mapper_config is validated as a string before being
saved to display.hardware (JSON callers could previously store an
object/list the matrix library can't use).
- .cursorrules: the Plugin Loading Process and File Organization
sections still said discovery scans plugins/ — now consistent with the
corrected overview (configured directory, default plugin-repos/).
- README.md: removed the stale '(except the core calendar)' claim — no
core calendar exists in src/ — and qualified the plugin inventory
(official plugins in the monorepo; third-party from their own repos).
- CONFIG_REFERENCE.md: hardware_mapping now shows the code fallback
(adafruit-hat-pwm) alongside the template value.
- PLUGIN_REGISTRY_SETUP_GUIDE.md: check_plugin.py takes --plugin, not a
positional id.
- scripts/fix_perms/fix_*.sh: exec bits set so the documented
'sudo ./...' invocations work.
- Guard tests hardened: template guard now catches multi-line
render_template() calls; widget guard parses actual <script> src
values and fails if the widgets dir goes missing; type hints and
docstrings added per repo coding guidelines.
Skipped with reasons (noted on the PR): limit_refresh_rate_hz 100-vs-90
is documented as intentional in CONFIG_REFERENCE.md; the psutil comment
already names the enforcing manifest; docs/archive/ findings are out of
scope per the docs policy (archive may rot).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: remove Cursor IDE tooling, consolidate its guidance into CLAUDE.md
The maintainer no longer uses Cursor. .cursorrules, .cursorignore and the
.cursor/ tree (rules, plugin templates, a parallel 751-line plugins
guide) are removed; measurement showed near-zero literal overlap risk —
the canonical content already lives in docs/. Unique guidance worth
keeping moved before deletion:
- CLAUDE.md gains the dev workflow (dev_plugin_setup.sh, dev_server.py,
run.py -e, check_plugin.py), the plugin-secrets namespacing contract,
and the no-draw_image()/paste-onto-PIL pitfall.
- PLUGIN_DEVELOPMENT_GUIDE.md absorbs the plugin version-management
rules (pre-push hook install, SKIP_TAG, version resolution order) that
its own text previously linked out to .cursorrules for.
- The one completed plan doc (.cursor/plans/) is archived to
docs/archive/ per the docs policy rather than deleted.
One of the deleted rule files (sports-managers.mdc) targeted
src/*_managers.py globs that have matched nothing since the plugin
migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: second-pass cleanup — dead installer branch, broken script, orphaned JS, misfiled test deps
- first_time_install.sh: removed the pip fallback branch that installed
from requirements_web_v2.txt — a file that has not existed since the
v2 web interface was removed (the branch always printed its own
'not found; skipping' warning).
- scripts/remove_plugin_backups.sh deleted: its PROJECT_ROOT resolved to
the repo's PARENT directory, and its verify_submodules() checks for
plugin submodules from an era before plugins moved to the store — it
could never have worked from its current location.
- plugins_manager.js: removed three functions with zero call sites
anywhere (addKeyValuePair, formatCommit, togglePasswordVisibility) —
verified against all templates, all JS, and the dynamic window[name]
dispatch sites, which resolve widget-registry keys only. Also replaced
base.html's misleading 'Legacy ... during migration' label: the file
is deliberately loaded last and provides the LIVE implementations of
seven window.* plugin actions that shadow same-named definitions in
app.js/app-shell.js.
- pytest/pytest-cov/pytest-mock moved from runtime requirements.txt to
requirements-test.txt (CI already installs both files; the installer's
line-by-line loop simply installs three fewer packages on devices; no
store plugin declares pytest). HOW_TO_RUN_TESTS.md updated.
- scripts/add_defaults_to_schemas.py and analyze_plugin_schemas.py
scanned the empty legacy plugins/ dir — now scan plugin-repos/.
Verified: fresh venv installs all four requirements files with pip check
clean and pytest available; bash -n on the installer; node --check on
the JS; widget/cache guard tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: correct semantically stale content across the user and developer guides
A second-pass content audit checked the guides' substantive claims
against the code (the first pass only fixed mechanical drift). Fixes:
- GETTING_STARTED: described booting a prebuilt SD image and seeing
default clock/weather plugins — neither exists. Now documents the real
install (Pi OS Lite + one-shot installer / first_time_install.sh) and
that displays come from the Plugin Store. Duration and ordering
instructions moved to the Rotation tab where the controls actually
live.
- WEB_INTERFACE_GUIDE: three whole tabs were undocumented (Rotation,
Backup & Restore, Tools) and the Display tab's Vegas Scroll section
was unmentioned. Fonts overrides are per display element (not per
plugin); Logs has an Auto-scroll checkbox (not a Pause button); the
aspirational keyboard-shortcut list and no-JS claim removed.
- TROUBLESHOOTING: the hand-written service-file template (wrong user,
wrong ExecStart, dropped the autostart gate) replaced with the real
systemd/ units + install scripts; recovery steps no longer copy
placeholder units verbatim; WiFi curl endpoint corrected to /api/v3/;
cache-clearing advice now targets the real cache locations.
- ADVANCED_FEATURES: removed a false claim that CacheManager has no
delete(); fixed two example snippets that raise TypeError
(BackgroundDataService and get_config_file_mode signatures); fixed
cache paths, a 5-minute TTL that is actually 1 hour, and the vegas
table now links the complete 26-key reference.
- EMULATOR_SETUP_GUIDE: documented run.py flags that don't exist
(--plugin/--test-plugins) removed in favor of dev_server.py and
check_plugin.py; shipped emulator config values corrected (browser
adapter default on :8888, not pygame).
- PLUGIN_QUICK_REFERENCE: drag-and-drop reordering is shipped, not
'not yet supported'; discovery-fallback and registry-repo claims
corrected. PLUGIN_API_REFERENCE: get_vegas_segment_width returns
panels, not pixels. CONTRIBUTING: the repo uses flake8/mypy/bandit
pre-commit hooks, not black/ruff, and tests need requirements-test.txt.
- SKIN_SYSTEM/DEVELOPER_QUICK_REFERENCE: stale module paths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix: raise the two new dependency floors past their CVEs, silence a deliberate re-export
All three of these were introduced by this PR, which is what makes them
worth fixing here rather than deferring.
`urllib3` and `jinja2` were added to the requirements so that direct
imports stop relying on transitives — right call, but both floors were
set to the version that introduced the API rather than a version that is
safe to install. `urllib3>=1.26.0` sits below roughly ten CVEs including
a decompression-bomb safeguard bypass, and `jinja2>=3.1.0` below five
including two sandbox breakouts. Raised to 2.7.0 and 3.1.6, which is what
a working device already runs, so no install is disturbed. The comments
now say the floor is a security floor, since the next person to read
"imported directly" would otherwise reasonably lower it again.
This is the same reasoning the PR already applied to werkzeug; these two
just missed it.
The `DateTimeEncoder` import in cache_manager is unused on purpose — the
canonical class moved to src.cache.disk_cache and this re-export keeps
the documented import path working. flake8 cannot see intent, so it gets
an explicit `# noqa: F401` rather than being removed and quietly breaking
anything importing it from here. Verified the re-export still resolves to
the same object and still serialises datetimes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix: address CodeRabbit re-review — installer robustness and doc lint
- install_dependencies_apt.py: an import-only check let Debian
Bookworm's python3-freetype 2.3.0 satisfy the freetype-py>=2.5.1 pin.
check_package_installed() now verifies the installed freetype-py
version, and an apt install that lands below the minimum falls through
to pip instead of counting as success.
- first_time_install.sh: the .web_deps_installed marker was created even
when the smart installer failed, so re-runs skipped installation with
dependencies missing. The marker is now created only on success.
- CONTRIBUTING.md: document installing the pre-commit CLI before
'pre-commit install' (the requirements files don't provide it).
- Doc lint: fence language on the on-demand cache example (MD040),
blockquote continuation in GETTING_STARTED (MD028), and the
suppress_adapter_load_errors key removed from the emulator debug
example to match the options table.
Skipped one finding with reason (noted on the PR): the per-plugin
display_duration field in PLUGIN_QUICK_REFERENCE's example is not
obsolete — BasePlugin.get_display_duration() reads it and
PLUGIN_CONFIG_CORE_PROPERTIES.md documents it as a core property;
display.display_durations is a per-mode override, not a replacement.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(plugins): make update scheduling atomic so update() cannot run twice at once
Closes#401.
`run_scheduled_updates()` decided whether to update a plugin with a
check-then-act sequence: `can_execute()` and `set_state(RUNNING)` were
separate calls with nothing between them, so two scheduler threads could
both observe ENABLED and both go on to call the same plugin's `update()`.
`update_all_plugins()` had the identical pattern.
Two schedulers really do run at once. The render loop calls
`_tick_plugin_updates()`, and Vegas mode fires its own `vegas-plugin-tick`
daemon thread that is never joined when `VegasModeCoordinator.play()`
returns — a slow `update()` still in flight overlaps the next tick from
the main loop. A plugin running `update()` twice concurrently is unsafe
unless it happens to be reentrant; shared mutable state, a non-thread-safe
HTTP session or cache all break.
The async path was already covered by the `_pending_lock` dedup in
`_enqueue_update`, so the live exposure was the synchronous kill-switch
path and `update_all_plugins()`. Both now claim the plugin through
`_reserve_for_update()`, which holds one lock across the eligibility
check, the due-time check and the RUNNING transition — and nothing more.
Holding it across `execute_update()` would serialize slow plugins behind
each other and reintroduce the render stall the async worker exists to
avoid.
The due-time check moved inside the lock deliberately. Left outside, a
thread that had already decided "due" could claim the plugin the instant
the winner finished, running `update()` twice within one interval.
Two supporting changes fall out of it:
- `_enqueue_update()` no longer sets RUNNING (the reservation did), and
hands the reservation back if the pending-dedup ever fires. Otherwise a
reserved-but-unqueued plugin would sit in RUNNING with nothing left to
release it, and `can_execute()` would refuse it forever.
- `_finish()` now clears the pending entry *before* flipping the state
back to ENABLED. The old order left a window where a scheduler saw
ENABLED, reserved the plugin, then had its enqueue silently dropped by
the dedup — harmless as a missed tick before, a stuck plugin once a
reservation is involved.
Regression suite added and enrolled in CI, along with
test_async_plugin_updates.py which was not previously run there. The
overlap tests delay `can_execute()` to hold every thread inside the
check-then-act gap: the real window is a couple of bytecodes wide, so a
plain hammering test passes against the unfixed scheduler and proves
nothing. With that delay the suite reports `update() ran 8x concurrently`
on both affected paths before the fix, and passes after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* test: fix a race in the reservation suite's own wait loops
`test_async_path_never_overlaps` failed in CI with "update() ran 0x
concurrently" — the test's bug, not the scheduler's. It polled
`plugin._active` to wait for the update to finish, but before the worker
picks the item up nothing is active yet, so the loop fell straight
through and asserted on a plugin that had never run.
Both async waits now key on `update_calls >= 1` as well, so they wait for
an update to have started *and* finished. The stranded-state test gets
the same guard for a second reason: ENABLED is also the starting state,
so without it that assertion passes vacuously on a plugin that was never
scheduled.
Verified over 12 consecutive local runs, 12 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(plugins): roll back the claim when dispatch fails
_enqueue_update() reserved the plugin and added it to the pending set,
then started the worker and queued the item. Thread.start() raises
RuntimeError when the OS refuses a new thread — not hypothetical on a Pi
under memory or thread pressure — and nothing is queued at that point to
release the plugin. It stayed RUNNING with a stale pending entry, so
can_execute() refused it for the rest of the process, and the exception
escaped run_scheduled_updates() and skipped every remaining plugin in
that tick.
That is the same stranded-RUNNING failure the reservation was introduced
to prevent, just reached through the dispatch rather than the dedup, so
it is handled the same way: discard the pending entry, hand the
reservation back, log the cause. Swallowed rather than raised so one
plugin failing to queue cannot abort the others' turn.
Both new tests fail against the un-rolled-back version — the second on
the escaping RuntimeError itself — and pass with it. 74 tests across the
reservation, async-update, plugin-system, health, Vegas-adapter and
controller-toggle suites still pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WexvwNDtWLVymGVqKD7BGk
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Around 11:00 EDT on 2026-08-04 ESPN's site.api began rejecting the agents
this repo sends. Every scoreboard that goes through the shared data
sources returned `403 Client Error: Forbidden` — standings, game
summaries and scoreboards alike. A device that had been running fine
logged 287 ESPN errors in a day.
The filter is not the familiar one. Probing site.api across agents and
libraries, using requests as the plugins do:
bare 'LEDMatrix/1.0' 403 (with or without Accept)
browser string 403 (no header rescues it)
'LEDMatrix/1.0 (+https://github.com/...)' 200
requests / urllib / curl defaults 200
So it rejects browser-style strings outright and bare custom tokens, and
accepts honest client tokens or an agent that identifies the client and
links to it. The instinct to "just send a browser User-Agent" is now
exactly backwards — that is the one thing guaranteed to stay blocked.
Both call sites here sent a bare token: `LEDMatrix/1.0` in the sports
data sources and `LEDMatrix-Common/1.0` in the API helper. The Accept
header the data sources already sent does not save it. Both now send an
agent carrying the project URL, which also gives ESPN someone to contact
rather than an anonymous token to rate-limit.
Verified on a live device: ESPN errors went from a steady stream to zero
across a restart, with live MLB games fetching again.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found while sweeping devpi for issues. Nine of 27 installed plugins were
reported degraded in the web UI -- including baseball-scoreboard and
f1-scoreboard -- for using a documented core feature.
The core reads three tuning keys out of each plugin's own config block:
vegas_width_pct and vegas_overflow (vegas_mode/plugin_adapter.py) and
vegas_max_width_screens (base_plugin.py). No plugin declares them, and 37 of
the 42 published config schemas set "additionalProperties": false -- so schema
validation reported them as violations.
That is not just log noise. _validate_config_schema_soft sets `degraded` in
the health tracker, which the web UI surfaces, so a user who tuned a core
Vegas setting saw the plugin marked broken.
The keys are stripped before validation. Fixing it plugin-side would mean 42
schema edits and 42 version bumps -- 42 store updates for a contract the core
owns.
Listed explicitly rather than matched on a `vegas_` prefix: vegas_mode is the
opposite case, plugin-owned and declared in schemas, and a prefix rule would
silently stop validating it.
Verified on devpi: degraded went 9 of 27 -> 0 of 27, schema-mismatch warnings
9 -> 0, 22 plugins still load, no tracebacks. 800 core unit tests pass,
8 of them new -- including that a genuine violation is still reported, so the
check has not been turned into a no-op, and that the caller's live config dict
is never mutated.
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat(store): evaluate compatible_versions, not just the floor
Closes the gap CodeRabbit surfaced on #427. `compatible_versions` is the
canonical compatibility contract -- schema/manifest_schema.json marks it
required, all 42 published manifests carry it -- and it is the only field that
can express an *upper* bound. `ledmatrix_min_version` is a floor and cannot
say "not compatible with 4.x".
The gate read only the floor, so a plugin declaring ["2.0.0 - 2.9.9"] would be
installed on 3.2.0 regardless of having said it stops at 2.x.
check() now evaluates both and the more restrictive wins. The array is a set of
alternatives (satisfying any one entry suffices), supporting every form the
schema permits: >=, <=, >, <, ~, ^, a bare exact version, and an inclusive
"A - B" range, with prerelease/build suffixes tolerated.
Refusal still requires evidence. Anything unparseable, absent, or below
TRUSTWORTHY_FLOOR resolves to compatible.
That last point needed a new strict parser. parse_semver is deliberately
lenient -- it strips non-digits and yields (0, 0, 0) for a string with no
numbers at all. Harmless for a floor (0.0.0 never blocks) but wrong for a
range, where the same leniency turned an unreadable spec into a *refusal*: a
manifest whose only entry was garbage got compared against 0.0.0 and refused.
Range specs are now shape-checked first, so garbage reads as "no evidence".
parse_semver itself is unchanged, since the loader depends on its behaviour.
Verified: 815 core unit tests pass, 18 of them new. Swept the real registry --
all 42 published manifests, at cores 1.0.0 / 2.0.0 / 3.1.0 / 3.2.0 / 4.0.0 --
and nothing is refused at any of them. The gate stays inert for shipped
plugins, which is the property that makes it safe to land ahead of B5.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* feat(store): protect the one population the sunset would break
The B6 sunset deletes each plugin's guarded-import fallback, so a plugin that
floors at 3.2.0 must never reach a core that lacks the 3.2.0 modules. The gate
could not stop that for the population most at risk.
A device installed from the v3.1.0 release reports __version__ = "1.0.0". The
gate treated anything below TRUSTWORTHY_FLOOR as "unknown, do not block" --
correct while every manifest floors at 2.0.0, because blocking would have
emptied the plugin store for those users. But after the sunset it hands them a
3.2.0-floored plugin with no fallback, which fails to load with one log line.
Nothing else in the system protects them: they cannot be told apart from a
genuine 1.0.0 install.
On an untrustworthy core the gate now refuses a floor ABOVE 2.0.0 and still
allows anything at or below it. A floor above the ecosystem baseline says the
plugin needs modules that arrived after 2.0.0, and a core reporting below that
-- whether it is the v3.1.0 release or something genuinely ancient -- will not
have them. Refusing leaves the user on the version they already run instead of
one that cannot load.
Measured against all 42 published manifests:
today (every manifest floors at 2.0.0)
core 1.0.0 / 2.0.0 / 3.1.0 / 3.2.0 / unparseable -> 0 of 42 refused
after B6 (same manifests floored at 3.2.0)
core 1.0.0 -> 38 refused, core 3.1.0 -> 38 refused, core 3.2.0 -> 0
So nobody loses the store today, and the sunset cannot reach a core that
cannot run it.
Two older tests asserted the previous "allow everything" behaviour; they now
express the new rule with a 2.0.0 floor, which is what their no-lockout intent
was actually about.
The 38-of-42 in that measurement surfaced a separate B6 trap, recorded here
because it will bite whoever raises the floors: four plugins (flights,
leaderboard, music, stocks) declare the floor as a TOP-LEVEL
`min_ledmatrix_version`, a third spelling, which declared_min_version checks
before the versions[] array. For those, editing versions[0] is a silent no-op
and the floor stays at 2.0.0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(store): address review — malformed manifests, and suffixed versions
Both CodeRabbit findings on #433 verified against the code and fixed.
1. Malformed manifest sections raised instead of degrading. `requires` as a
list hit AttributeError ('list' object has no attribute 'get') and
`versions` as a mapping hit KeyError: 0. Both reproduced.
This got worse with the sunset rule in the previous commit: that branch
resolves the floor for *every* manifest on an untrustworthy core, where the
old code returned early. One hand-edited or third-party file with the wrong
shape would have taken down the whole install path rather than just itself.
Container types are now validated and an unrecognised shape reads as "no
declared floor".
2. Prerelease and build metadata leaked into the version numbers. The digit
scrape parsed "3.2.0+build42" as (3, 2, 42) and "3.2.0-rc1" as (3, 2, 1) --
a release candidate ranking above its own release. Both fed reject
decisions, and the consequence was demonstrable: a plugin pinned to exactly
"3.2.0" refused a core running 3.2.0+build42, which is that same version.
The suggested remedy -- use the strict token parser -- would not have
fixed it. _parse_strict validates the shape but delegates the numbers to
parse_semver, so it returned the same (3, 2, 42). The bug is in the scrape,
so suffixes are now dropped before it. Prereleases compare equal to their
release rather than below it; full prerelease ordering is more than any
caller needs and equal is far closer to right than what it did before.
parse_semver is shared with PluginLoader, so its suite was re-run: unchanged,
and it only ever gets more correct here.
Verified: 839 core unit tests pass, 21 of them new -- six malformed shapes,
five suffixed forms, and the two demonstrated regressions. The real-registry
sweep is unchanged at 0 of 42 refused across cores 1.0.0, 3.1.0, 3.2.0,
3.2.0+build42 and an unparseable string.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 13:46:05 -04:00
820 changed files with 116487 additions and 65758 deletions
# Cursor Helper Files for LEDMatrix Plugin Development
This directory contains Cursor-specific helper files to assist with plugin development in the LEDMatrix project.
## Files Overview
### `.cursorrules`
Comprehensive rules file that Cursor uses to understand plugin development patterns, best practices, and workflows. This file is automatically loaded by Cursor and helps guide AI-assisted development.
Based on audit results showing 186 issues across 20 plugins.
## Overview
Three priority fixes identified from audit:
1.**Priority 1 (HIGH)**: Remove core properties from required array - will fix ~150 issues
2.**Priority 2 (MEDIUM)**: Verify default merging logic - will fix remaining required field issues
3.**Priority 3 (LOW)**: Calendar plugin schema cleanup - will fix 3 extra field warnings
## Priority 1: Remove Core Properties from Required Array
### Problem
Core properties (`enabled`, `display_duration`, `live_priority`) are system-managed but listed in schema `required` arrays. SchemaManager injects them into properties but doesn't remove them from `required`, causing validation failures.
### Solution
**File**: `src/plugin_system/schema_manager.py`
**Location**: `validate_config_against_schema()` method, after line 295
### Implementation Steps
1.**Add code to remove core properties from required array**:
```python
# After injecting core properties (around line 295), add:
# Remove core properties from required array (they're system-managed)
if "required" in enhanced_schema:
core_prop_names = list(core_properties.keys())
enhanced_schema["required"] = [
field for field in enhanced_schema["required"]
if field not in core_prop_names
]
```
2. **Add logging for debugging** (optional but helpful):
```python
if "required" in enhanced_schema and core_prop_names:
removed_from_required = [
field for field in enhanced_schema.get("required", [])
if field in core_prop_names
]
if removed_from_required and plugin_id:
self.logger.debug(
f"Removed core properties from required array for {plugin_id}: {removed_from_required}"
)
```
3. **Test the fix**:
- Run audit script: `python scripts/audit_plugin_configs.py`
- Expected: Issue count drops from 186 to ~30-40
- All "enabled" related errors should be eliminated
### Expected Outcome
- All 20 plugins should no longer fail validation due to missing `enabled` field
Some plugins have required fields with defaults that should be applied before validation. Need to verify the default merging happens correctly and handles nested objects.
### Solution
**File**: `web_interface/blueprints/api_v3.py`
**Location**: `save_plugin_config()` method, around lines 3218-3221
### Implementation Steps
1. **Review current default merging logic**:
- Check that `merge_with_defaults()` is called before validation (line 3220)
- Verify it's called after preserving enabled state but before validation
This guide provides comprehensive instructions for creating, running, and loading plugins in the LEDMatrix project.
## Table of Contents
1. [Plugin System Overview](#plugin-system-overview)
2. [Creating a New Plugin](#creating-a-new-plugin)
3. [Running Plugins](#running-plugins)
4. [Loading Plugins](#loading-plugins)
5. [Plugin Development Workflow](#plugin-development-workflow)
6. [Testing Plugins](#testing-plugins)
7. [Troubleshooting](#troubleshooting)
---
## Plugin System Overview
The LEDMatrix project uses a plugin-based architecture where all display functionality (except core calendar) is implemented as plugins. Plugins are dynamically loaded from the `plugins/` directory and integrated into the display rotation.
- **Automatic Version Management**: Version bumping is handled automatically via the pre-push git hook - no manual version bumping is required for normal development workflows
- **GitHub as Source of Truth**: Plugin store always fetches latest versions from GitHub (releases/tags/manifest/commit)
- **Pre-Push Hook**: Automatically bumps patch version and creates git tags when pushing code changes
- The hook is self-contained (no external dependencies) and works on any dev machine
- Installation: Copy the hook from LEDMatrix repo to your plugin repo:
- Each plugin needs: `manifest.json`, `config_schema.json`, `manager.py`,`requirements.txt`
- Each plugin needs: `manifest.json`, `config_schema.json`, and the entry point (`manager.py` by default);`requirements.txt` if it has dependencies. Required manifest fields: `docs/PLUGIN_API_REFERENCE.md#manifest-required-fields`
- Display dimensions: always read dynamically from `self.display_manager.matrix.width/height`
- Display dimensions: always read dynamically from `self.display_manager.width/height` — not `display_manager.matrix.width/height`, because `matrix` is `None` when hardware init fails (the properties fall back to the canvas size)
- Secrets: namespaced by plugin id in `config/config_secrets.json`, declared
via `"x-secret": true` in the plugin's config schema, and deep-merged into
the plugin's config dict at load time — plugins read them with plain
`config.get(...)`, never a separate accessor
## Dev Workflow
- Link a plugin for development: `./scripts/dev/dev_plugin_setup.sh link-github <name>` clones the `ledmatrix-plugins` monorepo into `~/.ledmatrix-dev-plugins/` and links its `plugins/<name>` under the manifest id (add a repo URL for a plugin with its own repo; or `link <name> <path>`); symlinks land in `plugins/` — set `plugin_system.plugins_directory` to `plugins` so discovery picks them up. Fork/location overrides: `dev_plugins.json` (from `dev_plugins.json.example`)
- Browser preview without the display loop: `python3 scripts/dev_server.py` → http://localhost:5001
- Full display in emulator mode: `python3 run.py -e` (or `EMULATOR=true python3 run.py`)
- Validate one plugin headlessly: `python3 scripts/check_plugin.py --plugin <id>`
- Soak a rig for frame timing (on the Pi, service running): `python3 scripts/frame_soak.py --preview` — late-frame rate across every scroller; see `docs/SCROLL_PERFORMANCE.md`
## Plugin Store Architecture
- Official plugins live in the `ledmatrix-plugins` monorepo (not individual repos)
-`plugins.json` registry at `https://raw.githubusercontent.com/ChuckBuilds/ledmatrix-plugins/main/plugins.json`
- Store manager (`src/plugin_system/store_manager.py`) handles install/update/uninstall
- Monorepo plugins are installed via ZIP extraction (no `.git` directory)
- Store manager (`PluginStoreManager` in `src/plugin_system/store_manager.py`) handles install/update/uninstall
- Monorepo plugins are installed without a `.git` directory: GitHub Trees API + raw downloads, falling back to ZIP extraction
- Update detection for monorepo plugins uses version comparison (manifest version vs registry latest_version)
- Plugin configs stored in `config/config.json`, NOT in plugin directories — safe across reinstalls
- Third-party plugins can use their own repo URL with empty `plugin_path`
## Skin System (visual overlays for sports scoreboards)
- Skins live in `skins/<skin-id>/` (skin.json + skin.py), NOT in plugin dirs — plugin reinstall deletes plugin dirs
- Core: `src/skin_system/` (ScoreboardSkin, SkinContext, runtime); hook: `SportsCore._render_game()` in `src/base_classes/sports.py`
- Skins render onto `ctx.canvas` only; fallback to built-in renderer on `False`/exception (3 strikes disables for session)
- View-model guaranteed keys are frozen (see `test/test_skin_system.py::TestViewModelContract`) — renaming keys in `_extract_game_details_common` or sport extractors breaks published skins
- Skins are NOT monorepo plugins: no manifest bump / update_registry.py needed
## Common Pitfalls
- paho-mqtt 2.x needs `callback_api_version=mqtt.CallbackAPIVersion.VERSION1` for v1 compat
- paho-mqtt 2.x requires a `CallbackAPIVersion` argument: `VERSION1` for code written against v1 callback signatures (the MQTT bridge uses `VERSION2`)
- BasePlugin uses `get_logger()` from `src.logging_config`, not standard `logging.getLogger()`
-`DisplayManager` has no `draw_image()` — paste onto the PIL image directly:
`self.display_manager.image.paste(img, (x, y))` then `update_display()`
(use a mask for transparency: `image.paste(rgba, (x, y), rgba)`)
- When modifying a plugin in the monorepo, you MUST bump `version` in its `manifest.json` and run `python update_registry.py` — otherwise users won't receive the update
-`src/pi5_matrix_support.py` hardcodes what the pinned `rpi-rgb-led-matrix-master` can drive on a Raspberry Pi 5 (`Rp1PioConfigSupported()` in `lib/rp1/rp1_pio_backend.cc`). Re-check it whenever the submodule is bumped: a stale rule blocks Pi 5 settings the new library supports, and a missing one lets the display service crash-loop. `src/matrix_support.py` holds the same kind of rules for every board (rows, chain length, mapping names, parallel per mapping) and needs the same re-check
Designed novice-first, with power tools kept within reach.
- **Primary: hobbyist builders.** People who assembled an LED matrix panel on a Raspberry Pi, often by following the install video, and are frequently new to Linux and the Pi. They set the display up once (panel size, timezone, WiFi), install and enable a few plugins, then come back occasionally to tweak what the panel shows. They usually reach the control panel from a phone or laptop on their home network, sometimes as an installed home-screen app.
- **Secondary: tinkerers and plugin developers.** Comfortable with SSH, `config.json`, and GitHub. They lean on the Config Editor, Logs, Cache, Operation History, Tools, GitHub-repo installs, and per-plugin config while building or debugging. Their tools must stay reachable without sitting in the novice's path.
## Product Purpose
LEDMatrix turns a Raspberry Pi and an RGB LED matrix panel into an information-rich display (clock, weather, calendar, sports scores, stocks, music, and more) through a plugin platform. The web control panel ("LED Matrix Control") is where the display gets configured, extended, and kept healthy.
Success means a builder gets from a freshly flashed Pi to a working, personalized display without needing a terminal, and can keep it running (updates, recovery, troubleshooting) the same way.
## Positioning
Four strengths define LEDMatrix, and future work must protect all of them:
1.**Plugin ecosystem.** The core ships only `starlark-apps` and `web-ui-info`; everything else comes from the built-in Plugin Store (the official `ledmatrix-plugins` monorepo), third-party GitHub repos, or Starlark (Tidbyt-style) apps. Each installed plugin gets its own configuration tab, generated from its schema.
2.**Runs on tiny Pis.** The UI is served by the same device that drives the matrix, on boards as small as the Pi Zero 2 W (512 MB), Pi 3/3B+, and the 1 GB Pi 4.
3.**Recovers without SSH.** WiFi access-point fallback with a captive setup page, backup & restore, in-UI updates, live logs, diagnostics, service control, and plugin health let users fix problems from the browser.
4.**Open and community-led.** GPL-3.0, a Discord community, and contributions welcome. The maintainer (ChuckBuilds) builds in public and openly relies on AI development tools.
## Operating Context
- **Access.** Served on the local network at `http://<pi-ip>:5000` by the `ledmatrix-web` service. It is installable as a PWA (`web_interface/static/v3/manifest.json`, short name "LEDMatrix").
- **First run.** When the Pi has no network it creates its own WiFi access point, so the captive setup page (`templates/v3/captive_setup.html`) may be the very first screen a user sees, on a phone, with no internet connection.
- **Navigation.**
- System tabs: Overview, General, WiFi, Schedule, Display, Rotation, Config Editor, Backup & Restore, Fonts, Logs, Cache, Operation History, Tools.
- A second row holds Plugin Manager (with the Plugin Store), Starlark Apps, and one tab per installed plugin.
- **Live data.** The Overview shows system stats (CPU, memory, temperature, power/throttling) and a live display preview, streamed over SSE.
- **Getting Started checklist.** The Overview's first-run checklist runs: set panel size → set timezone → install a plugin → enable it → configure it.
- **Development.** `python3 scripts/dev_server.py` gives a browser preview without the display loop; `python3 run.py -e` runs the full display in emulator mode.
## Capabilities and Constraints
- **Hard constraint: plugin UI compatibility.** Third-party plugins rely on JSON Schema (Draft-7) generated config forms, the widget registry (`static/v3/js/widgets/`), `x-secret` fields, and plugin web-UI actions. UI changes must keep these working.
- **Config storage.** Plugin configuration lives in `config/config.json` and secrets in `config/config_secrets.json`, never in plugin directories, so configs survive reinstalls.
- **Stack.** An existing Flask + HTMX + Alpine.js app with Jinja templates (`web_interface/templates/v3/`) and static JS/CSS (`web_interface/static/v3/`), with self-hosted vendor assets.
- **Open decisions** (offered during init, not adopted as constraints):
- Whether the UI must work fully offline, with no CDN fallbacks at runtime.
- Whether a Node/CSS build step is acceptable for contributors.
- Whether a formal accessibility standard (e.g. WCAG 2.2 AA) is a requirement.
## Brand Commitments
- **Names.** The product is "LEDMatrix" and the web UI is titled "LED Matrix Control". The maintainer brand is ChuckBuilds.
- **Voice.** Friendly, honest, and learning-in-public, as in the README.
- **App icons.** They live in `web_interface/static/v3/icons/`.
No other visual identity has been made binding.
## Evidence on Hand
- **Photos.** Real photographs of running displays are linked in `README.md` (clock, weather, calendar, NHL/MLB/NFL/NCAA, stocks, music).
- **Video.** YouTube install and walkthrough videos from ChuckBuilds.
- **Docs.** Extensive documentation in `docs/`, e.g. `WEB_INTERFACE_GUIDE.md`, `GETTING_STARTED.md`, `WIFI_NETWORK_SETUP.md`, `LOW_MEMORY_BOARDS.md`, `PLUGIN_STORE_GUIDE.md`.
- **Absences.** There are no testimonials, user counts, or benchmark figures. Do not fabricate them.
## Product Principles
1.**Novice path first, power one click away.** Default views serve the first-time builder, while advanced tools stay discoverable for tinkerers.
2.**Never strand the user at a terminal.** Every setup, recovery, and troubleshooting task has a browser path, including from the AP-mode captive page.
3.**Respect the Pi.** Every feature is paid for in memory and CPU on a Pi Zero 2 W that is also driving the display.
4.**The ecosystem is the product.** Plugins, including third-party ones, must feel first-class and keep working across core UI changes.
5.**Honest and welcoming.** Plain language, truthful status, and no overstated claims, in keeping with an open, community-built project.
@@ -50,7 +50,15 @@ I'm trying to be open to constructive criticism and support, as long as it's a r
<details>
<summary>Core Features</summary>
The following plugins are available inside of the LEDMatrix project. These modular, rotating Displays that can be individually enabled or disabled per the user's needs with some configuration around display durations, teams, stocks, weather, timezones, and more. Displays include:
LEDMatrix is a plugin platform: the displays below are plugins installed
from the built-in Plugin Store (web interface → Plugins), where each can be
individually enabled, ordered, and configured — display durations, teams,
stocks, weather, timezones, and more. The core repo ships with just two
bundled plugins (`starlark-apps` and `web-ui-info`); the official plugins
live in the [ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins)
monorepo and install with one click, and third-party plugins can be
installed from their own GitHub repositories. Displays available in the
@@ -132,30 +140,31 @@ The system supports live, recent, and upcoming game information for multiple spo
| This project can be finnicky! RGB LED Matrix displays are not built the same or to a high-quality standard. We have seen many displays arrive dead or partially working in our discord. Please purchase from a reputable vendor. |
### Raspberry Pi
- Raspberry Pi Zero's don't have enough processing power for this project.
- **Raspberry Pi 3B, 4, or 5**
-**Raspberry Pi 3B, 4, or 5** (a Pi Zero 2 W also works, with the limits described under the 1GB/low-memory bullet below; the original Pi Zero / Zero W doesn't have enough processing power for this project)
[Amazon Affiliate Link – Raspberry Pi 4 4GB RAM](https://amzn.to/4dJixuX)
[Amazon Affiliate Link – Raspberry Pi 4 8GB RAM](https://amzn.to/4qbqY7F)
- **Pi 5 users**: the installer automatically detects Pi 5 and builds the `rpi-rgb-led-matrix` library with RP1 support. If you previously installed on a Pi 4 and migrated the SD card, or if you see `mmap` errors in the logs, force a fresh library build:
- Pi 5 config: leave `rp1_rio` at `0` (PIO mode, default) and set `gpio_slowdown` to `1` or `2`.
- **1GB models (Pi 3B / 3B+) and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`.
- Pi 5 config: leave `rp1_rio` at `0` (PIO mode, default) and start `gpio_slowdown` at `1`, raising it a step at a time if the image flickers or shows garbage (see `gpio_slowdown` under Display Settings).
- **1GB models (Pi 3B / 3B+), the 512MB Pi Zero 2 W and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`. Once running, keep an eye on memory: see [docs/LOW_MEMORY_BOARDS.md](docs/LOW_MEMORY_BOARDS.md).
### RGB Matrix Bonnet / HAT
- [Adafruit RGB Matrix Bonnet/HAT](https://www.adafruit.com/product/3211) – supports one “chain” of horizontally connected displays
- [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular-pi1` as hardware mapping)*
- [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular` as hardware mapping)*
- [Electrodragon RGB HAT](https://www.electrodragon.com/product/rgb-matrix-panel-drive-board-raspberry-pi/) – supports up to 3 vertical “chains”
- [Seengreat Matrix Adapter Board](https://amzn.to/3KsnT3j) – single-chain LED Matrix *(use `regular` as hardware mapping)*
### LED Matrix Panels
(2x in a horizontal chain is recommended)
- [Adafruit 64×32](https://www.adafruit.com/product/2278) – designed for 128×32 but works with dynamic scaling on many displays (pixel pitch is user preference)
**Warning: Lately the Waveshare Panels have had different variations - only some are compatible with this project. I hope to identify what is different to fix it but so far there is a decent chance you get a mis-matched set of panels if you don't buy them all at once! **
- [Waveshare 64×32](https://amzn.to/3Kw55jK) - Does not require E addressable pad
- [Waveshare 96×48](https://amzn.to/4bydNcv) – higher resolution, requires soldering the **E addressable pad** on the [Adafruit RGB Bonnet](https://www.adafruit.com/product/3211) to “8” **OR** toggling the DIP switch on the Adafruit Triple LED Matrix Bonnet *(no soldering required!)*
> Amazon Affiliate Link – ChuckBuilds receives a small commission on purchases
- [Waveshare 96×48](https://amzn.to/4bydNcv) – higher resolution, requires soldering the **E addressable pad** on the [Adafruit RGB Bonnet](https://www.adafruit.com/product/3211) to “8” **OR** toggling the DIP switch on the Adafruit Triple LED Matrix Bonnet *(no soldering required!)*
- There are some Panels on Aliexpress that have worked fine for me, shop around! I think Adafruit is probably the "safest" but they do have some limitation on resolution and layout.
> Amazon Affiliate Links – ChuckBuilds receives a small commission on purchases
### Power Supply
- [5V 4A DC Power Supply](https://www.adafruit.com/product/658) (good for 2 -3 displays, depending on brightness and pixel density, you'll need higher amperage for more)
@@ -163,7 +172,7 @@ The system supports live, recent, and upcoming game information for multiple spo
## Optional but recommended mod for Adafruit RGB Matrix Bonnet
- By soldering a jumper between pins 4 and 18, you can run a specialized command for polling the matrix display. This provides better brightness, less flicker, and better color.
- If you do the mod, we will use the default config with led-gpio-mapping=adafruit-hat-pwm, otherwise just adjust your mapping in config.json to adafruit-hat
- The default config uses `hardware_mapping` `adafruit-hat`. If you do the mod, change it to `adafruit-hat-pwm` (Display settings in the web interface, or `config.json`)
- More information available: https://github.com/hzeller/rpi-rgb-led-matrix/tree/master?tab=readme-ov-file
@@ -319,6 +328,7 @@ This one-shot installer will automatically:
- Install required system packages (git, python3, build tools, etc.)
- Clone or update the LEDMatrix repository
- Run the complete first-time installation script
- Print the web interface address, then **reboot the Pi automatically** (your SSH session will disconnect; give it a few minutes to come back)
The installation process typically takes 10-30 minutes depending on your internet connection and Pi model. Pi 3B/3B+ and other 1GB boards land at the top of that range, because the C++ library is compiled serially to stay within available memory. All errors are reported explicitly with actionable fixes.
@@ -337,10 +347,10 @@ If you prefer to install manually or the one-shot installer doesn't work for you
ssh ledpi@ledpi
```
2. Update repositories, upgrade Raspberry Pi OS, and install prerequisites:
2. Update repositories, upgrade Raspberry Pi OS, and install git (`first_time_install.sh` installs the build dependencies itself: `python3-pip`, `python-dev-is-python3`, `build-essential`, `cmake`, `ninja-build` and the rest):
This single script installs services, dependencies, configures permissions and sudoers, and validates the setup.
It finishes by asking whether to reboot. If you run it non-interactively — piped, over a script, or with `-y` — there is no one to ask, so **it reboots immediately without prompting**. Pass `--no-reboot-prompt` to install without rebooting:
For most settings I recommend using the web interface:
Edit the project via the web interface at http://[IP ADDRESS or HOSTNAME]:5000 or http://ledpi:5000 .
@@ -380,7 +400,7 @@ If you need to manually edit your config file, you can follow the steps below:
<summary>Manual Config.json editing </summary>
1. **First-time setup**:
The previous "First_time_install.sh" script should've already copied the template to create your config.json:
The previous `first_time_install.sh` script should've already copied the template to create your config.json:
2. **Edit your configuration**:
```bash
@@ -417,7 +437,7 @@ I recommend using the web-ui "Quick Actions" to control the Display.
## Plugins
<details>
LEDMatrix uses a plugin-based architecture where all display functionality (except the core calendar) is implemented as plugins. All managers that were previously built into the core system are now available as plugins through the Plugin Store.
LEDMatrix uses a plugin-based architecture where all display functionality is implemented as plugins. All managers that were previously built into the core system are now available as plugins through the Plugin Store.
### Plugin Store
See the [Plugin Store documentation](https://github.com/ChuckBuilds/ledmatrix-plugins) for detailed installation instructions.
@@ -439,19 +459,9 @@ You can also install plugins directly from GitHub repositories:
See the [Plugin Store documentation](https://github.com/ChuckBuilds/ledmatrix-plugins) for detailed installation instructions.
For plugin development, check out the [Hello World Plugin](https://github.com/ChuckBuilds/ledmatrix-hello-world) repository as a starter template.
For plugin development, the `plugins/hello-world/` plugin in the [ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins) repository is a starter template.
### Visual Skins for Scoreboards
Want a different look for a sports scoreboard without forking the plugin?
**Skins** restyle the live/recent/upcoming screens while the plugin keeps
handling data, scheduling, caching, and vegas mode. Install one with
`git clone <skin repo> skins/<skin-id>`, select it in the plugin's config,
and you're done — see [docs/SKIN_SYSTEM.md](docs/SKIN_SYSTEM.md) (how it
works) and [docs/CREATING_SKINS.md](docs/CREATING_SKINS.md) (build your own,
including a ready-made Claude Code prompt).
2. **Built-in Managers Deprecated**: The built-in managers (hockey, football, stocks, etc.) are now deprecated and have been moved to the plugin system. **You must install replacement plugins from the Plugin Store** in the web interface instead. The plugin system provides the same functionality with better maintainability and extensibility.
**Built-in Managers Deprecated**: The built-in managers (hockey, football, stocks, etc.) are now deprecated and have been moved to the plugin system. **You must install replacement plugins from the Plugin Store** in the web interface instead. The plugin system provides the same functionality with better maintainability and extensibility.
</details>
## Detailed Information
@@ -466,6 +476,10 @@ If you are copying my exact setup, you can likely leave the defaults alone. Howe
The display settings are located in `config/config.json` under the `"display"` key and are organized into three main sections: `hardware`, `runtime`, and `display_durations`.
The defaults below are the values in `config/config.template.json`. They are what applies when you haven't set a key: on every load, LEDMatrix adds any key your `config.json` lacks from the template, so `DisplayManager`'s own fallbacks are never reached on a normal install.
The web UI and the config API refuse values the rgbmatrix library can't start with. If one is written into `config.json` by hand anyway, the display logs which setting it is (`Failed to initialize RGB Matrix` in `sudo journalctl -u ledmatrix`), runs in fallback mode, and the Display tab shows the message.
### Hardware Configuration (`display.hardware`)
These settings control the physical hardware configuration and how the matrix is driven.
@@ -475,15 +489,18 @@ These settings control the physical hardware configuration and how the matrix is
- **`rows`** (integer, default: 32)
- Number of LED rows (vertical pixels) in each panel
- Common values: 16, 32, 48, 64
- An even number from 8 to 64, the most the rgbmatrix library drives per panel
- Must match your physical panel configuration
- **`cols`** (integer, default: 64)
- Number of LED columns (horizontal pixels) in each panel
- Common values: 32, 64, 96, 128
- At least 16, with no upper limit
- Must match your physical panel configuration
- **`chain_length`** (integer, default: 2)
- Number of LED panels chained together horizontally
- 1 to 255 (the library's Python binding stores it in one byte); longer chains lower the refresh rate
- If you have 2 panels side-by-side, set to 2
- If you have 4 panels in a row, set to 4
- Total display width = `cols × chain_length`
@@ -492,68 +509,70 @@ These settings control the physical hardware configuration and how the matrix is
- Number of parallel chains (panels stacked vertically)
- Use 1 for a single row of panels
- Use 2 if you have panels stacked in two rows
- 1–3, and no more than your `hardware_mapping` has outputs: `regular` and `classic` have 3 (e.g. the Adafruit Triple LED Matrix Bonnet); `adafruit-hat`, `adafruit-hat-pwm`, `regular-pi1` and `classic-pi1` have 1. The library stops the display service outright on a mismatch, so it is refused
- Total display height = `rows × parallel`
#### Brightness and Visual Settings
- **`brightness`** (integer, 0-100, default: 90)
- **`brightness`** (integer, 1-100, default: 90)
- Display brightness level
- Lower values (0-50) are dimmer, higher values (50-100) are brighter
- Lower values (1-50) are dimmer, higher values (50-100) are brighter
- Recommended: 70-90 for indoor use, 90-100 for bright environments
- Very high brightness may cause distortion or require more power
- Specifies which GPIO pin mapping to use for your hardware
- **`"adafruit-hat-pwm"`**: Use this for Adafruit RGB Matrix Bonnet/HAT WITH the jumper mod (PWM enabled). This is the recommended setting for Adafruit hardware with the PWM jumper soldered.
- **`"adafruit-hat"`**: Use this for Adafruit RGB Matrix Bonnet/HAT WITHOUT the jumper mod (no PWM). Remove `-pwm` from the value if you did not solder the jumper.
- **`"regular"`**: Standard GPIO pin mapping for direct GPIO connections (Generic)
- **`"regular"`**: Standard GPIO pin mapping for direct GPIO connections (Generic). Also the right choice for the Adafruit Triple LED Matrix Bonnet
- **`"regular-pi1"`**: Standard GPIO pin mapping for Raspberry Pi 1 (older hardware or non-standard hat mapping)
- **`"classic"`** / **`"classic-pi1"`**: the library's original pin-outs, for old adapter boards wired to them. Not used by current HATs
- Any other name is refused. `compute-module` is only compiled in when the library is built with `ENABLE_WIDE_GPIO_COMPUTE_MODULE`, which the installer doesn't do. On a Raspberry Pi 5, `classic-pi1` isn't supported
- Choose the option that matches your specific hardware setup, if aren't sure try them all.
- Hardware pulsing (see `disable_hardware_pulsing`) needs the panel's OE line on GPIO 18, which `adafruit-hat-pwm` and `regular` provide and `adafruit-hat` does not
#### PWM (Pulse Width Modulation) Settings
These settings affect color fidelity and smoothness of color transitions:
- **`pwm_bits`** (integer, default: 9)
- Number of bits used for PWM (affects color depth)
- Higher values (9-11) = more color levels, smoother gradients
- Lower values (7-8) = fewer color levels, but may improve stability on some hardware
- Range: 1-11, recommended: 9-10
- **`pwm_bits`** (integer, 1-11, default: 9)
- Color depth per channel: how many brightness levels each LED gets
- Higher values (9-11) = more color levels, smoother gradients, lower refresh rate
- Lower values (7-8) = the subtlest shades are dropped for a higher refresh rate; `1` gives 8 colors
- Recommended: 9-10
- **`pwm_dither_bits`** (integer, default: 1)
- Additional dithering bits for smoother color transitions
- Helps reduce color banding in gradients
- Higher values (1-2) = smoother gradients but may impact performance
- Disables hardware pulsing (usually leave as false)
- Set to `true` only if you experience timing issues
- Most users should leave this as `false`
- `false` = the Pi's hardware PWM times each brightness pulse; `true` = software timing
- Leave `false` where possible. Software timing is less exact, so a row, or the whole panel, can briefly flash brighter
- Hardware pulsing needs the panel's OE line on GPIO 18 (`adafruit-hat-pwm`, `regular`, the Adafruit Triple LED Matrix Bonnet). With `adafruit-hat` the library uses software timing anyway
- It also needs the Pi's onboard sound driver (`snd_bcm2835`) disabled, which `first_time_install.sh` does. Set `true` only if you need the Pi's own audio
- **`inverse_colors`** (boolean, default: false)
- Inverts all colors (red becomes cyan, etc.)
@@ -561,9 +580,9 @@ These settings affect color fidelity and smoothness of color transitions:
- If the image is scrambled in a repeating pattern, try the value named after your panel first
- **`panel_type`** (string, default: `""`)
- Sends a start-up initialization sequence to driver chips that need one
- `""` = Standard (no initialization) — right for most panels, including FM6124 / FM6124D / FM6124DJ
- `"FM6126A"` or `"FM6127"` for panels with those chips; try `"FM6126A"` if the panel stays dark or lights only the first pixel on Standard
### Runtime Configuration (`display.runtime`)
These settings control runtime behavior and GPIO timing:
- **`gpio_slowdown`** (integer, default: 3)
- GPIO timing slowdown factor
- **Critical setting**: Must match your Raspberry Pi model for stability
- **Raspberry Pi 3**: Use 3
- **Raspberry Pi 4**: Use 4
- **Raspberry Pi 5**: Use 1–2 in PIO mode (`rp1_rio: 0`, the default); start with `1` and increase if you see flickering
- **Raspberry Pi Zero/1**: Use 1-2
- Incorrect values can cause display corruption, flickering, or system instability
- GPIO timing slowdown factor (0-10): slows GPIO writes so the panel electronics keep up. Higher is more reliable but lowers the refresh rate
- **Critical setting**: depends on your Raspberry Pi model and your panel
- **Raspberry Pi Zero/1**: 0-1
- **Raspberry Pi 2/3**: 1-3
- **Raspberry Pi 4**: 2-4 (the config template ships 3)
- **Raspberry Pi 5**: 1–3 in PIO mode (`rp1_rio: 0`, the default). Start at `1` (the library treats `0` as `1` there) and raise it a step at a time if the image flickers or shows garbage — chained panels are the likeliest to need it
- Panels on `row_address_type` 5 (SM5368 row drivers) can need 6-8 on a Pi 4
- Too low: garbage, flicker or rows jumping. Too high: a lower refresh rate
- If you experience issues, try adjusting this value up or down by 1
- **`rp1_rio`** (integer, 0 or 1, default: 0) — Raspberry Pi 5 only
- Which driver the Pi 5's RP1 chip uses: `0` = PIO (default, less CPU), `1` = RIO (registered I/O, can reach a higher refresh rate)
- In RIO mode the effect of `gpio_slowdown` is inverted: higher values may be faster
- Ignored on a Pi 0-4, and applied only if the installed rgbmatrix library supports it
@@ -645,7 +702,7 @@ Controls how long each display module stays visible in seconds before switching
- Some plugins can automatically adjust their display time based on content
- This setting limits how long they can extend (prevents one display from dominating)
- Example: If set to 60, a plugin can extend up to 60 seconds even if it requests longer
- Leave unset to use the default cap (typically 90 seconds)
- Leave unset to use the default cap (180 seconds; the web UI accepts 30-1800)
### Example Configuration
@@ -692,6 +749,14 @@ Controls how long each display module stays visible in seconds before switching
- Verify `hardware_mapping` matches your HAT/connection type
- Try adjusting `gpio_slowdown`
- Ensure your display doesn't need the E-Addressable line
- If it went blank right after a settings change, the Display tab shows a "simulation mode" banner, and `sudo journalctl -u ledmatrix` shows `Failed to initialize RGB Matrix` followed by the reason. When LEDMatrix refused the settings (for example more than 64 `rows`, `parallel` 2 on an `adafruit-hat` mapping, a misspelled `hardware_mapping`, or on a Raspberry Pi 5 a `row_address_type` other than 0 or 2), the message names each one: change them, save, and restart the display service. Otherwise the library itself failed, and its own message just before names the problem
- A repeating scramble points at `row_address_type` or `multiplexing`; a panel that stays dark, at `panel_type`
**Rows jump up and down, or the bottom row repeats other rows:**
- Raise `gpio_slowdown` a step at a time (SM5368 panels on `row_address_type` 5 can need 6-8 on a Pi 4)
**A row or the whole panel briefly flashes brighter:**
- Set `disable_hardware_pulsing` to `false` (needs the OE line on GPIO 18; see `hardware_mapping`)
**Colors are wrong or inverted:**
- Check `led_rgb_sequence` (try "GRB" if "RGB" doesn't work)
@@ -716,15 +781,21 @@ Controls how long each display module stays visible in seconds before switching
| `broadcast_logos/` | `news` and `odds-ticker` plugins |
| `static_images/` | Legacy examples referenced in the `static-image` plugin's docs; the plugin itself stores uploads under `assets/plugins/<plugin-id>/uploads/` |
| `plugins/` | Per-plugin uploaded files (`assets/plugins/<plugin-id>/uploads/`), served by the web interface |
Plugins resolve these paths relative to the LEDMatrix install directory, so
the directories are part of the de-facto plugin API even where no file in
this repo references them. New plugins should bundle their own assets or
use the per-plugin upload directory instead of adding top-level
| `priority` | `1` | Stored on each request (higher number = higher priority, per `FetchRequest`), but the service runs requests in submission order; it does not reorder by priority |
### Performance Impact
@@ -802,9 +924,9 @@ Enable background service per plugin in `config/config.json`:
The background data service is used by all of the sports scoreboard
Ownership, modes, sudo rules and the repair scripts are listed in
[PERMISSIONS.md](PERMISSIONS.md). This section covers the helpers code uses
to keep files shareable.
### Overview
LEDMatrix uses a dual-user architecture: the display service runs as root (hardware access), while the web interface runs as a non-privileged user. Centralized permission management ensures both can access necessary files.
@@ -875,6 +1004,7 @@ from src.common.permission_utils import (
ensure_file_permissions,
get_config_file_mode,
get_assets_file_mode,
get_assets_dir_mode,
get_plugin_file_mode,
get_cache_dir_mode
)
@@ -883,7 +1013,10 @@ from src.common.permission_utils import (
| `ledmatrix-web.service` | the installing user | [`start_web_conditionally.py`](../scripts/utils/start_web_conditionally.py) → [`web_interface/start.py`](../web_interface/start.py) (Flask, port 5000) | `install_service.sh`, [`install_web_service.sh`](../scripts/install/install_web_service.sh) |
| `ledmatrix-update-verify.path` / `.service` | the web user | Health check after an automatic update | the same installers, or [`src/auto_update_setup.py`](../src/auto_update_setup.py) at runtime |
| Preview viewer marker | `/tmp/led_matrix_preview_viewer` | web, while a preview is open | display: writes full-rate snapshots only while it is fresh |
| Timeouts | [`plugin_executor.py`](../src/plugin_system/plugin_executor.py) (`PluginExecutor`, 30 s default; a timed-out thread is abandoned, not killed) |
| Circuit breaker | [`plugin_health.py`](../src/plugin_system/plugin_health.py) (`PluginHealthTracker`: 3 consecutive failures open the circuit for 300 s) |
| Config schemas and defaults | [`schema_manager.py`](../src/plugin_system/schema_manager.py) |
| Install, update, uninstall | [`store_manager.py`](../src/plugin_system/store_manager.py) (`PluginStoreManager`), with its methods split across [`store_registry.py`](../src/plugin_system/store_registry.py) (registry, GitHub), [`store_install.py`](../src/plugin_system/store_install.py) and [`store_update.py`](../src/plugin_system/store_update.py) |
`data/auto_update_verify.request`. That file triggers
`ledmatrix-update-verify.path`, which runs the verifier as a separate unit
(so restarting the web service does not kill it). The verifier restarts
both services, waits for the web API to answer and the display service to
stay up, and on failure resets to the previous commit and restarts again.
Plugin updates run only after a verified core update. State is in
`data/auto_update_state.json` and `data/auto_update_pending.json`.
- **Startup validator.** `StartupValidator`
([`src/startup_validator.py`](../src/startup_validator.py)) runs twice in
`DisplayController.__init__`: config and cache directory first, then
enabled plugins once the plugin manager exists. It also warns when an
installed systemd unit differs from its template in `systemd/`. Results
are logged; startup continues either way.
## Where to start reading
| Task | Start with |
|---|---|
| Change rotation, durations or priorities | `DisplayController.run()` and `_get_display_duration()` in [`display_controller.py`](../src/display_controller.py) |
| Add a config key | [CONFIG_REFERENCE.md](CONFIG_REFERENCE.md), [`config/config.template.json`](../config/config.template.json), the tab's partial and `api_v3/config.py` |
| Change drawing or fonts | [`display_manager.py`](../src/display_manager.py), [`font_manager.py`](../src/font_manager.py), [`src/common/bdf_font.py`](../src/common/bdf_font.py) |
| Add a plugin-facing API | [`base_plugin.py`](../src/plugin_system/base_plugin.py) or [`src/common/`](../src/common/README.md); document it in [PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md) |
| Plugin install/update bugs | `PluginStoreManager` in [`store_manager.py`](../src/plugin_system/store_manager.py) |
| A plugin that won't load | `PluginManager.load_plugin()` and `PluginLoader.load_plugin()`; `python3 scripts/check_plugin.py --plugin <id>` |
| Add an API endpoint | the matching module in [`api_v3/`](../web_interface/blueprints/api_v3/) |
| Add a web UI tab or control | `templates/v3/base.html`, the tab's partial, `pages_v3.py` |
| Installer or permissions | [`first_time_install.sh`](../first_time_install.sh), [`scripts/install/`](../scripts/install/), [PERMISSIONS.md](PERMISSIONS.md) |
| Work without a Pi | [DEV_PREVIEW.md](DEV_PREVIEW.md), [EMULATOR_SETUP_GUIDE.md](EMULATOR_SETUP_GUIDE.md), [HOW_TO_RUN_TESTS.md](HOW_TO_RUN_TESTS.md) |
Every key in `config/config.json`, what it does, its default, and where the
code reads it. The file is created from `config/config.template.json` on
first run, and `ConfigManager._migrate_config()` merges any template keys
added by later releases into your existing config (your values are never
overwritten). Secrets live in `config/config_secrets.json` and are merged
into the config at load time.
Most settings are editable from the web interface; this page documents the
underlying keys for people editing `config.json` directly or writing
tooling against it.
## Top level
| Key | Type / default | Meaning | Read by |
|---|---|---|---|
| `web_display_autostart` | bool, `true` | Whether the web interface service starts with the system | `scripts/utils/start_web_conditionally.py` |
| `auto_update.enabled` | bool, `false` | Weekly automatic updates: LEDMatrix code first (health-checked, rolled back on failure), then installed plugins. Toggle in the General tab or install with `first_time_install.sh --enable-auto-update` | `web_interface/auto_update.py`, `src/auto_update_setup.py` (`is_enabled()`) |
| `timezone` | string, `"America/New_York"` | IANA timezone for schedules and displays | `ConfigManager.get_timezone()` |
| `target_fps` | int, `100` | Legacy "Scroll Frame Rate". Core scrolling no longer reads it: scroll frames are presented at `display.hardware.limit_refresh_rate_hz` divided by each scroll's frame hold, and speed comes from each plugin's scroll settings. Still exposed to plugins via `BasePlugin.global_config` | `src/plugin_system/base_plugin.py` |
| `location` | object | `city` / `state` / `country`. Supplies the **default** for a plugin's own `location_city` / `location_state` / `location_country` setting, so weather, radar and friends follow this device without being configured twice. A value saved on the plugin itself still overrides it. Starlark (Tidbyt) apps get the same treatment: a `Location` field left blank on the app renders at this city (geocoded once via Open-Meteo, coordinates cached permanently) instead of the app author's default, which is usually San Francisco. If the city can't be looked up (no match, or the geocoder is unreachable; retried after 30 minutes), the app keeps its own default. | `SchemaManager.apply_device_location()`, then plugins via merged config; `src/device_location.py` for Starlark apps |
Read by `DisplayController._check_schedule()` (`src/display_controller.py`).
Managed in the web UI under Schedule.
## `dim_schedule` — scheduled brightness dimming
Same shape as `schedule` (the template sets its `mode` to `"global"`), plus:
| Key | Type / default | Meaning |
|---|---|---|
| `dim_brightness` | int, `30` | Brightness percentage applied while the dim window is active |
Read by `DisplayController._check_dim_schedule()` (`src/display_controller.py`;
saved via `POST /api/v3/config/dim-schedule`). The display returns to
`display.hardware.brightness` outside the window.
## `display.hardware` — matrix panel hardware
All keys map to the corresponding `rpi-rgb-led-matrix` options and are read
in `DisplayManager._setup_matrix` (`src/display_manager.py`). Defaults are the
`config/config.template.json` values: `ConfigManager` adds any key missing from
`config.json` from the template on load, so `DisplayManager`'s own fallbacks
don't apply on a normal install.
The ranges are what the pinned rgbmatrix library and its Python binding accept
(`src/matrix_support.py`). The config API refuses anything else; a value
hand-edited into `config.json` makes the display log the setting and run in
fallback mode instead of starting the matrix.
| Key | Type / default |
|---|---|
| `rows` / `cols` | int, `32` / `64` — rows: even, 8–64; cols: at least 16 |
| `chain_length` | int, `2` — 1–255 (the Python binding stores it in one byte) |
| `parallel` | int, `1` — 1–3, and no more than `hardware_mapping` has outputs (`regular`, `classic`: 3; the others: 1) |
| `brightness` | int, `90` — 1–100 |
| `hardware_mapping` | string, `"adafruit-hat"` — `"adafruit-hat-pwm"`, `"adafruit-hat"`, `"regular"`, `"regular-pi1"`, `"classic"` or `"classic-pi1"` (case-insensitive; `compute-module` isn't in the installed build). A Pi 5 doesn't support `"classic-pi1"` |
| `disable_hardware_pulsing` | bool, `false` — `true` times brightness pulses in software (less exact); hardware pulsing needs the OE line on GPIO 18 and the Pi's onboard sound driver off |
| `inverse_colors` | bool, `false` |
| `show_refresh_rate` | bool, `false` — prints the refresh rate to stdout; draws nothing on the panel |
| `limit_refresh_rate_hz` | int, `100` — `0` = no cap; scroll timing assumes 100 Hz when `0` |
| `pixel_mapper_config` | string, `""` — e.g. `"U-mapper"` / `"Rotate:90"`; mappers that rotate or fold the chain change the display size plugins and the web preview see |
| `orientation` | string, `"normal"` — `"180"` rotates the rendered image 180° for panels physically mounted upside down (e.g. to move the Pi/wiring to a more convenient side); `"90"` / `"270"` for a panel on its side, swapping width and height; composed onto `pixel_mapper_config` as a trailing `Rotate:<degrees>` mapper, so it stays independent of any custom `pixel_mapper_config` value |
| `row_address_type` | int, `0` — non-standard panel row addressing: `1` AB, `2` direct row select, `3` ABC, `4` ABC shift + DE direct, `5` SM5368 / B707 row shift register (e.g. Waveshare 96x48 V2, with `led_rgb_sequence``"BGR"`). On a Pi 5 the library supports only `0` and `2`, and LEDMatrix enforces that (`src/pi5_matrix_support.py`) |
| `multiplexing` | int, `0` — 0–22, pixel wiring scheme for outdoor/specialty panels (names listed in the README) |
| `panel_type` | string, `""` — set to `"FM6126A"` or `"FM6127"` for panels needing init; FM6124 / FM6124D / FM6124DJ panels need none, so leave it `""` |
## `display.runtime`
| Key | Type / default | Meaning |
|---|---|---|
| `gpio_slowdown` | int, `3` | GPIO timing slowdown for faster Pis (0–10). On a Pi 5 in PIO mode start at `1` (`0` acts as `1`) and raise it if the image flickers or shows garbage. Panels on `row_address_type``5` (SM5368 row drivers) can need 6–8 on a Pi 4 — lower values make rows jump |
| `rp1_rio` | int, `0` | Pi 5 only: `0` = PIO (less CPU), `1` = RIO (higher refresh; `gpio_slowdown` effect inverted). Applied only if the installed matrix library supports it |
## `display.double_sided`
Drives `_LogicalMatrix` in `src/display_manager.py` — renders the same
logical image to multiple chained physical panels.
| `copies` | int, `2` | Number of physical copies in the chain |
| `axis` | `"horizontal"`, default | Axis along which panels are chained |
## `display` — other keys
| Key | Type / default | Meaning | Read by |
|---|---|---|---|
| `display_durations` | object, `{}` | Per-plugin display duration in seconds, keyed by plugin id (e.g. `"clock": 15`) | `DisplayController._get_display_duration()` (`src/display_controller.py`) |
| `plugin_rotation_order` | array, `[]` | Explicit rotation order of plugin ids; empty = all enabled plugins in discovery order | `DisplayController._apply_plugin_rotation_order()` (`src/display_controller.py`) |
| `use_short_date_format` | bool, `true` | Compact date rendering in sports scoreboards | Nothing since `src/base_classes` was removed; scoreboards read `display.use_short_date_format` from their own plugin config |
| `scan_order_compensation` | string, `"auto"` | `"auto"` shows one half of each panel a refresh behind while something scrolls at one frame per refresh, which removes the 1px step a 1:N-scan panel shows across its middle; `"off"` disables it. Applies only to layouts whose row order is known: plain or parallel chains, 0 or 180 degree orientation, `multiplexing` 0, `scan_mode` 0, and not in the emulator | `DisplayManager._setup_scan_order_compensation()` (`src/display_manager.py`, `src/scan_order.py`) |
| `dynamic_duration.max_duration_seconds` | int, optional | Cap for plugins that request dynamic display time | `DisplayController._get_global_dynamic_cap()` (`src/display_controller.py`) |
Read by `src/vegas_mode/config.py` (`VegasScrollConfig.from_config`). See
[ADVANCED_FEATURES.md](ADVANCED_FEATURES.md) for behavior details, including
[live content in the ticker](ADVANCED_FEATURES.md#live-content-in-the-ticker).
| Key | Type / default |
|---|---|
| `enabled` | bool, `false` |
| `scroll_speed` | int, `50` (px/s) |
| `separator_width` | int, `32` |
| `plugin_order` | array, `[]` |
| `excluded_plugins` | array, `[]` |
| `target_fps` | int, `125` |
| `buffer_ahead` | int, `2` |
| `intra_plugin_gap` | int, `8` |
| `render_width_pct` | int, `100` |
| `min_content_separation` | int, `24` |
| `min_cut_gap` | int, `6` |
| `continuous_scroll` | bool, `true` |
| `offscreen_prefetch` | bool, `true` — render every plugin's ticker content on the background thread, each on its own canvas. `false` restores handing canvas-bound plugins to the render thread, one pause at a time. Temporary; see [OFFSCREEN_RENDERING.md](OFFSCREEN_RENDERING.md) |
| `prefetch_gate` | bool, `true` — let that background thread run Python only while the render thread is waiting for the panel, so the render thread never waits for the GIL when a refresh comes round. Only takes effect with the rebuilt rgbmatrix binding (`scripts/build_rgbmatrix_nogil.sh`). See [OFFSCREEN_RENDERING.md](OFFSCREEN_RENDERING.md) |
| `switch_interval_ms` | float, `0` — experimental: shorten Python's GIL switch interval to this many ms while Vegas runs. `0` leaves the default (5 ms) alone |
| `smooth_scroll` | bool, `true` — move a whole number of pixels per panel refresh, locked to vsync. `scroll_speed` is snapped to the nearest speed the panel can show that way (at 95Hz: 95, 47.5, 31.7 px/s…), measured against the panel's real refresh rate once scrolling starts |
| `sub_pixel_blend` | bool, `false` — the older smoothing: advance by elapsed time and blend neighbouring pixel columns. Looks anti-aliased in the web preview but shimmers on the panel and is not locked to the refresh. Overrides `smooth_scroll` when on |
| `extend_threshold_screens` | float, `2.0` |
| `auto_trim` | bool, `true` |
| `trim_threshold` | int, `10` |
| `content_padding` | int, `8` |
| `min_plugin_width` | int, `8` |
| `lead_in_width` | int, `0` |
| `plugins_per_cycle` | int, `6` |
| `max_plugin_width_ratio` | float, `0.0` |
| `overflow_mode` | string, `"rotate"` |
| `dynamic_duration_enabled` | bool, `true` |
| `min_cycle_duration` | int, `60` |
| `max_cycle_duration` | int, `240` |
| `frame_based_scrolling` | bool, `true` — does not step or set a frame rate; motion is by elapsed time either way. When `true`, `scroll_speed` passes through a clamp of 0.1–5 px per `scroll_delay` (see next row) |
| `scroll_delay` | float, `0.02` — not a frame period. Only used with `frame_based_scrolling`: the applied speed is `clamp(scroll_speed × scroll_delay, 0.1, 5) / scroll_delay` px/s, so at `0.02` speeds under 5 px/s run at 5, and at `0.001` nothing runs slower than 100 px/s |
| `live_in_ticker` | bool, `false` — keep scrolling during live games instead of handing the display to a full-screen scoreboard |
| `live_weight` | int, `3` (1–10) — slots per cycle for a plugin with live content |
| `favorite_live_weight` | int, `5` (1–10) — slots per cycle when a plugin reports a favorite team is live |
## `sync` — multi-display synchronization
Read by `src/common/sync_manager.py` and `src/display_controller.py`.
| Key | Type / default | Meaning |
|---|---|---|
| `role` | `"standalone"` (default), `"leader"`, or `"follower"` | This device's role in a synced pair |
| `port` | int, `5765` | TCP port used for sync traffic |
| `follower_position` | `"left"` (default) or `"right"` | Which half of the combined image this follower renders (`src/display_controller.py`) |
## `plugin_system`
| Key | Type / default | Meaning |
|---|---|---|
| `plugins_directory` | string, `"plugin-repos"` | Where the Plugin Store installs plugins and the only directory the plugin loader scans. Read by `PluginManager` and `PluginStoreManager` (`src/plugin_system/`); editable under General settings |
| `auto_discover`, `auto_load_enabled`, `development_mode` | bool | **Unused.** Legacy keys, read by nothing and no longer in the template; older configs may still carry them. Plugins are always discovered, and every plugin with `enabled: true` is loaded — to keep a plugin installed but dormant, set its own `enabled` to `false`. Not shown in the web UI; may be left in or removed from config.json |
## Plugin config blocks
Every installed plugin stores its settings under a top-level key equal to
its plugin id (the template ships one for the bundled `web-ui-info`
plugin). The shape of each block is defined by that plugin's
`config_schema.json`; common keys are `enabled` and `display_duration`.
See [PLUGIN_CONFIG_CORE_PROPERTIES.md](PLUGIN_CONFIG_CORE_PROPERTIES.md).
## `config/config_secrets.json`
| Key | Meaning |
|---|---|
| `github.api_token` | Optional GitHub token the Plugin Store uses to avoid API rate limits (`src/plugin_system/store_registry.py`) |
| `<plugin-id>.*` | Secrets a plugin declares with `"x-secret": true` in its config schema; merged into that plugin's config at load time |
This document explains how the LEDMatrix project uses a multi-root workspace to manage plugins as separate Git repositories.
This document explains how to work on LEDMatrix and the official plugins side
by side, with one editor workspace and the plugins loaded straight from your
plugin checkout.
## Overview
The LEDMatrix project has been migrated from a git submodule implementation to a **multi-root workspace** implementation for managing plugins. This allows:
Official plugins live in a single repository,
[ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins), with one
directory per plugin under `plugins/`. There are no separate per-plugin
repositories. For development you clone that monorepo **next to** LEDMatrix
and symlink the plugin directories you are working on into LEDMatrix's
`plugins/` directory with `scripts/dev/dev_plugin_setup.sh`.
- ✅ Plugins to exist as independent Git repositories
- ✅ Updates to plugins without modifying the LEDMatrix project
- ✅ Easy development workflow with all repos in one workspace
- ✅ Plugin system discovers plugins via symlinks in `plugin-repos/`
- ✅ Plugin code stays in the monorepo checkout, with its own git history
- ✅ LEDMatrix discovers the plugins through symlinks in `plugins/`
(git-ignored), so the production `plugin-repos/` directory is untouched
- ✅ `LEDMatrix.code-workspace` opens both repositories in VS Code/Cursor
## Directory Structure
```text
/home/chuck/Github/
├── LEDMatrix/ # Main project
│ ├── plugin-repos/ # Symlinks to actual repos (managed automatically)
All plugin repositories are cloned to `/home/chuck/Github/` (parent directory of LEDMatrix) as regular Git repositories:
- `ledmatrix-clock-simple/`
- `ledmatrix-weather/`
- `ledmatrix-football-scoreboard/`
- etc.
### 2. Symlinks in plugin-repos/
The `LEDMatrix/plugin-repos/` directory contains symlinks pointing to the actual repositories in the parent directory. This allows the plugin system to discover plugins without modifying the project structure.
### 3. Multi-Root Workspace
The `LEDMatrix.code-workspace` file configures VS Code/Cursor to open all plugin repositories as separate workspace roots, allowing easy development across all repos.
## Setup Scripts
### Initial Setup
If you already have plugin repositories cloned, use the setup script:
Clone ledmatrix-plugins into the same parent directory as LEDMatrix (the
workspace file and `scripts/update_plugin_repos.py` look for
`../ledmatrix-plugins` relative to the LEDMatrix root):
countdown, birdnet-go, ledmatrix-music and odds-ticker. They arrive in bursts
("Whole group deferred; strip will extend as it drains") every minute or so,
12 fetches in five minutes. That is the "occasional pause" a viewer sees.
An 8-minute soak (`scripts/frame_soak.py --preview`) of the #628 build on
hdpi:
| late by | frames |
|---|---|
| 1 refresh | 238 |
| 2 | 32 |
| 3–5 | 30 |
| 6+ | 5 |
| freezes ≥ 250 ms | 2 (0.97 s total) |
The 3+ rows and the freezes are the pauses. The single-refresh row is a
separate problem: the blit is 6 ms of a 10 ms refresh, so there is little
slack. It is covered under *What this does not fix*.
## Why a plugin is canvas-bound
The plugin-facing canvas is a set of shared attributes on `DisplayManager`:
`image`, `draw`, `matrix`, and the `width`/`height` properties that read from
`matrix`. Three adapter paths (`src/vegas_mode/plugin_adapter.py`) need them,
and each returns `None` under `offscreen_only=True` so the plugin is queued for
the render thread:
1. **Display capture** (`_capture_display_content`): clear the canvas, call
`plugin.display()`, copy `display_manager.image`. Used by any plugin
without `get_vegas_content()` or a populated `scroll_helper`.
2. **Scroll-content generation** (`_trigger_scroll_content_generation`): a
ticker plugin whose `scroll_helper.cached_image` is empty is made to build
it by calling `display(force_clear=True)` or `_create_scrolling_display()`.
Both draw on the canvas.
3. **Narrowed rendering** (`DisplayManager.render_size`): swaps the shared
`matrix`, `image` and `draw` for a narrower set so the plugin lays out for
`render_width_pct`. The render thread would see the swap mid-frame.
The render thread keeps the canvas coherent only because nothing else touches
it at the same time. A background thread can't use it.
## The design: a per-thread render target
`capture_mode()` is already per-thread (#423 made its state a
`threading.local`, so a background capture no longer suppresses the render
loop's pushes). The same move applies to the canvas itself:
```python
with display_manager.offscreen(width=None, height=None) as surface:
plugin.display(force_clear=True)
content = surface.image.copy()
```
For the **calling thread only**, inside the block:
| accessor | resolves to |
|---|---|
| `display_manager.image`, `.draw` | the surface's own image and draw: a fresh black canvas, `fontmode = "1"` |
| `display_manager.matrix` | a logical proxy reporting the surface size, so `width`/`height` and plugins that read `matrix.width` follow it. Hardware calls through it (`SetImage`, `SwapOnVSync`, `Clear`, brightness writes) are inert. |
| `update_display()`, `clear()` | canvas-only: the block implies capture mode, which is already per-thread |
| `set_scrolling_state()`, `set_frame_hold()` | no-ops, so a plugin's `display()` cannot re-pace the live scroll. Today it can, when it is captured on the render thread. |
Every other thread sees the real canvas, unchanged. The render loop in
particular keeps presenting while a plugin draws elsewhere.
### Implementation sketch
- `image`, `draw` and `matrix` become properties over `_image`, `_draw` and
`_matrix`, plus a thread-local current surface. The getter returns the
surface's value when the calling thread has one, else the shared one; setters
mirror that. That costs about 0.1 µs per access, and `update_display()` reads
each a handful of times per frame. Every existing `self.image = ...` in
`DisplayManager` (`clear()`, setup, fallback) keeps working and becomes
thread-correct for free.
- `render_size()` is rebuilt on `offscreen()`: it creates or narrows the
calling thread's surface instead of swapping shared state.
- `offscreen()` nests and always restores on exit, including when the plugin
raises.
- `VisualDisplayManager` (the plugin test harness) gets the same method, for
parity.
### Adapter changes
- `get_content(offscreen_only=True)` stops returning `None` for the three
paths above. Each runs inside `display_manager.offscreen(render_width)`.
- `_capture_display_content` and `_trigger_scroll_content_generation` drop
their "copy the shared image, restore it afterwards" bookkeeping, since the
shared image is never touched.
- **Take the plugin's lock.**`PluginManager.get_plugin_lock()` keeps
`update()` and `display()` mutually exclusive in normal rotation, but Vegas
never takes it, so today's render-thread captures already race
`update()`. Off the render thread the adapter can afford to wait: blocking
acquire with a timeout (proposed 2 s). On timeout it keeps the cached segment
and tries again next group.
- `drain_deferred()` and the deferred queue are deleted. The only render-thread
fetch left is the inline fallback when no prepared group is ready, which in
practice is the first extension. Prefetching at start removes that too.
## Keeping live content fresh
Offscreen rendering is also what makes fresh sports scores possible. Today a
plugin's segment is drawn when its group is prefetched, and the strip carries
7,000–10,000 px of content ahead of the viewport (hdpi logs: "7153px still
ahead", "9842px ahead"). At ~100 px/s, a score drawn now reaches the screen
70–100 seconds later. When a plugin reports new data, Vegas only drops its
cache (`invalidate_pending_updates`), so the change is drawn on the plugin's
*next* turn, several minutes later. A segment already in the strip scrolls by
with the data it was drawn with.
That was the right trade while every redraw of a canvas-bound plugin stalled
the scroll. Off the render thread a redraw costs the scroll nothing, so the
strip can afford three things.
### 1. Refresh at the gate
Before a segment enters the viewport, check whether its plugin has updated
since the segment was drawn. If it has, redraw it offscreen and replace it
while it is still out of sight. Width changes are fine here, because
everything from that segment onward is still invisible.
The gate sits `lead` pixels ahead of the viewport's right edge:
`lead = max(one screen, speed × (render time + margin))`. The render time is
the plugin's own, measured on each render (sports cards take the longest,
hundreds of ms up to seconds per the prefetch notes). A plugin whose render
does not finish before its segment reaches the viewport keeps the old segment.
The scroll never waits for it.
Content is then at most `lead / speed` seconds old when it appears, a few
seconds instead of minutes, without changing how far ahead the rotation
fetches.
### 2. Replace ahead of the screen
When a plugin reports new data (the Vegas update tick already names them), any
of its segments that are **anywhere ahead of the viewport** are redrawn and
replaced straight away, not only at the gate. That covers the long stretch of
strip between prefetch and the gate.
### 3. Update on screen
A segment that is already **visible** is patched in place when the redrawn
version has the same geometry: the same total width, and the same width for
each card (a sports plugin returns one image per game, joined with
`intra_plugin_gap`). Scoreboard cards keep a fixed layout, so a score change
patches in and the digits update as the card scrolls past. The patch is a
pixel copy of one card (a 150×64 card is ~29 KB) applied by the render thread
between frames, so a frame never shows half of a patch.
When the geometry differs (a game added or dropped, a card that grew), the
visible part cannot change without a jump. Only the cards not yet on screen
are replaced, and only if the geometry up to that point is unchanged. Otherwise
the segment keeps its snapshot until it has scrolled off.
### Avoiding wasted work
- **Change detection.**`run_scheduled_updates_with_changes()` names a plugin
whenever its `update()` ran, not when its data changed. On hdpi
`clock-simple` and `ledmatrix-music` are named on every 4-second tick. A
redraw whose pixels hash the same as the segment's is discarded without a
swap.
- **Redraw on real updates only.** Vegas makes no API calls. Each plugin
fetches on its own schedule, and a redraw is triggered only when the
plugin's `update()` has run since its segment was drawn. On hdpi live
football, baseball and hockey poll every 30 s (live odds every 60 s,
everything else hourly), so a live sports card is redrawn once per poll.
- **Floor.** A plugin is redrawn at most once per
`vegas_scroll.refresh_min_interval` (proposed 10 s), and never while its
previous redraw is still running. The floor never holds back a sports card
polling every 30 s. It exists for chatty plugins: `clock-simple` updates
every second and `ledmatrix-music` polls every 2 s.
- **One worker.** Redraws go through the same background worker as prefetch,
one plugin at a time at `nice 10`, under the plugin's lock.
Data freshness is still bounded by each plugin's own fetch interval (how often
it polls live scores). Drawing faster cannot beat the data source.
### The strip becomes a list of segments
All three need the strip to be replaceable by segment. Today it is one
image (`ScrollHelper.cached_array`, 8,000–20,000 px wide, 1.5–3.8 MB), and
`append_content()` rebuilds the whole thing on the render thread for every
appended block. That is also a pause source.
Proposed `SegmentStrip`, used by Vegas in place of the single image:
- an ordered list of segments: plugin id, card boundaries, a pixel array, the
render time, and the plugin data version it was drawn from, plus its
x-offset in the strip;
- `visible(x, width)` assembles the viewport by slicing across at most a few
segments: the same ~100 KB copy per frame that slicing the single image
costs today;
- append and trim become O(block) list operations, not a copy of the strip;
- replace swaps one list entry and shifts the offsets of the segments after it
(dozens at most). A same-geometry patch copies pixels into the existing array.
Every mutation is prepared off the render thread and applied by the render
thread at a frame boundary, so the strip the render loop reads is never
half-changed.
### Multi-display sync
The follower renders from its own copy of the strip, offset from the leader's
scroll position. Today the leader sends that copy whole, and only in
`start_new_cycle()` (`send_scroll_image`), plus the scroll position every
frame. Continuous scroll, the default, extends and trims the strip without
starting a new cycle, and nothing sends those changes. From reading the code,
the follower therefore probably falls out of step after the first extension
already, before any of this design. That is untested; it needs a two-Pi rig.
With a segment strip, keeping the follower identical becomes **replaying the
leader's operations**:
- Every strip mutation (append, trim, replace, patch) is one operation in
strip coordinates. The leader applies it and sends the same operation to the
follower over the existing TCP channel. Segments are small: a card is ~29 KB
raw and compresses well.
- Operations on off-screen segments apply on arrival. A patch to a segment
that is on either panel carries an *apply at scroll position X* stamp a
couple of hundred milliseconds ahead. Both sides apply it when their scroll
position passes X, so both panels change on the same frame, within the
existing position-sync jitter.
- Each operation carries a sequence number. A follower that sees a gap (a
reconnect, a dropped message) asks for a full snapshot, which is today's
`send_scroll_image` path.
That also fixes the probable continuous-mode gap as a side effect, since
appends and trims become operations too. Until it is in place, fresh-content
updates are disabled while sync is active.
## Risks, and what was checked
1. **Plugins holding their own reference to the shared `draw` or `image`.**
They would keep drawing into the shared canvas, and routing by thread can't
redirect them. A grep of the 49 plugins installed on hdpi found none storing
`display_manager.draw` or `.image` in an attribute (a pattern search, so
indirect aliasing would slip past it). A plugin that did would
draw into an image nobody displays, which trims to a blank segment. That is
not corruption, and it is no worse than today.
2. **Plugins calling the matrix directly.** None in the audit. Inside
`offscreen()` the proxy makes it inert anyway.
3. **Font thread-safety.**`FontManager` shares font objects across plugins.
Measured on Pillow 12.3, two threads rendering text take 1.94× as long as
one, so text rendering holds the GIL and FreeType is never entered
concurrently. Re-check if Pillow changes that.
4. **Plugin thread-safety.**`display()` moves to the prefetch thread. The
plugin lock makes it exclusive with `update()`, which is more protection
than it has today. Threads a plugin starts itself are not covered, as today.
5. **The GIL.** Moving 40–600 ms of plugin rendering off the render thread
removes the pauses, but the work still needs the GIL. Pillow drawing holds
it, and a waiting thread only gets it back after the switch interval
(default 5 ms). Expect some single-refresh late frames while a prefetch
runs. Measure with the soak. A render process separate from plugin work
is the structural answer (the "native presenter" step). Two experiments
get most of the way first (results under Status, above):
- `vegas_scroll.switch_interval_ms` lowers the switch interval for a Vegas
run (1 ms is the obvious try), so the render thread waits at most that
long behind bytecode. It does nothing for a C call that keeps the GIL.
- `vegas_scroll.prefetch_gate` (`src/common/render_gate.py`) lets the
prefetch thread run Python only while the render thread is blocked in
`SwapOnVSync`, up to just before the refresh the swap returns on, and
parks it the rest of the time. That covers C calls too, since the gate is
checked before each one starts. It never parks the thread while it holds
a lock the render thread takes, and never for more than 50 ms. It needs
the rebuilt binding, which releases the GIL during the swap. On by
default.
## What this does not fix
- **The blit.** Copying a 512×64 frame into the matrix (`SetImage`) is ~6 ms at
8 PWM bits on a Pi 4, leaving ~4 ms of slack per refresh. That is the main
source of the single-refresh late frames. Holding frames for two refreshes
(≈50 px/s) doubles the budget. Cutting the blit itself is the native-presenter
step.
- **Live refreshes pushed from `update()`.** Some sports plugins call
`display()` and `update_display()` from inside `update()`, which runs on the
update worker and can push to the panel mid-Vegas. That is a separate
hazard. `offscreen()` gives a tool for it (run the update worker offscreen
while Vegas owns the panel), but it is out of scope here.
## Test plan
- **Unit, `DisplayManager`:** one thread inside `offscreen()` draws while
another reads `image`/`draw`/`matrix`/`width`/`height` and sees the real
canvas. Also: `update_display()` and `set_scrolling_state()` are inert inside;
`render_size()` narrows only the calling thread; nesting and exceptions
restore state.
- **Unit, adapter:** a stub display-capture plugin and a stub scroll-helper
plugin both return content with `offscreen_only=True`, and nothing is queued
for the render thread. The plugin lock is taken, and a timeout keeps the cached
segment.
- **Emulator integration:** a stub canvas-bound plugin whose `display()` sleeps
300 ms. The Vegas render loop never goes a frame without presenting (frame
timing recorder: zero freezes).
- **Unit, `SegmentStrip`:** the viewport assembled across segment boundaries
matches slicing one concatenated image, pixel for pixel. Append, trim,
replace-ahead and same-geometry patch each leave every other column
unchanged. A geometry-changing patch of a visible segment is refused.
- **Freshness:** a stub sports plugin whose score changes every second. The
score on screen is never older than `lead / speed` plus the plugin's fetch
interval. A visible card's digits change without the frame-timing recorder
seeing a late frame. An unchanged redraw is discarded.
- **Hardware:** an hdpi soak, A/B against the #628 build, alternating order.
Targets: no freezes, an empty 6+ bucket, the 3–5 bucket near zero, and the late
rate below 0.66%. Plus, for freshness: log each segment's age when it enters
the viewport, and compare the median and max before and after.
## Rollout
Three changes, each soaked on hdpi before the next:
1. **Offscreen rendering:**`offscreen()`, the adapter on the prefetch thread,
and the plugin lock. Removes the render-thread pauses.
2. **`SegmentStrip`:** Vegas's strip becomes a list of segments. Removes the
whole-strip copy on append. No visible behaviour change.
3. **Fresh content:** refresh at the gate, replace ahead, patch on screen,
Who owns what on an installed system, which privileged commands the web
interface may run, and how to repair ownership when it goes wrong. The
installer, [`first_time_install.sh`](../first_time_install.sh), sets all of
this up; this page describes the result.
## Users and groups
| Account | Used by | Why |
|---|---|---|
| `root` | `ledmatrix.service` (the display) | The LED matrix library needs direct GPIO access |
| The installing user (e.g. `ledpi`) | `ledmatrix-web.service`, `ledmatrix-update-verify.service` | A web server should not run as root |
| `ledmatrix` group | shared files | Members: the installing user, `root`, and `daemon` if it exists. Created by [`setup_cache.sh`](../scripts/install/setup_cache.sh) and the installer |
The installer also adds the web user to `systemd-journal` and `adm` so the
**Logs** tab can read the journal. Group changes apply after the user logs
in again (services pick them up on restart).
## Files and directories
| Path | Owner | Mode | Notes |
|---|---|---|---|
| Project directory | web user | dirs `755`, files `644`, `*.sh``755` | Set in the installer's "Normalize project file permissions" step |
| `config/` | web user | `2775` | |
| `config/config.json` | web user | `644` | Written by the web interface |
| `config/config_secrets.json` | web user : `ledmatrix` | `640` | Owned by the web user because the web interface writes it; root reads it regardless of mode |
| `plugin-repos/`, `plugins/` | web user | dirs `2775`, files `664` | The web interface installs and removes plugins |
| `assets/` | web user | dirs `755`, files `644` | Root writes downloaded logos regardless |
| `scripts/fix_perms/safe_plugin_rm.sh`, `safe_pip_install.sh` | `root:root` | `755` | Run as root through sudo, so the web user must not be able to edit them |
- `cp` of `/tmp/hostapd.conf` and `/tmp/dnsmasq.conf` to their fixed
destinations, and `rm -f /etc/dnsmasq.d/ledmatrix-captive.conf`
- `cp /tmp/ledmatrix-nm-dnsmasq.conf` to
`/etc/NetworkManager/dnsmasq-shared.d/ledmatrix-captive.conf`, and
`rm -f` of that file
**`iptables` is deliberately not granted.** The captive portal's rules are
built from the interface name and port, so a rule covering them would need a
trailing wildcard, and `iptables --modprobe=<path>` runs `<path>` as root: a
wildcard grant is a root shell for the web user. Doing it safely needs a
wrapper script that builds the rules itself, like `safe_plugin_rm.sh`. On a
stock Raspberry Pi OS image the default user's blanket `NOPASSWD` rule
(`/etc/sudoers.d/010_pi-nopasswd`) hides this gap.
### polkit
The same script installs `/etc/polkit-1/rules.d/10-ledmatrix-wifi.rules`,
which lets the web user perform any `org.freedesktop.NetworkManager.*`
action without authentication.
## Repair scripts
In [`scripts/fix_perms/`](../scripts/fix_perms/). Run them from the project
directory.
| Script | Run as | What it does | Notes |
|---|---|---|---|
| `fix_plugin_permissions.sh` | `sudo` | `plugins/` and `plugin-repos/` to `root:<user>`, dirs `2775`, files `664`; makes a `700` home directory `755` so root can traverse it | Safe. Group-writable, so the web user keeps write access |
| `fix_assets_permissions.sh` | `sudo` | `assets/` to `<user>:<group>`, mode `777` recursively | Works, but looser than the installer's `755`/`644` |
| `fix_cache_permissions.sh` | `sudo` | Runs [`setup_cache.sh`](../scripts/install/setup_cache.sh) for `/var/cache/ledmatrix` (`root:ledmatrix`, `2775`, files `660`), then makes `~/.ledmatrix_cache` (the fallback cache) `<user>:<group>` mode `777` | Safe. The `~/.ledmatrix_cache` mode is still `777` |
| `fix_web_permissions.sh` | the web user, **without**`sudo` | Resets project file ownership for the web user (it calls `sudo` itself), then makes `safe_plugin_rm.sh` and `safe_pip_install.sh``root:root``755` again and restores `config_secrets.json` to its owner, group `ledmatrix`, mode `640` | Refuses to run as root. It does not write sudoers rules |
| `safe_plugin_rm.sh`, `safe_pip_install.sh` | — | Called by the web interface through sudo | Not for manual use |
To reinstall the sudoers rules, run
`./scripts/install/configure_web_sudo.sh` (web rules) or
`./scripts/install/configure_wifi_permissions.sh` (WiFi rules and polkit) as
Three parts of core check `manifest.json`, each for a different set of
fields:
| Check | Fields | What happens when one is missing |
|---|---|---|
| JSON schema, [`schema/manifest_schema.json`](../schema/manifest_schema.json) | `id`, `name`, `version`, `author`, `entry_point`, `class_name`, `compatible_versions` | Install from URL logs a warning (`PluginStoreManager._validate_manifest_schema()`); nothing is refused |
| Plugin Store install, [`src/plugin_system/store_install.py`](../src/plugin_system/store_install.py) | `id`, `name`, `class_name`, `display_modes` | Install is refused. A registry install first tries to detect a missing `class_name` from the entry-point file |
| Plugin loader, [`src/plugin_system/plugin_loader.py`](../src/plugin_system/plugin_loader.py) | `class_name` | The plugin fails to load |
Defaults and other uses:
- `entry_point` defaults to `manager.py`; the store writes the default back
into the manifest on install.
- `compatible_versions` (a list of semver ranges such as `">=2.0.0"`) is how
the store decides whether a plugin can run on this core. An install is
refused only when the field excludes the running version
> The current implementation lives in `web_interface/app.py`,
> `web_interface/blueprints/api_v3.py`, and `web_interface/templates/v3/`.
> The user-facing description (Overview, Features, Form Generation
> Process) is still accurate.
## Overview
Each installed plugin now gets its own dedicated configuration tab in the web interface. This provides a clean, organized way to configure plugins without cluttering the main Plugins management tab.
Each installed plugin now gets its own dedicated configuration tab in the web interface. This provides a clean, organized way to configure plugins without cluttering the **Plugin Manager** tab.
## Features
@@ -20,24 +10,27 @@ Each installed plugin now gets its own dedicated configuration tab in the web in
- **JSON Schema-Based Forms**: Configuration forms are automatically generated based on each plugin's `config_schema.json`
- **Type-Safe Inputs**: Form inputs are created based on the JSON Schema type (boolean, number, string, array, enum)
- **Default Values**: All fields show current values or fallback to schema defaults
- **Reset Functionality**: Users can reset all settings to defaults with one click
- **Real-Time Validation**: Input constraints from JSON Schema are enforced (min, max, maxLength, etc.)
## User Experience
### Accessing Plugin Configuration
1. Navigate to the **Plugins** tab to see all installed plugins
1. Navigate to the **Plugin Manager** tab to see all installed plugins
2. Click the **Configure** button on any plugin card
3. You'll be automatically taken to that plugin's configuration tab
4. Alternatively, click directly on the plugin's tab button (marked with a puzzle piece icon)
4. Alternatively, click directly on the plugin's tab button in the second nav row
### Configuring a Plugin
1. Open the plugin's configuration tab
2. Modify settings using the generated form
3. Click **Save Configuration**
4. Restart the display service to apply changes
3. Click **Save Configuration**. The settings apply to the running display
without a restart: the display service reloads `config.json` when it
changes and calls the plugin's `on_config_change()`
The tab also has **Refresh** (reload the form), **Update** (update the
plugin) and **Uninstall** buttons.
### Plugin Manager vs Per-Plugin Configuration
@@ -52,22 +45,13 @@ Each installed plugin now gets its own dedicated configuration tab in the web in
### Requirements
To enable automatic configuration tab generation, your plugin must:
Every installed plugin gets a tab. To get a generated form in it, include a
`config_schema.json` file in the plugin's directory. The name is fixed: the
web interface finds the schema by that file name (`SchemaManager` in
`src/plugin_system/schema_manager.py`), and no manifest field points to it.
3. See confirmation: "Configuration saved for hello-world. Restart display to apply changes."
4. Restart the display service
3. See the confirmation notification. Plugin settings apply live: the
display service reloads `config.json` when it changes and passes the new
settings to the plugin's `on_config_change()`
## 🛠️ For Plugin Developers
@@ -105,19 +105,14 @@ Create `config_schema.json` in your plugin directory:
}
```
Reference it in `manifest.json`:
**Done!** The file name is fixed: the web interface looks for
`config_schema.json` in the plugin's directory; there is no manifest field
for it. Every installed plugin gets a tab; the schema is what turns it into a
form.
```json
{
"id": "my-plugin",
"icon": "fas fa-star", // Optional: add a custom icon!
"config_schema": "config_schema.json"
}
```
**Done!** Your plugin now has a configuration tab.
**Bonus:** Add an `icon` field for a custom tab icon! Use Font Awesome icons (`fas fa-star`), emoji (⭐), or custom images. See [PLUGIN_CUSTOM_ICONS.md](PLUGIN_CUSTOM_ICONS.md) for the full guide.
**Bonus:** an `icon` field in `manifest.json` names a Font Awesome class for
the tab (`"icon": "fas fa-star"`). See
[PLUGIN_CUSTOM_ICONS.md](PLUGIN_CUSTOM_ICONS.md).
## 🎨 Supported Input Types
@@ -171,12 +166,10 @@ User enters: `255, 0, 0`
### For Users
1. **Reset Anytime**: Use "Reset to Defaults" to restore original settings
2. **Navigate Back**: Switch to the **Plugin Manager** tab to see the
1. **Navigate Back**: Switch to the **Plugin Manager** tab to see the
full list of installed plugins
3. **Check Help Text**: Each field has a description explaining what it does
4. **Restart Required**: Remember to restart the display service from
**Overview** after saving
2. **Check Help Text**: Each field has a description explaining what it does
3. **No Restart Needed**: Saved plugin settings apply to the running display
### For Developers
@@ -189,18 +182,17 @@ User enters: `255, 0, 0`
## 🔧 Troubleshooting
### Tab Not Showing
- Check that `config_schema.json` exists
- Verify `config_schema` is in `manifest.json`
- Check that the plugin is installed and listed under **Plugin Manager**
- Refresh the page
- Check browser console for errors
### Settings Not Saving
- Ensure plugin is properly installed
- Restart the display service after saving
- Check that all required fields are filled
- Look for validation errors in browser console
### Form Looks Wrong
- Check that `config_schema.json` is in the plugin's directory
Plugins can specify custom icons that appear next to their name in the web interface tabs. This makes your plugin instantly recognizable and adds visual polish to the UI.
A plugin can name an icon for its tab in the web interface's second nav row
(next to **Plugin Manager**) with the `icon` field in `manifest.json`.
## Icon Types Supported
`GET /api/v3/plugins/installed` passes the manifest's `icon` through (a
non-string value comes back as `null`), and a plugin without one gets the
default puzzle piece.
The system supports three types of icons:
## Font Awesome classes only
### 1. Font Awesome Icons (Recommended)
`icon` is used verbatim as the CSS class of an `<i>` element
(`iconEl.className = plugin.icon || 'fas fa-puzzle-piece'` in
`web_interface/static/v3/js/app-shell.js` and the same fallback in
`app-early.js`). So it must be a Font Awesome class string. Emoji, image
paths and URLs are not supported: they would end up as a meaningless class
name and render nothing.
The web interface uses Font Awesome 6, giving you access to thousands of icons.
The web interface bundles Font Awesome Free 6
(`web_interface/static/v3/vendor/fontawesome/`), so any free `fas`, `far` or
`fab` icon works.
**Example:**
```json
{
"id": "my-plugin",
@@ -21,292 +30,33 @@ The web interface uses Font Awesome 6, giving you access to thousands of icons.
✅ **Tab & Header Icons** - Icons appear in both tab buttons and configuration page headers
## How It Works
### For Plugin Developers
Simply add an `icon` field to your plugin's `manifest.json`:
```json
{
"id": "my-plugin",
"name": "My Plugin",
"icon": "fas fa-star", // ← Add this line
"config_schema": "config_schema.json",
...
}
```
### Three Icon Types Supported
#### 1. Font Awesome Icons (Recommended)
```json
"icon": "fas fa-clock"
```
Best for: Professional, consistent UI appearance
#### 2. Emoji Icons (Fun!)
```json
"icon": "⏰"
```
Best for: Colorful, fun plugins; no setup needed
#### 3. Custom Images
```json
"icon": "/plugins/my-plugin/logo.png"
```
Best for: Unique branding; requires image file
## Implementation Details
### Frontend Changes (`templates/index_v2.html`)
**New Function: `getPluginIcon(plugin)`**
- Checks if plugin has `icon` field in manifest
- Detects icon type automatically:
- Contains `fa-` → Font Awesome
- 1-4 characters → Emoji
- Starts with URL/path → Custom image
- Otherwise → Default puzzle piece
**Updated Functions:**
- `generatePluginTabs()` - Uses custom icon for tab button
- `generatePluginConfigForm()` - Uses custom icon in page header
### Example Plugin Updates
**hello-world plugin:**
```json
"icon": "👋"
```
**clock-simple plugin:**
```json
"icon": "fas fa-clock"
```
## Code Example
Here's what the icon detection logic does. **Important:** Plugin manifests must be treated as untrusted input and require escaping/validation before rendering.
```javascript
// Helper function to escape HTML entities
function escapeHtml(text) {
const div = document.createElement('div');
div.textContent = text;
return div.innerHTML;
}
// Helper function to validate and sanitize image URLs
function isValidImageUrl(url) {
if (!url || typeof url !== 'string') {
return false;
}
// Only allow http, https, or relative paths starting with /
The LEDMatrix system has smart dependency installation that adapts based on who is running it. This guide explains how it works and potential pitfalls.
A plugin lists its Python packages in its `requirements.txt`. LEDMatrix
installs them for you when a plugin is installed, updated or loaded. This
guide explains where they end up and what to do when a plugin can't import a
package.
## How It Works
The rule to remember: **packages must be importable by `ledmatrix.service`,
which runs as root.** Anything installed only into another user's
`~/.local/` is invisible to it.
### Execution Context Detection
## Who Runs What
The plugin manager checks if it's running as root:
| `ledmatrix-web.service` (web UI) | the user who ran the installer (e.g. `ledpi`) | `User=__USER__` in `systemd/ledmatrix-web.service`, filled in by `scripts/install/install_service.sh` |
## How Dependencies Get Installed
### 1. Installing or updating a plugin from the web UI
The web interface is not root, so it installs through a narrow sudo helper:
This guide helps resolve issues with automatic plugin dependency installation in the LEDMatrix system.
This guide helps resolve problems installing a plugin's Python packages. For
how installation works, see the [Plugin Dependency Guide](PLUGIN_DEPENDENCY_GUIDE.md).
## Common Error Symptoms
@@ -10,109 +11,118 @@ ERROR: Could not install packages due to an OSError: [Errno 13] Permission denie
WARNING: The directory '/root/.cache/pip' or its parent directory is not owned or is not writable
```
### Context Mismatch
### Installed for the wrong user
The pip output shown after a web-UI install starts with:
```
WARNING: Installing plugin dependencies for current user (not root).
These will NOT be accessible to the systemd service.
[Root install unavailable (...); installed for the current process's user only.
Packages may not be visible to ledmatrix.service if it runs as a different
user — run scripts/install/configure_web_sudo.sh to fix this.]
```
### Plugin fails to load with `ModuleNotFoundError`
The display service can't see a package the plugin needs.
## Root Cause
Plugin dependencies must be installed in a context accessible to the LEDMatrix systemd service, which runs as root. Permission errors typically occur when:
Plugin packages must be importable by `ledmatrix.service`, which runs as
root. The web interface (`ledmatrix-web.service`) runs as the user who
installed LEDMatrix, so it installs through a sudo helper
(`scripts/fix_perms/safe_pip_install.sh`). Problems usually come from:
1. The pip cache directory has incorrect permissions
2. The process tries to install to user directories without proper permissions
3. Environment variables (like HOME) are not set correctly for the service context
1. The sudoers rule for that helpermissing, so the web UI installed the
packages for its own user only
2. Running `python3 run.py` by hand as a normal user, which installs missing
packages into `~/.local/`
3. pip's cache directory not being writable for root
## Solutions
### Solution 1: Use the Manual Installation Script (Recommended)
We provide a helper script that handles dependency installation correctly:
### Solution 1: Restore the sudo rule, then reinstall
```bash
# Run as root to install system-wide (for production)
This guide explains how to set up a development workflow for plugins that are maintained in separate Git repositories while still being able to test them within the LEDMatrix project.
> **Rendering guidance:** plugins should read the display size dynamically
> (`self.display_manager.matrix.width/height`) rather than hardcoding one
> panel. For plugins that want to *scale* their layout to any panel, the
> (`self.display_manager.width/height`) rather than hardcoding one
> panel. Don't read `display_manager.matrix.width/height`: `matrix` is
> `None` when hardware init fails, while the `width`/`height` properties
> fall back to the canvas size. For plugins that want to *scale* their layout to any panel, the
> opt-in adaptive layout system ([ADAPTIVE_LAYOUT.md](ADAPTIVE_LAYOUT.md))
> provides the shared helpers — fonts, images, and composite layouts that
> scale. Existing plugins keep their classic rendering unless they adopt
> those APIs; nothing migrates automatically.
> **Just want a different look for an existing sports scoreboard?** You may
> not need a plugin at all — a **skin** restyles the live/recent/upcoming
> rendering while the plugin keeps handling data, scheduling, caching, and
> vegas mode, in ~100 lines of drawing code. See
> [CREATING_SKINS.md](CREATING_SKINS.md).
## Overview
When developing plugins in separate repositories, you need a way to:
@@ -43,28 +39,45 @@ The solution uses **symbolic links** to connect plugin repositories to the `plug
## Quick Start
### 1. Link a Plugin from GitHub
Official plugins all live in one repository,
[ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins), with
one directory per plugin under `plugins/` (there are no per-plugin
`ledmatrix-<name>` repositories). The helper script links a plugin directory
from a checkout of that monorepo into LEDMatrix's `plugins/` directory.
The easiest way to link a plugin that's already on GitHub:
### 1. Link an Official Plugin
```bash
./scripts/dev/dev_plugin_setup.sh link-github music
# Clear errors older than 24 hours (the default), or all of them
curl -X POST http://localhost:5000/api/v3/errors/clear
curl -X POST -H 'Content-Type: application/json' -d '{"all": true}' \
http://localhost:5000/api/v3/errors/clear
```
The web interface is a separate process, so it reads a snapshot the display
service writes to the shared cache directory (`plugin_error_snapshot`): at most
every 10 seconds, and only when something changed. Expect the numbers to lag
by up to about 15 seconds, and to start from zero when the display service
restarts. `snapshot_available` is `false` until the display service has
reported. A clear is a request the display service applies within about 5
seconds; the API hides the cleared errors immediately. Details and response
shapes: [REST API reference](REST_API_REFERENCE.md#error-tracking).
### Error Patterns
When the same error occurs repeatedly (5+ times in 60 minutes), it's detected as a pattern and logged as a warning. This helps identify systemic issues.
> manager). Drift from current reality is called out inline.
This document provides a comprehensive overview of the plugin architecture implementation, consolidating details from multiple plugin-related implementation summaries.
## Executive Summary
The LEDMatrix plugin system transforms the project into a modular, extensible platform where users can create, share, and install custom displays through a GitHub-based store (similar to Home Assistant Community Store).
The LEDMatrix plugin system successfully transforms the project into a modular, extensible platform. The implementation provides:
- **For Users**: Easy plugin discovery, installation, and management
- **For Developers**: Clear plugin API and development tools
- **For Maintainers**: Smaller core codebase with community contributions
The system maintains full backward compatibility while enabling future growth through community-developed plugins. All major components are implemented, tested, and ready for production use.
---
*This document consolidates plugin implementation details from multiple phase summaries into a comprehensive technical overview.*
This guide explains how to set up and maintain your official plugin registry at [https://github.com/ChuckBuilds/ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins).
This page explains how the official plugin registry works and how a plugin
gets into it. The registry and the official plugins both live in one
its `SUBMISSION.md`, `VERIFICATION.md` and `docs/` are the authoritative
contributor guides.
## Overview
## How it fits together
Your plugin registry serves as a **central directory** that lists all official, verified plugins. The registry is just a JSON file; the actual plugins live in their own repositories.
## Repository Structure
```
```text
ledmatrix-plugins/
├── README.md # Main documentation
├── LICENSE # GPL-3.0
├── plugins.json # The registry file (main file!)
├── SUBMISSION.md # Guidelines for submitting plugins
├── VERIFICATION.md # Verification checklist
└── assets/ # Optional: screenshots, badges
└── screenshots/
├── plugins/
│ ├── clock-simple/ # one directory per official plugin
│ │ ├── manifest.json # source of truth for the plugin's version
│ │ ├── manager.py
│ │ ├── config_schema.json
│ │└── requirements.txt
│ └── ...
├── plugins.json # the registry the Plugin Store reads
└── update_registry.py # regenerates plugins.json from the manifests
**Note**: There's no need for version arrays or release tracking. The store queries GitHub for the latest commit details (date, branch, and short SHA) whenever metadata is requested.
[plugin_registry_template.json](plugin_registry_template.json) shows a
monorepo entry and a third-party entry.
## Step 2: Create Plugin Repositories
Don't edit `latest_version` or `last_updated` by hand for monorepo plugins:
`update_registry.py` in ledmatrix-plugins writes them from each plugin's
`manifest.json`.
Each plugin should have its own repository:
## Adding or changing an official plugin
### Example: Creating clock-simple Plugin
1. Add or edit `plugins/<your-plugin-id>/` in the monorepo, with the
3. **Testing**: Test installation and basic functionality
4. **Approval**: If accepted, merged and marked as verified
## After Approval
- Plugin appears in official store
- `verified: true` badge shown
- Included in plugin count
- Featured in README
## Updating Your Plugin
Whenever you push new commits to your plugin repository's default branch, the store will automatically surface the latest commit timestamp and short SHA. No release tagging or manifest version bumps are required.
The LEDMatrix Plugin Store allows you to discover, install, and manage display plugins for your LED matrix. Install curated plugins from the official registry or add custom plugins directly from any GitHub repository.
In the web interface, the **Plugin Store** is a section of the **Plugin
Manager** tab (below the installed plugins), followed by an **Install from
GitHub** section.
The Python examples below pass `plugins_dir="plugin-repos"`:
`PluginStoreManager()` defaults to `plugins`, but the web interface and the
plugin loader use `plugin_system.plugins_directory` from `config.json`
(`plugin-repos` by default).
---
## Quick Reference
### Install from Store
```bash
# Web UI: Plugin Store → Search → Click Install
# Web UI: Plugin Manager → Plugin Store section → Search → Click Install
# API:
curl -X POST http://your-pi-ip:5000/api/v3/plugins/install \
-H "Content-Type: application/json" \
@@ -19,7 +28,7 @@ curl -X POST http://your-pi-ip:5000/api/v3/plugins/install \
### Install from GitHub URL
```bash
# Web UI: Plugin Store → "Install from URL" → Paste URL
# Web UI: Plugin Manager → Install from GitHub → "Install Single Plugin" → Paste URL
# API:
curl -X POST http://your-pi-ip:5000/api/v3/plugins/install-from-url \
-H "Content-Type: application/json" \
@@ -57,7 +66,7 @@ The official plugin store contains curated, verified plugins that have been revi
**Via Web Interface:**
1. Open the web interface at http://your-pi-ip:5000
2. Navigate to the "Plugin Store" tab
2. Navigate to the "Plugin Manager" tab and scroll to the "Plugin Store" section
3. Browse or search for plugins
4. Click "Install" on the desired plugin
5. Wait for installation to complete
@@ -74,7 +83,7 @@ curl -X POST http://your-pi-ip:5000/api/v3/plugins/install \
```python
from src.plugin_system.store_manager import PluginStoreManager
store = PluginStoreManager()
store = PluginStoreManager(plugins_dir="plugin-repos")
success = store.install_plugin('clock-simple')
if success:
print("Plugin installed!")
@@ -90,10 +99,11 @@ Install any plugin directly from a GitHub repository, even if it's not in the of
**Via Web Interface:**
1. Open the web interface
2. Navigate to the "Plugin Store" tab
3. Find the "Install from URL" section
2. Navigate to the "Plugin Manager" tab
3. Find "Install Single Plugin" in the "Install from GitHub" section
4. Paste the GitHub repository URL (e.g., `https://github.com/user/ledmatrix-my-plugin`)
5. Click "Install from URL"
and optionally a branch
5. Click "Install"
6. Review the warning about unverified plugins
7. Confirm installation
8. Wait for installation to complete
@@ -110,7 +120,7 @@ curl -X POST http://your-pi-ip:5000/api/v3/plugins/install-from-url \
```python
from src.plugin_system.store_manager import PluginStoreManager
store = PluginStoreManager()
store = PluginStoreManager(plugins_dir="plugin-repos")
result = store.install_from_url('https://github.com/user/ledmatrix-my-plugin')
[Plugin: news] Scroll frame stats - 100.0 fps over 501 frames | median 10.00ms
p95 10.11ms max 12.03ms min 7.98ms | stalls 0 (0.0%) skips 0 (0.0%)
```
Reading it, on a 100 Hz panel:
A healthy median is the refresh period times the scroll's frame hold: 10 ms
for a hold of 1 (100 px/s), **20 ms for 50 px/s** (hold 2), 30 ms for 33.3 px/s.
A 20 ms median on a 50 px/s scroll is the hold doing its job, not missed
refreshes. The `Scroll configured:` log line gives the hold (`1px every 2
refreshes`).
| you see | it means |
|---|---|
| median = refresh period × hold, p95 within ~0.5 ms of it | healthy — locked to the panel |
| p95 or max a whole refresh period or more above that median | frames missing refreshes — per-frame work is overrunning, or a background thread is holding the GIL |
| non-zero **skips**, or a median *below* the expected one | **duplicate frames** — the swap was skipped because the image did not change, so the frame never waited on vsync. The scroller is advancing less than one pixel per frame, which a crisp fixed-step scroll never does; look for a plugin pacing off time or not passing the hold. |
| non-zero **stalls** | frames past 1.5× the median, which is the measure of judder that survives averaging |
`stalls` and `skips` are both counted against that window's own median, so they
stay meaningful on a panel running at any refresh rate.
To rank every scroller at once rather than reading lines one at a time:
If a plugin logs its scroll config **twice** with different modes, the second
line is what is running.
## Soaking a rig
The per-scroller lines above tell you *which* scroller misbehaves. The soak
answers the question a release has to answer for each rig: **over a long run,
how often did a moving frame reach the panel late?**
Every frame reaches the panel through `DisplayManager.update_display`, so it is
timed there once, whoever drew it -- Vegas, a ticker plugin, anything. The
render thread only appends a tuple; a worker thread aggregates and rewrites
`/dev/shm/ledmatrix_frame_stats.json` every 10 seconds (RAM, so no SD-card
wear). `src/common/frame_timing.py` has the details.
```bash
python3 scripts/frame_soak.py # 10 minutes, as the display is now
python3 scripts/frame_soak.py --preview # with the web preview open
python3 scripts/frame_soak.py --show # totals since the service started
python3 scripts/frame_soak.py --json a.json # keep the report to compare later
```
It runs as any user next to the display service and stops nothing. It needs
something to *scroll* during the run: a live game holding a static scoreboard
on screen gives no verdict. `--preview` keeps the web preview's viewer marker
fresh, which puts the preview's PNG encoding at full rate -- run it as the web
service's user.
| line | what it tells you |
|---|---|
| **Late frames** | Frames presented one or more refreshes after they were due: the panel showed the previous frame again, a visible hitch. **The pass/fail number**, 0.1% by default (`--max-late-pct`). Only intervals between two scrolling frames count, and a frame held for `frame_hold` refreshes is due `frame_hold` refreshes after the last. |
| **Freezes** | Gaps of 250 ms or more inside a scroll: recomposes, plugin handovers, blocking calls on the render thread. Reported but not failed on, because some are handovers between plugins rather than faults. A gap still counts when the display's scroll state went missing for one frame across it, as long as scrolling resumes within 1 s: both of that frame's intervals count. Two static frames in a row end the scroll. (The state expires after 2 s without scroll activity, and plugins can clear it from their own `display()`.) The late and early rates are over frames judged against a known refresh period, which the recorder adopts once two windows in a row agree on it. |
| **blit** | Copying the frame into the matrix canvas (`SetImage`). It grows with width × height ×`pwm_bits`: ~5.5 ms at 512×64 with 8 bits on a Pi 4. It is the biggest fixed cost, and it sets the refresh rates a rig can hold one pixel per refresh at. |
| **wait** | Time blocked in `SwapOnVSync`, i.e. the slack left in each refresh. A p50 near zero means the rig has no headroom and anything extra lands a frame late. |
| **work** | Everything else between two frames: drawing, scrolling, and waiting for the GIL. A wide gap between its p50 and p99 is another thread getting in the way. |
| **Binding** | `STOCK` means the rgbmatrix binding holds the GIL through the vsync wait, which starves every other thread. See *Rebuilding the binding*. |
The refresh rate is estimated from the frames themselves (swaps that block on
vsync can only land on refresh boundaries). Cross-check it with
`scroll_speeds.py --measure` if it looks wrong. It can read high on a rig where
nothing ever presented at the full refresh rate.
A soak is only meaningful against a fixed workload. Compare runs with the same
content and `--preview` setting, and alternate which build goes first when you
A/B two of them. A live-API workload drifts over time.
The soak says how often; the service's log says why. A scroll that presents no
frame for 250 ms logs `Render stall:` with the stack of the render thread and
the top of every other thread's, and whether the whole interpreter was blocked
(C code holding the GIL) rather than one thread. To see what is behind the
shorter hitches, run the service with `LEDMATRIX_STALL_WATCHDOG_MS=30`, which
dumps at three refreshes late instead: its extra polling costs a little GIL
time of its own, so do that on a diagnostic run, not a soak you are grading.
`LEDMATRIX_STALL_WATCHDOG=0` turns it off.
### Results: hdpi, 2026-09-24
Pi 4, 4×128×64 on one chain (512×64), `gpio_slowdown` 3, cap 120 Hz, the
GIL-releasing binding. Vegas mode with live content, 8-minute soaks with
`--preview`, run in the order shown so each build went both first and last.
It never starts or stops the service itself, so a crash in it cannot leave
the panel dark. It grades with the same recorder as the soak and prints the
same report, with the same exit status, except that **2** also means the run
could not be set up at all (no root, no panel, a fallback display), so a rig
that was never measured cannot pass by accident.
Two differences from the soak matter:
- **It measures the panel first.** Before scrolling it times bare swaps for a
few seconds to get the idle refresh rate, and seeds the recorder with it.
That is what catches a loop that never locked to the panel at all. The first
version of the bench announced its scrolling state once instead of every
frame; the state expired, the dirty-tracking skip fired mid-scroll, and the
loop free-ran at 827 fps. Graded against its own frames that looks perfectly
steady; graded against the panel's measured rate every frame is early, and
the run fails as NOT LOCKED. (The soak has no idle measurement, so it checks
the rate against `limit_refresh_rate_hz` instead: a "refresh" faster than
the cap cannot have been waiting for the panel.)
- **The stall watchdog prints to the terminal.** A frame held up for more than
250 ms prints the stack of what held it up, in the middle of the run.
Measured with the first version of the bench on hdpi (Pi 4, 512x64,
`pwm_bits` 8), two-minute runs at one pixel per refresh: 8 of 11,449 frames
late (0.070%), and with `--busy 2` 3 of 11,445 (0.026%). The render path and
the hardware pass on their own. Compare the soak results above, from the same
rig with the service running, for how much of the late rate comes from
everything else.
### The panel is slower while you are rendering into it
The bench prints two refresh rates, and they differ:
| | Pi 4, 512x64, `pwm_bits` 8 |
|---|---|
| idle, timing bare swaps | 100.4 Hz |
| while scrolling | 96.3 Hz |
Both are real. Driving an LED matrix is bit-banging on the same machine, so
`SetImage` over a 512x64 chain contends with the refresh itself and slows it.
The recorder therefore reads the rendering rate back from the frames: swaps
that block on vsync can only return on a refresh boundary, so the low end of
`interval / frame_hold` is the period. The idle figure is still printed,
because the gap between the two is itself a measure of how expensive a frame
is: **a rise in that gap is a render-cost regression even when nothing is
late.**
The practical consequence for config: set `limit_refresh_rate_hz` near the rate
the panel holds *while rendering*, not the idle rate and certainly not a cap it
can never reach. A cap well above the real rate makes `scroll_config` solve
speeds against a refresh that does not exist, which is where "3px every 4
refreshes" comes from.
### Bench-only counters
| line | meaning |
|---|---|
| `duplicate` | frames that advanced no pixels. A crisp fixed-step scroll should show none; any at all means the loop is presenting faster than the strip is moving. |
| `blank` | frames with no visible slice to draw: the helper had no content. Should be zero. |
| `restarts` | how many times the strip was scrolled through end to end. Informational: the bench restarts the strip where a plugin would hand over to the next one. |
`--json` writes the full report plus the panel geometry, the solved speed and
these counters, so two rigs (or one rig before and after a change) can be
compared without re-reading a terminal.
---
## A tear across the middle on fast scrolls
**Symptom:** while text scrolls, the top and bottom halves of the panel look
shifted sideways against each other along a horizontal line at mid-height, and
the shift grows with scroll speed. It shows most in Vegas mode at high speed.
**It is the panel's scan, not the software.** The measured panel, like most
64-row panels, is multiplexed 1:32 (some panels of the same size scan
differently, so check yours): it lights two rows at a time, one from each half
(row 0 with row 32, row 1 with row 33, …), stepping down both halves together
once per refresh. So row 31,
the last row of the top half, lights almost a whole refresh period after row 32
right below it. Your eye follows moving text, and moving content that lights at
different times lands in different places, so the two rows meet with an offset
of roughly
```
offset ≈ scroll speed × refresh period
```
Each frame reaches the panel whole (`SwapOnVSync` swaps complete frames between
refreshes); the shift is created inside a single refresh. Other panel heights
show it too, at the point where their two scan halves meet.
On the 2×128×64 chain above, which refreshes at about 130 Hz flat out
(7.7 ms per pass):
| scroll speed | offset at the midline |
|---|---|
| 50 px/s (Vegas default) | ~0.4 px |
| 100 px/s | ~0.8 px |
| 150 px/s | ~1.2 px, plainly visible |
### What the display does about it
At one pixel per refresh, the fastest crisp speed, the step is exactly one
refresh's worth of motion, so it can be cancelled: show one half of the panel
a refresh behind the other -- the half whose row at the seam lights at the
start of each refresh. The two rows either side of the seam then show the same
moment again. What is left is a
lean of one pixel per half from top to bottom, continuous across the panel,
which reads as nothing where the step read as a tear. `DisplayManager` does
this while something scrolls at one frame per refresh
(`display.scan_order_compensation`, `"auto"` by default, `"off"` to disable;
the geometry is in `src/scan_order.py`). The lagging rows come from the
previous frame the display presented, so it works for Vegas and every plugin
ticker without knowing how they scroll.
Checked on hdpi (4×128×64 on one chain, rotated 180, 2026-09-24) before it was
written: `scan_mode: 1` (interlaced) made the step vanish but turned moving
edges grainy, and halving the speed halved it, so it is the scan and not a torn
frame. With the compensation the step is gone at 90 px/s.
It is left off where the row order is unknown or the maths does not hold:
- **Slower speeds**, where each frame is held for two or more refreshes. The
offset there is half a pixel or less, and cancelling it would need a lag of
a fraction of a frame.
- **Other layouts:** pixel mappers other than a 0 or 180 degree rotation
(U-mapper, 90/270), non-zero `multiplexing`, interlaced `scan_mode`, and a
canvas remapped to another height (double-sided mode).
- **The emulator,** which has no scan order.
### When it cannot apply
Only a shorter scan period (a faster refresh) or a slower scroll. Measure what
the panel actually achieves first. The library prints the rate with a carriage
return and no newline, so read it from the raw journal:
```bash
# set display.hardware.show_refresh_rate to true (web UI, Display tab), restart, then:
@@ -7,9 +7,9 @@ becoming nine clients of a god class.
Nine plugins (`afl`, `baseball`, `basketball`, `football`, `hockey`, `lacrosse`,
`nrl`, `soccer`, `ufc`) each ship a ~3,000-line `sports.py` descended from this
repo's `src/base_classes/sports.py`. They have drifted into three lineages, and
only 28 of the 66 methods appearing across them are present in all nine. One
logical fix (the UTC start-time bug) cost 75 files.
repo's former `src/base_classes/sports.py` (since removed). They have drifted
into three lineages, and only 28 of the 66 methods appearing across them are
present in all nine. One logical fix (the UTC start-time bug) cost 75 files.
Merging everything into one base class would fix the duplication and create a
worse problem: a single 2,500-line class that all nine plugins inherit, where any
@@ -26,8 +26,8 @@ These are independent concerns. Conflating them is what produces god classes.
|---|---|
| Plugin loads on a core that predates a module | Guarded import with a bundled fallback (`try: from src.X import Y / except ModuleNotFoundError: from y import Y`) |
| Plugin loads on a core that predates a *method* | Capability probing — `hasattr(SportsCore, "_detect_stale_games")` — never a version comparison. The loader's compat check is advisory-only (it logs and continues), so probing is the real protection. |
| Core changes never break a plugin's rendering | The **view-model contract**: `_extract_game_details_common`returns a dict whose `GUARANTEED_KEYS`are frozen by `test/test_skin_system.py::TestViewModelContract`. Keys may be added, never renamed or removed. |
| A plugin can drop its bundled copy safely | The **sunset rule**: its manifest must floor `ledmatrix_min_version` at the first core release shipping the module (recorded in `CHANGELOG.md`) — *necessary but not sufficient*. Nothing enforces that floor today, so the copy also waits for the B6 gate below. |
| Core changes never break a plugin's rendering | The **view-model contract**: the game dict each plugin's`_extract_game_details_common`builds is read by the shared `src/common` renderers, so its keys may be added, never renamed or removed. |
| A plugin can drop its bundled copy safely | The **sunset rule**: its manifest must floor `ledmatrix_min_version` at the first core release shipping the module (recorded in `CHANGELOG.md`) — *necessary but not sufficient*. The store enforces that floor on every registry-managed install and on both supported update paths (sideloading via `install_from_url` is not gated), but a floor cannot reach a user who never updates, so the copy also waits for the B6 gate below. |
The core API is **additive-only**. A method the plugins call is never removed or
given a new required parameter; new behavior arrives as new methods with
@@ -67,23 +67,43 @@ This is the property the naive merge destroys, and it is enforced structurally:
## Layering
```
src/base_classes/sports/
__init__.py re-exports the public API (import path unchanged)
| `score_phrase(points, team_abbr)` | Celebration wording (`"GOOOOAAALLL!"` vs `"TOUCHDOWN!"`). `points` is the score delta, which sports with variable-value scores use to name the play | `"<abbr> SCORES!"` — only consulted when `CelebrationMixin` is present |
> no longer reads it: the crisp-speed ladder uses the panel refresh
> (`display_manager.refresh_hz`), and speed comes from
> `scroll_settings.scroll_speed` alone. See `docs/SCROLL_PERFORMANCE.md`.
## Phases
B0–B3 are merged and shipping in core 3.2.0. Everything that remains is
**rollout**, and it splits into three phases with very different risk profiles.
The original plan folded the last two together; they are separated here because
one of them is safe by construction and the other is not.
one of them cannot break a user on an old core and the other can.
| Phase | Scope | Status | Gate |
|---|---|---|---|
@@ -212,9 +248,9 @@ one of them is safe by construction and the other is not.
| **B1** | Promote the nine universal methods; convert `sports.py` → package | ✅ | Characterization suite green; no behavior change intended |
| **B2** | `CelebrationMixin` + rotation strategies as opt-in capabilities | ✅ | Non-adopters have zero new code in their MRO; strategies checked against verbatim plugin transcriptions |
| **B3** | Upstream the scroll **orchestration** layer as `src/common/sports_scroll.py`, reading `global_config['target_fps']` natively | ✅ | Content building stays per-sport |
| **B4** | Ship 3.2.0 *and* make version reporting trustworthy | ⏳ **next** | Tag, release, and `src.__version__` agree; compatibility gate merged |
| **B5** | Adoption — guarded core imports: three pilots, then the remaining six. **Bundled copies stay.** | after B4 | Per plugin: harness + goldens byte-identical, then a device soak |
| **B6** | Sunset — delete the bundled copies | **blocked** | B4's gate shipped *and* in users' hands (see below) |
| **B4** | Ship 3.2.0 *and* make version reporting trustworthy | ✅ | Released 2026-08-03; tag, release and `src.__version__` agree; compatibility gate merged (#428, #431, #433) |
| **B5** | Adoption — guarded core imports, all eight. **Bundled copies stay.** | ✅ | All eight adopted; harness byte-identical; see the B5 retrospective below — four shipped broken and were repaired in plugins #251 |
| **B6** | Sunset — delete the bundled copies | ✅ | Ran 2026-09-01, all eight. Floors at 3.2.0; the store refuses on all three routes in (#431/#433, #508, #510). See "B6 — what actually happened" |
### B4 — what "ship 3.2.0" actually requires
@@ -234,6 +270,20 @@ a floor can be trusted against, and today it is not:
compares the plugin's manifest version against the registry's
`latest_version` and nothing else.
*Fixed, in two parts.*`install_plugin` gained the gate in #431/#433, which
covers every registry-managed install and, through `_reinstall_with_rollback`,
the update path that re-downloads.
`update_plugin`'s git branch pulls in place and re-downloads nothing, so it
stayed ungated until `_gate_pulled_commit` closed it — checked after the pull
(the registry carries no floor field, so the incoming floor is unknowable
before it) and undone with `git reset --hard` to the pre-pull commit. That
route is rare in practice, since monorepo plugins install as archives; it was
closed because the sunset rule in the plugins repo's
`08-shared-sports-code.md` states as **condition 3** that the core enforces
the floor "at install/update time", and B6 rests on that being true rather
than merely written down. `install_from_url` — sideloading a plugin from a
URL — is still ungated.
So B4 is: tag and release 3.2.0; make the tag, the release, and `__version__`
agree, and keep them agreeing; reconsider the `< 2.0.0` skip; migrate manifests
from `ledmatrix_min` to `ledmatrix_min_version`; and add the install/update
@@ -262,13 +312,21 @@ default is the more restrictive.
`compatible_versions`. No manifest still carries it, so there is nothing to
migrate there.)
### B5 — adoption is safe by construction
### B5 — the *fallback* is safe by construction; the modern path is not
The heading matters, because the unqualified version of this claim is false and
this document proves it two sections down: four of the eight adopted plugins
shipped with scroll mode broken on a 3.2.0 core. What is safe by construction is
narrower than "adoption".
A plugin adopting core imports keeps its bundled copy and reaches it through the
guarded import (see the Upgradability table above). On a core that ships the
module the plugin uses core code; on one that doesn't it falls back and behaves
exactly as it does today. There is no version of this step that breaks a user,
which is why it does not wait for B6's gate.
guarded import (see the Upgradability table above). On a core that doesn't ship
the module the plugin falls back and behaves exactly as it does today. That
fallback compatibility — and only that — is safe by construction. On a core that
*does* ship the module, correctness is not automatic — object-level and scroll-mode validation
(building both classes and comparing, per the retrospective below) is required
to prove full behavior. There is no version of this step that breaks a user *on
an old core*, which is why it does not wait for B6's gate.
The hockey scroll-display pilot is **already validated**: adopted against a core
carrying 3.2.0, `scroll_display.py` went from 691 to 289 lines and all 16 harness
@@ -297,7 +355,8 @@ been in users' hands long enough that the population running a core without it
is small.** The bundled copies cost disk space; deleting them early costs
scoreboards, silently. That trade is not close.
Before the first sunset, add a **compatibility regression test**. It has to
Before the first sunset, add a **compatibility regression test**. **Built:**
core `test/test_sports_sunset_matrix.py` (#505). It has to
cover four cases, not one — B5's safety claim and B6's failure mode are
different propositions and only the second is obvious:
@@ -321,35 +380,145 @@ gate rather than trusting the failure to be noticed.
The same suite should exercise the install/update gate, since it is the other
half of the guarantee.
### B6 — what actually happened
**Ran 2026-09-01, across all eight scoreboards.** Held from 2026-08-05 to
2026-09-01 on the argument below, which is kept because the reasoning applies to
the next module, not because it is still in force.
**The hold, and why it lifted.** The stated gate was evidence of 3.2.0 uptake —
"a few months of it being the default download, or store-side install data".
That evidence never arrived and could not: the core updates by
`git pull --rebase`, so release-asset counts cannot measure uptake, and no
store-side telemetry exists. What changed instead is that the *risk* the gate
protected against was closed directly. The store now refuses a plugin whose
floor exceeds the running core on **all three** routes in:
| route | gated by |
|---|---|
| `install_plugin` — every path that re-downloads, `_reinstall_with_rollback` included | #431, #433 |
@@ -8,7 +8,7 @@ After running `first_time_install.sh`, SSH may become unavailable for the follow
**Primary Cause**: The WiFi monitor service (`ledmatrix-wifi-monitor`) automatically enables Access Point (AP) mode when it detects that the Raspberry Pi is not connected to WiFi. When AP mode is active:
- The Pi creates its own WiFi network: **LEDMatrix-Setup** (password: `ledmatrix123`)
- The Pi creates its own WiFi network: **LEDMatrix-Setup** (open, no password)
- The Pi's WiFi interface (`wlan0`) switches from client mode to AP mode
- **This disconnects the Pi from your original WiFi network**
- SSH becomes unavailable because the Pi is no longer on your network
@@ -20,7 +20,22 @@ The installation script:
- Installs and configures `dnsmasq` (DHCP server for AP mode)
- These services can interfere with normal WiFi client mode
### 3. Reboot After Installation
### 3. The Board Ran Out of Memory
On a 512MB or 1GB board, memory exhaustion stops `sshd` being able to fork a
session process. The connection is accepted and then closed immediately, before
any banner:
```text
kex_exchange_identification: Connection closed by remote host
```
The giveaway is that the board is otherwise healthy — ping is clean and the web
UI still responds — but nothing that needs to start a new process works, and
the panel is usually dark. Only a power cycle clears it. See
[LOW_MEMORY_BOARDS.md](LOW_MEMORY_BOARDS.md).
### 4. Reboot After Installation
If the script reboots the Pi (which it recommends), network services may restart in a different state, potentially triggering AP mode.
@@ -30,7 +45,7 @@ If the script reboots the Pi (which it recommends), network services may restart
1. **Find the AP Network**:
- Look for a WiFi network named **LEDMatrix-Setup** on your phone/computer
- Default password: `ledmatrix123`
- It is an open network: no password
2. **Connect to the AP**:
- Connect your device to the **LEDMatrix-Setup** network
@@ -38,7 +53,7 @@ If the script reboots the Pi (which it recommends), network services may restart
By default, Access Point (AP) mode is **not automatically enabled** after installation. AP mode must be manually enabled through the web interface when needed.
## Default Behavior
- **Auto-enable AP mode**: `false` (disabled by default)
- AP mode will **not** automatically activate when WiFi or Ethernet disconnects
- AP mode can only be enabled manually through the web interface
## Why Manual Enable?
This prevents:
- AP mode from activating unexpectedly after installation
- Network conflicts when Ethernet is connected
- SSH becoming unavailable due to automatic AP mode activation
- Unnecessary AP mode activation on systems with stable network connections
## Enabling AP Mode
### Via Web Interface
1. Navigate to the **WiFi** tab in the web interface
2. Click the **"Enable AP Mode"** button
3. AP mode will activate if:
- WiFi is not connected AND
- Ethernet is not connected
### Via API
```bash
# Enable AP mode
curl -X POST http://localhost:5001/api/v3/wifi/ap/enable
# Disable AP mode
curl -X POST http://localhost:5001/api/v3/wifi/ap/disable
```
## Enabling Auto-Enable (Optional)
If you want AP mode to automatically enable when WiFi/Ethernet disconnect:
### Via Web Interface
1. Navigate to the **WiFi** tab
2. Look for the **"Auto-enable AP Mode"** toggle or setting
The Background Data Service is a new feature that implements background threading for season data fetching to prevent blocking the main display loop. This significantly improves responsiveness and user experience during data fetching operations.
## Key Benefits
- **Non-blocking**: Season data fetching no longer blocks the main display loop
- **Immediate Response**: Returns cached or partial data immediately while fetching complete data in background
- **Configurable**: Can be enabled/disabled per sport with customizable settings
- **Thread-safe**: Uses proper synchronization for concurrent access
- **Retry Logic**: Automatic retry with exponential backoff for failed requests
- **Progress Tracking**: Comprehensive logging and statistics
## Architecture
### Core Components
1. **BackgroundDataService**: Main service class managing background threads
**You don't need to worry about these errors.** They are harmless and don't affect functionality. We've improved error suppression to hide them from the console.
## Error Types
### 1. Permissions-Policy Header Warnings
**Examples:**
```text
Error with Permissions-Policy header: Unrecognized feature: 'browsing-topics'.
Error with Permissions-Policy header: Unrecognized feature: 'run-ad-auction'.
Error with Permissions-Policy header: Origin trial controlled feature not enabled: 'join-ad-interest-group'.
```
**What they are:**
- Browser warnings about experimental/advertising features in HTTP headers
- These features are not used by our application
- The browser is just informing you that it doesn't recognize these policy features
**Why they appear:**
- Some browsers or extensions set these headers
- They're informational warnings, not actual errors
- They don't affect functionality at all
**Status:** ✅ **Harmless** - Now suppressed in console
### 2. HTMX insertBefore Errors
**Example:**
```javascript
TypeError: Cannot read properties of null (reading 'insertBefore')
at At (htmx.org@1.9.10:1:22924)
```
**What they are:**
- HTMX library timing/race condition issues
- Occurs when HTMX tries to swap content but the target element is temporarily null
- Usually happens during rapid content updates or when elements are being removed/added
**Why they appear:**
- HTMX dynamically swaps HTML content
- Sometimes the target element is removed or not yet in the DOM when HTMX tries to insert
- This is a known issue with HTMX in certain scenarios
**Impact:**
- ✅ **No functional impact** - HTMX handles these gracefully
- ✅ **Content still loads correctly** - The swap just fails silently and retries
- ✅ **User experience unaffected** - Users don't see any issues
**Status:** ✅ **Harmless** - Now suppressed in console
## What We've Done
### Error Suppression Improvements
1. **Enhanced HTMX Error Suppression:**
- More comprehensive detection of HTMX-related errors
- Catches `insertBefore` errors from HTMX regardless of format
- Suppresses timing/race condition errors
2. **Permissions-Policy Warning Suppression:**
- Suppresses all Permissions-Policy header warnings
- Includes specific feature warnings (browsing-topics, run-ad-auction, etc.)
- Prevents console noise from harmless browser warnings
3. **HTMX Validation:**
- Added `htmx:beforeSwap` validation to prevent some errors
- Checks if target element exists before swapping
- Reduces but doesn't eliminate all timing issues
## When to Worry
You should only be concerned about errors if:
1. **Functionality is broken** - If buttons don't work, forms don't submit, or content doesn't load
2. **Errors are from your code** - Errors in `plugins.html`, `base.html`, or other application files
3. **Network errors** - Failed API calls or connection issues
✅ **HTMX errors are caught and handled gracefully**
✅ **Permissions-Policy warnings are hidden**
✅ **Application functionality is unaffected**
## Technical Details
### HTMX insertBefore Errors
**Root Cause:**
- HTMX uses `insertBefore` to swap content into the DOM
- Sometimes the parent node is null when HTMX tries to insert
- This happens due to:
- Race conditions during rapid updates
- Elements being removed before swap completes
- Dynamic content loading timing issues
**Why It's Safe:**
- HTMX has built-in error handling
- Failed swaps don't break the application
- Content still loads via other mechanisms
- No data loss or corruption
### Permissions-Policy Warnings
**Root Cause:**
- Modern browsers support Permissions-Policy HTTP headers
- Some features are experimental or not widely supported
- Browsers warn when they encounter unrecognized features
**Why It's Safe:**
- We don't use these features
- The warnings are informational only
- No security or functionality impact
## Monitoring
If you want to see actual errors (not suppressed ones), you can:
1. **Temporarily disable suppression:**
- Comment out the error suppression code in `base.html`
- Only do this for debugging
2. **Check browser DevTools:**
- Look for errors in the Network tab (actual failures)
- Check Console for non-HTMX errors
- Monitor user reports for functionality issues
## Conclusion
**These errors are completely harmless and can be safely ignored.** They're just noise in the console that doesn't affect the application's functionality. We've improved the error suppression to hide them so you can focus on actual issues if they arise.
This will show you the complete log including what happened after dependency installation.
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.