Compare commits

...
77 Commits
Author SHA1 Message Date
ChuckandClaude Opus 5.5 ea09c0aba5 Merge origin/main into claude/deprecate-unused-plugin-api
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:49:55 -04:00
ChuckandClaude Opus 5.5 e1ce7189f1 fix(install): make one-shot retry() retry, and drop root grants on user files (#606)
retry() in one-shot-install.sh used `if ! "$@"; then status=$?`, where $? is
the status of the negation -- always 0. A failed command was never retried
and retry() reported success, so a failed `git clone` carried on until a
later check noticed the missing checkout. It now retries (3 attempts) and
returns the command's status. The two apt steps stay non-fatal: warning and
continuing is what they effectively did before, and making them fatal would
stop installs that work today. A clone that keeps failing stops the install,
as it already did, just sooner and with the one-shot's own error message.

Both installers granted the web user NOPASSWD root on display_controller.py,
start_display.sh and stop_display.sh. Those files are owned by the user after
Step 11's chown, so the grant let the web user rewrite them and run them as
root, and nothing ever ran them through sudo. Removed from both installers,
with a test that every project file granted as root is a root-owned
fix_perms helper.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:34:20 -04:00
ChuckandClaude Opus 5.5 342e9164b8 fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound

Each repeat of an error pattern appended every plugin in the time window to
the pattern's list again, so a plugin failing in a loop grew the display
process's memory without limit: 3,000 errors from three plugins reached 2.5
million entries. Keep the list unique.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): load a BDF font at its native size instead of PIL's default

FreeType rejects any size but a BDF strike's own, and FontManager answered
that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf
requested at 8 or 10px rendered as PIL's default font. Retry at the native
strike, as element_style already does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): plugin toggle failures no longer claim "operation in progress"

Every exception in POST /plugins/toggle was mapped to
PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation
is already in progress". Report the failure as what it is, and record the
plugin id in the operation history for form posts too (it read a `data`
variable that only the JSON path set).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): route plugin card clicks through handlePluginAction

The document-level delegation checked `typeof handlePluginAction`, which is
scoped inside the plugin-manager IIFE and so never visible to it. Every card
click took a copied fallback that stopped propagation (the grid's own
listener never ran), confirmed an uninstall twice, and sent Starlark app
uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>.
Expose the handler on window and delegate to it.

Also run every test/js/unit suite under pytest: they need only node, but CI
ran one of the eight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): apply Rotation durations, WiFi messages and Vegas settings

Three settings the web UI saves never reached the display:

- Rotation & Durations: display.display_durations was never read. Every
  plugin inherits get_display_duration() and the plugin was asked first. A
  saved value now wins. The page shows unsaved screens blank with the
  plugin's own duration as a placeholder, and saving a blank removes the
  override, so one save no longer pins every screen.
- WiFi status overlay: the controller looked for wifi_status.json one
  directory above the repo. Both sides now use
  wifi_manager.get_wifi_status_path(). The message is written by rename so
  the display never reads it half-written, and the resumed plugin redraws the
  whole panel afterwards.
- Vegas: nothing called coordinator.update_config(), so saved Vegas settings
  never reached a running scroll. They are now queued when
  display.vegas_scroll changes, and applied while Vegas is stopped too, so a
  disable then re-enable works. The follower's scroll-speed default (75) now
  matches VegasModeConfig's (50).

Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us
per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows
with each plugin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: keep affected_plugins order when serialized; guard non-Element targets

ErrorPattern.to_dict() ran the now-ordered list through set(), so
get_error_summary() listed plugins in an unstable order. The document-level
card-action listener called event.target.closest() without checking the
target is an Element.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:29:16 -04:00
claude[bot]andClaude Opus 5 0903f9055f docs(changelog): record #600 and #602 in the 3.5.0 section (#603)
* docs(changelog): record #600 in the 3.5.0 section

#600 merged into main while the release PR was open, so the 3.5.0 section
went in without it. Nothing in that PR touched the CHANGELOG, and no check
covers "everything merged since the last tag is written down", so tagging
v3.5.0 as main stands would ship the standings-endpoint fix undocumented.

The entry goes under Sports data, next to the other ESPN fetch changes, and
is written from the commit: what the old order did, why a college league's
200 defeated the 404 fallback, and what is now treated as routine.

No version change: 3.5.0 is not tagged yet, so this belongs in that section
rather than a new one. `scripts/check_release_version.py v3.5.0` still passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

* docs(changelog): record #602 in the 3.5.0 section

#602 merged into main after #601, the same way #600 merged during it, and
also touched no CHANGELOG. So the section was still a commit short of what
v3.5.0 will actually ship.

It gets its own "Installers" subsection rather than a line under "Small
fixes": a malformed drop-in in /etc/sudoers.d makes sudo refuse every command
for every user, which on a headless Pi is unrecoverable over SSH. That is not
a small fix, and someone reading the release notes to decide whether to update
should see it.

Written from the commit: what both installers did, what `visudo -c` now gates,
and the fixed /tmp path that mktemp replaced.

`scripts/check_release_version.py v3.5.0` still passes, and this branch is
rebased onto 967f3a05 so the section now covers every commit since v3.4.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 10:28:31 -04:00
ChuckandClaude Opus 5.5 47b56182e4 feat: deprecate unused plugin-facing methods for removal in 3.7.0
35 methods on CacheManager, DisplayManager, FontManager and PluginManager
have no caller in core, the ledmatrix-plugins monorepo or the registry's
third-party plugins, but plugins live elsewhere, so they stay for one
release. src.deprecation.deprecated logs a warning (and emits a
DeprecationWarning) the first time each is called in a process, naming the
release that removes it. The list and replacements are in CHANGELOG and
PLUGIN_API_REFERENCE's new Deprecated APIs section; a test pins the set.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-22 17:09:19 -04:00
claude[bot]andClaude 967f3a0567 fix(install): parse the sudoers rules before installing them (#602)
Both installers generated the ledmatrix_web rules and copied them straight
into /etc/sudoers.d without ever parsing them. Every rule is built from
`which` lookups, so an empty or surprising path produces a malformed
drop-in -- and a malformed file in /etc/sudoers.d makes sudo refuse every
command for every user. On a headless Pi that is unrecoverable over SSH.

first_time_install.sh now runs `visudo -c` on the generated file and, if it
does not parse, prints what visudo said and leaves the installed file
untouched rather than replacing it with a broken one. configure_web_sudo.sh
does the same before it offers the rules for confirmation.

first_time_install.sh also built the file at a fixed /tmp path as root;
mktemp now picks the name.

test/test_sudoers_is_validated.py renders the installer's own sudoers
heredoc and checks the result with visudo -- the check neither installer
had -- and asserts the install stays gated on it.


Claude-Session: https://claude.ai/code/session_01Dby94z9PV3zVM25fqGNXTt

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-21 16:08:29 -04:00
claude[bot]andClaude Opus 5 21c8a54f68 chore: prepare the 3.5.0 release (#601)
* chore: prepare the 3.5.0 release

Turns the CHANGELOG's Unreleased section into `## 3.5.0` and bumps
`src.__version__`, the value plugin `ledmatrix_min_version` floors compare
against. No behaviour change; nothing outside the CHANGELOG, `src/__init__.py`
and one docs line is touched.

The staged entries are reshaped into the `### ` subsections every released
section already uses, and the "new modules a plugin may import via `src.*`"
block moves to the top as the plugin-facing summary, the same shape as 3.4.0.
Its floor, written as "the release that ships this" while it was staged, is now
3.5.0, and `docs/SPORTS_UNIFICATION.md` says 3.5.0 for `sports_helpers.py`
instead of "(unreleased)".

Four merged changes had never been written down. They are added under the
subsection each belongs to, from the commits and their measurements:

- the idle back-off clamped to the next kickoff (#599)
- concurrent ESPN date chunks (#596)
- the three web routes that consulted plugin manifests before anything had
  discovered plugins, one of which wrote a plugin API key to config.json in
  plain text (#594)
- the cache permission fix and its systemd unit changes (#593), which get
  their own subsection

No tag and no release: `scripts/check_release_version.py v3.5.0` passes, so
tagging is a separate, deliberate step.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

* ci: let Claude Code Review run on PRs the Claude app opens

The review action refuses a workflow whose actor is a GitHub App unless the
app is named in `allowed_bots`, which this workflow never set:

  Actor is a GitHub App: claude[bot]
  Actor type: Bot
  Action failed with error: Workflow initiated by non-human actor: claude
  (type: Bot). Add bot to allowed_bots list or use '*' to allow all bots.

It aborts about two seconds in, before the diff is read, so the check is red
on every such PR and re-running cannot help: the actor does not change. Until
now no PR here had a bot author, so nothing tripped it.

`'claude'` rather than `'*'`: the action lowercases each entry and strips a
trailing `[bot]` before comparing it to the actor
(`isAllowedBot` in `src/github/validation/actor.ts`), so this admits
`claude[bot]` and no other app. `'*'` would admit any app that can trigger a
workflow here, with a prompt it controls — the action's own docs warn about
that on public repositories, and this one is public.

The write-permission check already allowed the app; `checkHumanActor` was the
only gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-21 16:08:05 -04:00
ChuckandClaude Opus 5 cf02538d2e fix(sports): ask the endpoint the league actually publishes for standings (#600)
* fix(sports): ask the endpoint the league actually publishes for standings

ESPNDataSource.fetch_standings tried /standings first regardless of league
and fell back to /rankings only on a 404. College leagues answer /standings
with a 200 that carries no poll, so the fallback never fired and the poll
came back empty every time. Nothing failed; the rank badge simply never
appeared, and anything keyed off rankings quietly did nothing.

Endpoints are now ordered by whether the league publishes a poll, a 200
that lacks the key counts as a miss so a league answering both still ends
up with whichever one carries the poll, and only a 404 is treated as
routine -- it is how a league says it has none. A connection error, a
timeout or an unparseable body is logged as an error again.

This is the implementation the football, baseball and hockey boards already
ship; core was the last copy still on the old one. Verified against live
ESPN: mens-college-basketball returns a populated rankings key where it
previously returned nothing, and nba still resolves from /standings alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(standings): stop the endpoint handler from swallowing its own bugs

Addresses both CodeRabbit findings on #600.

The handler caught `Exception`, so an AttributeError or TypeError raised
while *inspecting* the payload was indistinguishable from an endpoint that
failed. The loop would move on and, if the other endpoint had nothing
either, return {} -- silently dropping rankings for a league that has them.
That is the precise failure this function was written to fix, so the
handler was able to reintroduce it.

Only the request is guarded now. `requests.RequestException` covers the
transport failures and `ValueError` covers a body that will not parse;
payload inspection happens after the handler, where a bug surfaces instead
of being logged as a missing poll. A non-dict payload is treated as a miss
explicitly rather than by tripping over `.get`.

Tests: the fallback paths had no coverage -- the old single-endpoint code
would have passed the suite unchanged. Added order assertions for both
league kinds, a 200-without-a-poll fall-through, 404 and non-404 recovery,
a non-object payload, and a guard proving a bug is no longer swallowed.

`test_fetch_standings_returns_empty_on_error` faked a transport failure
with a bare `Exception`, which only passed because the handler caught
everything. It now raises ConnectionError, which is what actually happens.

Verified by mutation: restoring standings-first fails 5 tests, restoring
the catch-all fails the bug-not-swallowed guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 16:07:51 -04:00
ChuckandClaude Opus 5 81e1bc596f fix(sports): stop the idle back-off sleeping through a kickoff (#599)
* fix(sports): stop the idle back-off sleeping through a kickoff

A league with no live games backs its poll off as empty checks mount,
capped by live_idle_max_interval. The escalation counts empty looks and
nothing else, so a league three hours before kickoff is indistinguishable
from one three months out of season. Both reach the ceiling -- and the
ceiling then *is* the blind spot.

Measured on two rigs on 2026-09-19: gaps of up to 928s between looks, ten
of them at or above 900s. Reproduced in the wild on 2026-09-20, where an
unpatched rig sat for fifteen minutes with eight NFL games in progress and
had not noticed any of them. That is the "it doesn't pick up new live
games until I restart it" report -- restarting being the one thing that
forces an immediate look.

The clamp costs no extra request: the live fetch already downloads the
whole day's scoreboard, upcoming games included, so the earliest start
still ahead of us falls out of the payload the manager already has.
Before a kickoff the wait is shortened so it cannot run past it; just
after one, the live cadence is held for _KICKOFF_GRACE_SECONDS, because a
provider that has not yet flipped the status would otherwise look like
another empty check and escalate the back-off again, right when the game
is starting.

The grace window needed a second pass. A soak caught it as dead code: the
just-passed kickoff was replaced by the next fixture on the card the
instant it passed, `now < start` went true again, and the back-off
returned to its ceiling. Observed live -- the rig polled at 13:00:45,
found nothing because ESPN had not flipped the status, then went quiet for
a quarter of an hour. A kickoff inside the grace window is now kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(sports): pin absolute tolerances and correct a wrong grace expectation

pytest.approx defaults to a relative tolerance. On a unix timestamp that is
roughly 1790 seconds, so every kickoff assertion here was effectively
vacuous -- it called a kickoff half an hour away "equal". All seven now
pin abs=1.

That hid a wrong expectation. test_an_earlier_kickoff_still_wins_during_the_grace
asserted a game ten minutes out should displace one that kicked off moments
ago. It should not, and the code does not: while the grace holds, the wait
is the live cadence (30s), which is strictly tighter than clamping to the
nearer kickoff would give (~600s). Letting the candidate win would set a
ten-minute wait at the exact moment games are starting -- the dead grace
window this branch exists to fix.

The test now pins the real behaviour plus the safety property that makes it
correct, and is renamed to say what it checks.

Reported by CodeRabbit on the PR. The finding was right that code and test
disagreed; the suggested fix was the wrong way to resolve it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 18:25:33 -04:00
ChuckandClaude Opus 5 19686ab761 fix(web): plugin settings form shows schema defaults for unsaved keys (#597)
The server-rendered plugin settings partial rendered straight from the
saved config, so an option added in a plugin update (geochron 1.2.0's
show_date / show_date_line, default true) drew as an unchecked box, and
the save route's missing-checkbox handling then stored it as false.
Enum dropdowns likewise showed their first option instead of the default.

- _load_plugin_config_partial runs the stored section through
  prepare_plugin_config (as GET /plugins/config does) before masking
  secrets, so a secret's schema default is masked too.
- render_field falls back to the field's own default, covering children
  of objects that declare a default of their own (where the defaults
  extraction stops).
- The legacy-boolean parity test now compares against the config the
  plugin actually runs with (defaults included).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 11:09:25 -04:00
ChuckandClaude Opus 5 92f1960d00 perf(sports): fetch ESPN date chunks concurrently (#596)
* perf(sports): fetch ESPN date chunks concurrently

Since ESPN started rejecting `dates=YYYYMMDD-YYYYMMDD` on 2026-09-15, one
season request became a chunk per month -- and a month over the 500-event
cap becomes a request per day. A cold college-baseball season is about 130
requests, and they went out one at a time.

That is slower than the 20s budget `_update_plugins()` shares across every
plugin at startup, so scoreboards were logging `update() timed out` on
first run and being deferred to the scheduled tick with nothing on the
panel. Measured on a Pi 4 against live ESPN, March+April college baseball
(63 requests, 3101 events): 11.2s sequential, 1.6s concurrent. Over a whole
boot that moved football-scoreboard, ledmatrix-flights and birdnet-go
inside the budget -- 13 plugins deferred before, 10 after.

Chunks now go out six at a time, in two passes: months and edge days first,
then the days of any month that came back capped. Six keeps the shared
Session under requests' default pool_maxsize of 10, so no connection is
discarded. Merged events still follow `espn_date_chunks` order -- a capped
month's days are spliced back into its own slot -- so the payload does not
depend on which request won the race.

Request order is no longer significant, so the three tests that pinned it
compare the chunks as a set and keep asserting the merged event order,
which is the part callers actually see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): drop capped month payloads before fetching their days

Review of the concurrent chunk fetch found it raised the worst-case peak
memory more than the concurrency explains. The old loop discarded a month
that came back at the 500-event cap the moment it saw it; the rewrite kept
every capped month alive in `results`/`slots` until all of their day
requests had finished.

Measured on a Pi 4 fetching 20260201-20260531 college baseball (four capped
months, 5462 events), peak RSS growth over the call:

  sequential (main)               83 MB
  concurrent, months retained    121 MB  (+43)
  concurrent, one worker         108 MB  -- the retention alone was +25
  concurrent, months dropped      98-100 MB (+16)

docs/LOW_MEMORY_BOARDS.md puts a 1 GB Pi 3B+ at under 200 MB of headroom,
where running out makes the board unreachable until a power cycle, so the
difference matters. The remaining +16 MB is six responses parsing at once;
three workers saved about 6 MB more, within run-to-run noise, so the worker
count stays at six.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(sports): state what ESPN_CHUNK_WORKERS was measured to do, not more

The comment claimed the sequential fetch made scoreboards blow the 20s
startup update() timeout. A boot on this branch still deferred 12 plugins
and timed out baseball-scoreboard while its season fetches took 0.74s and
1.12s: the startup budget is spent on other per-plugin work. Say what was
measured -- 17.7s sequential, 2.6-3.3s concurrent -- and nothing else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 09:56:50 -04:00
ChuckandClaude Opus 5 116abb0daa fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service

BackgroundDataService always sent a season range first and, on a 400,
fell back to chunks without recording the rejection, so every background
season fetch spent a doomed request and live scoreboards learned nothing
from it (or it from them). The worker now consults and sets the same
6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight
to month/day chunks, and if every chunk fails the range is asked once for
a real error without re-spending the chunks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep plugin asset and action routes inside their directories

POST /plugins/assets/upload, GET /plugins/assets/list and POST
/plugins/assets/delete joined the request's plugin_id onto assets/plugins
unchecked, so '../../config' created, wrote, listed and deleted outside
it. #561 guarded only the route that serves the files. All three now go
through path_safety.resolve_under and answer 400 for anything but a
plain name, and delete only unlinks a metadata path that resolves into
that plugin's uploads directory.

PluginManager.get_plugin_directory refuses ids that are not one plain
path segment, so /plugins/action (which runs a manifest script from the
returned directory) and every other caller get the guard; the action
route also rejects such ids up front, covering its no-manager fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): report a no-op plugin update as already up to date

update_plugin() returns True both for a real update and for "nothing to
do" (a ZIP-installed monorepo plugin already at the registry version, a
bundled plugin). With no git commit to compare, POST /plugins/update
called every such success "updated successfully", so Check & Update All
counted most official plugins as updated on every run.

The route now reads what changed off the plugin itself (commit, else
manifest version, else last_updated) and returns data.update_status
(updated / up_to_date / local_only). The update-all toast is summarised
by PluginInstallManager.summarizeUpdateResults from that status, falling
back to the message for older servers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): scoreboard scroll speed no longer follows target_fps

sports_scroll computed the crisp speed ladder against the global
target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since
frame-locked presentation (#545) the helper steps a fixed number of whole
pixels per presented frame and the panel presents at its real refresh, so
the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a
50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s.

The ladder now uses the display manager's refresh_hz, then
display.hardware.limit_refresh_rate_hz, then the default. target_fps is
not consulted. Docstrings now say scroll_delay is ignored for pacing (no
behaviour change there) and describe the fixed-step model.

Tests: replace the tests that pinned target_fps as the ladder refresh and
described time-based stepping; assert speed independence from target_fps
(unit and end-to-end presented px/s against the real helper), that the
fixed per-frame step is applied, and that scroll_delay does not change
speed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): escape registry and upload values in plugin manager inline handlers

The store, saved-repository and custom-registry buttons built
onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone,
so a custom registry entry whose id contained ' closed the attribute and
added its own handler. One helper, jsStringAttr(), now HTML-escapes the
JSON literal for every one of those handlers, and the store View button
opens only http(s) repo links.

The live window.updateImageList (plugins_manager.js loads last, so its
copy wins over the file-upload widget's) wrote the uploaded file's
original name, path and ids into markup raw; they are escaped now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note plugin asset, action and inline handler guards

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): let the root pip wrapper install web_interface/requirements.txt

Update Code, the automatic update's health check and Install Base
Requirements install web_interface/requirements.txt through
safe_pip_install.sh, which only allowed the root requirements.txt. The
first commit changing that file would fail its dependency install, and
the automatic updater rolls back any update whose dependencies did not
install -- on every device, for every newer commit.

The wrapper now lists both core requirement files. Only their folders
are resolved, so a requirements.txt symlinked out of the project is
compared by its target and refused (previously the root file's own
symlink target was what got allowed). The updater's file list is a
named constant, and a test runs the real wrapper (pip stubbed) on
every file Update Code and the rollback install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): do not retry plugin requests that got an HTTP answer

PluginAPI.request wrapped everything that was not a structured error as
NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a
JSON error without error_code included. Check & Update All retries
NETWORK_ERROR, so those updates were re-sent five more times with
backoff, contrary to the #587 contract that an HTTP error response is
the server's answer.

NETWORK_ERROR now means only that fetch() rejected. Any HTTP response
without an error_code, or with a body that is not JSON, is API_ERROR
with the HTTP status attached. Tested against the shipped api_client.js.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): restart the stats window when an idle gap is dropped by size

#582 dropped an idle gap from the frame stats two ways: the reset_scroll()
sentinel, which also restarts the 5s window timer, and a size guard for
scrollers that never call reset_scroll(), which did not. On that path the
first real frame after the gap found the boundary overdue and logged a
stats line for a one-frame window. Both paths now share one seeding helper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): leave plugins alone when update_core's own rollback fails

update_core returns rollback_failed directly when a partial pull or an
update whose health check never started cannot be rolled back. run()
only held plugins back for 'verifying', so those devices still got new
plugin versions and a display restart on top of a core in an unknown
state -- the opposite of what the health-check path does, and of the
3.4.0 changelog (plugins are left alone if the rollback fails).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(api): make the REST reference match the api_v3 package

Every documented request body, query parameter and response shape was
re-checked against the handlers in web_interface/blueprints/api_v3/.
Fixes calls that failed as documented (repo_url, action_id/params,
files/image_id, font_file+font_family, ?font=, cache key,
auto_enable_ap_mode, plugin limit keys), removes the font-override
endpoints dropped in #566, corrects response shapes (plugins/config,
plugins/schema, health, metrics, operation history, github-status,
fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds
the 26 routes it omitted (backup, system auto-update/git, wifi radio,
starlark editor, MQTT bridge, status endpoints, skins).

Documents the merge semantics of partial JSON saves to /config/main and
/plugins/config and the dim-schedule POST accepting GET's days shape,
which land in the same change set. Replaces app.py line numbers and the
removed api_v3.py path with file and function names.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): remove the General-tab plugin system toggles that did nothing

plugin_system.auto_discover, auto_load_enabled and development_mode had
General-tab toggles whose help tips promised dormant plugins and verbose
logging, but nothing reads them: every enabled plugin is discovered and
loaded regardless. Remove the three toggles.

The keys stay tolerated in stored configs. The save handler now stores
a flag only when a client sends it; treating a missing key as an
unchecked box would otherwise rewrite all three to false on every
General-tab save, which still posts plugins_directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(scroll): remove dead code left by #523/#570

- Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read
  them since the numpy blend replaced the scipy path.
- Drop ScrollHelper._last_integer_position and frame_time_target, which
  were written but never read.
- Keep target_fps and set_target_fps() but document them as
  informational: nothing paces off them, yet ledmatrix-elections'
  test_scroll_pacing.py reads helper.target_fps back and third-party
  plugins may call the setter.
- Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to
  throttle", set_sub_pixel_scrolling's "default: True", and
  set_frame_based_scrolling's claim that it steps.

The plugins monorepo was grepped for every removed name; none is used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI

FONT_MANAGER.md told plugins to read display_manager.font_manager, which
does not exist, so a plugin following it failed to load with
AttributeError. The shared FontManager lives on the PluginManager and
BasePlugin._get_font_manager() returns it (with a fallback for harnesses).

Also removes the Fonts-tab override workflow and element-override panels
that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls

/plugins/store/search does not exist (404) and the list endpoint reads
query, not q. The registry guide's curl examples omitted the JSON
Content-Type, so the handlers saw an empty body and answered 400.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): use the shared core-key list in the last three private copies

StartupValidator warned "Plugin 'auto_update' is enabled but not found" on
every display start with auto-update or a dim schedule on; the reserved
plugin-id check missed auto_update, sync, location and the rest; and
ConfigManager's (uncalled) orphan cleanup would have deleted display,
schedule and auto_update. All three now read src/core_config_keys.py, which
also gains CORE_SECRETS_KEYS for the github/youtube secrets sections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): partial JSON saves to /config/main change only what they send

A JSON body with one field reset every checkbox in the sections it touched:
the MQTT bridge's brightness slider turned off disable_hardware_pulsing,
inverse_colors, show_refresh_rate and use_short_date_format, and a
timezone-only save turned off web-UI autostart and weekly auto-updates.
Missing-means-unchecked now applies only to form posts: form-encoded bodies
and the v3 forms, which mark themselves with a hidden __form_section input.

Also on the config routes:
- vegas_min/max_cycle_duration no longer match the generic *_duration rule,
  so they stop landing in display_durations and a blank one no longer
  rejects the whole Display save;
- saving from the Raw JSON editor calls start_setup_if_needed like the
  General form, so enabling auto-update there finishes its setup;
- the schedule and dim-schedule POSTs accept the per-day days.<day> shape
  their GETs return, as well as the flat form keys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): install plugin dependencies from the configured plugins directory

install_plugin_dependencies.sh scanned only plugins/, but the Plugin
Store installs into plugin_system.plugins_directory (default
plugin-repos), so the documented "Recommended" fix found 0 plugins on
every store install. It now reads plugins_directory from
config/config.json (relative to the project root or absolute, default
plugin-repos) and also scans plugins/ for dev symlinks, installing a
plugin reached through both only once.

With set -e alone, `pip ... | tee` took tee's exit status, so a failed
pip install was reported as success; set -o pipefail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: replace stale API names, line numbers and the api_v3.py path

- ADVANCED_FEATURES: StreamManager methods that exist
  (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the
  real on-demand status envelope ({status, data: {state, service}})
- app.py:199 / :144 / :607-619 line citations and
  web_interface/blueprints/api_v3.py (now a package) replaced with file and
  function names in ADVANCED_FEATURES, CONFIG_DEBUGGING,
  PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE,
  PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README
- CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use
  /config/raw/main to replace the file; describe where validation runs
- TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints
  usage)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): verify the web interface that actually ships, on port 5000

verify_installation.sh failed every healthy install: it required the
long-removed web_interface_v2.py and looked for a listener on port 5001,
while the web interface binds 5000 (web_interface/start.py). It now
checks the files ledmatrix-web.service runs (start_web_conditionally.py,
web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the
same 5001 port in its listen check, HTTP probe and printed URLs.

Port matches are anchored so :50001 no longer counts as :5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): one display-size contract: display_manager.width/height

CLAUDE.md (#580) says to read display_manager.width/height because
matrix is None when hardware init fails; the development guide, the
safety-harness doc and two DisplayManager docstrings still recommended
matrix.width/height. The bundled starlark-apps plugin read matrix.width
unguarded, so its magnify recommendation and frame scaling raised in
fallback mode (e.g. after the Pi 5 hardware refusal).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): make install_service.sh --help print usage instead of installing

install_service.sh parsed no arguments, so `sudo ./scripts/install/
install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md)
rewrote ledmatrix.service, ledmatrix-web.service and both update-verify
units and enabled/started them. It now handles -h/--help (usage, exit 0,
no changes) and rejects any other argument with exit 2 before doing
anything. Running it with no arguments, as first_time_install.sh does,
is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scroll): describe the fixed-step model and document frame_hold

Since #545 a crisp speed from scroll_config.configure() makes the helper
advance a fixed whole-pixel step per presented frame with no clock, and the
display manager's frame hold is part of the speed. The docs still described
the removed wall-clock model:

- scroll_config's module and configure() docstrings said speed is applied
  in time-based mode and that omitting the hold "falls back to fractional
  pixels"; omitting it actually runs the scroll frame_hold times too fast.
- SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both
  modes, and read a 20 ms stats median as missed refreshes although that
  is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step,
  the hold-dependent healthy median, that target_fps plays no part, and
  that a hand-added scroll_pixels_per_second loses to a schema-default pair.
- PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling)
  without frame_hold; it now documents the parameter (core 3.4.0) with a
  configure() + set_scrolling_state example.
- update_scroll_position/set_scroll_speed and set_scrolling_state
  docstrings say the same.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark target_fps legacy; describe what Vegas scroll_delay does

- General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after
  the sports_scroll fix nothing in core scrolling reads it. The field and
  its API validation stay so saved configs and plugins that read
  global_config['target_fps'] keep working. CONFIG_REFERENCE says the same.
- Vegas frame_based_scrolling/scroll_delay were described as frame-count
  stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based
  mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and
  still advances by elapsed time, so the applied speed is
  clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The
  config comments, render_pipeline comment and CONFIG_REFERENCE rows now
  say so. No behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(deps): describe how plugin dependencies are really installed

The guides said the web service runs as root, that installs pick --user
from os.geteuid(), and quoted a warning and a
PluginManager._install_plugin_dependencies() method that don't exist. The
web unit runs as the installing user; store installs go through
install_requirements_file() and sudo safe_pip_install.sh (root), with a
user-level fallback that says so, and load-time installs run in the
display service's own (root) interpreter.

Manual paths now use the configured plugins directory (plugin-repos/ by
default) instead of plugins/, which store installs no longer use, and
install_plugin_dependencies.sh is described as scanning that directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): count local changes one way for the preflight and the pull

The automatic update's preflight ignored mode-only changes and anything
whose status line contained plugins/ or plugin-repos/, then promised
"Automatic updates will not stash your changes". perform_core_update
used plain git status (modes count) and ignored only 'plugins/', then
ran 'git stash push -- :!plugins', which nothing ever pops. So an edit
to a bundled plugin under plugin-repos/, or the installer's chmods on
tracked scripts, passed the preflight and was stashed away for good.

- auto_update.local_changes() is the one predicate both use:
  core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/
  excluded by leading folder rather than substring (a core file under
  web_interface/static/v3/js/plugins/ now counts).
- Update Code's explicit stash leaves out both plugin folders; the
  pull's --autostash carries their edits and mode changes across and
  reapplies them.
- The automatic updater calls perform_core_update(stash_local_changes=
  False), which refuses instead of stashing edits that appeared after
  the preflight; update_core reports that as 'blocked'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): diagnostics follow the web autostart default and api_v3 package

#556 made a missing web_display_autostart mean "start" (only an explicit
false/off keeps the web interface down), but the diagnostics still said
otherwise: diagnose_web_ui.sh reported a missing key as "defaults to
false", diagnose_web_interface.sh said the web interface "will not start
unless this is set to true" and recommended enabling it, and
debug_web_manual.py printed False. Troubleshooting a down web UI pointed
users at a non-cause.

Both shell scripts now evaluate the setting with the launcher's own
autostart_enabled() (inline fallback if it cannot be imported) and report
on / off / not set (on) / unparseable config; debug_web_manual.py uses
the same function. They also check web_interface/blueprints/api_v3/
__init__.py: api_v3.py became a package in #553, so every healthy
checkout was reported as missing a file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(install): what install_service.sh installs; verify script port; no sudo for --help

install_service.sh installs and starts ledmatrix, ledmatrix-web and the
update-verify units, not only ledmatrix.service (systemd/README.md,
README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help'
as a harmless check; it now shows --help without sudo and warns what a
real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh
checks the web interface on port 5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note update-all, plugin system settings and script fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): size the preview after orientation and pixel mappers

display_geometry.physical_size claimed to give DisplayManager's answer but
only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured
after the library's pixel mappers, so a Rotate:90 / orientation 90 chain
previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for
128x64.

Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc
does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown
names leave it alone), and move the orientation composition here so
DisplayManager and the preview share it. The module docstring no longer
claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH.

Audit finding F18.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): refuse settings the rgbmatrix library aborts on, on every board

The library answers several settings with a NULL matrix or abort() rather
than an error, so the display service crash-looped (Restart=on-failure)
instead of reaching fallback mode: rows above 64, chain_length above 255
(uint8_t binding setter, documented as "no upper limit"), a misspelled
hardware_mapping, and parallel 2-3 on a single-output mapping, reachable
from the Display form on the default adafruit-hat(-pwm) mapping. #586 only
guarded the Pi 5 subset.

- src/matrix_support.py holds the rules for every board (Options::Validate
  ranges, binding integer types, mapping names and outputs from
  lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the
  API's numeric ranges.
- DisplayManager checks them before building options and raises
  MatrixSettingsRefused, so a hand-edited config falls back with a logged,
  reported reason. Emulator mode only warns.
- The config API refuses them with a 400 naming the setting; combinations
  are checked against stored values but reported only when the request
  sets a field involved.
- The hardware status file gains "cause" (settings/library/forced). The
  fallback log and Display banner give the Pi 5 rebuild hint only for a
  library failure instead of rebuild + gpio_slowdown advice for every
  failure; one Pi 5 slowdown recommendation (1-3, start at 1).
- The Display form offers classic/classic-pi1 and orientation 90/270 and
  renders any other stored mapping selected with a warning, so an
  unrelated save no longer rewrites them; the API accepts 90/270.

Audit findings F03, F16, F19, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(display): library limits, template defaults and Pi 5 slowdown

- rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs,
  classic/classic-pi1 mappings and orientation 90/270 documented.
- Defaults are the config.template.json values: config migration adds
  missing keys from the template, so the listed "code defaults" never
  applied.
- One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode,
  starting at 1.
- Troubleshooting describes the refused-settings fallback, and CHANGELOG
  corrects the Unreleased "no upper limit" entry.

Audit findings F19, F20, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py opens the panel with the service's options

--measure and --demo built RGBMatrixOptions from a private copy of the
display service's builder that had drifted: gpio_slowdown came from
display.hardware (default 2) instead of display.runtime (default 3), and
rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors,
pixel_mapper_config and orientation were skipped, with different defaults
(hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown
was measured -- or garbled -- in a setup the service never drives.

The option filling in DisplayManager._setup_matrix moves, unchanged, into
DisplayManager.apply_matrix_options(options, config), which _setup_matrix
calls and the script reuses (overriding only limit_refresh_rate_hz for
--measure). The script now loads the whole config rather than the hardware
block. Tests pin the script's options to the service's attribute for
attribute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py recommends keys the resolver honours

The ladder ended by telling users to set
display_options.scroll_pixels_per_second. scroll_config ranks that key
below the scroll_speed + scroll_delay pair, deliberately, and several
plugin schemas default the pair into config, so the advised key was
silently ignored (a schema-default 1/0.02 pair plus an advised 66 still
resolved to 50 px/s).

The advice is now the pair that selects the crisp speed exactly
(pixels_per_frame every frame_hold/refresh seconds), explains that the
pair outranks scroll_pixels_per_second, and gives the scoreboards'
per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the
printed pair over a schema-default pair and check it lands on the
advertised speed and hold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula

- SPORTS_UNIFICATION.md still presented honouring global target_fps as
  sports_scroll's added behaviour and its one user-visible gain; note that
  it was withdrawn because it had become a speed multiplier.
- ADVANCED_FEATURES.md gave Vegas scrolling as
  (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s
  by elapsed time, through a 0.1-5 px per scroll_delay clamp when
  frame_based_scrolling is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): scroll model fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(dev): link-github links plugins from the ledmatrix-plugins monorepo

link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git,
and those per-plugin repositories no longer exist: official plugins are
directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the
monorepo once into the dev directory, finds plugins/<name>,
plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and
links it under its manifest id. With an explicit repo URL it still links a
single-repository plugin as before.

dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a
fork), plus plugins_repo and plugins_branch; github_pattern, which was
documented but never read, is dropped and warned about. Ships
dev_plugins.json.example and git-ignores dev_plugins.json, both of which
the guide promised. Reading JSON falls back to python3 when jq is missing
(get_plugin_id silently returned nothing without jq).

update/status/list find the git checkout above a monorepo plugin
directory (its .git is not in the plugin dir), and update pulls a shared
checkout once. status no longer exits 1 when nothing is broken.

Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration,
workflow, store integration, hello-world link, submission), and the
nonexistent scripts/git-hooks/pre-push-plugin-version and
scripts/bump_plugin_version.py replaced with the real rule: bump the
manifest version and run update_registry.py. scripts/dev/README.md and
CLAUDE.md updated to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scripts): monorepo workspace layout; fix_perms and install READMEs

MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin;
setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into
plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the
workspace file opens LEDMatrix plus ../ledmatrix-plugins.

scripts/fix_perms/README.md listed cache directories
fix_cache_permissions.sh never touches and a 'ledmatrix' service user
that doesn't exist (also in scripts/install/README.md); adds
safe_pip_install.sh. install/README: install_service.sh installs the web
and update-verify units too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): keep the rollback's pip retries inside the unit time limit

The health check reinstalled the previous requirements by trying the
next bash path after any failure, including a 600 s pip timeout. Two
files, two paths: up to 40 minutes of pip alone, while systemd stops
ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the
rollback half-way and leaving the update 'verifying' until the web UI
calls it lost.

- Like permission_utils.install_requirements_file, only a sudo refusal
  moves on to the next bash; a pip that ran and failed or timed out is
  not repeated. The refusal wording is one list
  (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only
  verifier and pinned equal by a test.
- All reinstalls in one rollback share a 600 s budget.
- WORST_CASE_SECONDS adds up every timeout on the longest path (27.5
  min); a test holds it under the unit's TimeoutStartSec and that under
  the web UI's VERIFY_LOST_SECONDS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools

Plugin config was prepared differently depending on how it arrived:

- JSON POST /plugins/config built a partial body on schema defaults, so
  {"enabled": true} reset every other setting of the plugin. It now merges
  onto the stored section first, as the form path already did.
- Legacy-boolean normalization (#588) ran only at load: GET /plugins/config
  returned the raw boolean, posting it back failed validation, and hot
  reload handed plugins the raw section (a legacy dynamic_duration: true
  came back as a boolean). schema_manager.prepare_plugin_config (normalize,
  then defaults) is now used by PluginManager.load_plugin, both save paths,
  GET, the save notifications and DisplayController's hot-reload callback.
- The JSON save's filter kept only enabled/display_duration/live_priority
  and dropped a submitted skin, skin_options or vegas_* tuning key. There
  is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES,
  used by validation and by the save filter; PluginManager's
  CORE_OWNED_CONFIG_KEYS is its vegas subset.
- Plugin sections posted to /config/main were stored verbatim, including
  values /plugins/config rejects. They now go through the same preparation
  (_prepare_plugin_config_for_save, extracted from save_plugin_config), and
  a failing section rejects the whole save before anything is written.
- dev_server read only top-level defaults and let a schema enabled:false
  win; build_full_config shallow-merged overrides, dropping sibling
  defaults; the harness extracted defaults differently from the device.
  loading.build_config now uses the device's extraction and preparation,
  and dev_server, check_plugin, render_plugin and the harness all use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt-bridge): brightness changes apply live and touch nothing else

The display service's hot reload applies a saved brightness within a few
seconds, and /config/main no longer resets other display settings on a
brightness-only JSON body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): automatic update hardening

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI

It described web_interface_v2.py and index_v2.html (both gone), client-side
form generation, one POST per field with {key, value}, and 'no nested
objects'. The v3 UI renders plugin forms server-side from the schema
(pages_v3 partial + plugin_config.html macros, nested sections and
x-widgets), posts the whole form once, and save_plugin_config() merges onto
the stored section, validates, splits x-secret fields and notifies the
plugin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt): brightness saves apply via hot reload and leave other settings alone

The bridge README said brightness is applied on the display's next
restart; the display controller's config hot reload applies it within
seconds. It also now states that the bridge's partial JSON save changes
only brightness (the /config/main merge fix in this change set).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): don't log pip's output from the health check's reinstall

pip can echo a private index URL with embedded credentials;
permission_utils redacts it, the stdlib-only verifier cannot, so it
logs the exit code only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark the plugin_system toggles as unused legacy keys

auto_discover, auto_load_enabled and development_mode are read by
nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE
said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and
the REST reference listed them as live settings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): docs and developer tools group

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): legacy plugin-system toggles no longer count as a General save

auto_discover, auto_load_enabled and development_mode have left the General
form, so a post carrying only one of them is not a general-settings save and
must not treat web_display_autostart and auto_update as unchecked. The
plugin_system block itself is left as on main for the branch that reworks it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): config-save and plugin-config preparation fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(claude): re-check matrix_support.py rules when the library submodule is bumped

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address Codacy findings on the core audit PR

- plugin_manager.prepare_plugin_config: when the fallback legacy-boolean
  pass also fails, log a warning instead of a bare except/pass.
- api_client.js: request() refuses any endpoint that is not a plain path
  under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace,
  control characters) with INVALID_ENDPOINT before calling fetch(), and
  plugin ids are URL-encoded wherever they are put into a URL (also in the
  app-shell batch load).
- test_update_all.js: pins both against the shipped client.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): check endpoint control characters without a control-character regex

Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags
the \x00-\x1f range in checkEndpoint's regex. Test the char codes
instead; the endpoints refused are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(auto-update): make the seed script executable on disk, not only in the index

On Linux Repo.publish() commits with -a, which recorded scripts/run.sh
as 100644 upstream because the seed file was never chmod +x. The pull
then brought in the same mode the installer chmod had made locally, so
installer_chmod saw no mode change left to check. The updater was fine:
with the upstream commit at 100755 the --autostash carries the device's
chmod across. Verified under Linux (WSL, git 2.43): the old helper fails
exactly as CI did, the fixed one passes all 63 tests in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 16:37:29 -04:00
ChuckandClaude Opus 5 7e5967e160 fix(web): find installed plugins before anything has discovered them (#594)
The web process discovers plugins lazily: plugin_manifests is empty until
some endpoint calls discover_plugins(). Three routes consulted it without
discovering, so they misbehaved for as long as nothing else had run --
which, after every ledmatrix-web restart, is until someone opens the
dashboard:

- POST /display/on-demand/start answered 404 "Plugin <id> not found"
  (or "Mode <mode> not found"). Measured on a rig: 404 for over three
  minutes after a web restart, until GET /plugins/installed ran. The
  browser UI loads the plugin list first, so API-only callers (the Home
  Assistant MQTT bridge, scripts) are the ones who hit it.
- POST /plugins/toggle answered 404 "Plugin not found".
- POST /config/main did not recognise a plugin section, so it skipped
  secret separation and merged the section as-is: the plugin's API key
  was written to config.json in plain text instead of config_secrets.json.

Add _discovered_plugin_manifests(), which discovers when nothing has been
yet, and rescans once when a specific plugin id (or, for on-demand by
mode, a mode) is not found, so a plugin installed since the last scan is
found too. _installed_plugin_ids() now uses it.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 12:38:55 -04:00
ChuckandClaude Opus 5 f475038895 fix(cache): web UI can read what the display service caches again (#593)
* fix(cache): web UI can read what the display service caches again

ledmatrix-web.service carried CacheDirectory=ledmatrix. With User= set to
the installing user, systemd re-owns /var/cache/ledmatrix and everything
in it to that user and its primary group whenever the directory's owner
differs -- for a directory root created, on the first start. That erased
the root:ledmatrix setgid layout the installers set up, so every file the
display service (root) wrote afterwards was root:root 0660 and unreadable
by the web interface:

  WARNING - Permission denied loading cache for display_current_state ...

Since #547 install_service.sh renders the web unit from the template, so
every fresh install hit this. Measured on one rig: 392 unreadable files,
and the web UI's display status, on-demand state and plugin health empty.

Existing installs only receive `git pull`, never a reinstalled unit, so
the fix for them is in the code the root display service runs:

- DiskCache.set gives each file the directory's group (when the directory
  is group-writable) and 0660 on the open descriptor before the rename,
  independent of setgid. This also closes a window where a fresh file was
  visible as mkstemp's 0600.
- DiskCache.share_existing_files repairs files an older version left
  behind, once per process from the cleanup thread. It works through
  O_NOFOLLOW descriptors and skips hard links and other users' files: the
  directory is writable by the web user, and root must not be steered
  into changing a file outside it.

For new installs, the web unit drops CacheDirectory=/CacheDirectoryMode=,
and install_web_service.sh stops replacing an existing directory's
ledmatrix group with the user's group.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): on-demand and current-display status read the display's latest state

Found testing the cache-permission fix on a rig: once the web interface
could read display_on_demand_state at all, /display/on-demand/status kept
answering "active" for over 100 seconds while the file on disk said
"idle". Both status routes read the display service's keys through the
web process's memory tier, which serves the first copy it read for the
full max_age (120s). Read them with memory_ttl=0, as every other
cross-process reader (plugin health/metrics, the on-demand mailbox)
already does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): re-group the cache dir whenever the web user is outside its group

install_web_service.sh replaced an existing cache directory's group only
when it was root's. A directory in any other group the web user is not a
member of -- root:ledmatrix, for a user who is not in ledmatrix -- was left
alone, and every file root wrote there stayed unreadable to the web
interface. Replace the group whenever the installing user is not in it.

A directory whose group the user is already in (ledmatrix, or the user's
own group where CacheDirectory= left it) is still left as it is: re-grouping
a working directory strands the files already in it on the old group.

When the group does change and root-owned JSON files carrying the old group
are present, try-restart ledmatrix.service so DiskCache.share_existing_files
re-groups them through its symlink- and hard-link-safe path, rather than a
recursive chgrp.

Verified under WSL's systemd for seven directory states (user group,
ledmatrix member, ledmatrix non-member with and without root files,
root:root, missing, unnamed gid); the previous version left the non-member
case unchanged.

Addresses CodeRabbit review on #593.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 12:38:39 -04:00
ChuckandClaude Opus 5 1d51efe4c7 fix(backup): restore over existing files on hosts without os.chown (#592)
_copy_file() replaces each restored file and then carries the previous
owner across with os.chown. On Windows os.chown does not exist and
st_uid/st_gid are 0 rather than absent, so the ownership branch always
ran and raised AttributeError. That is not an OSError, so it escaped
every per-section handler in restore_backup(): a restore over any
existing config aborted at config.json and restored nothing.

Skip the ownership step where os.chown is missing, as
auto_update_setup.py already does. No change on POSIX.

test_restore_over_a_file_the_user_cannot_write simulates root-owned
files with chmod 0o444; on Windows that sets the read-only attribute,
which blocks any rename over the file, so it is skipped there. The
modes the app writes (0o644/0o640/0o600) replace fine on Windows.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 09:26:21 -04:00
ChuckandClaude Opus 5 7ae614aa35 fix(sports): recover from ESPN rejecting scoreboard date ranges (#591)
* fix(sports): recover from ESPN rejecting scoreboard date ranges

Since 2026-09-15 ESPN's site API answers `dates=YYYYMMDD-YYYYMMDD` with
400 "Failed to get events endpoint." for every sport. Single days, months
(`YYYYMM`) and season years still work. Every season and weeks-window fetch
in core failed, including the background service the scoreboards submit
their season schedules to.

src/common/espn_dates.py re-asks a rejected range as whole-month chunks
plus the leftover edge days, which tile the window exactly (a season is
8 requests, not 213). A month that comes back with exactly 500 events is
truncated (college baseball's March) and is re-asked day by day.

It also clamps `limit` to 500: above that ESPN truncates silently, e.g.
college football returns 25 of 68 games for one Saturday at limit=1000.

BackgroundDataService recovers rejected ranges on the worker thread and
advertises `handles_espn_date_ranges` so plugins can tell whether to hand
it a range. SportsCore, sports_shared, ESPNDataSource and APIHelper route
through the helper or the clamped limit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): ESPN date-range fallback and limit clamp

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): stop re-sending ESPN date ranges once one is rejected

Live scoreboards refresh every 30 seconds, and each refresh sent the range
first, got the 400, then fetched the chunks: three requests where one used
to do. After a rejection, ranges now go straight to chunks for six hours,
then the range is tried again so the workaround retires itself if ESPN
reverts. A single-day 400 does not set the memo, and when every chunk fails
the range request supplies the error without the chunks being fetched a
second time. Per-fetch chunk logging drops to debug; the rejection itself
stays a warning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): clamp limit only on ESPN scoreboard submissions

The background service is generic, and limit above 500 only truncates
scoreboards. /teams needs limit=1000 (college football has 762 teams and
limit=500 returns 500), so a teams submission must keep its limit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 17:12:28 -04:00
ChuckandClaude Opus 5 2082665252 fix(config): load config_secrets.json on hosts without os.geteuid (#590)
ensure_shared_group_ownership() - the chgrp self-heal ConfigManager runs
before reading config_secrets.json (#416) - looked up os.geteuid
unguarded. That name does not exist on Windows, and the AttributeError
is not an OSError, so it escaped the helper's best-effort handling and
every except clause in load_config(). Any Windows checkout with a
config/config_secrets.json got a ConfigError from every config load and
could not import web_interface.app.

That is what made test_update_all_plugins.py error at setup: its client
fixture imports web_interface.app. It was not state leaked between test
files - the trigger is whether the checkout has a secrets file.

Return early when os.geteuid or os.chown is missing. No change on POSIX.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 17:12:18 -04:00
ChuckandClaude Opus 5 f9b3d6ae52 fix(web): accept every panel size and row address type the rgbmatrix library does (#586)
* fix(web): accept every panel size and row address type the rgbmatrix library does

The Display form capped columns at 128 and chain length at 24, and its
submit handler (fixInvalidNumberInputs) rewrote anything larger to the cap,
so wide panels and long chains silently saved as the wrong size. The config
API checked none of the hardware numbers, so values the library rejects (odd
rows, parallel 4, PWM dither bits 3) saved and the matrix then refused to
start.

- Form limits now match the pinned library: rows even 8-64, cols >= 16 and
  chain_length >= 1 with no upper bound, parallel 1-3, PWM dither bits 0-2,
  PWM LSB nanoseconds 50-3000.
- save_main_config rejects out-of-range rows, cols, chain_length, parallel,
  brightness, scan_mode, pwm_bits, pwm_dither_bits, pwm_lsb_nanoseconds and
  gpio_slowdown with a 400.
- A stored gpio_slowdown or pwm_dither_bits of 0 renders as 0 instead of the
  default, so saving the tab no longer overwrites it.
- Row Address Type offers 5 (SM5368 / B707 row shift register). Verified on a
  Waveshare 96x48 V2 (24S-A1) on a Pi 4 with the Adafruit Triple LED Matrix
  Bonnet: rows 48, cols 96, row address type 5, BGR, GPIO slowdown 8.
- Help text and docs: FM6124-family panels use Panel Type Standard; on a Pi 5
  the library supports only row address types 0 and 2.

No change to the rpi-rgb-led-matrix submodule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): drop the rows cap and document every display setting accurately

Rows: no upper limit in the form or the API. Still even and at least 8. The
current rgbmatrix library rejects more than 64 per panel, so a larger value
saves but the matrix won't start; the help tip, README, config reference and
troubleshooting section all say so, and nothing here needs changing if the
library lifts the limit.

limit_refresh_rate_hz: the form accepts 0 (the library's "no cap"), a stored
0 no longer renders and re-saves as 120, and the API rejects negatives.

pwm_dither_bits stays 0-2: the library rejects 3 and 4, so the old form's
0-4 only ever let users save a config the display couldn't start with.

Docs and help tips, checked against the pinned library and its README:
- panel_type and rp1_rio get README entries
- show_refresh_rate prints to stdout; it never drew on the panel
- dither bits raise the refresh rate; the tip said they lowered it
- scan_mode is about interlacing at low refresh, not wrong colours
- disable_hardware_pulsing: hardware pulsing needs OE on GPIO 18 and the
  onboard sound driver off; software timing makes rows flash brighter
- gpio_slowdown guidance agrees between the README and the UI
- all 22 multiplexing values listed; every numeric setting states its range
- troubleshooting for a blank panel after a settings change, jumping rows
  and brightness flashes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): reject true and 5.5 for row_address_type and multiplexing

Both still went straight through int(), so a JSON true saved as 1 and 5.5
as 5. They now use the shared hardware range check like the other panel
fields. Review feedback on #586.

Also: the RP1 Backend tooltip said it is ignored on Pi 3/4 (it is ignored
on every model but the Pi 5), and the README gave the dynamic-duration
default cap as 90s; the code default is 180s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: refuse matrix settings a Raspberry Pi 5 can't drive

On a Pi 5 the pinned rgbmatrix library drives the panel through the RP1
chip, and that path supports only row address types 0 and 2, parallel 1-3
and the regular / regular-pi1 / classic / adafruit-hat(-pwm) mappings
(Rp1PioConfigSupported in lib/rp1/rp1_pio_backend.cc). For anything else
CreateFromOptions returns NULL; the Python binding doesn't check, so the
display process crashed on its first call into the matrix and systemd
restarted it into the same crash every 10 seconds.

- src/pi5_matrix_support.py: the rule and Pi 5 detection, matching the
  library's /proc/device-tree/model check
- DisplayManager raises before creating the matrix, so it is a logged init
  failure (reported by /api/v3/hardware/status) and fallback mode
- the config API rejects those settings on a Pi 5 when a request sets
  row_address_type, parallel or hardware_mapping
- the Display form offers only row address types 0 and 2 on a Pi 5, and
  warns when a stored value can't be used
- CLAUDE.md: re-check the rule whenever the submodule is bumped

Review feedback on #586.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 19:17:48 -04:00
ChuckandClaude Opus 5 9616a5a054 fix(web): auto_update and other core settings are not orphaned plugins (#589)
v3.4.0 shows "Plugin Config Warning - In config but not installed:
auto_update. Reinstall via the Plugin Store, or remove these entries from
config.json." auto_update is the core weekly-update setting from #581.
Reconciliation treated every top-level dict not in its private
_SYSTEM_CONFIG_KEYS list as a plugin id, and #581 could not know to extend
that list.

- Move core top-level keys into src/core_config_keys.py (CORE_CONFIG_KEYS)
  and use it in reconciliation. Tests fail if a config.template.json key or
  a key written by the general-settings save is missing from it.
- A secrets-file key only counts as a non-plugin when no installed plugin
  has that id. Plugin secrets are namespaced by id, so installed plugins
  with secrets were reported as missing from config on every run.
- still_unresolved() drops "not on disk" findings whose id is no longer a
  plugin entry in config, so a stored verdict clears without a restart.
- A plugin whose id is a core key is skipped with a warning, and the fix
  never writes a plugin stub over or in place of a core setting.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:19:39 -04:00
ChuckandClaude Opus 5 c200b5837d fix(plugins): normalize legacy boolean settings before schema validation (#588)
The news plugin's schema turned global.dynamic_duration from a boolean
into an {enabled, min_duration_seconds, ...} object. Installs that have
not saved the news settings since still hold `true`, so every start
logged "Plugin news config does not match its schema (loading anyway):
Field 'global.dynamic_duration': Expected type object, got bool" and
flagged news degraded.

The settings form already reads such a boolean as {"enabled": <bool>}
(render_nested_section in plugin_config.html) and the next save writes
the object. The loader did not. It now applies the same rule before
merging schema defaults and validating, so the defaults fill in the rest
of the object and the plugin receives it in the new shape.

The rule lives in schema_manager.legacy_bool_as_object /
normalize_legacy_booleans. It applies at any depth of nested objects
but not inside arrays, matching the form, and only to a real bool under
an object-typed property with an `enabled` child. Every other mismatch
still warns. A parity test renders the template macro against the helper
so the two cannot drift.

Nothing is written to config.json at load: the normalization is in
memory, and the next save of the plugin's settings persists the object.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:09:23 -04:00
ChuckandClaude Opus 5 9f2743471c fix(web): update-all skips Starlark apps and no longer misses plugins (#587)
* fix(web): update-all skips Starlark apps and no longer misses plugins

Check & Update All posted every entry from /plugins/installed to
POST /plugins/update, including the virtual starlark:<app_id> entries
that list installed Starlark apps. The store manager cannot find those,
so each answered 500 "plugin not found". Update-all now sends only
plugin ids (install_manager.js, and the older app-shell.js copy), and the
route answers a starlark: id with a 400 saying it is a Starlark app.

A request that got no HTTP answer was recorded as failed and never sent
again. On a device, a web-service restart mid-run killed the in-flight
request and refused the next one, stock-news, which was left on 2.6.2
with 2.8.0 available. Such requests are now re-sent with backoff
(about 30s) before being reported as failed. HTTP error answers are not
retried.

Tests: test/js/unit/test_update_all.js (run from pytest via
test/web_interface/test_update_all_plugins.py so CI covers it) and the
route contract for starlark: ids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): walk update-all retry delays without indexed lookup

Codacy's ESLint security/detect-object-injection rule flagged
retryDelays[attempt] as a High issue. The index was a bounded loop
counter over a fixed array, but shifting a per-plugin copy of the
schedule gives the same backoff without the pattern. No behaviour
change: test/js/unit/test_update_all.js (21) and
test/web_interface/test_update_all_plugins.py (7) pass unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:09:11 -04:00
ChuckandClaude Opus 5 fddb0e06db feat(web): honour x-display: hidden in plugin settings (#585)
* feat(web): honour x-display: hidden in plugin settings

Plugins keep deprecated and internal keys declared so stored configs keep
validating (weather api_key/radar_zoom, countdown's auto-generated row id),
but the settings form drew them as live controls.

A property marked "x-display": "hidden" -- or an object whose children are
all hidden -- now gets no control at any depth: top level, nested sections,
Advanced Settings (not counted either), array-table columns and the row
editor. A hidden top-level key is not reported in __rendered_section.

Saving never changes a hidden value. Plain and nested fields aren't posted,
so the save's deep merge keeps them; _set_missing_booleans_to_false skips
hidden booleans at every depth. A posted array row replaces the stored item,
so hidden row properties are carried as JSON-encoded hidden inputs and
decoded exactly on save (an id "1" stays a string). New rows get none.
JSON API saves are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): read hidden row keys without dynamic property access

Build the set of x-display: hidden item properties once and look values up
through Object.entries, instead of indexing objects by a variable key on
the lines this branch added (Codacy: object injection sink, 6 warnings).
Behaviour is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 15:47:06 -04:00
ChuckandClaude Opus 5 8360220809 feat(common): sports_helpers — the helpers all nine scoreboards carry identical copies of (#583)
* feat(common): sports_helpers, the helpers all nine scoreboards copy verbatim

Add src/common/sports_helpers.py: the helpers the scoreboard plugins'
sports.py carry byte-identical copies of (docstring-stripped AST, checked at
ledmatrix-plugins f09bff2), so a later plugins PR can delete its copies once
it floors on the core release that ships this.

- Free functions: clamp_window, clamp_seconds, logo_needs_refresh (lazy
  src.logo_downloader import, as in the plugins), spread_weighted_order,
  MIN_WINDOW_DAYS / MAX_WINDOW_DAYS. All nine plugins.
- SportsHelpersMixin (no __init__, stateless): _mode_customization,
  _setting_int, _reset_dwell_on_reentry, _next_switch_index,
  _spread_weighted_order (all nine), _odds_color and
  _upcoming_date_and_time_text (all but ufc), plus the _favorite_key seam
  from base_classes core.py for later phases.

A new module rather than more methods on sports_shared: a plugin that
deletes a copy and relies on an existing module having grown the method
fails at runtime with AttributeError on an older core, which neither the
loader nor check_min_core_version.py can see; a missing module fails at load.

Tests: behaviour for every helper, a derived host contract, and a parity
test that AST-compares every body against every plugin copy when
LEDMATRIX_PLUGINS points at a checkout (skipped otherwise).
test_common_is_hardware_free.py imports src.common and every sports_* module
with rgbmatrix blocked and scans src/common for module-level imports of
src.base_classes, src.display_manager and src.plugin_system (no existing
violations).

Nothing in core imports the new module; no behaviour change. CHANGELOG
Unreleased entry and a converging note in docs/SPORTS_UNIFICATION.md.
__version__ is not bumped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(common): address review on sports_helpers and the hardware-free test

- SportsHelpersMixin docstring and CHANGELOG: constructor-free, but it keeps
  lazy state on its host (_reset_dwell_on_reentry, _next_switch_index).
- test_common_is_hardware_free: the runtime check now filters every
  FORBIDDEN package, src.plugin_system included; the AST scan resolves
  relative imports against src.common, so `from .. import plugin_system`
  and `from ..plugin_system import x` are caught. Guard tests for both.
- Parity skip reason names the CI guard that runs the same comparison:
  ledmatrix-plugins scripts/check_sports_helpers_parity.py (#495).
- _odds_color: line-level pylint disable for a not-callable false positive
  (getter is None-checked); the AST is unchanged, parity still passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 15:32:21 -04:00
ChuckandClaude Opus 5 9e3f184d81 docs(changelog): complete 3.4.0 with weekly auto-updates and scroll timing fix (#584)
* docs(changelog): complete 3.4.0 with weekly auto-updates and scroll timing fix

#581 (weekly automatic updates) and #582 (scroll frame-stats idle gap)
merged after the 3.4.0 section was written in #580. v3.4.0 will be tagged
on main including both, so they belong in 3.4.0.

Also record src.common.font_layout (#539, #565), a src.* module plugins
may import that shipped in 3.4.0 but was never listed, and mark
display_geometry and auto_update_setup as core-internal.

Correct the 3.3.0 historical note: remote tags v3.3.0 (bc2dbf38) and
v3.3.1 (32d637a4) both report "3.3.0" and both ship sports_shared.py.
The "3.2.0" claim came from a stale local tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): cover the whole 3.4.0 release since v3.3.1

The 3.4.0 section listed plugin-facing API and per-element customization
but not the rest of what merged since v3.3.1. Group it under subheadings:
Install and updates, Scrolling, Plugins, Web interface, Tools and
security, Fixes, and put the existing customization block under its own
heading.

Omitted on purpose: #569 (fixes a regression and an editor race in the
unreleased per-element framework), #570 (no runtime change), and
test-only, refactor and dev-tooling PRs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 15:24:38 -04:00
ChuckandClaude Opus 5 869e36fb2f feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback

A General-tab toggle (off by default) checks for and installs LEDMatrix and
plugin updates once a week, overnight in the configured timezone.

- Pre-update checks skip (and report) instead of forcing: local edits or
  commits, merge/live rebase, no upstream, low disk, missing health check, or
  a version that was already rolled back. An abandoned rebase (HEAD back on a
  branch) is cleared, since it would otherwise block every pull.
- The pull reuses the Update Code path (now perform_core_update(), which
  reports dependency install failures as data).
- ledmatrix-update-verify.service, started via a .path unit from a request
  file, restarts the services from its own cgroup, requires them to come up
  and stay up, and otherwise resets to the previous commit and reinstalls the
  previous requirements. It runs a copy of the checker taken before the pull.
- No SSH needed: switching the toggle on restarts the display service, which
  (as root) installs the two units from the repo templates for the web user.
  first_time_install.sh installs them too and takes --enable-auto-update /
  LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh).
- Plugins update after the code passes its check; failures, blocks and
  rollbacks raise an Overview banner and show under the toggle.

Tested end to end on a Pi: web-UI setup, a good update, a broken web service
and a broken display (both rolled back), a blocked local edit, and an
abandoned rebase found on the device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(auto-update): address static-analysis findings

- Replace the subprocess.CompletedProcess the verifier fabricated for a
  command that could not start with a plain namedtuple; nothing is executed
  there, but the scanner flags any CompletedProcess built from variables.
- Mark the subprocess imports with the repo's standard B404 annotation (all
  calls are list-form argv, no shell).
- Mark the rollback-failed message as not SQL (B608 matched its wording).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): CI failures on Linux

- Keep the setup result when chown fails. CI runs as a non-root user, where
  chown to the web user raises; that discarded the result file, so the
  General tab would never learn whether setup worked. Regression test added.
- Register the two new /api/v3/system/auto-update routes in the URL map
  snapshot.
- Use utility classes app.css defines (space-y-1, hover:text-red-600).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): address review feedback

- Health check: a failed restart command no longer lets the check run
  against the still-running old process; it counts as a failure (and after a
  rollback, as a failed rollback). An unreadable restart count is never
  treated as stable, since a crash loop looks healthy between attempts.
- Installer writes the auto_update setting to a temp file and swaps it in,
  keeping mode and owner, so a running config watcher never reads a
  truncated config.json.
- Verify unit quotes its command-line paths (install folders with spaces);
  setup refuses folder names systemd would reinterpret (%, quotes,
  backslashes, control characters) and says so on the General tab.
- The auto-update status route no longer returns exception text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): keep error detail in the status route's 500

test_web_error_detail requires every 5xx handler to log the traceback and
return describe_exception(e), which redacts credentials, so failures are
diagnosable from the web UI. Dropping it for CodeQL broke that policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): dismiss route rejects non-object JSON with 400

A JSON array or scalar body made `.get('alert_id')` raise, returning 500.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): let the app-wide handler answer status-route errors

CodeQL (py/stack-trace-exposure, #709) flagged the route's own except,
which returned describe_exception(e). web_interface/app.py's error handler
already logs the traceback and returns the same redacted detail for any
unhandled exception, so the local copy is removed: same response, no new
exception-to-response flow, and test_web_error_detail's policy still holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:58:57 -04:00
ChuckandClaude Opus 5 d01da3bd9f fix(scroll): stop timing the idle gap between scrolls as a frame (#582)
ScrollHelper.last_frame_time was set once in __init__ and thereafter only
at the end of log_frame_rate(). Nothing re-armed it when a scroll began, so
the first frame of every scroll was timed against the last frame of the
*previous* one and the whole idle period between them was recorded as a
single frame.

Measured over 3 hours on a 256x64 Pi 4, that produced 31 windows reading

    Scroll frame stats - 0.0 fps over 1 frames | median 136776.02ms
    p95 136776.02ms max 136776.02ms min 136776.02ms | stalls 0 (0.0%)

and -- worse, because it is not obviously wrong -- put the same gap in the
max field of otherwise healthy windows, where the worst values were 537s
and 604s. It also counted as one stall per scroll start: at ~500 frames to
a window that is ~0.2%, against measured stall rates of 0.07-0.16%. The
stall rate is the number used to judge whether a scroll change worked, and
it was the same order of magnitude as its own artefact.

The first frame of a scroll has no predecessor, so it has no frame time.
last_frame_time is now None until one is rendered, and reset_scroll() puts
it back -- the same treatment last_update_time already gets three lines
above, for the same reason. reset_scroll() alone is not enough, because the
scrollers actually emitting these lines never call it, so a sample at or
past the 5s log interval is dropped as well: nothing that renders a scroll
takes that long over one frame. Seeding also restarts the window timer, or
the boundary is already overdue when the second frame arrives and every
scroll opens by reporting a window of exactly one frame. A window whose
samples were all dropped now logs nothing rather than reporting the gap.

docs/SCROLL_PERFORMANCE.md documented the diagnostic in terms of a
"Frame time: N ms" line that 6031e705 replaced with the aggregate, so its
grep matched nothing on any rig. The section now describes the line that is
actually emitted, reads duplicate frames off skips and a below-median
result rather than a 2ms mode, and adds a command that ranks every scroller
by p95 -- verified against 3 hours of journal, where it reproduces
src.base_odds_manager at p95 44.08ms against 10.19ms for the two scrollers
already on src/common/scroll_config.py.

requirements.txt still offered scipy for the sub-pixel interpolation path
deleted in #570. Installing it has no effect; the entry says so.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 18:44:52 -04:00
ChuckandClaude Opus 5 814c21de1c chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0 (#580)
* chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0

Skins: no current scoreboard plugin builds on src.base_classes, so the only
skin hook (SportsCore._render_game) never runs. The plugin schema endpoint no
longer injects the Visual Skin dropdown, the store hides and refuses
"type": "skin" registry entries, and GET /api/v3/skins reports
supported: false with a message. Stored skin config still loads and saves.
src/skin_system/ and its tests are unchanged apart from the support flag.

Docs: check_plugin.py/render_plugin.py examples use --plugin; document
BasePlugin.get_update_interval() and its interaction with the manifest
update_interval; CLAUDE.md drops the stale template line number and
recommends display_manager.width/height.

Preview size: new src/display_geometry.py holds the size computation and
defaults DisplayManager uses (double-sided applied, chain_length default 2).
The web preview, /display/current, Starlark magnify default, sync handshake
and two dev scripts use it.

Release: __version__ 3.4.0, CHANGELOG 3.4.0 section plus a 3.3.0 tag note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address CodeRabbit review on #580

- Preview fallbacks (SSE stream and /display/current) use logical_size({})
  (128x32, the shared default) instead of a hard-coded 128x64.
- display_geometry treats a non-mapping display/hardware block as missing,
  so a malformed config.json falls back to defaults instead of raising
  AttributeError (which turned the Starlark render into an HTTP 500).
- Docs: the static update interval falls back manifest -> plugin config
  -> 60s, in both the API reference and the architecture spec.

Not taken: validating double_sided copies against chain_length/parallel.
An orientation Rotate: or U-mapper pixel mapper decides which axis panels
lie on, so the counts would reject working setups (the existing
vertical-split test is one).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display_geometry): a non-finite hardware size raises ValueError, not OverflowError

CodeRabbit flagged the Starlark magnify default in
_standalone_render_starlark_app for truthy non-mapping display values. That
case was already handled by a9e1bd0b (_display/_hardware treat a non-mapping
block as missing, covered by test_non_mapping_display_config_uses_the_defaults),
and the magnify it produces from the 128x32 defaults is the same as from 64x32.

Checking the same path found one input that still escaped: Python's JSON
parser accepts Infinity, and int(inf) raises OverflowError, which neither the
Starlark path (TypeError, ValueError) nor the preview stream in app.py caught,
so a hand-edited "rows": Infinity returned HTTP 500. physical_size now raises
ValueError for it, matching its documented contract, so every caller's
existing fallback applies. DisplayManager already caught Exception.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:36:58 -04:00
ChuckandClaude Opus 5 914bf2002f fix(install): grant and harden safe_pip_install.sh in first_time_install.sh (#579)
first_time_install.sh granted the web user safe_plugin_rm.sh but not
safe_pip_install.sh, unlike scripts/install/configure_web_sudo.sh. On devices
set up only by the first-time installer, install_requirements_file could not
use the root wrapper and fell back to a user-level install that root-run
ledmatrix.service may not see.

Also harden both sudo-granted helpers to root:root 755. first_time_install.sh
never did this, and Step 11's project-wide chown to the user would undo it if
placed in Step 10, so it runs at the end of Step 11.1.

Add a test that parses the ledmatrix_web sudoers rules from both installers
and asserts they grant the same commands, and that every granted helper is
hardened (after the chown, in first_time_install.sh).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 17:45:03 -04:00
ChuckandClaude Opus 5 47afaaac2b fix(install): build rgbmatrix on ARMv6 Pi Zero / Pi 1, and stop locking users out of the submodule (#577)
The pinned rpi-rgb-led-matrix commit emits the ARMv7-only `dmb ishst`
instruction in lib/rp1/rp1_rio_backend.cc, guarded only by __arm__, so the
build fails on every ARMv6 board ("selected processor does not support
`dmb ishst' in ARM mode"). Bump the pin to upstream 1ee4f76, which merges
12d839f (guard on __ARM_ARCH >= 7) plus docs only.

The installer also needed two changes for that bump to reach anyone:

- git pull never moves an existing submodule checkout, so a device that
  already failed would keep building the broken commit. The build step now
  moves the checkout forward to the pin — never backward or sideways (a
  `git submodule update --remote` checkout is left alone), and never fatal.
- The submodule git commands ran as root on the user's clone (git's SUDO_UID
  exemption allows it), leaving .git/modules/rpi-rgb-led-matrix-master
  root-owned and the user unable to run git in it. They now run as the
  project directory's owner, and root-owned leftovers are handed back.
  Root-owned installs keep running as root.

test/test_install_rgb_checkout.py covers the non-root sync scenarios under
the installer's strict mode, checks that every called _helper is defined
before use, and pins the one-shot-install.sh -> first_time_install.sh
contract. Verified the tests fail on five deliberate mutations. Root/owner
scenarios were exercised manually under WSL Ubuntu, and the library was
cross-compiled for arm1176jzf-s at both pins (old: rp1_rio_backend.cc fails
at line 120; new: 16/16 sources compile).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 17:38:17 -04:00
ChuckandClaude Opus 5 6d1cbfb70b fix(web): plugin config page survives stored values the schema outgrew (#578)
Two stored shapes broke the config form:

* A scalar under a field that is now an object. News' dynamic_duration
  was a boolean and is becoming an object; render_nested_section did
  `key in true` and the whole page failed to render. Look into dicts
  only, and carry a legacy boolean over as the object's `enabled`, so
  the next save upgrades it without switching the feature off.
* A custom feed logo with a path but no id. The template always emitted
  an empty `logo.id` input, which the save route parsed to null, failing
  the id's string type on every save. Emit it only when there is an id,
  as custom-feeds.js already does.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 14:03:15 -04:00
ChuckandClaude Opus 5 8e6d7c280f test(element-style): cover the stateless element_color clamp path (#576)
#569 fixed _normalize_color and #572 covered the resolver path. That test's own
docstring notes the resolver "normalizes colour separately from element_color",
and the other path had no test: the stateless element_color(), which
src.common.sports_card delegates to and which every one of the nine scoreboard
plugins takes for each per-element colour it draws.

That is the path that regressed. element_color() moved here with the per-element
customization framework, the coercion rejected out-of-range components where the
reader it replaced clamped them, and a rejection reads as "not configured" -- so
one component over 255 painted the element white while the user's colour sat in
their config. Every scoreboard's test_element_text_colors.py failed on it, and
it took two plugin PRs red on CI to surface.

Six cases: clamping, in-range untouched, hex, unparseable fallback, missing
element, and agreement with sports_card.coerce_rgb. The last is the point --
the two shared readers disagreed about the same value, so this asserts against
coerce_rgb directly rather than restating the arithmetic, and any future move
of element_color has to keep them consistent.

Verified by mutation: restoring the rejecting coercion fails two of the six,
alongside the resolver test from #572.

Tests only; no source change.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:27:10 -04:00
ChuckandClaude Opus 5 dac71afedc fix(web): full-height plugin config form and full-width plugin card descriptions (#573)
* fix(web): let the plugin config form use the full page height

The form wrapper has carried `max-h-96 overflow-y-auto` since #145, but
the class was a no-op until #568 defined `.max-h-96` in app.css. That
silently capped the whole config form at 24rem with a nested scrollbar.
Drop the cap so the form flows naturally and the page scrolls.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): give installed plugin card descriptions the full card width

The enable/disable toggle was a flex sibling of the whole text column
(name, metadata, description), so it reserved its width for the full
height of the card body. Descriptions wrapped into a narrow strip,
leaving blank space under the toggle and making cards very tall.

Move the toggle into a header row with just the name and badges, and
render the metadata and description below at full width.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:27:00 -04:00
ChuckandClaude Opus 5 a6b3384032 fix(web): show the action script's error in the file-manager widgets (#574)
A failing plugin action returns a 400 whose JSON body carries the
script's own message, but both file-manager widgets threw it away:
plugin-file-manager's toggle always said "Toggle failed", and
json-file-manager's request helper threw "Server error 400" before
reading the body. That hid of-the-day's "Category ... not found in
config", which is why its toggles looked broken for no reason.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:26:46 -04:00
ChuckandClaude Opus 5 11bf39cd66 fix(web): two plugin config saves that always returned 400 (geochron, news) (#575)
* fix(web): render widget-less arrays of objects as a table, not comma text

An array of objects with no x-widget (geochron's `cities`) fell through to
the comma-separated text input. Jinja joined each item as a Python dict
repr, the save route read them back as a list of strings, and the schema
rejected them -- so every save of the plugin returned 400 "Configuration
validation failed", whatever setting was changed.

Default such arrays to the existing array-table widget, which already
edits arrays of objects and posts `field.N.key` inputs the save route
rebuilds into a list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): don't leave an empty object stub in array items on save

The unchecked-checkbox pass walked into every nested object of an array
item looking for booleans, creating it when absent. A news custom feed
with no logo came out with `logo: {}`, which fails the logo's
`required: [id, path]`, so every save of the news plugin returned 400.

Recurse into a scratch dict instead and attach it only if a boolean was
actually set in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:04:11 -04:00
ChuckandClaude Sonnet 5 bc60b41445 test(element-style): cover ElementStyleResolver's colour clamp path (#572)
* fix(colour): clamp out-of-range text_color components instead of dropping them

_normalize_color returned None for a triple with a component outside 0..255,
and None means "not configured" to element_color -- so configuring
[300, 0, 20] silently handed the element its *default* colour rather than red.
Every scoreboard reads its per-element colours through this path, so the bug
reached all eight.

It is also the odd one out: sports_card.coerce_rgb and
SportsShared._coerce_rgb both clamp, and core's own test is named
test_coerce_rgb_clamps_rather_than_rejecting. The rejecting normaliser arrived
with the shared readers in 82a65ad2 (#425) while the eight plugins' colour
tests kept asserting the clamping behaviour they had before, so the two sides
have disagreed ever since.

Clamped inline rather than delegating to coerce_rgb: sports_card already
imports element_style, so importing back would be circular.

Adds the core assertion whose absence let this drift -- element_color had no
test covering an out-of-range component.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(element-style): cover ElementStyleResolver's own colour clamp path

CodeRabbit flagged that the new sports_card clamp regression test only
exercises element_color(); ElementStyleResolver._resolve() normalizes
configured colours through a separate call to the same _normalize_color,
comparing against a schema/classic reference to decide user_forced_color.
Add a resolver-level case so a future regression in that path (e.g. going
back to rejecting out-of-range components instead of clamping) is caught
too.

Mutation-checked: fails if _normalize_color rejects instead of clamps.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KP6kWxjUtJi72c56GaMmC8

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:03:23 -04:00
ChuckandClaude Opus 5 f9b1f87e8d fix: clamp colour components, and let the style editor actually take over (#569)
* fix(element-style): clamp out-of-range colour components instead of rejecting

A regression this framework shipped. The eight scoreboards used to read their
colours through sports_card.coerce_rgb, which clamps; routing them through the
shared element_color sent them through _normalize_color, which rejected any
component outside 0..255 and fell back to the default. So a configured
[999, -5, 20] -- a typo'd bright red -- rendered white instead of (255, 0, 20).

Their own test_element_text_colors.py caught it: one case of nineteen, in all
eight plugins, failing only once the core change reached main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): the style editor takes over its own blocks -- and gets to at all

Two defects, both found by rendering the real partial in a browser rather than
by reading the code.

It was losing a race to its own fields. The hand-off guard asked "do any
fallback controls differ from their server-rendered defaults?" as a proxy for
"is someone editing this?". But the fallback holds this block's own font
fields, and the font-selector widget populates them on the same 50ms timer --
so a plain page load, with nobody touching anything, raced into "dirty" and the
editor removed itself, leaving the 701-line accordion form it exists to
replace. Measured: seven customization.*.font selects dirty ~60ms after
injection, clean again by 400ms. The question is whether a *person* typed, and
event.isTrusted answers exactly that; the listeners now go on synchronously,
because the edit worth protecting can happen before initWidget runs.

It took over too much. Taking over removed the whole fallback section, but a
customization block can hold more than styling -- football keeps
favorite_result_colors there -- so that removed the only UI those fields had,
and the editor also rendered them as an element, giving every row an "enabled"
and three colour columns. Core now marks the blocks it recognises as styling
(the compact declaration already did; hand-written adoption did not), the
widget renders only those, and the template drops only the children the widget
reports owning.

Verified on football's real schema: 28 rows across four mode tabs, columns
Element/Font/Size/Colour/X/Y, favorite_result_colors still editable with its
ten inputs, no duplicated field names, no console errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): style editor no longer drops layout-only fields it never rendered

CodeRabbit flagged elementKeys() in style-editor.js: render() claims the
whole customization.layout child as the widget's own (removing it from the
generic fallback renderer, since posting the same offset twice is worse),
but elementKeys() only listed keys that also have their own top-level style
block. A hand-written schema can put a key under layout that never got one
-- a logo, a timeout indicator, a possession arrow with a position but no
font or colour -- and that key's only control silently disappeared: no row
in the style editor's table (elementKeys never listed it) and no fallback
section either (layout was removed wholesale).

elementKeys() now appends any layout-declared key not already covered by a
style element, so table() renders a row for it (layout columns only, no
style columns) and the wholesale layout ownership claim stays truthful.

Verified against current code before fixing. New regression test
(test/js/unit/test_style_editor_element_keys.js, following this repo's
existing eval-extraction pattern for testing widget JS without a browser)
fails against the reverted function and passes with the fix; added to
run_all.js and the suite table in test/js/README.md.

Full pytest suite: 4887 passed, 62 skipped, 2 failed -- both the
pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias gap on this sandbox,
identical on origin/main, unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dpg3HLWohdCUdzz2QNHanm

* fix(web): style editor no longer strands leaf-valued layout fields

A prior fix on this PR made elementKeys() append any layout-only key with
no style block of its own (a logo, a timeout indicator, a possession
arrow), so table() draws a row for it instead of losing it when the
wholesale `layout` claim removes the generic fallback. That covers a
layout-only key shaped like an object (x_offset/y_offset, ...), because
columnsFor() only ever produced columns from a key's *sub-fields*.

It missed the case where the layout-only key's own value is itself a
leaf -- a plain "show_logo" boolean directly under layout, no x/y object
underneath. elementKeys() still lists it (any row: no matching column),
so it renders as an uneditable blank row and its only control -- the
generic fallback checkbox -- is still gone. Confirmed by executing the
real widget's render() against a synthetic schema in Node (a DOM-stub
harness, not committed): the field's name never appeared as an <input>.

columnsFor() now gives such a leaf key a column keyed to itself
('layout-leaf'), and elementRow() binds it to the leaf's own path
(customization.layout.<key>, matching the name the fallback would have
used) instead of leaving every cell blank.

New regression test (test/js/unit/test_style_editor_layout_leaf_columns.js,
following this PR's existing eval-extraction pattern) checks the leaf
column is produced, is self-keyed, doesn't duplicate, and that a schema
with no leaf-valued layout key is unaffected; wired into run_all.js and
the suite table in test/js/README.md.

test/js/run_all.js: 84 + 6 + 6 = all suites passed (jsdom unavailable
here, DOM suites skip as before). Python suite untouched by this change;
test_style_editor_extra_fields.py, test_style_editor_save_roundtrip.py
and the one PIL-dependent style_editor_takeover.py case fail identically
before this commit -- missing flask/PIL in this sandbox, not this PR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep layout-leaf style-editor columns distinct from name collisions

columnsFor() keyed a layout-only leaf field's column by its bare field
name. If an unrelated element's style block or another element's layout
axis block happened to declare a sub-field with that same name, the
`!seen.has(key)` guard skipped creating the leaf's column, silently
dropping its only control again -- the same failure the leaf-column fix
was meant to close, just reached through a name collision (CodeRabbit
review on 324a7ea).

Key layout-leaf columns under a namespaced id so they can never be
shadowed by an unrelated column sharing their name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): CSS.escape() the owned key before it becomes a selector

container.dataset.ownedKeys round-trips schema property keys through a
DOM dataset attribute, and the takeover handoff spliced each one
straight into '[data-child-key="' + k + '"]' with no escaping --
inconsistent with this codebase's own convention elsewhere
(plugin-file-manager.js, app-shell.js's escapeCssSelector) for building
a selector from a dynamic value. A key containing a quote or backslash
would break the selector or be steerable; Codacy's static analysis
flagged this pattern (1 high ErrorProne finding on PR #569, current
head at the time) as a new issue, though its dashboard is unreachable
from this sandbox (egress to app.codacy.com is blocked) and the
check-run API returned no detail text -- verified and fixed by reading
the diff directly rather than the tool's own description.

Added a source-assertion regression test alongside this file's
existing ones (this behavior lives in an inline script no Python test
executes).

Full suite: 4888 passed, 62 skipped, 2 failed -- both the pre-existing
Europe/Kiev/Asia/Calcutta tzdata-alias gap in this sandbox, identical
on origin/main, unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): every advertised layout offset gets a control in the style editor

The editor took the whole layout section over but matched offsets to style
rows by exact key. A hand-written schema's two blocks were never named alike --
football styles score_text but positions score -- so of football's eleven
positionable things only status_text had a control. Score, odds, both logos,
timeouts, possession, down-and-distance, date, time and records were options
the schema advertised and the renderer reads, reachable nowhere in the UI.

Core now resolves each style element's layout key through alias_keys, the map
the resolver already reads offsets with, and records it as x-layout-key. The
widget reads that rather than carrying a second copy of the rules, and posts
under the key the schema declares: football's own offset reader looks up
layout.score, so a value saved as layout.score_text would be kept and never
drawn. Layout entries no style element claims get an "Other positions" table
with its own columns, in every mode panel as well as the base one, in the order
the plugin declared them.

Verified in a browser against football's real schema: 92 of 92 layout fields
(23 base, 23 per mode) rendered exactly once under their declared names, none
posted under a style key, no duplicated field names, favorite_result_colors
still editable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:48:36 -04:00
ChuckandClaude Opus 5 7e580dc005 fix(wifi): make Connect work from the setup AP (#571)
* fix(wifi): make Connect work from the setup AP

Joining a network from LEDMatrix-Setup has to take the AP down first, which
drops the phone that sent the request. The connect endpoint answered only
after the attempt finished, so the browser never got a reply and the WiFi
tab's Connect button appeared to do nothing.

- /wifi/connect answers 202 immediately while the AP is active and connects
  in a background thread; the result (never the password) is reported via
  /wifi/status as last_connect_attempt. A second connect while one is
  pending gets 409.
- connect_to_network holds a /tmp flag for the attempt; the monitor daemon
  skips AP management while it is fresh. Previously the daemon's
  disconnected counter, accumulated over the whole AP session, re-enabled
  the AP on its next tick in the middle of the connect.
- The WiFi tab and captive setup page explain the handoff up front, and on
  reopening show why the last attempt failed. The wrong-password message
  now works: the route sets the error_type the captive page checks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(wifi): serialize connect attempts on both paths

Addresses CodeRabbit review on #571:

- Check for a pending attempt before branching on AP state. A background
  attempt takes the AP down long before it finishes, so a second click
  used to bypass the 409 and start a competing synchronous connect.
- Record pending for the synchronous (non-AP) path too, so two requests
  can't overlap and have the first clear the daemon's in-progress flag
  while the second is still connecting.
- Clear the pending state if the background thread fails to start, rather
  than refusing every later request until restart.
- Say the setup network returns "within a few minutes": a stale flag plus
  the daemon's grace period can take longer than one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:40:41 -04:00
ChuckandClaude Opus 5 d51f7ada14 chore(scroll): drop the dead sub-pixel path, and two dev-tooling papercuts (#570)
Three independent changes, none of which alter runtime behaviour.

1. Remove ScrollHelper._get_visible_portion_subpixel and
   _interpolate_subpixel (162 lines). get_visible_portion dispatches only to
   _blend_visible_portion, so _get_visible_portion_subpixel had no caller, and
   _interpolate_subpixel was reachable only from inside it -- a closed island.
   _blend_visible_portion's own docstring already records that the scipy path
   it replaced was dead; the replacement landed but the corpse stayed.

2. scripts/check_plugin.py: also search ../ledmatrix-plugins/plugins. The
   scoreboards live in the sibling checkout, so --all silently skipped every
   one of them and only --plugin-dir reached them.

3. .gitignore: ignore config/.config_secrets.json.tmp.*, which the suite
   leaves behind several of per run.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:42:50 -04:00
ChuckandClaude Opus 5 d1e821c625 fix(web): harden, polish and optimize the web UI per the Sept 2026 audit (#568)
* fix(web): harden, polish and optimize the web UI per the September 2026 audit

Works through docs/archive/WEB_UI_AUDIT_2026-09.md (health 8/20).

Implementation integrity (P0)
- app.css now defines every utility class the templates and JS use,
  including .hidden, so the ~145 JS show/hide toggles work. Button reset,
  and base component rules (.btn, .form-control) wrapped in :where() so
  utility classes on the same element win. New static-audit test fails
  when a used utility class has no rule.

Accessibility
- Focus rings render (the old ring rule referenced undefined variables);
  one :focus-visible outline everywhere; skip link; labelled nav landmarks.
- Shared dialog helper (js/utils/dialog.js): role/aria-modal, focus trap,
  Escape, focus return, applied to every modal.
- Named icon-only buttons and labelled ~70 form fields.
- Toasts announced once; errors persist >= 10s; one showNotification.
- Captive WiFi page: live region, timeouts, dark mode, 16px inputs.

Performance (Pi Zero 2 W)
- SSE streams and tab timers pause when hidden or off-tab; the display
  stream only runs while a preview is visible. app-shell.js deferred.
- Widget scripts served as one versioned bundle (/assets/widgets.js):
  52 -> 21 script tags, 66 -> 35 requests on first load.
- Stdlib gzip fallback when flask-compress is missing: first-load JS/CSS
  1358 KB -> 291 KB on the wire. SSE untouched.

Theming and responsive
- File managers, form fields and Fonts upload on theme tokens; bare
  inputs themed in dark mode; no more white surfaces.
- No horizontal overflow at 375px on any tab; 44px touch targets on
  coarse pointers; reduced-motion respected; header title truncates.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings on #568

- json-file-manager: focus-trap releases kept in a Map (no dynamic
  property access or delete; no value-returning forEach callback)
- notification / schedule-picker: style and day-label lookups via Map
- app.js: move the pending-queue assignment out of the expression
- diff_viewer / error_handler: named function declarations instead of
  arrow consts

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: check the OAuth widget ships in the widget bundle

base.html no longer tags widget scripts one by one; they load through
/assets/widgets.js. Assert the page requests the bundle and the bundle
contains google-oauth.js, which is what the test was protecting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): address review feedback on #568

- widget bundle version fingerprints every file (name, mtime_ns, size)
- gzip fallback appends Accept-Encoding to an existing Vary header
- dialog helper: releasing a non-top dialog no longer moves focus out of
  the dialog the user is in
- labels: file-upload targets its file input; fallback config fields get
  label for/id pairs; native color input has a fallback name
- utility audit also reads class names inside bound :class expressions

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): give the native color-picker input an accessible name

CodeRabbit flagged this on PR #568 as an outside-diff finding (never
posted inline, so it was missed in the round of fixes that addressed
the other 6 review comments). The <input type="color"> only carried a
title attribute; screen readers don't reliably announce title, and
there's no other label naming the control when showHexInput is false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings in app-shell.js

- drop the unused catch binding on the SSE JSON parse
- move the pending-notification queue assignment out of the expression

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): contain plugin widgets/ dir and bound style-editor retries

From CodeRabbit review on #568 (code that arrived with the main merge):
- serve_plugin_widget resolves widgets/ with resolve_under before
  resolving the manifest script under it, so a symlinked widgets
  directory can't become the containment base (CWE-22). New test.
- style-editor init stops polling after ~10s when the widget never
  registers and leaves the plain fallback fields in place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 09:42:24 -04:00
ChuckandClaude Opus 5 69d408b321 feat(core): one per-element display-customization framework, wired into the web UI (#566)
* fix(sports): rebuild un-shared faces through the pinned layout engine

unshare_element_fonts re-instantiates a duplicate font face so two
elements can be told apart by id(). It did so through bare
ImageFont.truetype, which takes PIL's default layout engine rather than
the one src/common/font_layout.py pins. Raqm and Basic disagree on
fractional advances -- that disagreement is the reason the pin exists,
having broken golden images across machines -- so a rebuilt face could
measure differently from the shared face it replaced, on any host where
Raqm is installed.

These were the only two call sites in src/ bypassing the pin.

The guard asserts that the rebuild goes through the pinned loader rather
than comparing engine values: where Raqm is absent, bare truetype returns
BASIC anyway, so an engine comparison passes whether or not the pin is
honoured. The first draft of this test did exactly that and passed with
the bug reintroduced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): drop the two dead client-side config-form renderers

generateConfigForm and generateSimpleConfigForm (580 lines) were defined
on the Alpine component and never called: server-side Jinja replaced them,
as pages_v3.py:641 records. Nothing in any template invokes them -- there
is no x-html in the templates and no bracket access on the component.

They carried their own x-widget dispatch, which made them an active trap:
the next person adding a widget would reasonably think both renderers
needed updating.

plugins/config_manager.js (PluginConfigManager, 133 lines) goes for the
same reason -- loaded on every page from base.html, referenced only by
itself and by an archived doc.

Kept, having checked them: widgets/example-color-picker.js is the worked
example docs/widget-guide.md points plugin authors at, and
widgets/plugin-loader.js is the client half of a documented feature
(manifest-declared plugin widgets) whose server route is missing --
soccer-scoreboard already ships a widgets/custom-leagues.js that this
loader is meant to fetch. That is an unfinished feature to complete, not
dead code to delete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): serve plugin-declared widgets, and actually ask for them

LEDMatrixWidgets.loadPluginWidget has always fetched
/static/plugin-widgets/<plugin>/<widget>.js, and docs/widget-guide.md has
always documented that path, but nothing served it. soccer-scoreboard has
shipped a 17KB widgets/custom-leagues.js since August that could never
load. Both halves were missing, not just the route:

- serve_plugin_widget serves the script from the plugin's widgets/
  directory as text/javascript. The manifest is the allowlist -- only a
  widget the plugin declares is reachable -- so installing a plugin does
  not publish everything it ships. Path handling mirrors the sibling
  serve_plugin_web_ui: allowlist regexes, os.path.basename, resolve() +
  relative_to() containment, and the ledmatrix- prefix fallback. The
  declared script name is guarded too, since it comes from the plugin
  rather than the request.

- The config form never requested one. Its x-widget dispatch is a
  hardcoded list of core widget names, so a plugin's own widget fell
  through to a plain text input. An unrecognised x-widget on a string
  field now asks ensureWidget() for it. The text input stays as the
  fallback and is removed only once the widget has actually rendered, so
  a missing or broken widget costs the user an editor rather than their
  configured value on the next save.

- manifest_schema.json gains "widgets", so the declaration is validated
  rather than merely tolerated by additionalProperties.

Verified in a browser against the real partial: a declared widget loads,
registers and renders, and its field posts exactly one value; a field
whose widget 404s keeps its text input and still posts its value.

Not addressed: loadPluginWidgetsFromManifest still has no caller. The
per-field ensureWidget path is lazier and is what the form now uses, so
that bulk helper is dead weight -- worth removing, but left alone here
rather than inventing a call site for it.

Known limitation, documented: only string-typed fields take this path.
object/array/boolean/number fields and enums are dispatched by the
template's own branches, which still only know core widgets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(element-style): a wrong-size BDF now keeps its font, not its size

BDF fonts are fixed-size bitmap strikes: FreeType accepts only the pixel
size baked into the file and raises for anything else. 32 of the 35
shipped fonts are BDF, so a size picked in the web UI usually is not a
valid strike -- and load_font caught that failure with its generic
"unloadable font" handler, which substitutes PressStart2P. Asking for
5x7.bdf at size 10 therefore rendered a completely different typeface,
silently.

It now falls back to the file's own native size instead, which is what
SportsCore._load_custom_font_from_element_config has always done. The
native size is read via FontManager._read_bdf_native_size rather than a
fourth copy of that parser, matching how core.py already delegates.

Also here, because they are the same code path:

- native_bdf_size() is exposed for the web UI, which needs to know when a
  size field can take effect at all. None means "free choice".
- ElementStyle.font_size now reports the size actually realised rather
  than the one requested. Callers lay out from it, and reserving space
  for a size nothing was drawn at is how this surfaces.
- The module font cache is a bounded LRU (256) instead of an unbounded
  dict. The display process runs for weeks and every config save can add
  a (font, size) pair; every other hot cache in the codebase is bounded
  this way.

Untouched configs are unaffected: the shipped classic fonts are the three
TTFs, so nothing was hitting the substitution path by default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): per-mode style and offset overrides

Lets one element be styled differently per situation -- a scoreboard's
live / upcoming / recent cards, weather's current / hourly / daily
screens -- under customization.modes.<mode>.

The mode is bound at construction rather than passed per call. That is
what makes this cheap to adopt: SportsUpcoming and SportsRecent are
already separate instances with distinct SKIN_MODE values, so binding
once makes every existing style()/offset_value() call site mode-aware
without editing any of them. A per-call mode argument exists for the rare
host that renders more than one mode.

The two layers answer different questions, deliberately:

- The base layer keeps the existing "differs from the schema default"
  rule, because the save flow writes the full default object into
  config.json whether or not the user touched it.
- A mode layer is pure override -- its fields default to None, so
  presence is intent. Nothing writes into it unasked, so there is nothing
  for the stricter rule to protect against.

None therefore means inherit, and has to stay distinct from 0: a mode
y_offset of 0 means "sit at the base position", not "no preference".
This is the same distinction scroll_card.switch_* draws with "inherit".

A malformed mode value falls back to the resolved base value rather than
to the caller's default -- caught by the degradation tests, which is what
they are for: resolving the mode first let one bad string in a mode block
silently discard a good base offset.

With no modes block, and for every existing caller, resolution is
unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): declare per-mode overrides in config_schema.json

A plugin adds "x-style-modes": ["live", "upcoming", "recent"] alongside
its x-style-elements declaration and gets a customization.modes.<mode>
group per mode, with every field of every declared element repeated as an
override.

Those override fields are typed nullable and default to null, which is
the whole trick. The save flow writes schema defaults into config.json
wholesale, so giving a mode field the base element's default would make
every mode a frozen copy of the base the first time a user pressed Save,
and the base would stop reaching them. Null means inherit. The mutation
test for this is explicit: with concrete defaults, a base font_size of 14
resolves as 10 with user_forced set.

min/max from the declaration carry into the mode blocks, so an
out-of-range override is rejected by validation rather than clamped
silently at render time.

Also: the emitted font field now carries "x-widget": "font-selector". The
widget already shipped and the config form already allowlisted it -- the
hint was simply never emitted, so the field rendered as a bare text box
that the user had to type a font filename into.

Verified through the real SchemaManager path -- load_schema, defaults
extraction, merge_with_defaults, validation, then resolution -- rather
than against a hand-built dict, since the thing at risk is what that
pipeline does to a null.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): render the config form from the schema the save route validates

The form read config_schema.json with a raw json.load while
api_v3.save_plugin_config went through SchemaManager. Those are not the
same schema: SchemaManager applies expand_style_elements, which turns a
compact customization.x-style-elements declaration into the per-element
blocks the form knows how to render.

Without it, that customization object has an x-style-elements key and no
"properties", so the template's object branch matched nothing and the
section rendered as empty space -- while saving still validated against
the expanded shape. of-the-day ships the compact form, so its
customization section has been invisible in the web UI.

pages_v3 gains a schema_manager the way it already has config_manager and
plugin_manager. use_cache=False matches the save route, so an edited
schema is not served stale during plugin development. The raw read stays
as a fallback for callers that register this blueprint without one.

Checked before making the change: load_schema does nothing here except
read, validate and expand -- inject_skin_selector is a separate method it
does not call -- so this is not a behaviour change for schemas without
the declaration.

The test pair renders the same compact schema with and without a
SchemaManager, so it documents exactly what was broken as well as what is
fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): style-editor widget -- a row per element instead of 65 accordions

Rendered element by element, a realistic scoreboard's customization block
is 65 nested sections, and reaching one per-mode font size takes five
levels of expanding. The widget collapses that to one compact row per
element -- font, size, colour, X, Y -- with a tab per declared mode.

It emits ordinary inputs under the same dotted names the generic renderer
would produce, so the save/validate/merge pipeline is untouched: no hidden
JSON blob and no new server-side parsing. It is driven entirely by the
schema block it is handed, so fields added to the schema later appear
without editing the widget. If it fails to load or throws, the generic
nested rendering it replaces is left in place.

Fixing two things the save path got wrong for nullable fields, found by
posting what the widget actually emits:

- The indexed-array recombiner (text_color.0/.1/.2 -> one list) compared
  the declared type to the string 'array', so a per-mode colour, typed
  ["array", "null"], was never reassembled and failed validation on save.
  _parse_form_value_with_schema had the same comparison.
- A blank nullable field became [] rather than None, which then failed the
  minItems the colour array declares. Null is the inherit sentinel, so it
  has to survive.

And two things the widget itself got wrong, found by looking at it:

- An unset base control fell back to the select's first option, so an
  untouched scoreboard claimed every element used 10x20.bdf -- and the
  size box then locked itself to that bitmap font's fixed size. Base
  controls now show the schema default; mode controls stay blank, because
  blank there means inherit.
- Elements arrived alphabetised (Detail and Odds above Score). Flask's
  JSON provider sorts keys, so declaration order has to be stated
  explicitly; expand_style_elements now emits x-propertyOrder, which the
  generic renderer already honoured too.

Size is disabled and shown as fixed for a bitmap font, using the
scalable/native_size the font catalog now reports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): visibility, alignment and scale per element

Completes the customization vocabulary: hide an element, align it, and
resize a logo, alongside the font/size/colour/offset that already existed.
All three per mode.

They resolve to "change nothing" until the user asks for something -- True,
None and 1.0 -- rather than to whatever the schema declares. That is the
same invariant the font fields keep: a caller that honours them still
renders an untouched config exactly as it did before they existed. A
schema default therefore does not count as a choice, which matters because
the save flow writes that default into config either way.

scale sits in the layout block with the offsets rather than in the element
block, because it is geometry: a logo has a scale and no font. The widget's
columns come from the schema, so a logo row shows visibility, offsets and
scale and no empty font cell.

Two bugs found by the tests rather than by reading:

- A nullable enum needs null in its enum list, not just in its type. The
  mode copy of `align` defaulted to null and then failed its own schema, so
  a plugin declaring any enum field with modes could not save at all. Six
  tests failed on this before any of them reached what they were testing.
- defaults_from_schema only ever extracted font/font_size/text_color, so
  the schema defaults for the new fields were invisible to the resolver and
  a declared default read as a user choice.

Widget: the table scrolls horizontally and pins the element-name column.
Nine columns do not fit the config panel, and clipping them hid the offsets
entirely while scrolling them made every row anonymous.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): resolve elements under the names plugins actually use

Two naming conventions collided as the scoreboards grew. Counted across
the published schemas: the style block names elements with a _text suffix
(score_text, status_text, detail_text), while the layout block mostly uses
the bare noun (score, date, time, odds) -- except status_text, which kept
the suffix in seven plugins and lost it in two. records vs record splits
seven to two the same way.

A lookup now tries the exact name first and then the spellings that mean
the same thing. Exact-first is what makes this inert for any config that
already matches; the aliases only decide cases that resolved to nothing
before.

This is also what makes migrating to the compact declaration form safe.
That form uses one key for both blocks, so a scoreboard adopting it asks
for layout.score_text while its users have layout.score saved -- without
the aliases, every offset they had dialled in would silently become 0.

Applies to the style block, the layout block, the schema defaults and the
per-mode overrides, since the drift shows up in all four.

Not attempting to canonicalise on write: renaming keys in config.json
would break the plugins still reading the old spelling from their own
bundled code, and the drift costs a dict miss rather than correctness.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(plugins): BasePlugin.styles -- per-element styling every plugin inherits

Adopting the element-style system meant repeating three things in every
plugin: a guarded import, finding its own config_schema.json, and
rebuilding the resolver when on_config_change swapped the config dict.
This is those three things once, on the class all 45 plugins already
inherit from.

    title = self.styles.style('title_text',
                              classic_font='PressStart2P-Regular.ttf',
                              classic_size=8, classic_color=(255, 255, 255))

The classic_* arguments are the adoption contract: with nothing configured
they come back verbatim, so a plugin that switches to this renders exactly
as before until a user changes something.

A plugin with one instance per display mode sets STYLE_MODE on the class
and every existing lookup becomes mode-aware without a call site changing
-- which is the point of binding the mode to the resolver rather than
passing it per call. styles_for() covers a plugin that renders several
modes from one instance.

Schema discovery reads the concrete class's own module rather than this
file, because this file lives in src/plugin_system where no plugin schema
exists -- the same trap SportsCore._config_schema_path documents. The
first mutation test for that passed anyway: an installed plugin's module
directory and its entry under plugins_dir are the same path, so the test
could not tell the two apart. The case where they diverge is a plugin
symlinked in for development, and the test now forces that shape.

Getting discovery wrong is silent rather than loud: with no schema the
resolver has no defaults to compare against, so every configured value
reads as a deliberate override and the plugin quietly stops honouring its
own shipped styling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): adopt hand-written customization blocks, and widen the font list

Nineteen plugins spell their style elements out longhand instead of
declaring them -- football's block is 701 lines for seven elements -- and
predate this system entirely. Core now recognises that shape, so they pick
up the row-per-element editor and the real font picker on a core update
rather than on a plugin release. Checked against every published schema:
21 plugins adopt, and the defaults of each still validate against the
schema generated for it.

Detection requires *every* field in a block to be one this system
understands. A looser "has at least one style field" rule sweeps in
baseball's `count`, which carries a text_color beside geometry that means
nothing here. That distinction took three attempts to test: the first two
assertions passed under both rules, because an over-eager rule leaves a
fontless block looking untouched and only surfaces as an extra row in the
editor.

The hardcoded font enum is replaced rather than extended. Football lists
five of the thirty-five installed fonts, which is why a font a user
uploads can never appear in one. It is not a curated safe set -- it omits
some twenty other faces that fit the declared size cap just as well -- it
is the fonts that happened to exist when it was written.

Widening it does need a guard, though, and not the one the schema already
has: a bitmap font ignores font_size and renders at its size baked into
the file, so `maximum: 16` cannot stop a 27px face. The picker now filters
out fixed-size fonts taller than the element's own declared ceiling, which
drops exactly the four that would overflow a 32px panel and keeps the
other thirty.

Per-mode overrides stay opt-in: core cannot invent a plugin's display
modes, so `x-style-modes` remains the one line that unlocks them. Their
layout half covers every positionable element rather than only those with
a style block -- the two namespaces do not line up in a hand-written
schema, and football positions six things (logos, timeouts, possession)
that have no style block at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): remove the two Fonts-tab panels that reported invented data

"Element Font Overrides" let a user configure an override, showed a
success toast, and changed nothing. All three endpoints behind it were
stubs -- GET returned a hardcoded {}, POST and DELETE returned success
without calling anything -- each marked "This would integrate with the
actual font system".

Wiring them to FontManager would not have fixed it. The machinery there is
real (_load_overrides/_save_overrides persist config/font_overrides.json,
resolve_font applies them, and the countdown plugin genuinely consumes
it), but the panel's element dropdown offered eleven invented keys --
nfl.live.score, clock.time, weather.current -- that no plugin has ever
read. An override saved against one of those would have persisted
correctly and still done nothing.

"Detected Manager Fonts" goes for the same reason. It claimed to show
"fonts currently in use by managers (auto-detected)"; its own comment said
"we'll simulate this", and it listed every font in the catalog with a
hardcoded usage_count of 1 -- the panel beside it, with fabricated
numbers attached.

Per-element font choice now lives in each plugin's own config editor,
against the elements that plugin actually has, and covers size, colour,
offsets, visibility, alignment and scale rather than family and size.

Kept: the font library (upload, preview, delete), which works, and
/fonts/tokens, which is a stub but genuinely feeds the preview's size
dropdown. FontManager's override methods are untouched -- countdown uses
them.

Verified in a browser with the tab's JS running: no console errors, 35
fonts listed, upload and preview intact. Removing the panel meant unwiring
it from populateFontSelects too, which would otherwise have bailed out
early on the missing select and left the preview dropdown empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(sports): one reader for element colours and layout offsets

There were two copies of the per-element colour read and three of the
layout-offset read. They had already drifted -- the scroll-card renderer
carries a comment about having ignored offsets its own schema advertised
-- and each new capability had to be added to all of them or silently work
in some places and not others.

All of them now go through src.element_style, which is what carries the
alias handling and the per-mode lookup. That lands immediately for the
nine plugins importing these modules: a scoreboard asking for `score_text`
offsets finds the `layout.score` its users configured, and a Live instance
resolves its own colours through SKIN_MODE without any call site passing a
mode.

_normalize_color learned "#RRGGBB" in the process. The scoreboards' own
readers have always accepted it, so the shared one had to, or consolidating
would have quietly dropped a form users' configs may hold. _coerce_offset
picked up the non-finite guard the scroll-card reader had and the other two
did not.

_get_layout_offset is promoted onto SportsCoreSharedMixin. Each plugin
still carries its own copy in its bundled sports.py, which wins by MRO --
so adopting this is a deletion in the plugin, and until that deletion
nothing changes for it.

Note for whoever runs the suite next: test_display_dirty_tracking.py is
order-dependent. Fifteen of its tests failed in one full run and passed in
the next with no change in between, and pass in isolation. Pre-existing,
unrelated to this, but it makes a full-run diff untrustworthy until it is
fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): record the element-style work under Unreleased

This file's own preamble asks for it: a plugin may delete its bundled
fallback copy of a core module only when its manifest floors on the first
release that shipped that module, which requires the additions to be
recorded here against a version.

Names a plugin can now import and floor on -- the stateless layout_offset
and element_color readers, alias_keys, native_bdf_size, the resolver's mode
binding, BasePlugin.styles, and the promoted
SportsCoreSharedMixin._get_layout_offset -- plus the schema and web-UI
changes, the four fixes and the three removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(fonts): log the BDF native-size read failure instead of swallowing it

The bdf-native-size lookup in get_fonts_catalog() caught any exception
and silently discarded it. Every other guarded read added in this PR
(the manifest parse in _declared_widget_script, the SchemaManager
fallback in _load_plugin_config_partial) logs before falling through
to the same degraded behavior. This one didn't, which is the shape a
silent-exception-swallow lint rule flags. Behavior is unchanged --
native_size still comes back None -- but a corrupt or unreadable BDF
file now leaves a trace.

Verified: font-related tests (140) and the full suite still pass,
with only the 2 pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias
failures already present on origin/main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address CodeRabbit findings on the style-editor/font-selector PR

- Fix _load_font_sized double-wrapping the (font, size) tuple on the
  missing-font path, which handed callers a tuple instead of a font.
- Fix _set_nested_value skipping an explicit None when the key already
  existed, which silently kept stale overrides when a user cleared a
  nullable per-mode field or blanked all channels of an indexed color.
- Preserve BDF scalable/native_size metadata through fetchFontCatalog's
  catalog-format mapping so maxFixedSize filtering actually applies.
- Stop caching an empty array on a failed font-catalog fetch so a later
  call can retry instead of being stuck with the failed result.
- Keep a saved font selected in the style editor even when it no longer
  fits a newly declared maxFixedSize, instead of silently deselecting it.
- Don't drop in-progress user edits to fallback fields when a plugin
  widget finishes loading asynchronously and takes over the form.
- Tighten the removed font-override endpoint test to assert 405, not
  just != 200.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): a partial save no longer switches off checkboxes it never showed

An HTML checkbox posts nothing when unchecked, so the save route walked the
schema and forced every boolean missing from the form to False. That is right
for the rendered form and wrong for every other caller: a script, the MQTT
bridge or a curl against the documented endpoint never rendered a checkbox, and
reading its silence as "all off" turns a one-field save into a mass disable.

Found on hardware. Posting four customization.* keys to a live device switched
off nfl.enabled, ncaa_fb.enabled and every display-mode toggle in one request.

The form now reports the top-level sections it drew (__rendered_section), and
inside those an absent checkbox still means unchecked -- including a section
whose only fields are checkboxes that are all off, which no heuristic could
recover. A post with no marker only touches objects it actually posted a field
from. Meta fields are dropped before form keys are treated as config paths,
because unknown keys are otherwise written straight into config.json.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(sports): resolve element colour by name, and honour visible/align/scale

Two of the three gaps this framework shipped with.

Colour by name. A draw resolved its colour by comparing the *identity* of the
font object it was handed, which cannot tell two elements apart when they share
a face -- so those draws went out white. Every bitmap font is in that case,
because a freetype.Face cannot be re-instantiated to un-share it, which is how
an element rendered in any of the 32 shipped BDF fonts silently lost a colour
its picker had offered all along. _draw_text_with_outline now takes
element="score_text" and reads the colour by name; the identity path remains
for un-annotated callers, but narrows before giving up -- one configured colour
among the sharers is the only thing the user can have meant.

Visible, align and scale. The resolver has understood these since the
framework landed and nothing consumed them: an element could be marked hidden
in the web UI and still render. Adds the stateless readers, the mixin
accessors, and a scale parameter on the one shared logo-sizing seam (keyed into
the cache, so two elements scaled differently cannot be served each other's
image). Naming an element in a draw also honours its visibility.

Untouched configs are unaffected: every new parameter defaults to today's
behaviour, and all ten affected plugins render pixel-identically to main across
every harness size.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): how to declare styleable elements; harden the widget's lookups

The plugin-author guide for the compact x-style-elements declaration -- what
each key does, how to read values back without breaking the "user-forced only
when it differs from the default" rule, and why a hand-written block needs no
changes to be adopted.

Also clears the static-analysis findings on style-editor.js. Every lookup in
that file is keyed by something out of a schema or a saved config, so a key of
__proto__ or constructor would walk the prototype chain and hand back a
function instead of a schema; reads now go through an own-property helper. The
panel registry became a list, and the flagged vars moved to their function
roots.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear the remaining static-analysis findings

Five, all on lines this branch touched.

The Python one is not a new defect: _set_missing_booleans_to_false's first
parameter was always named `config`, which shadows the `config` submodule
imported for its side effects at the bottom of this module. Editing the
signature simply put the existing warning on a changed line. The parameter is
the plugin's config dict, so `plugin_config` is what it should have been called
anyway; callers pass it positionally and are unaffected.

The JavaScript ones are the object-injection rule firing on reads keyed by
data. own() now goes through a property descriptor, so the one unavoidable
data-keyed read is no longer a computed member access; at() consumes its path
instead of indexing it; and the column set is a Map, which has no prototype to
pollute and needs no guarded reads at all.

Verified the widget still renders identically against football's real schema:
29 element rows, all four mode tabs, values populated, no console errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): drop the hasOwnProperty alias the descriptor read made redundant

own() now reads through Object.getOwnPropertyDescriptor, so the alias it used to call has no remaining reference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 11:53:50 -04:00
ChuckandClaude Opus 5 92ac231138 fix(fonts): load 4x6 on its pixel grid, from any working directory (#565)
* fix(fonts): load 4x6 on its pixel grid, from any working directory

`extra_small_font` loaded 4x6-font.ttf at 6, off the face's 7px grid.
Under `draw.fontmode = "1"` the mono rasteriser thresholds each glyph at
50% coverage, so every glyph lost its fourth column and deformed:
christmas-countdown rendered "UNTIL" as "VM1JL". The advance is 5px at
both sizes, so snapping to 7 reflows nothing.

- Sizes in DisplayManager._load_fonts go through crisp_size() instead of
  literals. crisp_size / FONT_PIXEL_GRID / FONT_NAME_ALIASES move to
  src/common/font_layout.py; sports_card re-exports them.
- Mirror the fix in VisualTestDisplayManager, the harness's fork of
  _load_fonts. Without it every golden is blessed at the old size.
- Resolve bundled font paths against the install root, not the cwd.
  FontManager._resolve_asset_path now delegates to
  font_layout.resolve_asset_path (kept by name; plugins probe for it).
- The startup banner's middle rung snaps to 7; the 5 rung stays off-grid
  on purpose (the only size that fits a dotted quad on 64px).
- loading.py reads all plugin JSON as UTF-8 (cp1252 on Windows aborted
  check_plugin.py on a 0x9d byte).
- check_plugin.py reports in ASCII and never dies on an unencodable char.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(fonts): resolve relative asset paths from the install root, not the cwd

resolve_asset_path checked os.path.exists(relative_path) unconditionally,
so a relative asset path was still resolved against the process cwd first
-- exactly the dependency this module exists to remove. An unrelated
working directory that happens to contain assets/fonts/4x6-font.ttf (a
stale checkout, a copied assets folder, another project) would shadow the
real bundled font instead of the install root ever being consulted.

Only an absolute path is now returned as-is; a relative path always
resolves against _INSTALL_ROOT first, matching the docstring's stated
contract. FontManager._resolve_asset_path delegates to this function, so
it's covered by the same fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 10:47:01 -04:00
ChuckandClaude Opus 5 772258f73e docs: add PRODUCT.md and September 2026 web UI audit (#567)
* docs: add PRODUCT.md product context for web UI design work

Captures durable product truth (users, positioning, operating context,
constraints, principles) so design passes on the web control panel share
one source. Open decisions (offline-only, CSS build step, WCAG target)
are recorded as undecided rather than adopted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: add PRODUCT.md and September 2026 web UI audit

PRODUCT.md captures durable product context (users, positioning,
operating context, constraints, principles) for web UI design work.
Open decisions (offline-only, CSS build step, WCAG target) are recorded
as undecided rather than adopted.

docs/archive/WEB_UI_AUDIT_2026-09.md records the technical audit of
web_interface/ (8/20): the hand-rolled Tailwind subset in app.css leaves
333 used utility classes undefined (including .hidden), focus rings never
render, modals lack dialog semantics, and SSE/polling never pause. Includes
a verified-and-rejected section so the cache-busting false positive is not
re-raised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 16:53:45 -04:00
ChuckandClaude Opus 5 9ad7528c9b fix(config): stop same-second backups overwriting each other (#564)
* fix(config): stop same-second backups overwriting each other

A backup's version is its identity. save_config_atomic() hands the path
back, rollback_config(backup_version=...) looks that version up, and the
paired secrets backup is found by reusing the same string.

The version was stamped at second granularity, so two saves inside the
same second produced the same filename and the second shutil.copy2()
silently overwrote the first backup. The path a caller was still holding
then pointed at different content, and rolling back to it restored the
wrong config. A user saving twice in quick succession lost a restore
point with no error.

list_backups() made it worse. It parsed the version off Path.stem, which
drops only the last dot-component, so for config.json.backup.20240101_120000
parts was ['config', 'json', 'backup'] and parts[-2] was 'json' -- never
'backup'. The filename branch was unreachable: every backup fell through
to the mtime fallback and reported a second-granularity restamp of its
mtime rather than the name on disk, so a unique filename alone would not
have been enough for rollback to find the right version.

Stamp microseconds, and never overwrite an existing backup -- on a
collision bump a -N suffix rather than lose a restore point. Parse the
version off the exact glob prefix so it round-trips with the filename,
still reading the legacy second-granularity format so restore points that
predate this keep working.

Two tests had encoded the bug:

  - test_multiple_config_changes asserted a rollback produced plugin1=45
    with plugin2=15, a state no single backup ever held -- 45 was only in
    the second backup, 15 only in the first. It passed because the two
    saves collided onto one file, so the first version resolved to the
    second's content. Corrected to the state that backup actually holds.

  - test_backup_rotation asserted against a hardcoded max of 3 while
    setUp configured 5, and still passed: every save in its loop collapsed
    onto a single filename, so there was only ever one backup to count and
    rotation was never exercised. It now asks the manager for its limit
    and overshoots it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): fold collision suffix into ordering, close backup-path race

_parse_backup_version() stripped any trailing "-segment" unconditionally,
so a collision-suffixed backup parsed to the exact same timestamp as its
sibling and list_backups() had no deterministic way to order them. Only
strip the suffix when it's numeric, and fold it back in as extra
microseconds so same-tick collisions sort newest-first reliably.

_create_backup() also checked backup_path.exists() before shutil.copy2(),
which two concurrent callers can both pass for the same path -- the second
copy2() then silently destroys the first call's restore point. Reserve
each path (config and, when configured, secrets) with exclusive file
creation instead of a check-then-copy, retrying on a real conflict.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 16:13:50 -04:00
ChuckandClaude Opus 5 6b3028ad58 test: isolate DisplayManager globals across modules, and name the failure (#563)
Follow-up to #562. That commit fixed the actual cause of the intermittent
15-test failure in test_display_dirty_tracking.py -- the emulator's fixed TCP
port 8888, a machine-wide singleton that a concurrent pytest process takes
away. This adds the two things that would have made it a five-minute
diagnosis instead of a long one, and closes the other door into the same
failure.

Confirmed the module is order-independent as it stands, on this checkout:

  pytest test/ -q, three times          115 failed / 4464 passed / 63 skipped,
                                        byte-identical failure sets, the
                                        module 21/21 passed each time
  module forced last (197 files first)  identical failure set
  module forced first                   identical failure set
  module after each of test_display_manager, test_display_controller,
    test_display_controller_vegas_tick, test_skin_system, test_sports_scroll,
    test_initial_update_budget, test_display_double_parity,
    test_initializing_screen                                     all pass
  four concurrent processes on the file                          21/21 each

And reproduced the original, to be sure the diagnosis in #562 is the whole
story. Holding 0.0.0.0:8888 from a separate process:

  HEAD's test/conftest.py        21 passed
  pre-#562 test/conftest.py      15 failed, 6 passed

The 15/6 split is not arbitrary: the six survivors are the only tests in the
file that never touch dm.matrix.

conftest.py: DisplayManager is a process-wide singleton and the RGBMatrix /
RGBMatrixOptions names it constructs through are module globals, bound once at
import. All three are shared by every test module in the run, so a module that
leaves an instance in _instance -- or leaves patch('src.display_manager.
RGBMatrix') standing -- changes what the NEXT module builds, invisibly, and
only in a full run. A module-scoped autouse fixture now resets the singleton
and restores either binding if a patch outlived its module. Module-scoped
rather than per-test so that files sharing one manager across their own tests
keep doing so; only the leak across the module boundary is cut. Autouse
fixtures are set up ahead of requested ones, so this is finalised after a
module's own DisplayManager fixture. Verified with a throwaway pair of probe
modules -- one leaks a patch and a singleton, the next asserts both are clean
-- which passed and were then removed.

test_display_dirty_tracking.py: _setup_matrix() swallows every construction
failure and falls back to matrix=None, so a broken environment arrived as
fifteen identical "'NoneType' object has no attribute 'SwapOnVSync'" errors
naming neither the fixture nor the cause. The fixture now fails once, and
says where to look; under a held port it reads

    DisplayManager fell back to matrix=None: RGBMatrix construction raised...
    Known causes: the emulator adapter losing a fixed TCP port to another
    process -- see pytest_configure in test/conftest.py -- or a
    patch('src.display_manager.RGBMatrix') leaked from an earlier test module.

with WinError 10048 in the captured log directly above it.

No regressions: full suite with both changes is 115 failed / 4464 passed /
63 skipped, failure set identical to the pre-change baseline. The 115 is the
pre-existing Windows-environment baseline (os.geteuid, POSIX modes, fcntl);
CI on Linux remains authoritative.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 16:13:38 -04:00
ChuckandClaude Opus 5 59997594ac test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths

Two shared-state problems, both of which show up as a permanently dirty
checkout or an unreproducible test failure.

test_display_dirty_tracking.py builds a real DisplayManager, whose
_snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web
UI reads. Every pytest process on the machine shares that one file, so two
concurrent runs -- CI shards, a second worktree, an agent running the suite
alongside -- overwrite each other's snapshot and the mtime assertions stop
meaning anything. The module fixture now points it at a session-unique temp
path; the individual tests that care still override it further.

To be clear about what this does and does not fix: this is a real shared-path
hazard, but it is NOT the cause of the intermittent 15-test failure in that
module. That turned out to be the emulator's fixed TCP port, fixed in the
follow-up commit. This change stands on its own merits.

web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json
and data/operation_history.json as the web interface runs, into a directory
that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig
that ever opened the web UI -- and every test run that constructs the app --
left three untracked files behind and a permanently dirty `git status`. Only
data/.gitkeep is tracked under data/, so the negation keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop the emulator binding a fixed port, so concurrent runs can't collide

This is the cause of the intermittent full-suite failures we have been chasing:
runs of identical code landing anywhere between 100 and 130 failures, while
every implicated test passed in isolation.

Six test modules set EMULATOR=true and build a real DisplayManager. The repo's
emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to
serve the dev preview. That port is a machine-wide singleton, so a second pytest
process -- a CI shard, another worktree, an agent running the suite alongside --
loses the bind. RGBMatrix construction then raises, DisplayManager catches it and
falls back to `self.matrix = None`, and every test that subsequently touches the
matrix dies with

    AttributeError: 'NoneType' object has no attribute 'SwapOnVSync'

which names neither a port nor a socket, and points at the wrong file entirely.
Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its
matrix-touching tests fail together or not at all -- the 15-test swing that made
the totals look random.

Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process
and running test_display_dirty_tracking.py:

    without this change    15 failed, 6 passed
    with this change       21 passed

The "raw" adapter renders in memory and binds nothing. Only display_adapter is
overridden, in a throwaway config written per pytest process; the repo's
emulator_config.json is untouched and `run.py -e` still opens the browser
preview on 8888. Nothing in the suite referenced the adapter, and the tests
wrap SwapOnVSync on the matrix object itself, so they are indifferent to what
sits underneath. allow_adapter_fallback is forced off -- falling back would
land us on the browser adapter and its fixed port, which is the whole problem.

CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set
to an absolute path: the previous behaviour depended on where pytest was invoked
from, and silently wrote a default config into whatever directory that was.

Verified no regressions: full suite on this branch and with origin/main's
versions of the touched files, same machine, back to back -- 115 failed /
4347 passed on both sides, zero failures unique to either. That 115 is the
pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell
scripts, Linux-only binaries); CI on Linux remains authoritative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: mark the shell entry points executable

Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh`
fails with "Permission denied" and only works if you know to prefix `bash`.
That one matters most: the web UI's own error hint, added in #560, tells users
to run exactly that path when a system action fails for want of passwordless
sudo, and following that instruction verbatim did not work.

All eleven carry a shebang and are invoked directly, never sourced. The two
sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately
left non-executable, which is what distinguishes a library from an entry point.

Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied
with `git update-index --chmod=+x` because this checkout is on Windows, where
core.fileMode is off and the working-tree bit is not tracked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:45:13 -04:00
ChuckandClaude Opus 5 f6367d63ae security: triage the CodeQL backlog — 129 alerts, three of them live (#561)
* fix(web): escape quotes in every HTML escaper, not just & < >

The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:

    value="${escapeHtml(v)}"   title="${escapeHtml(v)}"

so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.

It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.

app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.

cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `&#39;`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.

url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).

test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): stop request-supplied names from reaching paths outside their base

Three of the py/path-injection alerts were live, not lint:

* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
  whose resolved path *string-prefixed* the plugin directory. Flask's
  default converter forbids a slash but not dots, and
  get_plugin_directory('..') returned the parent of the plugins directory
  because it exists -- so every file under the project root then prefixed
  that directory, config/config_secrets.json included. The prefix check
  was also wrong on its own terms: with plugin dir "plugin-repos/foo",
  "../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
  start with "plugin-repos/foo".

* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
  body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
  file_id of "../../../../etc/something" deleted that file. This is the
  one finding in the batch that destroyed data rather than exposing it.

* POST /api/v3/cache/delete passed the body's key through
  CacheManager.clear_cache to DiskCache, which joined it as a filename and
  called os.remove. Same shape, same result. The guard goes in
  DiskCache.get_cache_path, the single choke point get/set/clear share, so
  every caller is covered rather than just this route. Real keys are the
  stems of files already flat in the cache directory -- that is how
  list_cache_files derives them -- so nothing legitimate is turned away.

The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.

Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.

test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): refuse a plugin id that is not a plain name, don't truncate it

pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.

Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(web): say what the plugin web_ui iframe actually is

The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.

Also drops the now-unused os/os.path imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): inline url-input's scheme guard at the previewLink.href sink

CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.

Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.

Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): address CodeRabbit findings on the CodeQL triage PR

- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
  credential-saving or connect flow runs. NetworkManager only accepts
  printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
  was previously saved/attempted before nmcli itself rejected it.

- web_interface/blueprints/api_v3/config.py: fail closed when the
  plugin config schema path can't be resolved under the plugins
  directory (e.g. a symlinked plugin dir). Previously this fell
  through with secret_fields left empty, so submitted credentials for
  that plugin were saved as ordinary, unencrypted configuration.

- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
  splicing the JSON day/column key into an inline oninput="..." handler
  string. escHtml() escapes quotes for a normal HTML attribute, but the
  browser HTML-decodes the attribute before running it as script, which
  undoes that escaping and lets a crafted column name (e.g. from an
  uploaded JSON file) break out of the JS string and execute. Cell
  edits now travel through data-day/data-col attributes read by one
  delegated 'input' listener instead.

  While in this file: fixed 6 pre-existing missing-')' typos on
  multi-line safeSetHTML(...) calls (already flagged by Biome in this
  PR's own CodeRabbit run as syntax errors blocking its lint pass).
  These predate this PR (present on main too) but made the whole file
  fail to parse in any JS engine, which is a bigger problem than the
  XSS finding itself and directly touches the same lines.

Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:32:03 -04:00
ChuckandClaude Opus 5 5137e86d16 feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab

PR #544's change, ported onto the api_v3 package split (#553). Identical
behaviour; only the placement of the new code differs.

The original added 508 lines to web_interface/blueprints/api_v3.py, which #553
deletes, so every hunk of it would conflict irreconcilably. Ported by AST:
26 new top-level items sorted to where the split puts each kind --

  __init__.py   2 imports, 11 constants, 7 helpers
  starlark.py   4 routes  (/starlark/editor/{apps,status,start,stop})
  misc.py       2 routes  (/integrations/mqtt-bridge{,/config})

Everything outside api_v3.py -- the Tools partial, the installer scripts, the
JS tests -- applied unchanged.

Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is
regenerated to match, which is exactly what test_api_v3_url_map.py is designed
to make you do when routes are added.

Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR
ships could not be run here -- node is not installed on this machine -- so
test/js/dom/test_tools_sections.js is unverified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST

The AST-based port of #544 onto the api_v3 package split dropped `time`
from starlark.py's import list. start_pixlet_editor() and
stop_pixlet_editor() both call time.time()/time.sleep() directly, so
every start (NameError building `state['started_at']`) and every stop
that has to wait out the EXIT trap crashed with a 500. No test caught
it because the route's own tests mock subprocess.Popen but never
actually invoked it before now.

Also carries over #544's later fix that this port branched before:
env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an
operator who had already pinned PIXLET_EDITOR_HOST to loopback,
forcing the unauthenticated `pixlet serve` process onto the LAN
regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as
api_v3.starlark.py's siblings already do for _pkg-owned names.

Both fixes route the shared _pkg.time reference the rest of the
package's route modules already use for anything a test might need to
patch, rather than a bare `import time` local to this file.

Ported the existing regression test from #544
(TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's
module layout (web_interface.blueprints.api_v3.starlark instead of the
old monolithic api_v3 module), which is what caught the NameError.

Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this
branch and on origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to
this change).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): clear the six lint errors this rebase introduced

All six were introduced by rebasing this branch onto the merged blueprint
split, not by the split itself. Confirmed by diffing pyflakes output against
main with line numbers normalised -- everything else it reports is present on
main too and is the package's deliberate re-export pattern.

starlark.py used _STARLARK_APPS_DIR three times without importing it (F821).
The rebase resolved an import-list conflict as a union of both sides, and that
symbol was on neither side of the conflict hunk, so it was silently lost. It is
defined in __init__.py and is now imported like its neighbours. This was the
only one of the six that would fail at runtime rather than merely lint.

__init__.py imported contextlib twice (F811): the cherry-pick added one next to
the existing import. Removed the duplicate; the original at line 19 is used.

__init__.py imported signal purely to re-export it to starlark.py, so pyflakes
saw it as unused (F401). signal is stdlib and does not need routing through the
blueprint package, so starlark.py imports it directly and __init__.py no longer
does. contextlib stays re-exported because this module genuinely uses it.

_read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule
this module imports at the bottom for its route side effects (F811). Renamed to
`settings`, with a comment saying why, since the name is otherwise the obvious
one to reach for.

Verified: pyflakes now reports nothing on this branch that main does not, the
package imports, all nine route modules load, and 117 routes register, matching
the pinned URL-map snapshot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply

Two CodeRabbit findings on the bridge settings endpoint, both of which returned
200 while doing something other than what the caller asked.

`request.get_json(silent=True) or {}` turned a missing or unparseable body --
and the JSON literals null, [] and false -- into an empty dict, which then
satisfied the isinstance(data, dict) guard on the very next line. The guard was
there to reject exactly those bodies. Dropping the `or {}` lets None fail it.

The same `or {}` on /errors/clear is left alone: its docstring documents the
body as optional, so an absent body legitimately means "use the defaults". The
difference is that saving settings has nothing sensible to do with no body.

`if data.get('clear_password'):` accepted any truthy value, and the string
"false" is truthy in Python -- so a client echoing the field back as a string
wiped a password it meant to keep. Now coerced through the package's existing
_coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else.

test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus
a missing body, and clear_password across truthy and falsy spellings. Verified
against the unfixed code -- reverting the body guard fails 5, reverting the
coercion fails 3.

Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials
when TLS is off (CWE-319). That is a policy decision about the feature rather
than a defect -- unencrypted MQTT on a trusted LAN is common and often
deliberate -- so it is raised on the PR for a maintainer call instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: work through the remaining review findings on the editor and bridge

allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses
the network in cleartext. Refused now rather than merely warned about -- but
refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal
deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults
false, and is coerced like the other booleans so the string "false" cannot
switch the guard off.

starlark.py:796 -- the supported service runs Flask threaded, so two start
requests could each see running=False, each launch an editor, and the second
state write replace the first PID, orphaning a process that holds the display
down with nothing recording it. The check-launch-write sequence now takes a
module-level lock.

starlark.py:848 -- if the state write failed the route returned success with an
editor running and no PID recorded: status and stop both reported no session
while the display stayed down until the timeout expired. It now terminates the
process group and returns an error.

starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so
nothing hands the display back, yet the response said "the display is
restarting". After an escalation the display is now restarted explicitly, and a
failure to do so returns an error naming the manual step instead of a success.

pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no
timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the
failure lands before the display is stopped rather than after.

pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface
address are three cases, not two. Collapsing the last two printed a URL saying
"localhost" whenever PIXLET_EDITOR_HOST named a LAN address.

tools.html:1254 -- escHtml does not encode single quotes, and the app id was
interpolated into an inline onclick="startPixletEditor('...')", so a directory
containing an apostrophe could break out of the JS string and run script. The
handler binds with addEventListener and reads the id from dataset, where it is
only ever parsed as an HTML attribute.

Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in
in both directions. The tools DOM suite gains three guards asserting the edit
buttons carry no inline onclick and pass the id via dataset -- those need jsdom
and did not run here, so CI verifies them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): log the traceback on the editor state-write failure

The 848 fix answers 500 when the session state cannot be written, and logged
that at error level -- but without exc_info, so the traceback never reached the
log. test_web_error_detail.py guards exactly this: a handler returning 5xx must
write an error-level record *with* the traceback and return the sanitized
detail, because checking that merely something was logged is too weak.

Caught by Core unit tests on the previous commit, not locally: the guard parses
every module under web_interface/blueprints/api_v3 as one source, so it only
fires once the whole package is read together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 11:32:12 -04:00
ant456 a3d505384d Add render_width/render_height support to Starlark Apps (#552) 2026-09-11 11:22:52 -04:00
ChuckandClaude Opus 5 bdb9a94033 refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package

web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181
functions in one module -- 9% of the core by line count and three times the
next largest file. It becomes a package of nine route modules grouped by path
segment, plus __init__.py for the shared imports, constants, Blueprint and
helpers.

Every route module decorates the SAME api_v3 Blueprint object, so endpoint
names stay api_v3.<function>, the URL map is unchanged and app.py is untouched.
Verified: 111 routes before, 111 after, byte-identical rules, endpoints and
methods, and every endpoint still on the one blueprint.

  plugins   3,867   config    1,178   starlark  692   system  619
  fonts       452   misc        398   wifi      361   display 326   backup 212
  __init__  1,787 (imports, constants, Blueprint, 56 helpers)

Two things the URL-map check could not catch, both found by running the suite:

1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one
   directory deeper made that resolve to web_interface/ instead of the project
   root. Nothing failed at import; it surfaced as ~110 tests failing with 404s
   and "installation script not found", because every path built from it was
   one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts
   PROJECT_ROOT/run.py exists so the next move cannot repeat it.

2. Module-attribute patching. Tests do
   monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route
   module that binds such a name by value never sees the patch. The shared code
   therefore stays in __init__.py rather than moving to a _common submodule --
   it has to live on the module the tests patch -- and the eleven names tests
   patch are read back through the package (_pkg.X) instead of bound by value.
   Those eleven were found by AST-scanning every setattr in the test tree, not
   by guessing; "time" is among them, used to drive a fake clock through the
   second-resolution credential-backup filenames.

Test changes are confined to what genuinely moved: patch targets that now name
the owning route module, imports of helpers, and six tests that scan the api_v3
source as a file and now read the package directory.

Full suite: 4,278 passed, 68 skipped, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(api-v3): address CodeRabbit findings from the blueprint-split review

Fixes to the api_v3 package split (PR #553), one per finding verified
against the actual code:

- __init__.py: _redact_credentials only blanked scalar values under a
  credential-named key; a bare list of secrets under such a key (e.g.
  tokens: ["a", "b"]) passed through untouched, since the list branch
  recursed with no memory that its key looked like a credential. Nested
  dicts still walk normally (a documented, tested behaviour -- a container
  like secrets: {api_key: ..., note: ...} is a section name, not a value to
  blank outright), but any value reached under a credential-shaped key is
  now actually blanked.

- __init__.py: the OAuth helper script's raw stderr/stdout went to
  logger.error unredacted (CWE-532) right next to a comment claiming this
  was deliberate; the HTTP response already used the existing redact_text
  helper. Routed the log line through the same helper.

- __init__.py / starlark.py: the standalone Starlark manifest fallback
  (used when the plugin instance isn't loaded) read-modified-wrote
  manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe
  (plugin-repos/starlark-apps/manager.py), which already holds an flock for
  the same file when the plugin is loaded. Added _starlark_manifest_lock,
  mirroring that pattern, and wrapped every standalone read-modify-write
  call site in it. The app-config update route also wrote config.json and
  the manifest as two separate, non-transactional writes (a second,
  distinct finding at the same call site); config.json is now rolled back
  if the manifest write that follows it fails.

- backup.py: restore options used bare bool() on values from the request,
  so {"restore_secrets": "false"} restored secrets anyway (bool("false") is
  True). Switched to the existing _coerce_to_bool helper already used for
  this exact purpose elsewhere in the package.

- config.py: an automated import-rewrite mangled four user-facing
  validation strings and their neighbouring comments -- "Invalid start
  time" had become "Invalid start _pkg.time" (and likewise for "end time")
  in both the schedule and dim-schedule per-day validation paths.

- display.py: `import _pkg.time as time_module` -- _pkg is a local alias
  for the package, not a real importable module, so this raised
  ModuleNotFoundError whenever a caller restarted an already-running
  display service via /display/on-demand/start, after the on-demand
  request was already written to cache. Fixed to `import time`. Audited
  the rest of the package for the same `_pkg.<module>` import mistake;
  every other `_pkg.` reference is a legitimate attribute read-through
  (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken
  import statement.

- fonts.py: validate_file_upload's max_size_mb parameter is silently
  unused by that helper (it only checks filename/extension) -- the font
  upload route saved arbitrarily large files as a result. Added the same
  seek-and-check pattern already used for the sibling .star upload.

- wifi.py: two ad hoc, inconsistent bool coercions. POST
  /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled
  it. POST /wifi/radio's enabled/force parsing recognized real bool and
  some strings but not int 1/0 (1 is True is False in Python). Factored one
  small _parse_bool_ish helper local to this file and used it at all three
  sites.

Not changed: the "unknown/misspelled restore option keys default to True"
half of the backup.py finding -- the file's own comment documents that a
missing key deliberately means "restore everything," matching the
already-existing JSON-parse-failure guard a few lines above it; only the
bool-coercion defect was a real bug.

Added or extended regression tests for every fix, following each area's
existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed
on both this branch and origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to this
change) -- no new failures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5

* fix(api-v3): reject unknown restore option keys

CodeRabbit's review of the blueprint split (#553) asked that
POST /backup/restore reject option keys outside RestoreOptions'
known set. The follow-up commit fixed the bool("false")-is-True
bug with _coerce_to_bool but never added the key check: a typo'd
or renamed key (e.g. "restoreSecrets") is silently ignored by
opts_dict.get(key, True), so the flag stays at its True default
and secrets get restored despite the caller's request saying
otherwise -- with no indication anything was wrong.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb

* fix(api-v3): address CodeRabbit findings on the blueprint split

- _redact_credentials: blank scalar descendants of objects reached
  through a credential-owned list (e.g. tokens: [{"value": "secret"}])
  regardless of field name -- the existing name-based walk only
  protected direct dict values under a credential key, not list items.
- wifi.py: reject enabled/force/auto_enable_ap_mode values
  _parse_bool_ish can't recognize (400) instead of silently treating
  them as False, which could disable Wi-Fi or the radio itself.
- Starlark manifest locking: lock a stable manifest.json.lock sidecar
  instead of manifest.json itself, in both the standalone route path
  (_starlark_manifest_lock) and the plugin path
  (StarlarkAppsPlugin._save_manifest / _update_manifest_safe).
  manifest.json is replaced by an atomic rename on every write, which
  swaps in a fresh inode; a lock held on the old inode does not
  exclude a second locker that opens the path afresh right after the
  rename and gets the new inode, so two writers could race despite
  each holding "a lock". A sidecar that no write ever touches always
  resolves to the same inode for every locker.

Skipped as stale: the "serialize the complete manifest
read-modify-write" finding at api_v3/__init__.py -- every standalone
handler that calls _write_starlark_manifest is already wrapped in
_starlark_manifest_lock() on this branch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): re-check reconciliation findings by the reconciler's own rules

Both CodeRabbit findings on the merge commit, verified against the code first.

Major, plugins.py: the stale-findings filter derived its own notion of "in
config" and "on disk", and both were looser than the reconciliation module's.
set(load_config()) also contains system keys, the secrets-file keys load_config()
merges in, and non-dict values; and any directory holding a manifest.json
counted as installed even when that manifest does not parse. Either looseness
clears a finding that is still true -- and a secrets key read as a plugin is the
precise bug the filter exists to stop reporting, so reintroducing that asymmetry
while re-checking was the wrong way round.

The two extractions now live in state_reconciliation.py as config_plugin_ids()
and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys()
alongside. _get_config_state() and _get_disk_state() use them too, so there is
one definition rather than two that can drift. _get_disk_state() re-reads each
manifest for version/name after taking membership from the shared extractor;
that costs one extra small read per plugin on a path that runs once per boot.

Minor, the new test: the fixture assigned api_v3.config_manager and
api_v3.plugin_manager directly. Those live on a module-level blueprint
singleton, so the mocks leaked into every later test that imports api_v3 --
pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr,
which restores them. This is the same pollution class that made an earlier test
in this session break seven unrelated ones, so it is worth getting right.

Five cases added for the parity itself: a secrets key, a system key and a
non-dict value must not clear an "installed but missing from config" finding,
and neither an unparseable manifest nor a .standalone-backup- directory may
count as installed. All five fail against the looser version.

Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and
CodeRabbit all pass. Codacy reads action_required on every commit of this
branch including the first, so it is pre-existing and not from this work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 10:07:38 -04:00
ChuckandClaude Opus 5 0ab95586fb fix(web): say when a system action failed for want of passwordless sudo (#560)
* fix(web): say when a system action failed for want of passwordless sudo

POSTing reboot_system to a Pi returns, in full:

  {"message": "Action failed; see logs for details", "status": "error"}

The cause is that the web interface runs unprivileged, and its
systemctl/reboot/journalctl calls only work once
scripts/install/configure_web_sudo.sh has granted NOPASSWD. first_time_install.sh
never invokes that script and no user-facing doc mentions it, so on a fresh
device every privileged action fails -- start_display, stop_display, the
autostart toggles, reboot, and the log viewer.

That last one closes the loop: "see logs for details" is unreachable advice
when journalctl is refused for the same reason. This is exactly the failure
src/web_interface/error_handler.py's describe_exception() was written to break,
and /system/action's exception handler was still discarding the cause instead
of using the helper the module already imports.

Two changes, no behaviour change when things work:

- The exception path now returns 'details': describe_exception(e), matching how
  the other handlers in this blueprint already report.
- A failure whose stderr or exception text is sudo refusing to prompt ("a
  password is required", "no tty present", "a terminal is required") reports
  what to do about it, naming configure_web_sudo.sh. Unrelated failures keep
  the generic message and their stderr, so a missing unit is not blamed on
  sudo.

Granting the sudo rights is left alone deliberately: auto-running a script that
hands out NOPASSWD is a security decision for the maintainer, not something to
slip into an installer. Making the refusal legible is the part that is
unambiguously an improvement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): apply the sudo hint on the on-demand start_display path too

start_display with a mode builds its own response and returns before the shared
nonzero-result path, so a recognized sudo refusal there reported only "Failed to
start display" and said nothing about the passwordless sudo that refused it --
the exact gap the rest of this PR closes everywhere else.

Raised by CodeRabbit on #560 and verified against the code before fixing: the
branch at api_v3.py:2058 does return early past the shared handler.

Three regression cases: the on-demand branch reports the sudo cause, keeps its
"Display started" message on success, and does not blame an unrelated failure on
sudo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:45:40 -04:00
ChuckandClaude Opus 5 39f27d285d fix(plugins): stop reconciliation inventing plugins and telling users to delete real config (#557)
On a device running four installed, configured, working plugins, the overview
banner read:

  Stale plugin config entries found: football-scoreboard, odds-ticker, data,
  ledmatrix-weather, starlark-apps. Remove them from config.json or reinstall
  via the Plugin Store.

Every claim in that sentence was wrong, and following its advice would have
deleted 4.9KB of working league settings. Four separate defects combined.

1. Secrets keys became phantom plugins. load_config() merges
   config_secrets.json into the config it returns, and the ignore list named
   only 'github' and 'youtube'. A 'data' key in that file therefore read as a
   plugin id and was reported as "in config but not on disk" forever. Read the
   secrets file's own top-level keys instead of hardcoding two of them.

2. The auto-fix clobbered real config. The handler for "on disk but not in
   config" assigned `config[plugin_id] = {'enabled': False}` unconditionally,
   so whenever detection was wrong it replaced a plugin's entire configuration
   with a stub. On the reported device it only failed to do so because the
   write hit EACCES. Now it refuses to overwrite an entry that already exists.

3. The banner gave backwards advice. plugin_missing_in_config ("on disk, not in
   config") and plugin_missing_on_disk ("in config, not on disk") are opposite
   problems, and both were rendered as "stale config entries ... remove them
   from config.json" -- which is correct for the second and destructive for the
   first. They are now reported separately, each with the advice that fits.

4. A stale verdict was served indefinitely. The result is a snapshot written
   once per run to a status file, and a run that fails to apply a fix also
   declares it will not retry. A condition that had since resolved kept being
   reported for hours. The status endpoint now re-checks stored findings
   against current state, dropping only what it can prove stale and keeping
   any kind it cannot re-verify.

The secrets-key lookup is deliberately fail-safe: an unreadable, absent,
malformed or non-path secrets location narrows the ignore set rather than
raising. An earlier revision let TypeError escape, which the broad handler in
_get_config_state() swallowed as "Error reading config state" -- emptying the
config state and making every downstream detection wrong. The existing
reconciliation tests caught it; there is now a regression test for it too.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:45:21 -04:00
ChuckandClaude Opus 5 aba96e25b3 chore: delete three functions nothing calls (#550)
src/base_classes/baseball.py       _get_baseball_display_text   45 lines
  src/web_interface/api_helpers.py   validate_request_params      22
  web_interface/blueprints/api_v3.py _validate_time_range         14

Each has exactly one occurrence across both repositories -- its own
definition. No decorator, no __all__, no getattr dispatch, nothing in
templates or JavaScript.

A fourth candidate was dropped after checking: _unshare_element_fonts in
src/common/sports_shared.py looked unreferenced, but eight scoreboard plugins
call SportsCore._unshare_element_fonts directly from their
test_element_text_colors.py, plus their own copies at runtime. It is live API.
The earlier reading came from a plugins checkout 84 commits behind main, which
is a good argument for re-verifying this kind of claim against a fresh tree
rather than trusting an earlier scan.

Full suite: 4,265 passed, 68 skipped.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:45:07 -04:00
ChuckandClaude Opus 5 577f5501a6 perf(plugins): stop re-deriving a display() signature the caller already cached (#549)
display_controller resolves once, and caches, whether a plugin's display()
takes a display_mode keyword -- self._plugin_accepts_display_mode, populated
right before the dispatch. It then handed the executor a
types.SimpleNamespace wrapping a closure, and execute_display() ran
inspect.signature() on that to work out the same thing.

Because the SimpleNamespace is rebuilt per call, the callable was new every
time, so nothing inside the executor could ever cache it either. Measured at
~39us per dispatch on a Pi 4, for a value the caller had a line earlier.

execute_display() now takes accepts_display_mode, falling back to inspecting
only when a caller does not pass it, so existing callers are unaffected.

Also documents two things that read as bugs and are not:

- execute_with_timeout()'s timeout is advisory. Nothing cancels the thread --
  Python cannot -- so on expiry the operation runs to completion in the
  background and only the caller gives up. A permanently hung plugin leaks a
  daemon thread per attempt. This is why callers holding a lock across the
  call must release it from inside the wrapped callable, as run()'s
  _release_display_lock already does.

- Only the first display() of each mode goes through the executor; the
  per-frame loops call display() directly. That is deliberate: a thread per
  frame would cost more than an advisory timeout buys. Both loops now say so,
  so the asymmetry does not read as an oversight.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:42:49 -04:00
ChuckandClaude Opus 5 dcd6e39c96 fix(web): report real disk usage and MemAvailable on the live status stream (#558)
The SSE status stream sent 'disk_used_percent': 0 as a literal, so every
consumer of the live view showed 0% disk no matter how full the card was.
/api/v3/system/status computed it correctly; the stream that the dashboard
actually watches did not. On a Pi with a modest SD card that is the warning a
user most needs, and it was guaranteed to never appear.

The stream also omitted memory_available_mb. /api/v3/system/status carries it
with a comment spelling out why it matters: MemAvailable accounts for
reclaimable page cache, so it is what separates a board reading 70% "used" that
is fine from one reading 70% that is about to fail fork(). A 1GB Pi 3B+ can sit
at either. The number that predicts the failure was missing from the live view.

An unreadable disk now reports None rather than 0. The UI already renders null
as '--'; a confident 0 reads as "plenty of room", which is worse than a blank.

Metric collection moves to web_interface/system_metrics.py, with no Flask or app
imports. That is not cosmetic: importing web_interface.app constructs the Flask
application and a CacheManager, and the latter claims the cache directory with a
cleanup thread. The first version of these tests imported the generator directly
and broke test_cache_cleanup_thread_ownership ("one thread per directory") plus
four starlark route tests through that side effect. Reading a CPU percentage
should not boot a web application, and testing it should not either.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:42:30 -04:00
ChuckandClaude Opus 5 ad5bc4b819 perf(sports): LRU-bound the decoded logo cache (#559)
SportsCore._logo_cache was a plain dict keyed by team abbreviation with no
eviction. Its entries are not file bytes but decoded RGBA thumbnails sized to
display*1.5 -- roughly 36KB on a 256x64 panel, more for wide wordmarks -- and
assets/sports/ncaa_logos ships 307 of them. A plugin that walked a full league
held the whole league resident: about 11-18MB per manager instance, and a league
runs three (live/recent/upcoming) that each keep their own cache, so the same
logos were duplicated across them.

On the 1GB Pi 3B+ this was measured on, one board was sitting at 439MB resident
with ~290MB available, so tens of megabytes of duplicated league logos is real
money. Bounded to 64 entries, which holds a full "other games" cycle (on the
order of 20 games, 40 teams) without thrashing while capping the cache well
below a 307-team league.

Eviction is LRU rather than clear-when-full, using the OrderedDict/popitem
pattern the neighbouring caches in this codebase already use (_IMAGE_CACHE_MAX,
_FIT_CACHE_MAX, _TEXT_WIDTH_CACHE_MAX). That ordering matters: the logos on
screen right now are precisely the ones that must not be discarded, so a cache
hit moves the entry to the end.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:42:18 -04:00
ChuckandClaude Sonnet 5 fb3b293ace fix(plugins): let a plugin ask to be polled faster while it has live content (#555)
* fix(plugins): let a plugin ask to be polled faster while it has live content

Reported: "the football plugin with live games only updates the live game in
progress if I restart the display."

The data path was never the problem. NFLLiveManager fetches ESPN with no cache,
SportsLive.update() refreshes current_game in place when the game IDs are
unchanged, and the scorebug redraws from the game dict every frame -- which is
why the reporter's logs look healthy.

The problem is cadence. _get_plugin_update_interval() read only the manifest's
static update_interval, football's manifest pins that to 60, and the plugin's
own live_update_interval (15s) was invisible to the scheduler. Measured on a rig
during the fourth quarter of the game in the report:

    23:21:49  23:22:50  23:23:50  23:24:50  23:25:50   <- exactly 60s apart

A clock and score up to a minute stale during a two-minute drill reads as a
frozen panel, and a restart is the one moment it is ever current.

A single static number cannot say "every 15 seconds while a game is on, every 15
minutes in July", and only the plugin knows which is true. get_update_interval()
lets it say so per tick; returning None means "no opinion" and the existing
manifest/config resolution applies, so every plugin that predates this is
unaffected.

Requests are clamped to MIN_DYNAMIC_UPDATE_INTERVAL (5s): a plugin returning 0
would otherwise be re-entered on every tick of the render loop, busy-waiting
against its own API. A hook that raises or returns a non-number is ignored
rather than propagated -- a scheduler that fails on one plugin's bug stops
updating all the others.

Deliberately NOT changed: the manifest still beats config in the static path.
That looked like the obvious fix -- user config being silently ignored -- until
checking a real rig, where football and baseball both carry update_interval 3600
in config against a manifest 60, and weather 1800 against 60. Those values are
stale precisely because nothing has been honouring them; making config win would
have slowed three plugins by 60x, turning a one-minute lag into an hour. The
dynamic hook makes the flip unnecessary. There is a test pinning the current
precedence with that reasoning attached.

Full suite: 4,283 passed, 68 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(plugins): drive the real scheduler, not just the interval resolver

test_plugin_dynamic_update_interval.py asserts that
_get_plugin_update_interval() returns the number the plugin asked for. That is
not the same claim as "the plugin gets updated more often", and the gap between
those two is exactly where the original bug lived: the plugin knew it wanted
15s, said so in live_update_interval, and nothing downstream acted on it.

So this ticks the real run_scheduled_updates() through a simulated hour and
counts dispatches. Against pre-fix core it reports "10 updates in 10 minutes of
a live game" -- the 60s manifest cadence, matching what was measured on a rig
during the reported game. Against the fix it reports ~40.

Also pins the regression that would be worse than the bug: an idle hour must
still be ~60 updates, not 240. Asking for the live interval year-round would
poll ESPN four times a minute all summer.

Scope note, since it is easy to over-read this fix: the *switch* display path
already refreshed the manager immediately before drawing, via
_try_manager_display() -> _ensure_manager_updated(), which honours the manager's
own 15s interval. So a switch-mode card was already <=15s stale at draw time
before this change. What this fixes is the background cadence, which is what
live-priority detection, Vegas content and scroll preparation all read.

Full suite: 4,288 passed, 68 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): reject bool and -inf hook results in dynamic interval

get_update_interval() ran bool through float() (bool is an int subclass,
so True/False became 1.0/0.0) and only checked for +inf, not -inf. Both
cases landed on the MIN_DYNAMIC_UPDATE_INTERVAL floor by coincidence
instead of falling back to the static/manifest interval as invalid
input should.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Tst9cied2ri9bH4QRWa6H

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:58 -04:00
ChuckandClaude Opus 5 8da13f02f8 chore: ignore team logos fetched at runtime (#551)
logo_downloader.py and LogoHelper write into assets/sports/<league>_logos/
whenever a plugin meets a team whose logo is not on disk. Those directories are
also tracked -- 209 NCAA logos and 153 soccer ones ship with the repo -- so
every rig accumulates untracked files nobody intended to commit. This checkout
had 62; hdpi shows the same.

The cost is not the files, it is that a permanently dirty `git status` trains
everyone to ignore the one signal that says a checkout is not what you think it
is. That is how a stale tree sat unnoticed on a rig for hours until a restart
surfaced four sports plugins that could no longer import.

Ignoring a directory does not untrack what is already in it, so the logos that
ship keep shipping -- verified: 209 and 153 still tracked, no deletions in the
diff. Only new downloads are hidden.

Adding a logo on purpose stays possible and is what the escape hatch in the
comment documents. It is also rare: the last deliberate addition was #415, four
named NCAA logos a plugin needed, and `git log` finds no other in a year. So the
common case is noise and the rare case is explicit, which is the right way round.

Untracked files: 62 -> 0.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:44 -04:00
ChuckandClaude Opus 5 28bc79566f fix(logo): remember a missing logo instead of re-warning every rotation (#548)
* fix(logo): remember a missing logo instead of re-warning every rotation

load_logo() stat'd the path and logged a WARNING on every call, and the
positive cache never covered it because a miss returns None and caches
nothing. A file that is simply not there therefore produced one warning per
rotation for as long as the process ran -- measured on a live rig at 114 lines
in 24 hours for a single missing ticker icon, for a file nobody was going to
add.

Misses are now remembered for 10 minutes: warn once, then return None without
touching the disk. Bounded rather than permanent because logo_downloader
writes logos at runtime, so a file that appears later must still be picked up
without a restart. Downloads through load_logo_with_download() clear the entry
outright -- load_logo() consults the miss record before it stats the disk, so
without that a freshly downloaded logo would stay invisible for the whole
window.

This is in the core rather than in ledmatrix-stocks, where it was found, so
every plugin that goes through LogoHelper gets it.

_cache_order stays a list. Swapping the pair for an OrderedDict would shave an
O(n) scan per cache hit, but n is capped at cache_size (100 by default) and
test_logo_helper.py pins the current structure; not worth the churn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(logo): make the miss TTL longer than the rotation it is meant to outlast

Deployed the previous commit to a live rig and measured it: no change at all.
"Logo not found for VOO" stayed at ~6 lines an hour, exactly the baseline.

The TTL was 600s and the display rotation is ~618s, so every recheck expired
just as the plugin came round again and the negative cache never once got to
suppress a warning. The fix was correct in shape and useless in practice,
which only measuring on the rig would show.

An hour instead. That is safe because the TTL is not the main way an entry
clears: load_logo_with_download() drops it the moment a download succeeds and
clear_cache() drops all of them. The TTL only covers a file that appeared some
other way -- someone copying one in by hand -- and waiting up to an hour for
that, or restarting, is a fair trade for not re-warning about a file nobody is
going to add.

The general lesson is in the comment: a TTL has to be long relative to the loop
that does the asking, not merely "a while".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:32 -04:00
ChuckandClaude Opus 5 12f3790994 fix(install): render the systemd units from their templates, not from heredocs (#547)
* fix(install): render the systemd units from their templates, not from heredocs

The installers carried their own inline copies of units that also exist as
templates under systemd/, and the copies drifted.

install_service.sh renders ledmatrix.service from the template correctly, then
wrote ledmatrix-web.service from a heredoc that predated it -- missing
Wants=network-online.target, RestartSec=10, SyslogIdentifier, CacheDirectory,
CacheDirectoryMode and Environment=USE_THREADING=1. install_web_service.sh had
a third copy, and install_wifi_monitor.sh a fourth, that one already differing
from its template (syslog where the template says journal).

startup_validator.py compares the installed unit against the template, so a
rig installed this way warned on every boot -- and the remedy the warning
names, "re-run scripts/install/install_service.sh", reinstalled the same stale
copy. The warning could never clear. Reproduced on a live rig running exactly
that unit.

All three installers now render systemd/*.service through the same placeholder
substitution. The template gains a __USER__ placeholder rather than hardcoding
User=root, because the web interface runs as whoever installed it.

That last point was a second, independent cause of a permanent warning: the
validator substituted a fixed "root", so any non-root install reported drift
forever. It now reads User= from the installed unit -- an install-time
decision, not something the template dictates -- and compares everything else
strictly. first_time_install.sh already reads the installed User= the same way.

Tests cover a non-root web unit not warning, a genuinely changed directive in
that unit still warning, the User= fallback, and a grep-based guard that no
installer under scripts/install/ contains an inline unit body. That guard is
what found the install_wifi_monitor.sh copy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(install): escape sed replacements, use mktemp, and make render failures fatal

Address CodeRabbit findings on install_service.sh, install_web_service.sh and
install_wifi_monitor.sh:

- Values interpolated into each script's sed expression (project root path,
  username) were not escaped, so a value containing &, \ or the | delimiter
  would corrupt the rendered systemd unit. Add a shared
  sed_escape_replacement() helper in the new scripts/install/lib_systemd_render.sh
  (sourced by all three scripts) and apply it to every sed replacement.
- install_service.sh rendered the main and web units to the predictable path
  /tmp/ledmatrix.service.tmp before installing them -- a symlink/TOCTOU race
  (CWE-377). Use mktemp for both, with a trap to clean up on exit.
- install_service.sh treated a missing template as a mere warning and then
  checked only whether a unit already existed at the destination before
  enabling/starting it, so a render failure could silently fall back to
  enabling a stale, previously-installed unit. Both unit blocks now exit
  non-zero on a missing template or a failed render.

Also rename the ambiguous loop variable `l` to `line` in
test/test_systemd_unit_drift.py (Ruff E741); ruff isn't wired into any CI
workflow in this repo today, so this isn't currently CI-blocking, but the
rename is trivial and correct regardless.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5

* test(install): cover sed_escape_replacement against sed-special characters

CodeRabbit asked for regression coverage using a project path containing an
ampersand; the earlier commits on this branch already fixed the escaping,
mktemp usage, and enable/start-on-fatal-render-failure findings, and the
l->line rename was already applied -- this closes the one remaining gap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:19 -04:00
ChuckandClaude Opus 5 2df273ecfc fix(web): default web_display_autostart to true, as the installer already does (#556)
start_web_conditionally.py read the flag with
`config_data.get("web_display_autostart", False)`, so a config that simply
lacked the key got no web interface. Both config/config.template.json and
first_time_install.sh ship the key as true, so the code default contradicted
the shipped default in two places: absence means an older or hand-edited
config, not a request to stay down.

The failure mode was silent in the worst way. The "not starting" path exits 0,
so `systemctl status ledmatrix-web` reported the unit as successfully started
while nothing was listening on the port, and the only trace was one journal
line saying the flag was "false or not set" -- which reads as a deliberate
setting rather than a missing key.

Also start the web interface when config.json is missing or unparseable,
instead of exiting. The web interface is how a config gets created and
repaired, so a broken config is exactly when the user needs it most; leaving
it down means there is no way back in. Only an explicit false/off disables
autostart now, and the disabled message says "explicitly disabled" so the
journal distinguishes a real setting from a default.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 17:42:00 -04:00
ChuckandClaude Opus 5 a29c84208e fix(scroll): advance whole pixels per frame, not per wall-clock second (#545)
* fix(scroll): advance whole pixels per frame, not per wall-clock second

Smooth motion is not a frame-rate property, and measuring it as one is why
this survived three rounds of fixes. odds-ticker's frame timing is excellent
-- 100.0 fps, 10.00ms median, 0% stalls, worst in-scroll frame 19.95ms -- and
it still visibly stuttered.

What the eye judges is whether the strip advances the same number of whole
pixels on every presented frame. update_scroll_position derived position from
scroll_speed * delta_time and get_visible_portion truncated it with int(), so
jitter in delta_time decided which side of a pixel boundary the position
landed on. The live windows show why that matters: a rock-steady 100.0 fps
whose individual frames still range 5.6ms to 15.2ms, which at 100 px/s is
0.57px to 1.44px of movement.

Run the measured frame times through the real helper and 5.8% of frames
advance 0 or 2 pixels instead of 1 -- about six hitches a second. A frame that
moves nothing followed by one that jumps two is exactly what micro-stutter
looks like.

It is worst at a crisp speed, which is the part that stings: at 100 px/s on a
100Hz panel the accumulator sits exactly on integer boundaries, so
sub-millisecond jitter flips it either way and the motion beats at around
50Hz. Snapping to the crisp ladder fixes the average and the wall clock then
throws away the per-frame uniformity the ladder was bought for.

So when scroll_config snaps to a crisp speed it now also puts the helper in
fixed-step mode: each presented frame advances exactly pixels_per_frame and no
clock is consulted. 100% of frames move by the same amount, whatever the
jitter.

This is only correct because SwapOnVSync blocks until the panel has taken the
frame, which makes the frame count a truer clock than time.time(). Before the
swap was locked to vsync it would have run at whatever speed the loop spun at.
Related: frame-based mode used to step discretely and was converted to
elapsed-time accumulation earlier in this series, because its threshold
comparison flipped on jitter. That was right for the code as it stood -- but
it treated the symptom, replacing a broken discrete step with a smooth-looking
accumulator instead of asking why a wall clock was involved at all.

Non-crisp speeds keep pacing off time, and set_scroll_speed() clears the fixed
step so a legacy caller changing speed is not silently ignored.

Trade-off worth naming: speed is now tied to the presentation rate rather than
to real time. If the loop cannot keep up with the panel the scroll runs slow
rather than jumping to catch up. That is the better failure -- uniform motion
at a slightly wrong speed beats correct average speed with a hitch six times a
second -- and a loop that cannot hit the resolved rate is a measurement
problem for the crisp ladder, not something to paper over with uneven steps.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(scroll): make the time-based pin actually pin something

Review caught that test_time_based_stepping_is_what_it_replaces could pass
against perfectly uniform motion, and it was right.

update_scroll_position sets last_update_time on its way through, so the very
first call sees a delta_time of zero and moves nothing in time-based mode.
_advances counted that synthetic frame, which put a guaranteed zero in every
histogram -- enough on its own to satisfy "uneven > 0". The test asserting the
defect exists would have passed after the defect was gone.

The first call is now primed and discarded, and the assertion is a proportion
rather than "more than zero": against these frame times the old path misses
roughly one frame in twenty, so 1% is well below the real rate and far above
anything a stray frame could produce.

Re-measured with the artefact removed, the numbers in the PR description are
unchanged: 5.85% of frames uneven before (114 zero-advance and 120 double
frames in 4000), 0.00% after.

Also fills in the docstrings the review flagged: everything in the new test
file, plus three pre-existing one-liners in scroll_config that the diff
touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:25:45 -04:00
ChuckandClaude Opus 5 1198615d19 test: install PyYAML so the starlark route tests can load the plugin (#546)
Tests has been red on main since #535. All 13 failures in
test/web_interface/test_starlark_pixlet_routes.py are the same
ModuleNotFoundError: No module named 'yaml'.

The test loads plugin-repos/starlark-apps/tronbyte_repository.py by path --
deliberately, "the way the blueprint does", since the core web blueprint
really does exec that plugin module -- and the plugin imports yaml.

Nothing is undeclared. The plugin's own requirements.txt already pins
PyYAML>=6.0.2, and on a real rig the plugin store installs it. CI installs
only requirements.txt and requirements-test.txt, so a core test that reaches
into a plugin gets none of the plugin's dependencies.

PyYAML goes in the test requirements rather than the core ones because it is
not a core dependency: nothing in src/ or web_interface/ imports yaml. This is
the same shape as the psutil entry directly above it -- a package the core does
not require, installed so a test can exercise a real path instead of a stub.

Verified locally: with yaml available the file goes from 13 failures to 64
passing. (One unrelated failure remains on Windows only, where os.geteuid does
not exist.)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:25:30 -04:00
ChuckandClaude Opus 5 f8e2e89edc refactor(sports): put the scoreboards on the shared scroll resolver (#542)
* refactor(sports): put the scoreboards on the shared scroll resolver

Eight sports scoreboards -- afl, baseball, basketball, football, hockey,
lacrosse, nrl, soccer -- scrolled through this module's own pacing while the
other eleven scrolling plugins went through src/common/scroll_config. Two
implementations of the same job, and this one was on the losing side of every
difference.

It never called set_scrolling_state. Two consequences, both of which this
release's work was about:

- The frame hold is applied through that call, so a speed the crisp ladder
  could render in whole pixels still presented a new frame every refresh.
- Core only runs deferred updates while nothing is scrolling. Believing
  nothing was, it ran blocking work in the middle of these scrolls.

The default is non-crisp today: scroll_speed 50.0 with scroll_delay 0.01 is
50 px/s, which on a 100Hz panel is half a pixel per refresh. That cannot
render as motion -- it alternates 0px and 1px steps and judders at a 50Hz
beat, on every scoreboard, out of the box. Resolved through the ladder it
stays 50 px/s and holds each frame for two refreshes: same speed, whole-pixel
motion.

The stepping disagreement that used to justify a separate module is gone.
scroll_config avoided frame-based mode because it stepped on a wall clock at
1/scroll_delay with scroll_delay set to the frame period, so the decision sat
on its own threshold and flipped on sub-millisecond jitter. That branch now
accumulates elapsed time, identical arithmetic to the time-based one, so the
two differ only in the units the speed arrives in.

What is NOT shared, and must not be: the two modules read identically-named
keys with different meanings. Here scroll_speed is px/SECOND and scroll_delay
only converts to px/frame; in scroll_config scroll_speed is px per STEP, so
px/s is speed/delay. Passing this module's settings dict to the resolver turns
50 px/s into 5000, clamped to 500 -- a tenfold speed-up everywhere. So
_get_scroll_settings keeps sole ownership of reading sports config, including
the league merging, and hands the resolver a plain px/s. A test pins that
specific number, because it is the mistake the refactor invites.

MIN/MAX_PIXELS_PER_FRAME are gone; the resolver bounds speed and the helper
clamps FPS. _resolve_target_fps stays, re-purposed: under the old model that
key was the rate frames were presented at, so it is the faithful translation
into the refresh the ladder is computed against, used when no hardware
refresh is configured.

Speed changes for panels that are not 100Hz: 50 px/s becomes 60 at 60Hz
(+20%) and 48 at 120Hz (-4%). At 100Hz it is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): drop the frame hold when a scroll times out, not just when it says so

set_scrolling_state(False) clears the hold. The other way a scroll ends is
is_currently_scrolling() deciding, after scroll_inactivity_threshold of
silence, that it is over -- which is what happens when the rotation moves on
mid-scroll or a plugin is torn down. That path cleared the flag and kept the
hold, so every later plugin, scrolling or static, was presented at refresh/N
by whoever scrolled last, until something called the explicit stop.

The method's own docstring already states the rule this breaks: the hold "must
not outlive the scroll that asked for it". The timeout was the exception it
did not cover.

Pre-existing, but reachable by three plugins before and eleven after the
sports scoreboards moved onto the shared resolver, so it belongs with that
change. The test ages the activity timestamp past the threshold rather than
sleeping.

Also adds scripts/sports_scroll_check.py. The sports scroll path is per-league
opt-in, so a rig showing static game cards never constructs a
SportsScrollDisplay and none of its pacing can be observed from a normal run
-- which is exactly what happened when this change was first put on hardware:
26 minutes, zero sports scroll lines. The script drives the path directly with
synthetic games and asserts the three things the resolver is meant to buy: the
speed lands on whole pixels, the hold is published, and it is released after.
It never starts or stops the display service, matching scroll_speeds.py, so a
crash here cannot leave the panel dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): refuse to grab the panel while the display service has it

The module docstring already said to stop ledmatrix first. Nothing enforced
it, and running the script against a live service is not a harmless mistake:
rpi-rgb-led-matrix configures GPIO directions and the hardware PWM inside
RGBMatrix(), and when the root check fails it calls exit() from C with no
cleanup. The service keeps rendering and swapping onto pins that have been
reconfigured underneath it, so the panel goes black while every diagnostic
says the display is healthy -- fresh framebuffer, every pixel lit, "RGB Matrix
initialized successfully", nothing in the log. A restart fixes it, once you
work out that is what happened.

Found the hard way: this is what took the panel down on the test rig, not the
change the script was written to verify.

--fallback skips the check, since it never opens the matrix. --force is there
for anyone who means it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(scripts): annotate the subprocess call the way this repo already does

Codacy fails a PR on one new issue, and bandit B404 fires on any subprocess
import. scripts/run_plugin_tests.py carries the same suppression with the same
justification -- list-form argv, no shell -- so this follows it rather than
inventing a second convention.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:13:04 -04:00
ChuckandClaude Opus 5 4423ec33d5 feat(plugins): search, filter and sort for Installed Plugins, on a shared ListFilter helper (#540)
* feat(plugins): add search, filter and sort to Installed Plugins, on a shared helper

The Installed Plugins grid had no way to narrow it down: no search, no way to
see only what's enabled, disabled, or out of date. On a rig with a couple dozen
plugins that means scrolling the whole grid to find one.

The two sections below it already solved this, twice, independently — the
Plugin Store and Starlark Apps carried a copy-paste fork of the same ~600 lines
(filter state, apply-filters-and-sort, page renderer, pagination strip,
active-filter badge, listener wiring). Rather than add a third copy, this
extracts the shared machinery and builds the new toolbar on it.

New: web_interface/static/v3/js/plugins/list_filter.js — ListFilter.create()
owns debounced search, filter axes, sort, the active-filter count, Clear, and
optional pagination/persistence. Callers keep their own card markup via a
`render` callback. Three control types cover every axis the page uses: pills
(new), select (store category, starlark author) and cycle (the tri-state
All -> Installed -> Not Installed button).

Installed Plugins gets a compact toolbar: search box, one-click All / Enabled /
Disabled / Updates pills, and a sort dropdown (A-Z, Z-A, updates first,
recently updated, category). Filters reset on load, so you never come back to a
mysteriously short list. No new CSS — this is the first consumer of the
.filter-pill rules already sitting unused in app.css.

renderInstalledPlugins() is split so it still publishes canonical state while
renderInstalledCards() draws only the visible subset; the filtered list is
never assigned to window.installedPlugins, which the toggle handler,
isStorePluginInstalled(), runUpdateAllPlugins() and the Alpine config tabs all
read as their source of truth. Toggling a plugin while filtered pins its card
so it doesn't vanish from under the cursor.

The Store and Starlark migrations are behaviour-preserving: same element ids,
same localStorage keys (storeSort/storePerPage, starlarkSort/starlarkPerPage),
same tri-state button markup, same pagination. Verified by differential tests
that run the old and new implementations side by side against identical
fixtures and compare every observable after each interaction. The only visible
change is the pagination attribute (data-store-page/data-starlark-page ->
data-list-page), which nothing outside its own click handler referenced.

Net -156 lines in plugins_manager.js while adding a feature.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): keep raw search text, and stop the store search refetching

Two review findings from CodeRabbit on #540.

Do not write the trimmed search value back into the input. setSearch() trimmed
before storing, and syncControls() then copied that trimmed value back over what
the user had typed. Pausing longer than the debounce after typing a space
deleted the space (and reset the caret), making multi-word terms effectively
untypable. The raw text is now kept alongside the trimmed one: filtering and
activeCount() still use the trimmed value, while the input keeps exactly what
was typed.

Remove the legacy #plugin-search / #plugin-category listeners in
initializePlugins(). They bound searchPluginStore as the event handler, so the
DOM event arrived as its `fetchCommitInfo` argument — always truthy, which
skipped the cached-filter fast path and refetched /api/v3/plugins/store/list
with commit info on every keystroke burst and category change. The store's
ListFilter controller already filters the cached list, which is what those two
controls should do. This double-binding predates this PR (the old code guarded
with _listenerSetup and _storeFilterInit, two different flags, so both sets
stayed live); it is fixed here because the refactor owns that wiring now.

Both fixes are covered by tests that fail without them: the trailing-space
regressions in the installed-plugins DOM suite, and a new whole-file jsdom test
that counts fetches while typing (1 request at init, 0 thereafter; previously
1 -> 2 -> 3 -> 5).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): build pagination via DOM APIs, drop computed member access

Addresses the five Codacy security findings, all in list_filter.js.

Pagination no longer assembles an HTML string (3 findings: 2 critical + 1 high,
"unsafe assignment to innerHTML"). The interpolated values were only page
integers and local class constants, so there was no injection path, but
concatenating markup into innerHTML is the pattern the scanners flag and
createElement is no less clear. Each button now also owns its click listener
directly instead of the container being re-queried afterwards, and the strip is
cleared with textContent = '' rather than by assigning empty markup. No
innerHTML assignment remains in the file.

haystack() now walks Object.entries(item) and keeps the configured fields,
instead of reading item[field] per field ("generic object injection sink").
Field order no longer drives the haystack order, which is irrelevant to the
substring test. matches() iterates controls with for...of instead of an index
("variable assigned to object injection sink").

The rendered pagination is unchanged: same buttons, labels, page numbers,
disabled states and classes. The old-vs-new differential tests now compare
pagination structurally (tag, text, page, disabled, sorted class list) rather
than as an HTML string, since building nodes legitimately serialises
differently — «/» as characters rather than &laquo;/&raquo;, disabled="" rather
than a bare attribute. That comparison is stronger than the string one it
replaces, and the real-DOM suite still drives the actual page buttons.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): keep configured field order when building the search haystack

The previous commit swapped item[field] for Object.entries(item) to clear a
static-analysis object-injection warning, and in doing so changed the order of
the haystack: entries follow the object's own key insertion order, not the
configured `fields` order. Since the values are concatenated, that order decides
which values end up adjacent, so a multi-word query spanning a field boundary
matched differently. For store fields [name, description, author, id, ...] and
API objects keyed {id, name, description, author, ...}, "bob plugin-01" matched
before and stopped matching after.

That contradicted the behaviour-preservation claim for the store and starlark
migrations, and the differential tests missed it because every fixture query was
a single word.

Values now come out of a Map built from Object.entries, iterated in `fields`
order: the original haystack is restored, and there is still no computed member
access for the analyser to flag.

Regression coverage for the ordering itself, at both levels:
  - unit: phrases spanning name->id and category->tags, plus the reverse
    (object-key) order asserted NOT to match
  - differential: the same class of query compared old-vs-new, with a guard that
    the phrase actually matches something so a mutual zero-result cannot pass
    vacuously

Verified both fail without this fix (3 unit, 2 differential) and pass with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(web): add JS suites for ListFilter and the plugin-manager grids

No JS toolchain exists in this repo, so these are plain node scripts with no
framework: each prints ok/FAIL lines and exits non-zero. `node test/js/run_all.js`
runs everything, skipping the DOM suites (rather than failing) when jsdom is
absent or nothing is listening, so it stays useful in a bare checkout.

  unit/test_list_filter.js    ListFilter search/filter/sort/count/sticky, driven
                              through the installed-plugins config eval'd
                              verbatim out of plugins_manager.js so the test
                              cannot drift from the real configuration
  unit/test_render_cards.js   renderInstalledCards markup, both empty states,
                              and escaping of hostile plugin metadata
  dom/test_installed_dom.js   the toolbar in a real DOM, including the HTMX
                              partial re-swap and a getComputedStyle check that
                              .filter-pill[data-active] matches what we emit
  dom/test_store_dom.js       store pagination, per-page, category, tri-state
                              Installed button, persistence across a re-boot
  dom/test_no_double_fetch.js loads the whole plugins_manager.js and counts
                              requests, so a keystroke cannot refetch the store

The DOM suites deliberately fetch the partial and the plugin data from a running
web interface instead of using fixtures, so a renamed element id or a changed
payload shape fails them loudly. Point them at a rig with a full plugin set when
it matters (BASE=http://host:5000); a dev box with two plugins installed passes
while exercising very little.

Several assertions exist to stop specific bugs recurring: trailing spaces
surviving the search debounce, a query spanning two adjacent search fields
(haystack field order is load-bearing), and window.installedPlugins staying at
full length while the grid is filtered. Others guard against passing vacuously —
counting only non-skeleton cards, and checking a search phrase matches something
before comparing two result sets.

The old-vs-new differential suites that verified the store and starlark
migrations are not included: they compared against the pre-refactor code, which
now exists only in git history.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 20:04:47 -04:00
ChuckandClaude Opus 5 26769ee37f fix(starlark): the store authenticated with a key nothing writes (#541)
#535 restored the thirteen routes, so the store stopped answering 404 --
and still would not load. Confirmed against a running device before
anything was changed: /repository/browse answers 200 with 1000 apps in
27s, so the routes are fine. Two things underneath them are not.

**The store never used the token the user configured.** The three
repository routes read `github_token` off config.json. Nothing writes
that key -- it is not in config.template.json, no setting offers it, and
it appears nowhere else in the codebase. The configured token goes to
config_secrets.json as `github.api_token`, which PluginStoreManager
loads and every other GitHub caller uses. So the store could never be
authenticated: 60 requests/hour, on the same per-IP budget 48 installed
plugins spend on update checks, while the 5000 the user had already
configured sat unused. On the device, /plugins/store/github-status
reported authenticated with a limit of 5000 at the same moment
/starlark/repository/browse reported 60, with 18 left. The store going
blank was that 60 running out.

**Every failure looked identical.** list_all_apps_cached turned any
listing failure -- rate limit, DNS, timeout, non-200 -- into an empty
app list, and the route sent that out as `status: success`, so a rate
limit and an empty repository drew the same blank grid with no error
anywhere. It now returns the reason, the route answers 502 with it, and
a failure is no longer cached as an empty repository for two hours.

The guard for a bad response was itself a crash: _make_request catches
`(json.JSONDecodeError, ValueError)` but `json` was never imported, so
evaluating the tuple raises NameError and the guard written for exactly
this case never ran. Reachable whenever something on the path answers
with HTML -- a captive portal, a proxy page, a DNS-hijacking router.

Seventeen handlers answered 5xx with no detail at all.
test_no_api_v3_handler_discards_its_exception is meant to prevent that
across api_v3, but it matched one exact message string, and all thirteen
Starlark routes wrote their own wording. The guard now keys on the shape
that matters: if it returns 5xx, it says why. The 15 pre-existing
non-Starlark functions are listed as a set that may shrink, never grow.

**The listing was capped at 1000 and did not say so.** The contents API
truncates a directory silently; tronbyt/apps has 1075 app directories,
so the store showed a truncated repository and looked complete doing it.
Now listed via the git trees API, which reports `truncated`, with the
contents API kept as a fallback.

Not addressed: the 27-second cold load -- 1075 manifests fetched five at
a time behind skeleton placeholders -- which is probably the largest part
of what "does not load" feels like, and wants its own change.

25 new tests.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:04:27 -04:00
ChuckandClaude Opus 5 968b953a51 fix(display): pin one text layout engine, and give the 5x7 BDF face a size (#539)
* fix(display): pin one text layout engine, and give the 5x7 face a size

Two ways a font could render differently on two machines running the same
code, both found while diagnosing four plugins whose golden images passed on
the machine that generated them and failed everywhere else.

**Layout engine.** `ImageFont.truetype` picks its engine at load time: Raqm
where the host Pillow was built with libraqm, Basic otherwise. The two round
fractional glyph advances differently. `PressStart2P-Regular.ttf` at 8px has
whole-pixel advances, so they agree — which is why most of the fleet matched
everywhere and hid this. `4x6-font.ttf` at 6px does not: glyph positions drift
cumulatively along a run, and the four plugins that draw body text in it
(geochron, of-the-day, christmas-countdown, ledmatrix-weather's almanac) are
exactly the four whose goldens travelled badly.

Every core font load now goes through `src/common/font_layout.load_truetype`,
which pins the Basic engine, so a render depends on the font file and the size
and nothing else. Basic gives up complex-script shaping and kerning pairs;
neither applies to bitmap-grid faces on an LED panel. Output is unchanged on a
host without libraqm.

**Zero font height.** `DisplayManager` built the 5x7 BDF face with
`freetype.Face(path)` and never called `set_char_size`, so `face.size.height`
stayed 0 and `get_font_height()` returned 0 for it — callers stacking rows by
`prev_y + prev_height + gap` drew two lines on top of each other. The
start-up line `Calendar font size: 0 pixels` has been printing the symptom all
along. `font_manager._load_bdf_font` already called `set_char_size`, so
whether measurement worked depended on which path loaded the face.

`DisplayManager` now sets it too, and `get_font_height()` falls back to the
strike the file declares rather than returning a zero line height.

Fixes ChuckBuilds/ledmatrix-plugins#397
Refs ChuckBuilds/ledmatrix-plugins#371, #375, #378, #391

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): give the startup banner a rung that fits a full address at 64px

CI caught what pinning the layout engine exposed rather than caused.
`_fitting_font` walks PressStart2P then 4x6 at 6px, and "255.255.255.255" --
the widest thing the startup banner ever shows -- measures 66px at 4x6/6px
against the 62 a 64x32 panel has to give. It used to squeak in only because
the measurement depended on which layout engine the host Pillow happened to
have; with the engine pinned it does not, so the rung the worst case actually
needs is now in the ladder instead of implied: 4x6 at 5px, which measures 51.

The fallback was wrong in the same place. When nothing in the ladder fit, it
returned `self.font` -- the *widest* option, and precisely how "Initializing"
came to run off the side of a 64px panel to begin with. It returns the
narrowest face that loaded now.

test/test_initializing_screen.py: 34 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): name the exceptions the BDF strike read can raise

Codacy flagged the try/except/pass. It was already narrow in intent -- a
malformed strike table on the measurement path must degrade to "size unknown"
rather than take the display down -- but a bare `except Exception: pass` says
neither of those things and hides a genuinely broken font behind a silent 8px
fallback. It now catches what reading `available_sizes` can actually raise and
logs which face failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: drop logo PNGs the render harness downloaded into the worktree

These are fetched at runtime by the logo cache; they are not source, and they
rode in on a `git add -A` while I was running check_plugin.py against this
branch. Nothing in the change needs them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 09:40:00 -04:00
ChuckandClaude Opus 5 e23f1f45d3 feat(starlark,on-demand): the third-party fixes worth taking, plus a Home Assistant MQTT bridge (#538)
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge

Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of
patches and services built while running this project on Starlark apps
under MQTT control. Its patches are whole-file copies taken against an
older tree, so applying them as written would revert #523's frame
pacing, #534's display() bool returns and the GitHub token masking in
plugins_manager.js. Three of its claimed fixes are already in main, and
its api_v3 Starlark routes are #535's. What follows is the rest --
verified against current code, and reimplemented where the patch's
approach did not hold up.

**On-demand display.** `pinned` reached the controller from the API, was
stored on it and republished in the status payload, but never narrowed
the rotation -- a pinned request still cycled every mode its plugin
owns. Right for a sports plugin, whose modes are views of one subject;
wrong for a plugin whose modes are unrelated, which is every Starlark
app. Now honoured, and it survives a restart.

Restarting while on-demand was active loaded *only* the on-demand
plugin, so normal rotation had nothing to return to for the life of the
process -- and a restart mid-session is routine, since that is how an
update is applied. The panel came back cycling one plugin's modes with
no way out but clearing the cache by hand. Every enabled plugin loads
now; on-demand still resumes on its saved mode.

Stop requests are exempt from the duplicate guards on purpose, so that a
second click stops a mode a race left running -- which means consuming
the mailbox is the only thing that ends one. It was never consumed, so
the same stop was re-read and re-processed on every poll, forever. Both
paths now share one compare-before-delete helper.

**Starlark rendering.** `extract_schema` parsed the source with a regex,
which can only see option lists written out literally: an app whose
dropdown is filled from a live API call inside `get_schema()` came back
empty, and the config form offered nothing to pick. Now runs `pixlet
schema`, which executes the app, and falls back to the parser when
Pixlet is absent, too old for the subcommand, or the app fails to run.
The third-party patch replaced the parser outright and hardcoded
/usr/local/bin/pixlet; this keeps the fallback and the binary search.

A `|` in a config value was dropped by a shell-metacharacter filter,
though the command is a list with no shell involved -- and apps do use
it as a separator inside one value. The key went missing silently and
the app rendered its own "not configured" screen with nothing to say
why. And a 0-byte render was reported as success: Pixlet exits 0 and
writes nothing when an app has no content, which read downstream as a
working app drawing a black panel.

**Starlark display.** `display()` ignored the mode it was called with,
so a specific app could not be addressed. It now accepts `display_mode`
-- which is the whole mechanism, since the controller inspects the
signature before passing it. Found while there: `_select_next_app` ran
only while `current_app` was unset, so with several apps installed the
first was picked once and shown forever while the rest were rendered on
schedule and never displayed. And `enable_scrolling` was missing, so
multi-frame apps were called once per rotation slot and never advanced
past frame one.

**GET /api/v3/display/modes.** Every mode that can be requested
on-demand, with the plugin that owns it. Nothing exposed this, so
anything driving the display from outside the web UI read each plugin's
manifest.json off disk and reimplemented PluginManager's fallbacks. It
also triggers discovery, which is otherwise lazy and normally happens
because a person opened the dashboard.

**integrations/mqtt_bridge.** Home Assistant control over MQTT
Discovery: a mode select, a stop button, power, brightness. Rewritten
against the API rather than the filesystem, so it needs no read access
to config.json and cannot drift from the web UI. paho-mqtt 2.x
VERSION2, TLS, an availability topic that is also the last will, and
secrets from the environment.

**Two opt-in extras.** A DNS single-request unit, for glibc's parallel
A/AAAA lookup stalling ~5s per name on routers that answer only the A
query -- which makes any plugin calling an external API slow and
Starlark apps, which have a render timeout, fail outright. And a Pixlet
config editor: a script you run and Ctrl+C rather than the third-party
version's always-on unauthenticated Flask service, since it stops the
display for the length of a session. Neither is installed by default.

Long Starlark app names now wrap instead of overflowing their card.

115 new tests across 5 files. Also unblocked
test_starlark_display_contract.py, which was silently skipping wherever
fcntl is absent. Whole suite: no new failures against main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(mqtt_bridge): the five issues Codacy flagged on this branch

All in the new bridge, all real:

  * requests floor was 2.31.0, which carries CVE-2024-35195,
    CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which
    is what the project's own requirements.txt already pins.
  * `import time` was never used.
  * `"mqtt_password": None` in DEFAULTS read as a hardcoded credential.
    It is the "no password configured" default; marked nosec B105, the
    convention used elsewhere in the repo.

Also dropped an unused `build_app` from the display-modes test imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: the review findings on this PR

Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and
is answered below.

**One bad config section blanked the whole mode list.**
`/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`,
so a non-dict under a plugin id -- a shape DisplayController already
guards, so it happens -- raised AttributeError mid-loop and answered 500
with no modes at all. Every MQTT bridge entity is built from that list.
Now skipped with a warning.

**The DNS scripts reported success they had not earned.** Three separate
paths: `resolvconf -u` failing was swallowed by `|| true`; the
systemd-resolved branch exited 0 without applying anything, so the
oneshot unit recorded success while the workaround was inactive; and the
installer's `|| echo` turned a failed start into "installation
complete." with exit 0. All three now fail loudly. `single-request` is a
glibc resolv.conf option with no resolved.conf equivalent, so on those
hosts the honest answer is that it cannot be applied.

A NetworkManager-generated resolv.conf is regenerated on connection
changes, not only at boot, and the unit is oneshot with RemainAfterExit
-- so the option can vanish mid-boot with nothing to put it back. Now
detected and stated plainly rather than implied to be permanent.

**`Before=` does not order a manual restart.** It only orders units
already in the same transaction, so `systemctl restart ledmatrix` could
bypass the fix. install_dns_fix.sh now writes a ledmatrix.service
drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround
failing should not stop the display.

**The Pixlet editor's `--lan` is gone.** `pixlet serve` has no
authentication, and a printed warning is not access control. Loopback
only, with the SSH port-forward in the header where the flag used to be
documented -- SSH does the authenticating and nothing is left listening.

**The MQTT example config now defaults to TLS** on 8883. The installer
copies it verbatim, and without TLS the broker password and every
command cross the network in cleartext. A plaintext broker is still
supported and documented, and the bridge warns once at startup when a
password is configured without TLS.

**Not taken: "the upstream Pixlet CLI has no `schema` subcommand."**
Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh`
installs `tronbyt/pixlet`, whose `cmd/schema.go` is
`schema [PATH]` -> JSON on stdout, built on
`runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is
exactly what extract_schema_via_pixlet calls. A binary without the
subcommand exits non-zero and falls back to the source parser, which is
already covered by a test.

**CodeQL stack-trace exposure: not taken either.** I removed `details`
first and that broke
test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception,
which enforces `describe_exception` across all ~75 handlers -- written
because a device with failing storage answered "see logs for details"
from the log viewer itself. describe_exception redacts credentials; the
trade-off is the project's and is already made. Restored, with the
reasoning in a comment.

11 new tests. Whole suite: no new failures against main, 4127 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 08:34:52 -04:00
ChuckandClaude Opus 5 793b988d33 fix(starlark): toggle the app id the list published, and store a relocatable star_file (#537)
The two review nitpicks left over from #535. Both are still on main after
that merge; the five findings alongside them landed with it.

**The toggle could not find what the list had just shown.**
`_starlark_virtual_plugins` publishes the raw manifest key as
`starlark:<key>`, and `_toggle_starlark_app` passed it back through
`_validate_and_sanitize_app_id`, which lowercases and rewrites every
character outside `[a-z0-9_]`. An app stored as `My-App` was listed as
`starlark:My-App` and looked up as `my_app`, so toggling an app the page
had drawn a moment earlier answered 404. Keys written by
`_install_star_file` are already sanitised, so this only shows up for
manifests written by the starlark-apps plugin itself or edited by hand.

`_validate_starlark_app_path` rejects traversal without rewriting, so it
is the check to use here -- listing and toggling now agree on one key.
The updater also uses `setdefault` rather than indexing: the app is
loaded but its on-disk entry need not exist, and `_update_manifest_safe`
does not catch `KeyError`, so that escaped as a 500 rather than writing
the entry.

**`star_file` was stored absolute.** Readers join it to the app's own
directory -- `_standalone_render_starlark_app` does `app_dir /
app_data.get('star_file', f'{app_id}.star')` -- so the key's default is a
bare filename and an absolute value gave it a second meaning. Since
`Path.__truediv__` discards the left side when the right is absolute,
the manifest was pinned to whatever PROJECT_ROOT installed it, and a
moved or redeployed install could not find its own file. Storing
`dest.name` matches the default and stays relocatable. Read paths are
unchanged, so manifests already holding an absolute path keep working.

7 new tests. Whole suite: no new failures against main, 4013 passed
against 4007.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 16:35:56 -04:00
Chuck 50258635a8 fix(starlark): restore the API routes #330 dropped (#535)
The Pixlet install button reported "Pixlet install failed: Resource not found" -- Flask's 404 handler, because the route did not exist. #253 added thirteen Starlark routes; #330 rewrote api_v3.py and dropped all of them, along with the `starlark:<app_id>` entries that surface installed apps in the plugins list and the toggle branch that enables them.

Restores all thirteen routes, the plugin-list entries and the toggle path, so Pixlet installs, the app store browses and installs, and an installed app can be managed like any other plugin.

Not a straight revert. Three error paths stopped returning exception text to the caller; the manifest write moved off a shared temp filename that two concurrent writers could interleave; both dynamic importers stopped leaving half-initialised modules in sys.modules; the config update rolls back when the save fails; the toggle checks that persistence succeeded; and the path check returns the validated path instead of a boolean so callers stop re-joining the raw value. New tests no longer reach GitHub.

Verified on a 256x64 Pi: Pixlet installs and runs (v0.53.1), the store lists 1000 apps, install/toggle/uninstall round-trip, and traversal and command-injection probes are rejected at every entry.

25 CodeQL alerts dismissed as verified false positives -- path-injection where traversal is blocked, and one list-form subprocess with no shell. Both classes already present on main.

Full core suite: 3981 passed.
2026-09-07 16:06:09 -04:00
ChuckandClaude Opus 5 c9289e3a1d fix(store): update_plugin silently did nothing for four installed plugins (#536)
install_plugin() deliberately renames a plugin's directory to the MANIFEST id
when it differs from the REGISTRY id, so registry `stocks` lands in
`ledmatrix-stocks/`. Every lookup in _find_plugin_path() is by directory name,
so update_plugin("stocks") found nothing, logged "Plugin not installed", and
returned False.

Nothing surfaced that to the user. Clicking update in the web UI was a no-op
with no error, and the plugin stayed on a stale version indefinitely. Four
installed plugins hit this on a real device -- leaderboard, music, stocks and
weather -- found because a scripted update of eleven plugins failed on exactly
those four.

Adds a manifest-id scan as the LAST step of the resolution chain, so the two
documented lookups above it (configured dir, then the sibling plugins/
fallback) keep their exact meaning and ordering. That ordering is pinned by
test_discovery_path_contract.py, which characterises the divergence between
the three resolvers on purpose; this extends the chain rather than reordering
it. Directories renamed aside with '.standalone-backup-' during an install or
rollback are skipped, since matching one would report a half-finished install
as a live plugin.

Also adds scripts/audit_render_path.py, which walks the call graph from
display() and reports blocking calls reachable from it. display() runs on the
render thread, so anything slow there stalls the panel; on a vsync-paced loop
a single 15ms call drops a frame and a network round trip freezes the marquee.
Two instances were already found the slow way, by reading frame-time
histograms -- odds-ticker reading the scoreboard cache per frame, and
soccer-scoreboard timing out inside update(). The audit finds that shape in
the source instead. It is a heuristic and says so: a hit behind an interval
check may be fine.

It currently flags 23 calls across six plugins. The clearest is
ledmatrix-music, whose display() falls back to an inline
requests.get(timeout=5) when album art has not been prefetched -- a deliberate
"show the art rather than go blank" tradeoff by its author, but up to five
seconds of frozen panel. Reported, not changed; that is its owner's call.

185 store tests pass. Three of the six new tests fail without the fix.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 15:55:46 -04:00
ChuckandClaude Opus 5 d12323e7f1 perf(scroll): pace frames to the panel — 44→100 fps, stalls 14% → 0.02% (#523)
* perf(scroll): pace frames to the panel, not to a fixed sleep

Scrolling ran at 44-46 fps on a 2x128x64 chain and 14-17% of frames took
41-53ms, which reads as judder. Four independent causes, each measured on
the hardware; details and the diagnostic recipe are in
docs/SCROLL_PERFORMANCE.md.

The high-FPS loop slept a flat 8ms after every render. display() has
already blocked on the panel's vsync by then, so that sleep was added to a
wait that had happened: ~4ms of render plus 8ms put each iteration at ~12ms
against a 10ms refresh grid, so every swap missed a refresh and the loop
settled at 50fps while asking for 125 -- with no headroom, so a further
14% of frames slipped again. It now sleeps only the remainder, with a 1ms
floor so plugin threads still get the GIL.

ScrollHelper stepped position on a wall clock at 1/scroll_delay steps per
second. Plugins set scroll_delay to the frame period, so that comparison
sat exactly on its own threshold: a frame arriving a hair early moved zero
pixels and rendered an identical frame, dirty-tracking skipped the swap, it
returned in ~2ms, and the beat repeated. No scroll_delay value tunes that
out -- a shorter delay trades stalled frames for periodic double-steps.
Both modes now accumulate elapsed time at the same configured speed, so
position stays proportional to real time.

Sub-pixel blending goes back to off by default. It renders a half-step by
mixing two adjacent columns, which on a coarse panel showing pixel-font
text alternates crisp and smeared frames and reads as shimmer -- visibly
worse than integer stepping on the hardware. Vegas mode still opts in.

disk_cache uses orjson when importable, falling back to the stdlib. Encoding
a ~1MB record drops from 14.8ms to 5.4ms end-to-end, and that work holds the
GIL while a marquee is on screen. display_manager also checksummed the whole
framebuffer twice per frame (dirty tracking, then the preview snapshot); the
snapshot now takes the checksum the caller already computed.

New src/common/scroll_config.py resolves scroll settings in one place. Five
ticker plugins each hand-rolled this and disagreed: odds-ticker ranked the
deprecated scroll_pixels_per_second above the documented scroll_speed/delay
pair, and because that key carries a schema default the documented settings
were dead for every user (ChuckBuilds/ledmatrix-plugins#408), while
ledmatrix-leaderboard read the same key only as a fallback. The resolver also
warns when a speed will not advance a whole number of pixels per refresh,
which is the property that actually determines whether a scroll looks smooth.

scripts/build_rgbmatrix_nogil.sh rebuilds the rgbmatrix binding so it
releases the GIL. Upstream declares SwapOnVSync without nogil, unlike
SetPixel/Clear/Fill beside it, so the render thread held the GIL for the
whole vsync wait and starved background threads into long uninterruptible
bursts. The script patches, builds and self-verifies into a scratch tree;
--install backs up the original and rolls back if the service does not come
back healthy.

Measured after: 100 fps locked, no stalls observed, render thread down from
51% to 19% of one core.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): keep the panel swap locked to vsync while scrolling

Dirty tracking skipped SwapOnVSync for byte-identical frames. That is the
right call for static content, but SwapOnVSync is also what paces the render
loop, so skipping it skips the wait for the panel: a duplicate frame returns
in ~8ms instead of ~10ms on a 100Hz panel, advances the strip only 0.8px
instead of 1.0px, and so makes the next frame more likely to repeat as well.
The effect sustains itself once it starts.

Measured over 20 minutes on a 2x128x64 chain, both scrollers configured
identically at 100 px/s:

    leaderboard   10ms x35, 11ms x3            (clean)
    odds-ticker   10ms x26, 8ms x7, 15ms x5    (~20% duplicates mid-scroll)

The duplicates were not end-of-cycle idling -- 38% of fast frames fell within
90s of a scroll completion against 35% of normal frames, a null result. The
trigger is per-frame work: odds does more of it, and more variably, so it is
first to land a frame that advances less than a whole pixel.

Pushing an identical frame costs one canvas copy. Falling out of vsync lock
costs smooth motion. Static content is untouched, because
is_currently_scrolling() expires on its own inactivity threshold -- covered
by test_stale_scrolling_state_stops_forcing_pushes so a plugin that stops
scrolling without saying so cannot pin the panel into always-push.

Also de-flakes test_snapshot_still_written_on_skip, which asserted a strict
mtime increase between two writes that can land in the same filesystem tick;
it failed about two runs in three on Windows regardless of the code under
test. The file is now backdated before the check.

156 tests pass on the Pi. Not yet confirmed by eye on the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): report the frame-time tail, and stop the row-major blit

Two problems, both found by looking at the panel rather than the metric.

The frame-stats line reported ONE instantaneous frame every 5 seconds --
about 1 frame in 500 -- printed beside a 100-frame average. Both hide exactly
the fault they are used to chase: a 2ms duplicate and a 21ms double-wait
average to precisely 10ms, so a ticker stalling on half its frames still
reports a healthy "Avg FPS: 100.0". That reading cost several rounds of
chasing the wrong layer. The line now aggregates every frame since the last
log and reports median, p95, max, min, and explicit stall and skip rates
(past 1.5x the median missed a refresh; under half never reached the panel,
because dirty tracking skipped the swap so the frame never waited on vsync).

On the hardware this now reads:

    leaderboard  100.0 fps over 501 frames | median 10.00ms p95 10.05ms
                 max 10.34ms | stalls 0 (0.0%) skips 0 (0.0%)

The binding rebuild's blit patch becomes opt-in (RGB_PATCH_BLIT=1, default
off). Reordering that loop to row-major changes what a torn frame looks like:
column-major tearing shows as a vertical seam, row-major as a horizontal split
between the panel's upper and lower halves. On a 1/32 scan panel that reads as
a one-pixel fold across the middle of every panel, which is what was reported
on hardware and what went away when the blit was reverted. All of the measured
gain comes from the SwapOnVSync change, so the risky half is simply not worth
taking; the header says so.

Also fixes --install resolving its paths against $HOME, which is /root under
sudo, so it looked in /root/rgbmatrix-nogil-build and died with "no built
module found" on a machine where the build had just succeeded. It now resolves
SUDO_USER's home. Both build paths are verified on the Pi: default yields one
GIL-release site, RGB_PATCH_BLIT=1 yields two.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(scroll): let users pick a crisp speed for their own panel

Whole-pixel motion was previously only available at multiples of the refresh
rate -- 100, 200, 300 px/s on a 100Hz panel. 100 px/s crosses a 256px panel in
2.6s, which is brisk for reading, and everything slower had to blend (blur) or
repeat frames unevenly (judder). There was no way to ask for 50 px/s and get
clean motion.

SwapOnVSync takes a framerate_fraction the display manager never passed. It
holds each frame for N panel refreshes; the panel keeps refreshing at its full
rate throughout, so holding costs nothing in flicker and only changes how often
a NEW image is presented. That turns 50 px/s into one whole pixel every second
refresh instead of half a pixel every refresh.

The crisp speeds are therefore refresh_hz / hold * pixels_per_frame, and that
ladder depends on the panel: a Pi Zero on a long chain has a different set of
good speeds from a Pi 4 on a short one. crisp_ladder() enumerates them and
solve_crisp() picks the best match for a requested speed.

solve_crisp weights motion quality rather than picking the numerically nearest
entry, which matters more than it sounds. Asked for 30 px/s, nearest-by-value
answers 28.6 -- 2px jumps at 14fps -- over 33.3, which is single-pixel motion
at 33fps and obviously better on the panel. The target is also clamped into the
ladder's range first, because relative error saturates near 1.0 for a target
far outside it and the quality penalty would otherwise answer "10000 px/s" with
the slowest entry.

configure() snaps to the ladder and applies the hold when given a display
manager. Without one the hold silently cannot happen and motion falls back to
fractional pixels, so it warns rather than failing quietly. set_frame_hold()
resets to 1 when scrolling stops, so one plugin's pacing cannot leak into
whatever is on screen next.

scripts/scroll_speeds.py is the user-facing part: it prints the ladder for the
configured rate, measures what the panel ACTUALLY manages (--measure, for
hardware that cannot reach its configured limit), highlights the nearest option
to a wanted speed, and demos one live. It never starts or stops the display
service itself -- doing that inside a script stranded the panel twice today.

Speeds below ~20 px/s remain stepped regardless. That is the pixel pitch, not a
software limit.

Also fixes the dirty-tracking test spy, which stubbed SwapOnVSync with a
single-argument function and would have masked the new call as a failed push,
and rewrites a configure() test that had started passing for the wrong reason:
it asserted a judder warning, which snapping now prevents, and was matching the
unrelated "hold could not be applied" warning instead.

183 tests pass on the Pi.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): tie the frame hold to the scroll, not the plugin

The hold applied in configure() never reached the panel. Plugins share one
display manager, and set_scrolling_state(False) -- fired whenever ANY other
plugin finishes its scroll -- reset the hold to 1. A hold set once at plugin
construction was therefore always gone by the time that plugin rendered.

The symptom was a log line that lied. ledmatrix-stocks reported

    Scroll configured: 50.0 px/s (1px every 2 refreshes = 50.0 fps, smooth)

while the panel measured 100.0 fps, median 10.00ms. Config, resolution and
snapping were all correct; only the pacing silently was not applied.

set_scrolling_state(is_scrolling, frame_hold=1) now carries it, so the hold
lives exactly as long as the scroll that asked for it. configure() reports the
value as ScrollSettings.frame_hold instead of applying it -- applying it behind
the caller's back could never have been right on a shared display manager.
Existing callers are unaffected; the default keeps one frame per refresh.

Verified on hardware: stocks at 50 px/s now measures

    50.0 fps over 251 frames | median 20.00ms p95 20.09ms | stalls 0 skips 0

20.00ms being exactly two refreshes, with the panel still refreshing at 100Hz
underneath so flicker is unchanged.

test_another_plugin_stopping_does_not_strand_a_hold pins the interaction that
broke this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll,cache): resolve CodeRabbit review on #523

Eight findings, all reproduced before fixing.

scroll_config.configure() read the refresh rate *after* resolve() had
already used it. resolve() fills in target_fps, pixels_per_frame and the
judder warning from that rate, so on a 60Hz panel every one of them
described 100Hz -- and with snap_to_crisp=False nothing downstream
corrected it, so set_target_fps() paced the helper to 100 FPS. The rate
is now settled first, and falls back to the global config rather than
straight to the default.

refresh_hz_from_config() used `(cfg.get("display") or {}).get(...)`,
which raises AttributeError when either level is truthy but not a
mapping -- out of a function whose whole contract is a rate or a default.

The frame-stats line reported the upper-middle sample as the median and
the 96th sorted sample as p95 of 100. Both are also thresholds (stalls
at 1.5x the median, skips at 0.5x), so the counts were biased too. The
arithmetic is now in frame_stats()/format_frame_stats(), testable
without a clock.

configure()'s docstring and docs/SCROLL_PERFORMANCE.md still said it
applies the frame hold and warns when it cannot. It deliberately does
neither since "tie the frame hold to the scroll, not the plugin"; a
caller following the old text would omit set_scrolling_state() and slow
snapped speeds would still present every refresh.

disk_cache had no policy for non-finite floats: orjson writes null,
the stdlib writes NaN/Infinity, and orjson then rejects those legacy
files so DiskCache.get deleted them as corrupt. One behaviour on both
paths now -- write null, keep legacy records readable. allow_nan=False
detects the values; the replacement walk runs only when there is one,
so the ordinary write path is byte-identical and pays nothing.

build_rgbmatrix_nogil.sh picked the build artifact with a glob piped to
`head -1`, which sorts cpython-311 ahead of cpython-313, so a stale .so
staged in from the source tree was installed as core.so while the GIL
check -- which reads the generated core.cpp, not the .so -- still passed.
It now requires the current interpreter's exact ABI name and fails
closed. Its systemctl calls were also unchecked under `set -uo pipefail`:
a failed stop left the old service running, the following start
succeeded as a no-op, and the health check reported SUCCESS for a
binding that was never loaded.

orjson floor raised to 3.11.6 for CVE-2025-67221 (unbounded recursion
in dumps); it covers the project's Python 3.10-3.13 range.

Adds test/test_cache_nonfinite_floats.py (14) plus regression tests in
test_scroll_config.py and test_scroll_helper.py. 9 of the cache tests
and 9 of the scroll_config tests fail against the pre-fix code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(harness): keep the visual double's signature tied to production

Moves set_scrolling_state's frame_hold into the test double here, where
DisplayManager gains it, rather than in #534 where it arrived a PR early.
CodeRabbit flagged the #534 version correctly: a double that accepts an
argument production does not lets the call pass every harness run and
raise TypeError on the panel, which is the one failure a safety harness
exists to prevent.

The drift has now gone both ways across two branches -- double behind
production on this branch, double ahead of it on #534 -- so it is pinned
instead of remembered. test_display_double_parity.py compares the two
signatures and fails with the direction of the drift named. It reads the
files with ast rather than importing them, because display_manager
imports rgbmatrix at module scope and this check should hold on a laptop
and in CI as well as on a Pi.

Plugins begin passing frame_hold in ledmatrix-plugins#462, which is why
production and the double both need it before that lands.

Full suite: 3889 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 13:37:54 -04:00
ChuckandClaude Opus 5 a0d3e64099 fix: ten defects found validating the whole plugin fleet on hardware (#534)
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers

Three independent fixes found while validating every plugin on a 256x64 rig.

FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf
ships with the core, so every plugin offering it logged "Font family
'tom_thumb' not found" (16 warnings per countdown render) and had to carry a
private loader to use a bundled font. Closes #524.

VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter
that DisplayManager gained, so any plugin passing it died with TypeError at
render time and failed every size. Nine plugins now make that call;
ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other
seven only passed because their scroll path was unreachable without data.
Closes #525.

api_v3 declared module-level config_manager/plugin_manager = None that nothing
ever assigned -- app.py sets the blueprint attributes, which the other 150+
call sites use. Three sites read the decoys, so /health reported the config
unreadable and the plugin system uninitialised (making "degraded" permanent and
unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on
every rig. The decoys are removed rather than assigned, so a bare name is now a
NameError at test time instead of a silent None. The same function's first-call
uptime was computed from two separate clock reads and came out negative.
Closes #529.

Verified on the rig: both previously-failing plugins render, the tom_thumb
warnings are gone, /health reports "healthy" with all three checks passing, and
/display/current reports the real 256x64.

* fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark

The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is
world-writable and sticky, and the display service runs as a different user
from the tooling, so a leftover temp owned by anyone else became unopenable
even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a
sticky directory. The preview and the health check's liveness proxy then froze
until someone deleted the file by hand; on the test rig that meant 23 hours of
a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with
cleanup on failure, matching the hardware-status write a few hundred lines
above. Closes #528.

_poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults
the in-memory TTL to max_age -- so the first request was pinned in memory for an
hour and every later poll returned that stale copy. No second on-demand request
was honoured until the service restarted, while the API kept returning 200.
get() already documents memory_ttl=0 for exactly this cross-process case.
The consumed request is also now deleted: leaving it on disk meant a restart
replayed the previous request, activated it, and ignored the one the caller had
just made. Closes #530.

starlark-apps returned None from display() when it has no app to show, which is
the state of every install without Pixlet and of a fresh one before any app is
added. The controller only skips on a boolean False, so that held a black panel
for the full display_duration instead of rotating on. Closes #456 (core side).

Verified on the rig: two consecutive on-demand requests with no restart between
them are both activated, where the second was previously dropped in silence.

* perf(harness): share one cache across a plugin's renders

_instantiate built a fresh MockCacheManager for every (size, mode), and that
mock is a per-instance in-memory dict, so each render was a cold start. A plugin
that fetches per game or per player re-fetched everything N times over --
baseball-scoreboard at one size took 840s for nine renders where the arithmetic
said ~72s, and at eight sizes it exceeded a 900s timeout.

The second and later renders also never exercised the cache-hit path, which is
what a running rig executes almost all of the time, so a caching regression
could not be caught here.

The cache is now built once per render_plugin_matrix call and threaded down.
The display manager stays per-render -- the bounds checking depends on that --
so only fetched data is shared.

Measured on the rig, same render counts and same goldens:
  tide-display         2s -> 1s   (32 renders)
  cricket-scoreboard  10s -> 3s   (24 renders)
No pass/fail change across tide-display, cricket-scoreboard, clock-simple,
geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages.

Closes #533.

* fix(scripts): run standalone plugin tests instead of collecting nothing

run_plugin_tests.py discovered every plugin test file and handed the lot to
pytest. Most plugin tests are standalone scripts -- module-level main() plus an
`if __name__ == "__main__"` guard, signalling through an exit code -- and pytest
collects zero items from those. The run printed how many files it had *found*,
then "no tests ran", and exited without executing any of them. On a rig with all
44 first-party plugins that is 151 of 248 files.

Files are now classified and each kind runs under the right runner: pytest for
real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1
fail convention ledmatrix-plugins' own runner established (a script that wants a
tty or an LED matrix is a skip, not a regression).

Before:
    $ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos
    Found 1 test file(s)
    collected 0 items
    no tests ran in 0.31s                      rc=0

After:
    Found 1 test file(s) -- 0 collectable, 1 standalone script(s)
    1 passed, 0 skipped, 0 failed (scripts)    rc=0

Verified across three shapes: countdown (1 script), jellyfin-now-playing and
pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files
split 4 collectable / 7 scripts, all seven of which had never run).

Closes #532.

Running the flights scripts for the first time also surfaced four genuinely
failing tests there, hidden by the mirror-image bug in the plugins repo's own
runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465.

* fix(harness): give an empty-looking mode a few frames before warning about it

check_plugin's "drew nothing but display() returned X" warning fired on a single
frame, rendered with force_clear=True, under a frozen clock. All three defeat a
scrolling plugin, whose first frame is legitimately its blank scroll-in buffer.
Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which
people stop reading a warning, which matters because the true positives are
real: a mode that draws nothing and does not return False holds a blank panel
for its whole display duration.

An apparently-empty frame is now re-driven for up to 48 more frames with
force_clear=False (force_clear means "reset the scroll", so repeating it would
redraw frame 1 for ever) and with the clock advancing -- freezegun's factory
where time is frozen, a real sleep where it is not, since scroll position is
usually a function of elapsed time. The first frame that draws content replaces
the result.

The clock is moved back afterwards. It is shared by every render in the matrix,
so time borrowed by the probe leaked into later modes and drifted their goldens
-- f1_upcoming picked up 5 spurious drifts before this was restored.

Measured on the rig:

                        empty warns          check
                        before  after
  f1-scoreboard            42      0    48 PASS / 0 FAIL, goldens intact
  ledmatrix-elections      16      0    16 PASS / 0 FAIL
  on-air                    8      8    true positive, kept
  nfl-draft                 8      8    true positive, kept
  clock-simple/geochron/    0      0    unchanged
  christmas-countdown

58 false positives gone, both true positives kept, no golden regressions. Cost
is confined to modes that really are blank: plugins that draw immediately are
unchanged (clock-simple and tide-display still 2s), while on-air -- eight
deliberately blank modes -- goes to 21s.

Closes #527.

* fix(harness): load nested schema defaults, and merge caller config at leaf level

load_config_defaults read only top-level properties. An object property carries
its defaults on its children, not on itself, so everything nested was dropped --
2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of
565. render_plugin_matrix's comment says the plugin then "behaves like a real
install", which for most of the fleet it did not.

_defaults_from_properties now recurses. merge_config deep-merges the caller's
config onto the result so an override lands at the leaf: a shallow merge would
let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard
every other nhl default, which is the same class of bug being fixed here.

Measured before/after across all 49 installed plugins on the rig: **no render
changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact.
Plugins already fall back to the same values internally via config.get(key,
default), so supplying them explicitly agrees with what they were doing. The
defaults really are arriving now:

  ufc-scoreboard        9 -> 87 defaults
  ledmatrix-flights    51 -> 95
  masters-tournament   10 -> 51
  cricket-scoreboard   22 -> 50
  tide-display         12 -> 18

and hockey-scoreboard, which used to load nhl.enabled=None, now gets
nhl.enabled=True with its full display_modes block.

Caveat worth carrying: the eight plugins with the most nested config
(soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of
the 2,386 dropped defaults, 68%) could not be measured. They import
src.common.sports_shared, which the test rig's core branch predates, so they
fail to load there identically before and after. Re-run this comparison against
a core that has that module before trusting the "nothing changed" result for
them; those are exactly the plugins whose renders should change most.

Closes #531.

* refactor: narrow the exception handlers this branch introduced

Codacy flagged the new code; it passes on other recent PRs, so the finding is
mine. Four of the five broad `except Exception` clauses I added were catching
far more than they needed to, which is the same shape as several bugs this
branch fixes -- hello-world's TypeError sat invisible for exactly this reason.

  freezer() / move_to() / tick()   -> (AttributeError, TypeError, ValueError)
  cache_manager.delete()           -> (OSError, AttributeError, KeyError)

The fifth stays broad and now says why: it wraps a call into a plugin's own
display(), which can raise anything, and the first frame has already rendered --
so a failure there must not turn a good result into an error.

Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty
warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows
7 golden drifts both before and after this branch, so it is not from these
changes -- its committed goldens predate #521's 1-bit text rendering.

* fix: resolve CodeRabbit review and Codacy findings on #534

CodeRabbit raised six; all six were real.

The test double had drifted ahead of production. VisualTestDisplayManager
accepted set_scrolling_state(frame_hold=...) while DisplayManager did not,
so such a call passed every harness run and would raise TypeError on the
panel -- the one failure a safety harness exists to prevent. frame_hold
belongs to the change that adds it to DisplayManager (#523), so it moves
there and the double matches main again.

The harness swallowed exceptions from re-rendered frames. _settle_loop
re-renders a mode that came back blank, to give a scroll time to draw;
returning silently on a crash meant a mode that renders one good frame
and then explodes was reported as passing. Recorded on result.error now,
keeping the captured frame so the failure stays inspectable.

starlark-apps display() returned True after _display_frame() failed, so
the controller held a dead frame for the whole display_duration instead
of rotating on. _display_frame now returns bool on all three paths.

run_plugin_tests.py used env.setdefault for PYTHONPATH and
LEDMATRIX_CORE, so an inherited value won and the subprocess imported a
different core than the one under test -- ledmatrix-plugins#467 exactly.
Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally.

The on-demand mailbox is polled after every frame, ~125x/second on a
scrolling mode, and the read is deliberately uncached, so it was that
many disk reads per second to find nothing. Floored at 250ms, which is
imperceptible for a web-UI click. Consuming it also deleted whatever was
present rather than what had just been processed, so a request posted
while the previous one was in flight was thrown away and never ran; the
delete is now keyed by request_id. That narrows the window rather than
closing it -- a true atomic claim needs a primitive the cache layer does
not offer, and the code says so rather than implying otherwise.

Codacy's 2 criticals were bandit B404/B603 on the subprocess call added
to run_plugin_tests.py. Fixed interpreter, argument list, no shell;
annotated with the repo's existing nosec convention. Bandit is clean on
the file.

Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py
(4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of
those fail against the pre-fix code. Full suite: 3961 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: satisfy Codacy's subprocess checks on the new test runner

Codacy runs Bandit and Opengrep (its Semgrep fork). The new
subprocess.run in scripts/run_plugin_tests.py trips three patterns, on
two different lines:

  Bandit   B404 on the import, B603 on the call
  Opengrep dangerous-subprocess-use-audit          on the run( line
           dangerous-subprocess-use-tainted-env-args on the argv line

A nosemgrep applies only to its own line, so the call line and the argv
line each need one; a single comment on the call covered neither rule
fully. Suppression is the right answer here rather than a rewrite: the
interpreter is sys.executable, the arguments are a list, and no shell is
involved, so there is nothing to word-split or expand.

Matches the pair the rest of the repo already uses for this shape --
permission_utils.py, plugin_loader.py, install_dependencies_apt.py.

Codacy: 0 new issues, up to standards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: leave visual_display_manager untouched so #523 can merge

The only change this branch made to that file was a docstring, and it
collided with #523's rewrite of the same method -- so #534 and #523 each
merged cleanly against main but conflicted with each other. Reverted to
main's text; #523 owns this method and adds frame_hold to it.

The note the docstring carried ('frame_hold arrives in #523') would have
been stale the moment #523 landed anyway. The parity test in #523 is
what actually keeps the two signatures honest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: add the Ruff suppression nosec/nosemgrep do not cover

Ruff reports S603 on the same call Bandit and Opengrep do, and none of
the three suppressions covers the others. Confirmed the precondition
first: path comes from discover_plugin_tests(), which globs test files
inside the repo, and the call is a fixed interpreter with a list argv
and no shell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 13:37:38 -04:00
ChuckClaude Opus 5coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
696acdbc7b feat(render_plugin): add --display-mode so multi-mode plugins can be rendered (#522)
* feat(render_plugin): add --display-mode so multi-mode plugins can be rendered

render_plugin.py always called plugin.display(force_clear=True) with no mode.
A plugin that declares one display mode is fine, but the sports scoreboards
declare three or more and keep their per-mode state on sub-managers; their
no-argument path selects nothing and returns False, so the render came out
blank with nothing to say why. Measured on nrl-scoreboard with identical
seeded state:

  live.display() directly                 True,  1892 lit pixels
  plugin.display(display_mode="nrl_live") True,  1892 lit pixels
  plugin.display()                        False,    0 lit pixels

--display-mode passes the requested mode through. It is only passed when
asked for, so the many plugins whose display() takes no display_mode keep
working untouched, and a plugin that declares modes but does not accept the
argument degrades to its default screen with a warning rather than a
TypeError.

This is what lets the plugin READMEs show a scoreboard at all, and it also
unblocks screens like birdnet_stats and the weather plugin's hourly, daily and
almanac modes, which could previously only be described in prose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Only fall back when plugin display rejects display_mode

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-09-06 17:28:22 -04:00
ChuckandClaude Opus 5 91d15a8943 fix(display): draw text 1-bit, so glyphs stay crisp on the LED grid (#521)
An LED panel has no partial brightness. PIL defaults ImageDraw's fontmode to
"L", which anti-aliases TrueType glyphs into a grey fringe the panel can only
round off -- a 4px glyph arrives smeared into 3px.

DisplayManager creates its shared `draw` in six places and set fontmode at
none of them, while _load_fonts loads extra_small_font as 4x6-font.ttf at
size 6. Measured at draw time, that face at that size puts 74% of its lit
pixels at partial coverage. Every plugin drawing small text through the
shared draw inherited the blur; geochron was the case that surfaced it.

The harness's VisualDisplayManager had the same gap, which mattered more than
it looks: goldens were recording anti-aliased text that production would not
produce, so the harness could not have caught this. Fixing only production
left geochron still blurry under the harness -- that is how the second site
was found.

Both are set to "1" so the harness renders what the panel renders.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 16:02:04 -04:00
Chuck 6bea1a7c21 fix(sports): say when the schema cannot be read, instead of failing silently (#520)
_schema_font_size swallowed every exception and cached an empty dict. That is
not cosmetic. With no schema, a configured font size can no longer be compared
against the schema default, so every size is treated as a deliberate user
choice and skips the snap to the font's pixel grid -- which renders
4x6-font.ttf at 6 instead of 7: a 3px-wide glyph instead of 4px.

That shipped. On a 256x64 panel it made the odds, the team records and the date
row hard to read, and it was found by a user counting pixels on a photo of the
panel rather than by anything here. The cause (_plugin_dir returning None under
the real plugin loader) is fixed in #519; this makes the same class of failure
audible next time:

    Orphan: could not read config_schema.json (FileNotFoundError: ...); every
    font size will be treated as user-chosen and will skip its pixel grid
    snap. Font sizes may render a pixel narrow.

The message names the consequence, not just the error, because the error alone
does not suggest "your fonts are a pixel narrow".

Logged rather than raised: an unreadable schema must not stop a plugin
rendering. The cache is built once per class (per schema path in sports_card),
so this cannot repeat per frame.

Scope deliberately small. An audit of the three shared modules found 23 handlers
that swallow and return a default, but all 23 catch specific types -- TypeError,
ValueError, ImportError -- turning bad config values into defaults, which is
what they are for. Of 77 broad handlers across the font and odds paths, 74
already log. Only these two were both broad and silent.
2026-09-04 16:01:50 -04:00
Chuck 0730d95200 fix(sports): let the plugin declare its own directory, don't deduce it (#519)
_plugin_dir() returned None on every device. The consequence was silent and
reached the panel:

    _plugin_dir()       -> None
    _schema_font_size() -> None for every element
    -> a configured size equal to the schema default stops looking like a
       default and is treated as a deliberate user choice
    -> the snap to the font's pixel grid is skipped
    -> 4x6-font.ttf renders at 6 instead of 7: 3px-wide glyphs, not 4px

On a 256x64 panel that made the odds, the team records and the date row hard to
read. Both `odds` and `detail` were affected -- anything resolving a
grid-snapped schema default was a pixel narrow.

Why it was invisible here. PluginLoader._namespace_plugin_modules renames every
bare module a plugin brought in (sports, game_renderer, ...) to
"_plg_<plugin_id>_<module>" and REMOVES the bare sys.modules entry, so two
plugins owning a module of the same name cannot collide. A class defined in
sports.py still reports __module__ == "sports", but sys.modules["sports"] is
gone, so walking the MRO for a module with a __file__ finds nothing.

Every test here imported plugins directly, which leaves the bare entry in
place, so the walk succeeded. The safety harness loads plugins its own way and
never reproduced it either. It was found by a user counting pixels on the
panel.

The directory is now declared by the plugin (_PLUGIN_DIR) and only deduced as
a fallback, for hosts that declare nothing -- the plugins' own probe harnesses
build classes with type().

Verified on hardware, which is the only place the original failure appeared:
before, the live service logged plugin_dir=None and 4x6-font.ttf@6 for all six
football managers; after, plugin_dir resolves and both odds and detail are @7.

Five regression tests, including the production shape: a class whose __module__
is absent from sys.modules still resolves via its declared directory, and the
precondition that the MRO walk alone returns None is pinned so the test keeps
meaning something if the fallback changes.
2026-09-03 17:01:30 -04:00
Chuck 32d637a446 fix(store): read the core version from disk, not from a stale import (#518)
* fix(store): read the core version from disk, not from a stale import

Updating the core to 3.3.0 and then updating plugins refused all eight sports
scoreboards:

    Refusing to install nrl-scoreboard: NRL Scoreboard supports LEDMatrix
    >=3.3.0, but this system is running 3.2.0.

while src/__init__.py on that machine read 3.3.0. Observed on hardware, not
theorised.

The gate ran `from src import __version__ as core_version`, which binds
whatever the process loaded at start. The plugin store's gate lives in the web
UI, a long-lived service of its own, and the update route deliberately restarts
nothing -- it replaces files on disk and asks the user to restart. Its prompt
named only the *display* service, so a user who followed it left the web
process holding the previous number.

Stale by exactly one release is the case that bites: every plugin flooring on
the release you just installed is refused, blaming a core version that is
already correct on disk. It reads as a broken plugin store. 3.3.0 is the first
release where this hits a whole family at once, since all eight scoreboards
floor there.

compatibility.current_core_version() reads the version from the file instead,
falling back to the imported value on any failure -- so it can only ever be as
correct as before, never worse. All four gate call sites use it: three in
store_manager (install, the git-pull update path, install_from_url) and one in
plugin_loader's advisory warning.

The restart prompt now names both services.

Twelve tests, including the hardware failure itself: a process holding 3.2.0
while disk says 3.3.0 refuses hockey, and reading fresh allows it. The inverse
is asserted too -- a genuinely old core still refuses, so the gate has not
become permissive. One test greps both modules for the old import-bound read;
reintroducing that line fails it, which is what stops this coming back.

Not changed: web_interface/__init__.py also imports __version__, but for
display rather than gating, and the API endpoint already reports a fresh
git describe.

* fix: drop the unused os import

Left over from a first draft that joined paths by hand before this used
pathlib. Flagged by CodeRabbit on #518; confirmed dead -- no os. reference
remains in the module.
2026-09-03 16:06:40 -04:00
353 changed files with 56450 additions and 15077 deletions
+6
View File
@@ -36,6 +36,12 @@ jobs:
uses: anthropics/claude-code-action@v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
# Review PRs opened by the Claude GitHub App. Without this the action
# aborts before reading the diff ("Workflow initiated by non-human
# actor"), so every such PR shows this check red. Named rather than
# '*': the allow-list is matched against the triggering actor, so
# this admits claude[bot] alone and no other app.
allowed_bots: 'claude'
plugin_marketplaces: 'https://github.com/anthropics/claude-code.git'
plugins: 'code-review@claude-code-plugins'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ github.event.pull_request.number }}'
+44 -3
View File
@@ -5,6 +5,9 @@ __pycache__/
# Secrets
config/config_secrets.json
# Atomic writes leave these behind when a save or a test is interrupted;
# the suite drops several per run.
config/.config_secrets.json.tmp.*
config/config.json
config/config.json.backup
config/wifi_config.json
@@ -36,11 +39,12 @@ htmlcov/
# Cache directory (root level only, not src/cache which is source code)
/cache/
# Development plugins directory
# Plugins are managed as separate repositories via multi-root workspace
# See docs/MULTI_ROOT_WORKSPACE_SETUP.md for details
# Development plugins directory: symlinks into a ledmatrix-plugins checkout
# See docs/PLUGIN_DEVELOPMENT_GUIDE.md and docs/MULTI_ROOT_WORKSPACE_SETUP.md
plugins/*
!plugins/.gitkeep
# Local settings for scripts/dev/dev_plugin_setup.sh (template: dev_plugins.json.example)
/dev_plugins.json
# Binary files and backups
bin/pixlet/
@@ -49,3 +53,40 @@ config/backups/
# Starlark apps runtime storage (installed .star files and cached renders)
/starlark-apps/
skin_renders/
# JS test deps (test/js)
node_modules/
package-lock.json
# Team logos fetched at runtime.
#
# src/logo_downloader.py and LogoHelper write into assets/sports/<league>_logos/
# whenever a plugin meets a team whose logo is not on disk. Those directories are
# also tracked -- 209 NCAA logos and 153 soccer ones ship with the repo -- so
# every rig accumulated untracked files it was never meant to commit and
# `git status` was permanently dirty. That noise is not harmless: it trains
# everyone to ignore the one signal that says a checkout is not what you think
# it is, which is how a stale tree sat unnoticed on a rig until a restart
# surfaced four dead sports plugins.
#
# Ignoring a directory does not untrack what is already in it, so the logos that
# ship keep shipping. Only new downloads are hidden.
#
# Adding a logo on purpose is rare and deliberate -- the last time was #415, four
# named NCAA logos a plugin needed, and there has been no other in a year. Do it
# with an explicit override:
# git add -f assets/sports/ncaa_logos/DUKE.png
assets/sports/*_logos/
assets/stocks/ticker_icons/
assets/stocks/crypto_icons/
# Plugin operation state written at runtime.
#
# web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json
# and data/operation_history.json as the web interface runs, into a directory that
# ships tracked (data/.gitkeep) and was otherwise unignored. So every rig that ever
# opened the web UI -- and every test run that constructs the app -- left three
# untracked files behind and a permanently dirty `git status`. Same reasoning as
# the logo rule above: a checkout that is always dirty is a checkout nobody reads.
data/*
!data/.gitkeep
+704
View File
@@ -17,8 +17,712 @@ release that ships it.
accepts both, but the store flags the old spelling as deprecated
(`store_manager.py`) and only the new one is in `schema/manifest_schema.json`.
## Unreleased
- `FontManager.get_font()` returns a BDF font at its native size when asked for
a size the file doesn't contain (5x7.bdf at 8 or 10px, say). It used to
return PIL's default font, a different typeface, so a plugin that relied on
that will now render the font it asked for.
- `src.wifi_manager.get_wifi_status_path()` — where WiFi status messages for
the display are written (`config/wifi_status.json`).
Deprecated, removed in 3.7.0 (each logs a warning on first use; see
`docs/PLUGIN_API_REFERENCE.md#deprecated-apis` for replacements). Nothing in
core, the monorepo or the registry's third-party plugins calls them:
- `CacheManager`: `has_data_changed`, `update_cache`, `setup_persistent_cache`,
`get_sport_live_interval`, `get_sport_key_from_cache_key`,
`get_background_cached_data`, `is_background_data_available`,
`record_cache_hit`, `record_cache_miss`, `record_fetch_time`,
`get_cache_metrics`, `log_cache_metrics`, `get_memory_cache_stats`.
- `DisplayManager`: `draw_weather_icon`, `draw_sun`, `draw_cloud`, `draw_rain`,
`draw_snow`, `draw_text_with_icons`, `get_scrolling_stats`.
- `FontManager`: `set_override`, `remove_override`, `get_overrides`,
`add_font`, `remove_font`, `validate_font`, `get_font_catalog`,
`get_available_fonts`, `get_size_tokens`, `get_performance_stats`,
`get_manager_fonts`, `get_detected_fonts`, `get_plugin_fonts`,
`unregister_plugin_fonts`.
- `PluginManager.get_enabled_plugins`.
## 3.5.0
New modules a plugin may import via `src.*` (floor on 3.5.0):
- `src/common/sports_helpers.py` — the helpers the scoreboards' `sports.py`
carry byte-identical copies of: `clamp_window`, `clamp_seconds`,
`logo_needs_refresh`, `spread_weighted_order` (+ `MIN_WINDOW_DAYS`,
`MAX_WINDOW_DAYS`), and `SportsHelpersMixin` with `_mode_customization`,
`_setting_int`, `_reset_dwell_on_reentry`, `_next_switch_index`,
`_spread_weighted_order`, `_odds_color`, `_upcoming_date_and_time_text` under
the plugins' names and signatures, plus the `_favorite_key` override point.
Constructor-free; keeps lazy state on its host (see the module docstring,
which also gives the host contract).
A new module rather than more methods on `sports_shared`: a plugin that
deletes a copy and leans on an older module having grown the method fails at
runtime with `AttributeError`, which no load-time check sees, while a missing
module fails at load. Nothing in core uses it yet.
- `test/test_common_is_hardware_free.py` — `src/common` must import without
`rgbmatrix` and never import `src.base_classes`, `src.display_manager` or
`src.plugin_system` at module level.
- `src/common/espn_dates.py` — `fetch_espn_scoreboard`,
`fetch_espn_date_chunks`, `espn_date_chunks`, `clamp_espn_limit`,
`ESPN_MAX_LIMIT`: fetch an ESPN scoreboard date range now that ESPN rejects
ranges (see Sports data below). Plugins bundle a copy of it.
### Config saves and plugin config preparation
- A JSON `POST /api/v3/config/main` changes only the keys it sends. The MQTT
bridge's brightness slider used to turn off `disable_hardware_pulsing`,
`inverse_colors`, `show_refresh_rate` and `use_short_date_format`, and a
timezone- or location-only save turned off web-UI autostart and weekly
automatic updates. Missing checkboxes still save as unchecked for the
settings forms (they now send a hidden `__form_section` field) and for
form-encoded posts.
- A partial JSON `POST /api/v3/plugins/config` merges onto the plugin's stored
settings instead of resetting everything it didn't send to the schema
defaults, and keeps a submitted `skin`, `skin_options`, `vegas_width_pct`,
`vegas_overflow` or `vegas_max_width_screens` (they were silently dropped).
- Plugin sections posted to `/config/main` are validated and prepared exactly
like `/plugins/config`; a value that endpoint rejects is rejected here too,
and nothing is saved.
- Legacy boolean settings (#588) are read as `{"enabled": ...}` objects
everywhere, not just when the plugin loads: `GET /plugins/config` returns
the object, posting it back saves, and hot reload hands plugins the same
shape (schema defaults included) they were constructed with.
`schema_manager.prepare_plugin_config` is the one implementation.
- A plugin's settings tab shows schema defaults for options its saved config
doesn't have yet. A boolean added with `"default": true` in a plugin update
(geochron 1.2.0's `show_date` and `show_date_line`) used to render unchecked,
and the next save of that tab stored it as `false`. Enum dropdowns likewise
showed their first option instead of the default. The partial now runs the
stored section through `prepare_plugin_config` like `GET /plugins/config`
(secrets are still masked, after the merge), and the form falls back to a
field's own `default` inside objects that declare a default of their own.
- `scripts/dev_server.py`, `check_plugin.py`, `render_plugin.py` and the plugin
harness build configs the way a device does: nested defaults are included,
a schema `enabled: false` no longer beats the forced `enabled: true` in the
dev server, and nested overrides such as `{"nhl": {"enabled": true}}` keep
the other defaults of that section.
- Clearing Vegas "Min/Max Cycle Time" no longer rejects the whole Display save,
and those fields no longer add junk entries to `display.display_durations`.
- Turning automatic updates on from the Raw JSON editor finishes their setup
like the General tab does, instead of waiting for the next display restart.
- `POST /config/schedule` and `/config/dim-schedule` accept the per-day
`days.<day>.{enabled,start_time,end_time}` shape their GETs return, as well
as the flat form keys.
- The startup check no longer warns that `auto_update` or `dim_schedule` is
"enabled but not found in plugins directory", and plugin ids that collide
with any core config section are flagged: the last private copies of the
core-key list now use `src/core_config_keys.py`.
### Sports data
- Since 2026-09-15 ESPN answers `dates=YYYYMMDD-YYYYMMDD` scoreboard queries
with `400 Bad Request` for every sport, so season schedules, the weeks window
and today's games all failed ("400 Client Error" from the NFL/NCAAFB managers
and `src.background_data_service`). A rejected range is now re-fetched as
whole months (`dates=YYYYMM`) plus the leftover days at each end, which cover
the window exactly: a football season is 8 requests. A month that returns
exactly 500 events is truncated and is re-fetched day by day.
- Scoreboard requests send `limit=500` at most. Above 500 ESPN silently returns
a short list: college football gave 25 of 68 games for one Saturday at the
`limit=1000` everything used to send.
- `BackgroundDataService.handles_espn_date_ranges` is `True`. Plugins check it
to decide whether to submit a season range to the service or fetch it
themselves on an older core.
- A league with no live games no longer backs its poll off past the next
kickoff. The escalation counted empty looks and nothing else, so a league
three hours before kickoff was indistinguishable from one out of season and
both reached `live_idle_max_interval`: measured gaps of up to 928 seconds,
and a rig that sat for a quarter of an hour with eight NFL games in progress
without noticing any of them. The wait is now clamped so it cannot run past
the earliest start still ahead, which the live fetch already downloads, so
it costs no extra request. Just after a kickoff the live cadence is held for
a grace window, because a provider that has not yet flipped the status would
otherwise read as another empty check and escalate the back-off again.
- ESPN date chunks are fetched six at a time (`ESPN_CHUNK_WORKERS`) in two
passes: months and edge days first, then the days of any month that came
back at the cap. A cold college-baseball season is about 130 requests, and
they went out one at a time; March and April measured on a Pi 4 (63
requests, 3101 events) went from 11.2s to 1.6s. Merged events still follow
`espn_date_chunks` order, so the payload does not depend on which request
won the race, and a capped month's payload is dropped before its days are
fetched, which keeps the peak memory of a four-capped-month fetch to about
16 MB over the sequential path rather than 43 MB — `docs/LOW_MEMORY_BOARDS.md`
puts a 1 GB Pi 3B+ at under 200 MB of headroom.
- `ESPNDataSource.fetch_standings` asks each league the endpoint that league
actually publishes. It tried `/standings` first whatever the league and fell
back to `/rankings` only on a 404, but college leagues answer `/standings`
with a 200 that carries no poll, so the fallback never fired: the rank badge
simply never appeared and anything keyed off rankings quietly did nothing.
Endpoints are now ordered by whether the league publishes a poll, and a 200
that lacks the key counts as a miss, so a league answering both still ends up
with whichever carries the poll. Only a 404 is routine — that is how a league
says it has none; a connection error, a timeout or an unparseable body is
logged as an error again, and a bug raised while inspecting the payload is no
longer swallowed as a missing poll. This is the implementation the football,
baseball and hockey boards already ship; core was the last copy on the old
one.
### Scrolling
- **Scoreboard scroll speed no longer changes with the General tab's "Scroll
Frame Rate" (`target_fps`).** Scoreboards on `src.common.sports_scroll`
computed their speed for that rate while the panel kept presenting at its
real refresh, so on a 100 Hz panel 60 ran a 50 px/s scoreboard at 100 px/s
and 200 ran it at 25 px/s. Speed now comes from `scroll_speed` and the panel
refresh only. The field is labelled legacy: nothing in core scrolling reads
it. Anyone who lowered it will see scoreboards scroll slower than before --
at the speed they configured.
- `scripts/scroll_speeds.py --measure` / `--demo` open the panel with the
display service's own options (`DisplayManager.apply_matrix_options`), so
`display.runtime.gpio_slowdown`, `rp1_rio`, `panel_type` and orientation are
honoured; the script used to read `gpio_slowdown` from `display.hardware`.
Its closing advice now gives the `scroll_speed` + `scroll_delay` pair
instead of `scroll_pixels_per_second`, which the resolver ignores whenever
the pair is present.
- The frame-stats log no longer opens a scroll with a one-frame window for
scrollers that never call `reset_scroll()`.
- Removed dead scroll code: the optional scipy import (`HAS_SCIPY`),
`ScrollHelper._last_integer_position` and `frame_time_target`.
`ScrollHelper.target_fps` / `set_target_fps()` remain, documented as
informational.
- Docs describe the fixed-step scroll model: `PLUGIN_API_REFERENCE.md`
documents `set_scrolling_state(..., frame_hold)` (omitting the hold runs a
scroll `frame_hold` times too fast), `SCROLL_PERFORMANCE.md` no longer reads a
held 20 ms frame as missed refreshes, and Vegas `frame_based_scrolling` /
`scroll_delay` are described as the speed clamp they are rather than frame
stepping. Scoreboard `scroll_delay` is documented as ignored for pacing.
### Web interface
- The plugin settings form honours `"x-display": "hidden"` in config schemas:
the property gets no control at any depth (top level, nested objects, array
rows, Advanced Settings), and saving the form never changes its stored value.
JSON API saves are unaffected. Lets plugins keep deprecated or internal keys
declared, e.g. countdown's row `id` and weather's `api_key` / `radar_zoom`.
See `docs/widget-guide.md`.
- Display settings no longer silently cut values on save: columns were capped
at 128, chain length at 24 and PWM LSB nanoseconds at 500. Columns have no
upper limit, chain length is 1–255 and rows must be even and 8–64 (see
"Display hardware settings the library refuses" below);
parallel is 1–3 and PWM dither bits 0–2, matching the library. A stored GPIO
slowdown, PWM dither bits or refresh-rate cap of 0 no longer shows (and
re-saves) as 3, 1 or 120, and the refresh cap accepts 0 (no cap). The config
API rejects out-of-range or non-integer `rows`, `cols`, `chain_length`,
`parallel`, `brightness`, `scan_mode`, `pwm_bits`, `pwm_dither_bits`,
`pwm_lsb_nanoseconds`, `limit_refresh_rate_hz`, `row_address_type`,
`multiplexing` and `gpio_slowdown` with a 400 (JSON `true` or `5.5` used to
save as 1 or 5) instead of saving a config the matrix refuses to start with.
- Display setting help tips and README / config-reference entries corrected
and completed: `panel_type` and `rp1_rio` are documented,
`show_refresh_rate` prints to the console rather than drawing on the panel,
PWM dither bits raise the refresh rate rather than lowering it, and every
numeric setting states its range.
- Row Address Type offers 5, the SM5368 / B707 row shift register. The
Waveshare 96x48 V2 panel (back silkscreen `24S-A1`) needs it with RGB
sequence BGR and, on a Pi 4, a GPIO slowdown of 6–8. Panels with FM6124
column drivers need no Panel Type.
- On a Raspberry Pi 5 the pinned rgbmatrix library can drive only row address
types 0 and 2, parallel 1–3 and the standard mappings. For anything else it
returns no matrix, which the Python binding doesn't catch, so the display
service crashed and restarted every 10 seconds. `DisplayManager` now refuses
those settings before creating the matrix (logged, reported by
`/api/v3/hardware/status`, fallback mode), the config API rejects them, and
the Display form offers only row address types 0 and 2 on a Pi 5. The rule
lives in `src/pi5_matrix_support.py` and must be re-checked when the
submodule is bumped.
- The Plugin Config Warning no longer lists core settings as plugins that are
"in config but not installed" (seen as `auto_update` on 3.4.0, where the
advice would have deleted the weekly-update setting). Core top-level config
keys now live in one list, `src/core_config_keys.py`, which reconciliation
uses and tests pin to `config.template.json` and the settings save endpoint.
A stored warning is also dropped once its entry is no longer a plugin in
config, so an old verdict clears without a restart.
- **Check & Update All** no longer sends installed Starlark apps
(`starlark:<app_id>` entries in `/plugins/installed`) to the plugin updater,
which answered each with a 500 "plugin not found". `POST /plugins/update`
now answers a `starlark:` id with a 400 saying it is a Starlark app. A
request that gets no HTTP answer (e.g. the web service restarting mid-run) is
re-sent with backoff instead of being counted as failed and skipped — that is
how a disabled plugin with an update waiting was silently left out.
- Three routes consulted the web process's plugin manifests without
discovering plugins first, so they misbehaved from every `ledmatrix-web`
restart until something else ran a discovery — in practice until someone
opened the dashboard, measured at over three minutes on one rig.
`POST /display/on-demand/start` and `POST /plugins/toggle` answered 404
"Plugin not found", and `POST /config/main` did not recognise a plugin
section, so it skipped secret separation and wrote the plugin's API key to
`config.json` in plain text instead of `config_secrets.json`. The routes now
discover when nothing has been discovered yet, and rescan once when a
specific plugin id (or, for on-demand by mode, a mode) is not found, so a
plugin installed since the last scan is found too.
### Security (request paths and inline handlers, siblings of #561)
- `POST /api/v3/plugins/assets/upload`, `GET .../assets/list` and
`POST .../assets/delete` validate `plugin_id` with `src/common/path_safety`
and answer 400 otherwise. A `plugin_id` of `../../config` used to create an
`uploads/` directory outside `assets/plugins`, write images and
`.metadata.json` there, list it, and delete whatever file a metadata entry
named. Delete now unlinks only a path that resolves inside that plugin's
uploads directory (any other entry is dropped without touching a file).
- `PluginManager.get_plugin_directory()` returns `None` for anything but a
plain name, so `POST /api/v3/plugins/action` can no longer run a manifest
script from a directory outside the plugins directory (`../elsewhere`); the
route also rejects such ids with 400.
- Plugin Store, saved-repository and custom-registry buttons escape registry
values for their inline `onclick` handlers (`jsStringAttr` in
`plugins_manager.js`). An entry id containing `'` used to close the attribute
and add its own script. The store's View button opens only `http(s)` links.
- The uploaded-images list escapes each file's original name, path and ids; a
name like `<img src=x onerror=...>.png` was inserted as markup.
### Display hardware settings the library refuses
- The rgbmatrix library answers several settings with no matrix or `abort()`
rather than an error, on every board, so the display service crash-looped
instead of falling back: rows above 64, `chain_length` above 255 (the Python
binding stores it in one byte; this was documented as "no upper limit"), a
misspelled `hardware_mapping`, and `parallel` 2–3 on a mapping with one output
(`adafruit-hat`, `adafruit-hat-pwm`, `regular-pi1`, `classic-pi1`) — the last
one reachable from the Display form on the default mapping. The config API
now refuses them with a 400 naming the setting, and `DisplayManager` refuses
a hand-edited one before creating the matrix: logged, fallback mode, reported
by `/api/v3/hardware/status`. The rules, including the Pi 5 ones, live in
`src/matrix_support.py` and must be re-checked when the submodule is bumped.
- `/api/v3/hardware/status` adds `cause`: `"settings"` when LEDMatrix refused
the config, `"library"` when the library failed. The Display tab banner and
the fallback log line give the Pi 5 rebuild hint only for a library failure;
they used to follow every failure with it and with GPIO slowdown advice.
- The Display form offers the `classic` and `classic-pi1` mappings and the
`90` / `270` orientations, and renders any other stored mapping selected with
a warning. With no option selected the browser posted the first one, so one
unrelated save rewrote those settings. The API accepts orientation `90` and
`270`, which `DisplayManager` already applied.
- The display size the web preview, Starlark magnify default and
`scripts/dev/vegas_audit.py` compute (`src/display_geometry.py`) now applies
`orientation` and `pixel_mapper_config` as the library does: `Rotate:90`
swaps width and height, `U-mapper` folds the chain.
- One Raspberry Pi 5 GPIO slowdown recommendation everywhere: 1–3 in PIO mode,
starting at 1. README and the config reference now describe the template
values as the defaults; the "code default" values they listed never apply,
because config migration fills missing keys from the template.
### Plugin system
- A plugin no longer starts with a schema warning and a degraded flag because
config.json still holds a boolean where its schema now has an object with an
`enabled` property (news' `global.dynamic_duration: true`). The loader reads
the boolean as `{"enabled": <bool>}` before merging schema defaults and
validating, the same rule the settings form already applies
(`legacy_bool_as_object` in `src/plugin_system/schema_manager.py`). Nothing
is written at load; the next save of that plugin's settings stores the object.
Other type mismatches still warn.
### Core
- `ConfigManager.load_config()` no longer raises on a host without the POSIX
ownership APIs. The self-heal that chgrp's `config_secrets.json` to the
shared group (added in #416) looked up `os.geteuid` unguarded; that name does
not exist on Windows, and the resulting `AttributeError` is not an `OSError`,
so it escaped the helper's own "best-effort" handling and every caller's.
Any Windows checkout with a `config/config_secrets.json` got a `ConfigError`
from every config load and could not `import web_interface.app` at all.
`ensure_shared_group_ownership()` now returns immediately when `os.geteuid`
or `os.chown` is missing. No behaviour change on the Pi.
- Restoring a backup on Windows no longer fails over files that already exist.
The restore carries each replaced file's owner across with `os.chown`, which
does not exist on Windows; the `AttributeError` escaped the per-file error
handling, so the restore stopped at `config.json` with nothing restored. The
ownership step is now skipped where `os.chown` is missing. No behaviour
change on the Pi.
### Cache permissions
- The web interface can read what the display service caches again.
`ledmatrix-web.service` carried `CacheDirectory=ledmatrix`, and systemd
re-owns `/var/cache/ledmatrix` and its contents to the unit's `User=`
whenever the directory's owner differs, which erased the `root:ledmatrix`
setgid layout the installers set up: every file the root display service
wrote afterwards was `root:root` 0660 and unreadable by the web interface
(392 unreadable files on one rig, with display status, on-demand state and
plugin health empty). Since #547 the web unit is rendered from its template
on every install, so every fresh install hit this.
`DiskCache.set` now gives each file the directory's group (when that
directory is group-writable) and 0660 on the open descriptor before the
rename, independent of setgid, which also closes a window where a fresh
file was visible as mkstemp's 0600. `DiskCache.share_existing_files`
repairs files an older version left behind, once per process, through
`O_NOFOLLOW` descriptors, skipping hard links and other users' files.
Existing installs only ever receive `git pull`, so that repair is the fix
for them; new installs also drop `CacheDirectory=` and
`CacheDirectoryMode=` from the web unit.
- `install_web_service.sh` replaces an existing cache directory's group
whenever the installing user is not in it. It used to replace only root's,
so a `root:ledmatrix` directory belonging to a user outside that group was
left alone and everything root wrote there stayed unreadable.
- `/display/on-demand/status` and the current-display status read the display
service's keys with `memory_ttl=0`, as every other cross-process reader
already does. They served the first copy the web process had read for the
full 120s `max_age`, so on-demand reported "active" for over 100 seconds
after the file on disk said "idle".
### Automatic updates and Update Code
- An update that changes `web_interface/requirements.txt` is no longer rolled
back on every auto-updating device. `safe_pip_install.sh` allowed only the
root `requirements.txt`, so the install Update Code and the health check run
for the web requirements was refused, and the health check rolls back any
update whose dependencies failed (Install Base Requirements failed the same
way). The wrapper now allows both core requirement files; a core requirement
file symlinked out of the project is refused.
- The automatic update's local-change check and Update Code now count changes
the same way (`auto_update.local_changes`): permission-only changes and
anything under `plugins/` or `plugin-repos/` don't count, and a core path
that merely contains `plugins/` does. Such edits used to pass the check and
then be stashed by the pull and never restored, despite "will not stash your
changes". The pull's `--autostash` now carries them across. Update Code
still stashes other edits; the automatic update refuses instead.
- When the automatic update's own rollback fails (a partial pull, or a health
check that never started), plugins are no longer updated and the display is
not restarted, as the 3.4.0 notes promised.
- The health check's dependency reinstall no longer retries pip failures or
timeouts with a second bash path, and all reinstalls share a 10-minute
budget, so a rollback finishes inside the unit's 30-minute limit instead of
being killed mid-way.
### Installers
- The generated `ledmatrix_web` sudoers rules are parsed before they are
installed. Both installers built the drop-in from `which` lookups and copied
it into `/etc/sudoers.d` without ever checking it, and a malformed file there
makes sudo refuse every command for every user — on a headless Pi, that is
unrecoverable over SSH. `first_time_install.sh` now runs `visudo -c` on the
generated file and, if it does not parse, prints what visudo said and leaves
the installed file untouched instead of replacing it with a broken one;
`configure_web_sudo.sh` does the same before offering the rules for
confirmation. `first_time_install.sh` also built that file at a fixed `/tmp`
path as root; `mktemp` now picks the name.
### Small fixes (update-all, plugin system settings, scripts)
- **Check & Update All** counts a plugin that had nothing to update as
"already up to date" instead of "updated". ZIP-installed monorepo plugins
(most official ones) already at the registry version were called "updated
successfully" on every run. `POST /plugins/update` now returns
`data.update_status` (`updated`, `up_to_date`, `local_only`).
- An update request that got an HTTP error answer without an `error_code`, or
a body that is not JSON (e.g. a reverse proxy's 502 page), is no longer
classified as `NETWORK_ERROR` and re-sent five times. Only a request that got
no HTTP answer is retried; the rest are `API_ERROR` with the HTTP status.
- The General tab no longer shows Auto Discover Plugins, Auto Load Enabled
Plugins or Development Mode. Nothing read `plugin_system.auto_discover`,
`auto_load_enabled` or `development_mode`: every enabled plugin was always
discovered and loaded. Stored values are kept, and saving the General tab no
longer rewrites them to `false`.
- `BackgroundDataService` shares the 6-hour "ESPN rejects date ranges" memo
with `fetch_espn_scoreboard`, so a background season fetch no longer spends a
doomed range request first once either path has seen a rejection.
- `scripts/install_plugin_dependencies.sh` installs from the configured
`plugin_system.plugins_directory` (default `plugin-repos`, where the Plugin
Store installs) and also scans `plugins/` for dev symlinks. It used to scan
only `plugins/` and find nothing. A failed `pip install` is now reported as a
failure instead of being hidden by `tee`.
- `scripts/verify_installation.sh` no longer fails a healthy install: it
checked for the removed `web_interface_v2.py` and port 5001. It and
`scripts/verify_web_ui.sh` now check port 5000, where the web interface
listens.
- `scripts/install/install_service.sh --help` prints usage and exits without
changes. It used to ignore the flag and reinstall and restart every service.
Unknown arguments are rejected before anything runs.
- `scripts/diagnose_web_ui.sh`, `scripts/diagnose_web_interface.sh` and
`scripts/debug/debug_web_manual.py` apply the launcher's own autostart rule
(only an explicit `web_display_autostart: false` keeps the web interface
down), so a missing key no longer shows as disabled. The shell scripts also
check `web_interface/blueprints/api_v3/`, which became a package, instead of
reporting `api_v3.py` as missing.
### Docs and developer tools
- `docs/REST_API_REFERENCE.md` rechecked against every handler: request
fields that made documented calls fail (`repo_url`, `action_id`/`params`,
`files`/`image_id`, `font_file`+`font_family`, `?font=`, cache `key`,
`auto_enable_ap_mode`, plugin limit keys) and response shapes are fixed, the
removed font-override endpoints are gone, and the 26 undocumented routes
(backup, git/auto-update, WiFi radio, Starlark editor, MQTT bridge, status
endpoints, skins) are listed. Store search is `/plugins/store/list?query=`.
- `FONT_MANAGER.md` no longer tells plugins to read
`display_manager.font_manager`, which does not exist; use
`plugin_manager.font_manager` / `BasePlugin._get_font_manager()`.
- Plugin docs, `DisplayManager` docstrings and the bundled `starlark-apps`
plugin now all read the display size from `display_manager.width/height`,
which works in fallback mode where `matrix` is `None`.
- `scripts/dev/dev_plugin_setup.sh link-github <name>` links the plugin from a
clone of the `ledmatrix-plugins` monorepo (per-plugin `ledmatrix-<name>`
repositories no longer exist). `dev_plugins.json` honours `github_user`,
`plugins_repo` and `plugins_branch`; `dev_plugins.json.example` ships and
`dev_plugins.json` is git-ignored. `update`/`status` handle monorepo links,
and `status` no longer exits 1 when nothing is broken.
- Rewritten for current behaviour: plugin dependency installation (web service
runs as the installing user and installs through `safe_pip_install.sh`),
`PLUGIN_CONFIG_ARCHITECTURE.md`, `MULTI_ROOT_WORKSPACE_SETUP.md`; stale
`app.py` line numbers, `api_v3.py` paths, StreamManager method names,
nonexistent version-bump scripts and `ledmatrix` service user references
removed.
## 3.4.0
Plugin-facing changes since 3.3.0 (tag `v3.3.1`) not covered further down:
- `BasePlugin.get_update_interval()` (#555) — return seconds to override the
manifest's `update_interval` at runtime (e.g. poll fast only while a game is
live), or `None` to keep it. Clamped to at least 5 seconds; a raising or
non-numeric return is ignored. Called every scheduling tick, so keep it
cheap. Older cores never call it. See `docs/PLUGIN_API_REFERENCE.md`.
- `src.common.scroll_config` (#523) — turns a plugin's scroll config into a
configured `ScrollHelper` in one place, replacing per-plugin resolution that
disagreed between tickers, and warns when a speed won't advance whole pixels
per panel refresh. Floor on 3.4.0 to import it.
- **Skins are marked unsupported.** No current scoreboard plugin builds on
`src.base_classes`, so the skin hook (`SportsCore._render_game`) never runs.
The web UI no longer shows the Visual Skin dropdown, the store hides and
refuses `"type": "skin"` entries, and `GET /api/v3/skins` reports
`"supported": false`. Saved `skin` config values still load and save.
`src/skin_system/` is unchanged.
- **Web preview size** now comes from `src/display_geometry.py`, the same
computation `DisplayManager` uses: double-sided setups preview one screen,
and a missing `chain_length` defaults to 2 everywhere (the Starlark magnify
default and the sync handshake used 1). The module is core-internal: plugins
keep reading `display_manager.width`/`height`.
- `src.common.font_layout` (#539, #565) — `load_truetype()` is
`ImageFont.truetype` with the layout engine pinned, so text lays out the same
whether or not the host's Pillow was built with libraqm; `crisp_size()` and
`FONT_PIXEL_GRID` give the size a bundled face renders on whole pixels at
(`sports_card` still re-exports them); `resolve_asset_path()` resolves
`assets/fonts/...` against the install root, not the working directory.
Floor on 3.4.0 to import it. Relatedly, `DisplayManager` now draws text
1-bit (#521), so golden images recorded against 3.3.x may need regenerating.
### Install and updates
**Weekly automatic updates (#581), off by default.** Switching on
*Automatically check for and install updates once a week* on the General tab
(or `first_time_install.sh --enable-auto-update` / `LEDMATRIX_AUTO_UPDATE=1`)
updates the core and then every installed plugin once a week, preferably 2–5 AM
local time. It follows the branch the checkout tracks — `main` on a standard
install — so a device gets whatever has merged there, not only tagged releases.
See `docs/WEB_INTERFACE_GUIDE.md`.
- The core step is skipped, with the reason shown, when the checkout has local
edits or commits, a rebase or merge is in progress, the branch has no
upstream, less than 300 MB is free, or that commit was already rolled back.
- After pulling, `ledmatrix-update-verify.service` restarts the services and
requires the web interface to answer and the display to stay up. If they
don't, or the new requirements fail to install, it resets to the previous
commit, reinstalls its requirements and restarts again. Anything but success
shows under the toggle and as a banner on Overview.
- Plugins update through the Plugin Store even when the core step is skipped,
fails or is rolled back. A plugin version whose `ledmatrix_min_version` is
above the device's core is held back, not installed. When the core did
update, plugins wait for its health check, and are left alone if that check
never reports or the rollback fails.
- No SSH is needed: switching the toggle on restarts the display service, which
installs the health-check units (`src/auto_update_setup.py`, core-internal
and not a plugin API).
Installer and service fixes:
- rgbmatrix builds on ARMv6 boards (Pi Zero, Pi 1); an existing checkout is
moved forward to the new pin and no longer left root-owned (#577).
- `first_time_install.sh` grants the web user `safe_pip_install.sh`, as
`configure_web_sudo.sh` already did, so plugin requirements install where
the display service can see them (#579).
- The web interface starts when `web_display_autostart` is missing or
`config.json` is unreadable; only an explicit `false` keeps it down (#556).
- Installers render every systemd unit from its `systemd/` template, so the
boot-time unit-drift warning can clear, non-root installs included (#547).
### Scrolling
- **Frame pacing (#523).** The loop waits only for the rest of each panel
refresh instead of a flat 8 ms: 44–46 fps → 100 fps, and slow frames 14% →
0.02%, on a 2×128×64 chain. Sub-pixel blending is off by default again (it
shimmered on pixel fonts; Vegas mode still opts in).
- **Whole-pixel steps (#545).** At a speed `scroll_config` can render in whole
pixels, every frame advances by exactly the same amount, removing about six
hitches a second. A loop that can't keep up now scrolls slightly slow rather
than jumping.
- The eight sports scoreboards scroll through `scroll_config` too (#542): the
default 50 px/s holds each frame for two refreshes instead of alternating
0 px and 1 px steps.
- **Frame stats ignore the pause between scrolls (#582).** The `Scroll frame
stats` log line counted the idle wait before each scroll as one frame,
inflating `max` and the stall rate. `docs/SCROLL_PERFORMANCE.md` now
describes the line actually logged.
### Plugins
- `FontManager` registers the bundled `tom_thumb` font, so plugins no longer
need a private loader (#534).
- The test harness's `set_scrolling_state()` accepts `frame_hold`, as
`DisplayManager`'s does (#534).
- A `display()` with nothing to draw should return `False`, the only value the
controller skips on; starlark-apps now does, rather than holding a black
panel (#534).
- Starlark apps may set `render_width`/`render_height` in their `config.json`
to render at their own canvas size instead of Pixlet's 64×32 (#552).
- `scripts/render_plugin.py --display-mode <mode>` renders one mode of a
multi-mode plugin; scoreboards previously rendered blank (#522).
- Scoreboards resolve their own directory under the real plugin loader
(declare `_PLUGIN_DIR`), so 4x6 text snaps to its 7px grid instead of
rendering a pixel narrow, and an unreadable schema is logged (#519, #520).
`DisplayManager` loads 4x6 on that grid too (#565).
- The 5x7 BDF face reports a real height, so rows stacked by
`get_font_height()` no longer overlap (#539).
- `LogoHelper` remembers a missing logo instead of warning every rotation
(#548), and the decoded sports logo cache is bounded (#559).
### Web interface
- Installed Plugins has search, All / Enabled / Disabled / Updates filters and
sort (#540).
- Hardened and polished per the September 2026 audit (#568): utility classes
such as `.hidden` actually exist, focus rings, labels and modal focus
trapping, dark theme throughout, no overflow at phone width, and background
streams pause when hidden, with first-load JS/CSS down from 1358 KB to 291 KB.
- WiFi Connect works from the LEDMatrix-Setup hotspot: the page is answered
before the hotspot drops, and reopening it shows why an attempt failed (#571).
- Pixlet install, the Starlark app store and app toggles work again (#535,
#537); the store uses the configured GitHub token and reports a rate limit
instead of drawing a blank grid (#541).
- Plugin config: geochron and news saves no longer always fail (#575), the page
survives stored values the schema outgrew (#578), the form uses the full page
height (#573), and file-manager widgets show the script's error (#574).
- The live status stream reports real disk usage and available memory (#558);
a system action refused for want of passwordless sudo says so and names
`configure_web_sudo.sh` (#560).
### Tools and security
- **CodeQL triage (#561):** 129 of 134 alerts fixed. Three were exploitable
path-handling flaws in the web interface and are closed; web UI escapers now
escape quotes, and URL fields refuse script schemes. Path checks share
`src/common/path_safety.py` (core-internal).
- **Home Assistant MQTT bridge** (`integrations/mqtt_bridge`, #538): mode
select, stop, power and brightness over MQTT Discovery.
- **Tools tab** manages the MQTT bridge and the Pixlet editor (#554); the
editor stays on loopback when `PIXLET_EDITOR_HOST` says so.
### Fixes
- Updating a plugin whose directory is named for its manifest id (leaderboard,
music, stocks, weather) silently did nothing (#536).
- Plugin reconciliation no longer reports working plugins as stale or replaces
their config with a stub, and the Overview banner advises each case correctly
(#557).
- Two config saves in the same second no longer share one backup, so rollback
restores the version asked for (#564).
- On-demand: a second request is honoured without a restart (#534), a pinned
request stays on its mode, and restarting mid-session loads every plugin
again (#538).
- `/health` and `/display/current` report real state, and the preview no longer
freezes on a leftover snapshot temp file (#534).
### Per-element display customization
**Per-element display customization, and the last mile of it into the web UI.**
A user can set the font, size, colour, position, visibility and alignment of
individual display elements per plugin -- and, where a plugin has display
modes, separately per mode.
New public API a plugin may import via `src.*` (floor on the release that
ships this):
- `src.element_style.layout_offset(config, element, axis, default, mode)` and
`element_color(config, element, default, mode)` — the stateless reads the
scoreboard helpers share. There were three copies of the offset read and two
of the colour read; these are the one implementation, and they carry the
element-name aliasing and the per-mode lookup.
- `src.element_style.alias_keys(element)` — the names one element may be stored
under. The style block names elements `score_text` while the layout block
says `score`, and `records`/`record` and `status_text`/`status` split seven
to two across the published schemas. A lookup tries the exact name first, so
this is inert for a config that already matches.
- `src.element_style.element_visible(config, element, default, mode)`,
`element_align(...)` and `element_scale(...)` — the stateless reads for the
three knobs the resolver already understood but no draw path consumed, so an
element could be marked hidden in the web UI and still render.
- `SportsCoreSharedMixin._draw_text_with_outline(..., element="score_text")` —
naming the element resolves its colour by name and honours its visibility
toggle. Without a name the colour is inferred from font-object identity,
which cannot separate two elements sharing a face; that is the case every
bitmap font is in, because a `freetype.Face` cannot be re-instantiated, and
it is how a BDF-rendered element silently lost a configured colour. Shared
faces now resolve when exactly one sharer has a colour set.
- `LogoHelper.load_logo(..., scale=)` — applies a user's image scale, and keys
the cache on the scaled box so two elements scaled differently cannot be
served each other's image.
- `src.element_style.native_bdf_size(font)` — the one pixel size a bitmap font
can render at, or None for a scalable one. The web UI needs this to know
whether a size control can take effect at all.
- `ElementStyleResolver(config, defaults, mode=...)` plus `visible`, `align`
and `scale` on `ElementStyle`. The mode binds to the resolver rather than
being passed per call, so a plugin with one instance per mode makes every
existing lookup mode-aware by setting one class attribute.
- `BasePlugin.styles` / `styles_for(mode)` / `STYLE_MODE` — the accessor every
plugin inherits, so adopting this is no longer a guarded import plus schema
discovery plus resolver invalidation in each plugin.
- `SportsCoreSharedMixin._get_layout_offset` — promoted from the plugins'
bundled copies. Each still carries its own, which wins by MRO, so adopting
it is a deletion.
Schema and web UI:
- A `customization` block is now rendered by a composite style editor: one row
per element rather than nested accordions, with a tab per declared mode.
Plugins that hand-wrote their style blocks get it without a plugin release;
`x-style-elements` and `x-style-modes` declare it compactly.
- Font fields become a real picker rather than a hardcoded `enum`, so a font
the user uploads is selectable. Bitmap fonts taller than the element's
declared size ceiling are filtered out, because a bitmap font ignores
`font_size` and renders at its own size.
- `/static/plugin-widgets/<plugin>/<widget>.js` serves a plugin's own web-UI
widgets. The client half and the docs already existed; nothing served them.
Fixed:
- A bitmap font asked for a size it has no strike for fell back to
*PressStart2P* — a different typeface — rather than to its own native size.
32 of the 35 shipped fonts are bitmap, so this was reachable for most font
choices.
- The plugin config form read `config_schema.json` directly while the save
route read it through `SchemaManager`. Only the latter expands a compact
`x-style-elements` declaration, so a plugin using that form had a
customization section that rendered as empty space.
- `unshare_element_fonts` rebuilt faces through bare `ImageFont.truetype`,
bypassing the layout engine `src/common/font_layout.py` pins. These were the
only two call sites in `src/` doing so.
- The form parser compared a schema type to a bare string, so a nullable field
(`["array", "null"]`) never had its indexed colour inputs recombined, and a
blank one became `[]` rather than null.
Removed:
- The Fonts tab's "Element Font Overrides" panel and its three endpoints. They
reported success and saved nothing, and the element keys the panel offered
(`nfl.live.score`, `clock.time`) are read by no plugin, so wiring them to the
real `FontManager` methods would still have changed nothing on the panel.
Per-element font choice now lives in each plugin's own config editor.
- "Detected Manager Fonts", which listed every installed font with a hardcoded
usage count.
- Two dead client-side config-form renderers in `app-shell.js` (~580 lines) and
the legacy `plugins/config_manager.js`, superseded by server-side rendering.
## 3.3.0
Historical note: tags `v3.3.0` and `v3.3.1` both report `__version__` "3.3.0" and both ship `src/common/sports_shared.py`, so a "3.3.0" floor always means a core with `sports_shared`.
**The release the sports scoreboards floor on to delete their bundled copies.**
3.2.0 shipped the unified sports library and made `ledmatrix_min_version`
enforceable; this ships the last three shared modules and completes the store
+9 -5
View File
@@ -6,7 +6,7 @@
- `config/config.json` — User plugin configuration (persists across plugin reinstalls)
- `plugin-repos/` — **Default** plugin install directory used by the
Plugin Store, set by `plugin_system.plugins_directory` in
`config.json` (default per `config/config.template.json:167`).
`config.json` (default per `config/config.template.json`).
Not gitignored.
- `plugins/` — Legacy/dev plugin location. Gitignored (`plugins/*`).
Used by `scripts/dev/dev_plugin_setup.sh` for symlinks. The plugin
@@ -23,14 +23,14 @@
- Each plugin needs: `manifest.json`, `config_schema.json`, `manager.py`, `requirements.txt`
- Plugin instantiation args: `plugin_id, config, display_manager, cache_manager, plugin_manager`
- Config schemas use JSON Schema Draft-7
- Display dimensions: always read dynamically from `self.display_manager.matrix.width/height`
- Display dimensions: always read dynamically from `self.display_manager.width/height` — not `display_manager.matrix.width/height`, because `matrix` is `None` when hardware init fails (the properties fall back to the canvas size)
- Secrets: namespaced by plugin id in `config/config_secrets.json`, declared
via `"x-secret": true` in the plugin's config schema, and deep-merged into
the plugin's config dict at load time — plugins read them with plain
`config.get(...)`, never a separate accessor
## Dev Workflow
- Link a plugin for development: `./scripts/dev/dev_plugin_setup.sh link-github <name>` (or `link <name> <path>`); symlinks land in `plugins/` — set `plugin_system.plugins_directory` to `plugins` so discovery picks them up
- Link a plugin for development: `./scripts/dev/dev_plugin_setup.sh link-github <name>` clones the `ledmatrix-plugins` monorepo into `~/.ledmatrix-dev-plugins/` and links its `plugins/<name>` under the manifest id (add a repo URL for a plugin with its own repo; or `link <name> <path>`); symlinks land in `plugins/` — set `plugin_system.plugins_directory` to `plugins` so discovery picks them up. Fork/location overrides: `dev_plugins.json` (from `dev_plugins.json.example`)
- Browser preview without the display loop: `python3 scripts/dev_server.py` → http://localhost:5001
- Full display in emulator mode: `python3 run.py -e` (or `EMULATOR=true python3 run.py`)
- Validate one plugin headlessly: `python3 scripts/check_plugin.py --plugin <id>`
@@ -45,9 +45,12 @@
- Plugin configs stored in `config/config.json`, NOT in plugin directories — safe across reinstalls
- Third-party plugins can use their own repo URL with empty `plugin_path`
## Skin System (visual overlays for sports scoreboards)
## Skin System (visual overlays for sports scoreboards) — NOT SUPPORTED YET
- Skins do not render with the current scoreboard plugins: the only hook is `SportsCore._render_game()` in `src/base_classes/sports/core.py`, and no current scoreboard plugin (monorepo or third-party registry) builds on `src.base_classes`
- So core doesn't offer them: no Visual Skin dropdown (`get_plugin_schema` skips `inject_skin_selector`), the store hides/refuses `"type": "skin"` entries, `GET /api/v3/skins` reports `"supported": false`. Switch: `SKINS_RENDER_SUPPORTED` in `src/skin_system/__init__.py`
- Stored `skin` / `skin_options` config values must keep loading and saving (base schema allows them; form saves deep-merge over the stored section)
- Skins live in `skins/<skin-id>/` (skin.json + skin.py), NOT in plugin dirs — plugin reinstall deletes plugin dirs
- Core: `src/skin_system/` (ScoreboardSkin, SkinContext, runtime); hook: `SportsCore._render_game()` in `src/base_classes/sports/core.py`
- Core: `src/skin_system/` (ScoreboardSkin, SkinContext, runtime); keep it and its tests
- Skins render onto `ctx.canvas` only; fallback to built-in renderer on `False`/exception (3 strikes disables for session)
- View-model guaranteed keys are frozen (see `test/test_skin_system.py::TestViewModelContract`) — renaming keys in `_extract_game_details_common` or sport extractors breaks published skins
- Validate skins headlessly: `python scripts/validate_skin.py --skin <id>`; docs: `docs/SKIN_SYSTEM.md`, `docs/CREATING_SKINS.md`
@@ -60,3 +63,4 @@
`self.display_manager.image.paste(img, (x, y))` then `update_display()`
(use a mask for transparency: `image.paste(rgba, (x, y), rgba)`)
- When modifying a plugin in the monorepo, you MUST bump `version` in its `manifest.json` and run `python update_registry.py` — otherwise users won't receive the update
- `src/pi5_matrix_support.py` hardcodes what the pinned `rpi-rgb-led-matrix-master` can drive on a Raspberry Pi 5 (`Rp1PioConfigSupported()` in `lib/rp1/rp1_pio_backend.cc`). Re-check it whenever the submodule is bumped: a stale rule blocks Pi 5 settings the new library supports, and a missing one lets the display service crash-loop. `src/matrix_support.py` holds the same kind of rules for every board (rows, chain length, mapping names, parallel per mapping) and needs the same re-check
+74
View File
@@ -0,0 +1,74 @@
# Product
<!-- impeccable:product-schema 1 -->
## Platform
web
## Users
Designed novice-first, with power tools kept within reach.
- **Primary: hobbyist builders.** People who assembled an LED matrix panel on a Raspberry Pi, often by following the install video, and are frequently new to Linux and the Pi. They set the display up once (panel size, timezone, WiFi), install and enable a few plugins, then come back occasionally to tweak what the panel shows. They usually reach the control panel from a phone or laptop on their home network, sometimes as an installed home-screen app.
- **Secondary: tinkerers and plugin developers.** Comfortable with SSH, `config.json`, and GitHub. They lean on the Config Editor, Logs, Cache, Operation History, Tools, GitHub-repo installs, and per-plugin config while building or debugging. Their tools must stay reachable without sitting in the novice's path.
## Product Purpose
LEDMatrix turns a Raspberry Pi and an RGB LED matrix panel into an information-rich display (clock, weather, calendar, sports scores, stocks, music, and more) through a plugin platform. The web control panel ("LED Matrix Control") is where the display gets configured, extended, and kept healthy.
Success means a builder gets from a freshly flashed Pi to a working, personalized display without needing a terminal, and can keep it running (updates, recovery, troubleshooting) the same way.
## Positioning
Four strengths define LEDMatrix, and future work must protect all of them:
1. **Plugin ecosystem.** The core ships only `starlark-apps` and `web-ui-info`; everything else comes from the built-in Plugin Store (the official `ledmatrix-plugins` monorepo), third-party GitHub repos, or Starlark (Tidbyt-style) apps. Each installed plugin gets its own configuration tab, generated from its schema.
2. **Runs on tiny Pis.** The UI is served by the same device that drives the matrix, on boards as small as the Pi Zero 2 W (512 MB), Pi 3/3B+, and the 1 GB Pi 4.
3. **Recovers without SSH.** WiFi access-point fallback with a captive setup page, backup & restore, in-UI updates, live logs, diagnostics, service control, and plugin health let users fix problems from the browser.
4. **Open and community-led.** GPL-3.0, a Discord community, and contributions welcome. The maintainer (ChuckBuilds) builds in public and openly relies on AI development tools.
## Operating Context
- **Access.** Served on the local network at `http://<pi-ip>:5000` by the `ledmatrix-web` service. It is installable as a PWA (`web_interface/static/v3/manifest.json`, short name "LEDMatrix").
- **First run.** When the Pi has no network it creates its own WiFi access point, so the captive setup page (`templates/v3/captive_setup.html`) may be the very first screen a user sees, on a phone, with no internet connection.
- **Navigation.**
- System tabs: Overview, General, WiFi, Schedule, Display, Rotation, Config Editor, Backup & Restore, Fonts, Logs, Cache, Operation History, Tools.
- A second row holds Plugin Manager (with the Plugin Store), Starlark Apps, and one tab per installed plugin.
- **Live data.** The Overview shows system stats (CPU, memory, temperature, power/throttling) and a live display preview, streamed over SSE.
- **Getting Started checklist.** The Overview's first-run checklist runs: set panel size → set timezone → install a plugin → enable it → configure it.
- **Development.** `python3 scripts/dev_server.py` gives a browser preview without the display loop; `python3 run.py -e` runs the full display in emulator mode.
## Capabilities and Constraints
- **Hard constraint: plugin UI compatibility.** Third-party plugins rely on JSON Schema (Draft-7) generated config forms, the widget registry (`static/v3/js/widgets/`), `x-secret` fields, and plugin web-UI actions. UI changes must keep these working.
- **Config storage.** Plugin configuration lives in `config/config.json` and secrets in `config/config_secrets.json`, never in plugin directories, so configs survive reinstalls.
- **Stack.** An existing Flask + HTMX + Alpine.js app with Jinja templates (`web_interface/templates/v3/`) and static JS/CSS (`web_interface/static/v3/`), with self-hosted vendor assets.
- **Terminology.** Plugin, Plugin Store, Starlark app, rotation, display duration, Vegas Scroll Mode, skin, on-demand, AP mode.
- **Open decisions** (offered during init, not adopted as constraints):
- Whether the UI must work fully offline, with no CDN fallbacks at runtime.
- Whether a Node/CSS build step is acceptable for contributors.
- Whether a formal accessibility standard (e.g. WCAG 2.2 AA) is a requirement.
## Brand Commitments
- **Names.** The product is "LEDMatrix" and the web UI is titled "LED Matrix Control". The maintainer brand is ChuckBuilds.
- **Voice.** Friendly, honest, and learning-in-public, as in the README.
- **App icons.** They live in `web_interface/static/v3/icons/`.
No other visual identity has been made binding.
## Evidence on Hand
- **Photos.** Real photographs of running displays are linked in `README.md` (clock, weather, calendar, NHL/MLB/NFL/NCAA, stocks, music).
- **Video.** YouTube install and walkthrough videos from ChuckBuilds.
- **Docs.** Extensive documentation in `docs/`, e.g. `WEB_INTERFACE_GUIDE.md`, `GETTING_STARTED.md`, `WIFI_NETWORK_SETUP.md`, `LOW_MEMORY_BOARDS.md`, `PLUGIN_STORE_GUIDE.md`.
- **Absences.** There are no testimonials, user counts, or benchmark figures. Do not fabricate them.
## Product Principles
1. **Novice path first, power one click away.** Default views serve the first-time builder, while advanced tools stay discoverable for tinkerers.
2. **Never strand the user at a terminal.** Every setup, recovery, and troubleshooting task has a browser path, including from the AP-mode captive page.
3. **Respect the Pi.** Every feature is paid for in memory and CPU on a Pi Zero 2 W that is also driving the display.
4. **The ecosystem is the product.** Plugins, including third-party ones, must feel first-class and keep working across core UI changes.
5. **Honest and welcoming.** Plain language, truthful status, and no overstated claims, in keeping with an open, community-built project.
+113 -61
View File
@@ -148,7 +148,7 @@ The system supports live, recent, and upcoming game information for multiple spo
```bash
sudo RPI_RGB_FORCE_REBUILD=1 ./first_time_install.sh
```
- Pi 5 config: leave `rp1_rio` at `0` (PIO mode, default) and set `gpio_slowdown` to `1` or `2`.
- Pi 5 config: leave `rp1_rio` at `0` (PIO mode, default) and start `gpio_slowdown` at `1`, raising it a step at a time if the image flickers or shows garbage (see `gpio_slowdown` under Display Settings).
- **1GB models (Pi 3B / 3B+) and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`.
@@ -463,13 +463,12 @@ For plugin development, check out the [Hello World Plugin](https://github.com/Ch
### Visual Skins for Scoreboards
Want a different look for a sports scoreboard without forking the plugin?
**Skins** restyle the live/recent/upcoming screens while the plugin keeps
handling data, scheduling, caching, and vegas mode. Install one with
`git clone <skin repo> skins/<skin-id>`, select it in the plugin's config,
and you're done — see [docs/SKIN_SYSTEM.md](docs/SKIN_SYSTEM.md) (how it
works) and [docs/CREATING_SKINS.md](docs/CREATING_SKINS.md) (build your own,
including a ready-made Claude Code prompt).
**Not supported yet.** Skins are meant to restyle a sports scoreboard's
live/recent/upcoming screens without forking the plugin, but the current
scoreboard plugins don't render them: a selected skin has no effect. The web
UI doesn't offer skin install or selection for that reason. The skin system
and its docs stay in place for when scoreboards adopt it; see
[docs/SKIN_SYSTEM.md](docs/SKIN_SYSTEM.md) for why.
2. **Built-in Managers Deprecated**: The built-in managers (hockey, football, stocks, etc.) are now deprecated and have been moved to the plugin system. **You must install replacement plugins from the Plugin Store** in the web interface instead. The plugin system provides the same functionality with better maintainability and extensibility.
</details>
@@ -486,6 +485,10 @@ If you are copying my exact setup, you can likely leave the defaults alone. Howe
The display settings are located in `config/config.json` under the `"display"` key and are organized into three main sections: `hardware`, `runtime`, and `display_durations`.
The defaults below are the values in `config/config.template.json`. They are what applies when you haven't set a key: on every load, LEDMatrix adds any key your `config.json` lacks from the template, so `DisplayManager`'s own fallbacks are never reached on a normal install.
The web UI and the config API refuse values the rgbmatrix library can't start with. If one is written into `config.json` by hand anyway, the display logs which setting it is (`Failed to initialize RGB Matrix` in `sudo journalctl -u ledmatrix`), runs in fallback mode, and the Display tab shows the message.
### Hardware Configuration (`display.hardware`)
These settings control the physical hardware configuration and how the matrix is driven.
@@ -495,15 +498,18 @@ These settings control the physical hardware configuration and how the matrix is
- **`rows`** (integer, default: 32)
- Number of LED rows (vertical pixels) in each panel
- Common values: 16, 32, 48, 64
- An even number from 8 to 64, the most the rgbmatrix library drives per panel
- Must match your physical panel configuration
- **`cols`** (integer, default: 64)
- Number of LED columns (horizontal pixels) in each panel
- Common values: 32, 64, 96, 128
- At least 16, with no upper limit
- Must match your physical panel configuration
- **`chain_length`** (integer, default: 2)
- Number of LED panels chained together horizontally
- 1 to 255 (the library's Python binding stores it in one byte); longer chains lower the refresh rate
- If you have 2 panels side-by-side, set to 2
- If you have 4 panels in a row, set to 4
- Total display width = `cols × chain_length`
@@ -512,68 +518,70 @@ These settings control the physical hardware configuration and how the matrix is
- Number of parallel chains (panels stacked vertically)
- Use 1 for a single row of panels
- Use 2 if you have panels stacked in two rows
- 1–3, and no more than your `hardware_mapping` has outputs: `regular` and `classic` have 3 (e.g. the Adafruit Triple LED Matrix Bonnet); `adafruit-hat`, `adafruit-hat-pwm`, `regular-pi1` and `classic-pi1` have 1. The library stops the display service outright on a mismatch, so it is refused
- Total display height = `rows × parallel`
#### Brightness and Visual Settings
- **`brightness`** (integer, 0-100, default: 90)
- **`brightness`** (integer, 1-100, default: 90)
- Display brightness level
- Lower values (0-50) are dimmer, higher values (50-100) are brighter
- Lower values (1-50) are dimmer, higher values (50-100) are brighter
- Recommended: 70-90 for indoor use, 90-100 for bright environments
- Very high brightness may cause distortion or require more power
#### Hardware Mapping
- **`hardware_mapping`** (string, default: "adafruit-hat-pwm")
- **`hardware_mapping`** (string, default: "adafruit-hat")
- Specifies which GPIO pin mapping to use for your hardware
- **`"adafruit-hat-pwm"`**: Use this for Adafruit RGB Matrix Bonnet/HAT WITH the jumper mod (PWM enabled). This is the recommended setting for Adafruit hardware with the PWM jumper soldered.
- **`"adafruit-hat"`**: Use this for Adafruit RGB Matrix Bonnet/HAT WITHOUT the jumper mod (no PWM). Remove `-pwm` from the value if you did not solder the jumper.
- **`"regular"`**: Standard GPIO pin mapping for direct GPIO connections (Generic)
- **`"regular"`**: Standard GPIO pin mapping for direct GPIO connections (Generic). Also the right choice for the Adafruit Triple LED Matrix Bonnet
- **`"regular-pi1"`**: Standard GPIO pin mapping for Raspberry Pi 1 (older hardware or non-standard hat mapping)
- **`"classic"`** / **`"classic-pi1"`**: the library's original pin-outs, for old adapter boards wired to them. Not used by current HATs
- Any other name is refused. `compute-module` is only compiled in when the library is built with `ENABLE_WIDE_GPIO_COMPUTE_MODULE`, which the installer doesn't do. On a Raspberry Pi 5, `classic-pi1` isn't supported
- Choose the option that matches your specific hardware setup, if aren't sure try them all.
- Hardware pulsing (see `disable_hardware_pulsing`) needs the panel's OE line on GPIO 18, which `adafruit-hat-pwm` and `regular` provide and `adafruit-hat` does not
#### PWM (Pulse Width Modulation) Settings
These settings affect color fidelity and smoothness of color transitions:
- **`pwm_bits`** (integer, default: 9)
- Number of bits used for PWM (affects color depth)
- Higher values (9-11) = more color levels, smoother gradients
- Lower values (7-8) = fewer color levels, but may improve stability on some hardware
- Range: 1-11, recommended: 9-10
- **`pwm_bits`** (integer, 1-11, default: 9)
- Color depth per channel: how many brightness levels each LED gets
- Higher values (9-11) = more color levels, smoother gradients, lower refresh rate
- Lower values (7-8) = the subtlest shades are dropped for a higher refresh rate; `1` gives 8 colors
- Recommended: 9-10
- **`pwm_dither_bits`** (integer, default: 1)
- Additional dithering bits for smoother color transitions
- Helps reduce color banding in gradients
- Higher values (1-2) = smoother gradients but may impact performance
- Range: 0-2, recommended: 1
- **`pwm_dither_bits`** (integer, 0-2, default: 1)
- Time-dithers the lowest color bits: their brightness comes from showing them on only some frames
- Raises the refresh rate; the cost is that dark shades can shimmer slightly
- `0` = steadiest dim colors, `2` = fastest
- The rgbmatrix library accepts only 0-2; a higher value stops the display starting
- **`pwm_lsb_nanoseconds`** (integer, default: 130)
- Least significant bit timing in nanoseconds
- Controls the base timing for PWM signals
- Lower values = faster PWM, higher values = slower PWM
- **`pwm_lsb_nanoseconds`** (integer, 50-3000, default: 130)
- On-time of the least significant color bit; each higher bit doubles it
- Lower values = higher refresh rate, but can cost color accuracy or add ghosting on some panels
- Higher values = less ghosting (faint trails behind bright text on black), lower refresh rate
- Typical range: 100-300 nanoseconds
- May need adjustment if you see flickering or color issues
#### Advanced Hardware Settings
- **`scan_mode`** (integer, default: 0)
- Panel scan mode (how rows are addressed)
- Common values: 0 (progressive), 1 (interlaced)
- Most panels use 0, but some require 1
- Check your panel datasheet if colors appear incorrect
- **`scan_mode`** (integer, 0-1, default: 0)
- Order the rows are refreshed in: `0` = progressive, `1` = interlaced
- Interlaced can look a little smoother when the refresh rate is very low, but usually shows a comb effect on anything moving
- Leave at `0` unless you are tuning a slow setup
- **`limit_refresh_rate_hz`** (integer, default: 100)
- Maximum refresh rate in Hz (frames per second)
- Caps the refresh rate for better stability
- Lower values (60-80) = more stable, less CPU usage
- Higher values (100-120) = smoother animations, more CPU usage
- Recommended: 80-100 for most setups
- Caps the panel refresh rate in Hz; `0` = no cap
- A steady cap reduces flicker caused by other activity on the Pi, and in camera recordings
- Scroll speeds are worked out against this value (against 100 Hz when it is `0`), so a cap the panel can actually hold keeps scrolling even
- Recommended: 80-120. `sudo python3 scripts/scroll_speeds.py --measure` reports the rate your panel really achieves
- **`disable_hardware_pulsing`** (boolean, default: false)
- Disables hardware pulsing (usually leave as false)
- Set to `true` only if you experience timing issues
- Most users should leave this as `false`
- `false` = the Pi's hardware PWM times each brightness pulse; `true` = software timing
- Leave `false` where possible. Software timing is less exact, so a row, or the whole panel, can briefly flash brighter
- Hardware pulsing needs the panel's OE line on GPIO 18 (`adafruit-hat-pwm`, `regular`, the Adafruit Triple LED Matrix Bonnet). With `adafruit-hat` the library uses software timing anyway
- It also needs the Pi's onboard sound driver (`snd_bcm2835`) disabled, which `first_time_install.sh` does. Set `true` only if you need the Pi's own audio
- **`inverse_colors`** (boolean, default: false)
- Inverts all colors (red becomes cyan, etc.)
@@ -581,9 +589,9 @@ These settings affect color fidelity and smoothness of color transitions:
- Set to `true` only if colors appear inverted
- **`show_refresh_rate`** (boolean, default: false)
- Displays the current refresh rate on the matrix (for debugging)
- Set to `true` to see FPS on the display
- Useful for troubleshooting performance issues
- Prints the live refresh rate to the console; nothing is drawn on the panel
- Readable when you stop the service and run `sudo python3 run.py` in a terminal; under the service the output is buffered
- `sudo python3 scripts/scroll_speeds.py --measure` is an easier way to see the real refresh rate
#### Advanced Panel Configuration (Advanced Users Only)
@@ -593,6 +601,7 @@ These settings are typically only needed for non-standard panels or custom confi
- Color channel order for your LED panel
- Common values: "RGB", "RBG", "GRB", "GBR", "BRG", "BGR"
- Most panels use "RGB", but some use "GRB" or other orders
- If red shows as blue, try "BGR" (the Waveshare 96x48 V2 needs it)
- Check your panel datasheet if colors appear wrong
- **`pixel_mapper_config`** (string, default: "")
@@ -607,35 +616,68 @@ These settings are typically only needed for non-standard panels or custom confi
- Set to `"180"` (or use the "Upside Down" option in the web UI's Display
settings) if the panel is mounted upside down — useful for optimizing
where the Raspberry Pi and wiring sit relative to the mounting location
- `"90"` and `"270"` are for a panel mounted on its side; they swap the
display's width and height
- Applied independently of `pixel_mapper_config` (appended as a trailing
`Rotate:180` mapper), so custom mapper configs keep working alongside it
`Rotate:<degrees>` mapper), so custom mapper configs keep working alongside it
- **`row_address_type`** (integer, default: 0)
- How rows are addressed on the panel
- Most panels use 0 (direct addressing)
- Some panels require 1 (AB addressing) or 2 (ABC addressing)
- 1 = AB-addressed, 2 = direct row select, 3 = ABC-addressed,
4 = ABC shift + DE direct (SM5266), 5 = SM5368 / B707 row shift register
- ABC panels (no E line, e.g. many 128x64 FM6124 panels) use 3
- Panels with SM5368 row drivers use 5 with `led_rgb_sequence` `"BGR"` —
e.g. the Waveshare 96x48 V2 (back silkscreen `24S-A1`; the V1, `24S-A2.1`,
uses the defaults). This is what Waveshare's `96X48_1_24_SM5368` panel
type sets in their library fork.
- SM5368 row drivers are timing-sensitive: if rows jump up and down or the
bottom row shows a copy of other rows, raise `gpio_slowdown`. On a Pi 4
with an Adafruit Triple LED Matrix Bonnet, 4 left rows jumping; 6–8 gave a
stable image.
- On a Raspberry Pi 5 the rgbmatrix library currently supports only 0 and 2
(and `parallel` 1-3). Anything else would crash the display service, so on
a Pi 5 the web UI offers only 0 and 2, the config API refuses the others,
and if one is set in `config.json` anyway the display logs why and runs in
fallback mode
- Check your panel datasheet if display appears corrupted
- **`multiplexing`** (integer, default: 0)
- Panel multiplexing type
- 0 = no multiplexing (standard panels)
- Higher values for panels with different multiplexing schemes
- Check your panel datasheet for the correct value
- **`multiplexing`** (integer, 0-22, default: 0)
- How pixels are wired on outdoor/specialty panels (P10, P8, P4 and P3 outdoor modules and similar) whose LEDs aren't laid out in straight rows
- `0` = direct (standard indoor panels)
- `1` Stripe, `2` Checkered, `3` Spiral, `4` ZStripe, `5` ZnMirrorZStripe,
`6` Coreman, `7` Kaler2Scan, `8` ZStripeUneven, `9` P10-128x4-Z,
`10` QiangLiQ8, `11` InversedZStripe, `12`–`14` P10Outdoor1R1G1B v1–v3,
`15` P10CoremanMapper, `16` P8Outdoor1R1G1B, `17` FlippedStripe,
`18` P10-32x16-HalfScan, `19` P10-32x16-QuarterScan, `20` P3Outdoor-64x64,
`21` DoubleZMultiplex, `22` P4Outdoor-80x40
- If the image is scrambled in a repeating pattern, try the value named after your panel first
- **`panel_type`** (string, default: `""`)
- Sends a start-up initialization sequence to driver chips that need one
- `""` = Standard (no initialization) — right for most panels, including FM6124 / FM6124D / FM6124DJ
- `"FM6126A"` or `"FM6127"` for panels with those chips; try `"FM6126A"` if the panel stays dark or lights only the first pixel on Standard
### Runtime Configuration (`display.runtime`)
These settings control runtime behavior and GPIO timing:
- **`gpio_slowdown`** (integer, default: 3)
- GPIO timing slowdown factor
- **Critical setting**: Must match your Raspberry Pi model for stability
- **Raspberry Pi 3**: Use 3
- **Raspberry Pi 4**: Use 4
- **Raspberry Pi 5**: Use 1–2 in PIO mode (`rp1_rio: 0`, the default); start with `1` and increase if you see flickering
- **Raspberry Pi Zero/1**: Use 1-2
- Incorrect values can cause display corruption, flickering, or system instability
- GPIO timing slowdown factor (0-10): slows GPIO writes so the panel electronics keep up. Higher is more reliable but lowers the refresh rate
- **Critical setting**: depends on your Raspberry Pi model and your panel
- **Raspberry Pi Zero/1**: 0-1
- **Raspberry Pi 2/3**: 1-3
- **Raspberry Pi 4**: 2-4 (the config template ships 3)
- **Raspberry Pi 5**: 1–3 in PIO mode (`rp1_rio: 0`, the default). Start at `1` (the library treats `0` as `1` there) and raise it a step at a time if the image flickers or shows garbage — chained panels are the likeliest to need it
- Panels on `row_address_type` 5 (SM5368 row drivers) can need 6-8 on a Pi 4
- Too low: garbage, flicker or rows jumping. Too high: a lower refresh rate
- If you experience issues, try adjusting this value up or down by 1
- **`rp1_rio`** (integer, 0 or 1, default: 0) — Raspberry Pi 5 only
- Which driver the Pi 5's RP1 chip uses: `0` = PIO (default, less CPU), `1` = RIO (registered I/O, can reach a higher refresh rate)
- In RIO mode the effect of `gpio_slowdown` is inverted: higher values may be faster
- Ignored on a Pi 0-4, and applied only if the installed rgbmatrix library supports it
### Display Durations (`display.display_durations`)
Controls how long each installed plugin stays visible in seconds before switching to the next one, keyed by plugin id.
@@ -668,7 +710,7 @@ Controls how long each installed plugin stays visible in seconds before switchin
- Some plugins can automatically adjust their display time based on content
- This setting limits how long they can extend (prevents one display from dominating)
- Example: If set to 60, a plugin can extend up to 60 seconds even if it requests longer
- Leave unset to use the default cap (typically 90 seconds)
- Leave unset to use the default cap (180 seconds; the web UI accepts 30-1800)
### Example Configuration
@@ -715,6 +757,14 @@ Controls how long each installed plugin stays visible in seconds before switchin
- Verify `hardware_mapping` matches your HAT/connection type
- Try adjusting `gpio_slowdown`
- Ensure your display doesn't need the E-Addressable line
- If it went blank right after a settings change, the Display tab shows a "simulation mode" banner, and `sudo journalctl -u ledmatrix` shows `Failed to initialize RGB Matrix` followed by the reason. When LEDMatrix refused the settings (for example more than 64 `rows`, `parallel` 2 on an `adafruit-hat` mapping, a misspelled `hardware_mapping`, or on a Raspberry Pi 5 a `row_address_type` other than 0 or 2), the message names each one: change them, save, and restart the display service. Otherwise the library itself failed, and its own message just before names the problem
- A repeating scramble points at `row_address_type` or `multiplexing`; a panel that stays dark, at `panel_type`
**Rows jump up and down, or the bottom row repeats other rows:**
- Raise `gpio_slowdown` a step at a time (SM5368 panels on `row_address_type` 5 can need 6-8 on a Pi 4)
**A row or the whole panel briefly flashes brighter:**
- Set `disable_hardware_pulsing` to `false` (needs the OE line on GPIO 18; see `hardware_mapping`)
**Colors are wrong or inverted:**
- Check `led_rgb_sequence` (try "GRB" if "RGB" doesn't work)
@@ -789,9 +839,11 @@ sudo ./scripts/install/install_service.sh
The script will:
- Detect your user account and home directory
- Install the service file with the correct paths
- Enable the service to start on boot
- Start the service immediately
- Install `ledmatrix.service` (display, runs as root), `ledmatrix-web.service`
(web interface, runs as your user) and the `ledmatrix-update-verify` units,
with the correct paths
- Enable them to start on boot
- Start them immediately
### Managing the Service
+3
View File
@@ -1,5 +1,8 @@
{
"web_display_autostart": true,
"auto_update": {
"enabled": false
},
"schedule": {
"enabled": false,
"mode": "per-day",
+6
View File
@@ -0,0 +1,6 @@
{
"dev_plugins_dir": "~/.ledmatrix-dev-plugins",
"github_user": "ChuckBuilds",
"plugins_repo": "ledmatrix-plugins",
"plugins_branch": "main"
}
+2 -2
View File
@@ -223,8 +223,8 @@ The harness already renders every plugin at a spread of sizes (now
including 96x48):
```bash
python scripts/check_plugin.py <plugin-dir> --sizes 64x32,128x32,96x48,128x64,256x64
python scripts/render_plugin.py <plugin-dir> --width 96 --height 48
python scripts/check_plugin.py --plugin <plugin-id> --sizes 64x32,128x32,96x48,128x64,256x64
python scripts/render_plugin.py --plugin <plugin-id> --width 96 --height 48
```
`BoundsCheckingDisplayManager` flags right/bottom overflow and now records
+36 -14
View File
@@ -377,9 +377,16 @@ Vegas mode consists of four core components working together to provide smooth 1
5. Compose into continuous stream with separators
**Key Methods:**
- `get_stream_content()` - Returns current stream content as PIL Image
- `advance_stream(pixels)` - Advances stream by N pixels
- `refresh_stream()` - Regenerates stream from current plugins
- `get_next_segment()` - Returns the next buffered `ContentSegment` (or `None`)
- `take_next_group(count=None, offscreen_only=False)` - Hands over the next
slice of the rotation as `(plugin_id, images)` groups
- `get_grouped_content_for_composition()` - Buffered images grouped by plugin
- `mark_plugin_updated(plugin_id)` / `process_updates()` - Refresh one
plugin's segment in place when its data changes
- `refresh()` - Re-read the plugin list and config
- `advance_cycle()` - Clear the active buffer when a scroll cycle completes
(`src/vegas_mode/stream_manager.py`)
#### 3. PluginAdapter
@@ -433,10 +440,14 @@ Vegas mode consists of four core components working together to provide smooth 1
- **Frame Rate Control:** Precise timing to maintain 125 FPS
- **Pre-rendered Content:** Plugins pre-render during update()
**Scroll Speed Calculation:**
**Scroll Speed Calculation:** motion is by elapsed time; `target_fps` paces
the render loop, not the speed.
```python
pixels_per_frame = (scroll_speed / target_fps)
scroll_position += pixels_per_frame * elapsed_time
# frame_based_scrolling: false
scroll_position += scroll_speed * elapsed_time # scroll_speed in px/s
# frame_based_scrolling: true (the default) -- not stepping, just a clamp
applied = clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay
scroll_position += applied * elapsed_time
```
#### Component Interactions
@@ -552,7 +563,8 @@ time when something is active.
### REST API Reference
The API is mounted at `/api/v3` (`web_interface/app.py:199`).
The API is mounted at `/api/v3` (the `api_v3` blueprint, registered in
`web_interface/app.py`). Full details: [REST_API_REFERENCE.md](REST_API_REFERENCE.md#display-control).
#### Start On-Demand Display
@@ -608,20 +620,30 @@ curl http://localhost:5000/api/v3/display/on-demand/status
# Response:
{
"active": true,
"plugin_id": "weather",
"mode": "weather",
"remaining": 25.5,
"pinned": false,
"status": "active"
"status": "success",
"data": {
"state": {
"active": true,
"plugin_id": "weather",
"mode": "weather",
"duration": 30,
"pinned": false,
"status": "running",
"last_updated": 1234567890.1
},
"service": {"active": true, "returncode": 0, "stdout": "active", "stderr": ""}
}
}
```
When nothing is running on demand, `data.state` is
`{"active": false, "status": "idle", "last_updated": null}`.
> There is no public Python on-demand API. The display controller's
> on-demand machinery is internal — drive it through the REST endpoints
> above (or the web UI buttons). The API handlers
> (`start_on_demand_display()` / `stop_on_demand_display()` in
> `web_interface/blueprints/api_v3.py`) write a request into the cache
> `web_interface/blueprints/api_v3/display.py`) write a request into the cache
> manager under the `display_on_demand_request` key, which
> `DisplayController._poll_on_demand_requests()`
> (`src/display_controller.py`) picks up. A separate
+14 -5
View File
@@ -250,14 +250,21 @@ WARNING - Plugin ID 'Football-Scoreboard' may conflict with 'football-scoreboard
## Checking Configuration via API
The API blueprint mounts at `/api/v3` (`web_interface/app.py:144`).
The API blueprint (`web_interface/blueprints/api_v3/`) is registered at
`/api/v3` in `web_interface/app.py`.
```bash
# Get full main config (includes all plugin sections)
# Get full main config (includes all plugin sections; credential-named
# fields are blanked in the response)
curl http://localhost:5000/api/v3/config/main
# Save updated main config
# Change some settings: only the keys you send are changed
curl -X POST http://localhost:5000/api/v3/config/main \
-H "Content-Type: application/json" \
-d '{"timezone": "America/Chicago", "brightness": 80}'
# Replace config.json wholesale (advanced)
curl -X POST http://localhost:5000/api/v3/config/raw/main \
-H "Content-Type: application/json" \
-d @new-config.json
@@ -269,8 +276,10 @@ curl "http://localhost:5000/api/v3/plugins/config?plugin_id=football-scoreboard"
```
> There is no dedicated `/config/plugin/<id>` or `/config/validate`
> endpoint — config validation runs server-side automatically when you
> POST to `/config/main` or `/plugins/config`. See
> endpoint. `POST /plugins/config` validates against the plugin's schema
> and rejects an invalid config with `400`; `POST /config/main` checks the
> individual fields it knows (display hardware values, durations, Vegas
> and sync settings). See
> [REST_API_REFERENCE.md](REST_API_REFERENCE.md) for the full list.
## Backup and Recovery
+36 -33
View File
@@ -17,7 +17,7 @@ tooling against it.
|---|---|---|---|
| `web_display_autostart` | bool, `true` | Whether the web interface service starts with the system | `scripts/utils/start_web_conditionally.py` |
| `timezone` | string, `"America/New_York"` | IANA timezone for schedules and displays | `ConfigManager.get_timezone()` |
| `target_fps` | int, `100` | Frame-rate ceiling for plugin rendering | `src/plugin_system/base_plugin.py`, `src/common/sports_scroll.py` |
| `target_fps` | int, `100` | Legacy "Scroll Frame Rate". Core scrolling no longer reads it: scroll frames are presented at `display.hardware.limit_refresh_rate_hz` divided by each scroll's frame hold, and speed comes from each plugin's scroll settings. Still exposed to plugins via `BasePlugin.global_config` | `src/plugin_system/base_plugin.py` |
| `location` | object | `city` / `state` / `country`. Supplies the **default** for a plugin's own `location_city` / `location_state` / `location_country` setting, so weather, radar and friends follow this device without being configured twice. A value saved on the plugin itself still overrides it. | `SchemaManager.apply_device_location()`, then plugins via merged config |
## `schedule` — display on/off hours
@@ -47,39 +47,44 @@ saved via `POST /api/v3/config/dim-schedule`). The display returns to
## `display.hardware` — matrix panel hardware
All keys map to the corresponding `rpi-rgb-led-matrix` options and are read
in `DisplayManager` (`src/display_manager.py`, ~lines 270–295).
in `DisplayManager._setup_matrix` (`src/display_manager.py`). Defaults are the
`config/config.template.json` values: `ConfigManager` adds any key missing from
`config.json` from the template on load, so `DisplayManager`'s own fallbacks
don't apply on a normal install.
The ranges are what the pinned rgbmatrix library and its Python binding accept
(`src/matrix_support.py`). The config API refuses anything else; a value
hand-edited into `config.json` makes the display log the setting and run in
fallback mode instead of starting the matrix.
| Key | Type / default |
|---|---|
| `rows` / `cols` | int, `32` / `64` |
| `chain_length` | int, `2` |
| `parallel` | int, `1` |
| `brightness` | int, `90` |
| `hardware_mapping` | string, `"adafruit-hat"` (code default `"adafruit-hat-pwm"`) |
| `scan_mode` | int, `0` |
| `pwm_bits` | int, `9` (code default 10) |
| `pwm_dither_bits` | int, `1` |
| `pwm_lsb_nanoseconds` | int, `130` (code default 150) |
| `disable_hardware_pulsing` | bool, `false` |
| `rows` / `cols` | int, `32` / `64` — rows: even, 8–64; cols: at least 16 |
| `chain_length` | int, `2` — 1–255 (the Python binding stores it in one byte) |
| `parallel` | int, `1` — 1–3, and no more than `hardware_mapping` has outputs (`regular`, `classic`: 3; the others: 1) |
| `brightness` | int, `90` — 1–100 |
| `hardware_mapping` | string, `"adafruit-hat"` — `"adafruit-hat-pwm"`, `"adafruit-hat"`, `"regular"`, `"regular-pi1"`, `"classic"` or `"classic-pi1"` (case-insensitive; `compute-module` isn't in the installed build). A Pi 5 doesn't support `"classic-pi1"` |
| `scan_mode` | int, `0` — `0` progressive, `1` interlaced |
| `pwm_bits` | int, `9` — 1–11 |
| `pwm_dither_bits` | int, `1` — 0–2 |
| `pwm_lsb_nanoseconds` | int, `130` — 50–3000 |
| `disable_hardware_pulsing` | bool, `false` — `true` times brightness pulses in software (less exact); hardware pulsing needs the OE line on GPIO 18 and the Pi's onboard sound driver off |
| `inverse_colors` | bool, `false` |
| `show_refresh_rate` | bool, `false` |
| `led_rgb_sequence` | string, `"RGB"` |
| `limit_refresh_rate_hz` | int, `100` (code default 90) |
| `pixel_mapper_config` | string, `""` — e.g. `"U-mapper"` / `"Rotate:90"` |
| `orientation` | string, `"normal"` — `"180"` rotates the rendered image 180° for panels physically mounted upside down (e.g. to move the Pi/wiring to a more convenient side); composed onto `pixel_mapper_config` as a trailing `Rotate:180` mapper, so it stays independent of any custom `pixel_mapper_config` value |
| `row_address_type` | int, `0` — non-standard panel row addressing |
| `multiplexing` | int, `0` — panel multiplexing scheme |
| `panel_type` | string, `""` — set to `"FM6126A"` or `"FM6127"` for panels needing init |
Where "code default" differs from the template value, the code default only
applies if the key is missing entirely from your config.
| `show_refresh_rate` | bool, `false` — prints the refresh rate to stdout; draws nothing on the panel |
| `led_rgb_sequence` | string, `"RGB"` — `"RGB"`, `"RBG"`, `"GRB"`, `"GBR"`, `"BRG"` or `"BGR"` |
| `limit_refresh_rate_hz` | int, `100` — `0` = no cap; scroll timing assumes 100 Hz when `0` |
| `pixel_mapper_config` | string, `""` — e.g. `"U-mapper"` / `"Rotate:90"`; mappers that rotate or fold the chain change the display size plugins and the web preview see |
| `orientation` | string, `"normal"` — `"180"` rotates the rendered image 180° for panels physically mounted upside down (e.g. to move the Pi/wiring to a more convenient side); `"90"` / `"270"` for a panel on its side, swapping width and height; composed onto `pixel_mapper_config` as a trailing `Rotate:<degrees>` mapper, so it stays independent of any custom `pixel_mapper_config` value |
| `row_address_type` | int, `0` — non-standard panel row addressing: `1` AB, `2` direct row select, `3` ABC, `4` ABC shift + DE direct, `5` SM5368 / B707 row shift register (e.g. Waveshare 96x48 V2, with `led_rgb_sequence` `"BGR"`). On a Pi 5 the library supports only `0` and `2`, and LEDMatrix enforces that (`src/pi5_matrix_support.py`) |
| `multiplexing` | int, `0` — 0–22, pixel wiring scheme for outdoor/specialty panels (names listed in the README) |
| `panel_type` | string, `""` — set to `"FM6126A"` or `"FM6127"` for panels needing init; FM6124 / FM6124D / FM6124DJ panels need none, so leave it `""` |
## `display.runtime`
| Key | Type / default | Meaning |
|---|---|---|
| `gpio_slowdown` | int, `3` | GPIO timing slowdown for faster Pis |
| `rp1_rio` | int, `0` | RP1 RIO mode on Pi 5 (applied only if the installed matrix library supports it) |
| `gpio_slowdown` | int, `3` | GPIO timing slowdown for faster Pis (0–10). On a Pi 5 in PIO mode start at `1` (`0` acts as `1`) and raise it if the image flickers or shows garbage. Panels on `row_address_type` `5` (SM5368 row drivers) can need 6–8 on a Pi 4 — lower values make rows jump |
| `rp1_rio` | int, `0` | Pi 5 only: `0` = PIO (less CPU), `1` = RIO (higher refresh; `gpio_slowdown` effect inverted). Applied only if the installed matrix library supports it |
## `display.double_sided`
@@ -134,8 +139,8 @@ Read by `src/vegas_mode/config.py` (`VegasScrollConfig.from_config`). See
| `dynamic_duration_enabled` | bool, `true` |
| `min_cycle_duration` | int, `60` |
| `max_cycle_duration` | int, `240` |
| `frame_based_scrolling` | bool, `true` — frame-count-based scroll stepping |
| `scroll_delay` | float, `0.02` — seconds between scroll updates (~50 FPS) |
| `frame_based_scrolling` | bool, `true` — does not step or set a frame rate; motion is by elapsed time either way. When `true`, `scroll_speed` passes through a clamp of 0.1–5 px per `scroll_delay` (see next row) |
| `scroll_delay` | float, `0.02` — not a frame period. Only used with `frame_based_scrolling`: the applied speed is `clamp(scroll_speed × scroll_delay, 0.1, 5) / scroll_delay` px/s, so at `0.02` speeds under 5 px/s run at 5, and at `0.001` nothing runs slower than 100 px/s |
| `live_in_ticker` | bool, `false` — keep scrolling during live games instead of handing the display to a full-screen scoreboard |
| `live_weight` | int, `3` (1–10) — slots per cycle for a plugin with live content |
| `favorite_live_weight` | int, `5` (1–10) — slots per cycle when a plugin reports a favorite team is live |
@@ -152,14 +157,12 @@ Read by `src/common/sync_manager.py` and `src/display_controller.py`.
## `plugin_system`
Read by the plugin loader/manager (`src/plugin_system/`).
| Key | Type / default | Meaning |
|---|---|---|
| `plugins_directory` | string, `"plugin-repos"` | Where the Plugin Store installs plugins |
| `auto_discover` | bool, `true` | Scan the plugins directory at startup |
| `auto_load_enabled` | bool, `true` | Load discovered plugins automatically |
| `development_mode` | bool, `false` | Development conveniences in the web UI (editable under General settings) |
| `plugins_directory` | string, `"plugin-repos"` | Where the Plugin Store installs plugins and the only directory the plugin loader scans. Read by `PluginManager` and `PluginStoreManager` (`src/plugin_system/`); editable under General settings |
| `auto_discover` | bool, `true` | **Unused.** Legacy key, read by nothing. Plugins are always discovered, and every plugin with `enabled: true` is loaded. Not shown in the web UI; may be left in or removed from config.json |
| `auto_load_enabled` | bool, `true` | **Unused.** Legacy key, read by nothing (see `auto_discover`). To keep a plugin installed but dormant, set its own `enabled` to `false` |
| `development_mode` | bool, `false` | **Unused.** Legacy key, read by nothing |
## Plugin config blocks
+17 -5
View File
@@ -1,5 +1,14 @@
# Creating Skins
> **Not supported yet: skins don't render with the current scoreboard
> plugins.** The only render hook is `SportsCore._render_game()` in
> `src/base_classes/sports/core.py`, and no current scoreboard (monorepo or
> third-party) builds on `src.base_classes`, so a skin you build here passes
> `validate_skin.py` but never appears on the matrix. The web UI and Plugin
> Store don't offer skins for that reason. Details:
> [SKIN_SYSTEM.md](SKIN_SYSTEM.md#status-not-supported-yet). The guide below
> stays accurate for the skin API itself.
A skin restyles a sports scoreboard (live / recent / upcoming) without
forking the plugin: the plugin keeps fetching data, scheduling, caching, and
doing vegas mode; your skin only draws. Architecture background:
@@ -19,7 +28,9 @@ panel sizes with **no hardware, no network, no running service**, saves PNGs
(plus 4x previews) to `skin_renders/`, and fails loudly on errors. Iterate:
edit → validate → look at the PNGs.
To see it on your matrix, add to your plugin's section in `config/config.json`:
To select it, add to your plugin's section in `config/config.json` (this is
stored and validated, but has no visible effect until a scoreboard uses the
skin hook — see the note at the top):
```json
"baseball-scoreboard": {
@@ -28,8 +39,8 @@ To see it on your matrix, add to your plugin's section in `config/config.json`:
}
```
or pick it from the **Visual Skin** dropdown in the web UI (it appears once a
matching skin is installed). `"skin"` also accepts a per-mode mapping:
The web UI's **Visual Skin** dropdown is hidden while skins are unsupported.
`"skin"` also accepts a per-mode mapping:
`{"live": "my-skin", "recent": "built-in"}`.
## The manifest (`skin.json`)
@@ -234,8 +245,9 @@ Tips that keep Claude (and you) out of trouble:
dev machine
Distribute by publishing the directory as a git repo (users
`git clone <repo> skins/<id>`), or submit it to the plugin registry as an
entry with `"type": "skin"` (see [SKIN_SYSTEM.md](SKIN_SYSTEM.md) §Distribution).
`git clone <repo> skins/<id>`). Registry entries with `"type": "skin"` are
hidden and refused by the Plugin Store while skins are unsupported (see
[SKIN_SYSTEM.md](SKIN_SYSTEM.md) §Distribution).
**Trust note:** a skin is Python running inside the display service — the
same trust level as a plugin. Review code before installing skins from
+38 -33
View File
@@ -12,10 +12,28 @@
The enhanced FontManager provides comprehensive font management for the LEDMatrix application with support for:
- Manager font registration and detection
- Plugin font management
- Manual font overrides via web interface
- Programmatic per-element font overrides
- Performance monitoring and caching
- Dynamic font discovery
## Getting the FontManager
There is one shared FontManager per display process. The display controller
creates it and hands it to the `PluginManager`, so a plugin reaches it
through its `plugin_manager`:
```python
class MyPlugin(BasePlugin):
def __init__(self, plugin_id, config, display_manager, cache_manager, plugin_manager):
super().__init__(plugin_id, config, display_manager, cache_manager, plugin_manager)
self.font_manager = self._get_font_manager()
```
`BasePlugin._get_font_manager()` returns `plugin_manager.font_manager`, or a
standalone FontManager when none is available (test harnesses, mocks).
`DisplayManager` has **no** `font_manager` attribute —
`display_manager.font_manager` raises `AttributeError`.
## Architecture
### Manager-Centric Design
@@ -40,8 +58,9 @@ Manager requests font → Check manual overrides → Apply manager choice → Ca
from src.font_manager import FontManager
class MyManager:
def __init__(self, config, display_manager, cache_manager):
self.font_manager = display_manager.font_manager # Access shared FontManager
def __init__(self, config, display_manager, cache_manager, plugin_manager):
self.display_manager = display_manager
self.font_manager = plugin_manager.font_manager # Shared FontManager
self.manager_id = "my_manager"
def display(self):
@@ -80,8 +99,9 @@ class MyManager:
```python
class AdvancedManager:
def __init__(self, config, display_manager, cache_manager):
self.font_manager = display_manager.font_manager
def __init__(self, config, display_manager, cache_manager, plugin_manager):
self.display_manager = display_manager
self.font_manager = plugin_manager.font_manager
self.manager_id = "advanced_manager"
# Define your font specifications
@@ -152,19 +172,13 @@ font = self.font_manager.resolve_font(
> URIs documented below are resolved relative to the plugin's
> install directory.
>
> The **Fonts** tab in the web UI that lists detected
> manager-registered fonts is still a **placeholder
> implementation** — fonts that managers register through
> `register_manager_font()` do not yet appear there. The
> programmatic per-element override workflow described in
> [Manual Font Overrides](#manual-font-overrides) below
> (`set_override()` / `remove_override()` / the
> `config/font_overrides.json` store) **does** work today and is
> the supported way to override a font for an element until the
> Fonts tab is wired up. If you can't wait and need a workaround
> right now, you can also just load the font directly with PIL
> (or `freetype-py` for BDF) inside your plugin's `manager.py`
> and skip the override system entirely.
> The web UI's **Fonts** tab lists, uploads, previews and deletes the
> font files in `assets/fonts/`. It does not show fonts registered
> through `register_manager_font()` and has no override editor (the
> override panels and `/api/v3/fonts/overrides` endpoints were removed).
> The programmatic override workflow in
> [Manual Font Overrides](#manual-font-overrides) below still works.
> Let users pick fonts through your plugin's own config schema.
### Plugin Font Registration
@@ -200,10 +214,10 @@ In your plugin's `manifest.json`:
### Using Plugin Fonts
```python
class PluginManager:
def __init__(self, config, display_manager, cache_manager, plugin_id):
self.font_manager = display_manager.font_manager
self.plugin_id = plugin_id
class MyPlugin(BasePlugin):
def __init__(self, plugin_id, config, display_manager, cache_manager, plugin_manager):
super().__init__(plugin_id, config, display_manager, cache_manager, plugin_manager)
self.font_manager = self._get_font_manager()
def display(self):
# Use plugin font (automatically namespaced)
@@ -219,17 +233,8 @@ class PluginManager:
## Manual Font Overrides
Users can override any font through the web interface:
1. Navigate to **Fonts** tab
2. View **Detected Manager Fonts** to see what's currently in use
3. In **Element Overrides** section:
- Select the element (e.g., "nfl.live.score")
- Choose a different font family
- Choose a different size
- Click **Add Override**
Overrides are stored in `config/font_overrides.json` and persist across restarts.
Overrides are set in code (there is no web UI or REST endpoint for them).
They are stored in `config/font_overrides.json` and persist across restarts.
### Programmatic Overrides
+4 -4
View File
@@ -83,10 +83,10 @@ You should see:
1. Open the **Display** tab
2. Set your matrix configuration:
- **Rows**: 32 or 64 (match your hardware)
- **Columns**: commonly 64 or 96; the web UI accepts any integer
in the 1–128 range, but 64 and 96 are the values the bundled
panel hardware ships with
- **Rows**: match your panel — commonly 32 or 64; any even number
from 8 to 64
- **Columns**: match your panel — commonly 64 or 96; at least 16,
with no upper limit
- **Chain Length**: Number of panels chained horizontally
- **Hardware Mapping**: usually `adafruit-hat-pwm` (with the PWM jumper
mod) or `adafruit-hat` (without). See the root README for the full list.
+5 -2
View File
@@ -59,9 +59,12 @@ sudo ./scripts/install/install_service.sh
After updating your scripts, verify they still work:
```bash
# Test installation scripts (if needed)
# Check the installation scripts are at their new paths
ls scripts/install/*.sh
sudo ./scripts/install/install_service.sh --help
./scripts/install/install_service.sh --help # prints usage only
# Note: running install_service.sh for real (with sudo, no --help)
# reinstalls, enables and restarts ledmatrix.service, ledmatrix-web.service
# and the update-verify units.
# Test permission scripts
ls scripts/fix_perms/*.sh
+81 -96
View File
@@ -1,169 +1,154 @@
# Multi-Root Workspace Setup Guide
This document explains how the LEDMatrix project uses a multi-root workspace to manage plugins as separate Git repositories.
This document explains how to work on LEDMatrix and the official plugins side
by side, with one editor workspace and the plugins loaded straight from your
plugin checkout.
## Overview
The LEDMatrix project has been migrated from a git submodule implementation to a **multi-root workspace** implementation for managing plugins. This allows:
Official plugins live in a single repository,
[ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins), with one
directory per plugin under `plugins/`. There are no separate per-plugin
repositories. For development you clone that monorepo **next to** LEDMatrix
and symlink its plugin directories into LEDMatrix's `plugin-repos/`, which is
where the plugin loader looks by default.
- ✅ Plugins to exist as independent Git repositories
- ✅ Updates to plugins without modifying the LEDMatrix project
- ✅ Easy development workflow with all repos in one workspace
- ✅ Plugin system discovers plugins via symlinks in `plugin-repos/`
- ✅ Plugin code stays in the monorepo checkout, with its own git history
- ✅ LEDMatrix discovers the plugins through symlinks in `plugin-repos/`
- ✅ `LEDMatrix.code-workspace` opens both repositories in VS Code/Cursor
## Directory Structure
```text
/home/chuck/Github/
├── LEDMatrix/ # Main project
│ ├── plugin-repos/ # Symlinks to actual repos (managed automatically)
│ │ ├── ledmatrix-clock-simple -> ../../ledmatrix-clock-simple
│ │ ├── ledmatrix-weather -> ../../ledmatrix-weather
~/Github/
├── LEDMatrix/ # Main project
│ ├── plugin-repos/ # Plugin directory the loader scans
│ │ ├── starlark-apps/ # Bundled with LEDMatrix (tracked in git)
│ │ ├── web-ui-info/ # Bundled with LEDMatrix (tracked in git)
│ │ ├── clock-simple -> ../../ledmatrix-plugins/plugins/clock-simple
│ │ ├── ledmatrix-weather -> ../../ledmatrix-plugins/plugins/ledmatrix-weather
│ │ └── ...
│ ├── LEDMatrix.code-workspace # Multi-root workspace configuration
│ ├── LEDMatrix.code-workspace # Opens LEDMatrix and ../ledmatrix-plugins
│ └── ...
├── ledmatrix-clock-simple/ # Plugin repository (actual git repo)
├── ledmatrix-weather/ # Plugin repository (actual git repo)
├── ledmatrix-football-scoreboard/ # Plugin repository (actual git repo)
└── ... # Other plugin repos
└── ledmatrix-plugins/ # Plugin monorepo (git repo)
├── plugins/
│ ├── clock-simple/
│ ├── ledmatrix-weather/
│ └── ...
├── plugins.json # Store registry
└── update_registry.py
```
## How It Works
### 1. Plugin Repositories
### 1. The plugin monorepo
All plugin repositories are cloned to `/home/chuck/Github/` (parent directory of LEDMatrix) as regular Git repositories:
Clone ledmatrix-plugins into the same parent directory as LEDMatrix (the
scripts below look for `../ledmatrix-plugins` relative to the LEDMatrix
root):
- `ledmatrix-clock-simple/`
- `ledmatrix-weather/`
- `ledmatrix-football-scoreboard/`
- etc.
```bash
cd ~/Github
git clone https://github.com/ChuckBuilds/ledmatrix-plugins.git
```
### 2. Symlinks in plugin-repos/
The `LEDMatrix/plugin-repos/` directory contains symlinks pointing to the actual repositories in the parent directory. This allows the plugin system to discover plugins without modifying the project structure.
`scripts/setup_plugin_repos.py` creates one symlink per plugin in
`LEDMatrix/plugin-repos/`, named after the plugin's manifest `id` and pointing
at `../ledmatrix-plugins/plugins/<dir>`.
### 3. Multi-Root Workspace
### 3. Multi-root workspace
The `LEDMatrix.code-workspace` file configures VS Code/Cursor to open all plugin repositories as separate workspace roots, allowing easy development across all repos.
`LEDMatrix.code-workspace` has two roots: LEDMatrix itself and
`../ledmatrix-plugins`.
## Setup Scripts
### Initial Setup
If you already have plugin repositories cloned, use the setup script:
```bash
cd /home/chuck/Github/LEDMatrix
cd ~/Github/LEDMatrix
python3 scripts/setup_plugin_repos.py
```
This script:
- Reads the workspace configuration
- Creates symlinks in `plugin-repos/` pointing to actual repos
- Verifies all links are created correctly
- Reads each `manifest.json` under `../ledmatrix-plugins/plugins/`
- Creates `plugin-repos/<id>` symlinks (relative) to those directories
- Leaves correct links alone, replaces links that point elsewhere, and skips
(does not overwrite) a real directory of the same name — for example a
plugin you installed from the Plugin Store. Remove that directory first if
you want the linked copy.
### Updating Plugins
To update all plugin repositories:
```bash
cd /home/chuck/Github/LEDMatrix
cd ~/Github/LEDMatrix
python3 scripts/update_plugin_repos.py
```
This script:
- Finds all plugins in the workspace
- Runs `git pull` on each repository
- Reports which plugins were updated
This runs `git pull` in `../ledmatrix-plugins` and prints the result. The
symlinks pick up the new code; restart the display to load it.
## Configuration
The plugin system is configured in `config/config.json`:
The loader reads plugins from `plugin_system.plugins_directory` in
`config/config.json`. The default is already right for this setup:
```json
{
"plugin_system": {
"plugins_directory": "plugin-repos",
"auto_discover": true,
"auto_load_enabled": true
"plugins_directory": "plugin-repos"
}
}
```
The `plugins_directory` points to `plugin-repos/`, which contains symlinks to the actual repositories.
## Workflow
### Daily Development
1. **Open Workspace**: Open `LEDMatrix.code-workspace` in VS Code/Cursor
2. **All Repos Available**: All plugin repos appear as separate folders in the workspace
3. **Edit Plugins**: Edit plugin code directly in their repositories
4. **Update Plugins**: Run `update_plugin_repos.py` to pull latest changes
2. **Edit Plugins**: Edit code under `ledmatrix-plugins/plugins/<plugin>/`
3. **Test**: `python3 run.py -e` (emulator) or
`python3 scripts/check_plugin.py --plugin <id>` from LEDMatrix
4. **Ship**: Bump `version` in the plugin's `manifest.json`, run
`python update_registry.py` in ledmatrix-plugins, commit there
### Adding New Plugins
1. **Clone Repository**: Clone the new plugin repo to `/home/chuck/Github/`
2. **Add to Workspace**: Add the plugin folder to `LEDMatrix.code-workspace`
3. **Create Symlink**: Run `setup_plugin_repos.py` to create the symlink
### Updating Individual Plugins
Since plugins are regular Git repositories, you can update them individually:
```bash
cd /home/chuck/Github/ledmatrix-weather
git pull origin master
```
Or update all at once:
```bash
cd /home/chuck/Github/LEDMatrix
python3 scripts/update_plugin_repos.py
```
## Benefits
1. **No Submodule Hassle**: No need to update `.gitmodules` or run `git submodule update`
2. **Independent Updates**: Update plugins independently without touching LEDMatrix
3. **Clean Separation**: Each plugin is a separate repository with its own history
4. **Easy Development**: Multi-root workspace makes it easy to work across repos
5. **Automatic Discovery**: Plugin system automatically discovers plugins via symlinks
1. Create `plugins/<your-plugin-id>/` in the monorepo checkout
2. Run `python3 scripts/setup_plugin_repos.py` in LEDMatrix to link it
## Troubleshooting
### Symlinks Not Working
If plugins aren't being discovered:
### Plugins not discovered
```bash
cd /home/chuck/Github/LEDMatrix
python3 scripts/setup_plugin_repos.py
cd ~/Github/LEDMatrix
ls -la plugin-repos/ # links present and not broken?
python3 scripts/setup_plugin_repos.py # recreate them
```
This will recreate all symlinks.
Also check that `plugin_system.plugins_directory` is `plugin-repos`.
### Missing Plugins
### "Monorepo plugins directory not found"
If a plugin is in the workspace but not found:
`setup_plugin_repos.py` expects the monorepo at `../ledmatrix-plugins`. Clone
it there (or symlink it there).
1. Check if the repo exists in `/home/chuck/Github/`
2. Check if the symlink exists in `plugin-repos/`
3. Run `setup_plugin_repos.py` to recreate symlinks
### Plugin updates not showing
### Plugin Updates Not Showing
If changes to plugins aren't appearing:
1. Verify the symlink points to the correct directory: `ls -la plugin-repos/ledmatrix-weather`
2. Check that you're editing in the actual repo, not a copy
3. Restart the LEDMatrix service if running
1. Verify the link target: `ls -la plugin-repos/<id>`
2. Check that you're editing the monorepo checkout, not a store-installed copy
3. Restart the LEDMatrix service (or `run.py`)
## Notes
- The `plugin-repos/` directory is tracked in git, but only contains symlinks
- Actual plugin code lives in `/home/chuck/Github/ledmatrix-*/`
- Each plugin repo can be updated independently via `git pull`
- The LEDMatrix project doesn't need to be updated when plugins change
- `plugin-repos/` is tracked in git only for the bundled plugins
(`starlark-apps`, `web-ui-info`). The symlinks you create are untracked
files; don't commit them.
- For linking a single plugin into `plugins/` instead (without a sibling
checkout), see `scripts/dev/dev_plugin_setup.sh` in the
[Plugin Development Guide](PLUGIN_DEVELOPMENT_GUIDE.md).
- When changing a plugin in the monorepo, bump its manifest `version` and run
`python update_registry.py`, or users won't receive the update.
+112 -8
View File
@@ -13,6 +13,7 @@ Complete API reference for plugin developers. This document describes all method
- [Display Manager](#display-manager)
- [Cache Manager](#cache-manager)
- [Plugin Manager](#plugin-manager)
- [Deprecated APIs](#deprecated-apis)
---
@@ -36,7 +37,11 @@ self.enabled # Boolean enabled status
#### `update() -> None`
Fetch/update data for this plugin. Called based on `update_interval` specified in the plugin's manifest.
Fetch/update data for this plugin. Called on the plugin's update interval:
the value `get_update_interval()` returns when it returns a number, otherwise
the static interval: the `update_interval` in the plugin's manifest, else
`update_interval` in the plugin's section of `config.json`, else 60 seconds
(see [`get_update_interval()`](#get_update_interval---optionalfloat) below).
**Example**:
```python
@@ -109,6 +114,46 @@ Called when plugin is enabled.
Called when plugin is disabled.
#### `get_update_interval() -> Optional[float]`
How often this plugin wants `update()` called right now, in seconds. The
manifest's `update_interval` is one static number; override this when the
right cadence depends on state only the plugin knows, e.g. poll every 15s
while a game is live and fall back to the manifest value otherwise.
**Returns**: seconds as a number, or `None` (the default) for no opinion.
How `PluginManager` (`_get_plugin_update_interval` in
`src/plugin_system/plugin_manager.py`) resolves the interval on each
scheduling tick:
1. It calls `get_update_interval()`. A number wins over everything below.
Values under `PluginManager.MIN_DYNAMIC_UPDATE_INTERVAL` (5 seconds) are
raised to it.
2. If the hook returns `None`, raises, or returns something that isn't a
finite number (a `bool`, a string, NaN, infinity), it is ignored and the
static interval applies: the manifest's `update_interval`, else
`update_interval` in the plugin's section of `config.json`, else 60
seconds.
The static value is cached per plugin until the plugin is loaded or
unloaded again, so editing `update_interval` in config takes effect on the
next reload. The hook's return value is never cached: it is called on every
tick of the display loop, so keep it to attribute reads (no config lookups,
no I/O, no locks a fetch might hold) and don't let it raise.
**Example**:
```python
def get_update_interval(self):
# Fast while something is live, manifest default otherwise.
if any(m.live_games for m in self._live_managers):
return self.config.get("live_update_interval", 15)
return None
```
Added in core 3.4.0; older cores never call it, so a plugin that relies on
it should floor `ledmatrix_min_version` at `3.4.0`.
#### `get_display_duration() -> float`
Get display duration for this plugin. Can be overridden for dynamic durations.
@@ -470,21 +515,59 @@ self.display_manager.draw_text_with_icons(
For plugins that implement scrolling content, use these methods to coordinate with the display system.
#### `set_scrolling_state(is_scrolling: bool) -> None`
#### `set_scrolling_state(is_scrolling: bool, frame_hold: int = 1) -> None`
Mark the display as scrolling or not scrolling. Call when scrolling starts/stops.
Mark the display as scrolling or not scrolling, and set this scroll's frame
pacing. Call it when a scroll starts (calling it on every scroll frame is fine)
and with `False` when it stops.
**Parameters**:
- `is_scrolling` (bool): True if currently scrolling, False otherwise
- `frame_hold` (int, default 1): how many panel refreshes each pushed frame is
held for (clamped to 1-255; ignored when `is_scrolling` is False, which
resets it to 1). Pass the `frame_hold` of the settings
`src.common.scroll_config.configure()` returned. Added in core 3.4.0.
**Why `frame_hold` matters**: `scroll_config.configure()` snaps the speed to
one the panel can show in whole pixels and sets the `ScrollHelper` to advance a
fixed number of pixels on every presented frame -- no clock is consulted. The
panel presents frames at its refresh rate divided by the hold, so the hold is
part of the speed. Omit it and a 50 px/s scroll (1px every 2nd refresh on a
100 Hz panel) runs at 100 px/s. The hold is not applied by `configure()`
because it must not outlive the scroll: plugins share one display manager.
**Example**:
```python
from src.common import scroll_config
from src.common.scroll_helper import ScrollHelper
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.scroll_helper = ScrollHelper(
self.display_manager.width, self.display_manager.height, self.logger)
# ...later, hand it content with self.scroll_helper.set_scrolling_image(img)
self.scroll_settings = scroll_config.configure(
self.scroll_helper,
plugin_config=self.config,
global_config=self.global_config,
display_manager=self.display_manager,
plugin_logger=self.logger,
)
def display(self, force_clear=False):
self.display_manager.set_scrolling_state(True)
# Scroll content...
self.display_manager.set_scrolling_state(False)
self.display_manager.set_scrolling_state(
True, frame_hold=self.scroll_settings.frame_hold)
self.scroll_helper.update_scroll_position()
self.display_manager.image = self.scroll_helper.get_visible_portion()
self.display_manager.update_display()
if self.scroll_helper.is_scroll_complete():
self.display_manager.set_scrolling_state(False)
```
Don't pace the loop with `time.sleep()`: `update_display()` blocks on the
panel's vsync, which is what paces a scroll. See `docs/SCROLL_PERFORMANCE.md`
for choosing a speed.
#### `is_currently_scrolling() -> bool`
Check if the display is currently in a scrolling state.
@@ -954,9 +1037,10 @@ if "weather" in enabled_plugins:
self.display_manager.update_display()
```
3. **Handle scrolling state**: If your plugin scrolls, use scrolling state methods
3. **Handle scrolling state**: If your plugin scrolls, use scrolling state methods,
passing the frame hold `scroll_config.configure()` returned
```python
self.display_manager.set_scrolling_state(True)
self.display_manager.set_scrolling_state(True, frame_hold=settings.frame_hold)
# Scroll content...
self.display_manager.set_scrolling_state(False)
```
@@ -991,3 +1075,23 @@ if "weather" in enabled_plugins:
- [Plugin Development Guide](PLUGIN_DEVELOPMENT_GUIDE.md) - Complete development guide
- [Advanced Plugin Development](ADVANCED_PLUGIN_DEVELOPMENT.md) - Advanced patterns and examples
---
## Deprecated APIs
These still work in 3.6 but log a warning the first time they are called
(`journalctl -u ledmatrix` shows which one), and are **removed in 3.7.0**.
Nothing in core, the official plugins or the third-party plugins in the
registry calls them.
| Object | Methods | Instead |
|---|---|---|
| `cache_manager` | `update_cache` | `set()` |
| `cache_manager` | `get_background_cached_data`, `is_background_data_available` | `get()` |
| `cache_manager` | `has_data_changed`, `setup_persistent_cache`, `get_sport_live_interval`, `get_sport_key_from_cache_key`, `record_cache_hit`, `record_cache_miss`, `record_fetch_time`, `get_cache_metrics`, `log_cache_metrics`, `get_memory_cache_stats` | no replacement |
| `display_manager` | `draw_weather_icon`, `draw_sun`, `draw_cloud`, `draw_rain`, `draw_snow`, `draw_text_with_icons` | draw your own icons (the weather plugin ships `WeatherIcons`) |
| `display_manager` | `get_scrolling_stats` | no replacement |
| `font_manager` | `get_font_catalog`, `get_available_fonts` | read `font_catalog` |
| `font_manager` | `set_override`, `remove_override`, `get_overrides`, `add_font`, `remove_font`, `validate_font`, `get_size_tokens`, `get_performance_stats`, `get_manager_fonts`, `get_detected_fonts`, `get_plugin_fonts`, `unregister_plugin_fonts` | no replacement |
| `plugin_manager` | `get_enabled_plugins` | check `enabled` on the entries in `plugin_manager.plugins` |
+20 -2
View File
@@ -8,7 +8,8 @@
> - Code paths reference `web_interface_v2.py`; the current web UI is
> `web_interface/app.py` with v3 Blueprint-based templates.
> - The example Flask routes use `/api/plugins/*`; the real API
> blueprint is mounted at `/api/v3` (`web_interface/app.py:199`).
> blueprint (`web_interface/blueprints/api_v3/`) is mounted at `/api/v3`
> in `web_interface/app.py`.
> - The default plugin location is `plugin-repos/` (configurable via
> `plugin_system.plugins_directory`), not `./plugins/`.
> - Example imports use `src/plugin_system/base_classes/*_plugin.py`;
@@ -189,7 +190,9 @@ class BasePlugin(ABC):
def update(self) -> None:
"""
Fetch/update data for this plugin.
Called based on update_interval in manifest.
Called every get_update_interval() seconds when that returns a
number, otherwise at the static interval: the manifest's
update_interval, else the plugin config's update_interval, else 60s.
"""
pass
@@ -204,6 +207,21 @@ class BasePlugin(ABC):
"""
pass
def get_update_interval(self) -> Optional[float]:
"""
Seconds until update() should run again, decided at runtime.
Return None (the default) to use the static interval.
PluginManager._get_plugin_update_interval calls this on every
scheduling tick. A number overrides the manifest and is clamped up
to PluginManager.MIN_DYNAMIC_UPDATE_INTERVAL (5s); None, a raise,
or a non-finite/non-numeric value falls back to the manifest's
update_interval, then the plugin config's update_interval, then
60s. The static value is cached until the plugin reloads; the hook
is not cached, so it must be cheap and must not raise.
"""
return None
def get_display_duration(self) -> float:
"""
Get the display duration for this plugin instance.
+4 -6
View File
@@ -67,9 +67,7 @@ The main configuration file (`config/config.json`) now contains only essential s
"time_format": "%I:%M %p"
},
"plugin_system": {
"plugins_directory": "plugin-repos",
"auto_discover": true,
"auto_load_enabled": true
"plugins_directory": "plugin-repos"
}
}
```
@@ -93,9 +91,9 @@ The main configuration file (`config/config.json`) now contains only essential s
#### 4. Plugin System
- **plugin_system**: Plugin system configuration
- **plugins_directory**: Directory where plugins are stored
- **auto_discover**: Automatically discover plugins
- **auto_load_enabled**: Automatically load enabled plugins
- **plugins_directory**: Directory where plugins are stored (the only one the loader scans)
- `auto_discover`, `auto_load_enabled`, `development_mode` may still appear in
older configs; nothing reads them (see [CONFIG_REFERENCE.md](CONFIG_REFERENCE.md#plugin_system))
## Plugin Configuration
+2 -1
View File
@@ -6,7 +6,8 @@
> in the "Implementation Details" section below still reference the
> pre-v3 file layout (`web_interface_v2.py`, `templates/index_v2.html`).
> The current implementation lives in `web_interface/app.py`,
> `web_interface/blueprints/api_v3.py`, and `web_interface/templates/v3/`.
> `web_interface/blueprints/api_v3/` (plugin config handlers in
> `plugins.py`), and `web_interface/templates/v3/`.
> The user-facing description (Overview, Features, Form Generation
> Process) is still accurate.
+134 -382
View File
@@ -11,427 +11,179 @@
### Component Overview
```
┌─────────────────────────────────────────────────────────────────┐
│ Web Browser │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Tab Navigation Bar │ │
│ │ [Overview] [General] ... [Plugins] [Plugin X] [Plugin Y]│ │
│ └─────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────┐ ┌──────────────────────────────────┐ │
│ │ Plugins Tab │ │ Plugin X Configuration Tab │ │
│ │ │ │ │ │
│ │ • Install │ │ Form Generated from Schema: │ │
│ │ • Update │ │ • Boolean → Toggle │ │
│ │ • Uninstall │ │ • Number → Number Input │ │
│ │ • Enable │ │ • String → Text Input │ │
│ │ • [Configure]──────→ • Array → Comma Input │ │
│ │ │ │ • Enum → Dropdown │ │
│ └─────────────────┘ │ │ │
│ │ [Save] [Back] [Reset] │ │
│ └──────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────────────┐
│ Web browser (templates/v3/base.html, Alpine.js + HTMX) │
│ │
│ Second nav row: one tab per installed plugin │
│ Clicking a tab: GET /v3/partials/plugin-config/<plugin_id> │
│ → server-rendered form swapped into the tab │
│ │
│ Save: hx-post="/api/v3/plugins/config?plugin_id=<id>" (form data) │
└──────────────────────────────────────────────────────────────────┘
│
│ HTTP API
▼
┌─────────────────────────────────────────────────────────────────┐
│ Flask Backend │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ /api/v3/plugins/installed │ │
│ │ • Discover plugins in plugins/ directory │ │
│ │ • Load manifest.json for each plugin │ │
│ │ • Load config_schema.json if exists │ │
│ │ • Load current config from config.json │ │
│ │ • Return combined data to frontend │ │
│ └───────────────────────────────────────────────────────┘ │
│ │
│ ┌───────────────────────────────────────────────────────┐ │
│ │ /api/v3/plugins/config │ │
│ │ • Receive key-value pair │ │
│ │ • Update config.json │ │
│ │ • Return success/error │ │
│ └───────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────────────┐
│ Flask (web_interface/app.py) │
│ │
│ pages_v3 blueprint (blueprints/pages_v3.py) │
│ _load_plugin_config_partial(plugin_id) │
│ • SchemaManager.load_schema() → config_schema.json │
│ • config.json section for the plugin │
│ • masks x-secret fields │
│ • renders partials/plugin_config.html (render_field macros) │
│ │
│ api_v3 blueprint (blueprints/api_v3/plugins.py) │
│ save_plugin_config() POST /api/v3/plugins/config │
│ get_plugin_config() GET /api/v3/plugins/config │
│ get_plugin_schema() GET /api/v3/plugins/schema │
│ reset_plugin_config() POST /api/v3/plugins/config/reset │
└──────────────────────────────────────────────────────────────────┘
│
│ File System
▼
┌─────────────────────────────────────────────────────────────────┐
│ File System │
│ │
│ plugins/ │
│ ├── hello-world/ │
│ │ ├── manifest.json ───┐ │
│ │ ├── config_schema.json ─┼─→ Defines UI structure │
│ │ ├── manager.py │ │
│ │ └── requirements.txt │ │
│ └── clock-simple/ │ │
│ ├── manifest.json │ │
│ └── config_schema.json ──┘ │
│ │
│ config/ │
│ └── config.json ────────────→ Stores configuration values │
│ { │
│ "hello-world": { │
│ "enabled": true, │
│ "message": "Hello!", │
│ ... │
│ } │
│ } │
└─────────────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────────────┐
│ Files │
│ plugin-repos/<id>/config_schema.json JSON Schema (Draft-7) │
│ config/config.json { "<id>": { ... } } │
│ config/config_secrets.json { "<id>": { secrets } } │
└──────────────────────────────────────────────────────────────────┘
```
The plugins directory is `plugin_system.plugins_directory` in
`config/config.json` (default `plugin-repos/`). Plugin configuration lives in
`config/config.json`, not in the plugin directory, so it survives reinstalls.
## Data Flow
### 1. Page Load Sequence
### 1. Rendering a plugin's tab
```
User Opens Web Interface
│
▼
DOMContentLoaded Event
│
▼
refreshPlugins()
│
▼
GET /api/v3/plugins/installed
│
├─→ For each plugin directory:
│ ├─→ Read manifest.json
│ ├─→ Read config_schema.json (if exists)
│ └─→ Read config from config.json
│
▼
Return JSON Array:
[{
id: "hello-world",
name: "Hello World",
config: { enabled: true, message: "Hello!" },
config_schema_data: {
properties: {
enabled: { type: "boolean", ... },
message: { type: "string", ... }
}
}
}, ...]
│
▼
generatePluginTabs(plugins)
│
├─→ For each plugin:
│ ├─→ Create tab button
│ ├─→ Create tab content div
│ └─→ generatePluginConfigForm(plugin)
│ │
│ ├─→ Read schema properties
│ ├─→ Get current config values
│ └─→ Generate HTML form inputs
│
▼
Tabs Rendered in UI
User opens the plugin's tab
│
▼
GET /v3/partials/plugin-config/<plugin_id> (pages_v3)
│
├─→ Load schema (SchemaManager, no cache)
├─→ Load config.json[<plugin_id>]
├─→ Mask "x-secret" values (fails closed if the schema is unusable)
└─→ render partials/plugin_config.html
│
└─→ render_field() per property, recursively:
boolean → toggle, number/integer → input or slider,
string → input / textarea / select (enum),
array → list or table widget,
object → collapsible nested section,
"x-widget" → a registered widget
(static/v3/js/widgets/, or one the plugin ships)
```
### 2. Configuration Save Sequence
Nested objects are supported: a nested field is posted with a dotted name
(e.g. `transition.type`).
### 2. Saving
```
User Modifies Form
│
▼
User Clicks "Save"
│
▼
savePluginConfiguration(pluginId)
│
├─→ Get form data
├─→ For each field:
│ ├─→ Get schema type
│ ├─→ Convert value to correct type
│ │ • boolean: checkbox.checked
│ │ • integer: parseInt()
│ │ • number: parseFloat()
│ │ • array: split(',')
│ │ • string: as-is
│ │
│ └─→ POST /api/v3/plugins/config
│ {
│ plugin_id: "hello-world",
│ key: "message",
│ value: "Hello, World!"
│ }
│
▼
Backend Updates config.json
│
▼
Return Success
│
▼
Show Notification
│
▼
Refresh Plugins
User clicks Save
│
▼
validatePluginConfigForm() (client-side checks)
│
▼
POST /api/v3/plugins/config?plugin_id=<id> (form data, all fields of the form)
│
▼
save_plugin_config() (api_v3/plugins.py)
├─→ Start from the stored config.json[<id>]
├─→ Apply form fields: dotted names → nested keys, "[]" checkbox
│ groups → lists, values coerced to the schema's types
├─→ Merge schema defaults for keys that are still missing
├─→ Validate against the schema (plus core per-plugin properties);
│ invalid → 400 with the validation errors, nothing saved
├─→ Split "x-secret" fields out; masked/blank secrets are dropped so
│ an untouched secret keeps its stored value
├─→ Deep-merge regular fields into config.json[<id>] (atomic save)
├─→ Merge secrets into config_secrets.json[<id>]
└─→ Call the loaded plugin's on_config_change() (and
on_enable/on_disable if "enabled" changed)
│
▼
One response for the whole form → notification in the UI
```
## Class and Function Hierarchy
The display service picks up the new config through its config hot reload
(ConfigService) without a restart.
### Frontend (JavaScript)
JSON clients can post `{"plugin_id": ..., "config": {...}}` instead; the keys
sent are merged onto the stored config the same way. See
[REST_API_REFERENCE.md](REST_API_REFERENCE.md#save-plugin-configuration).
```
Window Load
└── DOMContentLoaded
└── refreshPlugins()
├── fetch('/api/v3/plugins/installed')
├── renderInstalledPlugins(plugins)
└── generatePluginTabs(plugins)
└── For each plugin:
├── Create tab button
├── Create tab content
└── generatePluginConfigForm(plugin)
├── Read config_schema_data
├── Read current config
└── Generate form HTML
├── Boolean → Toggle switch
├── Number → Number input
├── String → Text input
├── Array → Comma-separated input
└── Enum → Select dropdown
### 3. Reset
User Interactions
├── configurePlugin(pluginId)
│ └── showTab(`plugin-${pluginId}`)
│
├── savePluginConfiguration(pluginId)
│ ├── Process form data
│ ├── Convert types per schema
│ └── For each field:
│ └── POST /api/v3/plugins/config
│
└── resetPluginConfig(pluginId)
├── Get schema defaults
└── For each field:
└── POST /api/v3/plugins/config
```
### Backend (Python)
```
Flask Routes
├── /api/v3/plugins/installed (GET)
│ └── api_plugins_installed()
│ ├── PluginManager.discover_plugins()
│ ├── For each plugin:
│ │ ├── PluginManager.get_plugin_info()
│ │ ├── Load config_schema.json
│ │ └── Load config from config.json
│ └── Return JSON response
│
└── /api/v3/plugins/config (POST)
└── api_plugin_config()
├── Parse request JSON
├── Load current config
├── Update config[plugin_id][key] = value
└── Save config.json
```
## File Structure
```
LEDMatrix/
│
├── web_interface_v2.py
│ └── Flask backend with plugin API endpoints
│
├── templates/
│ └── index_v2.html
│ └── Frontend with dynamic tab generation
│
├── config/
│ └── config.json
│ └── Stores all plugin configurations
│
├── plugins/
│ ├── hello-world/
│ │ ├── manifest.json ← Plugin metadata
│ │ ├── config_schema.json ← UI schema definition
│ │ ├── manager.py ← Plugin logic
│ │ └── requirements.txt
│ │
│ └── clock-simple/
│ ├── manifest.json
│ ├── config_schema.json
│ └── manager.py
│
└── docs/
├── PLUGIN_CONFIGURATION_TABS.md ← Full documentation
├── PLUGIN_CONFIG_TABS_SUMMARY.md ← Implementation summary
├── PLUGIN_CONFIG_QUICK_START.md ← Quick start guide
└── PLUGIN_CONFIG_ARCHITECTURE.md ← This file
```
`POST /api/v3/plugins/config/reset` replaces the plugin's section with the
schema defaults (keeping secrets unless `preserve_secrets` is false).
## Key Design Decisions
### 1. Dynamic Tab Generation
### 1. Server-side rendered forms
**Why**: Plugins are installed/uninstalled dynamically
**How**: JavaScript creates/removes tab elements on plugin list refresh
**Benefit**: No server-side template rendering needed
**Why**: One renderer for every plugin, no per-plugin frontend code
**How**: Jinja macros in `partials/plugin_config.html` walk the schema
**Benefit**: The settings search index is built from the same rendered HTML
(`/v3/settings/search-index`)
### 2. JSON Schema as Source of Truth
### 2. JSON Schema as source of truth
**Why**: Standard, well-documented, validation-ready
**How**: Frontend interprets schema to generate forms
**Benefit**: Plugin developers use familiar format
**Why**: Standard, well-documented, validation-ready
**How**: The same schema drives the form, the defaults and server-side validation
**Benefit**: Plugin developers use a familiar format
### 3. Individual Config Updates
### 3. Whole-form saves that merge
**Why**: Simplifies backend API
**How**: Each field saved separately via `/api/v3/plugins/config`
**Benefit**: Atomic updates, easier error handling
**Why**: A partial form (or a field the form doesn't show) must not wipe
stored values
**How**: The handler starts from the stored section and merges what was posted
**Benefit**: One request per save, atomic write
### 4. Type Conversion in Frontend
### 4. Secrets kept out of config.json
**Why**: HTML forms only return strings
**How**: JavaScript converts based on schema type before sending
**Benefit**: Backend receives correctly-typed values
### 5. No Nested Objects
**Why**: Keeps UI simple
**How**: Only flat property structures supported
**Benefit**: Easy form generation, clear to users
**Why**: `config.json` is shown in the raw editor and returned by the API
**How**: `"x-secret": true` fields go to `config_secrets.json`, which is
deep-merged back into the plugin's config at load time
**Benefit**: Plugins read secrets with plain `config.get(...)`
## Extension Points
### Adding New Input Types
### Custom input widgets
Location: `generatePluginConfigForm()` in `index_v2.html`
Set `"x-widget": "<name>"` on a property. Core widgets are in
`web_interface/static/v3/js/widgets/` (see its README); a plugin can ship its
own widget script, served from `/static/plugin-widgets/<plugin_id>/<name>.js`.
See [widget-guide.md](widget-guide.md).
```javascript
if (type === 'your-new-type') {
formHTML += `
<!-- Your custom input HTML -->
`;
}
```
### Custom actions
### Custom Validation
Buttons that run plugin scripts are declared in the manifest's
`web_ui_actions`. See [PLUGIN_WEB_UI_ACTIONS.md](PLUGIN_WEB_UI_ACTIONS.md).
Location: `savePluginConfiguration()` in `index_v2.html`
### Reacting to changes
```javascript
// Add validation before sending
if (!validateCustomConstraint(value, propSchema)) {
throw new Error('Validation failed');
}
```
Implement `on_config_change(new_config)` in the plugin (see
[PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md)).
### Backend Hook
## Where to Look
Location: `api_plugin_config()` in `web_interface_v2.py`
```python
# Add custom logic before saving
if plugin_id == 'special-plugin':
value = transform_value(value)
```
## Performance Considerations
### Frontend
- **Tab Generation**: O(n) where n = number of plugins (typically < 20)
- **Form Generation**: O(m) where m = number of config properties (typically < 10)
- **Memory**: Each plugin tab ~5KB HTML
- **Total Impact**: Negligible for typical use cases
### Backend
- **Schema Loading**: Cached after first load
- **Config Updates**: Single file write (atomic)
- **API Calls**: One per config field on save (sequential)
- **Optimization**: Could batch updates in single API call
## Security Considerations
1. **Input Validation**: Schema constraints enforced client-side (UX) and should be enforced server-side
2. **Path Traversal**: Plugin paths validated against known plugin directory
3. **XSS**: All user inputs escaped before rendering in HTML
4. **CSRF**: Flask CSRF tokens should be used in production
5. **File Permissions**: config.json requires write access
| Concern | File |
|---------|------|
| Tab partial loader | `web_interface/blueprints/pages_v3.py` (`_load_plugin_config_partial`) |
| Form template and field macros | `web_interface/templates/v3/partials/plugin_config.html` |
| Save / get / schema / reset handlers | `web_interface/blueprints/api_v3/plugins.py` |
| Schema loading, defaults, validation | `src/plugin_system/schema_manager.py` |
| Secret masking and splitting | `src/web_interface/secret_helpers.py` |
| Widgets | `web_interface/static/v3/js/widgets/` |
## Error Handling
### Frontend
- Network errors: Show notification, don't crash
- Schema errors: Graceful fallback to no config tab
- Type errors: Log to console, continue processing other fields
### Backend
- Invalid plugin_id: 400 Bad Request
- Schema not found: Return null, frontend handles gracefully
- Config save error: 500 Internal Server Error with message
## Testing Strategy
### Unit Tests
- `generatePluginConfigForm()` for each schema type
- Type conversion logic in `savePluginConfiguration()`
- Backend schema loading logic
### Integration Tests
- Full save flow: form → API → config.json
- Tab generation from API response
- Reset to defaults
### E2E Tests
- Install plugin → verify tab appears
- Configure plugin → verify config saved
- Uninstall plugin → verify tab removed
## Monitoring
### Frontend Metrics
- Time to generate tabs
- Form submission success rate
- User interactions (configure, save, reset)
### Backend Metrics
- API response times
- Config update success rate
- Schema loading errors
### User Feedback
- Are users finding the configuration interface?
- Are validation errors clear?
- Are default values sensible?
## Future Roadmap
### Phase 2: Enhanced Validation
- Real-time validation feedback
- Custom error messages
- Dependent field validation
### Phase 3: Advanced Inputs
- Color pickers for RGB arrays
- File upload for assets
- Rich text editor for descriptions
### Phase 4: Configuration Management
- Export/import configurations
- Configuration presets
- Version history/rollback
### Phase 5: Developer Tools
- Schema editor in web UI
- Live preview while editing schema
- Validation tester
- Unknown plugin or unreadable schema: the partial renders an error message
- Validation failure: `400` with `details` and `context.validation_errors`;
the form shows them and nothing is saved
- Save failure: `500` with an error message; config.json is written
atomically, so a failed save leaves the previous file intact
+110 -184
View File
@@ -2,234 +2,160 @@
## Overview
The LEDMatrix system has smart dependency installation that adapts based on who is running it. This guide explains how it works and potential pitfalls.
A plugin lists its Python packages in its `requirements.txt`. LEDMatrix
installs them for you when a plugin is installed, updated or loaded. This
guide explains where they end up and what to do when a plugin can't import a
package.
## How It Works
The rule to remember: **packages must be importable by `ledmatrix.service`,
which runs as root.** Anything installed only into another user's
`~/.local/` is invisible to it.
### Execution Context Detection
## Who Runs What
The plugin manager checks if it's running as root:
```python
running_as_root = os.geteuid() == 0
| Service | Runs as | Set by |
|---------|---------|--------|
| `ledmatrix.service` (display) | `root` | `systemd/ledmatrix.service` |
| `ledmatrix-web.service` (web UI) | the user who ran the installer (e.g. `ledpi`) | `User=__USER__` in `systemd/ledmatrix-web.service`, filled in by `scripts/install/install_service.sh` |
## How Dependencies Get Installed
### 1. Installing or updating a plugin from the web UI
The web interface is not root, so it installs through a narrow sudo helper:
1. `PluginStoreManager._install_dependencies()`
(`src/plugin_system/store_manager.py`) calls
`install_requirements_file()` (`src/common/permission_utils.py`).
2. That runs `sudo -n bash scripts/fix_perms/safe_pip_install.sh <plugin>/requirements.txt`.
The helper checks the path is the project's own `requirements.txt` or a
`requirements.txt` under `plugin-repos/` or `plugins/`, then runs
`python3 -m pip install --break-system-packages --ignore-installed -r ...`
**as root**, so the display service can import the packages.
3. The sudoers rule that allows this is written by the installer
(`first_time_install.sh`) or by `scripts/install/configure_web_sudo.sh`.
If sudo refuses (the rule isn't installed), `install_requirements_file()`
falls back to installing with the web process's own interpreter, as the web
user, and prefixes the pip output with a note like:
```
[Root install unavailable (...); installed for the current process's user only.
Packages may not be visible to ledmatrix.service if it runs as a different
user — run scripts/install/configure_web_sudo.sh to fix this.]
```
Based on this, it chooses the appropriate installation method:
Fix it by running `./scripts/install/configure_web_sudo.sh` as the web
user (not with `sudo`; it asks for your password itself), then
reinstall the plugin (or use the manual install below).
| Running As | Installation Method | Location | Accessible To |
|------------|-------------------|----------|---------------|
| **root** (systemd service) | System-wide (`--break-system-packages`) | `/usr/local/lib/python3.X/dist-packages/` | All users |
| **ledpi** or other user | User-specific (`--user`) | `~/.local/lib/python3.X/site-packages/` | Only that user |
The **Reinstall Plugin Deps** button on the web UI's Tools tab goes
through the same helper for every installed plugin.
### 2. Loading a plugin
When a plugin loads, `PluginLoader.install_dependencies()`
(`src/plugin_system/plugin_loader.py`) checks its `requirements.txt`. If the
requirements are already satisfied it does nothing; otherwise it runs
`python3 -m pip install --break-system-packages -r requirements.txt` with the
interpreter of the process doing the loading (retrying with
`--ignore-installed` when a system package without a pip RECORD file is in
the way).
In `ledmatrix.service` that process is root, so restarting the display
service installs anything missing system-wide:
```bash
sudo systemctl restart ledmatrix
```
If you run `python3 run.py` by hand as a normal user instead, pip cannot
write to the system site-packages and installs into your `~/.local/`. That
works for your manual run but not for the service.
## Common Scenarios
### ✅ Scenario 1: Normal Production Use (Recommended)
### Installing plugins from the web UI (recommended)
**What:** Services running via systemd
Use the **Plugin Manager** tab. Dependencies are installed as root through
the sudo helper and the display service can use them.
### Running the display manually for debugging
```bash
sudo systemctl start ledmatrix
sudo systemctl start ledmatrix-web
cd ~/LEDMatrix
sudo python3 run.py # same user as the service
```
- **Runs as:** root (configured in .service files)
- **Installs to:** System-wide
- **Result:** ✅ Works perfectly, all dependencies accessible
Running as your own user works for plugins whose packages are already
installed system-wide, but any *missing* package lands in `~/.local/`.
### ✅ Scenario 2: Web Interface Plugin Installation
### A plugin works when run manually but fails in the service
**What:** Installing/enabling plugins via web interface at `http://pi-ip:5000`
Its packages were installed for your user only. Install them as root (see
below) and restart the service.
- **Web service runs as:** root (ledmatrix-web.service)
- **Installs to:** System-wide
- **Result:** ✅ Works perfectly, systemd service can access them
## Manual Installation
### ✅ Scenario 3: Manual Testing as ledpi (Read-only)
**What:** Running display manually as ledpi to test/debug
### All plugins
```bash
# As ledpi user
cd /home/ledpi/LEDMatrix
python3 run.py
```
- **Runs as:** ledpi
- **Can import:** ✅ System-wide packages (installed by root)
- **Result:** ✅ Works! Can use existing plugins with root-installed dependencies
### ⚠️ Scenario 4: Manual Plugin Installation as ledpi (Problematic)
**What:** Enabling a NEW plugin and running manually as ledpi
```bash
# As ledpi user
cd /home/ledpi/LEDMatrix
# Edit config to enable new plugin
nano config/config.json
# Run display - will try to install new plugin dependencies
python3 run.py
```
**What Happens:**
1. Plugin manager runs as `ledpi`
2. Installs dependencies with `--user` flag
3. Dependencies go to `~/.local/lib/python3.X/site-packages/`
4. ⚠️ **Warning logged:** "Installing plugin dependencies for current user (not root)"
**Problem:**
- When systemd service restarts (as root), it **can't see** `~/.local/` packages
- Plugin will fail to load for the systemd service
**Solution:**
After testing, restart the service to install dependencies system-wide:
```bash
sudo ~/LEDMatrix/scripts/install_plugin_dependencies.sh
sudo systemctl restart ledmatrix
```
## Best Practices
The script installs every `requirements.txt` found in the plugins directory
configured by `plugin_system.plugins_directory` in `config/config.json`
(default `plugin-repos/`). Run it with `sudo` so the packages are installed
system-wide.
### For Production/Normal Use
### One plugin
1. **Always use the web interface** to install/enable plugins
2. **Or restart the systemd service** after config changes:
```bash
sudo systemctl restart ledmatrix
```
### For Development/Testing
1. **Read existing plugins:** Safe to run as `ledpi` - can import system packages
2. **Test new plugins:** Use sudo or restart service to install dependencies:
```bash
# Option 1: Run as root
sudo python3 run.py
# Option 2: Install deps manually
sudo pip3 install --break-system-packages -r plugins/my-plugin/requirements.txt
python3 run.py
# Option 3: Let service install them
sudo systemctl restart ledmatrix
```
## Warning Messages
### If you see this warning:
```
Installing plugin dependencies for current user (not root).
These will NOT be accessible to the systemd service.
For production use, install plugins via the web interface or restart the ledmatrix service.
```
**What it means:**
- You're running as a non-root user
- Dependencies were installed to your user directory only
- The systemd service won't be able to use this plugin
**What to do:**
```bash
# Restart the service to install dependencies system-wide
cd ~/LEDMatrix/plugin-repos/PLUGIN-NAME # or your configured plugins directory
sudo python3 -m pip install --break-system-packages --no-cache-dir -r requirements.txt
sudo systemctl restart ledmatrix
```
`--no-cache-dir` avoids errors about `/root/.cache/pip` not being writable.
## Troubleshooting
### Plugin works when I run manually but fails in systemd service
**Cause:** Dependencies installed to user directory (`~/.local/`) instead of system-wide
**Fix:**
```bash
# Check where package is installed
pip3 list -v | grep <package-name>
# If it shows ~/.local/, reinstall system-wide:
sudo pip3 install --break-system-packages <package-name>
# Or just restart the service:
sudo systemctl restart ledmatrix
```
### Permission denied when installing dependencies
**If you see errors like:**
```
ERROR: Could not install packages due to an OSError: [Errno 13] Permission denied: '/root/.local'
WARNING: The directory '/root/.cache/pip' or its parent directory is not owned or is not writable
```
**Quick Fix - Use the Helper Script:**
```bash
sudo /home/ledpi/LEDMatrix/scripts/install_plugin_dependencies.sh
sudo systemctl restart ledmatrix
```
Use one of the manual installs above (they pass `--no-cache-dir`).
**Manual Fix:**
```bash
# Install dependencies with --no-cache-dir to avoid cache permission issues
cd /home/ledpi/LEDMatrix/plugins/PLUGIN-NAME
sudo pip3 install --break-system-packages --no-cache-dir -r requirements.txt
sudo systemctl restart ledmatrix
```
**For more detailed troubleshooting, see:** [Plugin Dependency Troubleshooting Guide](PLUGIN_DEPENDENCY_TROUBLESHOOTING.md)
## Architecture Summary
```
┌─────────────────────────────────────────────────────────────┐
│ LEDMatrix Services │
├─────────────────────────────────────────────────────────────┤
│ │
│ ledmatrix.service (User=root) │
│ ledmatrix-web.service (User=root) │
│ ├── Install dependencies system-wide │
│ └── Accessible to all users │
│ │
├─────────────────────────────────────────────────────────────┤
│ │
│ Manual execution as ledpi │
│ ├── Can READ system-wide packages ✅ │
│ ├── WRITES go to ~/.local/ ⚠️ │
│ └── Not accessible to root service │
│ │
└─────────────────────────────────────────────────────────────┘
```
## Recommendations
1. **For end users:** Always use the web interface for plugin management
2. **For developers:** Be aware of the user context when testing
3. **For plugin authors:** Test with `sudo systemctl restart ledmatrix` to ensure dependencies install correctly
4. **For CI/CD:** Always run installation as root or use the service
## Helper Scripts
### Install Plugin Dependencies Script
Located at: `scripts/install_plugin_dependencies.sh`
This script automatically finds and installs dependencies for all plugins:
### Checking where a package is installed
```bash
# Run as root (recommended for production)
sudo /home/ledpi/LEDMatrix/scripts/install_plugin_dependencies.sh
# How the service sees it
sudo python3 -c "import package_name; print(package_name.__file__)"
# Make executable if needed
chmod +x /home/ledpi/LEDMatrix/scripts/install_plugin_dependencies.sh
# A path under /home/<user>/.local/ means it was installed for that user only
python3 -m pip show -f package_name
```
Features:
- Auto-detects all plugins with requirements.txt
- Uses correct installation method (system-wide vs user)
- Bypasses pip cache to avoid permission issues
- Provides detailed logging and error messages
For more, see the [Plugin Dependency Troubleshooting Guide](PLUGIN_DEPENDENCY_TROUBLESHOOTING.md).
## For Plugin Authors
1. Keep `requirements.txt` minimal and pin only what you need.
2. Test that it installs the way the Pi will install it:
```bash
sudo python3 -m pip install --break-system-packages --no-cache-dir -r requirements.txt
```
3. Note any `apt` packages your plugin needs in its README.
## Files to Reference
- Service configs: `ledmatrix.service`, `ledmatrix-web.service`
- Plugin manager: `src/plugin_system/plugin_manager.py`
- Installation script: `first_time_install.sh`
- Dependency installer: `scripts/install_plugin_dependencies.sh`
- Troubleshooting guide: `PLUGIN_DEPENDENCY_TROUBLESHOOTING.md`
- Service units: `systemd/ledmatrix.service`, `systemd/ledmatrix-web.service`
- Store installs: `src/plugin_system/store_manager.py` (`_install_dependencies`)
- Root install helper: `src/common/permission_utils.py` (`install_requirements_file`), `scripts/fix_perms/safe_pip_install.sh`
- Load-time installs: `src/plugin_system/plugin_loader.py` (`install_dependencies`)
- Sudo rules: `scripts/install/configure_web_sudo.sh`
- Manual installer: `scripts/install_plugin_dependencies.sh`
+77 -82
View File
@@ -1,6 +1,7 @@
# Plugin Dependency Installation Troubleshooting
This guide helps resolve issues with automatic plugin dependency installation in the LEDMatrix system.
This guide helps resolve problems installing a plugin's Python packages. For
how installation works, see the [Plugin Dependency Guide](PLUGIN_DEPENDENCY_GUIDE.md).
## Common Error Symptoms
@@ -10,109 +11,118 @@ ERROR: Could not install packages due to an OSError: [Errno 13] Permission denie
WARNING: The directory '/root/.cache/pip' or its parent directory is not owned or is not writable
```
### Context Mismatch
### Installed for the wrong user
The pip output shown after a web-UI install starts with:
```
WARNING: Installing plugin dependencies for current user (not root).
These will NOT be accessible to the systemd service.
[Root install unavailable (...); installed for the current process's user only.
Packages may not be visible to ledmatrix.service if it runs as a different
user — run scripts/install/configure_web_sudo.sh to fix this.]
```
### Plugin fails to load with `ModuleNotFoundError`
The display service can't see a package the plugin needs.
## Root Cause
Plugin dependencies must be installed in a context accessible to the LEDMatrix systemd service, which runs as root. Permission errors typically occur when:
Plugin packages must be importable by `ledmatrix.service`, which runs as
root. The web interface (`ledmatrix-web.service`) runs as the user who
installed LEDMatrix, so it installs through a sudo helper
(`scripts/fix_perms/safe_pip_install.sh`). Problems usually come from:
1. The pip cache directory has incorrect permissions
2. The process tries to install to user directories without proper permissions
3. Environment variables (like HOME) are not set correctly for the service context
1. The sudoers rule for that helper missing, so the web UI installed the
packages for its own user only
2. Running `python3 run.py` by hand as a normal user, which installs missing
packages into `~/.local/`
3. pip's cache directory not being writable for root
## Solutions
### Solution 1: Use the Manual Installation Script (Recommended)
We provide a helper script that handles dependency installation correctly:
### Solution 1: Restore the sudo rule, then reinstall
```bash
# Run as root to install system-wide (for production)
sudo /home/ledpi/LEDMatrix/scripts/install_plugin_dependencies.sh
cd ~/LEDMatrix
./scripts/install/configure_web_sudo.sh # as the web user, not with sudo
```
# After installation, restart the service
Then reinstall the plugin from the **Plugin Manager** tab, or click
**Reinstall Plugin Deps** on the **Tools** tab.
### Solution 2: Install every plugin's dependencies from the terminal
```bash
sudo ~/LEDMatrix/scripts/install_plugin_dependencies.sh
sudo systemctl restart ledmatrix
```
This script:
- Detects all plugins with requirements.txt files
- Installs dependencies with correct permissions
- Uses `--no-cache-dir` to avoid cache permission issues
- Provides detailed logging for troubleshooting
The script finds each `requirements.txt` in the plugins directory set by
`plugin_system.plugins_directory` in `config/config.json` (default
`plugin-repos/`), installs with `--no-cache-dir`, and reports what it found.
### Solution 2: Manual Installation per Plugin
If you need to install dependencies for a specific plugin:
### Solution 3: Install one plugin's dependencies
```bash
# Navigate to the plugin directory
cd /home/ledpi/LEDMatrix/plugins/PLUGIN-NAME
# Your configured plugins directory; plugin-repos/ by default
cd ~/LEDMatrix/plugin-repos/PLUGIN-NAME
# Install as root (system-wide)
sudo pip3 install --break-system-packages --no-cache-dir -r requirements.txt
sudo python3 -m pip install --break-system-packages --no-cache-dir -r requirements.txt
# Restart the service
sudo systemctl restart ledmatrix
```
### Solution 3: Fix Cache Directory Permissions
### Solution 4: Let the display service install them
If you specifically have cache permission issues:
When a plugin loads, the display service installs any missing requirements
itself, as root:
```bash
sudo systemctl restart ledmatrix
sudo journalctl -u ledmatrix -f # watch for "Installing dependencies for plugin ..."
```
### Solution 5: Fix pip cache permissions
```bash
# Option A: Skip the cache (recommended)
sudo pip3 install --no-cache-dir --break-system-packages -r requirements.txt
sudo python3 -m pip install --no-cache-dir --break-system-packages -r requirements.txt
# Option B: Fix cache permissions (if needed)
# Option B: Fix cache permissions
sudo mkdir -p /root/.cache/pip
sudo chown -R root:root /root/.cache
sudo chmod -R 755 /root/.cache
```
### Solution 4: Install via Web Interface
The web interface handles dependency installation correctly in the service context:
1. Access the web interface (`http://ledpi:5000` or `http://your-pi-ip:5000`)
2. Open the **Plugin Manager** tab (use the **Plugin Store** section to
find the plugin, or **Install from GitHub**)
3. Install the plugin through the web UI
4. The system automatically handles dependency installation in the
service context (which has the right permissions)
## Prevention
### For Plugin Developers
When creating plugins with dependencies:
1. **Keep requirements minimal**: Only include essential packages
2. **Test installation**: Verify your requirements.txt works with:
2. **Test installation** the way the Pi does it:
```bash
sudo pip3 install --break-system-packages --no-cache-dir -r requirements.txt
sudo python3 -m pip install --break-system-packages --no-cache-dir -r requirements.txt
```
3. **Document dependencies**: Note any system packages needed (via apt)
### For Users
1. **Use web interface**: Install plugins via the web UI when possible
2. **Install as root**: When using SSH/terminal, use sudo for plugin installations
3. **Restart service**: After manual installations, restart the ledmatrix service
1. **Use the web interface** to install plugins
2. **Use sudo** for installs from SSH/terminal
3. **Restart the service** after manual installations
## Technical Details
### How Dependency Installation Works
### Where installs happen
The `PluginManager._install_plugin_dependencies()` method:
1. Detects if running as root using `os.geteuid() == 0`
2. If root: Uses system-wide installation with `--break-system-packages --no-cache-dir`
3. If not root: Uses user installation with `--user --break-system-packages --no-cache-dir`
4. The `--no-cache-dir` flag prevents cache-related permission issues
- **Web UI install/update:** `PluginStoreManager._install_dependencies()`
→ `install_requirements_file()` in `src/common/permission_utils.py`, which
runs `sudo -n bash scripts/fix_perms/safe_pip_install.sh <requirements.txt>`.
The helper only accepts the project's `requirements.txt` or one under
`plugin-repos/` or `plugins/`, and runs
`pip install --break-system-packages --ignore-installed` as root. If sudo
refuses, it falls back to a pip install as the web user and says so.
- **Plugin load:** `PluginLoader.install_dependencies()` in
`src/plugin_system/plugin_loader.py` skips satisfied requirements and
otherwise runs `pip install --break-system-packages` with the loading
process's interpreter — root in `ledmatrix.service`.
### Why `--break-system-packages`?
@@ -120,26 +130,20 @@ Debian 12+ (Bookworm) and Raspberry Pi OS based on it implement PEP 668, which p
### Service Context
The ledmatrix.service runs as:
- **User**: root
- **WorkingDirectory**: /home/ledpi/LEDMatrix
- **Python**: /usr/bin/python3
- `ledmatrix.service` runs as **root** with `/usr/bin/python3`
- `ledmatrix-web.service` runs as **the installing user**
Dependencies must be installed in root's Python environment or system-wide to be accessible.
Dependencies must be installed system-wide (as root) to be visible to the
display service.
## Checking Installation
Verify dependencies are installed correctly:
```bash
# Check as root (how the service sees it)
sudo python3 -c "import package_name"
sudo python3 -c "import package_name; print(package_name.__file__)"
# List installed packages
pip3 list
# Check specific package
pip3 show package_name
# A path under /home/<user>/.local/ means a user-only install
python3 -m pip show -f package_name
```
## Getting Help
@@ -151,19 +155,11 @@ If you continue to experience issues:
sudo journalctl -u ledmatrix -f
```
2. Check pip logs (created by manual script):
2. Verify the plugin manifest and requirements (default plugins directory
shown):
```bash
cat /tmp/pip_install_*.log
```
3. Verify plugin manifest is correct:
```bash
cat /home/ledpi/LEDMatrix/plugins/PLUGIN-NAME/manifest.json
```
4. Check plugin requirements:
```bash
cat /home/ledpi/LEDMatrix/plugins/PLUGIN-NAME/requirements.txt
cat ~/LEDMatrix/plugin-repos/PLUGIN-NAME/manifest.json
cat ~/LEDMatrix/plugin-repos/PLUGIN-NAME/requirements.txt
```
## Related Documentation
@@ -171,4 +167,3 @@ If you continue to experience issues:
- [Plugin Dependency Guide](PLUGIN_DEPENDENCY_GUIDE.md)
- [Plugin Development Guide](PLUGIN_DEVELOPMENT_GUIDE.md)
- [Troubleshooting](TROUBLESHOOTING.md)
+114 -93
View File
@@ -3,18 +3,20 @@
This guide explains how to set up a development workflow for plugins that are maintained in separate Git repositories while still being able to test them within the LEDMatrix project.
> **Rendering guidance:** plugins should read the display size dynamically
> (`self.display_manager.matrix.width/height`) rather than hardcoding one
> panel. For plugins that want to *scale* their layout to any panel, the
> (`self.display_manager.width/height`) rather than hardcoding one
> panel. Don't read `display_manager.matrix.width/height`: `matrix` is
> `None` when hardware init fails, while the `width`/`height` properties
> fall back to the canvas size. For plugins that want to *scale* their layout to any panel, the
> opt-in adaptive layout system ([ADAPTIVE_LAYOUT.md](ADAPTIVE_LAYOUT.md))
> provides the shared helpers — fonts, images, and composite layouts that
> scale. Existing plugins keep their classic rendering unless they adopt
> those APIs; nothing migrates automatically.
> **Just want a different look for an existing sports scoreboard?** You may
> not need a plugin at all — a **skin** restyles the live/recent/upcoming
> rendering while the plugin keeps handling data, scheduling, caching, and
> vegas mode, in ~100 lines of drawing code. See
> [CREATING_SKINS.md](CREATING_SKINS.md).
> **Want a different look for an existing sports scoreboard?** Skins are
> meant for that, but they are **not supported yet**: the current scoreboard
> plugins don't render them (see [SKIN_SYSTEM.md](SKIN_SYSTEM.md#status-not-supported-yet)).
> For now, change the look through the plugin's own display settings or its
> code.
## Overview
@@ -43,28 +45,45 @@ The solution uses **symbolic links** to connect plugin repositories to the `plug
## Quick Start
### 1. Link a Plugin from GitHub
Official plugins all live in one repository,
[ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins), with
one directory per plugin under `plugins/` (there are no per-plugin
`ledmatrix-<name>` repositories). The helper script links a plugin directory
from a checkout of that monorepo into LEDMatrix's `plugins/` directory.
The easiest way to link a plugin that's already on GitHub:
### 1. Link an Official Plugin
```bash
./scripts/dev/dev_plugin_setup.sh link-github music
./scripts/dev/dev_plugin_setup.sh link-github football-scoreboard
```
This will:
- Clone `https://github.com/ChuckBuilds/ledmatrix-music.git` to `~/.ledmatrix-dev-plugins/ledmatrix-music`
- Create a symbolic link from `plugins/music` to the cloned repository
- Validate that the plugin has a proper `manifest.json`
- Clone `https://github.com/ChuckBuilds/ledmatrix-plugins.git` to
`~/.ledmatrix-dev-plugins/ledmatrix-plugins` (or `git pull` it if it is
already there)
- Find `plugins/football-scoreboard` in it (also accepted:
`plugins/ledmatrix-<name>`, or a plugin whose manifest `id` is the name)
- Validate that it has a `manifest.json`
- Create a symbolic link named after the plugin's manifest id, e.g.
`plugins/football-scoreboard` → `~/.ledmatrix-dev-plugins/ledmatrix-plugins/plugins/football-scoreboard`
### 2. Link a Local Plugin Repository
`link-github music` finds the monorepo's `plugins/ledmatrix-music` directory
and links it into LEDMatrix as `plugins/ledmatrix-music`, because
`ledmatrix-music` is that plugin's manifest id.
If you already have a plugin repository cloned locally:
To work from your fork of the monorepo, set `github_user` in
`dev_plugins.json` (see [Configuration](#configuration)).
### 2. Link a Local Plugin Directory
If you already have the monorepo (or a third-party plugin repository) cloned
locally:
```bash
./scripts/dev/dev_plugin_setup.sh link music ../ledmatrix-music
./scripts/dev/dev_plugin_setup.sh link hello-world ../ledmatrix-plugins/plugins/hello-world
```
This creates a symlink from `plugins/music` to your local repository path.
This creates a symlink from `plugins/hello-world` to that directory.
### 3. Check Status
@@ -77,13 +96,17 @@ See which plugins are linked and their git status:
### 4. Work on Your Plugin
```bash
cd plugins/music # Actually editing the linked repository
# Make your changes
cd plugins/football-scoreboard # Actually editing the monorepo checkout
# Make your changes, then bump "version" in manifest.json
git add .
git commit -m "feat: add new feature"
git push origin main
git commit -m "feat(football-scoreboard): add new feature"
git push # to your fork, then open a PR against ledmatrix-plugins
```
In the monorepo, every plugin change must bump `version` in the plugin's
`manifest.json` and run `python update_registry.py`, or users won't receive
the update.
### 5. Update Plugins
Pull latest changes from remote:
@@ -116,7 +139,7 @@ Links a local plugin repository to the plugins directory.
**Example:**
```bash
./scripts/dev/dev_plugin_setup.sh link football-scoreboard ../ledmatrix-football-scoreboard
./scripts/dev/dev_plugin_setup.sh link football-scoreboard ../ledmatrix-plugins/plugins/football-scoreboard
```
**Notes:**
@@ -129,23 +152,25 @@ Links a local plugin repository to the plugins directory.
Clones a plugin from GitHub and links it.
**Arguments:**
- `plugin-name`: The name of the plugin (will be the directory name in `plugins/`)
- `repo-url`: (Optional) Full GitHub repository URL. If omitted, constructs from pattern: `https://github.com/ChuckBuilds/ledmatrix-<plugin-name>.git`
- `plugin-name`: Without `repo-url`, the plugin to link from the monorepo: a
directory under `plugins/` (`<name>` or `ledmatrix-<name>`) or a manifest
id. The link is named after the plugin's manifest id. With `repo-url`, the
name of the link in `plugins/`.
- `repo-url`: (Optional) A plugin that has its own repository (e.g. a
third-party plugin). The repository root is linked.
**Examples:**
```bash
# Auto-construct URL from plugin name
./scripts/dev/dev_plugin_setup.sh link-github music
# Official plugin, from the ledmatrix-plugins monorepo
./scripts/dev/dev_plugin_setup.sh link-github stocks
# Use explicit URL
./scripts/dev/dev_plugin_setup.sh link-github stocks https://github.com/ChuckBuilds/ledmatrix-stocks.git
# Link from a different GitHub user
# Third-party plugin with its own repository
./scripts/dev/dev_plugin_setup.sh link-github custom-plugin https://github.com/OtherUser/custom-plugin.git
```
**Notes:**
- Repositories are cloned to `~/.ledmatrix-dev-plugins/` by default (configurable)
- The monorepo is cloned once and shared by every plugin you link from it
- If the repository already exists, it will be updated with `git pull` instead of re-cloning
- The cloned repository is preserved when you unlink the plugin
@@ -217,30 +242,28 @@ Updates plugin(s) by running `git pull` in their repositories.
### Custom Development Directory
By default, GitHub repositories are cloned to `~/.ledmatrix-dev-plugins/`. You can customize this by creating a `dev_plugins.json` file:
By default, GitHub repositories are cloned to `~/.ledmatrix-dev-plugins/`
and official plugins come from `ChuckBuilds/ledmatrix-plugins`. To change
either, copy `dev_plugins.json.example` (in the LEDMatrix root) to
`dev_plugins.json` and edit it. `dev_plugins.json` is git-ignored.
```json
{
"dev_plugins_dir": "/path/to/your/dev/plugins",
"github_user": "ChuckBuilds",
"github_pattern": "ledmatrix-",
"plugins": {
"music": {
"source": "github",
"url": "https://github.com/ChuckBuilds/ledmatrix-music.git",
"branch": "main"
}
}
"dev_plugins_dir": "~/.ledmatrix-dev-plugins",
"github_user": "your-github-user",
"plugins_repo": "ledmatrix-plugins",
"plugins_branch": "main"
}
```
**Configuration options:**
**Configuration options** (all optional):
- `dev_plugins_dir`: Where to clone GitHub repositories (default: `~/.ledmatrix-dev-plugins`)
- `github_user`: Default GitHub username for auto-constructing URLs
- `github_pattern`: Pattern for repository names (default: `ledmatrix-`)
- `plugins`: Plugin definitions (optional, for future auto-discovery features)
- `github_user`: Owner of the plugin monorepo that `link-github <name>` clones — set it to use your fork (default: `ChuckBuilds`)
- `plugins_repo`: Name of that monorepo (default: `ledmatrix-plugins`)
- `plugins_branch`: Branch to clone it at (default: the repository's default branch). Only applies when the clone is first made.
**Note:** Copy `dev_plugins.json.example` to `dev_plugins.json` and customize it. The `dev_plugins.json` file is git-ignored.
`github_pattern` from older versions of this guide is no longer used (the
script warns if it is set).
## Development Workflow
@@ -248,43 +271,46 @@ By default, GitHub repositories are cloned to `~/.ledmatrix-dev-plugins/`. You c
1. **Link your plugin for development:**
```bash
./scripts/dev/dev_plugin_setup.sh link-github music
./scripts/dev/dev_plugin_setup.sh link-github clock-simple
```
2. **Test in LEDMatrix:**
```bash
# Run LEDMatrix with your plugin
python run.py
# Run LEDMatrix with your plugin (emulator shown)
python3 run.py -e
```
3. **Make changes:**
```bash
cd plugins/music
cd plugins/clock-simple
# Edit files...
# Test changes...
```
4. **Commit to plugin repository:**
4. **Commit to the plugin repository:**
```bash
cd plugins/music # This is actually your repo
cd plugins/clock-simple # This is inside your monorepo checkout
# bump "version" in manifest.json, then from the monorepo root:
# python update_registry.py
git add .
git commit -m "feat: add new feature"
git push origin main
git commit -m "feat(clock-simple): add new feature"
git push
```
5. **Update from remote (if needed):**
```bash
./scripts/dev/dev_plugin_setup.sh update music
./scripts/dev/dev_plugin_setup.sh update clock-simple
```
6. **When done developing:**
```bash
./scripts/dev/dev_plugin_setup.sh unlink music
./scripts/dev/dev_plugin_setup.sh unlink clock-simple
```
### Working with Multiple Plugins
You can have multiple plugins linked simultaneously:
You can have multiple plugins linked simultaneously. Plugins linked from the
monorepo share one checkout:
```bash
./scripts/dev/dev_plugin_setup.sh link-github music
@@ -294,7 +320,7 @@ You can have multiple plugins linked simultaneously:
# Check status of all
./scripts/dev/dev_plugin_setup.sh status
# Update all at once
# Update all at once (the shared monorepo checkout is pulled once)
./scripts/dev/dev_plugin_setup.sh update
```
@@ -409,7 +435,7 @@ If you have conflicts when updating:
1. **Manually resolve in the plugin repository:**
```bash
cd ~/.ledmatrix-dev-plugins/ledmatrix-music
cd ~/.ledmatrix-dev-plugins/ledmatrix-plugins
git pull
# Resolve conflicts...
git add .
@@ -470,18 +496,19 @@ You can mix local and GitHub plugins:
The development workflow is separate from the plugin store installation:
- **Plugin Store:** Installs plugins to `plugins/` as regular directories
- **Development Setup:** Links plugin repositories as symlinks
- **Plugin Store:** Installs plugins as regular directories in the configured
plugins directory (`plugin-repos/` by default)
- **Development Setup:** Links plugin directories as symlinks in `plugins/`
If you install a plugin via the store, you can still link it for development:
The plugin loader scans only one directory, so while developing set
`plugin_system.plugins_directory` to `plugins` (see the note at the top of
this guide). If `plugins/` already holds a regular directory of the same
name, `link`/`link-github` offers to rename it to
`<name>.backup.<timestamp>` before linking.
```bash
# Store installs to plugins/music (regular directory)
# Link for development (will prompt to replace)
./scripts/dev/dev_plugin_setup.sh link-github music
```
When you unlink, the directory is removed. If you want to switch back to the store version, re-install it via the plugin store.
`unlink` removes only the symlink. To switch back to the store version, set
`plugins_directory` back to `plugin-repos` (or reinstall the plugin from the
store).
## API Reference
@@ -525,7 +552,7 @@ Want to create and share your own plugin? Here's everything you need to know.
- [Advanced Plugin Development](ADVANCED_PLUGIN_DEVELOPMENT.md) - Patterns and examples
2. **Start with a template**:
- Use the [Hello World plugin](https://github.com/ChuckBuilds/ledmatrix-hello-world) as a starting point
- Use the [Hello World plugin](https://github.com/ChuckBuilds/ledmatrix-plugins/tree/main/plugins/hello-world) as a starting point
- Or fork an existing plugin and modify it
3. **Follow the plugin structure**:
@@ -589,24 +616,16 @@ Your plugin must:
### Versioning Best Practices
- **Use semantic versioning**: `MAJOR.MINOR.PATCH` (e.g., `1.2.3`)
- **GitHub as source of truth**: the plugin store resolves versions in this
order: GitHub Releases → GitHub Tags → manifest from branch → git commit hash
- **Automatic version bumping**: install the self-contained pre-push hook in
your plugin repo and patch versions bump themselves on push (a git tag
`v{version}` is created and `manifest.json` staged automatically):
```bash
# From your plugin repository directory
cp /path/to/LEDMatrix/scripts/git-hooks/pre-push-plugin-version .git/hooks/pre-push
chmod +x .git/hooks/pre-push
```
Set `SKIP_TAG=1` in the environment to skip auto-tagging for one push.
- **Manual versioning**: only needed for major/minor bumps, CI pipelines that
bypass hooks, or forks without the hook — use
`scripts/bump_plugin_version.py`.
- **Registry stores no versions**: `plugins.json` holds only metadata (name,
description, repo URL).
- **Bump `version` in `manifest.json` by hand** for every change you ship.
There is no automatic version-bump hook or bump script.
- **Official (monorepo) plugins**: after bumping the manifest, run
`python update_registry.py` in the `ledmatrix-plugins` checkout. It copies
each manifest's version into `plugins.json` as `latest_version`, which is
what the store compares installed versions against. Without it, users
won't be offered the update.
- **Plugins in their own repository**: still bump the manifest `version`,
so users can see which version they run; tagging releases (`v1.2.3`) to
match is a good habit.
### Submitting to Official Registry
@@ -618,12 +637,14 @@ To have your plugin added to the official plugin store:
- Follows best practices
- Tested on Raspberry Pi hardware
2. **Create GitHub repository**:
- Repository name: `ledmatrix-<plugin-name>`
- Public repository
- Proper README.md with installation instructions
2. **Choose where it lives** (see `SUBMISSION.md` in
[ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins)):
- **In the monorepo (preferred):** fork ledmatrix-plugins, add
`plugins/<your-plugin-id>/`, and open a pull request
- **In your own public repository** (conventionally
`ledmatrix-<plugin-name>`), with a README that covers installation
3. **Contact maintainers**:
3. **Contact maintainers** (own-repository plugins):
- Open a GitHub issue in the [ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins) repository
- Or reach out on Discord: https://discord.gg/uW36dVAtcT
- Include: Repository URL, plugin description, why it's useful
+135
View File
@@ -0,0 +1,135 @@
# Per-element styling for plugin authors
Users want to change the font, size and colour of individual things on screen,
nudge them a few pixels, hide the ones they do not care about, and scale a logo
down. This is the one system that does that, and a plugin joins it by
**declaring elements in its `config_schema.json`** — not by writing a style
resolver, a font cache or a web form.
The short version:
```jsonc
"customization": {
"type": "object",
"title": "Display Customization",
"x-style-elements": {
"score_text": {
"title": "Score",
"font": { "default": "PressStart2P-Regular.ttf" },
"size": { "default": 10, "min": 4, "max": 16 },
"color": { "default": [255, 255, 255] },
"offsets": true,
"visible": true,
"align": true
},
"home_logo": { "title": "Home logo", "offsets": true, "scale": true }
},
"x-style-modes": ["live", "upcoming", "recent"]
}
```
That is the whole declaration. The core expands it into a full JSON Schema, the
web UI renders a compact style editor with a row per element, and the values
land in `config.json` under the keys you named.
## What each key does
| Key | Effect |
|---|---|
| `font` | Font picker listing every shipped **and user-uploaded** font. |
| `size` | Number field. `min`/`max` also cap which fixed-size fonts are offered. |
| `color` | Colour swatch; stored as `[r, g, b]`. |
| `offsets` | X/Y nudge, stored under `customization.layout.<element>`. |
| `visible` | Show/hide toggle. |
| `align` | `left` / `center` / `right`. |
| `scale` | Size multiplier, for logos and images. Also under `layout`. |
`x-style-modes` is optional. Declare it and every element gains a per-mode
override tab — a scoreboard can then style its live, upcoming and recent cards
separately. **A mode field left blank means "inherit", not zero.**
## Reading the values
Every plugin inherits `BasePlugin.styles`, which finds your `config_schema.json`
on its own:
```python
style = self.styles.style(
"score_text",
classic_font="PressStart2P-Regular.ttf", # what you shipped
classic_size=10,
classic_color=(255, 255, 255),
)
if style.visible:
draw.text((x + style.offset[0], y + style.offset[1]),
text, font=style.font, fill=style.color)
```
For a specific mode, use `self.styles_for("recent")`, or set
`STYLE_MODE = "recent"` on the class and keep calling `self.styles`.
### The one rule that matters
**Pass your shipped values as the `classic_*` arguments.** The resolver returns
them verbatim unless the user actually changed something, which is what keeps an
untouched install rendering byte-identically. It can tell the difference because
a value only counts as user-forced when it *differs from the schema default* —
the save path writes the full default object into `config.json` on every save,
so "present in config" proves nothing.
Never compare against the default yourself; that rule lives in exactly one place.
### Stateless readers
For helpers handed a config dict rather than a plugin instance:
```python
from src.element_style import (element_color, element_visible,
element_align, element_scale, layout_offset)
colour = element_color(config, "score_text", (255, 255, 255), mode)
shown = element_visible(config, "records", True, mode)
dy = layout_offset(config, "score", "y_offset", 0, mode)
```
## Sports scoreboards
`SportsCoreSharedMixin` wires most of this up already. Two things to know:
* **Name your draws.** `_draw_text_with_outline(..., element="score_text")`
resolves the colour by name *and* honours the visibility toggle. Without it
the colour has to be guessed from the identity of the font object, which
cannot tell two elements apart when they share a face — the case every
bitmap font is in.
* **Modes are free.** Live/upcoming/recent are separate instances, so setting
`SKIN_MODE` on each is enough; no call site passes a mode.
## Adopting an existing hand-written block
If your schema already spells out `font` / `font_size` / `text_color` per
element longhand, **you do not need to change anything**. The core recognises
that shape and upgrades it in place: the style editor, the real font picker
(including uploaded fonts) and per-mode overrides all appear on a core update.
Add `x-style-modes` if you want the mode tabs.
## Fonts, and why size is sometimes locked
32 of the 35 shipped fonts are fixed-strike BDF bitmaps: they render at exactly
one pixel size and ignore `font_size`. The picker knows which, and the editor
locks the size field to the native size and labels it `fixed`. A font too tall
for the `max` you declared is not offered at all.
Uploaded fonts (Fonts tab) land in `assets/fonts/` and appear in the picker
automatically.
## Checklist
1. Declare `x-style-elements` (and `x-style-modes` if you have modes).
2. Read through `self.styles`, passing your shipped values as `classic_*`.
3. Honour `style.visible`, `style.offset` and `style.scale` where they apply.
4. Confirm an untouched config renders identically:
`python scripts/check_plugin.py --plugin <id>`.
5. Monorepo plugins: bump `manifest.json` and run `python update_registry.py`.
See also: [docs/PLUGIN_CONFIGURATION_GUIDE.md](PLUGIN_CONFIGURATION_GUIDE.md),
[docs/FONT_MANAGER.md](FONT_MANAGER.md).
+2 -1
View File
@@ -127,7 +127,8 @@ git push origin v1.0.0
### REST API
The API is mounted at `/api/v3` (`web_interface/app.py:199`).
The API is mounted at `/api/v3` (the `api_v3` blueprint in
`web_interface/blueprints/api_v3/`, registered in `web_interface/app.py`).
```bash
# Install plugin from the registry
+3
View File
@@ -103,6 +103,7 @@ All plugins can be installed through the LEDMatrix web interface:
Or via API:
```bash
curl -X POST http://your-pi-ip:5000/api/v3/plugins/install \
-H "Content-Type: application/json" \
-d '{"plugin_id": "clock-simple"}'
```
@@ -153,6 +154,7 @@ Before submitting, ensure your plugin:
```bash
# Install via URL on your Pi
curl -X POST http://your-pi:5000/api/v3/plugins/install-from-url \
-H "Content-Type: application/json" \
-d '{"repo_url": "https://github.com/you/ledmatrix-your-plugin"}'
```
@@ -312,6 +314,7 @@ git push
# 2. Review using VERIFICATION.md checklist
# 3. Test installation:
curl -X POST http://pi:5000/api/v3/plugins/install-from-url \
-H "Content-Type: application/json" \
-d '{"repo_url": "https://github.com/contributor/plugin"}'
# 4. If approved, merge PR
+4 -5
View File
@@ -131,13 +131,13 @@ else:
**Via REST API:**
```bash
# Search by query
curl "http://your-pi-ip:5000/api/v3/plugins/store/search?q=hockey"
curl "http://your-pi-ip:5000/api/v3/plugins/store/list?query=hockey"
# Filter by category
curl "http://your-pi-ip:5000/api/v3/plugins/store/search?category=sports"
curl "http://your-pi-ip:5000/api/v3/plugins/store/list?category=sports"
# Filter by tags
curl "http://your-pi-ip:5000/api/v3/plugins/store/search?tags=nhl&tags=hockey"
curl "http://your-pi-ip:5000/api/v3/plugins/store/list?tags=nhl&tags=hockey"
```
**Via Python:**
@@ -351,8 +351,7 @@ All API endpoints return JSON with this structure:
| Method | Endpoint | Description |
|--------|----------|-------------|
| GET | `/api/v3/plugins/store/list` | List all plugins in store |
| GET | `/api/v3/plugins/store/search` | Search for plugins |
| GET | `/api/v3/plugins/store/list` | List plugins in store; `?query=`, `?category=`, `?tags=` search and filter |
| GET | `/api/v3/plugins/installed` | List installed plugins |
| POST | `/api/v3/plugins/install` | Install from registry |
| POST | `/api/v3/plugins/install-from-url` | Install from GitHub URL |
+4 -2
View File
@@ -45,6 +45,8 @@ Going deeper:
- [PLUGIN_CONFIG_QUICK_START.md](PLUGIN_CONFIG_QUICK_START.md) — minimal config you need
- [PLUGIN_CONFIGURATION_GUIDE.md](PLUGIN_CONFIGURATION_GUIDE.md) — schema design
- [PLUGIN_ELEMENT_STYLING.md](PLUGIN_ELEMENT_STYLING.md) — let users restyle,
move, hide and scale individual elements (per display mode, if you have them)
- [PLUGIN_CONFIGURATION_TABS.md](PLUGIN_CONFIGURATION_TABS.md) — multi-tab UI configs
- [PLUGIN_CONFIG_ARCHITECTURE.md](PLUGIN_CONFIG_ARCHITECTURE.md) — how the config system works
- [PLUGIN_CONFIG_CORE_PROPERTIES.md](PLUGIN_CONFIG_CORE_PROPERTIES.md) — properties every plugin honors
@@ -54,8 +56,8 @@ Going deeper:
- [ADVANCED_FEATURES.md](ADVANCED_FEATURES.md) — Vegas scroll, on-demand display,
cache management, background services, permissions
- [FONT_MANAGER.md](FONT_MANAGER.md) — font system
- [SKIN_SYSTEM.md](SKIN_SYSTEM.md) — skin architecture for sports scoreboards
- [CREATING_SKINS.md](CREATING_SKINS.md) — writing and validating a skin
- [SKIN_SYSTEM.md](SKIN_SYSTEM.md) — skin architecture for sports scoreboards (not supported yet: current scoreboards don't render skins)
- [CREATING_SKINS.md](CREATING_SKINS.md) — writing and validating a skin (same caveat)
## Reference
+882 -380
View File
File diff suppressed because it is too large Load Diff
+335
View File
@@ -0,0 +1,335 @@
# Scroll Performance
How scrolling is paced on this hardware, what was wrong with it, and how to
configure a plugin so its marquee is smooth.
Measured on a Raspberry Pi 4 driving a 2×128×64 chain (256×64 logical) at
`limit_refresh_rate_hz: 100`. Numbers below come from that panel.
| | before | after |
|---|---|---|
| scroll frame rate | 44–46 fps | **100 fps, locked** |
| frames ≥ 45 ms | 14–17% | none observed |
| dominant frame time | 20 ms | **10 ms** |
| disk cache write (~1 MB) | 14.8 ms | **5.4 ms** |
---
## The one rule that matters
**Motion is smooth when the strip advances a whole number of pixels per panel
refresh.**
Advancing one pixel per refresh on a 100 Hz panel gives 100 px/s. Slower crisp
speeds come from holding each frame for several refreshes -- 50 px/s is one
pixel every second refresh -- which is covered under *Choosing a speed* below.
A speed that lands on no such combination has to do one of two bad things:
- **blend** two adjacent columns to render a half-step — on pixel-font text
this alternates crisp and smeared frames and reads as shimmer, or as the
text jumping a pixel ahead of itself;
- **repeat** a frame — the strip stands still, then jumps, which reads as
judder.
Neither is tunable away. Pick a speed that divides evenly.
`src.common.scroll_config` solves this for you: `configure()` snaps a requested
speed to the nearest one the panel can actually show in whole pixels, and
`scripts/scroll_speeds.py` prints the full ladder for your hardware.
## Choosing a speed
The crisp speeds are not a fixed list -- they depend on how fast *your* panel
refreshes, which depends on its size, `pwm_bits`, `gpio_slowdown` and the Pi
model. A Pi Zero driving a long chain has a completely different set of good
speeds from a Pi 4 driving a short one.
```bash
# what can this panel do? (reads your configured refresh rate)
python3 scripts/scroll_speeds.py
# what does it ACTUALLY manage, rather than what is configured?
sudo systemctl stop ledmatrix
sudo python3 scripts/scroll_speeds.py --measure
sudo systemctl start ledmatrix
# highlight the closest option to the speed you want
python3 scripts/scroll_speeds.py --want 45
# try one on the panel
sudo systemctl stop ledmatrix
sudo python3 scripts/scroll_speeds.py --demo 50
sudo systemctl start ledmatrix
```
Sample ladder for a 100 Hz panel:
```
20.0 px/s (1px every 5 refreshes = 20.0 fps, slightly stepped)
25.0 px/s (1px every 4 refreshes = 25.0 fps, slightly stepped)
33.3 px/s (1px every 3 refreshes = 33.3 fps, smooth)
50.0 px/s (1px every 2 refreshes = 50.0 fps, smooth)
66.7 px/s (2px every 3 refreshes = 33.3 fps, smooth)
100.0 px/s (1px every 1 refresh = 100.0 fps, smooth)
```
### How a slow speed stays crisp
`SwapOnVSync(canvas, framerate_fraction)` holds each frame for N panel
refreshes. **The panel keeps refreshing at its full rate either way**, so
holding a frame costs nothing in flicker -- it only changes how often a *new*
image is presented. That is what allows 50 px/s to be one whole pixel every
second refresh, instead of half a pixel every refresh (which has no good
rendering, only a choice between blur and judder).
`scroll_config.configure()` snaps the requested speed to the nearest entry on
the ladder, sets the helper to advance that entry's whole-pixel step on every
presented frame (`ScrollHelper.set_pixels_per_frame`), and reports the hold
that speed needs. It does **not** apply the hold: the hold belongs to a scroll, not to a plugin's lifetime, and plugins
share one display manager -- one set at construction is reset the moment any
other plugin finishes scrolling. Apply it yourself when the scroll starts:
```python
settings = scroll_config.configure(
self.scroll_helper,
plugin_config=self.config,
global_config=self.global_config,
display_manager=self.display_manager, # supplies the panel refresh rate
)
# ...then, each time this plugin begins scrolling:
self.display_manager.set_scrolling_state(True, frame_hold=settings.frame_hold)
```
Passing `display_manager` only lets `configure` read the true refresh rate from
`display.hardware`, which a plugin config cannot see. Skipping the
`set_scrolling_state(True, frame_hold=...)` call is the mistake that matters.
The helper consults no clock in this mode -- it moves the fixed step once per
`update_scroll_position()` call, and `SwapOnVSync` is what paces those calls --
so without the hold the panel presents a new frame every refresh and the scroll
runs `frame_hold` times too fast: 50 px/s (hold 2) plays at 100 px/s.
Pass `snap_to_crisp=False` to keep an exact requested speed and accept the
artefacts. The helper then paces off elapsed time instead of stepping, and the
hold is 1.
The General tab's `target_fps` ("Scroll Frame Rate") plays no part in any of
this: frames are presented at the panel refresh divided by the hold.
Speeds slower than about 20 px/s are stepped no matter what, because a 1-pixel
advance at 20 fps is simply a coarse increment. That is the pixel pitch, not a
software limit; the only way to move in smaller increments is sub-pixel
blending, which this display does not tolerate (see above).
## Configuring a plugin
Use the shared resolver rather than reading config keys yourself:
```python
from src.common import scroll_config
settings = scroll_config.configure(
self.scroll_helper,
plugin_config=self.config,
global_config=self.global_config,
display_manager=self.display_manager,
plugin_logger=self.logger,
)
# each frame of a scroll (or at least when it starts):
self.display_manager.set_scrolling_state(True, frame_hold=settings.frame_hold)
```
It resolves every config shape in one place, applies the speed, and returns
what it did. Precedence, highest first:
1. `display_options.scroll_speed` + `scroll_delay` — **the recommended form**
2. `display.scroll_speed` + `scroll_delay` — deprecated shape
3. `scroll_speed` + `scroll_delay` at the root — legacy flat
4. `scroll_pixels_per_second` — deprecated
5. the global `display` block
6. the built-in default (100 px/s)
`scroll_speed` is pixels per frame and `scroll_delay` is the frame period in
seconds, so the pair means `scroll_speed / scroll_delay` px/s. The recommended
config for a 100 Hz panel:
```json
"display_options": { "scroll_speed": 1.0, "scroll_delay": 0.01 }
```
### Why the deprecated key ranks below the explicit pair
Because some plugins give `scroll_pixels_per_second` a **schema default**, and
schema defaults are merged into plugin config. Ranking it above the pair means
it is always present and always wins, so the documented settings become
unreachable. That is a real, shipped bug — see
[ledmatrix-plugins#408](https://github.com/ChuckBuilds/ledmatrix-plugins/issues/408).
The flip side: a `scroll_pixels_per_second` you add by hand is ignored whenever
the plugin's config also carries the pair, which it does whenever the pair has
a schema default. Set the speed through the pair instead.
The sports scoreboards (`src.common.sports_scroll`) are the exception to all of
the above: they read `scroll_settings.scroll_speed` per league as px/s directly,
and their `scroll_delay` is kept for compatibility but ignored for pacing.
If you are writing a plugin: do not give a deprecated key a schema default.
## What was actually wrong
Four independent faults, each found by measurement.
### 1. The frame loop slept on top of a wait it had already done
`display_controller.py` ran the high-FPS loop as `render → SwapOnVSync (blocks
to the panel's refresh) → time.sleep(0.008) → plugin ticks`. The sleep was
unconditional and added to a wait that had already happened. Render work
measured ~4 ms, so each iteration cost ~12 ms against a 10 ms refresh grid —
every swap missed a refresh and landed on the next one. The loop settled at
exactly 50 fps while asking for 125, with no headroom, so ~14% of frames
slipped a further refresh.
Now the loop sleeps only the remainder of the frame budget, with a 1 ms floor
so plugin threads still get the GIL.
### 2. `SwapOnVSync` held the GIL while blocking
The rgbmatrix binding declares it without `nogil` (unlike `SetPixel`, `Clear`
and `Fill` immediately above it in `cppinc.pxd`), so the render thread held the
GIL for the entire vsync wait — most of every frame. Background threads were
starved into long uninterruptible bursts; a 1.5 MB API response costs ~17 ms to
parse and ~18 ms to re-encode for the cache, and `json.raw_decode` cannot be
preempted mid-document. Those bursts are what the render loop then waited on.
Fixed by rebuilding the binding: `scripts/build_rgbmatrix_nogil.sh`.
### 3. Sub-pixel blending was wrong for this display
Enabling it made things worse, not better — see the rule at the top. It is off
by default and only Vegas mode opts in via `set_sub_pixel_scrolling(True)`.
### 4. Frame-based stepping raced the vsync clock
Frame-based mode gated motion on a wall clock at `1/scroll_delay` steps per
second. Plugins set `scroll_delay` to the frame period, which puts that
comparison exactly on its own threshold: a frame arriving a hair early moved
zero pixels and rendered an identical frame, which dirty-tracking skipped, so
it returned in ~2 ms and the beat repeated. No `scroll_delay` value tunes this
out — a shorter delay just trades stalled frames for periodic double-steps.
A crisp speed configured through `scroll_config` no longer consults a clock at
all. Once `SwapOnVSync` blocks until the panel has taken the frame, the frame
count is a truer clock than `time.time()`, so the helper advances a fixed whole
number of pixels per presented frame (`set_pixels_per_frame`) and the display
manager holds each frame for `frame_hold` refreshes. Every frame moves the eye
by the same amount.
The time-based path remains only for callers that set a speed directly or pass
`snap_to_crisp=False`. There, frame-based mode no longer steps either: it
advances by elapsed time at `scroll_speed / scroll_delay` px/s.
## Diagnosing a juddery scroller
**An average will lie to you.** A 2 ms duplicate frame and a 21 ms double-wait
mean exactly 10 ms, so a ticker stalling on half its frames still averages to a
healthy 100 fps. The stats line reports the tail for that reason — read the
percentiles, not the fps.
Every scroller emits one line every 5 seconds covering *every* frame in that
window, tagged with the plugin it came from:
```bash
journalctl -u ledmatrix --since "-10min" --no-pager | grep "Scroll frame stats"
```
```
[Plugin: news] Scroll frame stats - 100.0 fps over 501 frames | median 10.00ms
p95 10.11ms max 12.03ms min 7.98ms | stalls 0 (0.0%) skips 0 (0.0%)
```
Reading it, on a 100 Hz panel:
A healthy median is the refresh period times the scroll's frame hold: 10 ms
for a hold of 1 (100 px/s), **20 ms for 50 px/s** (hold 2), 30 ms for 33.3 px/s.
A 20 ms median on a 50 px/s scroll is the hold doing its job, not missed
refreshes. The `Scroll configured:` log line gives the hold (`1px every 2
refreshes`).
| you see | it means |
|---|---|
| median = refresh period × hold, p95 within ~0.5 ms of it | healthy — locked to the panel |
| p95 or max a whole refresh period or more above that median | frames missing refreshes — per-frame work is overrunning, or a background thread is holding the GIL |
| non-zero **skips**, or a median *below* the expected one | **duplicate frames** — the swap was skipped because the image did not change, so the frame never waited on vsync. The scroller is advancing less than one pixel per frame, which a crisp fixed-step scroll never does; look for a plugin pacing off time or not passing the hold. |
| non-zero **stalls** | frames past 1.5× the median, which is the measure of judder that survives averaging |
`stalls` and `skips` are both counted against that window's own median, so they
stay meaningful on a panel running at any refresh rate.
To rank every scroller at once rather than reading lines one at a time:
```bash
journalctl -u ledmatrix --since "-3h" --no-pager | grep "Scroll frame stats" \
| sed -E 's/.*- (\S+) - (\[Plugin: [^]]+\] )?Scroll.*median ([0-9.]+)ms p95 ([0-9.]+)ms.*/\1 \3 \4/' \
| awk '$2 < 1000 {n[$1]++; m[$1]+=$2; p[$1]+=$3} END {for (k in n)
printf "%-28s %5d windows median %6.2fms p95 %6.2fms\n", k, n[k], m[k]/n[k], p[k]/n[k]}' \
| sort -k7 -rn
```
The `$2 < 1000` guard drops windows whose median is a whole second or more.
Those are not frames. Until the idle-gap fix in `log_frame_rate()`, the first
frame of every scroll was timed against the end of the *previous* scroll, so
the gap between them was recorded as one enormous sample — it landed in the
`max` field of otherwise healthy windows and counted as one stall per scroll,
roughly 0.2% at 500 frames to a window, which is the same order as the real
stall rates it sat beside. Current builds emit none, but the guard costs
nothing and keeps the command honest against older journals.
A scroller whose p95 sits several times its median is the one to fix, and it is
usually the one doing the most per-frame work rather than the one configured
worst. Measured over 20 minutes with two scrollers set identically at 100 px/s,
the leaderboard held 10 ms flat while the odds ticker spent ~20% of its frames
on duplicates. Same settings, different render cost: odds does more per-frame
work, and more variably, so it is first to land a frame that advances less than
a whole pixel. Check the render path before the config.
Then confirm what the plugin actually loaded — config edits do not always reach
the running code:
```bash
journalctl -u ledmatrix --since "-5min" --no-pager | grep -iE "px/s|px/frame"
```
If a plugin logs its scroll config **twice** with different modes, the second
line is what is running.
## Rebuilding the binding
```bash
bash scripts/build_rgbmatrix_nogil.sh # build into a scratch dir
sudo bash scripts/build_rgbmatrix_nogil.sh --install
sudo bash scripts/build_rgbmatrix_nogil.sh --rollback
```
The build never touches the installed module. `--install` backs up the original
to `~/rgbmatrix-core.so.ORIGINAL` first, and rolls back automatically if the
service does not come back healthy. Requires `build-essential`; Cython is
installed into a cached venv under `~/.cache/ledmatrix-cython`.
Re-run it after upgrading `rpi-rgb-led-matrix`, since a library upgrade
replaces the patched binding.
## Faster JSON
`src/cache/disk_cache.py` uses `orjson` when it is importable and falls back to
the stdlib otherwise, so it is optional:
```bash
sudo pip3 install --break-system-packages orjson
```
Encoding is where it pays — about 7× on this hardware. Decoding gains far less
(~1.3× on large payloads) because the cost there is building Python objects,
not scanning text. That is also why moving parsing to a subprocess does not
help: `pickle.loads` of the same payload costs 8.1 ms against `json.loads` at
10.9 ms, so the work just moves rather than disappearing.
+48 -12
View File
@@ -1,5 +1,35 @@
# Skin System Architecture
## Status: not supported yet
**Skins don't render with the current scoreboard plugins.** The skin system
below works in isolation (it loads, validates and renders skins in
`scripts/validate_skin.py` and `test/test_skin_system.py`), but nothing on a
running display calls it:
- The only render hook is `SportsCore._render_game()` in
`src/base_classes/sports/core.py`.
- None of the current scoreboard plugins build on `src.base_classes`. The
official scoreboards in the `ledmatrix-plugins` monorepo, and the
third-party scoreboards in the plugin registry, carry their own sports and
rendering code (with the shared `src/common/sports_*` helpers) and never
reach `SportsCore._render_game()`.
So a skin can be dropped into `skins/` and named in a plugin's config, but the
scoreboard keeps drawing its built-in layout. Until a scoreboard adopts the
hook, core does not offer skins to users:
- The plugin config page shows no **Visual Skin** dropdown.
- The Plugin Store hides registry entries with `"type": "skin"` and refuses
to install one (`POST /api/v3/plugins/install` answers 400 with the reason).
- `GET /api/v3/skins` still lists what is in `skins/`, with
`"supported": false` and a `message`.
- A config that already contains `"skin"` / `"skin_options"` still loads,
validates and saves unchanged; the value is simply unused.
The rest of this document describes the design as built, for whoever wires a
scoreboard to it.
Skins are user-installable **visual overlays** for the sports scoreboards.
A skin replaces only the *look* of a scoreboard — the host plugin keeps doing
data fetching, scheduling, caching, dedup, live-priority takeover, and vegas
@@ -32,8 +62,10 @@ crashing) simply restores the built-in look.
## The render funnel
Every sports scoreboard (baseball, football, basketball, hockey — anything
built on the `src/base_classes/sports/` package, `core.py`) renders through exactly one seam:
A sports scoreboard built on the `src/base_classes/sports/` package
(`core.py`) renders through exactly one seam. No current scoreboard plugin is
built on it (see [Status](#status-not-supported-yet)), so for them this seam is
never reached:
`SportsCore._render_game(game, force_clear)`.
1. The mode class's `display()` (live, `SportsUpcoming`, `SportsRecent`)
@@ -135,22 +167,26 @@ Inside the plugin's own config section in `config/config.json`:
`"built-in"` means the stock renderer. Because this rides the plugin's config
section, it persists across plugin reinstalls like every other setting.
The web UI shows a **Visual Skin** dropdown for plugins that have matching
skins installed: `SchemaManager.inject_skin_selector` adds an enum to the
*served* schema only. Validation never sees the enum — so a config that
references an uninstalled skin stays valid (rendering just falls back), and
the currently-configured value is always kept selectable. `GET /api/v3/skins`
lists installed skins (optionally filtered by `?plugin_id=`).
`SchemaManager.inject_skin_selector` can add a **Visual Skin** enum to the
*served* schema for plugins with matching skins installed. While skins are
unsupported the plugin schema endpoint does not call it, so the dropdown is
not shown. Validation never sees the enum either way: the base schema allows
any `skin` value, so a config that references an uninstalled skin stays valid.
`GET /api/v3/skins` lists installed skins (optionally filtered by
`?plugin_id=`) and reports `"supported": false`.
## Distribution
- **Manual:** `git clone <skin repo> skins/<skin-id>` — that's the whole
install. No manifest bumps, no `update_registry.py`; skins are not monorepo
plugins.
- **Store:** registry entries with `"type": "skin"` install through the same
`plugins.json` pipeline; `PluginStoreManager` routes them to `skins/`,
validates `skin.json` (including the API major version) instead of
`manifest.json`, and never installs dependencies — skins are render-only
- **Store (disabled while unsupported):** registry entries with
`"type": "skin"` are hidden from the store list and refused on install.
`PluginStoreManager._install_skin_from_info` is kept: once
`SKINS_RENDER_SUPPORTED` in `src/skin_system/__init__.py` is true, such
entries install through the same `plugins.json` pipeline, land in `skins/`,
are validated against `skin.json` (including the API major version) instead
of `manifest.json`, and never install dependencies — skins are render-only
(stdlib + PIL + the provided context, no third-party packages in v1).
## Trust model
+30 -2
View File
@@ -80,11 +80,30 @@ src/base_classes/sports/
src/common/
sports_scroll.py SportsScrollDisplay / …Manager — scroll orchestration
(content building stays in the plugins)
sports_helpers.py clamp/logo/rotation free functions + SportsHelpersMixin
(3.5.0) — the helpers byte-identical in the
plugins' sports.py, and the _favorite_key seam
```
`from src.base_classes.sports import SportsCore` keeps working — the package
`__init__` re-exports, so the conversion is invisible to every existing importer.
### Converging on `src/common`
The scoreboards do not build on `src/base_classes`; their own `sports.py` copies
have moved past it. So shared code now lands in hardware-free `src/common`
modules taken from the plugin copies, each a **new module** rather than growth
on an existing one: a plugin that deletes a method copy and relies on an older
module having gained it fails at runtime with an `AttributeError`, while a
missing module fails at load, where the version checks can see it.
`sports_helpers.py` is the first (it holds `_favorite_key`, the override point
listed below, for later phases); its parity test compares every body against
the plugin copies when `LEDMATRIX_PLUGINS` points at a checkout, and
`test/test_common_is_hardware_free.py` keeps `src/common` free of
`rgbmatrix`, `src.base_classes` and `src.plugin_system`. How a plugin adopts a
module and drops its copy is documented in the plugins repo's
`docs/plugin-development/08-shared-sports-code.md`.
## Override points (the plugin-facing seam)
The base class calls these; plugins implement or override them. This table is the
@@ -199,6 +218,14 @@ The one behavior the upstreamed version adds is native
Part A threaded it through each copy by hand, and this makes that threading
legacy compatibility rather than the mechanism.
> **Superseded.** Once presentation became frame-locked (#545) the helper
> steps a fixed whole-pixel amount per presented frame and the panel presents
> at its own refresh, so honouring `target_fps` only turned it into a speed
> multiplier (60 doubled a scoreboard's speed, 200 halved it). `sports_scroll`
> no longer reads it: the crisp-speed ladder uses the panel refresh
> (`display_manager.refresh_hz`), and speed comes from
> `scroll_settings.scroll_speed` alone. See `docs/SCROLL_PERFORMANCE.md`.
## Phases
B0–B3 are merged and shipping in core 3.2.0. Everything that remains is
@@ -445,8 +472,9 @@ After adoption plus the frozen legacy copies it was 10,610; removing the dead
inline duplication (plugins #252) brought it to roughly 8,620. B6 would take it
to about 3,300 including the shared core module — some 2,400 fewer than before
this project started. **Until B6 runs, the adoption is net negative on disk**,
and its one delivered user-visible gain is that adopted plugins honour the
global `target_fps` instead of hardcoding ~100 FPS.
and its one delivered user-visible gain was that adopted plugins honoured the
global `target_fps` instead of hardcoding ~100 FPS (since withdrawn: see the
note under the B3 design above).
### Decision: stop adopting further modules until B6 closes
+3 -1
View File
@@ -158,9 +158,11 @@ This script will check:
- Python dependencies
- Configuration files
- File permissions
- Web interface availability
- Web interface availability (`ledmatrix-web` listening on port 5000)
- Network connectivity
Once it passes, the web interface is at `http://<pi-ip>:5000`.
## Quick Reference Commands
```bash
+2 -2
View File
@@ -273,7 +273,7 @@ sudo systemctl cat ledmatrix-web | grep User
1. **Verify file structure:**
```bash
ls -l web_interface/app.py
ls -l web_interface/blueprints/api_v3.py
ls -ld web_interface/blueprints/api_v3/
ls -l web_interface/blueprints/pages_v3.py
```
@@ -531,7 +531,7 @@ sudo systemctl cat ledmatrix-web | grep User
```bash
# Clear the cache with the helper script
sudo python3 scripts/utils/clear_cache.py
sudo python3 scripts/utils/clear_cache.py --clear-all
# Or remove files manually from the cache dir in use, e.g.:
sudo rm -rf /var/cache/ledmatrix/*
+50 -14
View File
@@ -96,6 +96,35 @@ Configure basic system settings:
- **Plugin System Settings** — including the `plugins_directory` (default
`plugin-repos/`) used by the plugin loader
- **Autostart** options for the display service
- **Automatic updates** — once a week, update LEDMatrix and every installed
plugin with a newer version. Off by default. Runs 2–5 AM local time when
possible, otherwise within a day of being due. The last result and next
check are shown under the toggle, and anything other than success raises a
banner on **Overview**.
- *Checks first:* the code update is skipped, with the reason shown, if
tracked files were edited locally (permission-only changes and edits under
`plugins/` or `plugin-repos/` don't count; the pull carries those across
and puts them back), the checkout has local commits, a
rebase/merge is in progress, the branch has no upstream, less than 300 MB
is free, or the newest version already failed once. A failed fetch is
retried the next day.
- *Health check and rollback:* after pulling, `ledmatrix-update-verify.service`
restarts the services and checks that the web interface responds and the
display (if it was running) stays up. If not — or if the new dependencies
failed to install — it resets to the previous commit, reinstalls the
previous dependencies and restarts again. A running display is restarted;
a stopped one stays stopped.
- *Plugins* update through the Plugin Store, which refuses versions that need
a newer LEDMatrix and restores the old copy when an install fails. When the
code changed, plugins wait until it passes its health check. Plugins are
not health-checked after updating.
- *Setup needs no SSH.* Turning the toggle on restarts the display service,
which installs the health check (`ledmatrix-update-verify.path` and
`.service`); the General tab shows when it is ready, or why setup failed.
Until then only plugins update. New installs set it up during
installation and can switch updates on with
`first_time_install.sh --enable-auto-update` (or `LEDMATRIX_AUTO_UPDATE=1`,
which `one-shot-install.sh` passes through), or at the installer's prompt.
Click **Save** to write changes to `config/config.json`. Most changes
require a display service restart from **Overview**.
@@ -105,18 +134,26 @@ require a display service restart from **Overview**.
Configure your LED matrix hardware:
**Matrix configuration:**
- `rows` — LED rows (typically 32 or 64)
- `cols` — LED columns (typically 64 or 96)
- `rows` — LED rows per panel (typically 32 or 64; even, at least 8 — the
current rgbmatrix library rejects more than 64)
- `cols` — LED columns per panel (typically 64 or 96; at least 16)
- `chain_length` — number of horizontally chained panels
- `parallel` — number of parallel chains
- `parallel` — number of parallel chains (1–3)
- `hardware_mapping` — `adafruit-hat-pwm` (with PWM jumper mod),
`adafruit-hat` (without), `regular`, or `regular-pi1`
- `gpio_slowdown` — must match your Pi model (3 for Pi 3, 4 for Pi 4, etc.)
- `brightness` — 0–100%
`adafruit-hat` (without), `regular` (direct wiring, and the Adafruit Triple
LED Matrix Bonnet), or `regular-pi1`
- `gpio_slowdown` — depends on your Pi and panel (roughly 1–3 on a Pi 3,
2–4 on a Pi 4); raise it if rows jump or the image is garbage
- `brightness` — 1–100%
- `pwm_bits`, `pwm_lsb_nanoseconds`, `pwm_dither_bits` — PWM tuning
- Dynamic Duration — global cap for plugins that extend their display
time based on content
The collapsed **Advanced Hardware & Display Options** section holds
multiplexing, panel type, row address type, scan mode, PWM tuning, the
refresh-rate cap and hardware pulsing. Every field has a help tip, and the
README's Display Settings section describes each one with its allowed range.
**Vegas Scroll Mode:** the Display tab also has a full Vegas Scroll
Mode section — enable toggle, scroll speed, separator width, dynamic
duration, and related settings — so you can configure Vegas mode
@@ -171,11 +208,11 @@ Manage fonts for your display:
- See font previews
- Check font sizes and styles
**Font Overrides:**
- Overrides are set per display *element* (e.g. a specific score or
clock text element), not per plugin
- Override default font choices for individual elements
- Preview font changes
**Font Preview:**
- Render sample text in any TTF/OTF font at a chosen size
Fonts used by a plugin are chosen in that plugin's own settings tab; the
Fonts tab has no per-element override editor.
**Delete Fonts:**
- Remove unused fonts
@@ -292,9 +329,8 @@ The web interface is built on a REST API that you can access programmatically:
http://your-pi-ip:5000/api/v3
```
The API blueprint mounts at `/api/v3` (see
`web_interface/app.py:199`). All endpoints below are relative to that
base.
The API blueprint (`web_interface/blueprints/api_v3/`) is registered at
`/api/v3` in `web_interface/app.py`.
**Common Endpoints:**
- `GET /api/v3/config/main` — Get main configuration
+164
View File
@@ -0,0 +1,164 @@
# Web UI Technical Audit — September 2026
Scope: `web_interface/` (Flask + HTMX + Alpine, `templates/v3/`, `static/v3/`).
Method: Impeccable design detector, code review (accessibility; performance,
theming, responsive), and a live pass on the running app at desktop and
375px mobile widths in light and dark themes. Severe claims were verified
against the live page; one was rejected (see below). No code was changed.
Product context: see [`PRODUCT.md`](../../PRODUCT.md).
## Health score: 8/20 (Poor)
| # | Dimension | Score | Key finding |
|---|-----------|-------|-------------|
| 1 | Accessibility | 2 | Focus rings never render; modals have no dialog semantics or focus management |
| 2 | Performance | 2 | ~1.2 MB JS (≈250 KB gzip) on every page; SSE streams and polling never pause |
| 3 | Responsive | 2 | Mobile drawer works; header title wraps to 3 lines; many ~24px touch targets |
| 4 | Theming | 1 | Tokens exist but hex dominates; dark mode is a class-by-class patch with leaks |
| 5 | Implementation integrity | 1 | Templates use Tailwind classes that don't exist in the stylesheet |
## Implementation integrity verdict: fail
There is no Tailwind build. `static/v3/app.css` is a hand-written subset of
Tailwind, while templates and JS are authored as if full Tailwind were loaded.
- **333 of 516 utility class names used have no CSS rule** (2,582 uses),
confirmed against the live stylesheets. Top offenders: `border` (250),
`text-gray-700` (183), `block` (148), `mr-1` (127), `hidden` (79),
`py-1`, `text-blue-600`, `px-2`, `text-center`, `hover:bg-blue-700`,
`divide-y`, `uppercase`, `font-mono`.
- **`.hidden` has never existed in `app.css`**, so the 145
`classList.add/remove/toggle('hidden')` calls across 26 files do nothing.
Visible proof: the header shows both the moon and sun theme icons.
- **15 classes are defined only under `[data-theme="dark"]`** (e.g.
`bg-blue-50`, `bg-red-50`, `bg-yellow-50`, `border-blue-200`,
`text-red-700`), so tinted notice boxes are unstyled in light mode.
- Visible damage: the Getting Started checklist (`partials/overview.html:96-117`)
renders native gray outset buttons in both themes; 32 visible buttons on
the Plugin Manager page render with default browser chrome; search icons
overlap inputs; error/diff modal backdrops are transparent
(`bg-gray-500 bg-opacity-75` undefined).
Other drift:
- Four competing `showNotification` definitions (`app.js:6`,
`app-shell.js:2464`, `widgets/notification.js:298`, `partials/fonts.html:249`)
— the winner depends on load order — plus 53 `alert()`/`confirm()` calls.
- At least four modal implementations (on-demand modal in `base.html:1032`,
Tailwind-UI style in `error_handler.js`/`diff_viewer.js`, `.jfm-*`/`.pfm-*`
with injected CSS, ad-hoc modals in `plugins_manager.js`).
- `.btn` mixed with ~90 hand-assembled color-utility button combos.
- SSE wiring duplicated in `app-shell.js:5-60` and `app.js:186-205`.
## Findings by severity
### P0
**Undefined utility layer** (above). Every show/hide toggle and every layout
built from missing classes silently fails; root cause of most visual bugs.
Fix: replace the hand-rolled subset with a real, purged Tailwind build
(with dark-mode variants), or at minimum define the high-use missing classes
(`hidden`, `border`, `block`, spacing/text utilities) and a button reset.
→ `/impeccable harden`
### P1
- **Focus rings never render.** `focus:ring-2` (`app.css:289-291`) references
`--tw-ring-inset` and `--tw-ring-offset-width`, which are never defined, so
the `box-shadow` is invalid. `focus:outline-none` (18 uses) does remove the
outline. `peer-focus:ring-4` has no rule, so the plugin enable toggle
(`plugins_manager.js:1586-1593`, `sr-only` checkbox) shows no focus.
WCAG 2.4.7.
- **Modals lack dialog semantics.** Only `json-file-manager.js` has
`role="dialog"`/`aria-modal`/Escape/initial focus; none trap focus or
return it. On-demand (`base.html:1032`), error (`error_handler.js:164-205`),
diff (`diff_viewer.js:211-214`), plugin file manager
(`plugin-file-manager.js:372-392, 576-599`), array-table editor
(`array-table.js:460-467`) have none of it. WCAG 2.1.2 / 4.1.2.
- **Unnamed controls.** ~16 icon-only buttons with no accessible name, e.g.
`base.html:1036`, `plugins.html:173,210`, `plugin_config.html:485,656`,
`number-input.js:102,129`, `text-input.js:120`, `date-picker.js:95`,
`time-picker.js:100`, `password-input.js:141`. ~115 of 245 form fields have
no label (hotspots: `plugin_config.html` 19, `starlark_config.html` 14,
`plugins.html` 11); confirmed live on the 11 store search/sort/filter inputs.
- **Captive WiFi setup page** (first-run surface): `#msg` status has no live
region, and `outline:none` is replaced by a 15%-alpha shadow
(`captive_setup.html:16,48`).
- **Background traffic never stops.** `/stream/stats` and `/stream/display`
SSE stay open on every tab (display frames push with no preview visible).
Tab timers keep running after leaving the tab (`display.html:1046` 5s,
`logs.html:222` 5s, `tools.html:987` 15s, `plugins_manager.js:1882` 15s,
update check `base.html:1196` 30min). Only `tools.html:999` checks
`visibilitychange`. Costly on a Pi Zero 2 W.
- **Page weight.** 47 script tags on every page, including all 33 widgets
(`base.html:984-1018`). `app-shell.js` (177 KB) is render-blocking
(`base.html:956`); `plugins_manager.js` is 277 KB.
- **Dark mode leaks.** `plugin-file-manager.js` (53 hex) and
`json-file-manager.js` (63 hex, e.g. `.jfm-modal-box{background:#fff}`)
inject CSS that ignores `data-theme`; `.form-control` hard-codes
`#fff`/`#111827` (`app.css:668-671`). `app.css` has 186 hex + 46 rgb
literals vs 94 `var(--…)` uses.
### P2
- Toasts: `role="alert"` inside an `aria-live="polite"` container
(`notification.js:78,156`) → double/assertive announcements; auto-dismiss 4s.
- `prefers-reduced-motion` covers 3 animations; ~106 `animate-pulse`/`fa-spin`
uses, `modalSlideIn`, and toast slides ignore it.
- Mobile: header title wraps to three lines and spills out of the header;
~33 plugin-card buttons are `text-xs px-2 py-1` (~24px); `#logs-container`
forced to 400/350px with `!important` (`app.css:399-411`).
- Logs panel contrast: `text-gray-400` on `bg-gray-900` ≈ 3.9:1
(`logs.html:75,86`).
- Three unnamed nested `<nav>` landmarks (`base.html:474,477,548`); no skip link.
- Plugin lists fully rebuilt via `innerHTML` on every filter change
(`plugins_manager.js:1554, 3784, 3993, 4389, 5905`); `logs.html:225` adds a
reflow-forcing resize listener on every partial load.
### P3
- No `loading="lazy"` on images; Font Awesome `font-display:block`.
- Unpinned `alpinejs@3.x.x` unpkg fallback (`base.html:241`).
- `widgets/example-color-picker.js` is not loaded anywhere.
- Detector: 3px accent stripe on `.plugin-card::before` (`app.css:721`).
## Verified and rejected
- **"Static assets are never cache-busted" (raised as P0): false.**
`app.py:491` (`@app.url_defaults add_static_version`) appends file mtime as
`?v=` to every static URL; the live HTML confirms it. The manual
`?v=20260307` on two script tags is merely redundant.
- Light-mode gray text contrast is mostly fine: `app.css` remaps grays darker
(4.8–10:1).
- Detector `gray-on-color` hits at `app.css:84,285` and `broken-image` hits
(JS-populated `src`) are not real rendered issues.
## What works
- Theme set before first paint, follows OS preference, `data-theme` + tokens.
- Mobile drawer: Escape closes it, focus returns to the hamburger, 44px rows.
- `aria-current="page"` on nav tabs; real `<header>` and `<main>`.
- Status colors always paired with text; nearly all images have alt text.
- `toggle-switch.js` uses `role="switch"`; vendor assets self-hosted.
## Open decisions (block the P0 fix approach)
Recorded as undecided in `PRODUCT.md`:
- Must the UI work fully offline (no CDN fallbacks)?
- Is a Node/CSS build step acceptable for contributors?
- Is WCAG 2.2 AA a formal requirement?
## Recommended order
1. **[P0] `/impeccable harden`** — fix the utility layer (real Tailwind build
or define missing classes + button reset).
2. **[P1] `/impeccable harden`** — focus-ring variables and `peer-focus`;
one shared accessible modal helper; name icon buttons and label fields;
live region on the captive page.
3. **[P1] `/impeccable optimize`** — pause SSE/timers on hidden tab or page;
load widget scripts on demand.
4. **[P1] `/impeccable colorize`** — move file-manager CSS and `.form-control`
onto theme tokens.
5. **[P2] `/impeccable adapt`** — header wrap, touch targets, log height.
6. **[P2] `/impeccable animate`** — reduced-motion alternatives.
7. **`/impeccable polish`** — final pass.
+2 -1
View File
@@ -10,7 +10,8 @@ plugin without breaking a size or screen you didn't think to test.
There is **no fixed set of supported panel sizes** — an RGB matrix build can be
any width/height and configuration (square, rectangle, 2×2, 4×4, 8×2, long
strips, tall stacks). Plugins are expected to read dimensions dynamically
(`self.display_manager.matrix.width/height`) and lay themselves out
(`self.display_manager.width/height` — not `matrix.width/height`, since
`matrix` is `None` when hardware init fails) and lay themselves out
accordingly, so a hardcoded coordinate or unscaled font shows up as a failure
here.
+87 -11
View File
@@ -281,11 +281,49 @@ Guidelines:
their own collapsible sections) and is safely ignored by older cores, so
adding it never breaks compatibility.
## Hiding Fields From the Form (`x-display: "hidden"`)
Add `"x-display": "hidden"` to a property that must stay in the schema but
should not appear as a control: a deprecated key kept so existing configs keep
validating, or an internal value such as an auto-generated row id.
```json
{
"properties": {
"radar_zoom": {
"type": "integer",
"default": 6,
"title": "Radar Zoom Level (deprecated)",
"x-display": "hidden"
}
}
}
```
What the core does with it:
- **Not rendered** at any depth: top-level fields, children of an object
section, and properties of array-of-object items (never a table column, even
if `x-columns` names it, and never in the row editor). A hidden field flagged
`x-advanced` is not listed or counted in Advanced Settings, and an object
whose children are all hidden draws no empty section. Hidden fields don't
show up in the settings search either, since it indexes the rendered form.
- **Stored value preserved on save.** Saving the form never changes a hidden
value. The unchecked-checkbox rule ignores a hidden boolean. Array rows carry
a hidden property's stored value through the form, so the value survives the
row being posted back; a new row gets no value (the plugin fills it in).
- **The API is unaffected.** A JSON save to `POST /api/v3/plugins/config` can
still set a hidden field.
Older cores ignore the flag and render the field as a normal control.
## Creating Custom Widgets
### Step 1: Create Widget File
Create a JavaScript file in your plugin directory. The recommended location is `widgets/[widget-name].js`:
Create a JavaScript file in your plugin's `widgets/` directory, named
`widgets/[widget-name].js`. The directory is not optional: it is the only
place the core will serve a widget from.
```javascript
// Ensure LEDMatrixWidgets registry is available
@@ -366,7 +404,29 @@ window.LEDMatrixWidgets.register('my-custom-widget', {
});
```
### Step 2: Reference Widget in Schema
### Step 2: Declare the Widget in `manifest.json`
The manifest is the allowlist. A widget is served only if the plugin declares
it, so shipping a file under `widgets/` does not by itself publish it:
```json
{
"widgets": [
{
"name": "my-custom-widget",
"script": "my-custom-widget.js",
"description": "What this widget is for"
}
]
}
```
`name` is what you use in `x-widget` and in the URL. `script` is optional and
defaults to `[name].js`; it must be a plain filename directly inside
`widgets/` (no paths). Both are validated against
`schema/manifest_schema.json`.
### Step 3: Reference Widget in Schema
In your plugin's `config_schema.json`:
@@ -383,15 +443,30 @@ In your plugin's `config_schema.json`:
}
```
### Step 3: Widget Loading
### Step 4: Widget Loading
The widget will be automatically loaded when the plugin configuration form is rendered. The system will:
The widget is loaded on demand when the plugin's configuration form renders a
field that references it. The system will:
1. Check if widget is registered in the core registry
2. If not found, attempt to load from plugin directory: `/static/plugin-widgets/[plugin-id]/[widget-name].js`
3. Render the widget using the registered `render` function
1. Check whether the widget is already registered in the core registry.
2. If not, fetch it from `/static/plugin-widgets/[plugin-id]/[widget-name].js`.
That route serves the declared `script` from your plugin's `widgets/`
directory, as `text/javascript`.
3. Render it by calling the `render` function your script registered.
**Note:** Currently, widgets are server-side rendered via Jinja2 templates. Custom widgets registered via the registry will have their handlers available, but full client-side rendering is a future enhancement.
The fetch uses a dynamic `import()`, so the file must parse as an ES module.
A plain IIFE does — modules are strict mode, so avoid sloppy-mode constructs.
**If the widget fails to load** (not declared, file missing, script throws, or
it never calls `register`), the field falls back to a plain text input holding
the current value. This is deliberate: a broken widget costs the user an
editor, not their configured value.
**Limitation:** the on-demand path applies to `string`-typed fields (the
default branch of the config-form renderer). Fields typed `object`, `array`,
`boolean`, `integer` or `number`, and fields whose `enum` is set, are
dispatched by the server-side template to its own built-in renderers, so a
plugin-supplied `x-widget` on one of those is ignored today.
## Widget API Reference
@@ -497,10 +572,11 @@ See [`web_interface/static/v3/js/widgets/example-color-picker.js`](../web_interf
- ✅ Plugin widget loading system implemented
**Current Behavior:**
- Widgets are server-side rendered via Jinja2 templates (existing behavior preserved)
- Core widgets are server-side rendered via Jinja2 templates (existing behavior preserved)
- Widget handlers are registered and available globally
- Custom widgets can be created and registered
- Full client-side rendering is a future enhancement
- Custom widgets can be created, declared in `manifest.json`, and are served
and rendered on demand for `string`-typed fields
- Plugin widgets on non-string fields are not dispatched yet (see Step 4)
**Backwards Compatibility:**
- All existing plugins using widgets continue to work without changes
+176 -19
View File
@@ -125,6 +125,72 @@ fi
# Get the home directory of the actual user
USER_HOME=$(eval echo ~$ACTUAL_USER)
# --- rpi-rgb-led-matrix checkout helpers -------------------------------------
# Run git as whoever owns the project directory. Run as root against a
# user-owned repo, git refuses it ("dubious ownership"), and anything it does
# create — such as .git/modules/<submodule> — ends up root-owned, locking the
# user out of their own checkout. A root-owned install keeps running as root.
_rgb_repo_owner() {
stat -c %U "$PROJECT_ROOT_DIR" 2>/dev/null || echo root
}
_git_as_repo_owner() {
local owner
owner=$(_rgb_repo_owner)
if [ "$(id -u)" = "0" ] && [ "$owner" != "root" ] && command -v sudo >/dev/null 2>&1; then
sudo -u "$owner" -H git "$@"
else
git "$@"
fi
}
# Earlier installer versions ran the submodule git commands as root, leaving
# root-owned files the repo owner (and so _git_as_repo_owner) cannot write.
_reclaim_rgb_checkout() {
local owner path
owner=$(_rgb_repo_owner)
if [ "$(id -u)" != "0" ] || [ "$owner" = "root" ]; then
return 0
fi
for path in "$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master" "$PROJECT_ROOT_DIR/.git/modules/rpi-rgb-led-matrix-master"; do
if [ -e "$path" ]; then
chown -R "$owner:" "$path" 2>/dev/null || true
fi
done
}
# `git pull` on the main repo never moves an existing submodule checkout, so a
# submodule bump (e.g. the ARMv6 build fix for Pi Zero/1) would never reach a
# device installed before it. Move the checkout forward to the pinned commit —
# but never backward or sideways: a user who ran `git submodule update --remote`
# is newer than the pin and is left alone. Never fatal.
_sync_rgb_submodule() {
local sub="$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master" pinned current
if [ ! -f "$PROJECT_ROOT_DIR/.gitmodules" ] || ! grep -q "rpi-rgb-led-matrix" "$PROJECT_ROOT_DIR/.gitmodules" \
|| [ ! -e "$sub/.git" ]; then
return 0
fi
if ! pinned=$(_git_as_repo_owner -C "$PROJECT_ROOT_DIR" rev-parse "HEAD:rpi-rgb-led-matrix-master" 2>/dev/null) \
|| [ -z "$pinned" ]; then
return 0
fi
current=$(_git_as_repo_owner -C "$sub" rev-parse HEAD 2>/dev/null) || current=""
if [ "$current" = "$pinned" ]; then
return 0
fi
if [ -n "$current" ] && _git_as_repo_owner -C "$sub" cat-file -e "${pinned}^{commit}" 2>/dev/null \
&& ! _git_as_repo_owner -C "$sub" merge-base --is-ancestor "$current" "$pinned" 2>/dev/null; then
echo "rpi-rgb-led-matrix-master is at ${current:0:7}, not behind the pinned ${pinned:0:7}; leaving it as is"
return 0
fi
echo "Updating rpi-rgb-led-matrix-master to the pinned commit ${pinned:0:7}..."
if ! _git_as_repo_owner -C "$PROJECT_ROOT_DIR" submodule update --init --recursive rpi-rgb-led-matrix-master; then
echo "⚠ Could not update rpi-rgb-led-matrix-master to the pinned commit; building the existing checkout"
fi
return 0
}
# --- end rpi-rgb-led-matrix checkout helpers ---------------------------------
# Determine the Project Root Directory (where this script is located)
PROJECT_ROOT_DIR=$(cd "$(dirname "$0")" && pwd)
@@ -154,6 +220,8 @@ SKIP_PERF=${LEDMATRIX_SKIP_PERF:-0}
SKIP_REBOOT_PROMPT=${LEDMATRIX_SKIP_REBOOT_PROMPT:-0}
SKIP_SWAP=${LEDMATRIX_SKIP_SWAP:-0}
BUILD_JOBS_OVERRIDE=${LEDMATRIX_BUILD_JOBS:-}
# Weekly automatic updates: 1 on, 0 off, empty = ask (interactive) or leave as is.
AUTO_UPDATE=${LEDMATRIX_AUTO_UPDATE:-}
usage() {
cat <<USAGE
@@ -168,12 +236,15 @@ Options:
--skip-swap Never add temporary swap for the C++ build
--build-jobs N Compile the C++ library with N parallel jobs
(default: scaled to available RAM)
--enable-auto-update Turn on weekly automatic updates (with health
check and automatic rollback)
--no-auto-update Leave weekly automatic updates off
-h, --help Show this help message and exit
Environment variables (same effect as flags):
LEDMATRIX_ASSUME_YES=1, RPI_RGB_FORCE_REBUILD=1, LEDMATRIX_SKIP_SOUND=1,
LEDMATRIX_SKIP_PERF=1, LEDMATRIX_SKIP_REBOOT_PROMPT=1,
LEDMATRIX_SKIP_SWAP=1, LEDMATRIX_BUILD_JOBS=N
LEDMATRIX_SKIP_SWAP=1, LEDMATRIX_BUILD_JOBS=N, LEDMATRIX_AUTO_UPDATE=1|0
Low-memory devices:
On a Pi with under 2GB of RAM the C++ build is limited to fewer parallel
@@ -191,6 +262,8 @@ while [ $# -gt 0 ]; do
--skip-perf) SKIP_PERF=1 ;;
--no-reboot-prompt) SKIP_REBOOT_PROMPT=1 ;;
--skip-swap) SKIP_SWAP=1 ;;
--enable-auto-update) AUTO_UPDATE=1 ;;
--no-auto-update) AUTO_UPDATE=0 ;;
--build-jobs)
shift
if [ $# -eq 0 ]; then echo "--build-jobs requires a number"; usage; exit 1; fi
@@ -797,6 +870,50 @@ else
echo "✓ Main config file already exists"
fi
# Weekly automatic updates (General tab -> Automatic Updates). Off unless asked
# for: --enable-auto-update / LEDMATRIX_AUTO_UPDATE=1, or "y" at the prompt when
# installing interactively. Only an explicit choice changes the setting, so
# re-running the installer with -y never switches it silently.
if [ -z "$AUTO_UPDATE" ] && [ "$ASSUME_YES" != "1" ] && [ -t 0 ]; then
read -p "Automatically check for and install LEDMatrix updates once a week, with automatic rollback if an update breaks something? (y/N): " -n 1 -r
echo
if [[ $REPLY =~ ^[Yy]$ ]]; then AUTO_UPDATE=1; else AUTO_UPDATE=0; fi
fi
if [ "$AUTO_UPDATE" = "1" ] || [ "$AUTO_UPDATE" = "0" ]; then
if python3 - "$PROJECT_ROOT_DIR/config/config.json" "$AUTO_UPDATE" <<'PY'
import json, os, sys, tempfile
path, enabled = sys.argv[1], sys.argv[2] == "1"
with open(path, encoding="utf-8") as f:
config = json.load(f)
if not isinstance(config.get("auto_update"), dict):
config["auto_update"] = {}
config["auto_update"]["enabled"] = enabled
# Written beside the original and swapped in whole: the display service's
# config watcher may be running and must never read a half-written file.
original = os.stat(path)
fd, tmp = tempfile.mkstemp(dir=os.path.dirname(os.path.abspath(path)), prefix=".config.")
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
json.dump(config, f, indent=4)
f.write("\n")
f.flush()
os.fsync(f.fileno())
os.chmod(tmp, original.st_mode & 0o777)
if hasattr(os, "chown"):
os.chown(tmp, original.st_uid, original.st_gid)
os.replace(tmp, path)
except BaseException:
if os.path.exists(tmp):
os.unlink(tmp)
raise
PY
then
if [ "$AUTO_UPDATE" = "1" ]; then echo "✓ Weekly automatic updates enabled"; else echo "✓ Weekly automatic updates off"; fi
else
echo "⚠ Could not set auto_update in config/config.json; turn it on from the General tab instead"
fi
fi
# Create config_secrets.json from template if missing
if [ ! -f "$PROJECT_ROOT_DIR/config/config_secrets.json" ]; then
if [ -f "$PROJECT_ROOT_DIR/config/config_secrets.template.json" ]; then
@@ -1038,8 +1155,9 @@ else
# so git clone doesn't fail with "destination path already exists".
_clone_rpi_rgb() {
rm -rf "$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master"
git clone https://github.com/hzeller/rpi-rgb-led-matrix.git rpi-rgb-led-matrix-master
_git_as_repo_owner clone https://github.com/hzeller/rpi-rgb-led-matrix.git rpi-rgb-led-matrix-master
}
_reclaim_rgb_checkout
if [ ! -d "$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master" ]; then
echo "rpi-rgb-led-matrix-master not found. Initializing git submodule..."
cd "$PROJECT_ROOT_DIR"
@@ -1047,7 +1165,7 @@ else
# Try to initialize submodule if .gitmodules exists
if [ -f "$PROJECT_ROOT_DIR/.gitmodules" ] && grep -q "rpi-rgb-led-matrix" "$PROJECT_ROOT_DIR/.gitmodules"; then
echo "Initializing rpi-rgb-led-matrix submodule..."
if ! retry git submodule update --init --recursive rpi-rgb-led-matrix-master; then
if ! retry _git_as_repo_owner submodule update --init --recursive rpi-rgb-led-matrix-master; then
echo "⚠ Submodule init failed, cloning directly from GitHub..."
retry _clone_rpi_rgb
fi
@@ -1066,12 +1184,14 @@ else
cd "$PROJECT_ROOT_DIR"
rm -rf rpi-rgb-led-matrix-master
if [ -f "$PROJECT_ROOT_DIR/.gitmodules" ] && grep -q "rpi-rgb-led-matrix" "$PROJECT_ROOT_DIR/.gitmodules"; then
retry git submodule update --init --recursive rpi-rgb-led-matrix-master
retry _git_as_repo_owner submodule update --init --recursive rpi-rgb-led-matrix-master
else
retry _clone_rpi_rgb
fi
fi
_sync_rgb_submodule
# Add temporary swap on low-memory devices so the compiler survives.
CURRENT_STEP="Prepare the low-memory build environment"
if [ "$LOWMEM_AVAILABLE" = "1" ] && [ "$SKIP_SWAP" != "1" ]; then
@@ -1272,7 +1392,7 @@ if [ -f "$PROJECT_ROOT_DIR/scripts/install/install_web_service.sh" ]; then
fi
fi
if [ ! -f "/etc/systemd/system/ledmatrix-web.service" ] || [ "$NEEDS_UPDATE" = true ]; then
if [ ! -f "/etc/systemd/system/ledmatrix-web.service" ] || [ ! -f "/etc/systemd/system/ledmatrix-update-verify.path" ] || [ "$NEEDS_UPDATE" = true ]; then
bash "$PROJECT_ROOT_DIR/scripts/install/install_web_service.sh"
# Ensure systemd sees any new/changed unit files
systemctl daemon-reload || true
@@ -1288,7 +1408,7 @@ echo ""
CURRENT_STEP="Harden systemd unit file permissions"
echo "Step 8.1: Setting systemd unit file permissions..."
echo "-----------------------------------------------"
for unit in "/etc/systemd/system/ledmatrix.service" "/etc/systemd/system/ledmatrix-web.service" "/etc/systemd/system/ledmatrix-wifi-monitor.service"; do
for unit in "/etc/systemd/system/ledmatrix.service" "/etc/systemd/system/ledmatrix-web.service" "/etc/systemd/system/ledmatrix-wifi-monitor.service" "/etc/systemd/system/ledmatrix-update-verify.service" "/etc/systemd/system/ledmatrix-update-verify.path"; do
if [ -f "$unit" ]; then
chown root:root "$unit" || true
chmod 644 "$unit" || true
@@ -1384,6 +1504,9 @@ echo "------------------------------------------------"
# Create sudoers configuration for the web interface
echo "Creating sudoers configuration..."
SUDOERS_FILE="/etc/sudoers.d/ledmatrix_web"
# A predictable name in a world-writable directory is a symlink target;
# root writes the rules here, so let mktemp pick the name.
SUDOERS_TMP=$(mktemp "${TMPDIR:-/tmp}/ledmatrix_web_sudoers.XXXXXX")
# Get command paths
PYTHON_PATH=$(which python3)
@@ -1394,7 +1517,7 @@ BASH_PATH=$(which bash)
JOURNALCTL_PATH=$(which journalctl 2>/dev/null || true)
# Create sudoers content
cat > /tmp/ledmatrix_web_sudoers << EOF
cat > "$SUDOERS_TMP" << EOF
# LED Matrix Web Interface passwordless sudo configuration
# This allows the web interface user to run specific commands without a password
@@ -1412,13 +1535,13 @@ $ACTUAL_USER ALL=(ALL) NOPASSWD: $SYSTEMCTL_PATH is-active ledmatrix.service
$ACTUAL_USER ALL=(ALL) NOPASSWD: $SYSTEMCTL_PATH start ledmatrix-web.service
$ACTUAL_USER ALL=(ALL) NOPASSWD: $SYSTEMCTL_PATH stop ledmatrix-web.service
$ACTUAL_USER ALL=(ALL) NOPASSWD: $SYSTEMCTL_PATH restart ledmatrix-web.service
$ACTUAL_USER ALL=(ALL) NOPASSWD: $PYTHON_PATH $PROJECT_ROOT_DIR/display_controller.py
$ACTUAL_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT_DIR/start_display.sh
$ACTUAL_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT_DIR/stop_display.sh
$ACTUAL_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT_DIR/scripts/fix_perms/safe_plugin_rm.sh *
# Install a requirements.txt as root via vetted helper, so packages are visible
# to root-run ledmatrix.service (not just the web interface's own user).
$ACTUAL_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT_DIR/scripts/fix_perms/safe_pip_install.sh *
EOF
if [ -n "$JOURNALCTL_PATH" ]; then
cat >> /tmp/ledmatrix_web_sudoers << EOF
cat >> "$SUDOERS_TMP" << EOF
# NOEXEC, because these rules end in a wildcard and journalctl starts a pager
# when its output is a terminal. From that pager (less) a "!sh" is a root
# shell -- the standard journalctl escalation. The web interface always passes
@@ -1432,17 +1555,38 @@ $ACTUAL_USER ALL=(ALL) NOPASSWD:NOEXEC: $JOURNALCTL_PATH -t ledmatrix *
EOF
fi
if [ -f "$SUDOERS_FILE" ] && cmp -s /tmp/ledmatrix_web_sudoers "$SUDOERS_FILE"; then
echo "Sudoers configuration already up to date"
rm /tmp/ledmatrix_web_sudoers
# Never install rules we have not parsed. A malformed drop-in in
# /etc/sudoers.d makes sudo refuse every command for every user, which on a
# headless Pi leaves no way in at all. If the rules do not parse, say so and
# keep whatever is already installed.
SUDOERS_VALID=1
if command -v visudo >/dev/null 2>&1; then
if ! visudo -c -f "$SUDOERS_TMP" >/dev/null 2>&1; then
SUDOERS_VALID=0
echo "⚠ The generated sudoers rules did not parse:" >&2
visudo -c -f "$SUDOERS_TMP" >&2 || true
echo "⚠ Leaving $SUDOERS_FILE unchanged. The web interface cannot control" >&2
echo " the display service until this is fixed." >&2
fi
else
echo "Installing/updating sudoers configuration..."
cp /tmp/ledmatrix_web_sudoers "$SUDOERS_FILE"
chmod 440 "$SUDOERS_FILE"
rm /tmp/ledmatrix_web_sudoers
echo "⚠ visudo not found; installing the sudoers rules unvalidated"
fi
echo "✓ Passwordless sudo access configured"
if [ "$SUDOERS_VALID" = "0" ]; then
rm -f "$SUDOERS_TMP"
elif [ -f "$SUDOERS_FILE" ] && cmp -s "$SUDOERS_TMP" "$SUDOERS_FILE"; then
echo "Sudoers configuration already up to date"
rm -f "$SUDOERS_TMP"
else
echo "Installing/updating sudoers configuration..."
cp "$SUDOERS_TMP" "$SUDOERS_FILE"
chmod 440 "$SUDOERS_FILE"
rm -f "$SUDOERS_TMP"
fi
if [ "$SUDOERS_VALID" = "1" ]; then
echo "✓ Passwordless sudo access configured"
fi
echo ""
CURRENT_STEP="Configure WiFi management permissions"
@@ -1623,6 +1767,19 @@ chmod 755 "$PROJECT_ROOT_DIR/scripts/install/install_service.sh" "$PROJECT_ROOT_
# Re-apply special permissions for config directory (lost during normalization)
chmod 2775 "$PROJECT_ROOT_DIR/config" || true
# Harden the sudo-granted helper scripts: root-owned, not writable by the web
# user (matches scripts/install/configure_web_sudo.sh). The sudoers rules in
# Step 10 run these as root, so a user-owned copy is a root shell for whoever
# can edit it. This must come after Step 11's project-wide chown to
# $ACTUAL_USER, which would otherwise hand them straight back.
for helper in safe_plugin_rm.sh safe_pip_install.sh; do
HELPER_PATH="$PROJECT_ROOT_DIR/scripts/fix_perms/$helper"
if [ -f "$HELPER_PATH" ]; then
chown root:root "$HELPER_PATH" || echo "⚠ Could not set ownership on $HELPER_PATH"
chmod 755 "$HELPER_PATH" || echo "⚠ Could not set permissions on $HELPER_PATH"
fi
done
echo "✓ Project file permissions normalized"
echo ""
+1
View File
@@ -0,0 +1 @@
bridge_config.json
+123
View File
@@ -0,0 +1,123 @@
# Home Assistant MQTT Bridge
Control the matrix from Home Assistant: force any plugin or mode on demand,
turn the display on and off, and set brightness — as real HA entities, not
hand-written `mqtt.publish` calls.
The bridge owns no display logic. It subscribes to one command topic and
turns each message into a call against the same `api_v3` routes the web UI
uses, so behaviour lives in one place. It talks to the API over HTTP only —
no filesystem access — so it can run on the Pi or anywhere that can reach
the web interface.
## What appears in Home Assistant
On connect the bridge publishes [MQTT Discovery](https://www.home-assistant.io/integrations/mqtt/#mqtt-discovery)
config, so the matrix shows up under **Settings → Devices & Services → MQTT**
with no YAML:
| Entity | Does |
|---|---|
| `select.ledmatrix_display_mode` | Every mode across enabled plugins. Choosing one force-displays it. |
| `button.ledmatrix_stop_display` | Back to normal rotation. |
| `switch.ledmatrix_power` | Starts/stops the display service. |
| `number.ledmatrix_brightness` | 0–100. |
State is read back from the API every 30 seconds, so the entities also track
changes made from the web UI or an on-demand window expiring on its own.
All four share an availability topic that is the bridge's MQTT last will:
if the bridge dies, HA greys the controls out rather than leaving them
looking live but inert.
## Raw commands
For anything the entities do not cover, publish JSON to the command topic
(`ledmatrix/command` by default):
```jsonc
// Force a mode. plugin_id is optional — the bridge fills it in from
// /api/v3/display/modes.
{"action": "display", "mode": "nfl_live"}
// duration is seconds; pinned holds this one mode instead of rotating
// through every mode the plugin owns. Pin Starlark apps, where each mode
// is an unrelated widget; leave a sports plugin unpinned so live/recent/
// upcoming still cycle.
{"action": "display", "plugin_id": "starlark-apps", "mode": "aquarium",
"duration": 300, "pinned": true}
{"action": "stop_display"}
{"action": "power", "state": "on"}
{"action": "brightness", "value": 75}
// Re-read the mode list and re-publish discovery, after installing a plugin
{"action": "refresh"}
```
Every command publishes its outcome to `<command_topic>/status`, and current
state to `<command_topic>/state`.
## Requirements
- A LEDMatrix install with its web interface reachable (default `http://localhost:5000`)
- An MQTT broker that Home Assistant is also connected to
- Python 3 with `paho-mqtt` 2.x and `requests`
## Install
```bash
sudo ./scripts/install/install_mqtt_bridge.sh
```
That copies `bridge_config.example.json` to `bridge_config.json` on first
run, installs the dependencies, and enables `ledmatrix-mqtt-bridge.service`.
Edit the config with your broker details and re-run it.
```json
{
"mqtt_host": "192.168.1.10",
"mqtt_port": 8883,
"mqtt_username": "ledmatrix",
"mqtt_password": null,
"mqtt_topic": "ledmatrix/command",
"mqtt_tls": true,
"ledmatrix_api_base": "http://localhost:5000"
}
```
**TLS is on by default.** Without it the broker password and every display
command cross the network in cleartext. If your broker only listens on plain
1883 — which the Mosquitto add-on does out of the box — set `"mqtt_tls": false`
and `"mqtt_port": 1883`. The bridge logs a warning at startup when a password
is configured without TLS.
`bridge_config.json` is gitignored. Any key can also be supplied through the
environment as `LEDMATRIX_MQTT_<KEY>` (`LEDMATRIX_MQTT_MQTT_PASSWORD`, say),
which keeps a broker password out of a file on disk — put it in a systemd
drop-in with `Environment=` or `EnvironmentFile=` instead.
Set `mqtt_tls: true` for a broker with TLS. `mqtt_tls_insecure` skips
certificate verification and exists only for a self-signed broker on a
trusted LAN; it logs a warning when used.
To run it in the foreground while setting things up:
```bash
python3 integrations/mqtt_bridge/ledmatrix_mqtt_bridge.py --config integrations/mqtt_bridge/bridge_config.json
```
## Notes
- Only one thing can be on-demand at a time — the same constraint the web UI has.
- Forcing a mode restarts the display service, so the panel blanks for a moment.
- The mode list comes from `/api/v3/display/modes`, which triggers plugin
discovery itself. Discovery is lazy and normally happens because somebody
opened the dashboard; without that endpoint a bridge that never does would
see an empty list.
- Brightness writes `display.hardware.brightness` by posting
`{"brightness": N}` to `/api/v3/config/main`. A JSON save changes only the
keys it sends, so the other display settings are left as they were. The
display service's config hot reload notices the change within a few seconds
and applies it without a restart (unless `LEDMATRIX_HOT_RELOAD=false`; while
a dim schedule is dimming the panel, the dim level wins until the dim period
ends).
@@ -0,0 +1,13 @@
{
"mqtt_host": "192.168.1.10",
"mqtt_port": 8883,
"mqtt_username": "ledmatrix",
"mqtt_password": null,
"mqtt_client_id": "ledmatrix-mqtt-bridge",
"mqtt_topic": "ledmatrix/command",
"mqtt_tls": true,
"ledmatrix_api_base": "http://localhost:5000",
"request_timeout": 15,
"on_demand_duration": null,
"log_level": "INFO"
}
@@ -0,0 +1,548 @@
#!/usr/bin/env python3
"""Control a LEDMatrix display from Home Assistant over MQTT.
The bridge owns no display logic. It subscribes to one command topic and
turns each message into a call against the same api_v3 routes the web UI
uses, so behaviour stays in one place and this stays a translation layer.
On connect it publishes Home Assistant MQTT Discovery config, so a matrix
appears in HA as real entities rather than something you drive with
`mqtt.publish` by hand:
select.ledmatrix_display_mode every mode across enabled plugins;
choosing one force-displays it
button.ledmatrix_stop_display back to normal rotation
switch.ledmatrix_power the display service, on or off
number.ledmatrix_brightness 0-100
Anything the entities do not cover is still reachable by publishing JSON
to the command topic:
{"action": "display", "mode": "nfl_live"}
{"action": "display", "plugin_id": "starlark-apps", "mode": "aquarium",
"duration": 300, "pinned": true}
{"action": "stop_display"}
{"action": "power", "state": "on" | "off"}
{"action": "brightness", "value": 75}
{"action": "refresh"} re-publish discovery after installing a plugin
Every command publishes its result to <command_topic>/status.
Run it with `python3 ledmatrix_mqtt_bridge.py [--config PATH]`, or install
ledmatrix-mqtt-bridge.service.
"""
from __future__ import annotations
import argparse
import json
import logging
import os
import signal
import sys
import threading
from typing import Any, Callable, Dict, List, Optional
import requests
logger = logging.getLogger("ledmatrix-mqtt-bridge")
DISCOVERY_PREFIX = "homeassistant"
DEVICE_ID = "ledmatrix"
DEVICE_INFO = {
"identifiers": [DEVICE_ID],
"name": "LEDMatrix",
"manufacturer": "ChuckBuilds",
"model": "LEDMatrix Display",
}
DEFAULTS = {
"mqtt_host": "localhost",
"mqtt_port": 1883,
"mqtt_username": None,
"mqtt_password": None, # nosec B105 - "no password configured", not a credential
"mqtt_client_id": "ledmatrix-mqtt-bridge",
"mqtt_topic": "ledmatrix/command",
"mqtt_tls": False,
"mqtt_tls_insecure": False,
"ledmatrix_api_base": "http://localhost:5000",
"request_timeout": 15,
"on_demand_duration": None,
"log_level": "INFO",
}
class ConfigError(Exception):
"""The bridge cannot start with the configuration it was given."""
def load_config(path: str) -> Dict[str, Any]:
"""Read bridge_config.json, overlaid on DEFAULTS.
Every value may also come from the environment as LEDMATRIX_MQTT_<KEY>,
which is how a password stays out of a file that has to be world-readable
for the service user.
"""
config = dict(DEFAULTS)
if os.path.isfile(path):
with open(path, encoding="utf-8") as handle:
try:
loaded = json.load(handle)
except json.JSONDecodeError as err:
raise ConfigError(f"{path} is not valid JSON: {err}") from err
if not isinstance(loaded, dict):
raise ConfigError(f"{path} must contain a JSON object")
config.update(loaded)
else:
logger.warning("No config file at %s - using defaults and environment", path)
for key in DEFAULTS:
env_value = os.environ.get(f"LEDMATRIX_MQTT_{key.upper()}")
if env_value is not None:
config[key] = env_value
for key in ("mqtt_port", "request_timeout"):
try:
config[key] = int(config[key])
except (TypeError, ValueError) as err:
raise ConfigError(f"{key} must be a whole number, got {config[key]!r}") from err
for key in ("mqtt_tls", "mqtt_tls_insecure"):
config[key] = str(config[key]).lower() in ("1", "true", "yes", "on")
if config.get("mqtt_password") == "REPLACE_WITH_YOUR_ACTUAL_MQTT_PASSWORD":
raise ConfigError(
"mqtt_password is still the example placeholder - set a real password, "
"or remove the key if your broker allows anonymous connections")
return config
class LEDMatrixClient:
"""The api_v3 calls the bridge needs, and nothing else.
Everything goes through the HTTP API rather than the filesystem, so the
bridge does not have to live on the Pi, does not need read access to
config.json, and cannot drift from the web UI's own behaviour.
"""
def __init__(self, api_base: str, timeout: int = 15,
session: Optional[requests.Session] = None):
self.api_base = api_base.rstrip("/")
self.timeout = timeout
self.session = session or requests.Session()
def _call(self, method: str, path: str, **kwargs) -> Dict[str, Any]:
url = f"{self.api_base}/api/v3{path}"
response = self.session.request(method, url, timeout=self.timeout, **kwargs)
try:
body = response.json()
except ValueError:
body = {}
if response.status_code >= 400 or body.get("status") == "error":
message = body.get("message") or f"HTTP {response.status_code}"
raise RuntimeError(f"{method} {path} failed: {message}")
return body.get("data", body)
def list_modes(self) -> List[Dict[str, Any]]:
"""Every display mode that can be force-displayed, newest discovery.
/display/modes triggers plugin discovery itself, which matters because
discovery is lazy: a bridge that never opens the dashboard would
otherwise see nothing at all.
"""
return self._call("GET", "/display/modes").get("modes", [])
def display_status(self) -> Dict[str, Any]:
return self._call("GET", "/display/on-demand/status")
def start_on_demand(self, mode: str, plugin_id: Optional[str] = None,
duration: Optional[int] = None, pinned: bool = False) -> Dict[str, Any]:
payload: Dict[str, Any] = {"mode": mode, "pinned": pinned}
if plugin_id:
# find_plugin_for_mode only sees modes declared in a static
# manifest, so a plugin whose modes are generated -- each installed
# Starlark app is one -- 404s when plugin_id is omitted. Sending it
# skips that lookup. /display/modes reports it for every mode.
payload["plugin_id"] = plugin_id
if duration:
payload["duration"] = int(duration)
return self._call("POST", "/display/on-demand/start", json=payload)
def stop_on_demand(self) -> Dict[str, Any]:
return self._call("POST", "/display/on-demand/stop", json={})
def set_power(self, on: bool) -> Dict[str, Any]:
action = "start_display" if on else "stop_display"
return self._call("POST", "/system/action", json={"action": action})
def get_brightness(self) -> Optional[int]:
config = self._call("GET", "/config/main")
value = config.get("display", {}).get("hardware", {}).get("brightness")
try:
return int(value)
except (TypeError, ValueError):
return None
def set_brightness(self, value: int) -> Dict[str, Any]:
return self._call("POST", "/config/main", json={"brightness": int(value)})
class CommandHandler:
"""Turns one decoded MQTT payload into one API call.
Kept free of MQTT so it can be tested against a fake client: the failure
modes worth pinning are all in here (an unknown mode, an out-of-range
brightness, a mode name that needs its plugin_id attached).
"""
def __init__(self, client: LEDMatrixClient, default_duration: Optional[int] = None):
self.client = client
self.default_duration = default_duration
self._modes_by_name: Dict[str, Dict[str, Any]] = {}
def refresh_modes(self) -> List[Dict[str, Any]]:
modes = self.client.list_modes()
self._modes_by_name = {m["mode"]: m for m in modes}
# Home Assistant's select shows labels, so accept them back as well --
# otherwise picking "Simple Clock" in a dashboard is not a mode name.
for entry in modes:
self._modes_by_name.setdefault(entry.get("name") or entry["mode"], entry)
return modes
@property
def known_modes(self) -> List[Dict[str, Any]]:
return list({id(v): v for v in self._modes_by_name.values()}.values())
def handle(self, payload: Dict[str, Any]) -> Dict[str, Any]:
action = payload.get("action")
handlers: Dict[str, Callable[[Dict[str, Any]], Dict[str, Any]]] = {
"display": self._display,
"stop_display": lambda _p: self._ok(self.client.stop_on_demand()),
"power": self._power,
"brightness": self._brightness,
"refresh": lambda _p: self._ok({"modes": len(self.refresh_modes())}),
}
handler = handlers.get(action)
if handler is None:
return self._error(f"Unknown action {action!r}; expected one of "
f"{', '.join(sorted(handlers))}")
try:
return handler(payload)
except (requests.RequestException, RuntimeError) as err:
logger.error("Command %s failed: %s", action, err)
return self._error(str(err))
def _display(self, payload: Dict[str, Any]) -> Dict[str, Any]:
mode = payload.get("mode")
plugin_id = payload.get("plugin_id")
if not mode and not plugin_id:
return self._error("display requires 'mode' or 'plugin_id'")
known = self._modes_by_name.get(mode) if mode else None
if known is None and mode and not plugin_id:
# One retry against a fresh listing: a plugin installed since the
# last refresh is the common reason a valid mode looks unknown.
self.refresh_modes()
known = self._modes_by_name.get(mode)
if known is not None:
mode = known["mode"]
plugin_id = plugin_id or known.get("plugin_id")
duration = payload.get("duration", self.default_duration)
pinned = bool(payload.get("pinned", False))
result = self.client.start_on_demand(
mode=mode, plugin_id=plugin_id, duration=duration, pinned=pinned)
return self._ok(result, mode=mode, plugin_id=plugin_id)
def _power(self, payload: Dict[str, Any]) -> Dict[str, Any]:
state = str(payload.get("state", "")).strip().lower()
if state not in ("on", "off"):
return self._error("power requires 'state' of 'on' or 'off'")
return self._ok(self.client.set_power(state == "on"), state=state)
def _brightness(self, payload: Dict[str, Any]) -> Dict[str, Any]:
raw = payload.get("value")
try:
value = int(float(raw))
except (TypeError, ValueError):
return self._error(f"brightness requires a number, got {raw!r}")
if not 0 <= value <= 100:
return self._error(f"brightness must be between 0 and 100, got {value}")
return self._ok(self.client.set_brightness(value), value=value)
@staticmethod
def _ok(result: Any, **extra) -> Dict[str, Any]:
return {"status": "success", "result": result, **extra}
@staticmethod
def _error(message: str) -> Dict[str, Any]:
return {"status": "error", "message": message}
def discovery_messages(command_topic: str, state_topic: str, availability_topic: str,
mode_labels: List[str]) -> List[Dict[str, Any]]:
"""The retained MQTT Discovery configs, as {topic, payload} pairs.
Pure, so the entity shapes can be asserted without a broker. Every entity
shares one availability topic, which is also the bridge's last will -- HA
then shows the matrix as unavailable when the bridge dies, instead of
leaving stale controls that silently do nothing.
"""
common = {
"device": DEVICE_INFO,
"availability_topic": availability_topic,
"payload_available": "online",
"payload_not_available": "offline",
}
return [
{
"topic": f"{DISCOVERY_PREFIX}/select/{DEVICE_ID}/display_mode/config",
"payload": {
**common,
"name": "Display Mode",
"unique_id": f"{DEVICE_ID}_display_mode",
"command_topic": command_topic,
"command_template": '{"action": "display", "mode": "{{ value }}"}',
"state_topic": state_topic,
"value_template": "{{ value_json.mode }}",
"options": mode_labels,
"icon": "mdi:view-dashboard",
},
},
{
"topic": f"{DISCOVERY_PREFIX}/button/{DEVICE_ID}/stop_display/config",
"payload": {
**common,
"name": "Stop Display",
"unique_id": f"{DEVICE_ID}_stop_display",
"command_topic": command_topic,
"payload_press": '{"action": "stop_display"}',
"icon": "mdi:stop",
},
},
{
"topic": f"{DISCOVERY_PREFIX}/switch/{DEVICE_ID}/power/config",
"payload": {
**common,
"name": "Power",
"unique_id": f"{DEVICE_ID}_power",
"command_topic": command_topic,
"payload_on": '{"action": "power", "state": "on"}',
"payload_off": '{"action": "power", "state": "off"}',
"state_topic": state_topic,
"value_template": "{{ 'ON' if value_json.power else 'OFF' }}",
"state_on": "ON",
"state_off": "OFF",
"icon": "mdi:power",
},
},
{
"topic": f"{DISCOVERY_PREFIX}/number/{DEVICE_ID}/brightness/config",
"payload": {
**common,
"name": "Brightness",
"unique_id": f"{DEVICE_ID}_brightness",
"command_topic": command_topic,
"command_template": '{"action": "brightness", "value": {{ value }}}',
"state_topic": state_topic,
"value_template": "{{ value_json.brightness }}",
"min": 0,
"max": 100,
"step": 1,
"icon": "mdi:brightness-6",
},
},
]
def warn_if_cleartext(config: Dict[str, Any]) -> bool:
"""Say so, once, when a broker password is going over an unencrypted link.
The shipped example has TLS on, so reaching here means somebody turned it
off deliberately -- which is legitimate (the Mosquitto add-on is plaintext
on 1883) but should not be silent when there is a password to lose. Returns
whether it warned, so the decision is testable without a broker.
"""
if config.get("mqtt_tls") or not config.get("mqtt_password"):
return False
logger.warning(
'mqtt_tls is off and a password is set: the broker password and every '
'command are sent unencrypted. Set "mqtt_tls": true (port 8883 on most '
'brokers) unless this is a trusted, isolated network.')
return True
def read_state(client: LEDMatrixClient) -> Dict[str, Any]:
"""The state every entity reads, so HA opens on real values.
Each field is fetched independently: a matrix with its display service
stopped still has a brightness worth showing, and one unreachable field
should not blank the rest.
"""
state: Dict[str, Any] = {"power": False, "mode": None, "brightness": None}
try:
status = client.display_status()
state["power"] = bool(status.get("service", {}).get("active"))
on_demand = status.get("state", {})
if on_demand.get("active"):
state["mode"] = on_demand.get("mode")
except (requests.RequestException, RuntimeError) as err:
logger.debug("Could not read display status: %s", err)
try:
state["brightness"] = client.get_brightness()
except (requests.RequestException, RuntimeError) as err:
logger.debug("Could not read brightness: %s", err)
return state
class Bridge:
"""MQTT wiring around CommandHandler."""
def __init__(self, config: Dict[str, Any]):
self.config = config
self.command_topic = config["mqtt_topic"]
self.status_topic = f"{self.command_topic}/status"
self.state_topic = f"{self.command_topic}/state"
self.availability_topic = f"{self.command_topic}/availability"
self.client = LEDMatrixClient(config["ledmatrix_api_base"], config["request_timeout"])
self.handler = CommandHandler(self.client, config.get("on_demand_duration"))
self._stop = threading.Event()
self._mqtt = None
# -- MQTT callbacks (paho-mqtt 2.x VERSION2 signatures) ------------------
def _on_connect(self, client, _userdata, _flags, reason_code, _properties=None):
if getattr(reason_code, "is_failure", reason_code != 0):
logger.error("MQTT connection refused: %s", reason_code)
return
logger.info("Connected to MQTT broker; subscribing to %s", self.command_topic)
client.subscribe(self.command_topic, qos=1)
client.publish(self.availability_topic, "online", qos=1, retain=True)
# Re-publish on every reconnect, not just the first connect: a broker
# restart drops retained discovery configs, and HA would otherwise be
# left with entities it can no longer describe.
self.publish_discovery()
self.publish_state()
def _on_message(self, _client, _userdata, message):
try:
payload = json.loads(message.payload.decode("utf-8"))
except (UnicodeDecodeError, json.JSONDecodeError) as err:
logger.warning("Ignoring unparseable message on %s: %s", message.topic, err)
self._publish(self.status_topic, {"status": "error", "message": f"bad payload: {err}"})
return
if not isinstance(payload, dict):
self._publish(self.status_topic,
{"status": "error", "message": "payload must be a JSON object"})
return
logger.info("Command: %s", payload)
result = self.handler.handle(payload)
self._publish(self.status_topic, result)
# The API applies changes asynchronously (the controller polls its
# mailbox), so read state back rather than assuming the command took.
self.publish_state()
# -- publishing ---------------------------------------------------------
def _publish(self, topic: str, payload: Any, retain: bool = False) -> None:
if self._mqtt is None:
return
body = payload if isinstance(payload, str) else json.dumps(payload)
self._mqtt.publish(topic, body, qos=1, retain=retain)
def publish_discovery(self) -> None:
try:
modes = self.handler.refresh_modes()
except (requests.RequestException, RuntimeError) as err:
logger.error("Could not list display modes: %s", err)
modes = self.handler.known_modes
labels = sorted({m.get("name") or m["mode"] for m in modes})
for message in discovery_messages(self.command_topic, self.state_topic,
self.availability_topic, labels):
self._publish(message["topic"], message["payload"], retain=True)
logger.info("Published discovery for %d display mode(s)", len(labels))
def publish_state(self) -> None:
self._publish(self.state_topic, read_state(self.client), retain=True)
# -- lifecycle ----------------------------------------------------------
def run(self) -> int:
try:
import paho.mqtt.client as mqtt
except ImportError:
logger.error("paho-mqtt is not installed: pip install -r requirements.txt")
return 1
# VERSION2 is the current callback API. The compatibility note in
# CLAUDE.md is about code written against the v1 signatures; this file
# is written against v2 and requires paho-mqtt >= 2.0.
self._mqtt = mqtt.Client(
mqtt.CallbackAPIVersion.VERSION2,
client_id=self.config["mqtt_client_id"])
if self.config.get("mqtt_username"):
self._mqtt.username_pw_set(self.config["mqtt_username"],
self.config.get("mqtt_password"))
if self.config.get("mqtt_tls"):
self._mqtt.tls_set()
if self.config.get("mqtt_tls_insecure"):
logger.warning("TLS certificate verification is disabled (mqtt_tls_insecure)")
self._mqtt.tls_insecure_set(True)
else:
warn_if_cleartext(self.config)
self._mqtt.will_set(self.availability_topic, "offline", qos=1, retain=True)
self._mqtt.on_connect = self._on_connect
self._mqtt.on_message = self._on_message
logger.info("Connecting to %s:%s", self.config["mqtt_host"], self.config["mqtt_port"])
try:
self._mqtt.connect(self.config["mqtt_host"], self.config["mqtt_port"], keepalive=60)
except OSError as err:
logger.error("Could not reach the MQTT broker: %s", err)
return 1
self._mqtt.loop_start()
try:
while not self._stop.wait(30):
# HA is told the truth about state that changed outside the
# bridge -- somebody using the web UI, or an on-demand window
# expiring on its own.
self.publish_state()
finally:
self._publish(self.availability_topic, "offline", retain=True)
self._mqtt.loop_stop()
self._mqtt.disconnect()
return 0
def stop(self, *_args) -> None:
logger.info("Shutting down")
self._stop.set()
def main(argv: Optional[List[str]] = None) -> int:
parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
parser.add_argument(
"--config",
default=os.path.join(os.path.dirname(os.path.abspath(__file__)), "bridge_config.json"),
help="Path to bridge_config.json (default: alongside this script)")
args = parser.parse_args(argv)
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s - %(levelname)s - %(name)s - %(message)s")
try:
config = load_config(args.config)
except ConfigError as err:
logger.error("%s", err)
return 1
logging.getLogger().setLevel(str(config.get("log_level", "INFO")).upper())
bridge = Bridge(config)
signal.signal(signal.SIGTERM, bridge.stop)
signal.signal(signal.SIGINT, bridge.stop)
return bridge.run()
if __name__ == "__main__":
sys.exit(main())
@@ -0,0 +1,5 @@
# Floors are security floors, not API floors. requests < 2.33.0 carries
# CVE-2024-35195, CVE-2024-47081 and CVE-2026-25645; matches the pin in the
# project's own requirements.txt.
paho-mqtt>=2.0.0,<3.0.0
requests>=2.33.0,<3.0.0
+84 -26
View File
@@ -194,6 +194,15 @@ class StarlarkAppsPlugin(BasePlugin):
Each installed app becomes a dynamic display mode.
"""
#: Starlark apps are animations: a .webp render carries per-frame delays
#: and _display_frame advances at most one frame per call. The controller
#: reads this attribute to decide whether a mode needs its high-FPS loop;
#: without it display() was called once per rotation slot, so a multi-frame
#: app showed a single frame and never moved. static-image is force-run at
#: high FPS for the same reason (GIFs), but that plugin is special-cased by
#: name in the controller and this one has to declare it.
enable_scrolling = True
def __init__(self, plugin_id: str, config: Dict[str, Any],
display_manager, cache_manager, plugin_manager):
"""Initialize the Starlark Apps plugin."""
@@ -211,6 +220,8 @@ class StarlarkAppsPlugin(BasePlugin):
# App storage
self.apps_dir = self._get_apps_directory()
self.manifest_file = self.apps_dir / "manifest.json"
# A dedicated, never-replaced file to flock -- see _update_manifest_safe.
self.manifest_lock_file = self.apps_dir / "manifest.json.lock"
self.apps: Dict[str, StarlarkApp] = {}
# Display state
@@ -228,7 +239,7 @@ class StarlarkAppsPlugin(BasePlugin):
# Calculate optimal magnification based on display size
self.calculated_magnify = self._calculate_optimal_magnify()
if self.calculated_magnify > 1:
self.logger.info(f"Display size: {self.display_manager.matrix.width}x{self.display_manager.matrix.height}, "
self.logger.info(f"Display size: {self.display_manager.width}x{self.display_manager.height}, "
f"recommended magnify: {self.calculated_magnify}")
# Load installed apps
@@ -312,8 +323,8 @@ class StarlarkAppsPlugin(BasePlugin):
Recommended magnify value (1-8)
"""
try:
display_width = self.display_manager.matrix.width
display_height = self.display_manager.matrix.height
display_width = self.display_manager.width
display_height = self.display_manager.height
# Tronbyte native resolution
NATIVE_WIDTH = 64
@@ -351,8 +362,8 @@ class StarlarkAppsPlugin(BasePlugin):
Dictionary with recommendation details
"""
try:
display_width = self.display_manager.matrix.width
display_height = self.display_manager.matrix.height
display_width = self.display_manager.width
display_height = self.display_manager.height
NATIVE_WIDTH = 64
NATIVE_HEIGHT = 32
@@ -555,7 +566,8 @@ class StarlarkAppsPlugin(BasePlugin):
def _save_manifest(self, manifest: Dict[str, Any]) -> bool:
"""
Save apps manifest to file with file locking to prevent race conditions.
Acquires exclusive lock on manifest file before writing to prevent concurrent modifications.
Acquires exclusive lock on the manifest lock sidecar before writing to
prevent concurrent modifications.
"""
temp_file = None
lock_fd = None
@@ -563,9 +575,14 @@ class StarlarkAppsPlugin(BasePlugin):
# Create parent directory if needed
self.manifest_file.parent.mkdir(parents=True, exist_ok=True)
# Open manifest file for locking (create if doesn't exist, don't truncate)
# Use os.open with O_CREAT | O_RDWR to create if missing, but don't truncate
lock_fd = os.open(str(self.manifest_file), os.O_CREAT | os.O_RDWR, 0o644)
# Lock the sidecar file, not manifest_file itself: manifest_file is
# replaced by an atomic rename below, which swaps in a fresh inode
# a second locker's fresh os.open() would pick up unguarded. The
# sidecar is never written to or renamed over, so it always
# resolves to the same inode for every locker (see
# _update_manifest_safe and web_interface's _starlark_manifest_lock,
# which must lock this same file for the guarantee to hold).
lock_fd = os.open(str(self.manifest_lock_file), os.O_CREAT | os.O_RDWR, 0o644)
# Acquire exclusive lock on manifest file BEFORE creating temp file
# This serializes all writers and prevents concurrent races
@@ -621,8 +638,9 @@ class StarlarkAppsPlugin(BasePlugin):
# Create parent directory if needed
self.manifest_file.parent.mkdir(parents=True, exist_ok=True)
# Open manifest file for locking (create if doesn't exist, don't truncate)
lock_fd = os.open(str(self.manifest_file), os.O_CREAT | os.O_RDWR, 0o644)
# Lock the sidecar file, not manifest_file itself -- see the
# comment in _save_manifest for why.
lock_fd = os.open(str(self.manifest_lock_file), os.O_CREAT | os.O_RDWR, 0o644)
# Acquire exclusive lock for entire read-modify-write cycle
fcntl.flock(lock_fd, fcntl.LOCK_EX)
@@ -681,38 +699,62 @@ class StarlarkAppsPlugin(BasePlugin):
if app.is_enabled() and app.should_render(current_time):
self._render_app(app, force=False)
def display(self, force_clear: bool = False) -> None:
def display(self, display_mode: Optional[str] = None, force_clear: bool = False) -> bool:
"""
Display current Starlark app.
This method is called during the display rotation.
Displays frames from the currently active app.
`display_mode` names the app to show when it matches an installed
app_id. The controller passes the mode it is rotating to and inspects
this signature to decide whether to, so accepting it is what lets a
specific app be addressed -- including by an on-demand request pinned
to one app. Anything else (the plugin id itself, when the plugin
exposes no per-app modes) falls through to normal rotation.
Returns False when there is no app to show -- which is the state of
every install without Pixlet, and of a fresh one before any app is
added. The display controller only skips a mode on a boolean False
(it checks isinstance(result, bool)), so returning None held a black
panel for the full display_duration instead of rotating on.
"""
try:
if force_clear:
self.display_manager.clear()
# If no current app, try to select one
if not self.current_app:
if display_mode and display_mode in self.apps:
self.current_app = self.apps[display_mode]
elif force_clear or not self.current_app:
# Advance on entry to the mode. _select_next_app only ran when
# current_app was unset, so the first enabled app was picked
# once and then shown forever -- every other installed app was
# rendered on schedule and never displayed. force_clear is the
# controller's "we just switched to you" signal (it is reset
# immediately after this call), so one app gets each turn.
self._select_next_app()
if not self.current_app:
# No apps available
self.logger.debug("No Starlark apps to display")
return
return False
# Render app if needed
if not self.current_app.frames:
success = self._render_app(self.current_app, force=True)
if not success:
self.logger.error(f"Failed to render app: {self.current_app.app_id}")
return
return False
# Display current frame
self._display_frame()
# Display current frame. The result is propagated: a failed frame
# update is not a displayed frame, and returning True regardless
# told the controller the mode had rendered, so it held the dead
# frame for the whole display_duration instead of rotating on.
return self._display_frame()
except Exception as e:
self.logger.error(f"Error displaying Starlark app: {e}")
return False
def _select_next_app(self) -> None:
"""Select the next enabled app for display."""
@@ -763,15 +805,25 @@ class StarlarkAppsPlugin(BasePlugin):
magnify = self._get_effective_magnify()
self.logger.debug(f"Using magnify={magnify} for {app.app_id}")
# Filter out LEDMatrix-internal timing keys before passing to pixlet
INTERNAL_KEYS = {'render_interval', 'display_duration'}
# Optional native render size for an app whose own declared canvas
# differs from Pixlet's 64x32 default -- without this an app
# declaring a wider native canvas got half its own content
# clipped at render time, before magnify ever got a chance to
# scale anything.
render_width = app.config.get("render_width")
render_height = app.config.get("render_height")
# Filter out LEDMatrix-internal timing/sizing keys before passing to pixlet
INTERNAL_KEYS = {'render_interval', 'display_duration', 'render_width', 'render_height'}
pixlet_config = {k: v for k, v in app.config.items() if k not in INTERNAL_KEYS}
success, error = self.pixlet.render(
star_file=str(app.star_file),
output_path=str(app.cache_file),
config=pixlet_config,
magnify=magnify
magnify=magnify,
width=render_width,
height=render_height
)
if not success:
@@ -800,8 +852,8 @@ class StarlarkAppsPlugin(BasePlugin):
# Scale frames if needed
if self.config.get("scale_output", True):
width = self.display_manager.matrix.width
height = self.display_manager.matrix.height
width = self.display_manager.width
height = self.display_manager.height
# Get scaling method from config
scale_method_str = self.config.get("scale_method", "nearest")
@@ -835,10 +887,13 @@ class StarlarkAppsPlugin(BasePlugin):
self.logger.error(f"Error loading frames for {app.app_id}: {e}")
return False
def _display_frame(self) -> None:
"""Display the current frame of the current app."""
def _display_frame(self) -> bool:
"""Display the current frame of the current app.
:returns: whether a frame actually reached the display manager.
"""
if not self.current_app or not self.current_app.frames:
return
return False
try:
current_time = time.time()
@@ -856,8 +911,11 @@ class StarlarkAppsPlugin(BasePlugin):
)
self.current_app.last_frame_time = current_time
return True
except Exception as e:
self.logger.error(f"Error displaying frame: {e}")
return False
def install_app(self, app_id: str, star_file_path: str, metadata: Optional[Dict[str, Any]] = None, assets_dir: Optional[str] = None) -> bool:
"""
+124 -11
View File
@@ -218,7 +218,9 @@ class PixletRenderer:
star_file: str,
output_path: str,
config: Optional[Dict[str, Any]] = None,
magnify: int = 1
magnify: int = 1,
width: Optional[int] = None,
height: Optional[int] = None
) -> Tuple[bool, Optional[str]]:
"""
Render a .star file to WebP output.
@@ -228,6 +230,21 @@ class PixletRenderer:
output_path: Where to save WebP output
config: Configuration dictionary to pass to app
magnify: Magnification factor (default 1)
width: Optional native render width in pixels. Previously
there was no way to tell Pixlet to render at anything
other than its own default (64), relying entirely on
magnify to scale up afterward -- fine for apps designed
at that native size, but wrong for an app whose own
declared canvas size is genuinely different (confirmed
on real hardware, 2026-09-06, with an imported app
declaring width=128: rendering at the default 64 and
then magnifying silently clipped half the app's own
content before scaling ever happened, rather than
producing a correctly-sized image). Passed through as
Pixlet's own -w flag when provided; omitted (Pixlet's
default) otherwise, preserving existing behavior for
every other app.
height: Same as width, for Pixlet's -t flag.
Returns:
Tuple of (success: bool, error_message: Optional[str])
@@ -264,10 +281,18 @@ class PixletRenderer:
else:
value_str = str(value)
# Validate value doesn't contain dangerous shell metacharacters
# Block: backticks, $(), pipes, redirects, semicolons, ampersands, null bytes
# Allow: most printable chars including spaces, quotes, brackets, braces
if re.search(r'[`$|<>&;\x00]|\$\(', value_str):
# Validate value doesn't contain dangerous shell metacharacters.
# Kept as defence in depth only: cmd is a list and there is no
# shell=True below, so nothing here is ever interpreted by a
# shell. That made the list worth trimming rather than growing
# -- "|" is a normal character inside a config value, and apps
# do use it as a separator (a PennDOT sign id is
# "I-476 North|175659"). Blocking it dropped the whole key
# silently, and the app then rendered its own "not configured"
# screen with nothing to say why.
# Block: backticks, $(), redirects, semicolons, ampersands, null bytes
# Allow: most printable chars including spaces, quotes, brackets, braces, pipes
if re.search(r'[`$<>&;\x00]|\$\(', value_str):
logger.warning(f"Skipping config value with unsafe shell characters for key {key}: {value_str}")
continue
@@ -279,6 +304,10 @@ class PixletRenderer:
"-o", output_path,
"-m", str(magnify)
])
if width is not None:
cmd.extend(["-w", str(width)])
if height is not None:
cmd.extend(["-t", str(height)])
# Build sanitized command for logging (redact sensitive values)
sanitized_cmd = [self.pixlet_binary, "render", star_file]
@@ -286,6 +315,10 @@ class PixletRenderer:
config_keys = list(config.keys())
sanitized_cmd.append(f"[{len(config_keys)} config entries: {', '.join(config_keys)}]")
sanitized_cmd.extend(["-o", output_path, "-m", str(magnify)])
if width is not None:
sanitized_cmd.extend(["-w", str(width)])
if height is not None:
sanitized_cmd.extend(["-t", str(height)])
logger.debug(f"Executing Pixlet: {' '.join(sanitized_cmd)}")
# Execute rendering
@@ -299,13 +332,21 @@ class PixletRenderer:
)
if result.returncode == 0:
if os.path.isfile(output_path):
logger.debug(f"Successfully rendered: {star_file} -> {output_path}")
return True, None
else:
if not os.path.isfile(output_path):
error = "Rendering succeeded but output file not found"
logger.error(error)
return False, error
# Pixlet exits 0 and writes a 0-byte file when the app renders
# nothing -- an app whose config leaves it with no content to
# show does exactly that. Treating existence alone as success
# handed the caller a file with no frames in it, which reads
# downstream as a working app that draws a black panel.
if os.path.getsize(output_path) == 0:
error = "Rendering produced an empty (0-byte) file - the app rendered no content"
logger.error(error)
return False, error
logger.debug(f"Successfully rendered: {star_file} -> {output_path}")
return True, None
else:
error = f"Pixlet failed (exit {result.returncode}): {result.stderr}"
logger.error(error)
@@ -319,11 +360,76 @@ class PixletRenderer:
logger.exception("Rendering exception")
return False, "Rendering failed - see logs for details"
#: Schema extraction runs an app's own get_schema(), which may make a
#: network call. Short enough that a hung app does not stall an upload,
#: long enough for a real API round trip on a slow connection.
SCHEMA_TIMEOUT = 20
def extract_schema_via_pixlet(self, star_file: str) -> Optional[Dict[str, Any]]:
"""Ask Pixlet itself for the app's schema, or None if it cannot say.
`pixlet schema` executes get_schema() instead of reading it, which is
the only way to see options an app computes at runtime -- a dropdown
whose choices come from a live API call has no option list anywhere in
the source for the regex parser below to find, so that parser reports
an empty dropdown and the config form offers nothing to pick.
Pixlet's own field keys are remapped to the ones the rest of this
plugin and the config UI already use ("typeOf"/"desc"), so the two
extractors return the same shape and callers cannot tell them apart.
"""
if not self.pixlet_binary:
return None
try:
result = subprocess.run(
[self.pixlet_binary, "schema", star_file],
capture_output=True, text=True, timeout=self.SCHEMA_TIMEOUT,
cwd=self._get_safe_working_directory(star_file),
)
except subprocess.TimeoutExpired:
logger.warning(
"pixlet schema timed out after %ss for %s - get_schema() may be "
"making a slow network call", self.SCHEMA_TIMEOUT, star_file)
return None
except (subprocess.SubprocessError, OSError) as e:
logger.warning(f"Could not run pixlet schema for {star_file}: {e}")
return None
if result.returncode != 0:
# Not an error worth failing on: older Pixlet builds have no
# `schema` subcommand at all, and the source parser still works.
logger.debug(
"pixlet schema exited %d for %s: %s",
result.returncode, star_file, (result.stderr or '').strip()[:300])
return None
try:
schema = json.loads(result.stdout)
except (json.JSONDecodeError, ValueError) as e:
logger.warning(f"pixlet schema returned unparseable output for {star_file}: {e}")
return None
if not isinstance(schema, dict) or not isinstance(schema.get("schema"), list):
logger.warning(f"pixlet schema returned an unexpected shape for {star_file}")
return None
for field in schema["schema"]:
if not isinstance(field, dict):
continue
if "type" in field and "typeOf" not in field:
field["typeOf"] = field.pop("type")
if "description" in field and "desc" not in field:
field["desc"] = field.pop("description")
return schema
def extract_schema(self, star_file: str) -> Tuple[bool, Optional[Dict[str, Any]], Optional[str]]:
"""
Extract configuration schema from a .star file by parsing source code.
Extract configuration schema from a .star file.
Supports:
Prefers `pixlet schema`, which runs the app and therefore sees options
it computes at runtime. Falls back to parsing the source when Pixlet is
unavailable, too old to have the subcommand, or the app fails to run --
that parser handles:
- Static field definitions (location, text, toggle, dropdown, color, datetime)
- Variable-referenced dropdown options
- Graceful degradation for unsupported field types
@@ -337,6 +443,13 @@ class PixletRenderer:
if not os.path.isfile(star_file):
return False, None, f"Star file not found: {star_file}"
schema = self.extract_schema_via_pixlet(star_file)
if schema is not None:
logger.debug(
"Extracted schema with %d field(s) from %s via pixlet schema",
len(schema.get('schema', [])), star_file)
return True, schema, None
try:
# Read .star file
with open(star_file, 'r', encoding='utf-8') as f:
+112 -28
View File
@@ -5,6 +5,7 @@ Handles interaction with the Tronbyte apps repository on GitHub.
Fetches app listings, metadata, and downloads .star files.
"""
import json
import logging
import time
import requests
@@ -49,6 +50,13 @@ class TronbyteRepository:
self.base_url = "https://api.github.com"
self.raw_url = "https://raw.githubusercontent.com"
# Why the last GitHub API call failed, in words a user can act on.
# _make_request used to log the reason and return a bare None, so
# every caller up the stack knew only that "something" went wrong --
# which is how an exhausted rate limit reached the store page as an
# empty grid with no explanation.
self.last_error: Optional[str] = None
self.session = requests.Session()
if github_token:
self.session.headers.update({
@@ -70,29 +78,50 @@ class TronbyteRepository:
Returns:
JSON response or None on error
"""
self.last_error = None
try:
response = self.session.get(url, timeout=timeout)
if response.status_code == 403:
# Rate limit exceeded
logger.warning("[Tronbyte Repo] GitHub API rate limit exceeded")
if response.status_code in (403, 429):
# 403 is both "rate limited" and "forbidden"; the remaining
# counter is what tells them apart, and the difference matters
# to whoever reads the message -- one is fixed by waiting or
# adding a token, the other is not.
remaining = response.headers.get('X-RateLimit-Remaining')
if remaining == '0':
self.last_error = (
"GitHub API rate limit exceeded"
f" ({response.headers.get('X-RateLimit-Limit', '?')} requests/hour"
f"{'' if self.github_token else ', unauthenticated'})."
" Add a GitHub token in settings, or wait for the limit to reset."
)
else:
self.last_error = f"GitHub refused the request ({response.status_code})"
logger.warning(f"[Tronbyte Repo] {self.last_error}")
return None
elif response.status_code == 404:
self.last_error = "Not found on GitHub"
logger.warning(f"[Tronbyte Repo] Resource not found: {url}")
return None
elif response.status_code != 200:
self.last_error = f"GitHub API error {response.status_code}"
logger.error(f"[Tronbyte Repo] GitHub API error: {response.status_code}")
return None
return response.json()
except requests.Timeout:
self.last_error = "Timed out reaching GitHub"
logger.error(f"[Tronbyte Repo] Request timeout: {url}")
return None
except requests.RequestException as e:
self.last_error = f"Network error reaching GitHub: {e.__class__.__name__}"
logger.error(f"[Tronbyte Repo] Request error: {e}", exc_info=True)
return None
except (json.JSONDecodeError, ValueError) as e:
# Reachable whenever something on the path answers with HTML --
# a captive portal, a proxy error page, a DNS-hijacking router.
self.last_error = "GitHub returned a response that was not JSON"
logger.error(f"[Tronbyte Repo] JSON parse error for {url}: {e}", exc_info=True)
return None
@@ -125,6 +154,61 @@ class TronbyteRepository:
logger.error(f"[Tronbyte Repo] Network error fetching raw file {file_path}: {e}", exc_info=True)
return None
def _list_app_dirs_via_trees(self) -> Optional[List[Dict[str, Any]]]:
"""App directories via the git trees API, or None on failure.
The contents API caps a directory listing at 1000 entries and says
nothing about having truncated it, so the store showed the first 1000
apps of a repository that has more and looked complete while doing it.
The trees API caps far higher and sets `truncated` when it does, at
the cost of one extra call to resolve the `apps` tree.
"""
repo = f"{self.base_url}/repos/{self.REPO_OWNER}/{self.REPO_NAME}"
root = self._make_request(f"{repo}/git/trees/{self.DEFAULT_BRANCH}")
if not isinstance(root, dict):
return None
apps_sha = next(
(e.get('sha') for e in root.get('tree', []) or []
if e.get('path') == self.APPS_PATH and e.get('type') == 'tree'),
None)
if not apps_sha:
self.last_error = f"No '{self.APPS_PATH}' directory in the repository"
return None
tree = self._make_request(f"{repo}/git/trees/{apps_sha}")
if not isinstance(tree, dict):
return None
if tree.get('truncated'):
logger.warning(
"[Tronbyte Repo] GitHub truncated the app tree; the listing is incomplete")
return [
{'id': e['path'], 'path': f"{self.APPS_PATH}/{e['path']}", 'url': None}
for e in tree.get('tree', []) or []
if e.get('type') == 'tree' and e.get('path') and not e['path'].startswith('.')
]
def _list_app_dirs_via_contents(self) -> Optional[List[Dict[str, Any]]]:
"""App directories via the contents API. Capped at 1000 entries."""
url = f"{self.base_url}/repos/{self.REPO_OWNER}/{self.REPO_NAME}/contents/{self.APPS_PATH}"
data = self._make_request(url)
if data is None:
return None
if not isinstance(data, list):
self.last_error = "GitHub returned an unexpected listing format"
return None
return [
{'id': item.get('name'), 'path': item.get('path'), 'url': item.get('url')}
for item in data
if item.get('type') == 'dir' and item.get('name')
and not item['name'].startswith('.')
]
def list_apps(self) -> Tuple[bool, Optional[List[Dict[str, Any]]], Optional[str]]:
"""
List all available apps in the repository.
@@ -132,26 +216,17 @@ class TronbyteRepository:
Returns:
Tuple of (success, apps_list, error_message)
"""
url = f"{self.base_url}/repos/{self.REPO_OWNER}/{self.REPO_NAME}/contents/{self.APPS_PATH}"
data = self._make_request(url)
if data is None:
return False, None, "Failed to fetch repository contents"
if not isinstance(data, list):
return False, None, "Invalid response format"
# Filter directories (apps)
apps = []
for item in data:
if item.get('type') == 'dir':
app_id = item.get('name')
if app_id and not app_id.startswith('.'):
apps.append({
'id': app_id,
'path': item.get('path'),
'url': item.get('url')
})
apps = self._list_app_dirs_via_trees()
if apps is None:
# Fall back rather than fail: the contents API was what shipped,
# so a trees-only outage should not take the store down with it.
trees_error = self.last_error
logger.warning(
f"[Tronbyte Repo] Trees listing failed ({trees_error}); "
"falling back to the contents API")
apps = self._list_app_dirs_via_contents()
if apps is None:
return False, None, self.last_error or trees_error or "Failed to fetch repository contents"
logger.info(f"Found {len(apps)} apps in repository")
return True, apps, None
@@ -267,14 +342,22 @@ class TronbyteRepository:
'categories': _apps_cache['categories'],
'authors': _apps_cache['authors'],
'count': len(_apps_cache['data']),
'cached': True
'cached': True,
'error': None,
}
# Fetch directory listing (1 GitHub API call)
# Fetch directory listing (a small number of GitHub API calls)
success, app_dirs, error = self.list_apps()
if not success or not app_dirs:
logger.error(f"Failed to list apps for bulk fetch: {error}")
return {'apps': [], 'categories': [], 'authors': [], 'count': 0, 'cached': False}
# Returning an empty list here used to read downstream as "the
# repository has no apps", and the route reported that as a
# success -- so a rate limit, a DNS failure and an empty
# repository were all drawn as the same blank grid. Hand the
# reason back instead and let the caller surface it.
reason = error or "No apps found in the repository"
logger.error(f"Failed to list apps for bulk fetch: {reason}")
return {'apps': [], 'categories': [], 'authors': [],
'count': 0, 'cached': False, 'error': reason}
logger.info(f"Bulk-fetching manifests for {len(app_dirs)} apps...")
@@ -341,7 +424,8 @@ class TronbyteRepository:
'categories': categories,
'authors': authors,
'count': len(apps_with_metadata),
'cached': False
'cached': False,
'error': None,
}
def download_star_file(self, app_id: str, output_path: Path, filename: Optional[str] = None) -> Tuple[bool, Optional[str]]:
+5
View File
@@ -7,3 +7,8 @@ freezegun>=1.2,<2 # deterministic time for golden-image tests
psutil>=6.0.0,<8.0.0 # optional at runtime; installed for tests so the
# /system/status endpoint's real path is exercised
mypy>=1.5.0,<2.0.0 # static type checking (also pinned in .pre-commit-config.yaml)
PyYAML>=6.0.2,<7.0.0 # not a core dependency: test_starlark_pixlet_routes loads
# plugin-repos/starlark-apps/tronbyte_repository.py the way
# the blueprint does, and that plugin imports yaml. The
# plugin declares it in its own requirements.txt, which the
# store installs on a real rig but CI never does.
+20 -4
View File
@@ -43,10 +43,14 @@ packaging>=23.0,<27.0
# full feature set, or skip them for a minimal install.
# ───────────────────────────────────────────────────────────────────────
#
# scipy — sub-pixel interpolation in
# src/common/scroll_helper.py for smoother
# scrolling. Falls back to a simpler shift algorithm.
# pip install 'scipy>=1.10.0,<2.0.0'
# scipy — nothing, as of #570. It was listed for the sub-pixel
# interpolation path in src/common/scroll_helper.py, but
# get_visible_portion never consulted HAS_SCIPY, so that
# path was dead before it was deleted. The blend that
# replaced it is numpy-only. Do not install it expecting
# smoother scrolling: sub-pixel blending is off by default
# because it reads worse on a coarse panel, not because it
# is missing a library. See docs/SCROLL_PERFORMANCE.md.
#
# psutil — per-plugin resource monitoring in
# src/plugin_system/resource_monitor.py. The monitor
@@ -55,6 +59,18 @@ packaging>=23.0,<27.0
# range as a hard dependency — keep the two in sync.
# pip install 'psutil>=6.0.0,<7.0.0'
#
# orjson — faster JSON for the disk cache
# (src/cache/disk_cache.py). Encoding a ~1MB cache
# record drops from ~12ms to ~1.6ms on a Pi 4, which
# matters because that work holds the GIL and stalls
# the render thread mid-scroll. Falls back to the
# stdlib json when missing — see docs/SCROLL_PERFORMANCE.md.
# The 3.11.6 floor is CVE-2025-67221: orjson.dumps did not
# limit recursion on deeply nested documents, and the disk
# cache encodes payloads parsed straight from third-party
# APIs. 3.11.6 covers the Python range above.
# pip install 'orjson>=3.11.6,<4.0'
#
# Flask-Limiter — request rate limiting in web_interface/app.py
# (accidental-abuse protection, not security). The
# web interface starts without rate limiting when
+24
View File
@@ -257,6 +257,30 @@
},
"description": "Web UI action definitions"
},
"widgets": {
"type": "array",
"description": "Custom web-UI widgets this plugin provides. Each entry is served at /static/plugin-widgets/<plugin-id>/<name>.js from the plugin's widgets/ directory; only declared widgets are served. Reference one from config_schema.json with \"x-widget\": \"<name>\".",
"items": {
"type": "object",
"required": ["name"],
"properties": {
"name": {
"type": "string",
"pattern": "^[a-zA-Z0-9_-]{1,64}$",
"description": "Widget name, as used in x-widget and in the URL."
},
"script": {
"type": "string",
"pattern": "^[a-zA-Z0-9_-]{1,64}\\.js$",
"description": "Filename inside widgets/. Defaults to <name>.js."
},
"description": {
"type": "string",
"description": "Human-readable summary shown to plugin authors."
}
}
}
},
"ledmatrix_version": {
"type": "string",
"description": "Deprecated: Use compatible_versions instead. LEDMatrix version this plugin targets"
+159
View File
@@ -0,0 +1,159 @@
#!/usr/bin/env python3
"""Find blocking work reachable from a plugin's render path.
`display()` runs on the render thread. Anything slow reached from it stalls the
panel for its whole duration, and on a vsync-paced loop that is immediately
visible: a single 15ms call on a 100Hz panel drops a frame, and a network round
trip freezes the marquee outright.
This has bitten twice already. odds-ticker called `_has_live_games()` every
frame, whose slow path read the scoreboard cache from disk and parsed JSON per
enabled league -- one stalled frame every few minutes. soccer-scoreboard timed
out inside `update()` during a cache refresh. Both were found by staring at
frame-time histograms, which is a slow way to find a bug that is visible in the
source.
The audit walks the call graph from `display()` through same-class `self.*`
methods and reports anything that reaches a known-blocking API. It is a
heuristic, not a proof: it cannot see through indirection, and a hit is not
automatically a bug -- a call guarded by an interval check may be fine. It is a
list of places worth a human look.
python3 scripts/audit_render_path.py # all plugins
python3 scripts/audit_render_path.py --dir plugin-repos # a specific tree
python3 scripts/audit_render_path.py --plugin odds-ticker
"""
from __future__ import annotations
import argparse
import ast
import sys
from pathlib import Path
#: Calls that can block for longer than a frame. Matched on the attribute or
#: function name, so `requests.get`, `self.session.get` and a bare `get` on a
#: requests-ish object all register.
BLOCKING = {
"get": "network or cache read",
"post": "network",
"put": "network",
"request": "network",
"urlopen": "network",
"read": "I/O",
"open": "file I/O",
"load": "JSON/file parse",
"loads": "JSON parse",
"dump": "file write",
"dumps": "serialise",
"sleep": "sleep on the render thread",
"run": "subprocess",
"check_output": "subprocess",
"connect": "network",
"download_logo": "network",
"_fetch": "fetch",
}
#: Names that make a hit far more likely to matter.
HIGH_SIGNAL = ("requests", "urllib", "session", "cache_manager", "subprocess",
"socket", "http")
RENDER_ENTRY = "display"
class Analyzer:
def __init__(self, tree: ast.AST):
self.methods: dict[str, ast.FunctionDef] = {}
for node in ast.walk(tree):
if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)):
self.methods.setdefault(node.name, node)
def calls_in(self, fn: ast.AST):
"""(self-method names called, blocking hits) inside one function."""
self_calls, hits = set(), []
for node in ast.walk(fn):
if not isinstance(node, ast.Call):
continue
func = node.func
if isinstance(func, ast.Attribute):
name = func.attr
base = ast.unparse(func.value) if hasattr(ast, "unparse") else ""
if base == "self" and name in self.methods:
self_calls.add(name)
continue
if name in BLOCKING:
hits.append((name, base, BLOCKING[name], node.lineno))
elif isinstance(func, ast.Name) and func.id in BLOCKING:
hits.append((func.id, "", BLOCKING[func.id], node.lineno))
return self_calls, hits
def reachable_from(self, entry: str, max_depth: int = 3):
"""Blocking hits reachable from `entry`, with the path that reaches them."""
if entry not in self.methods:
return []
found, seen = [], set()
stack = [(entry, [entry], 0)]
while stack:
name, path, depth = stack.pop()
if name in seen or depth > max_depth:
continue
seen.add(name)
self_calls, hits = self.calls_in(self.methods[name])
for hit in hits:
found.append((path, hit))
for callee in sorted(self_calls):
stack.append((callee, path + [callee], depth + 1))
return found
def audit_file(path: Path):
try:
tree = ast.parse(path.read_text(encoding="utf-8"))
except (OSError, SyntaxError):
return []
analyzer = Analyzer(tree)
return analyzer.reachable_from(RENDER_ENTRY)
def main():
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
root = Path(__file__).resolve().parent.parent
ap.add_argument("--dir", default=str(root / "plugin-repos"))
ap.add_argument("--plugin", help="only this plugin directory")
ap.add_argument("--all-hits", action="store_true",
help="include low-signal hits (open/read/load on locals)")
args = ap.parse_args()
base = Path(args.dir)
if not base.is_dir():
sys.exit("not a directory: %s" % base)
plugins = [base / args.plugin] if args.plugin else sorted(
d for d in base.iterdir() if d.is_dir())
total = 0
for plugin in plugins:
rows = []
for src in sorted(plugin.glob("*.py")):
if src.name.startswith("test_"):
continue
for path, (name, s_base, why, lineno) in audit_file(src):
signal = any(h in s_base.lower() for h in HIGH_SIGNAL)
if not signal and not args.all_hits:
continue
rows.append((src.name, lineno, "->".join(path),
("%s.%s" % (s_base, name)) if s_base else name, why))
if rows:
total += len(rows)
print("\n%s" % plugin.name)
for fname, lineno, chain, call, why in sorted(rows):
print(" %s:%-5d %-34s via %s" % (fname, lineno, call + " (" + why + ")", chain))
print("\n%d blocking call(s) reachable from display() across %d plugin(s)"
% (total, len(plugins)))
print("Heuristic: a hit guarded by an interval check may be fine. Look, do "
"not assume.")
if __name__ == "__main__":
main()
+330
View File
@@ -0,0 +1,330 @@
#!/usr/bin/env bash
#
# Rebuild the rgbmatrix Python binding so it releases the GIL.
#
# WHY THIS EXISTS
# ---------------
# The upstream binding declares FrameCanvas::SwapOnVSync WITHOUT `nogil`
# (cppinc.pxd), unlike SetPixel/Clear/Fill on the lines just above it.
# SwapOnVSync blocks until the panel's next vertical sync -- up to a full
# refresh period on every frame -- so the render thread was holding the GIL
# for most of every frame. Background threads (API fetches, JSON parsing,
# image decode) were starved into long uninterruptible bursts, which in turn
# made the render loop miss refreshes.
#
# Measured on a Pi 4 driving a 2x128x64 chain at limit_refresh_rate_hz=100:
#
# before ~44 fps average, 14-17% of frames 41-53ms
# after 100 fps, median 10.00ms, p95 10.05ms, 0% stalls
#
# The per-pixel blit (SetPixelsPillow) can also release the GIL and walk the
# Pillow buffer row-major, but that is OFF by default and you almost certainly
# want to leave it that way. Row-major changes what a partially-written frame
# looks like: column-major tearing shows as a vertical seam, row-major tearing
# shows as a horizontal split between the panel's upper and lower halves. On a
# 1/32 scan panel that reads as a one-pixel "fold" across the middle of every
# panel -- reported on hardware, and it went away when the blit was reverted.
# Enable with RGB_PATCH_BLIT=1 only if you have measured that you need it;
# essentially all of the gain above comes from the SwapOnVSync change alone.
#
# SAFETY
# ------
# Builds into a scratch directory; touches the installed module only in the
# --install step, and backs up the original first. Roll back at any time with:
#
# sudo bash scripts/build_rgbmatrix_nogil.sh --rollback
#
# USAGE
# bash scripts/build_rgbmatrix_nogil.sh # build only
# sudo bash scripts/build_rgbmatrix_nogil.sh --install
# sudo bash scripts/build_rgbmatrix_nogil.sh --rollback
#
set -uo pipefail
# Resolve the invoking user's home, not root's. --install runs under sudo,
# where $HOME is /root, so every default path below pointed somewhere the
# build had never written and the install died with "no built module found".
if [ -n "${SUDO_USER:-}" ]; then
OWNER_HOME="$(getent passwd "$SUDO_USER" | cut -d: -f6)"
fi
OWNER_HOME="${OWNER_HOME:-$HOME}"
SRC_TREE="${RGB_SRC_TREE:-$OWNER_HOME/LEDMatrix/rpi-rgb-led-matrix-master}"
BUILD_DIR="${RGB_BUILD_DIR:-$OWNER_HOME/rgbmatrix-nogil-build}"
VENV="${RGB_CYTHON_VENV:-$OWNER_HOME/.cache/ledmatrix-cython}"
BACKUP="${RGB_BACKUP:-$OWNER_HOME/rgbmatrix-core.so.ORIGINAL}"
PATCH_BLIT="${RGB_PATCH_BLIT:-0}"
die() { echo "FATAL: $*" >&2; exit 1; }
# This script runs under `set -uo pipefail` -- no -e -- so an unchecked
# systemctl failure is silently ignored. That matters most for `stop`: leaving
# the old service running means cp overwrites a module the running process has
# mapped, the following `start` succeeds as a no-op, and the health check sees
# an active unit and reports SUCCESS for a binding that was never loaded.
# A machine with no ledmatrix.service at all is a normal build host, so that
# case is skipped rather than treated as a failure.
service_present() { systemctl cat ledmatrix.service >/dev/null 2>&1; }
service_do() {
local verb="$1"
if ! service_present; then
echo " (no ledmatrix.service installed - skipping $verb)"
return 0
fi
systemctl "$verb" ledmatrix || die "systemctl $verb ledmatrix failed"
}
py_site() {
python3 -c 'import rgbmatrix, os; print(os.path.dirname(rgbmatrix.__file__))' 2>/dev/null
}
# The extension filename the interpreter that builds -- and then loads -- this
# module actually uses, e.g. core.cpython-313-aarch64-linux-gnu.so. The build
# venv is made with --system-site-packages from python3, so the two agree;
# falling back keeps --install working when the venv has been cleaned up.
abi_name() {
local py="$VENV/bin/python"
[ -x "$py" ] || py=python3
"$py" -c \
'import sysconfig; print("core" + sysconfig.get_config_var("EXT_SUFFIX"))' \
2>/dev/null
}
# Exactly the current interpreter's artifact, never merely the first one that
# sorts. Staging copies $SRC_TREE wholesale, so a core.cpython-*.so left in the
# source tree by an earlier build comes along for the ride; build_ext --inplace
# only ever overwrites the current ABI's name, and a glob piped to `head -1`
# sorts cpython-311 ahead of cpython-313. That installed a stale, unpatched
# module as core.so while the GIL check below -- which reads the freshly
# generated core.cpp, not the .so -- still reported success.
abi_so() {
local name path
name="$(abi_name)" || return 1
[ -n "$name" ] || return 1
path="$BUILD_DIR/bindings/python/rgbmatrix/$name"
[ -f "$path" ] || return 1
printf '%s\n' "$path"
}
do_rollback() {
local dst; dst="$(py_site)"
[ -n "$dst" ] || die "could not locate the installed rgbmatrix package"
[ -f "$BACKUP" ] || die "no backup at $BACKUP"
service_do stop
cp -a "$BACKUP" "$dst/core.so" || die "restore failed"
find "$dst" -name __pycache__ -type d -exec rm -rf {} + 2>/dev/null
service_do start
echo "rolled back to the original core.so"
exit 0
}
do_install() {
local so dst
so="$(abi_so)"; [ -n "$so" ] || die "no built module found - run the build first"
dst="$(py_site)"; [ -n "$dst" ] || die "could not locate the installed rgbmatrix package"
if [ ! -f "$BACKUP" ]; then
cp -a "$dst/core.so" "$BACKUP" || die "could not back up the original"
echo "backed up original core.so -> $BACKUP"
else
echo "backup already present at $BACKUP (keeping the true original)"
fi
service_do stop
cp "$so" "$dst/core.so" || die "install failed"
find "$dst" -name __pycache__ -type d -exec rm -rf {} + 2>/dev/null
service_do start
echo "waiting 25s for the display to come back..."
sleep 25
local healthy=1
systemctl is-active --quiet ledmatrix || healthy=0
if journalctl -u ledmatrix --since "40 sec ago" --no-pager \
| grep -qiE "Traceback|ImportError|Segmentation fault|undefined symbol"; then
healthy=0
fi
if [ "$healthy" = "1" ]; then
echo "SUCCESS - running on the rebuilt binding"
else
echo "UNHEALTHY - rolling back"
cp -a "$BACKUP" "$dst/core.so" \
|| echo "ROLLBACK FAILED: could not restore $BACKUP -> $dst/core.so" >&2
if service_present && ! systemctl restart ledmatrix; then
echo "ROLLBACK FAILED: ledmatrix did not restart - the display is" \
"down; restore manually with 'sudo bash $0 --rollback'" >&2
fi
journalctl -u ledmatrix --since "90 sec ago" --no-pager | tail -25
exit 1
fi
exit 0
}
case "${1:-}" in
--rollback) do_rollback ;;
--install) do_install ;;
"" ) ;;
*) die "unknown option: $1" ;;
esac
# ---------------------------------------------------------------- build ----
[ -d "$SRC_TREE" ] || die "matrix source tree not found at $SRC_TREE (set RGB_SRC_TREE)"
command -v g++ >/dev/null || die "g++ not installed (apt install build-essential)"
echo "==> staging a scratch copy at $BUILD_DIR"
rm -rf "$BUILD_DIR"
cp -r "$SRC_TREE" "$BUILD_DIR" || die "copy failed"
# Drop any extension artifacts that came across from the source tree. Nothing
# downstream should be able to pick one up, and build_ext --inplace can decide
# a copied .so is already up to date and skip the compile entirely.
find "$BUILD_DIR/bindings/python/rgbmatrix" -maxdepth 1 \
-name 'core*.so' -delete 2>/dev/null
echo "==> patching the bindings to release the GIL"
python3 - "$BUILD_DIR" "$PATCH_BLIT" <<'PYEOF' || die "patch failed"
import io
import sys
base = sys.argv[1] + "/bindings/python/rgbmatrix/"
patch_blit = len(sys.argv) > 2 and sys.argv[2] == "1"
# --- declare SwapOnVSync as nogil ---------------------------------------
p = base + "cppinc.pxd"
s = io.open(p, encoding="utf-8").read()
OLD_DECL = " FrameCanvas *SwapOnVSync(FrameCanvas*, uint8_t)\n"
NEW_DECL = " FrameCanvas *SwapOnVSync(FrameCanvas*, uint8_t) nogil\n"
if OLD_DECL in s:
io.open(p, "w", encoding="utf-8", newline="\n").write(s.replace(OLD_DECL, NEW_DECL, 1))
print(" cppinc.pxd: SwapOnVSync declared nogil")
elif NEW_DECL in s:
print(" cppinc.pxd: already nogil")
else:
sys.exit("could not find the SwapOnVSync declaration")
# --- release the GIL across the vsync wait ------------------------------
p = base + "core.pyx"
s = io.open(p, encoding="utf-8").read()
OLD_SWAP = (
" def SwapOnVSync(self, FrameCanvas newFrame, uint8_t framerate_fraction = 1):\n"
" return __createFrameCanvas("
"self.__matrix.SwapOnVSync(newFrame.__canvas, framerate_fraction))\n"
)
NEW_SWAP = (
" def SwapOnVSync(self, FrameCanvas newFrame, uint8_t framerate_fraction = 1):\n"
" # Blocks until the panel's next vertical sync. Holding the GIL\n"
" # across that wait starves every other Python thread for most of\n"
" # each frame. Pointers are hoisted into C locals so the blocking\n"
" # call itself needs no Python state.\n"
" cdef cppinc.RGBMatrix* matrix = self.__matrix\n"
" cdef cppinc.FrameCanvas* frame = newFrame.__canvas\n"
" cdef uint8_t fraction = framerate_fraction\n"
" cdef cppinc.FrameCanvas* swapped\n"
" with nogil:\n"
" swapped = matrix.SwapOnVSync(frame, fraction)\n"
" return __createFrameCanvas(swapped)\n"
)
if OLD_SWAP in s:
s = s.replace(OLD_SWAP, NEW_SWAP, 1)
print(" core.pyx: SwapOnVSync releases the GIL")
elif "swapped = matrix.SwapOnVSync(frame, fraction)" in s:
print(" core.pyx: SwapOnVSync already patched")
else:
sys.exit("could not find the SwapOnVSync body")
# --- optional: release the GIL across the blit --------------------------
OLD_BLIT = (
" buffer = get_pillow_buffer(image_capsule)\n"
"\n"
" for col in range(max(0, -xstart), min(width, frame_width - xstart)):\n"
" for row in range(max(0, -ystart), min(height, frame_height - ystart)):\n"
" pixel = buffer[row][col]\n"
" r = (pixel ) & 0xFF\n"
" g = (pixel >> 8) & 0xFF\n"
" b = (pixel >> 16) & 0xFF\n"
" my_canvas.SetPixel(xstart+col, ystart+row, r, g, b)\n"
)
NEW_BLIT = (
" buffer = get_pillow_buffer(image_capsule)\n"
"\n"
" # Bounds hoisted so the blit needs no Python state and can run\n"
" # without the GIL: it touches only a C buffer and a C++ canvas.\n"
" # NOTE: row-major order makes a torn frame show as a horizontal\n"
" # split across the panel's halves. See the header before enabling.\n"
" cdef int col_start = max(0, -xstart)\n"
" cdef int col_end = min(width, frame_width - xstart)\n"
" cdef int row_start = max(0, -ystart)\n"
" cdef int row_end = min(height, frame_height - ystart)\n"
"\n"
" with nogil:\n"
" for row in range(row_start, row_end):\n"
" for col in range(col_start, col_end):\n"
" pixel = buffer[row][col]\n"
" r = (pixel ) & 0xFF\n"
" g = (pixel >> 8) & 0xFF\n"
" b = (pixel >> 16) & 0xFF\n"
" my_canvas.SetPixel(xstart+col, ystart+row, r, g, b)\n"
)
if patch_blit:
if OLD_BLIT in s:
s = s.replace(OLD_BLIT, NEW_BLIT, 1)
print(" core.pyx: pixel blit releases the GIL, row-major")
elif "for row in range(row_start, row_end):" in s:
print(" core.pyx: blit already patched")
else:
sys.exit("could not find the SetPixelsPillow loop")
else:
print(" core.pyx: blit left unpatched (RGB_PATCH_BLIT=1 to enable)")
io.open(p, "w", encoding="utf-8", newline="\n").write(s)
PYEOF
echo "==> building librgbmatrix.a (this takes a few minutes)"
nice -n 10 make -C "$BUILD_DIR/lib" -j2 >/dev/null 2>&1 \
|| die "library build failed - rerun 'make -C $BUILD_DIR/lib' to see why"
[ -f "$BUILD_DIR/lib/librgbmatrix.a" ] || die "librgbmatrix.a was not produced"
echo "==> preparing Cython"
[ -d "$VENV" ] || python3 -m venv --system-site-packages "$VENV" || die "venv failed"
"$VENV/bin/pip" install --quiet cython || die "cython install failed"
cat > "$BUILD_DIR/bindings/python/setup.py" <<'EOF'
from setuptools import setup, Extension
from Cython.Build import cythonize
core = Extension(
"rgbmatrix.core",
sources=["rgbmatrix/core.pyx", "rgbmatrix/shims/pillow.c"],
include_dirs=["../../include", "rgbmatrix/shims"],
extra_objects=["../../lib/librgbmatrix.a"],
language="c++",
extra_compile_args=["-O3", "-Wall", "-fno-exceptions", "-std=c++11"],
extra_link_args=["-lrt", "-lm", "-lpthread"],
)
setup(name="rgbmatrix",
ext_modules=cythonize([core], language_level="3str",
compiler_directives={"binding": False}))
EOF
echo "==> compiling the extension"
( cd "$BUILD_DIR/bindings/python" && "$VENV/bin/python" setup.py build_ext --inplace ) \
>/dev/null 2>&1 || die "extension build failed"
SO="$(abi_so)" || true
[ -n "$SO" ] || die "no .so produced - expected $(abi_name) in $BUILD_DIR/bindings/python/rgbmatrix"
# Verify the GIL really is released before anyone installs this.
EXPECTED=1; [ "$PATCH_BLIT" = "1" ] && EXPECTED=2
PAIRS=$(grep -c "PyEval_SaveThread\|Py_UNBLOCK_THREADS" \
"$BUILD_DIR/bindings/python/rgbmatrix/core.cpp")
[ "$PAIRS" -ge "$EXPECTED" ] \
|| die "generated C++ has $PAIRS GIL-release sites, expected >= $EXPECTED"
echo
echo "BUILT: $SO"
echo " ($PAIRS GIL-release site(s) in the generated C++)"
echo
echo "Install with: sudo bash $0 --install"
echo "Roll back with: sudo bash $0 --rollback"
+30 -2
View File
@@ -35,6 +35,31 @@ sys.path.insert(0, str(PROJECT_ROOT))
os.environ['EMULATOR'] = 'true'
def _make_output_encoding_safe() -> None:
"""Stop an unencodable character from killing the run.
This script's own report is ASCII, but it echoes text it does not control
-- plugin ids, mode names and exception messages -- and a Windows console
is cp1252, which cannot encode most of what a plugin might put there. The
default 'strict' error handler turns that into a UnicodeEncodeError from
inside `print`, so a rendering run that had already succeeded exited
non-zero with a traceback instead of printing its results.
'replace' degrades the offending character to '?' and keeps going; the
encoding itself is left alone so output still matches the terminal.
"""
for stream in (sys.stdout, sys.stderr):
try:
stream.reconfigure(errors='replace')
except (AttributeError, ValueError, OSError):
# Not a reconfigurable TextIOWrapper (redirected, wrapped by a
# test harness). Nothing to do -- this is best-effort hardening.
pass
_make_output_encoding_safe()
from src.logging_config import get_logger # noqa: E402
from src.plugin_system.testing.loading import ( # noqa: E402
build_full_config, find_plugin_dir, load_harness_spec, load_manifest,
@@ -53,6 +78,9 @@ logger = get_logger("[Check Plugin]")
DEFAULT_SEARCH_DIRS = [
str(PROJECT_ROOT / 'plugins'),
str(PROJECT_ROOT / 'plugin-repos'),
# The scoreboards live in the sibling ledmatrix-plugins checkout, not
# in this repo. Without this, --all silently skips every one of them.
str(PROJECT_ROOT.parent / 'ledmatrix-plugins' / 'plugins'),
]
@@ -178,7 +206,7 @@ def print_report(all_results: Dict[str, List[RenderResult]]) -> bool:
status = "PASS"
detail = ""
if r.golden_checked:
detail = " (golden ✓)"
detail = " (golden ok)"
if r.update_error is not None:
detail += f" (update warn: {r.update_error})"
if r.fill_checked and r.fill_ok is None and r.fill_extent:
@@ -196,7 +224,7 @@ def print_report(all_results: Dict[str, List[RenderResult]]) -> bool:
status, detail = "FAIL", f" overflow bbox={r.overflow}"
elif r.golden_ok is False:
status = "FAIL"
detail = f" golden drift: {r.golden_diff_pixels}px (max Δ={r.golden_max_delta})"
detail = f" golden drift: {r.golden_diff_pixels}px (max delta={r.golden_max_delta})"
elif r.fill_ok is False:
ex, ey = r.fill_extent or (0.0, 0.0)
status = "FAIL"
+7 -2
View File
@@ -54,8 +54,13 @@ def main():
config = config_manager.load_config()
print(" ✅ Config loaded")
autostart = config.get('web_display_autostart', False)
print(f" 🔧 web_display_autostart: {autostart}")
# Same rule ledmatrix-web.service applies: only an explicit false/off
# keeps the web interface down; a missing key means on.
sys.path.insert(0, str(project_root / 'scripts' / 'utils'))
from start_web_conditionally import autostart_enabled
raw = config.get('web_display_autostart', '(not set, defaults to on)')
state = 'starts' if autostart_enabled(config) else 'will NOT start'
print(f" 🔧 web_display_autostart: {raw} (web interface {state})")
except Exception as e:
print(f" ❌ Config check failed: {e}")
traceback.print_exc()
+10
View File
@@ -12,9 +12,19 @@ This directory contains scripts and utilities for development and testing.
### Plugin Development Setup
```bash
# Official plugin: clones ChuckBuilds/ledmatrix-plugins (once) and links
# its plugins/<plugin-name> into plugins/
./scripts/dev/dev_plugin_setup.sh link-github <plugin-name>
# Plugin with its own repository
./scripts/dev/dev_plugin_setup.sh link-github <plugin-name> <repo-url>
```
Set `plugin_system.plugins_directory` to `plugins` so the loader finds the
links. To use a fork or another clone location, copy
`dev_plugins.json.example` to `dev_plugins.json`. Details:
[docs/PLUGIN_DEVELOPMENT_GUIDE.md](../../docs/PLUGIN_DEVELOPMENT_GUIDE.md).
### Running Emulator
```bash
./scripts/dev/run_emulator.sh
+193 -75
View File
@@ -10,8 +10,14 @@ PROJECT_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
PLUGINS_DIR="$PROJECT_ROOT/plugins"
CONFIG_FILE="$PROJECT_ROOT/dev_plugins.json"
DEFAULT_DEV_DIR="$HOME/.ledmatrix-dev-plugins"
GITHUB_USER="ChuckBuilds"
GITHUB_PATTERN="ledmatrix-"
# Official plugins live in one monorepo: <github_user>/<plugins_repo>, one
# directory per plugin under plugins/. Both can be overridden in
# dev_plugins.json (e.g. to work from a fork).
DEFAULT_GITHUB_USER="ChuckBuilds"
DEFAULT_PLUGINS_REPO="ledmatrix-plugins"
GITHUB_USER="$DEFAULT_GITHUB_USER"
PLUGINS_REPO="$DEFAULT_PLUGINS_REPO"
PLUGINS_BRANCH=""
# Colors for output
RED='\033[0;31m'
@@ -37,18 +43,55 @@ log_error() {
echo -e "${RED}[ERROR]${NC} $1"
}
# Print a top-level string field of a JSON file, or nothing if it is absent.
# Uses jq when installed, else python3.
json_field() {
local file="$1"
local key="$2"
if command -v jq >/dev/null 2>&1; then
jq -r --arg k "$key" '.[$k] // empty | select(type == "string")' "$file" 2>/dev/null || true
elif command -v python3 >/dev/null 2>&1; then
python3 - "$file" "$key" <<'PY' 2>/dev/null || true
import json, sys
try:
with open(sys.argv[1], encoding="utf-8") as f:
value = json.load(f).get(sys.argv[2])
except Exception:
value = None
if isinstance(value, str):
print(value)
PY
fi
}
# Load configuration file
load_config() {
DEV_PLUGINS_DIR="$DEFAULT_DEV_DIR"
if [[ -f "$CONFIG_FILE" ]]; then
DEV_PLUGINS_DIR=$(jq -r '.dev_plugins_dir // "'"$DEFAULT_DEV_DIR"'"' "$CONFIG_FILE" 2>/dev/null || echo "$DEFAULT_DEV_DIR")
# Expand ~ in path
DEV_PLUGINS_DIR="${DEV_PLUGINS_DIR/#\~/$HOME}"
else
DEV_PLUGINS_DIR="$DEFAULT_DEV_DIR"
local value
value=$(json_field "$CONFIG_FILE" dev_plugins_dir)
[[ -n "$value" ]] && DEV_PLUGINS_DIR="$value"
value=$(json_field "$CONFIG_FILE" github_user)
[[ -n "$value" ]] && GITHUB_USER="$value"
value=$(json_field "$CONFIG_FILE" plugins_repo)
[[ -n "$value" ]] && PLUGINS_REPO="$value"
value=$(json_field "$CONFIG_FILE" plugins_branch)
[[ -n "$value" ]] && PLUGINS_BRANCH="$value"
if [[ -n "$(json_field "$CONFIG_FILE" github_pattern)" ]]; then
log_warn "dev_plugins.json: github_pattern is no longer used (official plugins are in the $PLUGINS_REPO monorepo)"
fi
fi
# Expand ~ in path
DEV_PLUGINS_DIR="${DEV_PLUGINS_DIR/#\~/$HOME}"
mkdir -p "$DEV_PLUGINS_DIR"
}
# Top level of the git checkout containing a path, or nothing.
# A monorepo plugin is a subdirectory, so its .git is not in the plugin dir.
git_root_of() {
git -C "$1" rev-parse --show-toplevel 2>/dev/null || true
}
# Validate plugin structure
validate_plugin() {
local plugin_path="$1"
@@ -63,7 +106,7 @@ validate_plugin() {
get_plugin_id() {
local plugin_path="$1"
if [[ -f "$plugin_path/manifest.json" ]]; then
jq -r '.id // empty' "$plugin_path/manifest.json" 2>/dev/null || echo ""
json_field "$plugin_path/manifest.json" id
fi
}
@@ -176,50 +219,104 @@ clone_from_github() {
return 0
}
# Clone a repository into DEV_PLUGINS_DIR, or update the existing clone.
# Prints the clone's path on stdout (log output goes to stderr).
ensure_clone() {
local repo_url="$1"
local branch="${2:-}"
local repo_name
repo_name=$(basename "$repo_url" .git)
local target_dir="$DEV_PLUGINS_DIR/$repo_name"
if [[ -d "$target_dir" ]]; then
log_info "Repository already exists at $target_dir" >&2
if [[ -d "$target_dir/.git" ]]; then
log_info "Updating repository..." >&2
(cd "$target_dir" && git pull --rebase) >&2 || true
fi
else
if ! clone_from_github "$repo_url" "$target_dir" "$branch" >&2; then
return 1
fi
fi
echo "$target_dir"
}
# Find a plugin's directory inside a monorepo clone: plugins/<name>,
# plugins/ledmatrix-<name>, or the directory whose manifest id is <name>.
find_monorepo_plugin() {
local repo_dir="$1"
local name="$2"
local candidate
for candidate in "$repo_dir/plugins/$name" "$repo_dir/plugins/ledmatrix-$name"; do
if [[ -f "$candidate/manifest.json" ]]; then
echo "$candidate"
return 0
fi
done
for candidate in "$repo_dir"/plugins/*/; do
candidate="${candidate%/}"
[[ -f "$candidate/manifest.json" ]] || continue
if [[ "$(get_plugin_id "$candidate")" == "$name" ]]; then
echo "$candidate"
return 0
fi
done
return 1
}
# Link plugin from GitHub
link_github_plugin() {
local plugin_name="$1"
local plugin_name="${1:-}"
local repo_url="${2:-}"
if [[ -z "$plugin_name" ]]; then
log_error "Usage: $0 link-github <plugin-name> [repo-url]"
exit 1
fi
load_config
# Construct repo URL if not provided
if [[ -z "$repo_url" ]]; then
repo_url="https://github.com/${GITHUB_USER}/${GITHUB_PATTERN}${plugin_name}.git"
log_info "Using default GitHub URL: $repo_url"
fi
# Determine target directory name from URL
local repo_name=$(basename "$repo_url" .git)
local target_dir="$DEV_PLUGINS_DIR/$repo_name"
# Check if already cloned
if [[ -d "$target_dir" ]]; then
log_info "Repository already exists at $target_dir"
if [[ -d "$target_dir/.git" ]]; then
log_info "Updating repository..."
(cd "$target_dir" && git pull --rebase) || true
fi
else
# Clone the repository
if ! clone_from_github "$repo_url" "$target_dir"; then
if [[ -n "$repo_url" ]]; then
# A plugin with its own repository (e.g. a third-party plugin): the
# repository root is the plugin.
local target_dir
if ! target_dir=$(ensure_clone "$repo_url"); then
exit 1
fi
if ! validate_plugin "$target_dir"; then
log_error "Cloned repository does not appear to be a valid plugin"
exit 1
fi
link_plugin "$plugin_name" "$target_dir"
return
fi
# Validate plugin structure
if ! validate_plugin "$target_dir"; then
log_error "Cloned repository does not appear to be a valid plugin"
# Official plugins: clone the monorepo once, link plugins/<dir> from it.
repo_url="https://github.com/${GITHUB_USER}/${PLUGINS_REPO}.git"
log_info "Using plugin monorepo: $repo_url"
local repo_dir
if ! repo_dir=$(ensure_clone "$repo_url" "$PLUGINS_BRANCH"); then
exit 1
fi
# Link the plugin
link_plugin "$plugin_name" "$target_dir"
local plugin_dir
if ! plugin_dir=$(find_monorepo_plugin "$repo_dir" "$plugin_name"); then
log_error "No plugin named '$plugin_name' in $repo_dir/plugins"
log_info "Plugins are the directory names under $repo_dir/plugins, or their manifest ids"
exit 1
fi
# Link under the manifest id: that is the name the plugin loader and
# config.json use, and it can differ from the directory name
# (plugins/ledmatrix-music has id ledmatrix-music, not music).
local link_name
link_name=$(get_plugin_id "$plugin_dir")
[[ -n "$link_name" ]] || link_name=$(basename "$plugin_dir")
if [[ "$link_name" != "$plugin_name" ]]; then
log_info "Linking as '$link_name' (the plugin's manifest id)"
fi
link_plugin "$link_name" "$plugin_dir"
}
# Unlink a plugin
@@ -274,7 +371,7 @@ list_plugins() {
echo " → $target"
# Check git status if it's a git repo
if [[ -d "$target/.git" ]]; then
if [[ -n "$(git_root_of "$target")" ]]; then
local branch=$(cd "$target" && git rev-parse --abbrev-ref HEAD 2>/dev/null || echo "unknown")
local status=$(cd "$target" && git status --porcelain 2>/dev/null | head -1)
if [[ -n "$status" ]]; then
@@ -327,7 +424,7 @@ check_status() {
echo -e "${GREEN}✓${NC} ${BLUE}$plugin_name${NC}"
echo " Path: $target"
if [[ -d "$target/.git" ]]; then
if [[ -n "$(git_root_of "$target")" ]]; then
local branch=$(cd "$target" && git rev-parse --abbrev-ref HEAD 2>/dev/null || echo "unknown")
local remote=$(cd "$target" && git remote get-url origin 2>/dev/null || echo "no remote")
local commits_behind=$(cd "$target" && git rev-list --count HEAD..@{upstream} 2>/dev/null || echo "0")
@@ -360,9 +457,13 @@ check_status() {
done
echo "Summary:"
echo " ${GREEN}Clean: $clean_count${NC}"
echo " ${YELLOW}Needs attention: $dirty_count${NC}"
[[ $broken_count -gt 0 ]] && echo -e " ${RED}Broken: $broken_count${NC}"
echo -e " ${GREEN}Clean: $clean_count${NC}"
echo -e " ${YELLOW}Needs attention: $dirty_count${NC}"
# An if, not `[[ ]] &&`: as the function's last command a false test made
# `status` exit 1 whenever nothing was broken.
if [[ $broken_count -gt 0 ]]; then
echo -e " ${RED}Broken: $broken_count${NC}"
fi
}
# Update plugin(s)
@@ -384,39 +485,46 @@ update_plugins() {
fi
local target=$(get_symlink_target "$plugin_name")
if [[ ! -d "$target/.git" ]]; then
local root
root=$(git_root_of "$target")
if [[ -z "$root" ]]; then
log_error "Plugin repository is not a git repository: $target"
exit 1
fi
log_info "Updating $plugin_name from $target"
(cd "$target" && git pull --rebase)
log_info "Updating $plugin_name from $root"
(cd "$root" && git pull --rebase)
log_success "Updated $plugin_name"
else
# Update all linked plugins
# Update all linked plugins. Plugins linked from the monorepo share one
# checkout, which is pulled once.
log_info "Updating all linked plugins..."
local updated=0
local failed=0
local pulled_roots=" "
for item in "$PLUGINS_DIR"/*; do
[[ -e "$item" ]] || continue
[[ -d "$item" ]] || continue
local name=$(basename "$item")
[[ "$name" =~ ^\.|^_ ]] && continue
if is_symlink "$item"; then
local target=$(get_symlink_target "$name")
if [[ -d "$target/.git" ]]; then
log_info "Updating $name..."
if (cd "$target" && git pull --rebase); then
log_success "Updated $name"
updated=$((updated + 1))
else
log_error "Failed to update $name"
failed=$((failed + 1))
fi
local root
root=$(git_root_of "$target")
[[ -n "$root" ]] || continue
[[ "$pulled_roots" == *" $root "* ]] && continue
pulled_roots="$pulled_roots$root "
log_info "Updating $root (for $name)..."
if (cd "$root" && git pull --rebase); then
log_success "Updated $root"
updated=$((updated + 1))
else
log_error "Failed to update $root"
failed=$((failed + 1))
fi
fi
done
@@ -439,7 +547,12 @@ Commands:
link-github <plugin-name> [repo-url]
Clone and link a plugin from GitHub
If repo-url is not provided, uses: https://github.com/${GITHUB_USER}/${GITHUB_PATTERN}<plugin-name>.git
Without repo-url: clones (or updates) the official plugin monorepo,
https://github.com/${DEFAULT_GITHUB_USER}/${DEFAULT_PLUGINS_REPO}.git, and links its
plugins/<plugin-name> (or plugins/ledmatrix-<plugin-name>) under the
plugin's manifest id
With repo-url: clones a plugin that has its own repository and links
the repository root
unlink <plugin-name>
Remove symlink for a plugin (preserves repository)
@@ -458,25 +571,30 @@ Commands:
Show this help message
Examples:
# Link a local plugin
$0 link music ../ledmatrix-music
# Link from GitHub (auto-detects URL)
$0 link-github music
# Link from GitHub with custom URL
$0 link-github stocks https://github.com/ChuckBuilds/ledmatrix-stocks.git
# Link an official plugin from the monorepo
$0 link-github football-scoreboard
# Link a plugin from a local monorepo checkout
$0 link hello-world ../ledmatrix-plugins/plugins/hello-world
# Link a third-party plugin from its own repository
$0 link-github my-plugin https://github.com/OtherUser/ledmatrix-my-plugin.git
# Check status
$0 status
# Update all plugins
$0 update
Configuration:
Create dev_plugins.json in project root to customize:
Copy dev_plugins.json.example to dev_plugins.json (git-ignored) to customize:
- dev_plugins_dir: Where to clone GitHub repos (default: ~/.ledmatrix-dev-plugins)
- plugins: Plugin definitions (optional, for auto-discovery)
- github_user: Owner of the plugin monorepo, e.g. your fork (default: ${DEFAULT_GITHUB_USER})
- plugins_repo: Name of the plugin monorepo (default: ${DEFAULT_PLUGINS_REPO})
- plugins_branch: Branch to clone the monorepo at (default: its default branch)
Symlinks are created in plugins/. Set plugin_system.plugins_directory to
"plugins" in config/config.json so the plugin loader discovers them.
EOF
}
+2 -6
View File
@@ -72,12 +72,8 @@ def load_main_config(path: Path) -> Dict[str, Any]:
def display_size_from_config(config: Dict[str, Any]) -> tuple:
"""Derive the logical ticker size the way DisplayManager does."""
hw = config.get('display', {}).get('hardware', {})
cols = int(hw.get('cols', 64))
chain = int(hw.get('chain_length', 1))
rows = int(hw.get('rows', 32))
parallel = int(hw.get('parallel', 1))
return cols * chain, rows * parallel
from src.display_geometry import logical_size
return logical_size(config)
def enabled_plugin_ids(config: Dict[str, Any]) -> List[str]:
+24 -17
View File
@@ -32,6 +32,8 @@ os.environ['EMULATOR'] = 'true'
from flask import Flask, render_template, request, jsonify
from src.common.path_safety import resolve_under, safe_path_component
app = Flask(__name__, template_folder=str(Path(__file__).parent / 'templates'))
logger = logging.getLogger(__name__)
@@ -118,7 +120,8 @@ def find_plugin_dir(plugin_id: str) -> Optional[Path]:
one of the plugin search dirs, so a crafted id can never name a path
outside them.
"""
if not isinstance(plugin_id, str) or not _SAFE_PLUGIN_ID_RE.match(plugin_id):
plugin_id = safe_path_component(plugin_id)
if not plugin_id or not _SAFE_PLUGIN_ID_RE.match(plugin_id):
return None
from src.plugin_system.plugin_loader import PluginLoader
loader = PluginLoader()
@@ -139,17 +142,18 @@ def find_plugin_dir(plugin_id: str) -> Optional[Path]:
def load_config_defaults(plugin_dir: 'str | Path') -> Dict[str, Any]:
"""Extract default values from config_schema.json."""
schema_path = Path(plugin_dir) / 'config_schema.json'
if not schema_path.exists():
"""Extract default values from config_schema.json.
The same extraction a device and the plugin harness use
(src/plugin_system/testing/loading.py), nested defaults included.
"""
from src.plugin_system.testing.loading import (
load_config_defaults as _load_config_defaults,
)
schema_path = resolve_under(plugin_dir, 'config_schema.json')
if schema_path is None or not schema_path.exists():
return {}
with open(schema_path, 'r') as f:
schema = json.load(f)
defaults: Dict[str, Any] = {}
for key, prop in schema.get('properties', {}).items():
if 'default' in prop:
defaults[key] = prop['default']
return defaults
return _load_config_defaults(schema_path.parent)
# --------------------------------------------------------------------------
@@ -175,8 +179,8 @@ def api_plugin_schema(plugin_id):
if not plugin_dir:
return jsonify({'error': f'Plugin not found: {plugin_id}'}), 404
schema_path = plugin_dir / 'config_schema.json'
if not schema_path.exists():
schema_path = resolve_under(plugin_dir, 'config_schema.json')
if schema_path is None or not schema_path.exists():
return jsonify({'schema': {'type': 'object', 'properties': {}}})
with open(schema_path, 'r') as f:
@@ -300,10 +304,13 @@ def _parse_render_request(data):
with open(manifest_path, 'r') as f:
manifest = json.load(f)
# Build config: schema defaults + user overrides
config = {'enabled': True}
config.update(load_config_defaults(trusted_dir))
config.update(data.get('config', {}))
# Build config the way a device would: schema defaults under a forced
# enabled, with the user's overrides deep-merged on top
from src.plugin_system.testing.loading import build_config
overrides = data.get('config') or {}
if not isinstance(overrides, dict):
raise ValueError('config must be a JSON object')
config = build_config(trusted_dir, overrides)
return trusted_dir, manifest, config, data.get('mock_data', {}), data.get('skip_update', False)
+55 -11
View File
@@ -20,6 +20,38 @@ PROJECT_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
cd "$PROJECT_DIR"
# Report web_display_autostart the way scripts/utils/start_web_conditionally.py
# (what ledmatrix-web.service runs) decides it: only an explicit false/off keeps
# the web interface down; a missing key or an unreadable config starts it.
# Prints "on <value>", "off <value>", "default" (key not set) or "unreadable".
web_autostart_state() {
(cd "$1" && python3 - 2>/dev/null <<'PY'
import json, os, sys
sys.path.insert(0, os.path.join(os.getcwd(), "scripts", "utils"))
try:
from start_web_conditionally import autostart_enabled
except Exception:
def autostart_enabled(config):
value = config.get("web_display_autostart", True)
if isinstance(value, str):
return value.strip().lower() not in ("off", "false", "no", "0")
return bool(value)
try:
with open(os.path.join("config", "config.json"), encoding="utf-8") as f:
config = json.load(f)
except Exception:
config = None
if not isinstance(config, dict):
print("unreadable")
elif "web_display_autostart" not in config:
print("default")
else:
raw = json.dumps(config["web_display_autostart"])
print(("on " if autostart_enabled(config) else "off ") + raw)
PY
) || echo "unknown"
}
echo -e "${BLUE}━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━${NC}"
echo -e "${BLUE}1. SERVICE STATUS${NC}"
echo -e "${BLUE}━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━${NC}"
@@ -41,14 +73,26 @@ if [ -f "$PROJECT_DIR/config/config.json" ]; then
echo -e "${GREEN}✓ Config file found${NC}"
# Check web_display_autostart setting
AUTOSTART=$(grep -o '"web_display_autostart"[[:space:]]*:[[:space:]]*[a-z]*' "$PROJECT_DIR/config/config.json" | grep -o '[a-z]*$')
if [ "$AUTOSTART" == "true" ]; then
echo -e "${GREEN}✓ web_display_autostart: true${NC}"
else
echo -e "${YELLOW}⚠ web_display_autostart: ${AUTOSTART:-not set}${NC}"
echo -e "${YELLOW} Web interface will not start unless this is set to true${NC}"
fi
AUTOSTART=$(web_autostart_state "$PROJECT_DIR")
case "$AUTOSTART" in
on\ *)
echo -e "${GREEN}✓ web_display_autostart: ${AUTOSTART#on }${NC}"
;;
off\ *)
echo -e "${YELLOW}⚠ web_display_autostart: ${AUTOSTART#off }${NC}"
echo -e "${YELLOW} Web interface will not start with this value${NC}"
;;
default)
echo -e "${GREEN}✓ web_display_autostart: not set (defaults to on)${NC}"
;;
unreadable)
echo -e "${YELLOW}⚠ config.json could not be parsed (the web interface still starts so it can be repaired)${NC}"
;;
*)
echo -e "${YELLOW}⚠ web_display_autostart: could not be evaluated (python3 unavailable?)${NC}"
;;
esac
else
echo -e "${RED}✗ Config file not found at: $PROJECT_DIR/config/config.json${NC}"
fi
@@ -63,7 +107,7 @@ declare -a REQUIRED_FILES=(
"web_interface/app.py"
"web_interface/start.py"
"web_interface/requirements.txt"
"web_interface/blueprints/api_v3.py"
"web_interface/blueprints/api_v3/__init__.py"
"web_interface/blueprints/pages_v3.py"
"scripts/utils/start_web_conditionally.py"
)
@@ -134,8 +178,8 @@ if ! sudo systemctl is-active --quiet ledmatrix-web; then
echo " sudo systemctl start ledmatrix-web"
fi
if [ "$AUTOSTART" != "true" ]; then
echo -e "${YELLOW}→ Enable web_display_autostart in config/config.json${NC}"
if [ "${AUTOSTART%% *}" = "off" ]; then
echo -e "${YELLOW}→ Set web_display_autostart to true in config/config.json (or remove it; missing means on)${NC}"
fi
if [ "$ALL_FILES_OK" = false ]; then
+53 -12
View File
@@ -22,6 +22,38 @@ fi
PROJECT_DIR="${HOME}/LEDMatrix"
# Report web_display_autostart the way scripts/utils/start_web_conditionally.py
# (what ledmatrix-web.service runs) decides it: only an explicit false/off keeps
# the web interface down; a missing key or an unreadable config starts it.
# Prints "on <value>", "off <value>", "default" (key not set) or "unreadable".
web_autostart_state() {
(cd "$1" && python3 - 2>/dev/null <<'PY'
import json, os, sys
sys.path.insert(0, os.path.join(os.getcwd(), "scripts", "utils"))
try:
from start_web_conditionally import autostart_enabled
except Exception:
def autostart_enabled(config):
value = config.get("web_display_autostart", True)
if isinstance(value, str):
return value.strip().lower() not in ("off", "false", "no", "0")
return bool(value)
try:
with open(os.path.join("config", "config.json"), encoding="utf-8") as f:
config = json.load(f)
except Exception:
config = None
if not isinstance(config, dict):
print("unreadable")
elif "web_display_autostart" not in config:
print("default")
else:
raw = json.dumps(config["web_display_autostart"])
print(("on " if autostart_enabled(config) else "off ") + raw)
PY
) || echo "unknown"
}
echo "1. Checking service status..."
echo "------------------------------"
if systemctl is-active --quiet ledmatrix-web 2>/dev/null || sudo systemctl is-active --quiet ledmatrix-web 2>/dev/null; then
@@ -47,16 +79,25 @@ echo "3. Checking configuration file..."
echo "------------------------------"
if [ -f "${PROJECT_DIR}/config/config.json" ]; then
echo -e "${GREEN}✓ Config file exists${NC}"
AUTOSTART=$(grep -o '"web_display_autostart":\s*\(true\|false\)' "${PROJECT_DIR}/config/config.json" | grep -o '\(true\|false\)' || echo "not found")
if [ "$AUTOSTART" = "true" ]; then
echo -e "${GREEN}✓ web_display_autostart is set to TRUE${NC}"
elif [ "$AUTOSTART" = "false" ]; then
echo -e "${RED}✗ web_display_autostart is set to FALSE (web UI won't start!)${NC}"
echo " Fix: Edit config.json and set 'web_display_autostart': true"
else
echo -e "${YELLOW}⚠ web_display_autostart setting not found (defaults to false)${NC}"
echo " Fix: Add 'web_display_autostart': true to config.json"
fi
AUTOSTART=$(web_autostart_state "$PROJECT_DIR")
case "$AUTOSTART" in
on\ *)
echo -e "${GREEN}✓ web_display_autostart is ${AUTOSTART#on } (web UI starts)${NC}"
;;
off\ *)
echo -e "${RED}✗ web_display_autostart is ${AUTOSTART#off } (web UI won't start!)${NC}"
echo " Fix: Edit config.json and set 'web_display_autostart': true"
;;
default)
echo -e "${GREEN}✓ web_display_autostart is not set (defaults to on; web UI starts)${NC}"
;;
unreadable)
echo -e "${YELLOW}⚠ config.json could not be parsed (the web UI still starts so it can be repaired)${NC}"
;;
*)
echo -e "${YELLOW}⚠ Could not evaluate web_display_autostart (python3 unavailable?)${NC}"
;;
esac
else
echo -e "${RED}✗ Config file NOT FOUND at ${PROJECT_DIR}/config/config.json${NC}"
fi
@@ -82,7 +123,7 @@ FILES_TO_CHECK=(
"web_interface/start.py"
"web_interface/app.py"
"web_interface/requirements.txt"
"web_interface/blueprints/api_v3.py"
"web_interface/blueprints/api_v3/__init__.py"
"web_interface/blueprints/pages_v3.py"
)
@@ -175,7 +216,7 @@ echo "Diagnostic Summary"
echo "=========================================="
echo ""
echo "Most common issues:"
echo " 1. web_display_autostart is false or missing in config.json"
echo " 1. web_display_autostart is set to false in config.json (a missing key means on)"
echo " 2. Service not enabled or not started"
echo " 3. Missing dependencies (Flask, etc.)"
echo " 4. Import errors in web_interface/app.py"
+15 -7
View File
@@ -7,8 +7,10 @@ install as the wrong user, after a manual file copy that didn't preserve
ownership, or after a permissions-related error from the display or
web service.
Most of these scripts require `sudo` since they touch directories
owned by the `ledmatrix` service user or by `root`.
Most of these scripts require `sudo` since they touch directories owned
by `root` (the display service's user) or by the user you installed
LEDMatrix as (the web service's user). There is no dedicated `ledmatrix`
system user.
## Scripts
@@ -16,11 +18,12 @@ owned by the `ledmatrix` service user or by `root`.
permissions on the `assets/` tree so plugins can download and cache
team logos, fonts, and other static content.
- **`fix_cache_permissions.sh`** — Fixes permissions on every cache
directory the project may use (`/var/cache/ledmatrix/`,
`~/.cache/ledmatrix/`, `/opt/ledmatrix/cache/`, project-local
`cache/`). Also creates placeholder logo subdirectories used by the
sports plugins.
- **`fix_cache_permissions.sh`** — Creates (if missing) and fixes
permissions on `/var/cache/ledmatrix/` and `~/.ledmatrix_cache/` of the
user running `sudo`, and creates
`/var/cache/ledmatrix/placeholder_logos/` for the sports plugins. It does
not touch the cache manager's other fallbacks (`/opt/ledmatrix/cache`,
`$TMPDIR/ledmatrix_cache`).
- **`fix_plugin_permissions.sh`** — Fixes ownership on the plugins
directory so both the root display service and the web service user
@@ -31,6 +34,11 @@ owned by the `ledmatrix` service user or by `root`.
systemd journal access, and the sudoers entries the web interface
needs to control the display service.
- **`safe_pip_install.sh`** — Installs a `requirements.txt` as root
after checking it is the project's own or one under `plugin-repos/` or
`plugins/`. Used by the web interface (via sudo) to install plugin
dependencies where `ledmatrix.service` can import them.
- **`safe_plugin_rm.sh`** — Validates that a plugin removal path is
inside an allowed base directory before deleting it. Used by the web
interface (via sudo) when a user clicks **Uninstall** on a plugin —
+20 -7
View File
@@ -1,6 +1,7 @@
#!/bin/bash
# safe_pip_install.sh — Install a requirements.txt as root after validating
# that the resolved path is the project's own requirements.txt or a plugin's
# that the resolved path is one of the project's own requirements files
# (requirements.txt, web_interface/requirements.txt) or a plugin's
# requirements.txt under plugin-repos/ or plugins/.
#
# This script is intended to be called via sudo from the web interface, so
@@ -25,9 +26,17 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
PROJECT_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
# Allowed locations (resolved, no trailing slash):
# - the project's own requirements.txt
# - the project's own requirements files. Update Code, the automatic
# update's health check and Install Base Requirements install both, so
# one missing here is refused on every device -- and the automatic
# updater rolls back any update that changes it.
# Only their folders are resolved: resolving the files too would follow
# a requirements.txt symlinked out of the project and allow its target.
# - any requirements.txt under plugin-repos/ or plugins/
ALLOWED_EXACT="$(realpath --canonicalize-missing "$PROJECT_ROOT/requirements.txt")"
ALLOWED_EXACT=(
"$(realpath --canonicalize-missing "$PROJECT_ROOT")/requirements.txt"
"$(realpath --canonicalize-missing "$PROJECT_ROOT/web_interface")/requirements.txt"
)
ALLOWED_BASES=(
"$(realpath --canonicalize-missing "$PROJECT_ROOT/plugin-repos")"
"$(realpath --canonicalize-missing "$PROJECT_ROOT/plugins")"
@@ -43,9 +52,13 @@ if [ "$(basename "$RESOLVED_TARGET")" != "requirements.txt" ]; then
fi
ALLOWED=false
if [ "$RESOLVED_TARGET" = "$ALLOWED_EXACT" ]; then
ALLOWED=true
else
for EXACT in "${ALLOWED_EXACT[@]}"; do
if [ "$RESOLVED_TARGET" = "$EXACT" ]; then
ALLOWED=true
break
fi
done
if [ "$ALLOWED" = false ]; then
for BASE in "${ALLOWED_BASES[@]}"; do
if [[ "$RESOLVED_TARGET" == "$BASE/"* ]]; then
ALLOWED=true
@@ -56,7 +69,7 @@ fi
if [ "$ALLOWED" = false ]; then
echo "DENIED: $RESOLVED_TARGET is not an allowed requirements.txt location" >&2
echo "Allowed: $ALLOWED_EXACT, or any requirements.txt under: ${ALLOWED_BASES[*]}" >&2
echo "Allowed: ${ALLOWED_EXACT[*]}, or any requirements.txt under: ${ALLOWED_BASES[*]}" >&2
exit 2
fi
+9 -5
View File
@@ -7,14 +7,18 @@ This directory contains scripts for installing and configuring the LEDMatrix sys
- **`one-shot-install.sh`** - Single-command installer; clones the
repo, checks prerequisites, then runs `first_time_install.sh`.
Invoked via `curl ... | bash` from the project root README.
- **`install_service.sh`** - Installs the main LED Matrix display service (systemd)
- **`install_web_service.sh`** - Installs the web interface service (systemd)
- **`install_service.sh`** - Installs, enables and starts the display
service (`ledmatrix.service`), the web interface service
(`ledmatrix-web.service`) and the update-verify units (systemd)
- **`install_web_service.sh`** - Installs only the web interface service
and the update-verify units (systemd)
- **`install_wifi_monitor.sh`** - Installs the WiFi monitor daemon service
- **`setup_cache.sh`** - Sets up persistent cache directory with proper permissions
- **`configure_web_sudo.sh`** - Configures passwordless sudo access for web interface actions
- **`configure_wifi_permissions.sh`** - Grants the `ledmatrix` user
the WiFi management permissions needed by the web interface and
the WiFi monitor service
- **`configure_wifi_permissions.sh`** - Grants the web interface's user
(the user who runs the script, i.e. the one you installed LEDMatrix as;
there is no `ledmatrix` system user) the passwordless `nmcli` and related
WiFi permissions the web interface needs
- **`migrate_config.sh`** - Migrates configuration files to new formats (if needed)
- **`debug_install.sh`** - Diagnostic helper used when an install
fails; collects environment info and recent logs
+13 -9
View File
@@ -111,13 +111,6 @@ TEMP_SUDOERS="/tmp/ledmatrix_web_sudoers_$$"
echo "$WEB_USER ALL=(ALL) NOPASSWD:NOEXEC: $JOURNALCTL_PATH -t ledmatrix *"
fi
# Required: python3, bash
# NOTE: display_controller.py/start_display.sh/stop_display.sh live at the
# project root, not under scripts/install/ (where this script lives) —
# must use PROJECT_ROOT here, not PROJECT_DIR.
echo "$WEB_USER ALL=(ALL) NOPASSWD: $PYTHON_PATH $PROJECT_ROOT/display_controller.py"
echo "$WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/start_display.sh"
echo "$WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/stop_display.sh"
echo ""
echo "# Allow web user to remove plugin directories via vetted helper script"
echo "# The helper validates that the target path resolves inside plugin-repos/ or plugins/"
@@ -130,6 +123,19 @@ TEMP_SUDOERS="/tmp/ledmatrix_web_sudoers_$$"
echo "$WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $SAFE_PIP_INSTALL_PATH *"
} > "$TEMP_SUDOERS"
# Never offer to install rules we have not parsed. A malformed drop-in in
# /etc/sudoers.d makes sudo refuse every command for every user.
if command -v visudo >/dev/null 2>&1; then
if ! visudo -c -f "$TEMP_SUDOERS" >/dev/null 2>&1; then
echo ""
echo "✗ The generated sudoers rules did not parse:" >&2
visudo -c -f "$TEMP_SUDOERS" >&2 || true
echo "Nothing was changed." >&2
rm -f "$TEMP_SUDOERS"
exit 1
fi
fi
echo ""
echo "Generated sudoers configuration:"
echo "--------------------------------"
@@ -142,8 +148,6 @@ echo "- Start/stop/restart the ledmatrix service"
echo "- Enable/disable the ledmatrix service"
echo "- Check service status"
echo "- View system logs via journalctl"
echo "- Run display_controller.py directly"
echo "- Execute start_display.sh and stop_display.sh"
echo "- Reboot and shutdown the system"
echo "- Remove plugin directories (for update/uninstall when root-owned files block deletion)"
echo "- Install plugin/base requirements.txt as root (so ledmatrix.service can see them)"
+86
View File
@@ -0,0 +1,86 @@
#!/bin/bash
# DNS single-request fix installation script.
#
# Optional. Install this only if plugins that call external APIs (Starlark
# apps, weather, sports, music) are timing out or feel slow to first paint
# while the network is otherwise fine. See the header of
# scripts/utils/apply_dns_single_request.sh for what it changes and why.
set -e
PROJECT_ROOT_DIR=$(cd "$(dirname "$0")/../.." && pwd)
SERVICE_NAME="ledmatrix-dns-fix"
UNIT_SRC="$PROJECT_ROOT_DIR/systemd/$SERVICE_NAME.service"
UNIT_DEST="/etc/systemd/system/$SERVICE_NAME.service"
DROPIN_DIR="/etc/systemd/system/ledmatrix.service.d"
if [ "$EUID" -eq 0 ]; then
SYSTEMCTL_CMD="systemctl"
SUDO=""
else
SYSTEMCTL_CMD="sudo systemctl"
SUDO="sudo"
fi
echo "Installing LED Matrix DNS fix service"
echo "Project root directory: $PROJECT_ROOT_DIR"
if [ ! -f "$UNIT_SRC" ]; then
echo "✗ Missing unit file: $UNIT_SRC"
exit 1
fi
chmod +x "$PROJECT_ROOT_DIR/scripts/utils/apply_dns_single_request.sh"
echo "Installing $UNIT_DEST..."
sed "s|__PROJECT_ROOT_DIR__|$PROJECT_ROOT_DIR|g" "$UNIT_SRC" \
| $SUDO tee "$UNIT_DEST" > /dev/null
# Order ledmatrix.service after the fix. `Before=` in the unit itself only
# orders units already in the same transaction, so a plain
# `systemctl restart ledmatrix` would not wait for it -- and since this fix is
# opt-in, ledmatrix.service cannot carry the dependency in the repo.
# Wants=, not Requires=: a DNS workaround failing should not stop the display.
echo "Installing the ledmatrix.service ordering drop-in..."
$SUDO mkdir -p "$DROPIN_DIR"
printf '[Unit]\nWants=%s.service\nAfter=%s.service\n' "$SERVICE_NAME" "$SERVICE_NAME" \
| $SUDO tee "$DROPIN_DIR/10-dns-fix.conf" > /dev/null
$SYSTEMCTL_CMD daemon-reload
$SYSTEMCTL_CMD enable "$SERVICE_NAME.service"
# Do not mask a failure here. The unit exits non-zero when it could not apply
# the option -- a systemd-resolved host, an unwritable resolv.conf, a failed
# `resolvconf -u` -- and reporting "installation complete" over that would
# leave the operator believing a workaround is active when it is not.
START_STATUS=0
$SYSTEMCTL_CMD start "$SERVICE_NAME.service" || START_STATUS=$?
echo ""
if grep -qs "^options single-request$" /etc/resolv.conf; then
echo "✓ 'options single-request' is active in /etc/resolv.conf"
elif [ "$START_STATUS" -ne 0 ]; then
echo "✗ The DNS fix could not be applied on this host."
echo " The service reported why:"
echo " journalctl -u $SERVICE_NAME -n 20"
echo ""
echo " The unit is installed and will try again on the next boot. Nothing"
echo " else about your install has changed."
exit "$START_STATUS"
else
echo "⚠ 'options single-request' is not in /etc/resolv.conf yet."
echo " Check what the service reported:"
echo " journalctl -u $SERVICE_NAME -n 20"
fi
echo ""
echo "DNS fix installation complete."
echo ""
echo "Useful commands:"
echo " sudo systemctl status $SERVICE_NAME # Check status"
echo " sudo journalctl -u $SERVICE_NAME -n 50 # View logs"
echo " sudo systemctl disable --now $SERVICE_NAME # Undo the service"
echo " sudo rm $DROPIN_DIR/10-dns-fix.conf # Undo the ordering drop-in"
echo " # then remove the 'options single-request' line from /etc/resolv.conf"
echo ""
+64
View File
@@ -0,0 +1,64 @@
#!/bin/bash
# Home Assistant MQTT bridge installation script.
#
# Optional. Installs integrations/mqtt_bridge as a service so Home Assistant
# can force display modes, toggle power and set brightness over MQTT.
# See integrations/mqtt_bridge/README.md.
set -e
PROJECT_ROOT_DIR=$(cd "$(dirname "$0")/../.." && pwd)
BRIDGE_DIR="$PROJECT_ROOT_DIR/integrations/mqtt_bridge"
SERVICE_NAME="ledmatrix-mqtt-bridge"
UNIT_SRC="$PROJECT_ROOT_DIR/systemd/$SERVICE_NAME.service"
UNIT_DEST="/etc/systemd/system/$SERVICE_NAME.service"
if [ "$EUID" -eq 0 ]; then
SYSTEMCTL_CMD="systemctl"
SUDO=""
else
SYSTEMCTL_CMD="sudo systemctl"
SUDO="sudo"
fi
echo "Installing LED Matrix MQTT bridge"
echo "Project root directory: $PROJECT_ROOT_DIR"
if [ ! -f "$BRIDGE_DIR/bridge_config.json" ]; then
cp "$BRIDGE_DIR/bridge_config.example.json" "$BRIDGE_DIR/bridge_config.json"
chmod 600 "$BRIDGE_DIR/bridge_config.json"
echo ""
echo "⚠ Created $BRIDGE_DIR/bridge_config.json from the example."
echo " Edit it with your broker details, then re-run this script."
echo " The service will refuse to start until the placeholder password is replaced."
echo ""
fi
echo "Installing Python dependencies..."
python3 -m pip install -r "$BRIDGE_DIR/requirements.txt" 2>/dev/null \
|| python3 -m pip install --break-system-packages -r "$BRIDGE_DIR/requirements.txt"
echo "Installing $UNIT_DEST..."
sed "s|__PROJECT_ROOT_DIR__|$PROJECT_ROOT_DIR|g" "$UNIT_SRC" \
| $SUDO tee "$UNIT_DEST" > /dev/null
$SYSTEMCTL_CMD daemon-reload
$SYSTEMCTL_CMD enable "$SERVICE_NAME.service"
$SYSTEMCTL_CMD restart "$SERVICE_NAME.service" || true
echo ""
if $SYSTEMCTL_CMD is-active --quiet "$SERVICE_NAME.service" 2>/dev/null; then
echo "✓ MQTT bridge is running"
echo " The matrix should appear in Home Assistant under Settings > Devices > MQTT."
else
echo "⚠ MQTT bridge is not running. Check the logs:"
echo " sudo journalctl -u $SERVICE_NAME -n 50"
fi
echo ""
echo "Useful commands:"
echo " sudo systemctl status $SERVICE_NAME"
echo " sudo journalctl -u $SERVICE_NAME -f"
echo " sudo systemctl disable --now $SERVICE_NAME # Undo"
echo ""
+114 -35
View File
@@ -3,6 +3,41 @@
# Exit on error
set -e
usage() {
cat <<'USAGE'
Usage: sudo ./scripts/install/install_service.sh [-h|--help]
Installs (or reinstalls) the LEDMatrix systemd units from the templates in
systemd/, then enables and starts them:
- ledmatrix.service main display (runs as root)
- ledmatrix-web.service web interface (runs as the invoking user)
- ledmatrix-update-verify.service automatic-update health check
- ledmatrix-update-verify.path
Existing unit files in /etc/systemd/system are overwritten. The script takes
no other options; run it with no arguments to install.
Options:
-h, --help Show this help and exit without changing anything.
USAGE
}
# Parse arguments before touching anything: this script rewrites and restarts
# services, so an unrecognised option must not fall through to a full install.
for arg in "$@"; do
case "$arg" in
-h|--help)
usage
exit 0
;;
*)
echo "ERROR: unknown option: $arg" >&2
usage >&2
exit 2
;;
esac
done
# Get the actual user who invoked sudo
if [ -n "$SUDO_USER" ]; then
ACTUAL_USER="$SUDO_USER"
@@ -16,20 +51,39 @@ USER_HOME=$(eval echo ~$ACTUAL_USER)
# Determine the Project Root Directory (parent of scripts/install/)
PROJECT_ROOT_DIR=$(cd "$(dirname "$0")/../.." && pwd)
# shellcheck source=scripts/install/lib_systemd_render.sh
source "$PROJECT_ROOT_DIR/scripts/install/lib_systemd_render.sh"
echo "Installing LED Matrix Display Service for user: $ACTUAL_USER"
echo "Using home directory: $USER_HOME"
echo "Project root directory: $PROJECT_ROOT_DIR"
# Create a temporary service file for the main display with the correct paths
# Assuming ledmatrix.service template exists and uses /home/ledpi as a placeholder for user home
# Render the main display unit from its template. The display service runs as
# root (it needs GPIO), so __USER__ is always root here -- unlike the web unit
# below, which runs as whoever installed it.
#
# A missing template or a failed render is fatal: falling through would leave
# whatever unit already sits at /etc/systemd/system/ledmatrix.service (from a
# previous install) untouched, and the enable/start step below would then
# silently reuse that stale unit instead of the one this run was asked to
# install.
if [ -f "$PROJECT_ROOT_DIR/systemd/ledmatrix.service" ]; then
sed "s|/home/ledpi|$USER_HOME|g; s|__PROJECT_ROOT_DIR__|$PROJECT_ROOT_DIR|g; s|__USER__|root|g" "$PROJECT_ROOT_DIR/systemd/ledmatrix.service" > /tmp/ledmatrix.service.tmp
ESCAPED_PROJECT_ROOT_DIR=$(sed_escape_replacement "$PROJECT_ROOT_DIR")
MAIN_UNIT_TMP=$(mktemp)
trap 'rm -f "$MAIN_UNIT_TMP"' EXIT
if ! sed "s|__PROJECT_ROOT_DIR__|$ESCAPED_PROJECT_ROOT_DIR|g; s|__USER__|root|g" \
"$PROJECT_ROOT_DIR/systemd/ledmatrix.service" > "$MAIN_UNIT_TMP"; then
echo "ERROR: failed to render ledmatrix.service from its template." >&2
exit 1
fi
# Copy the service file to the systemd directory
sudo cp /tmp/ledmatrix.service.tmp /etc/systemd/system/ledmatrix.service
sudo cp "$MAIN_UNIT_TMP" /etc/systemd/system/ledmatrix.service
# Clean up
rm /tmp/ledmatrix.service.tmp
rm -f "$MAIN_UNIT_TMP"
trap - EXIT
else
echo "WARNING: ledmatrix.service template not found at $PROJECT_ROOT_DIR/systemd/ledmatrix.service. Main display service not configured."
echo "ERROR: ledmatrix.service template not found at $PROJECT_ROOT_DIR/systemd/ledmatrix.service." >&2
exit 1
fi
@@ -48,42 +102,67 @@ fi
# === LEDMatrix Web Interface service (ledmatrix-web.service) ===
echo "Installing LEDMatrix Web Interface service (ledmatrix-web.service)..."
WEB_SERVICE_FILE_CONTENT=$(cat <<EOF
[Unit]
Description=LED Matrix Web Interface (Conditional Start)
After=network.target
# Wants=ledmatrix.service
# After=network.target ledmatrix.service
# Rendered from systemd/ledmatrix-web.service, the same template
# install_web_service.sh uses. This was an inline heredoc until it drifted from
# the template: it had lost Wants=network-online.target, RestartSec,
# SyslogIdentifier and Environment=USE_THREADING. Because
# src/startup_validator.py compares the installed unit against the template,
# every boot warned "re-run install_service.sh" -- and doing so reinstalled the
# same stale copy, so the warning could never clear.
#
# As with the main unit above, a missing template or a failed render is
# fatal -- otherwise the enable/start check below would fall back to
# whatever unit (possibly stale) already exists at the destination path.
if [ -f "$PROJECT_ROOT_DIR/systemd/ledmatrix-web.service" ]; then
ESCAPED_ACTUAL_USER=$(sed_escape_replacement "$ACTUAL_USER")
WEB_UNIT_TMP=$(mktemp)
trap 'rm -f "$WEB_UNIT_TMP"' EXIT
if ! sed "s|__PROJECT_ROOT_DIR__|$ESCAPED_PROJECT_ROOT_DIR|g; s|__USER__|$ESCAPED_ACTUAL_USER|g" \
"$PROJECT_ROOT_DIR/systemd/ledmatrix-web.service" > "$WEB_UNIT_TMP"; then
echo "ERROR: failed to render ledmatrix-web.service from its template." >&2
exit 1
fi
sudo cp "$WEB_UNIT_TMP" /etc/systemd/system/ledmatrix-web.service
rm -f "$WEB_UNIT_TMP"
trap - EXIT
else
echo "ERROR: ledmatrix-web.service template not found at $PROJECT_ROOT_DIR/systemd/ledmatrix-web.service." >&2
exit 1
fi
[Service]
Type=simple
ExecStart=/usr/bin/python3 ${PROJECT_ROOT_DIR}/scripts/utils/start_web_conditionally.py
WorkingDirectory=${PROJECT_ROOT_DIR}
StandardOutput=journal
StandardError=journal
User=${ACTUAL_USER}
Restart=on-failure
# Environment="PYTHONUNBUFFERED=1"
[Install]
WantedBy=multi-user.target
EOF
)
# Write the new service file
echo "$WEB_SERVICE_FILE_CONTENT" | sudo tee /etc/systemd/system/ledmatrix-web.service > /dev/null
# Health check / rollback units for automatic updates; see install_web_service.sh.
for VERIFY_UNIT in ledmatrix-update-verify.service ledmatrix-update-verify.path; do
if [ -f "$PROJECT_ROOT_DIR/systemd/$VERIFY_UNIT" ]; then
VERIFY_UNIT_TMP=$(mktemp)
if sed "s|__PROJECT_ROOT_DIR__|$ESCAPED_PROJECT_ROOT_DIR|g; s|__USER__|$ESCAPED_ACTUAL_USER|g" "$PROJECT_ROOT_DIR/systemd/$VERIFY_UNIT" > "$VERIFY_UNIT_TMP"; then
sudo cp "$VERIFY_UNIT_TMP" "/etc/systemd/system/$VERIFY_UNIT"
else
echo "WARNING: failed to render $VERIFY_UNIT; automatic code updates will stay paused." >&2
fi
rm -f "$VERIFY_UNIT_TMP"
fi
done
echo "Reloading systemd daemon for web service..."
sudo systemctl daemon-reload
echo "Enabling ledmatrix-web.service to start on boot..."
sudo systemctl enable ledmatrix-web.service
if [ -f "/etc/systemd/system/ledmatrix-web.service" ]; then
echo "Enabling ledmatrix-web.service to start on boot..."
sudo systemctl enable ledmatrix-web.service
echo "Starting ledmatrix-web.service..."
sudo systemctl start ledmatrix-web.service
if [ -f /etc/systemd/system/ledmatrix-update-verify.path ]; then
echo "Enabling ledmatrix-update-verify.path (automatic update health check)..."
sudo systemctl enable --now ledmatrix-update-verify.path || echo "WARNING: could not enable ledmatrix-update-verify.path; automatic code updates will stay paused" >&2
fi
echo "LEDMatrix Web Interface service (ledmatrix-web.service) installation complete."
echo "It will start based on the 'web_display_autostart' setting in config/config.json."
echo "Starting ledmatrix-web.service..."
sudo systemctl start ledmatrix-web.service
echo "LEDMatrix Web Interface service (ledmatrix-web.service) installation complete."
echo "It will start based on the 'web_display_autostart' setting in config/config.json."
else
echo "Skipping enable/start for ledmatrix-web.service as it was not configured."
fi
# === End of LEDMatrix Web Interface service ===
+72 -45
View File
@@ -17,6 +17,9 @@ fi
# Determine the Project Root Directory (parent of scripts/install/)
PROJECT_ROOT_DIR=$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)
# shellcheck source=scripts/install/lib_systemd_render.sh
source "$PROJECT_ROOT_DIR/scripts/install/lib_systemd_render.sh"
echo "Installing for user: $ACTUAL_USER"
echo "Project root directory: $PROJECT_ROOT_DIR"
@@ -26,63 +29,80 @@ if [ "$EUID" -ne 0 ]; then
exit 1
fi
# Generate the service file dynamically with the correct paths
echo "Generating service file with dynamic paths..."
WEB_SERVICE_FILE_CONTENT=$(cat <<EOF
[Unit]
Description=LED Matrix Web Interface Service
After=network-online.target
Wants=network-online.target
# Render the unit from systemd/ledmatrix-web.service. That template is the
# only description of the unit; this script used to carry its own heredoc copy,
# and install_service.sh a third, which is how the installed unit on real rigs
# ended up missing RestartSec and SyslogIdentifier while
# src/startup_validator.py warned about drift on every boot.
TEMPLATE="$PROJECT_ROOT_DIR/systemd/ledmatrix-web.service"
if [ ! -f "$TEMPLATE" ]; then
echo "ERROR: unit template not found at $TEMPLATE"
exit 1
fi
[Service]
Type=simple
User=${ACTUAL_USER}
WorkingDirectory=${PROJECT_ROOT_DIR}
Environment=USE_THREADING=1
ExecStart=/usr/bin/python3 ${PROJECT_ROOT_DIR}/scripts/utils/start_web_conditionally.py
Restart=on-failure
RestartSec=10
StandardOutput=syslog
StandardError=syslog
SyslogIdentifier=ledmatrix-web
# Automatically create and manage cache directory
CacheDirectory=ledmatrix
CacheDirectoryMode=0775
[Install]
WantedBy=multi-user.target
EOF
)
# Write the service file to systemd directory
echo "Writing service file to /etc/systemd/system/ledmatrix-web.service"
echo "$WEB_SERVICE_FILE_CONTENT" > /etc/systemd/system/ledmatrix-web.service
ESCAPED_PROJECT_ROOT_DIR=$(sed_escape_replacement "$PROJECT_ROOT_DIR")
ESCAPED_ACTUAL_USER=$(sed_escape_replacement "$ACTUAL_USER")
sed "s|__PROJECT_ROOT_DIR__|$ESCAPED_PROJECT_ROOT_DIR|g; s|__USER__|$ESCAPED_ACTUAL_USER|g" \
"$TEMPLATE" > /etc/systemd/system/ledmatrix-web.service
# Ensure cache directory exists with proper permissions
# This is a fallback for older systemd versions that don't support CacheDirectory
# Systemd 239+ will automatically create it via CacheDirectory directive
# Health check and rollback for the web UI's automatic updates. Its own unit so
# it survives the web service restart it performs; never enabled -- the web
# interface starts it after an update. Without it, automatic code updates
# stay paused rather than running with nothing to undo them.
for VERIFY_UNIT in ledmatrix-update-verify.service ledmatrix-update-verify.path; do
VERIFY_TEMPLATE="$PROJECT_ROOT_DIR/systemd/$VERIFY_UNIT"
if [ -f "$VERIFY_TEMPLATE" ]; then
echo "Writing unit file to /etc/systemd/system/$VERIFY_UNIT"
sed "s|__PROJECT_ROOT_DIR__|$ESCAPED_PROJECT_ROOT_DIR|g; s|__USER__|$ESCAPED_ACTUAL_USER|g" \
"$VERIFY_TEMPLATE" > "/etc/systemd/system/$VERIFY_UNIT"
chmod 644 "/etc/systemd/system/$VERIFY_UNIT"
else
echo "WARNING: $VERIFY_TEMPLATE not found; automatic code updates will stay paused"
fi
done
# Shared cache directory. The display service (root) and this web service both
# write here and read each other's files, which are created 0660, so the two
# share it through the directory's group: ledmatrix when the installing user
# is in it (first_time_install.sh / setup_cache.sh set that up), otherwise the
# user's own group. setgid makes new files inherit that group.
#
# An existing directory keeps its group whenever the web user can read through
# it -- ledmatrix, or the user's own group where systemd's old CacheDirectory=
# left it -- because re-grouping a working directory strands every file already
# in it on the old group. Only a group the user is not in (root's, or ledmatrix
# for a user outside it) is replaced. This used to force the user's group on
# every run, replacing the ledmatrix group setup_cache.sh had set.
echo "Setting up cache directory..."
CACHE_DIR="/var/cache/ledmatrix"
USER_GROUPS=$(id -nG "$ACTUAL_USER" 2>/dev/null | tr ' ' '\n')
if printf '%s\n' "$USER_GROUPS" | grep -qx ledmatrix; then
CACHE_GROUP="ledmatrix"
else
CACHE_GROUP=$(id -gn "$ACTUAL_USER" 2>/dev/null || echo root)
fi
if [ ! -d "$CACHE_DIR" ]; then
mkdir -p "$CACHE_DIR"
# Set group ownership to allow both root and web user access
# Try to use ACTUAL_USER's group, fallback to root if that fails
if getent group "$ACTUAL_USER" > /dev/null 2>&1; then
chown root:"$ACTUAL_USER" "$CACHE_DIR" 2>/dev/null || chown root:root "$CACHE_DIR"
else
chown root:root "$CACHE_DIR"
fi
chmod 775 "$CACHE_DIR"
chown root:"$CACHE_GROUP" "$CACHE_DIR" 2>/dev/null || true
echo "✓ Cache directory created: $CACHE_DIR"
else
# Ensure permissions are correct
chmod 775 "$CACHE_DIR" 2>/dev/null || true
# Try to set group ownership if possible
if getent group "$ACTUAL_USER" > /dev/null 2>&1; then
chown root:"$ACTUAL_USER" "$CACHE_DIR" 2>/dev/null || true
DIR_GROUP=$(stat -c %G "$CACHE_DIR" 2>/dev/null)
if ! printf '%s\n' "$USER_GROUPS" | grep -qx "$DIR_GROUP"; then
if chgrp "$CACHE_GROUP" "$CACHE_DIR" 2>/dev/null; then
echo "✓ Cache directory group changed from $DIR_GROUP to $CACHE_GROUP"
# Files already there keep the old group. The display service
# re-groups its own files when it starts (DiskCache.share_existing_files,
# which refuses symlinks and hard links); a recursive chgrp here
# would not. try-restart does nothing if the service is not running.
if find "$CACHE_DIR" -maxdepth 1 -name '*.json' -user root ! -group "$CACHE_GROUP" -print -quit 2>/dev/null | grep -q .; then
systemctl try-restart ledmatrix.service 2>/dev/null || true
fi
fi
fi
echo "✓ Cache directory exists: $CACHE_DIR"
fi
chmod 2775 "$CACHE_DIR" 2>/dev/null || true
# Reload systemd to recognize the new service
echo "Reloading systemd..."
@@ -92,6 +112,13 @@ systemctl daemon-reload
echo "Enabling ledmatrix-web.service..."
systemctl enable ledmatrix-web.service
# The path unit is what starts the health check after an automatic update.
if [ -f /etc/systemd/system/ledmatrix-update-verify.path ]; then
echo "Enabling ledmatrix-update-verify.path..."
systemctl enable --now ledmatrix-update-verify.path || \
echo "WARNING: could not enable ledmatrix-update-verify.path; automatic code updates will stay paused"
fi
# Start the service
echo "Starting ledmatrix-web.service..."
systemctl start ledmatrix-web.service
+13 -21
View File
@@ -18,6 +18,9 @@ USER_HOME=$(eval echo ~$ACTUAL_USER)
# Determine the Project Root Directory (parent of scripts/install/)
PROJECT_ROOT_DIR=$(cd "$(dirname "$0")/../.." && pwd)
# shellcheck source=scripts/install/lib_systemd_render.sh
source "$PROJECT_ROOT_DIR/scripts/install/lib_systemd_render.sh"
echo "Installing LED Matrix WiFi Monitor Service for user: $ACTUAL_USER"
echo "Using home directory: $USER_HOME"
echo "Project root directory: $PROJECT_ROOT_DIR"
@@ -64,30 +67,19 @@ if [ ${#MISSING_PACKAGES[@]} -gt 0 ]; then
echo "✓ Package installation completed"
fi
# Create service file with correct paths
# Render the unit from systemd/ledmatrix-wifi-monitor.service rather than
# inlining a second copy here. The copy this replaced had already drifted --
# it wrote StandardOutput/StandardError=syslog where the template says journal.
echo ""
echo "Creating systemd service file..."
SERVICE_FILE_CONTENT=$(cat <<EOF
[Unit]
Description=LED Matrix WiFi Monitor Daemon
After=network.target
Wants=network.target
TEMPLATE="$PROJECT_ROOT_DIR/systemd/ledmatrix-wifi-monitor.service"
if [ ! -f "$TEMPLATE" ]; then
echo "ERROR: unit template not found at $TEMPLATE"
exit 1
fi
[Service]
Type=simple
User=root
WorkingDirectory=$PROJECT_ROOT_DIR
ExecStart=/usr/bin/python3 $PROJECT_ROOT_DIR/scripts/utils/wifi_monitor_daemon.py --interval 30
Restart=on-failure
RestartSec=10
StandardOutput=syslog
StandardError=syslog
SyslogIdentifier=ledmatrix-wifi-monitor
[Install]
WantedBy=multi-user.target
EOF
)
ESCAPED_PROJECT_ROOT_DIR=$(sed_escape_replacement "$PROJECT_ROOT_DIR")
SERVICE_FILE_CONTENT=$(sed "s|__PROJECT_ROOT_DIR__|$ESCAPED_PROJECT_ROOT_DIR|g; s|__USER__|root|g" "$TEMPLATE")
if [ "$EUID" -eq 0 ]; then
echo "$SERVICE_FILE_CONTENT" | tee /etc/systemd/system/ledmatrix-wifi-monitor.service > /dev/null
+27
View File
@@ -0,0 +1,27 @@
#!/bin/bash
#
# Shared helper for rendering systemd unit templates via sed.
#
# Sourced by install_service.sh, install_web_service.sh and
# install_wifi_monitor.sh so all three escape sed replacement text the same
# way instead of carrying three copies of the same fix.
# sed_escape_replacement VALUE
#
# Print VALUE escaped for safe use as the replacement side of `sed
# s|pattern|replacement|`. Every one of these scripts builds its sed
# expression by interpolating a shell variable (a path, a username, ...)
# straight into the replacement text. sed gives three characters special
# meaning there: backslash (escape character), & (whole match) and the
# delimiter itself (here `|`). A value containing any of them -- e.g. a
# username or path with an `&`, a literal backslash, or a `|` -- would
# otherwise corrupt the rendered unit file instead of being substituted
# literally. Escape the backslash first so the later escapes aren't
# double-escaped.
sed_escape_replacement() {
local value="$1"
value="${value//\\/\\\\}"
value="${value//&/\\&}"
value="${value//|/\\|}"
printf '%s' "$value"
}
View File
+23 -13
View File
@@ -65,15 +65,14 @@ retry() {
local delay_seconds=5
local status
while true; do
# Run command in a context that disables errexit so we can capture exit code
# This prevents errexit from triggering before status=$? runs
if ! "$@"; then
status=$?
else
status=0
fi
if [ $status -eq 0 ]; then
# The condition of an if doesn't trip errexit, and in the else branch
# $? is the command's own exit status. (This used to be `if ! "$@";
# then status=$?`, where $? is the status of the negation -- always 0 --
# so a failure never retried and was reported as success.)
if "$@"; then
return 0
else
status=$?
fi
if [ $attempt -ge $max_attempts ]; then
print_error "Command failed after $attempt attempts: $*"
@@ -259,22 +258,32 @@ main() {
# Update package list first. first_time_install.sh is told the lists are
# already fresh so it does not repeat this a minute later.
# A refresh that still fails after retries (say one unreachable mirror)
# only warns: that is what this step effectively did before retry()
# could report a failure, and making it fatal would stop installs that
# work today.
if [ "$EUID" -eq 0 ]; then
retry apt-get update -qq
retry apt-get update -qq || print_warning "apt-get update failed; continuing with the existing package lists"
else
retry sudo apt-get update -qq
retry sudo apt-get update -qq || print_warning "apt-get update failed; continuing with the existing package lists"
fi
export LEDMATRIX_APT_UPDATED=1
# Install git and curl (needed for cloning and the script itself)
if ! command -v git >/dev/null 2>&1 || ! command -v curl >/dev/null 2>&1; then
print_warning "git or curl not found, installing..."
# Not fatal here, for the same reason: without git the clone below
# fails and stops the install with its own error.
if [ "$EUID" -eq 0 ]; then
retry apt-get install -y git curl
retry apt-get install -y git curl || true
else
retry sudo apt-get install -y git curl
retry sudo apt-get install -y git curl || true
fi
if command -v git >/dev/null 2>&1 && command -v curl >/dev/null 2>&1; then
print_success "git and curl installed"
else
print_warning "Could not install git and curl"
fi
print_success "git and curl installed"
else
print_success "git and curl already installed"
fi
@@ -408,6 +417,7 @@ main() {
# which would silently reinstate the duplicate apt update.
sudo -E env TMPDIR=/tmp LEDMATRIX_ASSUME_YES=1 \
LEDMATRIX_APT_UPDATED="${LEDMATRIX_APT_UPDATED:-0}" \
LEDMATRIX_AUTO_UPDATE="${LEDMATRIX_AUTO_UPDATE:-}" \
bash ./first_time_install.sh -y </dev/null
fi
INSTALL_EXIT_CODE=$?
Regular → Executable
View File
+53 -6
View File
@@ -3,6 +3,8 @@
# Use this if automatic dependency installation fails
set -e
# A failed pip must fail the `pip ... | tee` pipeline below, not be hidden by tee.
set -o pipefail
# Colors for output
RED='\033[0;31m'
@@ -28,15 +30,50 @@ echo ""
# Get the directory where this script is located
SCRIPT_DIR="$( cd "$( dirname "${BASH_SOURCE[0]}" )" && pwd )"
LEDMATRIX_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
PLUGINS_DIR="$LEDMATRIX_DIR/plugins"
CONFIG_FILE="$LEDMATRIX_DIR/config/config.json"
# The Plugin Store installs into plugin_system.plugins_directory from
# config.json (default plugin-repos), resolved against the project root like
# the display and web services do. plugins/ is also scanned: it holds the
# symlinks scripts/dev/dev_plugin_setup.sh creates.
CONFIGURED_DIR="plugin-repos"
if [ -f "$CONFIG_FILE" ] && command -v python3 >/dev/null 2>&1; then
CONFIGURED_DIR="$(LEDMATRIX_CONFIG_FILE="$CONFIG_FILE" python3 -c '
import json, os
try:
with open(os.environ["LEDMATRIX_CONFIG_FILE"], encoding="utf-8") as f:
value = (json.load(f).get("plugin_system") or {}).get("plugins_directory")
except Exception:
value = None
print(value if isinstance(value, str) and value.strip() else "plugin-repos")
' 2>/dev/null)" || CONFIGURED_DIR="plugin-repos"
[ -n "$CONFIGURED_DIR" ] || CONFIGURED_DIR="plugin-repos"
fi
case "$CONFIGURED_DIR" in
/*) PLUGINS_DIR="$CONFIGURED_DIR" ;;
*) PLUGINS_DIR="$LEDMATRIX_DIR/$CONFIGURED_DIR" ;;
esac
DEV_PLUGINS_DIR="$LEDMATRIX_DIR/plugins"
echo "LEDMatrix directory: $LEDMATRIX_DIR"
echo "Plugins directory: $PLUGINS_DIR"
echo "Plugins directory: $PLUGINS_DIR (plugin_system.plugins_directory)"
SCAN_DIRS=()
PLUGINS_DIR_REAL=""
if [ -d "$PLUGINS_DIR" ]; then
SCAN_DIRS+=("$PLUGINS_DIR")
PLUGINS_DIR_REAL="$(cd "$PLUGINS_DIR" && pwd -P)"
fi
if [ -d "$DEV_PLUGINS_DIR" ] && [ "$(cd "$DEV_PLUGINS_DIR" && pwd -P)" != "$PLUGINS_DIR_REAL" ]; then
echo "Also scanning dev plugins: $DEV_PLUGINS_DIR"
SCAN_DIRS+=("$DEV_PLUGINS_DIR")
fi
echo ""
# Check if plugins directory exists
if [ ! -d "$PLUGINS_DIR" ]; then
# Check if a plugins directory exists
if [ ${#SCAN_DIRS[@]} -eq 0 ]; then
echo -e "${RED}Error: Plugins directory not found at $PLUGINS_DIR${NC}"
echo "Install a plugin from the Plugin Store first, or check plugin_system.plugins_directory in $CONFIG_FILE"
exit 1
fi
@@ -47,12 +84,21 @@ echo ""
PLUGINS_FOUND=0
PLUGINS_INSTALLED=0
PLUGINS_FAILED=0
SEEN_PLUGIN_PATHS=" "
for plugin_dir in "$PLUGINS_DIR"/*/ ; do
for scan_dir in "${SCAN_DIRS[@]}"; do
for plugin_dir in "$scan_dir"/*/ ; do
if [ -d "$plugin_dir" ]; then
plugin_name=$(basename "$plugin_dir")
requirements_file="$plugin_dir/requirements.txt"
# A dev symlink can point at a plugin already scanned; install it once.
real_plugin_dir="$(cd "$plugin_dir" && pwd -P)"
case "$SEEN_PLUGIN_PATHS" in
*" $real_plugin_dir "*) continue ;;
esac
SEEN_PLUGIN_PATHS="$SEEN_PLUGIN_PATHS$real_plugin_dir "
if [ -f "$requirements_file" ]; then
PLUGINS_FOUND=$((PLUGINS_FOUND + 1))
echo -e "${GREEN}Found plugin: ${plugin_name}${NC}"
@@ -79,6 +125,7 @@ for plugin_dir in "$PLUGINS_DIR"/*/ ; do
fi
fi
done
done
# Summary
echo ""
+23 -1
View File
@@ -52,6 +52,10 @@ def main() -> int:
parser.add_argument('--height', type=int, default=32, help='Display height (default: 32)')
parser.add_argument('--skip-update', action='store_true',
help='Skip calling update() (render display only)')
parser.add_argument('--display-mode', default=None,
help='Display mode to render, for plugins that declare '
'more than one in their manifest (e.g. nrl_live). '
'Omitted, the plugin picks its own default.')
args = parser.parse_args()
@@ -141,8 +145,26 @@ def main() -> int:
except Exception as e:
logger.warning("update() raised: %s — continuing to display()", e)
# A plugin that declares several display modes usually renders nothing
# useful without being told which one to draw: the scoreboards keep their
# state on per-mode sub-managers and their no-argument path returns False.
# Only pass the argument when asked for, so the many plugins whose display()
# takes no display_mode keep working untouched.
try:
plugin_instance.display(force_clear=True)
if args.display_mode:
try:
plugin_instance.display(display_mode=args.display_mode,
force_clear=True)
except TypeError as error:
if ("unexpected keyword argument" not in str(error)
or "display_mode" not in str(error)):
raise
logger.warning(
"%s.display() does not accept display_mode; rendering its "
"default screen instead", args.plugin)
plugin_instance.display(force_clear=True)
else:
plugin_instance.display(force_clear=True)
logger.debug("display() completed")
except Exception as e:
logger.error("Error in display(): %s", e)
+98 -8
View File
@@ -6,6 +6,9 @@ Discovers and runs tests for LEDMatrix plugins.
Supports both unittest and pytest.
"""
import os
import re
import subprocess # nosec B404 - list-form argv only, no shell # nosemgrep
import sys
import argparse
from pathlib import Path
@@ -68,6 +71,79 @@ def _find_tests_in_dir(directory: Path) -> list:
return sorted(set(test_files))
def _is_script_style(path) -> bool:
"""True when a test file is a standalone script, not a pytest module.
Most plugin tests are written as `def main()` plus an `if __name__ ==
"__main__"` guard and signal through an exit code. pytest collects zero
items from those, so handing them to pytest printed "no tests ran" and this
runner reported success over work it had not done -- 151 of 248 files on a
fully populated rig.
"""
try:
src = Path(path).read_text(encoding="utf-8", errors="replace")
except OSError:
return False
has_pytest_items = re.search(r"^\s*(def test_|class Test|async def test_)", src, re.M)
has_main_guard = "__main__" in src and "__name__" in src
return bool(has_main_guard and not has_pytest_items)
def run_script_tests(test_files: list, verbose: bool = False) -> int:
"""Run standalone test scripts, honouring the 0 pass / 2 skip / 1 fail
convention that ledmatrix-plugins' own runner established.
Scripts opt into skipping by printing "SKIP: <reason>" and exiting 2 --
a script that needs a tty or an LED matrix is not a regression.
"""
env = dict(os.environ)
# Prepend rather than setdefault. An inherited PYTHONPATH -- a developer's
# shell, a tox run, another checkout -- otherwise wins outright, and the
# subprocess imports a different copy of the core than the one under test.
# That is exactly the failure ledmatrix-plugins#467 describes, and it is
# invisible: the tests pass or fail against a tree nobody meant to test.
inherited = env.get("PYTHONPATH")
env["PYTHONPATH"] = (f"{PROJECT_ROOT}{os.pathsep}{inherited}"
if inherited else str(PROJECT_ROOT))
env["LEDMATRIX_CORE"] = str(PROJECT_ROOT)
passed = skipped = failed = 0
failures = []
for path in test_files:
try:
# Fixed interpreter (sys.executable) plus a test path this script
# discovered by globbing the repo; argument list, no shell, so
# nothing is word-split or expanded. Same suppression pair the
# rest of the repo uses for this shape (see permission_utils.py).
proc = subprocess.run( # noqa: S603 # nosec B603 - no shell invoked (list-form argv) # nosemgrep
[sys.executable, str(path)], # nosemgrep
cwd=str(Path(path).parent),
capture_output=True, text=True, env=env,
stdin=subprocess.DEVNULL, timeout=300,
)
rc = proc.returncode
tail = " | ".join((proc.stdout or proc.stderr or "").strip().splitlines()[-2:])[:200]
except subprocess.TimeoutExpired:
rc, tail = 1, "timed out after 300s"
if rc == 0:
passed += 1
label = "pass"
elif rc == 2:
skipped += 1
label = "SKIP"
else:
failed += 1
label = "FAIL"
failures.append(f"{Path(path).name}: exit {rc} | {tail}")
if verbose or rc != 0:
print(f" [{label}] {Path(path).name}" + (f" -- {tail}" if rc != 0 else ""))
print(f"\n{passed} passed, {skipped} skipped, {failed} failed (scripts)")
for f in failures:
print(f" - {f}", file=sys.stderr)
return 1 if failed else 0
def run_unittest_tests(test_files: list, verbose: bool = False) -> int:
"""
Run tests using unittest.
@@ -186,11 +262,16 @@ def main():
print("No test files found in plugins directory")
return 0
print(f"Found {len(test_files)} test file(s)")
scripts = [f for f in test_files if _is_script_style(f)]
modules = [f for f in test_files if f not in scripts]
print(f"Found {len(test_files)} test file(s)"
+ (f" -- {len(modules)} collectable, {len(scripts)} standalone script(s)"
if scripts else ""))
for test_file in test_files:
print(f" - {test_file}")
print()
# Determine runner
runner = args.runner
if runner == 'auto':
@@ -199,12 +280,21 @@ def main():
runner = 'pytest'
except ImportError:
runner = 'unittest'
# Run tests
if runner == 'pytest':
return run_pytest_tests(test_files, args.verbose, args.coverage)
else:
return run_unittest_tests(test_files, args.verbose)
# Standalone scripts cannot be collected by pytest or unittest -- run them
# as the scripts they are. Doing this rather than silently collecting zero
# items is the whole point: this runner used to report success having
# executed nothing.
rc = 0
if scripts:
rc |= run_script_tests(scripts, args.verbose)
if modules:
if runner == 'pytest':
rc |= run_pytest_tests(modules, args.verbose, args.coverage)
else:
rc |= run_unittest_tests(modules, args.verbose)
return rc
if __name__ == '__main__':
+267
View File
@@ -0,0 +1,267 @@
#!/usr/bin/env python3
"""Show and try the scroll speeds your panel can display cleanly.
Motion looks smooth when the strip advances a WHOLE number of pixels per panel
refresh. Anything else has to blend two columns (which on pixel-font text reads
as shimmer) or repeat frames unevenly (which reads as judder). So the speeds
worth using are not arbitrary -- they are
refresh_hz / frame_hold * pixels_per_frame
for whole numbers of frame_hold and pixels_per_frame, and that ladder depends
on how fast YOUR panel actually refreshes. A Pi Zero driving a big chain will
have a completely different set of good speeds from a Pi 4 driving a small one.
# what can this panel do? (no hardware needed, uses your configured rate)
python3 scripts/scroll_speeds.py
# measure what the panel ACTUALLY manages, rather than what is configured
sudo systemctl stop ledmatrix
sudo python3 scripts/scroll_speeds.py --measure
sudo systemctl start ledmatrix
# what would a 60Hz panel offer?
python3 scripts/scroll_speeds.py --hz 60
# try one on the panel
sudo systemctl stop ledmatrix
sudo python3 scripts/scroll_speeds.py --demo 50
sudo systemctl start ledmatrix
This script never starts or stops the display service itself -- that is left to
you, so a crash here can never leave the panel dark.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
import time
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from src.common import scroll_config # noqa: E402
CONFIG = Path(__file__).resolve().parent.parent / "config" / "config.json"
def load_config():
"""The whole config.json, or {} when it is missing or unreadable."""
try:
with open(CONFIG, encoding="utf-8") as handle:
config = json.load(handle)
except (OSError, ValueError):
return {}
return config if isinstance(config, dict) else {}
def hardware_of(config):
return (config.get("display") or {}).get("hardware") or {}
def build_options(config, refresh_override=None):
"""The matrix options the display service would use for this config.
Built by DisplayManager.apply_matrix_options, not a copy of it, so the
measurement and the demo drive the panel exactly as the service does
(runtime gpio_slowdown, rp1_rio, panel_type, orientation, defaults).
``refresh_override`` replaces limit_refresh_rate_hz; 0 means uncapped.
"""
from src.display_manager import DisplayManager, RGBMatrixOptions
options = DisplayManager.apply_matrix_options(RGBMatrixOptions(), config)
if refresh_override is not None:
options.limit_refresh_rate_hz = int(refresh_override)
return options
def open_matrix(config, refresh_override=None):
"""Construct the matrix, or explain why it will not open."""
if os.geteuid() != 0:
sys.exit("this needs root for GPIO access - rerun with sudo")
try:
from src.display_manager import RGBMatrix
except ImportError as exc:
sys.exit("could not load the display stack ({}); is rgbmatrix "
"installed on this machine?".format(exc))
try:
return RGBMatrix(options=build_options(config, refresh_override))
except Exception as exc: # pragma: no cover - hardware dependent
sys.exit(
"could not open the panel ({}).\n"
"If the display service is running it owns the GPIO - stop it first:\n"
" sudo systemctl stop ledmatrix".format(exc)
)
def measure_refresh(config, seconds=6.0):
"""Actual refresh rate, by running uncapped and timing the swaps.
SwapOnVSync blocks until the panel's next refresh, so an unthrottled loop
runs at exactly the panel's rate. This is what an older Pi or a longer
chain will really give you, as opposed to whatever limit_refresh_rate_hz
optimistically asks for.
"""
matrix = open_matrix(config, refresh_override=0)
canvas = matrix.CreateFrameCanvas()
canvas = matrix.SwapOnVSync(canvas) # discard the first, it includes setup
frames = 0
started = time.perf_counter()
while time.perf_counter() - started < seconds:
canvas = matrix.SwapOnVSync(canvas)
frames += 1
measured = frames / (time.perf_counter() - started)
matrix.Clear()
return measured
def demo(config, target, seconds):
"""Scroll text at the crisp speed nearest `target`."""
from PIL import Image, ImageDraw, ImageFont
from src.common.font_layout import load_truetype
hz = scroll_config.refresh_hz_from_config(config)
choice = scroll_config.solve_crisp(target, hz)
print("asked for {:.0f} px/s -> {}".format(target, choice.describe()))
matrix = open_matrix(config)
canvas = matrix.CreateFrameCanvas()
W, H = canvas.width, canvas.height
font = None
for path, size in (
(str(Path(__file__).resolve().parent.parent / "assets/fonts/PressStart2P-Regular.ttf"), 16),
("/usr/share/fonts/truetype/dejavu/DejaVuSansMono-Bold.ttf", 26),
):
try:
font = load_truetype(path, size)
break
except OSError:
continue
if font is None:
font = ImageFont.load_default()
text = " {:.0f} px/s *** THE QUICK BROWN FOX JUMPS OVER THE LAZY DOG ***".format(
choice.pixels_per_second)
box = ImageDraw.Draw(Image.new("RGB", (8, 8))).textbbox((0, 0), text, font=font)
tw, th = box[2] - box[0], box[3] - box[1]
reps = max(2, (W * 3) // max(tw, 1) + 1)
strip = Image.new("RGB", (tw * reps, H), (0, 0, 0))
draw = ImageDraw.Draw(strip)
for i in range(reps):
draw.text((i * tw, (H - th) // 2 - box[1]), text, font=font, fill=(255, 210, 60))
offset = 0
frames = 0
started = time.time()
while time.time() - started < seconds:
window = strip.crop((offset, 0, offset + W, H))
if window.width < W:
whole = Image.new("RGB", (W, H), (0, 0, 0))
head = strip.crop((offset, 0, strip.width, H))
whole.paste(head, (0, 0))
whole.paste(strip.crop((0, 0, W - head.width, H)), (head.width, 0))
window = whole
canvas.SetImage(window)
canvas = matrix.SwapOnVSync(canvas, choice.frame_hold)
offset = (offset + choice.pixels_per_frame) % strip.width
frames += 1
elapsed = time.time() - started
print(" {} frames in {:.1f}s = {:.1f} fps = {:.1f} px/s actual".format(
frames, elapsed, frames / elapsed, frames * choice.pixels_per_frame / elapsed))
matrix.Clear()
def print_ladder(hz, highlight=None):
print("")
print("Whole-pixel scroll speeds at {:.1f}Hz refresh".format(hz))
print("(the panel refreshes at {:.0f}Hz for every one of these - holding a "
"frame costs no flicker)".format(hz))
print("")
for entry in scroll_config.crisp_ladder(hz):
if entry.pixels_per_second > hz * 3:
break
mark = " <-- nearest to {:.0f}".format(highlight) if (
highlight is not None
and entry.pixels_per_second == scroll_config.solve_crisp(highlight, hz).pixels_per_second
) else ""
print(" " + entry.describe() + mark)
print("")
print_config_advice(scroll_config.solve_crisp(highlight if highlight else hz / 2, hz))
def config_advice(choice):
"""The config that selects ``choice``, in the keys the resolver honours.
Tickers take a ``scroll_speed`` (px per step) + ``scroll_delay`` (seconds)
pair, and scroll_config ranks that pair ABOVE ``scroll_pixels_per_second``
-- deliberately, because some plugins give the flat key a schema default.
Many schemas default the pair too, so a flat key added by hand is usually
ignored. Advise the pair: pixels_per_frame every frame_hold/refresh
seconds is exactly the crisp speed.
"""
pair = {
"scroll_speed": choice.pixels_per_frame,
"scroll_delay": round(choice.frame_hold / choice.refresh_hz, 6),
}
scoreboard = {"scroll_speed": round(choice.pixels_per_second, 2)}
return pair, scoreboard
def print_config_advice(choice):
pair, scoreboard = config_advice(choice)
print("To use {:.1f} px/s, set it where the plugin keeps its scroll speed.".format(
choice.pixels_per_second))
print("Tickers take a scroll_speed (px per step) + scroll_delay (seconds) pair:")
print(' "display_options": {}'.format(json.dumps(pair)))
print("(some plugins keep the pair at the top level or under \"display\").")
print("The pair outranks scroll_pixels_per_second, which is ignored whenever the")
print("pair is present -- and schema defaults usually put it there.")
print("Sports scoreboards take pixels per second per league instead:")
print(' "scroll_settings": {}'.format(json.dumps(scoreboard)))
def main():
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--hz", type=float,
help="refresh rate to compute the ladder for (default: your config)")
ap.add_argument("--measure", action="store_true",
help="measure the panel's real refresh rate (needs root, service stopped)")
ap.add_argument("--demo", type=float, metavar="PXPS",
help="scroll text at the crisp speed nearest this (needs root)")
ap.add_argument("--seconds", type=float, default=15.0, help="demo duration")
ap.add_argument("--want", type=float, metavar="PXPS",
help="highlight the entry nearest this speed")
args = ap.parse_args()
config = load_config()
configured = float(hardware_of(config).get("limit_refresh_rate_hz") or 0)
if args.demo is not None:
demo(config, args.demo, args.seconds)
return
if args.measure:
measured = measure_refresh(config)
print("measured panel refresh: {:.1f}Hz".format(measured))
if configured:
print("configured limit_refresh_rate_hz: {:.0f}".format(configured))
if measured < configured * 0.95:
print(" -> the panel cannot reach the configured rate; the ladder")
print(" below uses what it actually manages")
print_ladder(measured, args.want)
return
hz = args.hz or configured or scroll_config.DEFAULT_REFRESH_HZ
if not args.hz and not configured:
print("no limit_refresh_rate_hz in config; assuming {:.0f}Hz".format(hz))
print("run with --measure to find your panel's real rate")
print_ladder(hz, args.want)
if __name__ == "__main__":
main()
+196
View File
@@ -0,0 +1,196 @@
#!/usr/bin/env python3
"""Drive a sports scoreboard scroll on the panel and report what it did.
The eight sports scoreboards scroll through ``src/common/sports_scroll.py``,
and that path is per-league opt-in: a rig showing static game cards never
constructs a SportsScrollDisplay at all, so nothing about its pacing can be
observed from a normal run. This drives it directly, with synthetic games, so
the pacing can be measured without changing anyone's configuration.
What it checks is what the shared resolver is supposed to buy:
* the requested speed lands on a whole number of pixels per refresh
* the frame hold that makes that true is published to the display manager
* frames actually arrive at the interval the hold implies
sudo systemctl stop ledmatrix
sudo python3 scripts/sports_scroll_check.py --seconds 20
sudo systemctl start ledmatrix
Like scripts/scroll_speeds.py, this never starts or stops the display service
itself -- that is left to the caller, so a crash here cannot leave the panel
dark.
"""
from __future__ import annotations
import argparse
import json
import statistics
import subprocess # nosec B404 - list-form argv only, no shell # nosemgrep
import sys
import time
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
from PIL import Image # noqa: E402
from src.common.sports_scroll import SportsScrollDisplay # noqa: E402
from src.display_manager import DisplayManager # noqa: E402
class _Check(SportsScrollDisplay):
"""A scoreboard whose cards are plain blocks -- pacing is what matters."""
SCROLL_LEAGUE_KEYS = ("nfl",)
def prepare_scroll_content(self, games, game_type, leagues, rankings_cache=None):
width = self.display_height * 2
cards = []
for i, _ in enumerate(games):
card = Image.new("RGB", (width, self.display_height), (0, 0, 0))
shade = 40 + (i * 37) % 180
for x in range(2, width - 2):
for y in range(2, self.display_height - 2):
card.putpixel((x, y), (shade, 90, 220 - shade // 2))
cards.append(card)
self._current_games = list(games)
self._current_game_type = game_type
self._current_leagues = list(leagues)
self.scroll_helper.create_scrolling_image(content_items=cards, item_gap=24)
return bool(cards)
class _HoldSpy:
"""Records what the scroll publishes, without changing what it does."""
def __init__(self, display_manager):
self.dm = display_manager
self.calls = []
self._real = display_manager.set_scrolling_state
def __enter__(self):
def spy(is_scrolling, frame_hold=1):
self.calls.append((is_scrolling, frame_hold))
return self._real(is_scrolling, frame_hold=frame_hold)
self.dm.set_scrolling_state = spy
return self
def __exit__(self, *exc):
self.dm.set_scrolling_state = self._real
return False
MESSAGE = """ledmatrix is running and owns the panel's GPIO.
Stop it first, or this run can leave the display dark:
sudo systemctl stop ledmatrix
sudo python3 scripts/sports_scroll_check.py
sudo systemctl start ledmatrix
Use --fallback to check the pacing logic without the panel, or --force if
you really mean it."""
def _refuse_if_the_service_is_running(force):
"""Refuse to touch the panel while ledmatrix has it.
rpi-rgb-led-matrix configures GPIO directions and the hardware PWM inside
RGBMatrix(), and on the root check it calls exit() from C -- no cleanup.
Do that while the service is driving those same pins and the panel goes
dark while the service carries on rendering happily: fresh framebuffer,
every pixel lit, "RGB Matrix initialized successfully", nothing in the log.
A restart brings it back, but only once you work out that is what happened.
The module docstring says to stop the service first. This makes it true.
"""
if force:
return
try:
active = subprocess.run( # nosec B603 B607 - hardcoded systemctl args # nosemgrep
["systemctl", "is-active", "ledmatrix"],
capture_output=True, text=True).stdout.strip()
except OSError:
return # not a systemd box; nothing to protect
if active == "active":
sys.exit(MESSAGE)
def main():
ap = argparse.ArgumentParser(
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--seconds", type=float, default=20.0)
ap.add_argument("--speed", type=float, default=None,
help="px/s to request; default is the module's own")
ap.add_argument("--games", type=int, default=6)
ap.add_argument("--force", action="store_true",
help="run even though the display service is up. It owns "
"the GPIO; expect a dark panel until you restart it.")
ap.add_argument("--fallback", action="store_true",
help="run without the panel. Driving the real matrix needs "
"root; this checks everything except the vsync pacing "
"-- what speed resolves to, that the hold is published, "
"and that it is released afterwards.")
args = ap.parse_args()
if not args.fallback:
_refuse_if_the_service_is_running(args.force)
root = Path(__file__).resolve().parent.parent
config = json.loads((root / "config" / "config.json").read_text(encoding="utf-8"))
display_manager = DisplayManager(config, force_fallback=args.fallback)
settings = {} if args.speed is None else {
"nfl": {"scroll_settings": {"scroll_speed": args.speed}}}
display = _Check(display_manager, settings, global_config=config)
resolved = display._scroll_settings
print("resolved: %s" % resolved.describe())
print("frame hold: %d refresh(es) per frame" % resolved.frame_hold)
if resolved.warning:
print("warning: %s" % resolved.warning)
display.prepare_scroll_content(
[{"id": "g%d" % i} for i in range(args.games)], "live", ["nfl"])
gaps, drawn = [], 0
last = None
with _HoldSpy(display_manager) as spy:
started = time.perf_counter()
while time.perf_counter() - started < args.seconds:
if not display.display_scroll_frame():
break
now = time.perf_counter()
if last is not None:
gaps.append((now - last) * 1000.0)
last = now
drawn += 1
display.clear()
if not gaps:
sys.exit("no frames were drawn -- the scroll never started")
gaps.sort()
expected = 1000.0 * resolved.frame_hold / (resolved.crisp.refresh_hz
if resolved.crisp else 100.0)
print("\n%d frames in %.1fs -> %.1f fps" % (
drawn, args.seconds, drawn / args.seconds))
print("frame gap median %.2fms p95 %.2fms max %.2fms (hold implies %.2fms)"
% (statistics.median(gaps), gaps[int(len(gaps) * 0.95)], gaps[-1], expected))
holds = {h for on, h in spy.calls if on}
print("published while scrolling: frame_hold=%s" % (sorted(holds) or "NOTHING"))
print("released on clear: %s" % any(not on for on, _ in spy.calls))
print("display manager hold now: %d (1 means released)"
% getattr(display_manager, "_frame_hold", -1))
if not holds:
sys.exit("FAIL: the scroll never told the core it was scrolling")
if holds != {resolved.frame_hold}:
sys.exit("FAIL: published %s but resolved %d" % (holds, resolved.frame_hold))
print("\nOK: the resolved hold reached the panel and was released after")
if __name__ == "__main__":
main()
+30
View File
@@ -9,6 +9,8 @@ This directory contains utility scripts for maintenance and system operations.
- **`wifi_monitor_daemon.py`** - Background daemon that monitors WiFi/Ethernet connection and manages access point mode
- **`cleanup_venv.sh`** - Cleans up Python virtual environment files
- **`clear_python_cache.sh`** - Clears Python cache files (__pycache__, *.pyc, etc.)
- **`pixlet_config_editor.sh`** - Opens Pixlet's own config UI for one installed Starlark app
- **`apply_dns_single_request.sh`** - Adds `options single-request` to the resolver (run by `ledmatrix-dns-fix.service`)
## Usage
@@ -25,3 +27,31 @@ This script is typically called by the systemd service (`ledmatrix-web.service`)
### WiFi Monitor Daemon
This daemon is typically run as a systemd service (`ledmatrix-wifi-monitor.service`) and automatically manages WiFi access point mode based on network connectivity.
### Pixlet Config Editor
Run it when you want Pixlet's own config form for a Starlark app -- live
render preview, cascading dropdowns -- rather than the LEDMatrix one.
```bash
./scripts/utils/pixlet_config_editor.sh # list installed apps
./scripts/utils/pixlet_config_editor.sh penndot_signs # edit, on localhost:8080
```
Deliberately not a service. It stops the display for the length of the
session and `pixlet serve` listens with no authentication, so it should only
be running while you are actually editing. It backs the config up first and
restarts the display on exit, however it exits.
It binds loopback only, with no flag to change that: anything that can reach
`pixlet serve` can rewrite the app's config, and a printed warning is not
access control. To edit from another machine, forward the port -- SSH does the
authenticating and nothing is left listening on the LAN:
```bash
ssh -L 8080:localhost:8080 pi@ledpi.local
```
### Apply DNS Single-Request Fix
Installed and run by `ledmatrix-dns-fix.service`; see `systemd/README.md`.
Safe to run by hand (`sudo ./scripts/utils/apply_dns_single_request.sh`) and
idempotent.
+94
View File
@@ -0,0 +1,94 @@
#!/bin/bash
#
# Add `options single-request` to the system resolver configuration.
#
# glibc's getaddrinfo() sends the A and AAAA queries for a name in
# parallel on one socket. Some routers answer the A query and drop the
# AAAA one, so the resolver waits out its full timeout -- about five
# seconds -- before returning an address that was already available.
# Disabling IPv6 in the kernel does not help: the resolver still asks.
#
# `single-request` makes it send the two queries one after the other,
# which those routers answer correctly. Anything on the matrix that
# calls an external API pays that five seconds per lookup otherwise, and
# a Starlark app with a render timeout will simply fail instead.
#
# Idempotent, and safe to run on a machine that does not need it. Run by
# ledmatrix-dns-fix.service on every boot, because whatever manages
# resolv.conf regenerates it and drops the option again.
#
# Usage: sudo ./scripts/utils/apply_dns_single_request.sh
set -eu
OPTION="options single-request"
RESOLVCONF_TAIL="/etc/resolvconf/resolv.conf.d/tail"
RESOLV_CONF="/etc/resolv.conf"
log() { echo "[dns-single-request] $*"; }
already_applied() {
grep -qs "^${OPTION}\$" "$1"
}
# resolvconf regenerates /etc/resolv.conf from these fragments, so the
# tail file is the only place an addition survives. Prefer it when the
# directory exists, whether or not resolvconf has run yet.
if [ -d "$(dirname "$RESOLVCONF_TAIL")" ]; then
if already_applied "$RESOLVCONF_TAIL"; then
log "already present in $RESOLVCONF_TAIL"
else
echo "$OPTION" >> "$RESOLVCONF_TAIL"
log "added to $RESOLVCONF_TAIL"
fi
# Only a missing resolvconf is ignorable. If it is present and the
# regeneration fails, /etc/resolv.conf still lacks the option, and
# reporting success would be a lie.
if command -v resolvconf >/dev/null 2>&1; then
if ! resolvconf -u; then
log "resolvconf -u failed; $RESOLV_CONF was not regenerated"
exit 1
fi
fi
fi
# systemd-resolved owns its stub file and rewrites anything appended to it,
# and `single-request` is a glibc resolv.conf option with no resolved.conf
# equivalent -- so there is nothing this script can do here. Exit non-zero:
# the unit would otherwise record success while the workaround is inactive,
# which is the failure mode this whole script exists to avoid.
if [ -L "$RESOLV_CONF" ] && readlink -f "$RESOLV_CONF" | grep -q "systemd"; then
log "$RESOLV_CONF is managed by systemd-resolved."
log "'options single-request' is a glibc resolv.conf option and has no"
log "resolved.conf equivalent, so it cannot be applied on this host."
log "If external API calls are slow, the workaround is to stop using the"
log "systemd-resolved stub (see 'man systemd-resolved', NSS/resolv.conf modes)."
exit 1
fi
if already_applied "$RESOLV_CONF"; then
log "already present in $RESOLV_CONF"
exit 0
fi
if [ ! -w "$RESOLV_CONF" ] && [ -e "$RESOLV_CONF" ]; then
log "cannot write $RESOLV_CONF (run with sudo?)"
exit 1
fi
# A NetworkManager-generated resolv.conf is regenerated on every connection
# change, not only at boot -- and this unit is oneshot with RemainAfterExit,
# so it will not re-run within the same boot to put the option back. Say so
# rather than implying the fix is permanent. Nothing is silently swallowed:
# the append below still happens and still works until the next renewal.
if grep -qs "Generated by NetworkManager" "$RESOLV_CONF" \
&& [ ! -d "$(dirname "$RESOLVCONF_TAIL")" ]; then
log "NOTE: $RESOLV_CONF is generated by NetworkManager and has no"
log "resolvconf tail directory to write to. The option is being added, but"
log "NetworkManager will drop it on the next connection renewal, and this"
log "unit does not run again until the next boot. If lookups go slow again"
log "before a reboot, re-run this script."
fi
echo "$OPTION" >> "$RESOLV_CONF"
log "added to $RESOLV_CONF"
+315
View File
@@ -0,0 +1,315 @@
#!/usr/bin/env python3
"""Check that an automatic LEDMatrix update left the device working; roll it back if not.
Started by the web interface's weekly updater (web_interface/auto_update.py)
through ledmatrix-update-verify.service, right after it pulls new code. It has
to run outside the web service: checking the update means restarting that
service, and a check running inside it would be killed by its own restart.
It must not be the code it is checking, either. The updater copies this file
to data/auto_update_verifier.py *before* pulling and the unit runs that copy,
so a broken update cannot break its own rollback. Standard library only for
the same reason: the rollback cannot depend on packages the update changed.
The updater leaves data/auto_update_pending.json:
{"status": "pending", "old_head": ..., "new_head": ...,
"display_was_active": bool, "dependency_failures": [...], "created_at": ...}
This moves its status to "verifying" and then to one of "success",
"rolled_back" or "rollback_failed", with "reason" and "detail" saying why.
The web interface reports that outcome and raises a banner for anything but
success.
"""
import json
import os
import subprocess # nosec B404 - list-form argv only, no shell # nosemgrep
import sys
from collections import namedtuple
import tempfile
import time
import traceback
import urllib.error
import urllib.request
from pathlib import Path
PENDING_NAME = 'auto_update_pending.json'
REQUIREMENT_FILES = ('requirements.txt', 'web_interface/requirements.txt')
WEB_HEALTH_URL = 'http://127.0.0.1:5000/api/v3/system/version'
#: How long the services get to come up after a restart...
HEALTH_TIMEOUT_SECONDS = 180
#: ...and how long they must then stay up. Restart=on-failure makes a crash
#: loop look healthy between attempts, so a single "is-active" proves nothing.
STABLE_SECONDS = 45
POLL_SECONDS = 5
WEB_CHECK_TIMEOUT_SECONDS = 5
SYSTEMCTL_QUERY_TIMEOUT_SECONDS = 10
RESTART_TIMEOUT_SECONDS = 90
GIT_TIMEOUT_SECONDS = 60
GIT_RESET_TIMEOUT_SECONDS = 120
PIP_TIMEOUT_SECONDS = 600
#: All of a rollback's dependency reinstalls together. A pip that times out
#: or fails is not retried: systemd stops this unit at TimeoutStartSec, and a
#: rollback killed half-way leaves the update reported as still verifying.
PIP_BUDGET_SECONDS = 600
#: sudoers matches the exact command line, so bash is named by path, the same
#: candidates src/common/permission_utils.install_requirements_file tries...
BASH_CANDIDATES = ('/usr/bin/bash', '/bin/bash')
#: ...and, like it, moves to the next one only when sudo refused the command
#: line (permission_utils.SUDO_REFUSAL_PHRASES), never after pip itself ran.
SUDO_REFUSAL_PHRASES = ('a password is required', 'is not allowed to run', 'no tty present')
#: The longest one health check can take: restart and wait, roll back
#: (diff, reset, reinstalls), restart and wait again. A wait's last poll can
#: start just before its deadline and run every query to its timeout.
_WAIT_WORST_SECONDS = (HEALTH_TIMEOUT_SECONDS + STABLE_SECONDS + WEB_CHECK_TIMEOUT_SECONDS
+ 2 * SYSTEMCTL_QUERY_TIMEOUT_SECONDS + POLL_SECONDS)
WORST_CASE_SECONDS = (2 * (2 * RESTART_TIMEOUT_SECONDS + _WAIT_WORST_SECONDS)
+ GIT_TIMEOUT_SECONDS + GIT_RESET_TIMEOUT_SECONDS + PIP_BUDGET_SECONDS)
#: What a command that could not run at all reports: its callers only read
#: these three fields, the same ones a completed subprocess has.
_Failed = namedtuple('_Failed', 'returncode stdout stderr')
def pending_path(project_root):
return Path(project_root) / 'data' / PENDING_NAME
def read_pending(path):
try:
with open(path, 'r', encoding='utf-8') as f:
data = json.load(f)
return data if isinstance(data, dict) else None
except (OSError, ValueError):
return None
def write_pending(path, data):
path = Path(path)
path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp = tempfile.mkstemp(dir=str(path.parent), prefix='.auto_update_pending_')
try:
with os.fdopen(fd, 'w', encoding='utf-8') as f:
json.dump(data, f, indent=2)
os.replace(tmp, path)
except BaseException:
try:
os.unlink(tmp)
except OSError:
pass
raise
def _web_responds(url=WEB_HEALTH_URL):
try:
with urllib.request.urlopen(url, timeout=WEB_CHECK_TIMEOUT_SECONDS) as resp: # nosec B310 - fixed loopback URL
return resp.status == 200
except (urllib.error.URLError, OSError, ValueError):
return False
def _short(sha):
return (sha or 'unknown')[:7]
class Verifier:
def __init__(self, project_root, run=subprocess.run, sleep=time.sleep,
clock=time.monotonic, web_responds=_web_responds, log=None):
self.project_root = Path(project_root)
self.pending_file = pending_path(project_root)
self.run = run
self.sleep = sleep
self.clock = clock
self.web_responds = web_responds
self.log = log or (lambda msg: print(f'[auto-update-verify] {msg}', flush=True))
def _run(self, args, timeout=GIT_TIMEOUT_SECONDS):
try:
return self.run(args, cwd=str(self.project_root), capture_output=True,
text=True, timeout=timeout)
except (subprocess.SubprocessError, OSError) as e:
return _Failed(returncode=1, stdout='', stderr=str(e))
# -- services ---------------------------------------------------------
def service_active(self, unit):
return self._run(['systemctl', 'is-active', unit],
timeout=SYSTEMCTL_QUERY_TIMEOUT_SECONDS).stdout.strip() == 'active'
def restart_count(self, unit):
out = self._run(['systemctl', 'show', '-p', 'NRestarts', '--value', unit],
timeout=SYSTEMCTL_QUERY_TIMEOUT_SECONDS).stdout.strip()
return int(out) if out.isdigit() else None
def restart(self, unit):
result = self._run(['sudo', '-n', 'systemctl', 'restart', f'{unit}.service'],
timeout=RESTART_TIMEOUT_SECONDS)
if result.returncode != 0:
self.log(f'restarting {unit} failed: {(result.stderr or "").strip()}')
return result.returncode == 0
def restart_services(self, display):
"""Restart what should be running. False if any restart command failed."""
ok = True
# A display the user had stopped stays stopped.
if display:
ok = self.restart('ledmatrix') and ok
return self.restart('ledmatrix-web') and ok
def wait_healthy(self, display):
"""None once the services are up and stay up, else what went wrong."""
deadline = self.clock() + HEALTH_TIMEOUT_SECONDS + STABLE_SECONDS
healthy_since = baseline = None
web = disp = False
count_known = True
while self.clock() < deadline:
web = self.web_responds()
disp = self.service_active('ledmatrix') if display else True
restarts = self.restart_count('ledmatrix') if display else None
# Without a restart count a crash loop looks healthy between
# attempts, so an unreadable count never counts as stable.
count_known = not display or restarts is not None
if web and disp and count_known and (healthy_since is None or restarts == baseline):
if healthy_since is None:
healthy_since, baseline = self.clock(), restarts
elif self.clock() - healthy_since >= STABLE_SECONDS:
return None
else:
healthy_since = None
self.sleep(POLL_SECONDS)
problems = []
if not web:
problems.append('the web interface did not respond')
if not disp:
problems.append('the display service did not stay running')
if web and disp and not count_known:
problems.append("the display service's restart count could not be read")
return '; '.join(problems) or 'the display service kept restarting'
# -- rollback ---------------------------------------------------------
def changed_requirements(self, old, new):
result = self._run(['git', 'diff', '--name-only', old, new])
# If the diff is unavailable, reinstall both rather than guess.
changed = set(result.stdout.split()) if result.returncode == 0 else set(REQUIREMENT_FILES)
return [rel for rel in REQUIREMENT_FILES if rel in changed]
def install_requirements(self, rel, deadline=None):
"""Install one requirements file through the root wrapper, by ``deadline``."""
wrapper = self.project_root / 'scripts' / 'fix_perms' / 'safe_pip_install.sh'
req = self.project_root / rel
if not req.exists():
return True
for bash in BASH_CANDIDATES:
timeout = PIP_TIMEOUT_SECONDS
if deadline is not None:
timeout = min(timeout, deadline - self.clock())
if timeout <= 0:
self.log(f'no time left to reinstall {rel}')
return False
result = self._run(['sudo', '-n', bash, str(wrapper), str(req)], timeout=timeout)
if result.returncode == 0:
return True
# Only a refused command line is worth the next candidate. A pip
# that ran and failed, or timed out, would just do it again.
if not any(phrase in (result.stderr or '') for phrase in SUDO_REFUSAL_PHRASES):
# Not pip's output: it can echo an index URL's credentials.
self.log(f'reinstalling {rel} failed (exit {result.returncode})')
return False
return False
def rollback(self, pending):
"""Reset to the previous commit and its dependencies. Returns (ok, detail)."""
old, new = pending.get('old_head'), pending.get('new_head')
if not old:
return False, 'the commit to roll back to is unknown'
requirements = self.changed_requirements(old, new) if new else list(REQUIREMENT_FILES)
# --hard: the updater refuses to run with local edits to tracked core
# files (web_interface/auto_update.local_changes), so outside the
# plugin folders the only thing this discards is the update. Edits
# under plugins/ and plugin-repos/, which that check leaves to the
# pull's --autostash, are reset along with it.
result = self._run(['git', 'reset', '--hard', old], timeout=GIT_RESET_TIMEOUT_SECONDS)
if result.returncode != 0:
return False, (f'"git reset --hard {old}" failed: '
f'{(result.stderr or result.stdout or "").strip()}')
deadline = self.clock() + PIP_BUDGET_SECONDS
failed = [rel for rel in requirements if not self.install_requirements(rel, deadline)]
if failed:
return True, ('reinstalling the previous dependencies from ' + ', '.join(failed)
+ ' failed; run Install Base Requirements from the Tools tab')
return True, ''
# -- the check itself -------------------------------------------------
def _finish(self, pending, status, reason=None, detail=None):
pending.update({'status': status, 'reason': reason, 'detail': detail or None,
'finished_at': time.time()})
write_pending(self.pending_file, pending)
self.log(' '.join(p for p in (status, reason or '', detail or '') if p))
def verify(self):
pending = read_pending(self.pending_file)
if not pending or pending.get('status') != 'pending':
self.log('no update is waiting to be verified')
return 0
pending['status'] = 'verifying'
write_pending(self.pending_file, pending)
display = bool(pending.get('display_was_active'))
dependency_failures = pending.get('dependency_failures') or []
if dependency_failures:
# Never restart onto code whose packages did not install.
reason = 'installing its dependencies failed (' + ', '.join(dependency_failures) + ')'
elif not self.restart_services(display):
# The old process may still be answering; checking it would pass
# an update that never started.
reason = 'restarting the services failed'
else:
reason = self.wait_healthy(display)
if reason is None:
self._finish(pending, 'success')
return 0
self.log(f'update to {_short(pending.get("new_head"))} is unhealthy ({reason}); '
f'rolling back to {_short(pending.get("old_head"))}')
ok, detail = self.rollback(pending)
if not ok:
self._finish(pending, 'rollback_failed', reason, detail)
return 1
still = (self.wait_healthy(display) if self.restart_services(display)
else 'restarting the services failed')
if still:
self._finish(pending, 'rollback_failed', reason,
f'still unhealthy after rolling back: {still}'
+ (f'; {detail}' if detail else ''))
return 1
self._finish(pending, 'rolled_back', reason, detail)
return 0
def main(argv):
if len(argv) != 2:
print('usage: auto_update_verify.py PROJECT_ROOT', file=sys.stderr)
return 2
verifier = Verifier(Path(argv[1]))
try:
return verifier.verify()
except Exception as e:
traceback.print_exc()
# Whatever happened, the web interface must not be left thinking the
# check is still running.
try:
pending = read_pending(verifier.pending_file) or {}
if pending.get('status') in ('pending', 'verifying'):
pending.update({'status': 'rollback_failed', 'reason': 'the health check crashed',
'detail': str(e), 'finished_at': time.time()})
write_pending(verifier.pending_file, pending)
except OSError:
pass
return 1
if __name__ == '__main__':
sys.exit(main(sys.argv))
Regular → Executable
View File
View File
+223
View File
@@ -0,0 +1,223 @@
#!/bin/bash
#
# Edit an installed Starlark app's config in Pixlet's own config UI.
#
# `pixlet serve` runs the app for real, so its form has working cascading
# dropdowns and option lists fetched live -- useful for an app whose choices
# only exist at runtime, or when you want to see the render change as you
# type. The LEDMatrix config form now reads the same runtime schema (see
# PixletRenderer.extract_schema_via_pixlet), so reach for this when you want
# Pixlet's live preview, not because the normal form is missing options.
#
# Deliberately a script you run and then Ctrl+C, not a service: it stops the
# display for the length of the session, and `pixlet serve` listens on a port
# with no authentication. Nothing here should be listening when you are not
# actually editing.
#
# Usage:
# ./scripts/utils/pixlet_config_editor.sh # list installed apps
# ./scripts/utils/pixlet_config_editor.sh <app_id> # edit
#
# Binds the LAN by default, matching the web interface, which already serves
# 0.0.0.0:5000 with no authentication -- anything that can reach this can
# already reconfigure the display there. `pixlet serve` has no authentication
# either, so treat both the same way: fine on a home network, not on an open
# one. Override the bind and the session length with:
#
# PIXLET_EDITOR_HOST=127.0.0.1 ./scripts/utils/pixlet_config_editor.sh <app>
# PIXLET_EDITOR_TIMEOUT=600 ./scripts/utils/pixlet_config_editor.sh <app>
#
# For loopback-only editing from another machine, forward the port instead:
#
# ssh -L 8080:localhost:8080 pi@ledpi.local
#
# The session always ends by itself after PIXLET_EDITOR_TIMEOUT seconds
# (default 30 minutes). The display is stopped while editing, so a session
# left open would otherwise leave the panel dark indefinitely -- the timeout
# is what makes it safe to start one from the web interface.
set -eu
PROJECT_ROOT_DIR=$(cd "$(dirname "$0")/../.." && pwd)
APPS_DIR="$PROJECT_ROOT_DIR/starlark-apps"
PORT="${PIXLET_EDITOR_PORT:-8080}"
# LAN by default; see the header for why, and how to force loopback.
BIND_HOST="${PIXLET_EDITOR_HOST:-0.0.0.0}"
# Hard stop, so the display cannot be left off by a forgotten session.
EDITOR_TIMEOUT="${PIXLET_EDITOR_TIMEOUT:-1800}"
APP_ID="${1:-}"
list_apps() {
if [ -d "$APPS_DIR" ]; then
find "$APPS_DIR" -maxdepth 1 -mindepth 1 -type d -printf ' %f\n' 2>/dev/null | sort
fi
}
if [ -z "$APP_ID" ]; then
echo "Usage: $0 <app_id>"
echo ""
echo "Installed apps:"
list_apps || true
[ -n "$(list_apps)" ] || echo " (none found in $APPS_DIR)"
exit 1
fi
APP_DIR="$APPS_DIR/$APP_ID"
if [ ! -d "$APP_DIR" ]; then
echo "No such app: $APP_ID"
echo ""
echo "Installed apps:"
list_apps
exit 1
fi
STAR_FILE=$(find "$APP_DIR" -maxdepth 1 -iname "*.star" | head -1)
if [ -z "$STAR_FILE" ]; then
echo "No .star file found in $APP_DIR"
exit 1
fi
# Same search order the plugin itself uses: the bundled binary for this
# architecture first, then PATH -- so this works on an install that never put
# pixlet on PATH.
find_pixlet() {
local arch bundled
case "$(uname -s)-$(uname -m)" in
Linux-aarch64|Linux-arm64) arch="pixlet-linux-arm64" ;;
Linux-x86_64|Linux-amd64) arch="pixlet-linux-amd64" ;;
Darwin-arm64) arch="pixlet-darwin-arm64" ;;
Darwin-x86_64) arch="pixlet-darwin-amd64" ;;
*) arch="" ;;
esac
bundled="$PROJECT_ROOT_DIR/bin/pixlet/$arch"
if [ -n "$arch" ] && [ -x "$bundled" ]; then
echo "$bundled"
return 0
fi
command -v pixlet 2>/dev/null || return 1
}
PIXLET_BIN=$(find_pixlet) || {
echo "Pixlet not found. Install it with:"
echo " ./scripts/download_pixlet.sh"
exit 1
}
# find_pixlet supports Darwin, so this script has to as well. macOS ships no
# timeout(1); GNU coreutils installs it as gtimeout. Resolve whichever exists
# and fail here with instructions rather than at the invocation far below,
# where the failure would land after the display has already been stopped.
find_timeout() {
local candidate
for candidate in timeout gtimeout; do
if command -v "$candidate" >/dev/null 2>&1; then
command -v "$candidate"
return 0
fi
done
return 1
}
TIMEOUT_BIN=$(find_timeout) || {
echo "Neither 'timeout' nor 'gtimeout' was found on PATH."
echo "This script needs one to bound the editing session."
echo "On macOS, install GNU coreutils:"
echo " brew install coreutils"
exit 1
}
CONFIG_FILE="$APP_DIR/config.json"
if [ -f "$CONFIG_FILE" ]; then
cp "$CONFIG_FILE" "$CONFIG_FILE.backup"
echo "Backed up existing config to $CONFIG_FILE.backup"
else
echo "{}" > "$CONFIG_FILE"
fi
DISPLAY_WAS_RUNNING=false
if systemctl is-active --quiet ledmatrix 2>/dev/null; then
DISPLAY_WAS_RUNNING=true
fi
# Restart the display however this exits -- Ctrl+C, an error, or pixlet
# dying on its own. Leaving the panel dark because the editor crashed is the
# failure worth guarding against.
cleanup() {
echo ""
# Kill the serve child explicitly. `timeout` is started with --foreground so
# it shares this script's process group (without that it makes its own, and
# a group signal aimed at this script would orphan pixlet with the port
# still bound). Belt and braces: signal the recorded pid too, because a
# group signal only reaches it while the group is shared.
if [ -n "${SERVE_PID:-}" ] && kill -0 "$SERVE_PID" 2>/dev/null; then
kill -TERM "$SERVE_PID" 2>/dev/null || true
for _ in 1 2 3 4 5 6 7 8 9 10; do
kill -0 "$SERVE_PID" 2>/dev/null || break
sleep 0.3
done
kill -KILL "$SERVE_PID" 2>/dev/null || true
fi
if [ "$DISPLAY_WAS_RUNNING" = true ]; then
echo "Restarting the display service..."
sudo systemctl restart ledmatrix || echo "⚠ Could not restart ledmatrix - do it by hand"
fi
echo "Your config as it was before this session: $CONFIG_FILE.backup"
}
trap cleanup EXIT INT TERM
if [ "$DISPLAY_WAS_RUNNING" = true ]; then
echo "Stopping the display service so it does not read config.json mid-write..."
sudo systemctl stop ledmatrix
fi
# Wildcard, loopback and an explicit interface address are three different
# cases. Collapsing the last two into "localhost" printed a URL pointing at the
# user's own machine whenever PIXLET_EDITOR_HOST named a LAN address.
case "$BIND_HOST" in
0.0.0.0|::|"") REACH_HOST="$(hostname).local" ;;
127.0.0.1|::1|localhost) REACH_HOST="localhost" ;;
*) REACH_HOST="$BIND_HOST" ;;
esac
echo ""
echo "Editing: $APP_ID"
echo "App file: $STAR_FILE"
echo "URL: http://$REACH_HOST:$PORT/"
echo ""
if [ "$BIND_HOST" = "0.0.0.0" ]; then
echo "Reachable on the LAN, and pixlet serve has no authentication -- the"
echo "same footing as the web interface on port 5000. Set"
echo "PIXLET_EDITOR_HOST=127.0.0.1 to keep it to this machine."
else
echo "Listening on $BIND_HOST only. From another machine, forward the port:"
echo " ssh -L $PORT:localhost:$PORT $(whoami)@$(hostname)"
fi
echo ""
echo "Changes save straight to the real config as you make them."
echo "Press Ctrl+C when finished - the display restarts automatically."
echo "This session stops on its own after ${EDITOR_TIMEOUT}s regardless."
echo ""
cd "$APP_DIR"
# `timeout` owns the hard stop rather than the caller: the trap above restarts
# the display however this exits, so a session that outlives the person who
# started it still gives the panel back. Exit 124 is timeout's own code for
# "expired", which is a normal end here, not a failure.
# --foreground: stay in this script's process group so one signal reaches the
# whole session. Backgrounded + `wait` so the EXIT trap can run while the child
# is still alive; a foreground child would leave bash waiting on it instead.
"$TIMEOUT_BIN" --foreground "$EDITOR_TIMEOUT" "$PIXLET_BIN" serve "$(basename "$STAR_FILE")" \
--host "$BIND_HOST" \
--port "$PORT" \
--no-browser \
--saveconfig "$CONFIG_FILE" &
SERVE_PID=$!
status=0
wait "$SERVE_PID" || status=$?
if [ "$status" -eq 124 ]; then
echo "Session reached its ${EDITOR_TIMEOUT}s limit."
status=0
fi
exit "$status"
+30 -12
View File
@@ -74,23 +74,41 @@ def install_dependencies():
print(f"Failed to install dependencies: {e}")
return False
#: String spellings that turn autostart OFF. Anything else -- including the key
#: being absent entirely -- leaves it on.
DISABLED_STRINGS = ("off", "false", "no", "0")
def autostart_enabled(config_data):
"""Whether to bring the web interface up. Defaults to True.
config.template.json and first_time_install.sh both ship
``web_display_autostart`` as true, so a config that lacks the key is an
older or hand-edited one rather than a request to stay down. Defaulting to
False meant any such config silently got no web interface -- and because
the "not starting" path exits 0, systemd reported the unit as successfully
started while nothing was listening. Only an explicit false/off disables it.
"""
value = config_data.get("web_display_autostart", True)
if isinstance(value, str):
return value.strip().lower() not in DISABLED_STRINGS
return bool(value)
def main():
try:
with open(CONFIG_FILE, 'r') as f:
config_data = json.load(f)
except FileNotFoundError:
print(f"Config file {CONFIG_FILE} not found. Web interface will not start.")
sys.exit(0) # Exit gracefully, don't start
except Exception as e:
print(f"Error reading config file {CONFIG_FILE}: {e}. Web interface will not start.")
sys.exit(1) # Exit with error, service might restart depending on config
# The web interface is how a config gets created and repaired, so a
# missing one is the case where the user needs it most.
print(f"Config file {CONFIG_FILE} not found. Starting the web interface so it can be configured.")
config_data = {}
except (json.JSONDecodeError, OSError) as e:
print(f"Error reading config file {CONFIG_FILE}: {e}. Starting the web interface anyway so the config can be repaired.")
config_data = {}
autostart_enabled = config_data.get("web_display_autostart", False)
# Handle both boolean True and string "on"/"true" values
is_enabled = (autostart_enabled is True) or (isinstance(autostart_enabled, str) and autostart_enabled.lower() in ("on", "true", "yes", "1"))
if is_enabled:
if autostart_enabled(config_data):
print("Configuration 'web_display_autostart' is enabled. Starting web interface...")
# Only install dependencies if not already done during first-time setup
@@ -116,7 +134,7 @@ def main():
print(f"Failed to exec web interface: {e}")
sys.exit(1) # Failed to start
else:
print("Configuration 'web_display_autostart' is false or not set. Web interface will not be started.")
print("Configuration 'web_display_autostart' is explicitly disabled. Web interface will not be started.")
sys.exit(0) # Exit gracefully, service considered successful
if __name__ == '__main__':
+8 -6
View File
@@ -27,6 +27,8 @@ sys.path.insert(0, str(PROJECT_ROOT))
from PIL import Image, ImageDraw, ImageFont # noqa: E402
from src.common.font_layout import load_truetype # noqa: E402
FIXTURES_DIR = PROJECT_ROOT / "src" / "skin_system" / "fixtures"
MODES = ("live", "recent", "upcoming")
SPORTS = ("baseball", "basketball", "football", "hockey")
@@ -52,12 +54,12 @@ class FixtureHost:
try:
press = str(PROJECT_ROOT / "assets/fonts/PressStart2P-Regular.ttf")
small = str(PROJECT_ROOT / "assets/fonts/4x6-font.ttf")
fonts['score'] = ImageFont.truetype(press, 10)
fonts['time'] = ImageFont.truetype(press, 8)
fonts['team'] = ImageFont.truetype(press, 8)
fonts['status'] = ImageFont.truetype(small, 6)
fonts['detail'] = ImageFont.truetype(small, 6)
fonts['rank'] = ImageFont.truetype(press, 10)
fonts['score'] = load_truetype(press, 10)
fonts['time'] = load_truetype(press, 8)
fonts['team'] = load_truetype(press, 8)
fonts['status'] = load_truetype(small, 6)
fonts['detail'] = load_truetype(small, 6)
fonts['rank'] = load_truetype(press, 10)
except IOError:
default = ImageFont.load_default()
for key in ('score', 'time', 'team', 'status', 'detail', 'rank'):
+14 -9
View File
@@ -121,19 +121,24 @@ fi
echo ""
# 5. Check web interface
# ledmatrix-web.service runs scripts/utils/start_web_conditionally.py, which
# starts web_interface/start.py; the app binds port 5000 (web_interface/start.py).
echo "=== Web Interface ==="
if [ -f "$PROJECT_ROOT/web_interface_v2.py" ]; then
check_pass "web_interface_v2.py exists"
else
check_fail "web_interface_v2.py is missing"
fi
WEB_PORT=5000
for web_file in scripts/utils/start_web_conditionally.py web_interface/start.py web_interface/app.py; do
if [ -f "$PROJECT_ROOT/$web_file" ]; then
check_pass "$web_file exists"
else
check_fail "$web_file is missing"
fi
done
# Check if web service is listening
if systemctl is-active --quiet ledmatrix-web.service 2>/dev/null; then
if netstat -tuln 2>/dev/null | grep -q ":5001" || ss -tuln 2>/dev/null | grep -q ":5001"; then
check_pass "Web interface is listening on port 5001"
if netstat -tuln 2>/dev/null | grep -qE ":${WEB_PORT}([^0-9]|$)" || ss -tuln 2>/dev/null | grep -qE ":${WEB_PORT}([^0-9]|$)"; then
check_pass "Web interface is listening on port $WEB_PORT"
else
check_warn "Web service is running but port 5001 may not be listening"
check_warn "Web service is running but port $WEB_PORT may not be listening"
fi
else
check_warn "Web service is not running (cannot check port)"
@@ -204,7 +209,7 @@ if [ "$ALL_PASSED" = true ]; then
echo -e "${GREEN}Installation verification PASSED${NC}"
echo ""
echo "Next steps:"
echo "1. Access the web interface at: http://$(hostname -I | awk '{print $1}'):5001"
echo "1. Access the web interface at: http://$(hostname -I | awk '{print $1}'):$WEB_PORT"
echo "2. Check service status: sudo systemctl status ledmatrix.service"
echo "3. View logs: journalctl -u ledmatrix.service -f"
exit 0
+25 -21
View File
@@ -8,6 +8,10 @@ echo "Web UI Verification"
echo "=========================================="
echo ""
# The web interface binds port 5000 (web_interface/start.py).
WEB_PORT=5000
PORT_PATTERN=":${WEB_PORT}([^0-9]|$)"
# Colors
GREEN='\033[0;32m'
RED='\033[0;31m'
@@ -32,25 +36,25 @@ else
fi
echo ""
# 2. Check if port 5001 is listening
echo "2. Checking if port 5001 is listening..."
# 2. Check if port $WEB_PORT is listening
echo "2. Checking if port $WEB_PORT is listening..."
if command -v ss >/dev/null 2>&1; then
if ss -tuln 2>/dev/null | grep -q ":5001"; then
echo -e "${GREEN}✓${NC} Port 5001 is listening"
if ss -tuln 2>/dev/null | grep -qE "$PORT_PATTERN"; then
echo -e "${GREEN}✓${NC} Port $WEB_PORT is listening"
echo ""
echo "Active connections on port 5001:"
ss -tuln | grep ":5001"
echo "Active connections on port $WEB_PORT:"
ss -tuln | grep -E "$PORT_PATTERN"
else
echo -e "${RED}✗${NC} Port 5001 is NOT listening"
echo -e "${RED}✗${NC} Port $WEB_PORT is NOT listening"
fi
elif command -v netstat >/dev/null 2>&1; then
if netstat -tuln 2>/dev/null | grep -q ":5001"; then
echo -e "${GREEN}✓${NC} Port 5001 is listening"
if netstat -tuln 2>/dev/null | grep -qE "$PORT_PATTERN"; then
echo -e "${GREEN}✓${NC} Port $WEB_PORT is listening"
echo ""
echo "Active connections on port 5001:"
netstat -tuln | grep ":5001"
echo "Active connections on port $WEB_PORT:"
netstat -tuln | grep -E "$PORT_PATTERN"
else
echo -e "${RED}✗${NC} Port 5001 is NOT listening"
echo -e "${RED}✗${NC} Port $WEB_PORT is NOT listening"
fi
else
echo -e "${YELLOW}⚠${NC} Cannot check port (ss/netstat not available)"
@@ -59,15 +63,15 @@ echo ""
# 3. Test HTTP connection
echo "3. Testing HTTP connection..."
if curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:5001 > /dev/null 2>&1; then
HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:5001 2>/dev/null)
if curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:$WEB_PORT > /dev/null 2>&1; then
HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:$WEB_PORT 2>/dev/null)
if [ "$HTTP_CODE" = "200" ] || [ "$HTTP_CODE" = "302" ] || [ "$HTTP_CODE" = "301" ]; then
echo -e "${GREEN}✓${NC} Web interface is responding (HTTP $HTTP_CODE)"
else
echo -e "${YELLOW}⚠${NC} Web interface responded with HTTP $HTTP_CODE"
fi
else
echo -e "${RED}✗${NC} Cannot connect to web interface on port 5001"
echo -e "${RED}✗${NC} Cannot connect to web interface on port $WEB_PORT"
fi
echo ""
@@ -79,7 +83,7 @@ if [ -n "$IP_ADDRESSES" ]; then
echo ""
echo "Access web interface at:"
for ip in $IP_ADDRESSES; do
echo " http://$ip:5001"
echo " http://$ip:$WEB_PORT"
done
else
echo -e "${YELLOW}⚠${NC} Could not determine IP address"
@@ -118,12 +122,12 @@ if systemctl is-active --quiet ledmatrix-web.service 2>/dev/null; then
SERVICE_RUNNING=true
fi
if (ss -tuln 2>/dev/null | grep -q ":5001") || (netstat -tuln 2>/dev/null | grep -q ":5001"); then
if (ss -tuln 2>/dev/null | grep -qE "$PORT_PATTERN") || (netstat -tuln 2>/dev/null | grep -qE "$PORT_PATTERN"); then
PORT_LISTENING=true
fi
if curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:5001 > /dev/null 2>&1; then
HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:5001 2>/dev/null)
if curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:$WEB_PORT > /dev/null 2>&1; then
HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" --max-time 5 http://localhost:$WEB_PORT 2>/dev/null)
if [ "$HTTP_CODE" = "200" ] || [ "$HTTP_CODE" = "302" ] || [ "$HTTP_CODE" = "301" ]; then
HTTP_RESPONDING=true
fi
@@ -134,7 +138,7 @@ if [ "$SERVICE_RUNNING" = true ] && [ "$PORT_LISTENING" = true ] && [ "$HTTP_RES
echo ""
echo "You can access it at:"
for ip in $IP_ADDRESSES; do
echo " http://$ip:5001"
echo " http://$ip:$WEB_PORT"
done
exit 0
elif [ "$SERVICE_RUNNING" = false ]; then
@@ -145,7 +149,7 @@ elif [ "$SERVICE_RUNNING" = false ]; then
echo " sudo systemctl enable ledmatrix-web.service # to start on boot"
exit 1
elif [ "$PORT_LISTENING" = false ]; then
echo -e "${RED}✗ Service is running but port 5001 is not listening${NC}"
echo -e "${RED}✗ Service is running but port $WEB_PORT is not listening${NC}"
echo ""
echo "Check logs for errors:"
echo " sudo journalctl -u ledmatrix-web.service -f"
+8 -3
View File
@@ -1,5 +1,10 @@
# skins/
> **Not supported yet.** The current scoreboard plugins don't render skins,
> so a skin placed here and selected in config has no effect, and the web UI
> and Plugin Store don't offer them. See
> [docs/SKIN_SYSTEM.md](../docs/SKIN_SYSTEM.md#status-not-supported-yet).
User-installable **visual skins** for the sports scoreboards. Each
subdirectory is one skin:
@@ -10,10 +15,10 @@ skins/<skin-id>/
preview.png # optional
```
- Install a skin: `git clone <skin repo> skins/<skin-id>` (or via the Plugin
Store for registry entries with `"type": "skin"`).
- Install a skin: `git clone <skin repo> skins/<skin-id>`. The Plugin Store
refuses registry entries with `"type": "skin"` while skins don't render.
- Select it: set `"skin": "<skin-id>"` in the plugin's section of
`config/config.json`, or use the web UI's Visual Skin dropdown.
`config/config.json`. The web UI no longer shows a Visual Skin dropdown.
- Build one: start from `example-classic-baseball/` and read
[docs/CREATING_SKINS.md](../docs/CREATING_SKINS.md). Validate with
`python scripts/validate_skin.py --skin <skin-id>`.
+1 -1
View File
@@ -4,5 +4,5 @@ LEDMatrix Display System
Core source package for the LED Matrix Display project.
"""
__version__ = "3.3.0"
__version__ = "3.5.0"
+221
View File
@@ -0,0 +1,221 @@
"""Install the automatic-update health check, from the display service.
The weekly updater (web_interface/auto_update.py) will not update LEDMatrix
code unless ledmatrix-update-verify.path and .service are installed: they
restart the services after an update and roll it back if the device is
unhealthy. Installing units takes root and the web interface is not root, and
"SSH in and run an installer" means most people never get updates with a
safety net.
The display service already runs this repository's code as root, so it
installs them -- but only while the user has automatic updates turned on, only
these two units, rendered from the repository's templates for the web
interface's own user, and it reports what happened in
data/auto_update_setup.json for the General tab. It grants nothing new: the
units run as the web user, who can already change the code this process runs.
Called at display startup; the web interface restarts the display service
when the toggle is switched on, so setup happens straight away. Refreshing a
unit whose template changed happens the same way, which is why this compares
content rather than only checking that the files exist.
"""
import json
import logging
import os
import re
import subprocess # nosec B404 - list-form argv only, no shell # nosemgrep
import tempfile
import time
from pathlib import Path
logger = logging.getLogger(__name__)
PROJECT_ROOT = Path(__file__).resolve().parent.parent
SYSTEMD_DIR = Path('/etc/systemd/system')
SERVICE_UNIT = 'ledmatrix-update-verify.service'
PATH_UNIT = 'ledmatrix-update-verify.path'
UNITS = (SERVICE_UNIT, PATH_UNIT)
WEB_UNIT = 'ledmatrix-web.service'
RESULT_REL = Path('data') / 'auto_update_setup.json'
_USER_RE = re.compile(r'^[a-z_][a-z0-9_-]{0,31}$')
class SetupError(Exception):
"""A reason setup cannot proceed, worded for the General tab."""
def _is_root():
return hasattr(os, 'geteuid') and os.geteuid() == 0
def _lookup_ids(user):
try:
import pwd
entry = pwd.getpwnam(user)
return entry.pw_uid, entry.pw_gid
except (ImportError, KeyError):
return None
def _directive(text, key):
match = re.search(rf'^{key}=(.*)$', text or '', re.M)
return match.group(1).strip() if match else None
def _read(path):
try:
return Path(path).read_text(encoding='utf-8')
except OSError:
return None
def is_enabled(config):
return bool((config.get('auto_update') or {}).get('enabled', False))
class UpdateHelperSetup:
def __init__(self, project_root=PROJECT_ROOT, systemd_dir=SYSTEMD_DIR, run=subprocess.run,
is_root=_is_root, lookup_ids=_lookup_ids, clock=time.time):
self.project_root = Path(project_root)
self.systemd_dir = Path(systemd_dir)
self.run = run
self.is_root = is_root
self.lookup_ids = lookup_ids
self.clock = clock
self.result_file = self.project_root / RESULT_REL
self._web_ids = None
def _systemctl(self, *args):
return self.run(['systemctl', *args], capture_output=True, text=True, timeout=60)
def _check(self, result, what):
if result.returncode != 0:
raise SetupError(f'"{what}" failed: {(result.stderr or result.stdout or "").strip()}')
def path_active(self):
try:
return self._systemctl('is-active', PATH_UNIT).stdout.strip() == 'active'
except (subprocess.SubprocessError, OSError):
return False
def ensure(self, config):
"""Install or refresh the units while automatic updates are on.
Returns the result recorded for the General tab, or None when there
was nothing to do (updates off, or not a systemd host at all).
"""
if not is_enabled(config) or not self.systemd_dir.is_dir():
return None
try:
changed = self._install()
except SetupError as e:
return self._report('failed', str(e))
except (OSError, subprocess.SubprocessError) as e:
return self._report('failed', f'Could not install the update health check: {e}')
if changed:
return self._report('installed', 'Installed the update health check.')
return self._report('installed', 'The update health check is installed.', quiet=True)
def _install(self):
if not self.is_root():
raise SetupError('The display service is not running as root, so it cannot install the '
'update health check. Run "sudo ./scripts/install/install_web_service.sh" once.')
web_text = _read(self.systemd_dir / WEB_UNIT)
if web_text is None:
raise SetupError('The web interface service (ledmatrix-web.service) is not installed.')
user = _directive(web_text, 'User') or 'root'
ids = self.lookup_ids(user) if _USER_RE.match(user) else None
if ids is None:
raise SetupError(f'The web interface runs as "{user}", which is not a usable account.')
self._web_ids = ids
workdir = _directive(web_text, 'WorkingDirectory')
if not workdir or Path(workdir).resolve() != self.project_root.resolve():
raise SetupError(f'The web interface service runs from {workdir or "an unknown folder"}, '
f'not {self.project_root}.')
# Spaces are fine -- the templates quote every command-line path --
# but systemd expands % specifiers, and a quote, backslash or line
# break would be reinterpreted in a unit file. (On Windows, where the
# tests also run, a backslash is the path separator, not a name.)
root_text = str(self.project_root)
unsafe = set('%"') | ({'\\'} if os.sep == '/' else set())
if any(ch in unsafe or ord(ch) < 32 for ch in root_text):
raise SetupError(f'LEDMatrix is installed in {root_text!r}, a folder name systemd cannot use '
'in a unit file. Move it to a path without %, quotes, backslashes or '
'control characters.')
rendered = {}
for name in UNITS:
template = _read(self.project_root / 'systemd' / name)
if template is None:
raise SetupError(f'The unit template systemd/{name} is missing.')
rendered[name] = (template.replace('__PROJECT_ROOT_DIR__', str(self.project_root))
.replace('__USER__', user))
# The templates are ordinary repository files. Whatever they say,
# this root process only installs a service that runs as the web user
# and a path unit that starts exactly that service.
if _directive(rendered[SERVICE_UNIT], 'User') != user:
raise SetupError(f'systemd/{SERVICE_UNIT} does not run as the web interface user; '
'refusing to install it.')
if _directive(rendered[PATH_UNIT], 'Unit') != SERVICE_UNIT:
raise SetupError(f'systemd/{PATH_UNIT} does not start {SERVICE_UNIT}; refusing to install it.')
changed = [name for name in UNITS if _read(self.systemd_dir / name) != rendered[name]]
for name in changed:
self._write_unit(self.systemd_dir / name, rendered[name])
if changed:
self._check(self._systemctl('daemon-reload'), 'systemctl daemon-reload')
self._check(self._systemctl('enable', PATH_UNIT), f'systemctl enable {PATH_UNIT}')
self._check(self._systemctl('restart', PATH_UNIT), f'systemctl restart {PATH_UNIT}')
elif not self.path_active():
self._check(self._systemctl('enable', '--now', PATH_UNIT), f'systemctl enable --now {PATH_UNIT}')
changed = [PATH_UNIT]
if not self.path_active():
raise SetupError(f'{PATH_UNIT} did not start; see "journalctl -u {PATH_UNIT}".')
return bool(changed)
def _write_unit(self, path, text):
fd, tmp = tempfile.mkstemp(dir=str(path.parent), prefix=f'.{path.name}.')
try:
with os.fdopen(fd, 'w', encoding='utf-8') as f:
f.write(text)
os.chmod(tmp, 0o644)
os.replace(tmp, path)
except BaseException:
try:
os.unlink(tmp)
except OSError:
pass
raise
def _report(self, status, message, quiet=False):
previous = None
try:
previous = json.loads(_read(self.result_file) or 'null')
except ValueError:
pass
if quiet and isinstance(previous, dict) and previous.get('status') == status:
return previous # nothing new; don't rewrite it on every boot
result = {'status': status, 'message': message, 'at': self.clock()}
(logger.info if status == 'installed' else logger.warning)("Automatic update setup: %s", message)
try:
self.result_file.parent.mkdir(parents=True, exist_ok=True)
fd, tmp = tempfile.mkstemp(dir=str(self.result_file.parent), prefix='.auto_update_setup_')
with os.fdopen(fd, 'w', encoding='utf-8') as f:
json.dump(result, f, indent=2)
os.chmod(tmp, 0o644)
if self._web_ids and hasattr(os, 'chown'):
# Only root can give the file away; the result is readable
# (0644) either way, so a failed chown must not lose it.
try:
os.chown(tmp, *self._web_ids)
except OSError:
pass
os.replace(tmp, self.result_file)
except OSError as e:
logger.warning("Could not record automatic update setup result: %s", e)
return result
def ensure_update_helper(config):
return UpdateHelperSetup().ensure(config)
+70 -8
View File
@@ -25,6 +25,14 @@ from enum import Enum
import queue
from concurrent.futures import ThreadPoolExecutor
from src.cache_manager import CacheManager
from src.common.espn_dates import (
RANGE_RETRY_SECONDS,
_note_range_rejected,
_ranges_known_rejected,
clamp_espn_limit,
fetch_espn_date_chunks,
parse_espn_date_range,
)
# Configure logging
logger = logging.getLogger(__name__)
@@ -98,7 +106,12 @@ class BackgroundDataService:
This service manages a pool of background threads to fetch data asynchronously,
with intelligent caching, retry logic, and progress tracking.
"""
# Plugins feature-detect this. A core without it sends season ranges to
# ESPN as-is and gets 400s since 2026-09-15, so plugins fetch those
# ranges themselves instead of submitting them here.
handles_espn_date_ranges = True
def __init__(self, cache_manager: CacheManager, max_workers: int = 3, request_timeout: int = 30):
"""
Initialize the background data service.
@@ -247,6 +260,12 @@ class BackgroundDataService:
logger.debug(f"Cache hit for {sport} {year} data")
return request_id
# limit above 500 makes an ESPN *scoreboard* return a truncated list
# (src/common/espn_dates.py). Other endpoints need more: /teams has 762
# college-football teams, so only scoreboards are clamped.
if url.split('?', 1)[0].rstrip('/').endswith('/scoreboard'):
params = clamp_espn_limit(params)
# Create fetch request
request = FetchRequest(
id=request_id,
@@ -254,7 +273,7 @@ class BackgroundDataService:
year=year,
cache_key=cache_key,
url=url,
params=params or {},
params=dict(params or {}),
headers={**self.default_headers, **(headers or {})},
timeout=timeout or self.request_timeout,
max_retries=max_retries,
@@ -338,12 +357,38 @@ class BackgroundDataService:
logger.info(f"Starting background fetch for {request.sport} {request.year}")
# Perform HTTP request with retry logic
response = self._make_request_with_retry(request)
response.raise_for_status()
# Parse response
data = response.json()
# ESPN stopped accepting dates=YYYYMMDD-YYYYMMDD on 2026-09-15 and
# answers 400 for every sport. Re-ask in months and days rather
# than let a whole season fail. See src/common/espn_dates.py.
# The "ranges are rejected" memo is shared with
# fetch_espn_scoreboard(): once either path has seen a range
# rejected, the other skips the doomed range request too.
is_range = parse_espn_date_range(request.params.get("dates")) is not None
data = None
chunks_tried = False
if is_range and _ranges_known_rejected():
data = self._fetch_in_date_chunks(request)
# Every chunk failed: ask for the range itself below so the
# failure carries a real HTTP error, without re-spending chunks.
chunks_tried = data is None
if data is None:
# Perform HTTP request with retry logic
response = self._make_request_with_retry(request)
if is_range and response.status_code == 400 and not chunks_tried:
_note_range_rejected()
logger.warning(
"ESPN rejected the date range %s (400); fetching it as "
"month/day chunks, and fetching ranges that way for the "
"next %d hours",
request.params.get("dates"), RANGE_RETRY_SECONDS // 3600,
)
data = self._fetch_in_date_chunks(request)
if data is None:
response.raise_for_status()
else:
response.raise_for_status()
data = response.json()
# Validate data structure
if not isinstance(data, dict):
@@ -519,6 +564,23 @@ class BackgroundDataService:
"""
result.data = None
def _fetch_in_date_chunks(self, request: FetchRequest) -> Optional[Dict[str, Any]]:
"""Re-fetch a rejected ``YYYYMMDD-YYYYMMDD`` range as month/day chunks.
None means the request was not a day range, or every chunk failed; the
caller then re-raises the original 400 instead of caching an empty
season. See src/common/espn_dates.py.
"""
logger.info("Recovering %s %s from a rejected date range", request.sport, request.year)
return fetch_espn_date_chunks(
self.session,
request.url,
params=request.params,
headers=request.headers,
timeout=request.timeout,
logger=logger,
)
def _make_request_with_retry(self, request: FetchRequest) -> requests.Response:
"""
Make HTTP request with retry logic and exponential backoff.
+4 -1
View File
@@ -530,12 +530,15 @@ def _copy_file(src: Path, dst: Path) -> None:
os.chmod(tmp_path, existing_mode)
else:
shutil.copymode(src, tmp_path)
if existing_owner is not None:
if existing_owner is not None and hasattr(os, 'chown'):
# Replacing a file creates a new inode owned by whoever is running,
# which would silently move a root-owned config to the web user.
# Carry the previous owner across when the OS permits it — only
# root can hand a file to another user, so this is best-effort and
# a plain restore as the web user simply keeps its own ownership.
# os.chown does not exist on Windows (where st_uid/st_gid are just
# 0); looking it up there raises AttributeError, which no caller
# catches, so every restore over an existing file aborted.
try:
os.chown(tmp_path, existing_owner[0], existing_owner[1])
except (OSError, PermissionError):
-45
View File
@@ -37,51 +37,6 @@ class Baseball(SportsCore):
self.data_source = ESPNDataSource(logger)
self.sport = "baseball"
def _get_baseball_display_text(self, game: Dict) -> str:
"""Get baseball-specific display text."""
try:
display_parts = []
# Inning information
if self.show_innings:
inning = game.get("inning", "")
if inning:
display_parts.append(f"Inning: {inning}")
# Outs information
if self.show_outs:
outs = game.get("outs", 0)
if outs is not None:
display_parts.append(f"Outs: {outs}")
# Bases information
if self.show_bases:
bases = game.get("bases", "")
if bases:
display_parts.append(f"Bases: {bases}")
# Count information
if self.show_count:
strikes = game.get("strikes", 0)
balls = game.get("balls", 0)
if strikes is not None and balls is not None:
display_parts.append(f"Count: {balls}-{strikes}")
# Pitcher/Batter information
if self.show_pitcher_batter:
pitcher = game.get("pitcher", "")
batter = game.get("batter", "")
if pitcher:
display_parts.append(f"Pitcher: {pitcher}")
if batter:
display_parts.append(f"Batter: {batter}")
return " | ".join(display_parts) if display_parts else ""
except Exception as e:
self.logger.error(f"Error getting baseball display text: {e}")
return ""
def _is_baseball_game_live(self, game: Dict) -> bool:
"""Check if a baseball game is currently live."""
try:
+70 -36
View File
@@ -10,6 +10,7 @@ from typing import Dict, List
import requests
import logging
from datetime import datetime
from src.common.espn_dates import fetch_espn_scoreboard
class DataSource(ABC):
"""Abstract base class for data sources."""
@@ -71,10 +72,10 @@ class ESPNDataSource(DataSource):
now = datetime.now()
formatted_date = now.strftime("%Y%m%d")
url = f"{self.base_url}/{sport}/{league}/scoreboard"
response = self.session.get(url, params={"dates": formatted_date, "limit": 1000}, headers=self.get_headers(), timeout=15)
response.raise_for_status()
data = response.json()
data = fetch_espn_scoreboard(
self.session, url, params={"dates": formatted_date, "limit": 1000},
headers=self.get_headers(), timeout=15, logger=self.logger,
)
events = data.get('events', [])
# Filter for live games
@@ -99,10 +100,10 @@ class ESPNDataSource(DataSource):
"limit": 1000
}
response = self.session.get(url, headers=self.get_headers(), params=params, timeout=15)
response.raise_for_status()
data = response.json()
data = fetch_espn_scoreboard(
self.session, url, params=params,
headers=self.get_headers(), timeout=15, logger=self.logger,
)
events = data.get('events', [])
self.logger.debug(f"Fetched {len(events)} scheduled games for {sport}/{league}")
@@ -113,35 +114,68 @@ class ESPNDataSource(DataSource):
return []
def fetch_standings(self, sport: str, league: str) -> Dict:
"""Fetch standings from ESPN API."""
# Try standings endpoint first (for professional leagues like NFL, NBA, etc.)
try:
url = f"{self.base_url}/{sport}/{league}/standings"
response = self.session.get(url, headers=self.get_headers(), timeout=15)
response.raise_for_status()
data = response.json()
self.logger.debug(f"Fetched standings for {sport}/{league}")
"""Fetch standings, or the poll for leagues that have one.
Order matters and used to be wrong. College leagues publish a poll at
/rankings and a records table at /standings; professional leagues have
only /standings. The old code tried /standings first and fell back to
/rankings only on a 404 -- but college /standings answers 200, so the
fallback never fired and college rankings came back empty forever.
Nothing failed; the AP rank badge simply never appeared, and anything
else keyed off rankings quietly did nothing.
A 200 that lacks the key is treated as a miss, so a league answering
both endpoints still ends up with whichever one actually carries a poll.
"""
league_name = (league or "").lower()
wants_poll = "college" in league_name or "ncaa" in league_name
endpoints = ["rankings", "standings"] if wants_poll else ["standings", "rankings"]
for endpoint in endpoints:
url = f"{self.base_url}/{sport}/{league}/{endpoint}"
# Only the request is guarded. Inspecting the payload happens
# below, outside the handler, so that a bug in this method cannot
# be mistaken for an endpoint that failed -- that mistake would
# silently drop rankings for a league that has them, which is the
# exact failure this function was written to fix.
try:
response = self.session.get(
url, headers=self.get_headers(), timeout=15
)
response.raise_for_status()
data = response.json()
except (requests.RequestException, ValueError) as e:
status = getattr(getattr(e, "response", None), "status_code", None)
# Only a 404 is routine -- it is how a league says "no poll
# here". Everything else is worth an error, and `status is
# None` covers the ones that matter most: ConnectionError,
# Timeout, a body that would not parse. Silencing those left a
# board that could not reach ESPN with one debug line, and the
# ranked filter running on an empty table.
if status != 404:
self.logger.error(
f"Error fetching {endpoint} from ESPN for "
f"{sport}/{league}: {e}"
)
continue
if not isinstance(data, dict):
# A list or a bare string is not something the callers can
# read. Treat it as a miss so the other endpoint still gets a
# turn, but say so -- this means ESPN changed shape.
self.logger.error(
f"Unexpected {endpoint} payload for {sport}/{league}: "
f"got {type(data).__name__}, expected an object"
)
continue
if endpoint == "rankings" and not data.get("rankings"):
continue
self.logger.debug(f"Fetched {endpoint} for {sport}/{league}")
return data
except Exception as e:
# If standings doesn't exist, try rankings (for college sports)
if hasattr(e, 'response') and hasattr(e.response, 'status_code') and e.response.status_code == 404:
try:
url = f"{self.base_url}/{sport}/{league}/rankings"
response = self.session.get(url, headers=self.get_headers(), timeout=15)
response.raise_for_status()
data = response.json()
self.logger.debug(f"Fetched rankings for {sport}/{league}")
return data
except Exception:
# Both endpoints failed - standings/rankings may not be available for this sport/league
self.logger.debug(f"Standings/rankings not available for {sport}/{league} from ESPN API")
return {}
else:
# Non-404 error - log at debug level since standings are optional
self.logger.debug(f"Error fetching standings from ESPN for {sport}/{league}: {e}")
return {}
self.logger.debug(
f"Standings/rankings not available for {sport}/{league} from ESPN API"
)
return {}
class MLBAPIDataSource(DataSource):
+38 -17
View File
@@ -8,13 +8,16 @@ import os
import tempfile
import time
from abc import ABC, abstractmethod
from collections import OrderedDict
from datetime import datetime, timedelta
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
import pytz
from src.common.espn_dates import fetch_espn_scoreboard
import requests
from PIL import Image, ImageDraw, ImageFont
from src.common.font_layout import load_truetype
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
@@ -166,7 +169,14 @@ class SportsCore(ABC):
self.session.mount("https://", adapter)
self.session.mount("http://", adapter)
self._logo_cache = {}
# LRU-bounded: entries are decoded RGBA thumbnails, not file bytes.
# Each is up to display_width*1.5 x display_height*1.5 -- about 36KB on
# a 256x64 panel, more for wide wordmarks. The key is a team
# abbreviation and assets/sports/ncaa_logos alone ships 307 of them, so
# an unbounded dict here held the whole league: ~11-18MB per manager
# instance, and a league runs three (live/recent/upcoming) each with
# its own cache. That is real money on a 1GB Pi.
self._logo_cache: "OrderedDict[str, Image.Image]" = OrderedDict()
# Font caches for _load_custom_font_from_element_config: per-frame
# callers (font-ladder walks) resolve the same (name, size) over and
@@ -449,12 +459,12 @@ class SportsCore(ABC):
press_start = self._resolve_font_path("PressStart2P-Regular.ttf")
four_by_six = self._resolve_font_path("4x6-font.ttf")
try:
fonts['score'] = ImageFont.truetype(press_start, 10)
fonts['time'] = ImageFont.truetype(press_start, 8)
fonts['team'] = ImageFont.truetype(press_start, 8)
fonts['status'] = ImageFont.truetype(four_by_six, 6) # Using 4x6 for status
fonts['detail'] = ImageFont.truetype(four_by_six, 6) # Added detail font
fonts['rank'] = ImageFont.truetype(press_start, 10)
fonts['score'] = load_truetype(press_start, 10)
fonts['time'] = load_truetype(press_start, 8)
fonts['team'] = load_truetype(press_start, 8)
fonts['status'] = load_truetype(four_by_six, 6) # Using 4x6 for status
fonts['detail'] = load_truetype(four_by_six, 6) # Added detail font
fonts['rank'] = load_truetype(press_start, 10)
self.logger.info("Successfully loaded fonts")
except OSError:
# Name the directory we searched: the usual cause is an install
@@ -559,11 +569,17 @@ class SportsCore(ABC):
draw.text((x + dx, y + dy), text, font=font, fill=outline_color)
draw.text((x, y), text, font=font, fill=fill)
#: Decoded logos to keep. A scroll of "other games" shows on the order of
#: 20 games (40 teams), so this holds a full cycle without thrashing while
#: capping the cache well below a 307-team league.
_LOGO_CACHE_MAX = 64
def _load_and_resize_logo(self, team_id: str, team_abbrev: str, logo_path: Path, logo_url: str | None ) -> Optional[Image.Image]:
"""Load and resize a team logo, with caching and automatic download if missing."""
self.logger.debug(f"Logo path: {logo_path}")
if team_abbrev in self._logo_cache:
self.logger.debug(f"Using cached logo for {team_abbrev}")
self._logo_cache.move_to_end(team_abbrev)
return self._logo_cache[team_abbrev]
try:
@@ -603,6 +619,8 @@ class SportsCore(ABC):
max_height = int(self.display_height * 1.5)
logo.thumbnail((max_width, max_height), Image.Resampling.LANCZOS)
self._logo_cache[team_abbrev] = logo
while len(self._logo_cache) > self._LOGO_CACHE_MAX:
self._logo_cache.popitem(last=False)
return logo
except Exception as e:
@@ -827,9 +845,11 @@ class SportsCore(ABC):
formatted_date_yesterday = yesterday.strftime("%Y%m%d")
# Fetch todays games only
url = f"https://site.api.espn.com/apis/site/v2/sports/{self.sport}/{self.league}/scoreboard"
response = self.session.get(url, params={"dates": f"{formatted_date_yesterday}-{formatted_date}", "limit": 1000}, headers=self.headers, timeout=10)
response.raise_for_status()
data = response.json()
data = fetch_espn_scoreboard(
self.session, url,
params={"dates": f"{formatted_date_yesterday}-{formatted_date}", "limit": 1000},
headers=self.headers, timeout=10, logger=self.logger,
)
events = data.get('events', [])
self.logger.info(f"Fetched {len(events)} todays games for {self.sport} - {self.league}")
@@ -852,9 +872,10 @@ class SportsCore(ABC):
end_date = now + timedelta(weeks=1)
date_str = f"{start_date.strftime('%Y%m%d')}-{end_date.strftime('%Y%m%d')}"
url = f"https://site.api.espn.com/apis/site/v2/sports/{self.sport}/{self.league}/scoreboard"
response = self.session.get(url, params={"dates": date_str, "limit": 1000},headers=self.headers, timeout=10)
response.raise_for_status()
data = response.json()
data = fetch_espn_scoreboard(
self.session, url, params={"dates": date_str, "limit": 1000},
headers=self.headers, timeout=10, logger=self.logger,
)
immediate_events = data.get('events', [])
if immediate_events:
@@ -1031,7 +1052,7 @@ class SportsCore(ABC):
if os.path.exists(font_path):
# Try loading as TTF first (works for both TTF and some BDF files with PIL)
if font_path.lower().endswith('.ttf'):
font = ImageFont.truetype(font_path, font_size)
font = load_truetype(font_path, font_size)
self.logger.debug(f"Loaded font: {font_name} at size {font_size}")
self._font_cache[cache_key] = font
return font
@@ -1049,7 +1070,7 @@ class SportsCore(ABC):
# correct one: the newer copies call truetype() on a BDF at
# any size (which simply fails) or refuse BDF outright.
try:
font = ImageFont.truetype(font_path, font_size)
font = load_truetype(font_path, font_size)
self.logger.debug(f"Loaded BDF font: {font_name} at size {font_size}")
self._font_cache[cache_key] = font
return font
@@ -1061,7 +1082,7 @@ class SportsCore(ABC):
self._bdf_native_size_cache[font_path] = native_size
if native_size and native_size != font_size:
try:
font = ImageFont.truetype(font_path, native_size)
font = load_truetype(font_path, native_size)
self.logger.debug(
f"Loaded BDF font: {font_name} at its native size {native_size} "
f"(requested {font_size} isn't a valid strike for this file)"
@@ -1089,7 +1110,7 @@ class SportsCore(ABC):
_resolve_font_family_alias(base_default))
try:
if os.path.exists(default_font_path):
font = ImageFont.truetype(default_font_path, font_size)
font = load_truetype(default_font_path, font_size)
else:
self.logger.warning("Default font not found, using PIL default")
font = ImageFont.load_default()
+250 -27
View File
@@ -5,7 +5,9 @@ Handles persistent disk-based caching with atomic writes and error recovery.
"""
import json
import math
import os
import stat
import time
import tempfile
import logging
@@ -14,6 +16,13 @@ import zlib
from typing import Dict, Any, Optional, Protocol
from datetime import datetime
from src.common.path_safety import safe_path_component
try: # optional: large speedup on the cache write path, see _dumps below
import orjson
except ImportError: # pragma: no cover - exercised on hosts without the wheel
orjson = None
# How old an abandoned write's temp file must be before the sweep removes it.
# A real write holds its temp file for milliseconds, so an hour is far beyond
# any in-flight write while still clearing the same day's debris. Deliberately
@@ -40,13 +49,156 @@ class CacheStrategyProtocol(Protocol):
class DateTimeEncoder(json.JSONEncoder):
"""JSON encoder that handles datetime objects."""
"""JSON encoder that handles datetime objects.
Retained for the stdlib fallback path and for any caller importing it.
"""
def default(self, obj: Any) -> Any:
if isinstance(obj, datetime):
return obj.isoformat()
return super().default(obj)
def _datetime_default(obj: Any) -> Any:
"""Serialise datetimes exactly as DateTimeEncoder did."""
if isinstance(obj, datetime):
return obj.isoformat()
raise TypeError(f"Object of type {type(obj).__name__} is not JSON serializable")
def _replace_nonfinite(obj: Any) -> Any:
"""Non-finite floats -> None, matching what ``orjson.dumps`` writes.
Only reached once a strict pass has proved there is something to replace,
so the ordinary write path never pays for this walk.
"""
if isinstance(obj, float):
return obj if math.isfinite(obj) else None
if isinstance(obj, dict):
return {k: _replace_nonfinite(v) for k, v in obj.items()}
if isinstance(obj, (list, tuple)):
return [_replace_nonfinite(v) for v in obj]
return obj
# NON-FINITE FLOATS
# -----------------
# JSON has no NaN or Infinity. The stdlib emits them anyway as an extension;
# orjson refuses to and writes null. That divergence is not acceptable in a
# cache whose files outlive the decision of which encoder is installed, so the
# policy here is one behaviour on both paths:
#
# writing non-finite floats become null, whichever encoder is in use
# reading files already on disk that carry the stdlib's NaN/Infinity
# tokens stay readable, whichever encoder is in use
#
# Without the write half, installing orjson silently changed cached values.
# Without the read half, installing orjson turned every legacy record holding a
# NaN into a "corrupted cache file" that DiskCache.get logged as an error and
# deleted. Both halves are covered by test/test_cache_nonfinite_floats.py.
if orjson is not None:
# Encoding the cache record dominated the background fetch worker: on a
# Pi 4, stdlib json.dumps runs ~12ms per MB and holds the GIL for all of
# it, which stalls the render thread mid-scroll. orjson measures ~7x
# faster on the same payloads (11.9ms -> 1.6ms for 985KB). Decoding gains
# far less (~1.3x on large payloads) because the cost there is building
# the Python objects, not scanning the text, but it is still free to take.
#
# OPT_NON_STR_KEYS: stdlib json coerces int/float dict keys to strings;
# orjson raises without this, and cache records do carry numeric keys.
# OPT_PASSTHROUGH_DATETIME: orjson would otherwise emit its own RFC 3339
# form for datetimes instead of calling default(). Routing them through
# _datetime_default keeps byte-for-byte parity with the records already
# on disk.
_DUMPS_OPTS = orjson.OPT_NON_STR_KEYS | orjson.OPT_PASSTHROUGH_DATETIME
def _dumps(data: Any) -> bytes:
return orjson.dumps(data, default=_datetime_default, option=_DUMPS_OPTS)
def _loads(raw: bytes) -> Any:
try:
return orjson.loads(raw)
except orjson.JSONDecodeError:
# Legacy record written by the stdlib path, carrying NaN or
# Infinity. Genuinely malformed files raise again from here, as
# json.JSONDecodeError, which is what DiskCache.get expects.
return json.loads(raw)
else:
def _dumps(data: Any) -> bytes:
try:
return json.dumps(data, cls=DateTimeEncoder,
allow_nan=False).encode("utf-8")
except ValueError:
# allow_nan=False is what detects the non-finite values; the walk
# runs only now that we know there is one to replace.
return json.dumps(_replace_nonfinite(data), cls=DateTimeEncoder,
allow_nan=False).encode("utf-8")
def _loads(raw: bytes) -> Any:
return json.loads(raw)
# SHARING CACHE FILES BETWEEN THE TWO SERVICES
# --------------------------------------------
# The display service runs as root and the web interface as the installing
# user, and the web interface reads records only the display writes
# (display_current_state, display_on_demand_state, plugin_metrics:*). Files are
# written 0660, so the web interface can read one only through its group.
#
# The installers rely on the directory's setgid bit to set that group. That is
# not something the cache can count on: systemd's CacheDirectory=, which
# ledmatrix-web.service carried until Sept 2026, re-owns the directory and
# everything in it to the web user and its primary group whenever the
# directory's owner does not match, and the setgid layout never survives that.
# From then on every file root creates is root:root 0660, unreadable by the web
# interface. Measured on one rig: 365 such files, and the web UI's display
# status, on-demand state and plugin health all silently empty.
#
# So a cache file takes its group from the directory explicitly, whether or
# not setgid is set. Only a group-writable directory counts as shared: that
# group can already replace any file in it, so reading them grants nothing new.
#
# Everything here works on an open descriptor, never a path. The directory is
# writable by the web user, so between a path check and a path operation that
# user could put a symlink in the file's place, and root would then chown and
# chmod whatever it points at.
_CACHE_FILE_MODE = 0o660
def _shared_group(directory: str) -> Optional[int]:
"""The group a cache file in ``directory`` should carry, if it is shared."""
try:
st = os.stat(directory)
except OSError:
return None
if not st.st_mode & stat.S_IWGRP:
return None
return st.st_gid
def _share_open_file(fd: int, group: Optional[int]) -> None:
"""Make an open cache file readable by the other service. Best effort."""
fchmod = getattr(os, 'fchmod', None) # absent on Windows before 3.13
if fchmod is not None:
try:
fchmod(fd, _CACHE_FILE_MODE)
except OSError:
pass
fchown = getattr(os, 'fchown', None) # absent on Windows
if fchown is None or group is None:
return
try:
if os.fstat(fd).st_gid != group:
fchown(fd, -1, group)
except OSError:
# Not a member of the directory's group and not root: nothing to do,
# and the file keeps the group it was created with.
pass
class DiskCache:
"""Manages persistent disk-based cache."""
@@ -70,16 +222,30 @@ class DiskCache:
def get_cache_path(self, key: str) -> Optional[str]:
"""
Get the path for a cache file.
The key becomes a filename, so it has to be one. Keys reach this
method from the web API -- POST /api/v3/cache/delete passes the
request body's ``key`` straight through CacheManager.clear_cache to
os.remove -- and a key of ``../../../../etc/whatever`` named a file
well outside the cache directory. Every real key is the stem of a
file already sitting flat in cache_dir (that is how list_cache_files
derives them), so rejecting anything with a path component turns
away only inputs that could never have been written here.
Args:
key: Cache key
Returns:
Path to cache file or None if cache is disabled
Path to cache file, or None if cache is disabled or the key is
not a usable filename
"""
if not self.cache_dir:
return None
return os.path.join(self.cache_dir, f"{key}.json")
safe_key = safe_path_component(key)
if safe_key is None:
self.logger.warning("Rejected unsafe cache key %r", key)
return None
return os.path.join(self.cache_dir, f"{safe_key}.json")
def get(self, key: str, max_age: Optional[int] = 300) -> Optional[Dict[str, Any]]:
"""
@@ -99,8 +265,8 @@ class DiskCache:
try:
with self._lock:
with open(cache_path, 'r', encoding='utf-8') as f:
record = json.load(f)
with open(cache_path, 'rb') as f:
record = _loads(f.read())
# Determine record timestamp (prefer embedded, else file mtime)
record_ts = None
@@ -189,12 +355,12 @@ class DiskCache:
# write path below, and cache files are machine-read only — indenting
# them just multiplied the bytes written to the SD card.
try:
payload = json.dumps(data, cls=DateTimeEncoder)
payload = _dumps(data)
except (TypeError, ValueError) as e:
self.logger.warning("Cache data for key '%s' not serializable: %s", key, e)
return
digest = zlib.adler32(payload.encode('utf-8'))
digest = zlib.adler32(payload)
try:
# Atomic write to avoid partial/corrupt files
@@ -242,15 +408,14 @@ class DiskCache:
# wear source (dozens of fsyncs/min on API-heavy
# installs) for data that can be re-downloaded.
try:
with os.fdopen(fd, 'w', encoding='utf-8') as tmp_file:
with os.fdopen(fd, 'wb') as tmp_file:
tmp_file.write(payload)
# Before the rename, not after: mkstemp
# creates the file 0600, and a reader that
# opened it in between was refused.
_share_open_file(tmp_file.fileno(), _shared_group(tmp_dir))
os.replace(tmp_path, cache_path)
self._write_digests[key] = digest
# Set proper permissions: 660 (rw-rw----) for group-readable cache files
try:
os.chmod(cache_path, 0o660) # nosec B103 - intentional; web UI and service share a group
except OSError:
pass # Non-critical if chmod fails
finally:
if os.path.exists(tmp_path):
try:
@@ -260,14 +425,10 @@ class DiskCache:
else:
# Fallback: direct write (not atomic, but better than failing)
try:
with open(cache_path, 'w', encoding='utf-8') as cache_file:
with open(cache_path, 'wb') as cache_file:
cache_file.write(payload)
_share_open_file(cache_file.fileno(), _shared_group(tmp_dir))
self._write_digests[key] = digest
# Set proper permissions: 660 (rw-rw----) for group-readable cache files
try:
os.chmod(cache_path, 0o660) # nosec B103 - intentional; web UI and service share a group
except OSError:
pass # Non-critical if chmod fails
self.logger.debug("Wrote cache for %s directly (non-atomic)", key)
except (IOError, OSError, PermissionError) as write_error:
# If direct write also fails, try fallback location
@@ -290,13 +451,9 @@ class DiskCache:
# is a different path, so future sets must keep
# retrying the primary location.
fallback_path = os.path.join(fallback_dir, os.path.basename(cache_path))
with open(fallback_path, 'w', encoding='utf-8') as tmp_file:
with open(fallback_path, 'wb') as tmp_file:
tmp_file.write(payload)
# Set proper permissions: 660 (rw-rw----) for group-readable cache files
try:
os.chmod(fallback_path, 0o660) # nosec B103 - intentional; web UI and service share a group
except OSError:
pass # Non-critical if chmod fails
_share_open_file(tmp_file.fileno(), _shared_group(fallback_dir))
self.logger.debug("Cache wrote to fallback location: %s", fallback_path)
return # Successfully wrote to fallback, exit gracefully
except (IOError, OSError, PermissionError) as e2:
@@ -353,6 +510,72 @@ class DiskCache:
def get_cache_dir(self) -> Optional[str]:
"""Get the cache directory path."""
return self.cache_dir
def share_existing_files(self) -> int:
"""Give cache files already on disk the group set() now gives new ones.
set() fixes every file it writes from now on; this repairs the ones an
older version left behind as root:root, which the web interface cannot
read until each key happens to be rewritten -- and some, like a
plugin's metrics, may not be for a long time. Meant to run once per
process, off the startup path.
Only this process's own regular files are touched, and each one through
a descriptor opened with O_NOFOLLOW and checked for a single link: the
directory is writable by the web user, and a root process must not be
steered into changing a file outside it.
Returns:
Number of files whose group or mode was changed.
"""
fchown = getattr(os, 'fchown', None)
geteuid = getattr(os, 'geteuid', None)
nofollow = getattr(os, 'O_NOFOLLOW', None)
if not self.cache_dir or fchown is None or geteuid is None or nofollow is None:
return 0
group = _shared_group(self.cache_dir)
if group is None:
return 0
euid = geteuid()
changed = 0
try:
entries = list(os.scandir(self.cache_dir))
except OSError as e:
self.logger.debug("Could not scan %s to share cache files: %s", self.cache_dir, e)
return 0
for entry in entries:
if not entry.name.endswith('.json'):
continue
try:
st = entry.stat(follow_symlinks=False)
except OSError:
continue
if (not stat.S_ISREG(st.st_mode) or st.st_uid != euid
or (st.st_gid == group and stat.S_IMODE(st.st_mode) == _CACHE_FILE_MODE)):
continue
try:
fd = os.open(entry.path, os.O_RDONLY | nofollow | getattr(os, 'O_NONBLOCK', 0))
except OSError:
continue
try:
st = os.fstat(fd)
if not stat.S_ISREG(st.st_mode) or st.st_uid != euid or st.st_nlink != 1:
continue
_share_open_file(fd, group)
st = os.fstat(fd)
if st.st_gid == group and stat.S_IMODE(st.st_mode) == _CACHE_FILE_MODE:
changed += 1
except OSError:
continue
finally:
os.close(fd)
if changed:
self.logger.info(
"Made %d cache file(s) in %s readable by the directory's group "
"(gid %d) so the web interface can read them",
changed, self.cache_dir, group)
return changed
@staticmethod
def _is_orphaned_temp(filename: str) -> bool:

Some files were not shown because too many files have changed in this diff Show More