Compare commits

..
303 Commits
Author SHA1 Message Date
ChuckandClaude Opus 5.5 4ddc3a3620 chore: prepare the 3.8.0 release (#709)
* chore(deprecation): remove the 35 APIs deprecated for 3.8.0

The usage scan (docs/DEPRECATIONS_3.8.md, regenerated 2026-10-01 and
committed here) finds no call or override of any of them in the 46
monorepo plugins or the 8 third-party plugins plugins.json lists; the
only core callers were other deprecated methods removed alongside.

- CacheManager: 13 methods, plus the private helpers only
  has_data_changed used (_has_*_changed, _is_market_open).
- DisplayManager: 7 methods, plus WEATHER_COLORS and the private
  _draw_sun/_cloud/_rain/_snow/_storm helpers only the icon methods used.
- FontManager: 14 methods, plus size_tokens, _save_overrides and
  _clear_plugin_font_cache. font_overrides and _load_overrides stay:
  resolve_font() still applies config/font_overrides.json.
  performance_stats stays: get_font() keeps it and tests read it.
- PluginManager.get_enabled_plugins.

test_deprecation.py pins only the two 3.9.0 markers now; the scanner
tests run against a stand-in core instead of the real markers. The
memory-tier tests read stats through log_memory_cache_stats() and the
component, and the test of the removed _clear_plugin_font_cache goes.
Docs drop the removed methods' reference entries; the Deprecated APIs
table becomes "Removed in 3.8.0". CHANGELOG gains a Removed section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(deprecation): drop the test harness's copies of the removed icon methods

VisualTestDisplayManager still drew weather icons that DisplayManager no
longer has, so a plugin's visual tests could pass on calls that raise
AttributeError on the real display. Its draw_sun/draw_cloud/draw_rain/
draw_snow/draw_weather_icon/draw_text_with_icons, WEATHER_COLORS and the
private helpers go, with the tests that exercised them. The CHANGELOG's
Deprecations entries no longer say nothing is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: prepare the 3.8.0 release

Bumps src.__version__ to 3.8.0 and turns Unreleased into ## 3.8.0, with a
summary and a New modules list (vegas_elements, testing.vegas, sports_vegas;
display_watchdog, plugin_catalog, plugin_runtime) for plugins flooring on
3.8.0. Adds the CHANGELOG line #701's second commit lacked (blocks laid out
off the render thread).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: 3.8.0 also ships what landed on main since the prep

#703, #705, #706 and #693 merged after this branch was cut; their CHANGELOG
entries now sit under 3.8.0. The summary and New modules list name them
(sports consolidation stage 4's four modules, src/ipc, field_model), stage
4's section says to floor on 3.8.0, and src/common/README.md marks its four
modules 3.8.0 instead of Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 11:07:02 -04:00
ChuckandClaude Opus 5.5 2601cb4cbb chore(deprecation): remove the 35 APIs deprecated for 3.8.0 (#708)
* chore(deprecation): remove the 35 APIs deprecated for 3.8.0

The usage scan (docs/DEPRECATIONS_3.8.md, regenerated 2026-10-01 and
committed here) finds no call or override of any of them in the 46
monorepo plugins or the 8 third-party plugins plugins.json lists; the
only core callers were other deprecated methods removed alongside.

- CacheManager: 13 methods, plus the private helpers only
  has_data_changed used (_has_*_changed, _is_market_open).
- DisplayManager: 7 methods, plus WEATHER_COLORS and the private
  _draw_sun/_cloud/_rain/_snow/_storm helpers only the icon methods used.
- FontManager: 14 methods, plus size_tokens, _save_overrides and
  _clear_plugin_font_cache. font_overrides and _load_overrides stay:
  resolve_font() still applies config/font_overrides.json.
  performance_stats stays: get_font() keeps it and tests read it.
- PluginManager.get_enabled_plugins.

test_deprecation.py pins only the two 3.9.0 markers now; the scanner
tests run against a stand-in core instead of the real markers. The
memory-tier tests read stats through log_memory_cache_stats() and the
component, and the test of the removed _clear_plugin_font_cache goes.
Docs drop the removed methods' reference entries; the Deprecated APIs
table becomes "Removed in 3.8.0". CHANGELOG gains a Removed section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(deprecation): drop the test harness's copies of the removed icon methods

VisualTestDisplayManager still drew weather icons that DisplayManager no
longer has, so a plugin's visual tests could pass on calls that raise
AttributeError on the real display. Its draw_sun/draw_cloud/draw_rain/
draw_snow/draw_weather_icon/draw_text_with_icons, WEATHER_COLORS and the
private helpers go, with the tests that exercised them. The CHANGELOG's
Deprecations entries no longer say nothing is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:45:28 -04:00
ChuckandClaude Opus 5.5 695ff92009 feat(ipc): display control socket, stage 1 - on-demand with acks (#706)
The display serves a control socket (/run/ledmatrix/control.sock) carrying versioned JSON commands, one per line, each answered. Stage 1 covers on-demand start, stop and status; commands are queued on the socket thread and applied on the render thread through the mailbox's own handler, and the web interface falls back to the file mailbox when the socket is unavailable. Protocol and security model: docs/IPC_CONTROL_SOCKET.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:33:02 -04:00
ChuckandClaude Opus 5.5 74696d2108 fix(display): routine rotation log lines are DEBUG (30% fewer journal lines) (#693)
"Processing mode", "display() returned False" and "No content to display" repeated what "Switching to mode" already logs on every rotation; they are now DEBUG. On ledpi this cut the display's journal lines by about 30%; the measured SD-write saving is small (within noise).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:21:22 -04:00
ChuckandClaude Opus 5.5 3ad0438e75 feat(common): sports consolidation stage 4 -- the identical sweep (plugin host, live scroll, display rules, font path) (#705)
Moves the code every scoreboard plugin carries identically into core: src.common.sports_plugin_host, sports_live_scroll, sports_display_rules and sports_font_path, with unit tests and a parity test against the ledmatrix-plugins copies (LEDMATRIX_PLUGINS).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:10:22 -04:00
ChuckandClaude Opus 5.5 c6701ac00d feat(web): ES-module page lifecycle and one schema field model (stage 1) (#703)
Adds a native ES-module layer to the web UI (core/boot, registry, api, facade; window.LEDMatrix as the one global), a page lifecycle that the Cache tab is converted to as the reference, text/javascript serving and revalidation for unversioned module requests, and src/plugin_system/field_model.py with a parity test against the render_field macro. Also: the cache page toggles its grey 'Not configured' style instead of only adding it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 10:00:26 -04:00
ChuckandClaude Opus 5.5 795834811f perf(scroll): extend and trim the Vegas strip in place (#701)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): build the strip's PIL image only when something reads it

Every Vegas strip extension rebuilt ScrollHelper.cached_image from
cached_array in full, twice (append, then trim), on the render thread:
Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a
Pi 4 (measured on ledpi), about two thirds of an extension's render-thread
cost. Nothing on the frame path reads the image's pixels; every frame is cut
from the array.

cached_image is now a property. append_content and drop_scrolled_prefix
defer it; the first read builds it from the array it started with and keeps
it only if the strip has not changed meanwhile, so a sync push racing an
extension cannot leave a stale image cached. Assigning cached_image stores
exactly what was assigned, as before. has_strip() says whether there is a
strip without building its image; the helper's frame path, Vegas and the
adapter's scroll-cache invalidation use it. The strip is also no longer held
in memory twice.

In Vegas the image is now built only by a multi-display sync push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements -- a plugin API for content that changes while it scrolls

Vegas bakes each plugin's pictures into one strip, so a card already on its
way across the panel keeps what it showed when it was drawn. This adds the
API and bookkeeping for content that can be updated in place; the worker
that redraws and swaps it follows separately. No shipped plugin implements
the hook yet, so nothing changes for users.

Plugin API (core 3.8.0), all no-ops by default:
- BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version,
  live, refresh_hz)]: named, fixed-width pieces of Vegas content.
- BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free
  redraw for content that changes with time.
- BasePlugin.notify_vegas_data_changed(): data that lands outside update().
- src/plugin_system/vegas_elements.py (VegasElement, re-exported from
  base_plugin).

Core:
- PluginAdapter asks a plugin that implements the hook for elements on the
  background fetch only (under its lock, on its own canvas); every other
  path keeps get_vegas_content(). Live elements are pinned (padded with
  content_padding, never trimmed), tagged with their key, digest and data
  epoch in Image.info so the existing cache and group plumbing carry them
  unchanged, and untagged if a width budget crops them.
- RenderPipeline records where each live element lands (ElementRecord), in
  absolute strip columns a trim does not move; the block-start arithmetic
  is shared with the STATIC markers.
- PluginManager update listeners (add/remove_update_listener,
  notify_data_changed): told the moment update() completes, not at the
  next ~4s Vegas poll. The coordinator uses one to move each plugin's data
  epoch on.
- vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval,
  live_lead_screens; per-plugin core-owned vegas_live. Live elements are
  off under multi-display sync, in swap mode and with offscreen_prefetch off.
- scripts/check_plugin.py checks the element contract
  (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub
  is a working example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements update in place while they scroll

One background worker (src/vegas_mode/live_worker.py) redraws a plugin's
live elements when its data epoch moves on (update listener) or on their
refresh_hz, nearest the screen first, and hands changed pixels lock-free to
the render thread, which copies them into the strip between frames
(RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most
four patches or two screens of bytes a frame, no drawing or locks there.
The worker takes over group prefetch once a live element is placed, runs
inside the render gate, and is supervised. Update tick 1s while live
elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md
describes what was built and why SegmentStrip was not needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(sports): live Vegas cards for the scoreboards (shared layer)

One live element per game, drawn only when what the card shows changes, so
a score changes on a card already crossing the panel. The shared part, so
each scoreboard adopts it in a few lines:

- src/common/sports_vegas.py: game_key, game_fingerprint (the whole game
  dict, frozen: no drawn field can be missed), dedupe_games, VegasCardCache,
  StickyOdds (odds a live poll left out stay drawn), finished_games /
  with_finished_games (a game that just went final keeps its card, after its
  league's live games; one a heuristic only judged over keeps its live
  state, so a tied end of regulation never shows FINAL early).
- SportsScrollDisplay.make_vegas_renderer() is the override point;
  build_vegas_elements() and SportsScrollDisplayManager
  .get_vegas_elements_for() do the rest. A card's version includes its
  teams' ranks, which the renderer draws from the rankings cache.
- SportsLiveSharedMixin._record_finished_game() / finished_games_snapshot():
  held for FINISHED_GAME_TTL after it leaves the live list.

A sport that does not implement make_vegas_renderer keeps its ordinary Vegas
content, so no scoreboard changes until it opts in.

scripts/render_plugin.py --vegas renders a plugin's Vegas block as the
ticker lays it out, and --timeline stacks it at successive moments as
the ticker would update it in place; the join is now
render_pipeline.join_plugin_rows().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): keep live games in the ticker by default

display.vegas_scroll.live_in_ticker now defaults to true: through a live
game the marquee keeps running and the live scoreboard takes extra turns in
it -- its cards updating in place while they scroll -- instead of the ticker
giving way to the full-screen scoreboard.

The new default would reach nobody on its own: every existing config holds
an explicit false copied from the template (there was no control for it),
and the template merge only adds missing keys. ConfigManager therefore turns
a stored false on once, with a backup, and records live_in_ticker_migrated
so a false chosen afterwards stays. The marker is never in the template.

A "Keep live games in the ticker" checkbox under Vegas mode sets it. Tests
that pin the full-screen takeover now say live_in_ticker=false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(dev): preview a plugin's Vegas strip in the dev server

The dev server's View selector gains "Vegas strip (live elements)" and
"Vegas strip (plain Vegas content)": the plugin's block of the Vegas ticker,
laid out by the ticker's own code (render_vegas_strip, as render_plugin.py
--vegas uses), with its live elements listed. /api/render takes
"vegas": "live" | "plain"; the display view is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): extend and trim the Vegas strip in place

Every strip extension rebuilt the whole strip (np.concatenate: 2-2.6 ms for
a 10-14k px strip at 512x64 on a Pi 4) and every trim copied what was left
(1.2-1.8 ms), on the render thread. With ~4 ms of slack per refresh, every
extension frame on hdpi missed its refresh (5/5 in each soak run).

The strip now lives in a buffer with spare room; cached_array is a view of
its live columns. An append writes only the new columns (~0.2 ms), a trim
only moves the view's start, and the one full copy happens when the buffer
is reallocated (STRIP_SPARE_FACTOR 3: about once every two strip-lengths
scrolled). A cached_array set from outside -- the multi-display follower's
read-only one, create_scrolling_image's -- is never written through, and a
new strip lets the old buffer go. last_copy_bytes says what was copied, and
the Vegas frame-timing attribution reports that instead of the whole strip.

test_scroll_helper_in_place.py: the buffer is reused and only new columns
copied, trims copy nothing, reallocation when the room runs out, outside
arrays untouched, and random appends/trims/patches/scrolling checked frame
by frame against the old copying strip (mutation-checked).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sports): a default _determine_game_type on SportsScrollDisplay

render_vegas_card looked the method up with getattr and a None default, which
static analysis (Codacy) reports as calling something that may not be
callable. The base class now has the default -- the card type from the game's
state -- and the plugins that define their own override it as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(scroll): no assert in _extended_strip

An assert vanishes under python -O (Codacy); a real check says the same.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: review follow-ups on the shared live-card layer

- The reused Vegas renderer always gets the current rankings, empty
  included, so ranks cleared since are not kept drawn.
- render_plugin.py: --timeline refuses --no-live (a timeline shows live
  elements changing), --timeline/--no-live need --vegas, and the Vegas
  paths create the output's directory like the display path does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(vegas): lay a group's blocks out before the extension, off the render thread

#701 cut the strip copy, but the frame after an extension was still late:
the render thread also joined each plugin's rows (separation_gap measures
every pair) and pasted the blocks into an addition image, ~37 ms on hdpi
(Pi 4, 512x64) against ~3.75 ms of slack.

The thread that fetched the group now does that as each member arrives:
RenderPipeline.prepare_group_member joins the rows and turns the block into
pixels (the prefetch thread and the live worker, under the render gate).
extend_scroll_content takes those blocks, and ScrollHelper.append_content
writes items -- images or RGB arrays -- straight into the strip's spare room,
blanking only the gaps. A member that was not prepared (an inline fetch) is
joined at the extension as before; the strip is identical either way.

On hdpi the extension's render-thread work goes from 37.5 ms to 3.2 ms p50
in place (6.9 ms when the buffer is reallocated). The helper's per-append
INFO line, a duplicate of the pipeline's, is now DEBUG.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 09:48:47 -04:00
ChuckandClaude Opus 5.5 dfd67c7c8b feat(dev): preview a plugin's Vegas strip in the dev server (#700)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): build the strip's PIL image only when something reads it

Every Vegas strip extension rebuilt ScrollHelper.cached_image from
cached_array in full, twice (append, then trim), on the render thread:
Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a
Pi 4 (measured on ledpi), about two thirds of an extension's render-thread
cost. Nothing on the frame path reads the image's pixels; every frame is cut
from the array.

cached_image is now a property. append_content and drop_scrolled_prefix
defer it; the first read builds it from the array it started with and keeps
it only if the strip has not changed meanwhile, so a sync push racing an
extension cannot leave a stale image cached. Assigning cached_image stores
exactly what was assigned, as before. has_strip() says whether there is a
strip without building its image; the helper's frame path, Vegas and the
adapter's scroll-cache invalidation use it. The strip is also no longer held
in memory twice.

In Vegas the image is now built only by a multi-display sync push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements -- a plugin API for content that changes while it scrolls

Vegas bakes each plugin's pictures into one strip, so a card already on its
way across the panel keeps what it showed when it was drawn. This adds the
API and bookkeeping for content that can be updated in place; the worker
that redraws and swaps it follows separately. No shipped plugin implements
the hook yet, so nothing changes for users.

Plugin API (core 3.8.0), all no-ops by default:
- BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version,
  live, refresh_hz)]: named, fixed-width pieces of Vegas content.
- BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free
  redraw for content that changes with time.
- BasePlugin.notify_vegas_data_changed(): data that lands outside update().
- src/plugin_system/vegas_elements.py (VegasElement, re-exported from
  base_plugin).

Core:
- PluginAdapter asks a plugin that implements the hook for elements on the
  background fetch only (under its lock, on its own canvas); every other
  path keeps get_vegas_content(). Live elements are pinned (padded with
  content_padding, never trimmed), tagged with their key, digest and data
  epoch in Image.info so the existing cache and group plumbing carry them
  unchanged, and untagged if a width budget crops them.
- RenderPipeline records where each live element lands (ElementRecord), in
  absolute strip columns a trim does not move; the block-start arithmetic
  is shared with the STATIC markers.
- PluginManager update listeners (add/remove_update_listener,
  notify_data_changed): told the moment update() completes, not at the
  next ~4s Vegas poll. The coordinator uses one to move each plugin's data
  epoch on.
- vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval,
  live_lead_screens; per-plugin core-owned vegas_live. Live elements are
  off under multi-display sync, in swap mode and with offscreen_prefetch off.
- scripts/check_plugin.py checks the element contract
  (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub
  is a working example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements update in place while they scroll

One background worker (src/vegas_mode/live_worker.py) redraws a plugin's
live elements when its data epoch moves on (update listener) or on their
refresh_hz, nearest the screen first, and hands changed pixels lock-free to
the render thread, which copies them into the strip between frames
(RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most
four patches or two screens of bytes a frame, no drawing or locks there.
The worker takes over group prefetch once a live element is placed, runs
inside the render gate, and is supervised. Update tick 1s while live
elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md
describes what was built and why SegmentStrip was not needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(sports): live Vegas cards for the scoreboards (shared layer)

One live element per game, drawn only when what the card shows changes, so
a score changes on a card already crossing the panel. The shared part, so
each scoreboard adopts it in a few lines:

- src/common/sports_vegas.py: game_key, game_fingerprint (the whole game
  dict, frozen: no drawn field can be missed), dedupe_games, VegasCardCache,
  StickyOdds (odds a live poll left out stay drawn), finished_games /
  with_finished_games (a game that just went final keeps its card, after its
  league's live games; one a heuristic only judged over keeps its live
  state, so a tied end of regulation never shows FINAL early).
- SportsScrollDisplay.make_vegas_renderer() is the override point;
  build_vegas_elements() and SportsScrollDisplayManager
  .get_vegas_elements_for() do the rest. A card's version includes its
  teams' ranks, which the renderer draws from the rankings cache.
- SportsLiveSharedMixin._record_finished_game() / finished_games_snapshot():
  held for FINISHED_GAME_TTL after it leaves the live list.

A sport that does not implement make_vegas_renderer keeps its ordinary Vegas
content, so no scoreboard changes until it opts in.

scripts/render_plugin.py --vegas renders a plugin's Vegas block as the
ticker lays it out, and --timeline stacks it at successive moments as
the ticker would update it in place; the join is now
render_pipeline.join_plugin_rows().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): keep live games in the ticker by default

display.vegas_scroll.live_in_ticker now defaults to true: through a live
game the marquee keeps running and the live scoreboard takes extra turns in
it -- its cards updating in place while they scroll -- instead of the ticker
giving way to the full-screen scoreboard.

The new default would reach nobody on its own: every existing config holds
an explicit false copied from the template (there was no control for it),
and the template merge only adds missing keys. ConfigManager therefore turns
a stored false on once, with a backup, and records live_in_ticker_migrated
so a false chosen afterwards stays. The marker is never in the template.

A "Keep live games in the ticker" checkbox under Vegas mode sets it. Tests
that pin the full-screen takeover now say live_in_ticker=false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(dev): preview a plugin's Vegas strip in the dev server

The dev server's View selector gains "Vegas strip (live elements)" and
"Vegas strip (plain Vegas content)": the plugin's block of the Vegas ticker,
laid out by the ticker's own code (render_vegas_strip, as render_plugin.py
--vegas uses), with its live elements listed. /api/render takes
"vegas": "live" | "plain"; the display view is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sports): a default _determine_game_type on SportsScrollDisplay

render_vegas_card looked the method up with getattr and a None default, which
static analysis (Codacy) reports as calling something that may not be
callable. The base class now has the default -- the card type from the game's
state -- and the plugins that define their own override it as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: review follow-ups on the shared live-card layer

- The reused Vegas renderer always gets the current rankings, empty
  included, so ranks cleared since are not kept drawn.
- render_plugin.py: --timeline refuses --no-live (a timeline shows live
  elements changing), --timeline/--no-live need --vegas, and the Vegas
  paths create the output's directory like the display path does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:37:15 -04:00
ChuckandClaude Opus 5.5 7f06cc9c3b feat(vegas): keep live games in the ticker by default (#699)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): build the strip's PIL image only when something reads it

Every Vegas strip extension rebuilt ScrollHelper.cached_image from
cached_array in full, twice (append, then trim), on the render thread:
Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a
Pi 4 (measured on ledpi), about two thirds of an extension's render-thread
cost. Nothing on the frame path reads the image's pixels; every frame is cut
from the array.

cached_image is now a property. append_content and drop_scrolled_prefix
defer it; the first read builds it from the array it started with and keeps
it only if the strip has not changed meanwhile, so a sync push racing an
extension cannot leave a stale image cached. Assigning cached_image stores
exactly what was assigned, as before. has_strip() says whether there is a
strip without building its image; the helper's frame path, Vegas and the
adapter's scroll-cache invalidation use it. The strip is also no longer held
in memory twice.

In Vegas the image is now built only by a multi-display sync push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements -- a plugin API for content that changes while it scrolls

Vegas bakes each plugin's pictures into one strip, so a card already on its
way across the panel keeps what it showed when it was drawn. This adds the
API and bookkeeping for content that can be updated in place; the worker
that redraws and swaps it follows separately. No shipped plugin implements
the hook yet, so nothing changes for users.

Plugin API (core 3.8.0), all no-ops by default:
- BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version,
  live, refresh_hz)]: named, fixed-width pieces of Vegas content.
- BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free
  redraw for content that changes with time.
- BasePlugin.notify_vegas_data_changed(): data that lands outside update().
- src/plugin_system/vegas_elements.py (VegasElement, re-exported from
  base_plugin).

Core:
- PluginAdapter asks a plugin that implements the hook for elements on the
  background fetch only (under its lock, on its own canvas); every other
  path keeps get_vegas_content(). Live elements are pinned (padded with
  content_padding, never trimmed), tagged with their key, digest and data
  epoch in Image.info so the existing cache and group plumbing carry them
  unchanged, and untagged if a width budget crops them.
- RenderPipeline records where each live element lands (ElementRecord), in
  absolute strip columns a trim does not move; the block-start arithmetic
  is shared with the STATIC markers.
- PluginManager update listeners (add/remove_update_listener,
  notify_data_changed): told the moment update() completes, not at the
  next ~4s Vegas poll. The coordinator uses one to move each plugin's data
  epoch on.
- vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval,
  live_lead_screens; per-plugin core-owned vegas_live. Live elements are
  off under multi-display sync, in swap mode and with offscreen_prefetch off.
- scripts/check_plugin.py checks the element contract
  (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub
  is a working example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements update in place while they scroll

One background worker (src/vegas_mode/live_worker.py) redraws a plugin's
live elements when its data epoch moves on (update listener) or on their
refresh_hz, nearest the screen first, and hands changed pixels lock-free to
the render thread, which copies them into the strip between frames
(RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most
four patches or two screens of bytes a frame, no drawing or locks there.
The worker takes over group prefetch once a live element is placed, runs
inside the render gate, and is supervised. Update tick 1s while live
elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md
describes what was built and why SegmentStrip was not needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(sports): live Vegas cards for the scoreboards (shared layer)

One live element per game, drawn only when what the card shows changes, so
a score changes on a card already crossing the panel. The shared part, so
each scoreboard adopts it in a few lines:

- src/common/sports_vegas.py: game_key, game_fingerprint (the whole game
  dict, frozen: no drawn field can be missed), dedupe_games, VegasCardCache,
  StickyOdds (odds a live poll left out stay drawn), finished_games /
  with_finished_games (a game that just went final keeps its card, after its
  league's live games; one a heuristic only judged over keeps its live
  state, so a tied end of regulation never shows FINAL early).
- SportsScrollDisplay.make_vegas_renderer() is the override point;
  build_vegas_elements() and SportsScrollDisplayManager
  .get_vegas_elements_for() do the rest. A card's version includes its
  teams' ranks, which the renderer draws from the rankings cache.
- SportsLiveSharedMixin._record_finished_game() / finished_games_snapshot():
  held for FINISHED_GAME_TTL after it leaves the live list.

A sport that does not implement make_vegas_renderer keeps its ordinary Vegas
content, so no scoreboard changes until it opts in.

scripts/render_plugin.py --vegas renders a plugin's Vegas block as the
ticker lays it out, and --timeline stacks it at successive moments as
the ticker would update it in place; the join is now
render_pipeline.join_plugin_rows().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): keep live games in the ticker by default

display.vegas_scroll.live_in_ticker now defaults to true: through a live
game the marquee keeps running and the live scoreboard takes extra turns in
it -- its cards updating in place while they scroll -- instead of the ticker
giving way to the full-screen scoreboard.

The new default would reach nobody on its own: every existing config holds
an explicit false copied from the template (there was no control for it),
and the template merge only adds missing keys. ConfigManager therefore turns
a stored false on once, with a backup, and records live_in_ticker_migrated
so a false chosen afterwards stays. The marker is never in the template.

A "Keep live games in the ticker" checkbox under Vegas mode sets it. Tests
that pin the full-screen takeover now say live_in_ticker=false.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sports): a default _determine_game_type on SportsScrollDisplay

render_vegas_card looked the method up with getattr and a None default, which
static analysis (Codacy) reports as calling something that may not be
callable. The base class now has the default -- the card type from the game's
state -- and the plugins that define their own override it as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: review follow-ups on the shared live-card layer

- The reused Vegas renderer always gets the current rankings, empty
  included, so ranks cleared since are not kept drawn.
- render_plugin.py: --timeline refuses --no-live (a timeline shows live
  elements changing), --timeline/--no-live need --vegas, and the Vegas
  paths create the output's directory like the display path does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 08:27:16 -04:00
ChuckandClaude Opus 5.5 a12be7c3c5 feat(sports): live Vegas cards for the scoreboards (shared layer) (#698)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): build the strip's PIL image only when something reads it

Every Vegas strip extension rebuilt ScrollHelper.cached_image from
cached_array in full, twice (append, then trim), on the render thread:
Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a
Pi 4 (measured on ledpi), about two thirds of an extension's render-thread
cost. Nothing on the frame path reads the image's pixels; every frame is cut
from the array.

cached_image is now a property. append_content and drop_scrolled_prefix
defer it; the first read builds it from the array it started with and keeps
it only if the strip has not changed meanwhile, so a sync push racing an
extension cannot leave a stale image cached. Assigning cached_image stores
exactly what was assigned, as before. has_strip() says whether there is a
strip without building its image; the helper's frame path, Vegas and the
adapter's scroll-cache invalidation use it. The strip is also no longer held
in memory twice.

In Vegas the image is now built only by a multi-display sync push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements -- a plugin API for content that changes while it scrolls

Vegas bakes each plugin's pictures into one strip, so a card already on its
way across the panel keeps what it showed when it was drawn. This adds the
API and bookkeeping for content that can be updated in place; the worker
that redraws and swaps it follows separately. No shipped plugin implements
the hook yet, so nothing changes for users.

Plugin API (core 3.8.0), all no-ops by default:
- BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version,
  live, refresh_hz)]: named, fixed-width pieces of Vegas content.
- BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free
  redraw for content that changes with time.
- BasePlugin.notify_vegas_data_changed(): data that lands outside update().
- src/plugin_system/vegas_elements.py (VegasElement, re-exported from
  base_plugin).

Core:
- PluginAdapter asks a plugin that implements the hook for elements on the
  background fetch only (under its lock, on its own canvas); every other
  path keeps get_vegas_content(). Live elements are pinned (padded with
  content_padding, never trimmed), tagged with their key, digest and data
  epoch in Image.info so the existing cache and group plumbing carry them
  unchanged, and untagged if a width budget crops them.
- RenderPipeline records where each live element lands (ElementRecord), in
  absolute strip columns a trim does not move; the block-start arithmetic
  is shared with the STATIC markers.
- PluginManager update listeners (add/remove_update_listener,
  notify_data_changed): told the moment update() completes, not at the
  next ~4s Vegas poll. The coordinator uses one to move each plugin's data
  epoch on.
- vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval,
  live_lead_screens; per-plugin core-owned vegas_live. Live elements are
  off under multi-display sync, in swap mode and with offscreen_prefetch off.
- scripts/check_plugin.py checks the element contract
  (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub
  is a working example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements update in place while they scroll

One background worker (src/vegas_mode/live_worker.py) redraws a plugin's
live elements when its data epoch moves on (update listener) or on their
refresh_hz, nearest the screen first, and hands changed pixels lock-free to
the render thread, which copies them into the strip between frames
(RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most
four patches or two screens of bytes a frame, no drawing or locks there.
The worker takes over group prefetch once a live element is placed, runs
inside the render gate, and is supervised. Update tick 1s while live
elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md
describes what was built and why SegmentStrip was not needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(sports): live Vegas cards for the scoreboards (shared layer)

One live element per game, drawn only when what the card shows changes, so
a score changes on a card already crossing the panel. The shared part, so
each scoreboard adopts it in a few lines:

- src/common/sports_vegas.py: game_key, game_fingerprint (the whole game
  dict, frozen: no drawn field can be missed), dedupe_games, VegasCardCache,
  StickyOdds (odds a live poll left out stay drawn), finished_games /
  with_finished_games (a game that just went final keeps its card, after its
  league's live games; one a heuristic only judged over keeps its live
  state, so a tied end of regulation never shows FINAL early).
- SportsScrollDisplay.make_vegas_renderer() is the override point;
  build_vegas_elements() and SportsScrollDisplayManager
  .get_vegas_elements_for() do the rest. A card's version includes its
  teams' ranks, which the renderer draws from the rankings cache.
- SportsLiveSharedMixin._record_finished_game() / finished_games_snapshot():
  held for FINISHED_GAME_TTL after it leaves the live list.

A sport that does not implement make_vegas_renderer keeps its ordinary Vegas
content, so no scoreboard changes until it opts in.

scripts/render_plugin.py --vegas renders a plugin's Vegas block as the
ticker lays it out, and --timeline stacks it at successive moments as
the ticker would update it in place; the join is now
render_pipeline.join_plugin_rows().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sports): a default _determine_game_type on SportsScrollDisplay

render_vegas_card looked the method up with getattr and a None default, which
static analysis (Codacy) reports as calling something that may not be
callable. The base class now has the default -- the card type from the game's
state -- and the plugins that define their own override it as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: review follow-ups on the shared live-card layer

- The reused Vegas renderer always gets the current rankings, empty
  included, so ranks cleared since are not kept drawn.
- render_plugin.py: --timeline refuses --no-live (a timeline shows live
  elements changing), --timeline/--no-live need --vegas, and the Vegas
  paths create the output's directory like the display path does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 21:32:07 -04:00
ChuckandClaude Opus 5.5 f4bda50710 feat(vegas): live elements update in place while they scroll (#697)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): build the strip's PIL image only when something reads it

Every Vegas strip extension rebuilt ScrollHelper.cached_image from
cached_array in full, twice (append, then trim), on the render thread:
Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a
Pi 4 (measured on ledpi), about two thirds of an extension's render-thread
cost. Nothing on the frame path reads the image's pixels; every frame is cut
from the array.

cached_image is now a property. append_content and drop_scrolled_prefix
defer it; the first read builds it from the array it started with and keeps
it only if the strip has not changed meanwhile, so a sync push racing an
extension cannot leave a stale image cached. Assigning cached_image stores
exactly what was assigned, as before. has_strip() says whether there is a
strip without building its image; the helper's frame path, Vegas and the
adapter's scroll-cache invalidation use it. The strip is also no longer held
in memory twice.

In Vegas the image is now built only by a multi-display sync push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements -- a plugin API for content that changes while it scrolls

Vegas bakes each plugin's pictures into one strip, so a card already on its
way across the panel keeps what it showed when it was drawn. This adds the
API and bookkeeping for content that can be updated in place; the worker
that redraws and swaps it follows separately. No shipped plugin implements
the hook yet, so nothing changes for users.

Plugin API (core 3.8.0), all no-ops by default:
- BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version,
  live, refresh_hz)]: named, fixed-width pieces of Vegas content.
- BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free
  redraw for content that changes with time.
- BasePlugin.notify_vegas_data_changed(): data that lands outside update().
- src/plugin_system/vegas_elements.py (VegasElement, re-exported from
  base_plugin).

Core:
- PluginAdapter asks a plugin that implements the hook for elements on the
  background fetch only (under its lock, on its own canvas); every other
  path keeps get_vegas_content(). Live elements are pinned (padded with
  content_padding, never trimmed), tagged with their key, digest and data
  epoch in Image.info so the existing cache and group plumbing carry them
  unchanged, and untagged if a width budget crops them.
- RenderPipeline records where each live element lands (ElementRecord), in
  absolute strip columns a trim does not move; the block-start arithmetic
  is shared with the STATIC markers.
- PluginManager update listeners (add/remove_update_listener,
  notify_data_changed): told the moment update() completes, not at the
  next ~4s Vegas poll. The coordinator uses one to move each plugin's data
  epoch on.
- vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval,
  live_lead_screens; per-plugin core-owned vegas_live. Live elements are
  off under multi-display sync, in swap mode and with offscreen_prefetch off.
- scripts/check_plugin.py checks the element contract
  (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub
  is a working example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements update in place while they scroll

One background worker (src/vegas_mode/live_worker.py) redraws a plugin's
live elements when its data epoch moves on (update listener) or on their
refresh_hz, nearest the screen first, and hands changed pixels lock-free to
the render thread, which copies them into the strip between frames
(RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most
four patches or two screens of bytes a frame, no drawing or locks there.
The worker takes over group prefetch once a live element is placed, runs
inside the render gate, and is supervised. Update tick 1s while live
elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md
describes what was built and why SegmentStrip was not needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 21:07:31 -04:00
ChuckandClaude Opus 5.5 56947298d6 a plugin API for content that changes while it scrolls (#696)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): build the strip's PIL image only when something reads it

Every Vegas strip extension rebuilt ScrollHelper.cached_image from
cached_array in full, twice (append, then trim), on the render thread:
Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a
Pi 4 (measured on ledpi), about two thirds of an extension's render-thread
cost. Nothing on the frame path reads the image's pixels; every frame is cut
from the array.

cached_image is now a property. append_content and drop_scrolled_prefix
defer it; the first read builds it from the array it started with and keeps
it only if the strip has not changed meanwhile, so a sync push racing an
extension cannot leave a stale image cached. Assigning cached_image stores
exactly what was assigned, as before. has_strip() says whether there is a
strip without building its image; the helper's frame path, Vegas and the
adapter's scroll-cache invalidation use it. The strip is also no longer held
in memory twice.

In Vegas the image is now built only by a multi-display sync push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(vegas): live elements -- a plugin API for content that changes while it scrolls

Vegas bakes each plugin's pictures into one strip, so a card already on its
way across the panel keeps what it showed when it was drawn. This adds the
API and bookkeeping for content that can be updated in place; the worker
that redraws and swaps it follows separately. No shipped plugin implements
the hook yet, so nothing changes for users.

Plugin API (core 3.8.0), all no-ops by default:
- BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version,
  live, refresh_hz)]: named, fixed-width pieces of Vegas content.
- BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free
  redraw for content that changes with time.
- BasePlugin.notify_vegas_data_changed(): data that lands outside update().
- src/plugin_system/vegas_elements.py (VegasElement, re-exported from
  base_plugin).

Core:
- PluginAdapter asks a plugin that implements the hook for elements on the
  background fetch only (under its lock, on its own canvas); every other
  path keeps get_vegas_content(). Live elements are pinned (padded with
  content_padding, never trimmed), tagged with their key, digest and data
  epoch in Image.info so the existing cache and group plumbing carry them
  unchanged, and untagged if a width budget crops them.
- RenderPipeline records where each live element lands (ElementRecord), in
  absolute strip columns a trim does not move; the block-start arithmetic
  is shared with the STATIC markers.
- PluginManager update listeners (add/remove_update_listener,
  notify_data_changed): told the moment update() completes, not at the
  next ~4s Vegas poll. The coordinator uses one to move each plugin's data
  epoch on.
- vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval,
  live_lead_screens; per-plugin core-owned vegas_live. Live elements are
  off under multi-display sync, in swap mode and with offscreen_prefetch off.
- scripts/check_plugin.py checks the element contract
  (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub
  is a working example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 20:59:16 -04:00
ChuckandClaude Opus 5.5 596809acc3 perf(scroll): build the strip's PIL image only when something reads it (#695)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(scroll): build the strip's PIL image only when something reads it

Every Vegas strip extension rebuilt ScrollHelper.cached_image from
cached_array in full, twice (append, then trim), on the render thread:
Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a
Pi 4 (measured on ledpi), about two thirds of an extension's render-thread
cost. Nothing on the frame path reads the image's pixels; every frame is cut
from the array.

cached_image is now a property. append_content and drop_scrolled_prefix
defer it; the first read builds it from the array it started with and keeps
it only if the strip has not changed meanwhile, so a sync push racing an
extension cannot leave a stale image cached. Assigning cached_image stores
exactly what was assigned, as before. has_strip() says whether there is a
strip without building its image; the helper's frame path, Vegas and the
adapter's scroll-cache invalidation use it. The strip is also no longer held
in memory twice.

In Vegas the image is now built only by a multi-display sync push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 20:50:34 -04:00
ChuckandClaude Opus 5.5 77862b631b perf(timing): say which render-thread work a late frame followed (#694)
* perf(timing): say which render-thread work a late frame followed

The soak already says how often a moving frame reached the panel late, but
not what the render thread was doing just before it. Vegas does two kinds of
work there between frames -- building its strip (compose, extend) and, with
live elements, patching changed pixels into it -- and deciding whether either
is affordable needs their own numbers.

- FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame.
  Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind;
  aggregate() still takes frames without ops. The file schema is unchanged.
- Vegas tags compose and every strip extension (with the bytes it copied).
- frame_soak prints an "after work" table: frames, late %, freezes and MB
  moved per kind, only when something tagged its work.
- render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes /
  --patch-every / --patch-where (in-place column writes, as a live element
  update does) and --extend-every-screens / --extend-width (append + trim on
  a fixed cadence that holds the strip's width).

No runtime behaviour changes: this is the measurement gate for live Vegas
elements.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note the frame-op attribution and bench modes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 20:40:39 -04:00
ChuckandClaude Opus 5.5 7804ea8f69 feat(update): stable/beta update channel; stable follows release tags (#684)
Adds auto_update.channel: stable follows the newest vX.Y.Z release tag
(detached HEAD; pre-releases and other tags ignored), beta follows main as
before. Nothing ever moves a device backwards: a checkout newer than the
newest release keeps following main (or stays put when detached) until a
release contains its commit. Legacy configs migrate to stable when they
reach a release. Update Code, the weekly updater's preflight, and the
verifier's rollback (back to old_ref: branch or detached release) all
honour the channel. General tab Update Channel select, GET/POST
/api/v3/system/update-channel, release-aware Overview banner and Tools git
panel. New installs default to stable.

Rig fix (ledpi): /system/check-update reports update_available: false when
the channel's action is none (a detached HEAD newer than the newest
release), matching Update Code; the Tools panel no longer calls every
detached HEAD "a release".

Merged with main through #687 (heartbeat verifier, #683 login, #688
plugin_catalog, #685 Tailwind build).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:35:26 -04:00
ChuckandClaude Opus 5.5 64c7289593 feat(display): systemd watchdog and heartbeat for a frozen render loop (#687)
If the render loop gets stuck inside a plugin's display(), ledmatrix.service
stays active and the panel stays frozen. This adds a way to detect that.

- src/display_watchdog.py (standard library only) sends sd_notify over
  $NOTIFY_SOCKET and writes /run/ledmatrix/display-heartbeat.json. Only the
  render thread counts: beats from other threads are ignored.
- ledmatrix.service: WatchdogSec=120, NotifyAccess=main,
  RuntimeDirectory=ledmatrix (0755), RestartSteps=4 and
  RestartMaxDelaySec=2min. It stays Type=simple. run.py widens the watchdog
  to 15 min for start-up, and load_plugin() does the same on the render
  thread. The loop arms after its first frame.
- /api/v3/health adds checks.display_loop: running, stalled (no heartbeat
  for over 60s, which makes the status degraded) or not_reported. With web
  login on, a caller who is not logged in still gets only healthy/degraded,
  and a stall degrades that answer.
- The update verifier requires a fresh heartbeat from the restarted display
  when the display it replaced was writing one. A frozen panel is rolled
  back.
- Existing installs get the systemd watchdog only after install_service.sh
  is re-run. The heartbeat works right away.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 11:15:31 -04:00
ChuckandClaude Opus 5.5 b09434a418 refactor(plugins): the display publishes plugin runtime state; retire plugin_state.json (#690)
Stage 2 of the web plugin catalog, after #688.

- The display publishes a plugin runtime snapshot (plugin_runtime.py) to
  the shared cache: per plugin loaded, lifecycle state, a short redacted
  error summary, the version it loaded and when, plus published_at /
  stale_after / running. Written on change (throttled to 10 s; the
  RUNNING/ENABLED flip of an ordinary update is not a change) and once a
  minute otherwise; cleanup() publishes running: false.
- The web reads it back and restores loaded / state / error_info in
  /api/v3/plugins/installed (plus loaded_version, loaded_at and
  data.runtime). Only a live snapshot counts; stale, stopped or missing
  answers null and says which.
- data/plugin_state.json is retired: every reader and writer moved to
  config + disk (desired) or the snapshot (observed). Nothing in it was
  non-derivable, so nothing is migrated and an existing file is left
  unread. The web-side PluginStateManager (state_manager.py) is removed;
  the display's plugin_state.PluginStateManager is the only state machine.
- StateReconciliation compares config + disk with the snapshot, reporting
  enabled-but-not-loaded and older-version-loaded as no_action findings.
- Backups list installed manifests with enabled from config.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 10:48:14 -04:00
ChuckandClaude Opus 5.5 7ab6fb1aff refactor(web): read plugins through a PluginCatalog; only the display runs them (#688)
The web process built its own PluginManager and loaded plugins into itself:
store installs and updates loaded or reloaded a web-side copy, and config
saves and enable/disable called on_config_change, on_enable and on_disable
on it. None of that reached the panel, and /plugins/installed reported
runtime state from those copies.

- Add PluginCatalog (src/plugin_system/plugin_catalog.py): manifests,
  directories, display modes, installed version, schema and config reads,
  with no way to run a plugin. app.py and both blueprints use it; the
  plugin_manager blueprint attribute is gone.
- Remove every lifecycle call from the web routes. Config changes already
  reach the display through ConfigService (on_config_change) and the
  enabled-set reconcile.
- Health and metrics readers move to api_v3.health_tracker /
  resource_monitor. /plugins/installed reports loaded/state/error_info as
  null (the display does not publish them) and enabled by the display's
  rule.
- Store install, update and uninstall answer restart_required when the
  running display will not pick the change up by itself
  (display_restart_required). The restart banner follows the flag via
  window.noteRestartRequired instead of the /config/main URL heuristic;
  /config/main now sends restart_required: true.
- The one remaining in-process import of plugin code (Starlark helper
  modules, oauth_flow action scripts) goes through
  _import_plugin_code_in_web_process() until a web-entry contract.
- /plugins/installed reports vegas_participation (from #682) from the
  user's setting or the manifest, with vegas_participation_source; when
  only the plugin's code decides it, null with source 'runtime', since the
  web process no longer has plugin instances to ask.
- Check & Update All keeps its restart flags when the final list refresh
  fails, and asks for a restart when an enabled plugin's first request got
  no answer and the re-sent one found it up to date.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 10:39:44 -04:00
ChuckandClaude Opus 5.5 ba6eccb489 build(web): generate the UI's Tailwind CSS with the pinned standalone CLI (#685)
Replaces the hand-written Tailwind subset in app.css with a real, purged
Tailwind build: scripts/build_css.py runs the pinned, SHA-256-checked
standalone Tailwind CLI (no Node), the generated tailwind.css and
plugin-frame.css are committed, and CI fails when they are stale. The Pi
never builds anything. The login page (#683) now links tailwind.css too,
and the load-order test covers every template that links app.css.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:38:16 -04:00
ChuckandClaude Opus 5.5 5ea0d511dc feat(store): read ledmatrix_min_version, aliases and commit from the registry (#686)
The store reads three optional registry fields: ledmatrix_min_version
(an incompatible install/update is refused before any download, with a
"Needs LEDMatrix X+" card badge), aliases (update/uninstall/reinstall by
registry id find a plugin installed under its manifest id, with registry
proof only), and commit (shown and linked on the store card). An older
plugins.json behaves as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:31:16 -04:00
ChuckandClaude Opus 5.5 1c928b2033 feat(vegas): one declared participation per plugin (scroll | pause | exclude) (#682)
A plugin takes part in Vegas mode in one declared way: 'scroll', 'pause'
or 'exclude', resolved from the user's vegas_participation setting, the
manifest field, then the legacy hooks, so no plugin changes behaviour.
The stream manager decides inclusion and pauses through it; the installed
plugins API and the Vegas plugin-order list report it. Deprecates
get_supported_vegas_modes, get_vegas_segment_width and vegas_panel_count
for removal in 3.9.0, and regenerates docs/DEPRECATIONS_3.8.md to include
them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:30:37 -04:00
ChuckandClaude Opus 5.5 b2df0fda1b docs(sports): record the ufc round-break verify outcome (refuted) (#692)
ESPN sends the break between rounds as STATUS_END_OF_ROUND with
displayClock "-", not "0:00", so the shared game-over rule never
drops a five-round fight at the round 4 break. Verified against
recorded payloads in ChuckBuilds/ledmatrix-plugins#580, which pins it.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:20:30 -04:00
ChuckandClaude Opus 5.5 15c61def67 fix(web): make the SSE streams' 200/min rate limit actually apply (#691)
app.py called limiter.limit("200 per minute")(stream_x) after the routes
were registered and discarded the result. flask-limiter 3.x enforces a
decorated limit in the wrapper limit() returns, and marks the original
function so the before_request middleware skips it, so the streams had
no limit at all -- not even the 1000/min default. Register the wrapper
as the view instead.

The new test (skipped without flask-limiter) reconnects to each stream
201 times and expects the last to get a 429; it fails on the old code.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:12:23 -04:00
ChuckandClaude Opus 5.5 c8a0ddcf7b chore(deprecation): retarget the 35 deprecations to 3.8.0 and add a usage scan (#681)
Moves the 35 @deprecated markers from 3.7.0 (already shipped with them in
place) to 3.8.0, and adds scripts/plugin_api_usage.py plus the generated
docs/DEPRECATIONS_3.8.md: who still calls or overrides each deprecated
method across core, the monorepo and third-party plugins. Removes nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:11:36 -04:00
ChuckandClaude Opus 5.5 9fe23af432 docs(sports): reconcile-then-promote roadmap, drift report and report-only CI job (#680)
Rewrites the roadmap in docs/SPORTS_UNIFICATION.md for the
reconcile-then-promote decision (stages 0-3 recorded as done), and adds the
sports drift report script with a report-only CI job.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:11:02 -04:00
ChuckandClaude Opus 5.5 c0d97e4867 fix(plugins): one hung plugin no longer stops every plugin from updating (#677)
The update worker no longer blocks forever on a plugin whose display()
never returns. It waits at most PLUGIN_LOCK_TIMEOUT (5s) for a plugin's
lock, then skips that plugin's update (a report-only "busy skip" in
health) and keeps updating every other plugin. display() frames are timed
(slow calls logged and counted; calls past the executor timeout recorded as
hangs), a hung update() is recorded, and on_config_change() now runs under
the plugin lock or is deferred to the worker. The plugin-facing API is
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:10:41 -04:00
ChuckandClaude Opus 5.5 e3c85cece6 feat(web): optional web login and API tokens, off by default (stacked on #674) (#683)
Optional web login, off by default: a device that sets no password behaves
exactly as before. Set under General > Security; then every page and API
route needs a session login or an API token (Authorization: Bearer).
Loopback, the Wi-Fi setup flow in AP mode, static files, captive-portal
probes and a reduced /api/v3/health stay open. Secrets live in the web_auth
section of config_secrets.json and no API returns them.
scripts/reset_web_password.py turns login off. Stacked on #674.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:09:40 -04:00
ChuckandClaude Opus 5.5 ba38a83c2c fix(web): refuse cross-site state-changing requests (Origin/Referer check) (#674)
The web interface refuses state-changing requests (POST/PUT/PATCH/DELETE)
whose Origin (or, without one, Referer) is not the host they were sent to,
or is null: 403 CROSS_SITE_REQUEST (web_interface/origin_guard.py). Any
website a LAN user visited could otherwise make their browser POST a plain
form to the Pi. /api/v3/system/action also refuses form-encoded and
text/plain bodies (415) unless sent by HTMX. Clients that send no Origin or
Referer (curl, requests, Home Assistant, the MQTT bridge) are unaffected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 09:02:01 -04:00
ChuckandClaude Opus 5.5 c3a7a110c4 fix(display): on-demand loads a disabled plugin live instead of failing (#678)
* fix(web): on-demand no longer restarts a running display service

POST /display/on-demand/start treated start_service (default true, sent by
"Preview on display", the on-demand dialog and the MQTT bridge) as
"restart": with the service running it ran systemctl stop, slept 1.5s and
started it again. Every request cold-started the display process -- every
plugin reloaded, panel blank -- to deliver a request the running process
already reads from the cache mailbox every ON_DEMAND_POLL_INTERVAL (0.25s),
including mid-dwell, mid-screen and mid-Vegas. The restart bought nothing:
startup only restores a session the display saved itself
(display_on_demand_config), so the new request arrived through the same
mailbox either way.

start_service now means "start it if it is not running". The stop route
coerces stop_service to a boolean so "false" no longer stops the service.
test_api_v3_on_demand_restart.py pinned the old restart path; it now pins
the replacement. Docs updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): on-demand loads a disabled plugin live instead of failing

The display process only loads enabled plugins, so an on-demand request for
a disabled one -- "Preview on display" offers it on every config page, with a
note that the plugin will be enabled for the preview -- failed with
invalid-mode. Nothing enabled it short of a restart, and the on-demand route
no longer restarts the service.

_activate_on_demand now loads an installed-but-not-running plugin through
the live-enable path (load_plugin + _register_loaded_plugin), with a new
load_plugin(force_enabled=True) so the instance runs enabled while
config.json keeps saying disabled. The plugin is tracked in
_on_demand_loaded_plugins, and the main loop unloads it through
_unregister_plugin once on-demand moves off it (stop, expiry, another
request, or a failed request that ends the session) -- right after its own
poll, where no display() is on the stack. A failed load publishes status
error with load-failed. A plugin enabled during the session stays loaded.

A session restored after a restart uses the same tracking instead of
setting enabled in the config dict config_manager caches, so its plugin is
unloaded when the session ends rather than staying loaded until the next
restart. Ending a session no longer resumes the rotation onto a plugin that
is about to be unloaded, which a restored session did.

Also: a stop sent while on-demand is inactive clears a failed request's
error, instead of /display/on-demand/status reporting status: error until
the state aged out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 17:59:12 -04:00
ChuckandClaude Opus 5.5 6047eb5e4e test: stop the suite reinstalling plugins into the real plugin-repos/ (#679)
Any test that imported web_interface.app and sent a request fired the app's
startup reconciliation, which runs against the checkout's real config.json
and plugin-repos/ and reinstalls every configured-but-missing plugin from the
live store. A full Windows run left basketball-scoreboard, calendar,
football-scoreboard, leaderboard and ledmatrix-stocks untracked in
plugin-repos/ (not gitignored) from that daemon thread.

test/conftest.py now installs an import hook that sets the app's run-once
_reconciliation_started latch as the module finishes executing, so lazy
imports, module-level imports and reloads all start disarmed.
StateReconciliation's own tests are unaffected. A regression test pins that
a request to the imported app launches no reconciliation thread.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 17:47:32 -04:00
ChuckandClaude Opus 5.5 c4c46d3ba7 fix(web): on-demand no longer restarts a running display service (#676)
POST /display/on-demand/start treated start_service (default true, sent by
"Preview on display", the on-demand dialog and the MQTT bridge) as
"restart": with the service running it ran systemctl stop, slept 1.5s and
started it again. Every request cold-started the display process -- every
plugin reloaded, panel blank -- to deliver a request the running process
already reads from the cache mailbox every ON_DEMAND_POLL_INTERVAL (0.25s),
including mid-dwell, mid-screen and mid-Vegas. The restart bought nothing:
startup only restores a session the display saved itself
(display_on_demand_config), so the new request arrived through the same
mailbox either way.

start_service now means "start it if it is not running". The stop route
coerces stop_service to a boolean so "false" no longer stops the service.
test_api_v3_on_demand_restart.py pinned the old restart path; it now pins
the replacement. Docs updated.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 14:42:58 -04:00
ChuckandClaude Opus 5.5 da9a999102 chore: prepare the 3.7.0 release (#673)
Bumps src.__version__ to 3.7.0 and turns Unreleased (#672: sports_celebration,
sports_fetch and sports_card_wrappers) into ## 3.7.0; src/common/README.md and
docs/SPORTS_UNIFICATION.md say 3.7.0 for the three modules.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 12:51:24 -04:00
ChuckandClaude Opus 5.5 1e4c890d59 feat(common): sports_celebration, sports_fetch and sports_card_wrappers, promoted from the scoreboards (sports consolidation stage 3) (#672)
Three new hardware-free modules holding code the scoreboard plugins carry
as identical copies (executable AST, docstrings stripped, checked across
every carrying plugin at ledmatrix-plugins 30455671). The bodies are the
plugins'; the changes are type annotations for the mypy ratchet, the
colour helpers losing their leading underscore as public free functions,
and two comments that described the plugins' files.

- src/common/sports_celebration.py: SportsCelebrationMixin, the score/win
  takeover drawn by afl, football, hockey, nrl and soccer
  (_draw_celebration_layout and the palette, backdrop, scenery, confetti,
  crest and _fit_font steps, with their class constants), plus the colour
  helpers (logo_palette, lift_color, cap_luminance, mix_color, ...). Only
  the drawing: _start_celebration, _check_for_goal/_check_for_score,
  _check_for_win and display() differ between the plugins and stay there.
- src/common/sports_fetch.py: SportsFetchMixin, the four SportsCore methods
  identical in all nine scoreboards: _fetch_season_directly,
  _background_fetches_espn_ranges, _needs_previous_day and
  _wants_live_odds, with _LOOKBACK_CUTOFF_HOUR and _LIVE_ODDS_LOOKAHEAD.
  _get_timezone, _extract_game_details and _fetch_data are as identical
  and stay behind, for the reasons sports_shared gives (a per-plugin
  import; the abstract contract); so does SportsUpcoming.__init__, since
  no src/common mixin has a constructor.
- src/common/sports_card_wrappers.py: SportsCardWrappersMixin, the
  seventeen sports_card delegations the eight game renderers carry (15 in
  all eight, 2 in all but football, whose own versions override them).
  _schema_font_size/_resolve_font_size look identical but read each
  plugin's own _SCHEMA_PATH, so they stay.

Each mixin has no __init__ and creates no attributes (the host contract is
declared as annotations only), defines no name the mixins beside it
define, and documents the attributes it reads; a host-contract test
parses each and fails on an undocumented read. A method kept on a
plugin's class wins over the mixin's.

Tests: behaviour ported from the plugins' celebration, odds, lookback and
date-range tests against stub hosts carrying exactly the contract, with
crests drawn by the test (test_sports_celebration.py, test_sports_fetch.py,
test_sports_card_wrappers.py), and test_sports_stage3_parity.py, which with
LEDMATRIX_PLUGINS set compares every body with every plugin copy that is
left (58 pass against the plugins today; a copy that is gone counts as
adopted). All three modules are on the mypy ratchet, in
src/common/README.md, the CHANGELOG's Unreleased section and
SPORTS_UNIFICATION's module table. Nothing in core uses them yet.

Full suite: the same 67 failing test ids as main (Windows-only), 77 more
passing.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 12:38:40 -04:00
ChuckandClaude Opus 5.5 7f96075076 chore: prepare the 3.6.2 release (#671)
Bumps src.__version__ to 3.6.2 and turns Unreleased (#670, the favourite
check's false "season has finished" for list-calendar competitions between
rounds) into ## 3.6.2.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 11:55:48 -04:00
ChuckandClaude Opus 5.5 439013b18c fix(common): favourite check no longer calls the Europa League finished between matchdays (#670)
On 2026-09-29 ESPN's uefa.europa scoreboard still showed the 17 September
matchday, so every event was past. Its calendar is a "list" of rounds
(League Phase to 30 Jan 2027, then the knockout rounds to the final), not
a match-day whitelist, and the league's season type is a soccer id rather
than 2/3, so neither 3.6.1 rule applied and the check said the season had
finished.

When every event is past, a round in a list calendar that has not started
yet now draws no conclusion. Only a round's start date counts: end dates
are padded past the last game (AFL's Grand Final round still had a day to
run three days after the Grand Final), and rounds in an offseason phase
(college football's All-Star week) are skipped. Season end dates are still
ignored, so PLL (season to 2027-01-01) stays "finished", as do the World
Cup and AFL. Of 28 live ESPN scoreboards only uefa.europa's message changes.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 11:49:36 -04:00
ChuckandClaude Opus 5.5 e5bbfa2ae3 chore: prepare the 3.6.1 release (#669)
Bumps src.__version__ to 3.6.1 and records #667 (the favourite check's false
"season has finished") under ## 3.6.1; #667 had no CHANGELOG entry. Plugins
that drop their bundled favourite-check copy floor on 3.6.1.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:59:20 -04:00
ChuckandClaude Opus 5.5 db49275075 fix(common): favourite check no longer calls a started postseason a finished season (#667)
* fix(common): favourite check no longer calls a started postseason a finished season

The day after a regular season ends, ESPN's default scoreboard still
returns that last regular-season day, while leagues[0].season has moved
to Postseason. All events were in the past, so the check logged "the
season has finished" for MLB on 2026-09-29 while the upcoming manager in
the same process was showing TB's wild-card games.

When every event is past and the league is in a later in-season phase
(regular season or postseason) than all of the returned events, draw no
conclusion. The offseason is excluded, so a genuinely finished season is
still reported as finished.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(common): favourite check reports the next matchday between soccer rounds

Between matchdays ESPN's soccer scoreboard keeps showing the last one, so
every event is in the past and in the league's current phase, which the
postseason rule does not cover; the check said the Premier League season
had finished on 2026-09-29 (last games 20 September, next 10 October).
When the league calendar is a "day" whitelist, its entries are days with
games, so a future one is used as the next fixture. MLB's day calendar is
a blacklist and is not read that way; PLL's whitelist has no future days
and is still reported as finished.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:52:38 -04:00
ChuckandClaude Opus 5.5 a11412dabb chore: prepare the 3.6.0 release (#666)
Turns the CHANGELOG's Unreleased section into ## 3.6.0 and bumps
src.__version__, the value plugin ledmatrix_min_version floors compare
against. 3.6.0 ships the two modules from #665 (favorite_team_check,
sports_timezone); nothing else has changed since 3.5.0. src/common/README.md
and docs/SPORTS_UNIFICATION.md say 3.6.0 for them instead of Unreleased.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 09:28:21 -04:00
ChuckandClaude Opus 5.5 fe5bed2886 feat(common): favorite_team_check and sports_timezone, promoted from the scoreboards (sports consolidation stage 2) (#665)
* feat(common): favorite_team_check and sports_timezone, promoted from the scoreboards (sports consolidation stage 2)

Two new hardware-free modules, taken from files the scoreboard plugins carry
as copies:

- src/common/favorite_team_check.py: FavoriteTeamCheck(logger, leagues), the
  seven byte-identical <sport>_favorite_check.py copies. Same code; the only
  additions are two type annotations (for the mypy ratchet).
- src/common/sports_timezone.py: resolve_timezone_name(), resolve_timezone(),
  system_timezone_name(), from the ten <sport>_timezone.py copies. They
  differed only in the plugin label named in the nothing-resolved warning and
  the write-back-bug values, which become keyword-only arguments
  (plugin_label, writeback_fixed_in). Same resolution order and log text.

Tests are ported from the plugins' own (test_favorite_check.py,
test_schedule_note_uses_game_dates.py, test_timezone_resolution.py; the
timezone ones run once per plugin's values and pin the exact warning text).
Both modules are on the mypy ratchet, in src/common/README.md, the CHANGELOG's
Unreleased section and SPORTS_UNIFICATION's module table. Nothing in core uses
them yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(common): bdf_font and json_body shipped in 3.5.0

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(common): annotate the favourite check's deliberate except/pass for Bandit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 09:05:10 -04:00
ChuckandClaude Opus 5.5 5b30052b59 docs(changelog): fold Unreleased into 3.5.0 for the release (#664)
* docs(changelog): fold Unreleased into 3.5.0 for the release

Every Unreleased entry (#605-#663) moves into the 3.5.0 section, grouped
with the existing 3.5.0 areas; new groups for Display and Vegas, Plugin
error reporting, Wi-Fi, Fonts and Removed. "## Unreleased" stays as an
empty heading.

Module list: add src/common/json_body.py (espn_dates imports it with a
fallback) and src/common/bdf_font.py; list the other modules new since
v3.4.0 as core-internal; add the new names in existing modules
(handles_espn_date_ranges, register_plugin_fonts(plugin_dir),
forget_manager_fonts). Record the src.common and plugin_system modules
#608 deleted.

Add entries for merged PRs that had none: #604, #605, #606, #607, #608,
#609, #613, #616, #618, #622, #625, #628, #630, #633. Note that three
scripts named in older 3.5.0 entries were later deleted by #607.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): list the hardware-free test under developer tools, not plugin modules

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 16:05:22 -04:00
ChuckandClaude Opus 5.5 8557eff88a fix(plugins): put a (re)loading plugin's directory first on sys.path (#663)
Plugins import their own files by bare name (`from sports import ...`),
which resolves to the first directory on sys.path that has the file. The
loader added a plugin's directory only if it was missing, so on a reload --
a live re-enable from the web UI -- the plugin's directory stayed behind
every plugin loaded since, and its bare imports found their files first.

Seen on ledpi: re-enabling UFC with hockey running failed with "cannot
import name '_status_is_final' from 'sports'" (it got hockey's sports.py).
A loading plugin's directory is now always moved to the front. Every
scoreboard ships its own sports.py, so any of them was exposed on reload.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 15:02:01 -04:00
ChuckandClaude Opus 5.5 b8c01c69fb ci: mypy ratchet -- keep type-clean modules clean (71 modules, 536 -> 442 errors) (#661)
* ci: mypy ratchet -- keep type-clean modules clean

mypy-clean.txt lists the 71 modules under src/ that type-check clean;
scripts/check_types.py runs mypy (--follow-imports=silent) on exactly
those files and fails on any error or a missing/unsorted/duplicate entry.
A new "Type check (mypy ratchet)" CI job runs it with mypy 1.20.2 and
pinned stubs; the manual pre-commit mypy hook now runs the same script
(a local hook, so mypy sees the installed requirements like CI does).

35 modules were made clean with annotation-only fixes: hints, typing.cast,
TYPE_CHECKING imports, implicit-Optional defaults made explicit, and
annotations widened (never guards removed) where mypy called a defensive
isinstance check unreachable. No runtime behaviour change.

mypy.ini: numpy and orjson are treated as Any (follow_imports=skip, also
for stubs). numpy 2.3+ stubs use 3.12 `type` statements that mypy won't
parse at python_version 3.10, and orjson is optional, so seeing its stubs
made the result depend on whether it was installed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: annotate check_types.py's list-form mypy subprocess

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 15:01:35 -04:00
ChuckandClaude Opus 5.5 e6e0a16140 ci: run the web UI DOM test suites; fix two stale suites (#660)
A new "Web UI JS tests" job installs jsdom, starts the web interface in
emulator mode and runs test/js/run_all.js with REQUIRE_DOM=1, which makes a
DOM suite that can't run a failure rather than a silent skip. (The unit
suites were already covered through pytest.)

Two suites failed against main when run for real:
- test_tools_sections rendered the Tools partial without LEDEscape, which
  base.html's app-early.js defines; it now installs it in beforeParse, and
  supplies two sample Starlark apps (one id with a quote) when the server
  has none, instead of assuming a device with apps and Pixlet.
- test_store_dom assumed the live registry had at most 48 plugins; it now
  checks pagination whichever side of 48 it is.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 15:01:08 -04:00
ChuckandClaude Opus 5.5 989eae9405 refactor(plugins): split PluginStoreManager into mixins (#659)
* refactor(plugins): split PluginStoreManager into mixins

src/plugin_system/store_manager.py (2,977 lines) keeps the class, its
shared state, locks, the uninstall registry, directory lookup and
uninstall; its methods are split by area into:
- store_registry.py (_RegistryMixin): registry, GitHub metadata, search,
  manifest validation
- store_install.py (_InstallMixin): install paths and dependencies
- store_update.py (_UpdateMixin): updates, rollback, local git state

Pure move: all 56 members are byte-identical (checked with ast) and the
assembled class has exactly the same attributes as before (checked at
runtime). PluginStoreManager is imported from store_manager.py as before.
Tests that patched shared modules (subprocess, requests, tempfile, shutil)
through store_manager now reach them through the module whose code they
exercise; a source-text contract test reads all store_*.py modules.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: annotate findings the split moved into new store modules

subprocess imports and a list-form git clone (no shell), and the config
template's placeholder token string -- existing code that Codacy reported
as new because it moved. Annotated with the repo's nosec/nosemgrep style.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: annotate the default-branch git clone the split moved

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 15:00:45 -04:00
ChuckandClaude Opus 5.5 d469fe39d2 fix(odds): don't return the cached no-odds marker as odds (#662)
A game ESPN had no odds for is cached as {"no_odds": True}, so it isn't
re-requested on every update. On the next update get_odds() returned that
marker from the cache as if it were odds: a truthy dict that callers took
to mean the game had some. It's still a cache hit (its ttl decides when to
ask again), but get_odds() now returns None for it -- on the cache hit and
in the stale-cache fallback after a failed fetch -- as the plugins' bundled
copies already did.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 14:59:40 -04:00
ChuckandClaude Opus 5.5 09103a8a7d refactor(web): split api_v3/plugins.py by area (#658)
* refactor(web): split api_v3/plugins.py by area

web_interface/blueprints/api_v3/plugins.py (3,285 lines) becomes:
- plugins.py: installed list, enable/disable, plugin actions
- plugin_store.py: install, update, uninstall, store, saved repositories
- plugin_config.py: config get/save, schema, reset
- plugin_assets.py: asset uploads and plugin static files
- plugin_health.py: health, metrics, limits
- plugin_operations.py: operation history, state reconciliation
- plugin_calendar.py: calendar credentials and auth

Pure move: all 44 functions and 38 route decorators are byte-identical
(checked with ast), URLs and endpoint names are unchanged (url-map test).
Each module imports only what it uses. Tests and config.py that reached
into plugins.py for moved names now import from the new module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): keep exception text out of calendar responses; annotate moved code

The split made scanners report existing findings in the moved code as new:
- CodeQL: the calendar auth and calendar-list routes returned exception
  text (redacted, but still derived from the exception). Both now log the
  exception and return a fixed message pointing at the log.
- MD5 in the asset upload only makes a filename unique: usedforsecurity=False.
- pickle reads/writes the calendar plugin's own OAuth token (as before):
  annotated. Token-status labels and a log line naming the secrets path are
  false positives: annotated with the repo's nosec/nosemgrep convention.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): name uploaded assets with SHA-256 instead of MD5

The hash only makes an uploaded image's filename unique. Codacy flags MD5
even with usedforsecurity=False, and SHA-256 does the job as well; existing
files keep their names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): keep the redacted exception detail in calendar errors

Reverts the calendar part of 5e695b7c. The project's policy
(test_no_api_v3_handler_discards_its_exception) is that an API error
carries the redacted exception detail -- describe_exception runs it
through the credential redactor -- so a failure is diagnosable from the web
UI. Dropping it for CodeQL broke that; CodeQL can't see the redaction, so
its two alerts here are false positives, like the existing ones on main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 14:58:53 -04:00
ChuckandClaude Opus 5.5 724673ba0b fix: one retry layer for background fetches; CI installs the web requirements; one Discord invite (#657)
- BackgroundDataService: the session adapter retried connection errors 3x
  inside each attempt of the service's own retry loop (up to 16 connection
  attempts per request on a dead network). The adapter no longer retries;
  ESPN date chunks, which bypass the loop and skip a failed chunk, get a
  small connection retry of their own (_ConnectionRetryingSession).
- CI installs web_interface/requirements.txt. The brotli header test now
  checks its intent (core never hand-sets br; requests may advertise it when
  a decoder is installed) instead of failing whenever brotli is present.
- Every Discord link uses the LEDMatrix server's invite (RdrC37rEag).

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 11:08:43 -04:00
ChuckandClaude Opus 5.5 6cfcf2e384 fix(web): plugin dir resolver in routes, nmcli AP detection, daemon config reload, upload safety, BDF preview (#655)
* fix(web): plugin dir resolver in routes, nmcli AP detection, daemon config reload, upload safety

- Route plugin lookups (installed list, update, recorded version, config
  form, web UI pages) through the plugin manager's resolver so plugins in
  ledmatrix-<id> directories work.
- Captive-portal detection also sees the nmcli fallback AP (cached).
- WiFi monitor daemon re-reads wifi_config.json when its mtime changes.
- Drop the AP check in disconnect_from_network that could never fire.
- LED status file per WiFiManager; config path falls back to this checkout.
- BDF font preview via src.common.bdf_font.
- Asset uploads validate every file before saving; metadata and calendar
  credentials written atomically; no absolute path in the response;
  asset delete answers 400 for a missing body.
- Coerce string booleans in plugin toggle, on-demand start and AP force.
- SSE broadcaster clears its thread handle before exiting.
- start.py log filter handles every exc_info form.
- Cleanups: unused plugins/fonts partial work, duplicate backup catch-alls,
  raw-config error helper, update-route tidy, redundant imports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): request BDF font previews now that the server renders them

The Fonts tab skipped the preview request for .bdf files because the server
used to refuse them; /fonts/preview now draws BDF with the shared loader.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): take the update route's plugin directory from a directory listing

CodeQL flagged the path built from the request's plugin_id (the id was
already validated with safe_path_component, which CodeQL doesn't model; the
same flow on main is alerts 738/739). The directory is now the entry of
plugins_dir matched by name, so nothing built from user input reaches the
filesystem; an id with nothing installed goes to the store manager, which
reports it not found as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): read the blueprint's plugin_manager defensively in _plugin_directory

_get_plugin_version now goes through _plugin_directory, which read
api_v3.plugin_manager directly; the attribute exists only once the app sets
it, so test_path_traversal_guards::test_a_real_manifest_is_read failed
when run on its own (order-dependent in the full suite).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:42:07 -04:00
ChuckandClaude Opus 5.5 c00bf5e8e6 fix(plugin-system): unload/update race, failed-load cleanup, limits validation, schema lookup, install rollback (#653)
* fix(plugin-system): unload/update race, failed-load module cleanup, limits validation, schema lookup, install rollback, op-queue dedupe

- unload_plugin takes the per-plugin lock (5s bounded) before cleanup(),
  and an update() that finishes after its plugin was unloaded no longer
  sets the state back to ENABLED.
- A load that fails after import drops plugin_<id> and its submodules
  and forgets its manager fonts, so a fixed plugin reloads new code.
- Resource limits are validated as non-negative numbers: 400 at
  POST /plugins/limits, bad cached records ignored with one warning.
  Route docstrings note health/metrics reset and limits only change the
  web process's view.
- SchemaManager.get_schema_path resolves each search dir via
  resolve_plugin_dir (manifest id, ledmatrix-<id>) before the literal
  paths; plugins/ still before plugin-repos/. Misses cached 30s and
  logged once at DEBUG.
- install_from_url sets an existing copy aside and restores it if the
  move fails, under the per-plugin reinstall lock.
- Operation queue refuses a second pending op for a plugin and trims
  _operations with history.
- get_vegas_render_width reads display_manager.width first.
- get_logger in store/schema/health/resource/saved_repositories;
  UTF-8 reads in store_manager and state_manager.
- Docs: update_interval precedence (manifest over config) stated where
  users are told to set it in config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): build the limits 400 message from the field name, not an exception

CodeQL flagged str(e) flowing into the response. invalid_limit_field()
returns the offending field without raising, and limits_from_dict uses it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:41:40 -04:00
ChuckandClaude Opus 5.5 0e9e2cabba fix(web): widget cache-busting, dead frontend code, and dependency pins (#656)
- Plugin-supplied widgets load as /static/plugin-widgets/...js?v=<plugin
  version>, so an update isn't hidden behind the year-long immutable cache.
- Fire-and-forget loadInstalledPlugins() calls catch the rejection it has
  already reported, so the global handler no longer adds a second toast.
- Timezone picker renders again when the General partial is re-injected.
- Remove dead code: executePluginAction's six plugin-id fallbacks and
  [DEBUG] logging, window.currentPluginConfig and every read of it, the
  file-upload JSON delete branch, unused PluginAPI / PluginInstallManager /
  PluginStateManager helpers, loadPluginWidgetsFromManifest, the stale
  install_manager.js and LEDVisibility fallbacks, error_handler.js's global
  escapeHtml, 13 unused CSS rules, and stale comments/no-op returns.
- pytz < 2027, psutil < 7 in requirements-test.txt, pytest-cov < 8.
- Pin anthropics/claude-code-action to the commit v1 resolves to.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:40:52 -04:00
ChuckandClaude Opus 5.5 65d82580bc test(starlark): cover the review fixes #535 shipped without tests (#650)
Six Starlark fixes are on main via #535 and #537, but a follow-up commit
carrying tests for half of them was pushed to fix/starlark-pixlet-install
six minutes after #535 merged, so those tests never landed. This ports
them onto the api_v3 package split:

- a failed toggle write answers 500, and a loaded app is not flipped in
  memory when the manifest write fails
- each manifest writer gets its own temp file; concurrent writes leave
  readable JSON; no temp files are left behind
- a failed dynamic import of tronbyte_repository / pixlet_renderer does
  not stay cached in sys.modules
- a failed save_config() leaves config and timing untouched and does not
  re-render; a successful save still applies

It also logs when the timing update to the manifest is not persisted.
_update_manifest_safe answers False rather than raising, so the existing
except branch never saw that failure.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:40:32 -04:00
ChuckandClaude Opus 5.5 f6c0fe55d9 fix(core): font zip cache, monotonic timers, resolver back-off, and other core/common fixes (#654)
* fix(core): font zip cache, monotonic timers, resolver back-off, and other core/common fixes

- font_manager: a .zip font URL is served as its extracted font after a
  restart (the cached-file check returned the archive first); downloads
  use requests with a 30s timeout into a temp file + os.replace.
- api_helper / sync_manager: rate-limit and heartbeat/leader timeouts use
  time.monotonic(); last_request_time and the status file's ts stay
  wall-clock. set_on_new_cycle docstring no longer claims core uses it.
- logo_helper: the placeholder uses the same scaled box as a real logo.
- permission_utils: one _sudo_bash_candidates() helper (with the sudoers
  exact-argv rationale) shared by sudo_remove_directory, which now retries
  the next bash path on a sudo refusal, and install_requirements_file.
- dynamic_team_resolver: failed/empty fetch backs off 5 min; duplicate
  INFO log and contradictory docstring example fixed.
- element_style: scale default looked up through element aliases.
- background_data_service: cache-hit callback runs outside the lock.
- config_arrays: union-aware type check (["array","null"]); stale
  dotToNested() reference removed.
- auto_update_setup: non-dict auto_update reads as off; temp result file
  unlinked when the write fails.
- exceptions: constructors copy the caller's context dict.
- logging_config: StructuredFormatter json.dumps(default=str).
- error_aggregator: removed unused export_path/export_to_file/_auto_export.
- Docstrings: validate_file_upload max_size_mb, raise_on_errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sync): retry the status-file rename like the other atomic writers

On Windows os.replace can fail with "Access is denied" while a scanner
briefly holds the target open; config_manager_atomic._replace already
retries that (and re-raises at once on other platforms). The sync status
writer called os.replace directly, which made
test_concurrent_writers_each_use_their_own_temp_file flaky on Windows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:40:16 -04:00
ChuckandClaude Opus 5.5 6f45ff5e63 fix(display): thread-safety for deferred updates, BDF faces and follower image; one refresh default (#652)
- DisplayManager.defer_update()/process_deferred_updates(): one lock around
  every queue mutation (appends from the update thread were lost to the
  render thread's filter/slice reassignments); callables run outside it.
- FontManager and element_style no longer cache BDF freetype.Face objects
  process-wide (load_bdf_face caches them per thread); element_style's LRU
  is locked against get/move_to_end vs eviction races.
- limit_refresh_rate_hz default is one constant, DEFAULT_REFRESH_LIMIT_HZ =
  100 (the template's), for the library options, refresh_hz, the matrix
  guard, Vegas and scroll_config. Previously a missing key capped the panel
  at 90 while pacing assumed 100.
- Sync follower: the TCP thread queues the leader's scroll image; the render
  thread swaps image/array/width in between frames.
- update_display() error log rate-limited (traceback first, then once a
  minute with a count); swallowed DisplayController exceptions log at DEBUG.
- Root display_controller.py runs run.py via runpy.
- stream_manager: correct the RLock release comments; merge duplicate if.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:39:40 -04:00
ChuckandClaude Opus 5.5 1e62677257 fix(vegas): pause for STATIC plugins where their turn falls in the strip (#651)
The static trigger peeked at the front of StreamManager's segment buffer,
which continuous scrolling (the default) never advances -- it extends the
strip with take_next_group() -- so the same first segment was examined on
every frame. A STATIC plugin paused the scroll only if it was first, once,
at startup; otherwise it scrolled past as ordinary content. Swap mode had
the same problem for any STATIC plugin not first in its cycle.

The render pipeline now records a marker (strip column, plugin id) for
each STATIC plugin where the strip is built -- composition and every
extension -- shifts the markers when the scrolled prefix is trimmed, and
clears them on reset. The coordinator pauses when the scroll reaches the
next marker: a tuple comparison per frame instead of a lock, a plugin
lookup and a get_vegas_display_mode() call. take_next_group() no longer
renders STATIC plugins' content. The pause calls display() under the
plugin lock and is timed with the monotonic clock.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 10:39:14 -04:00
ChuckandClaude Opus 5.5 224847cebc fix(display): stop the run loop spinning when no mode has anything to show (#649)
* fix(display): stop the run loop spinning when no mode has anything to show

A mode whose display() reports nothing rotates to the next at once, with no
dwell. With every enabled mode empty (only a sports plugin in its
off-season, say) the loop went round with no sleep: on ledpi, 169% CPU and
~1,800 "No content" log lines every 10 seconds. After one full rotation of
empty passes it now pauses EMPTY_ROTATION_PAUSE (1s) per pass, servicing
plugin updates and returning early on on-demand or schedule changes; live
priority is still checked at the top of every pass, and the streak resets
as soon as any mode shows something.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): restart the loop if the empty-rotation pause starts on-demand; per-rotation streak

- an on-demand request serviced during the pause returned early into the
  on-demand branch, which advanced past the mode just requested; the loop
  now restarts when the pause changed the mode, on-demand state or schedule
- the streak is reset when the rotation changes (on-demand start/stop, a
  plugin enabled or disabled), so a streak from one rotation can't make
  another pause before its own modes are tried
- docstring: live content is picked up within the pause, not "at once"

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 09:07:08 -04:00
ChuckandClaude Opus 5.5 aeaeaa4e94 chore: make contributor tooling work; fix drifted docs (#648)
- mypy.ini parses again (multi-line exclude and inline value comments made
  mypy reject the file); the mypy pre-commit hook is manual-only until the
  ~500 existing errors in src/ are paid down, and CONTRIBUTING says so
- .gitignore: ignore all of config/ except the templates (ytm_auth.json and
  others weren't ignored)
- .gitattributes: LF for .sh and .service
- claude-code-review: skip fork PRs, which have no secrets
- check_system_compatibility.sh: 3.13 supported, <3.10 an error
- docs/scripts drift: emulator guide, README API Metrics, route count,
  docs index, scripts README; pyflakes nits in dev scripts

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:26:55 -04:00
ChuckandClaude Opus 5.5 76f5d8a336 fix(web): seven web UI bugs, and remove dead plugins_manager.js helpers (#647)
- Operation History: the plugin filter lists the installed plugin ids
  instead of one option, "plugins" (Object.keys of {plugins: [...]}).
- Ctrl/Cmd+S submits the active tab's first visible form with
  requestSubmit() (validation and onsubmit guards run) instead of a bare
  Event on the first form in the document; skipped inside a modal dialog
  and on tabs without a form.
- Overview "Check Updates" confirms like "Update Code", takes its button
  explicitly (no implicit global event) and shows the server's message.
  Both, and the Tools tab git pull, raise the restart-pending banner on
  restart_required.
- Tools: toolsAction and diagnostics show the server's error message;
  only a non-JSON body falls back to HTTP <status>.
- Installed list after uninstall: PluginAPI writes clear the throttler's
  GET cache, a forced loadInstalledPlugins clears it too, and the
  post-uninstall reload goes through refreshInstalledPlugins().
- Plugin widgets load from /static/plugin-widgets/ only (the other two
  paths have no route).
- Raw JSON editor escapes the parse error; slider escapes value/min/max/step.
- Removed the unreferenced array-of-objects and key-value helpers from
  plugins_manager.js, the textarea auto-resize and Ctrl+R handlers in
  app.js, and a redundant ?v= on the plugins_manager.js script tag.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:26:27 -04:00
ChuckandClaude Opus 5.5 7eb7a58d0c fix: web UI and src.common bugs (wifi wrong-password, plugin icon, starlark toggle, API caching, scroll/logo/font helpers) (#646)
- wifi: keep the "wrong_password:" prefix through the restore/AP fallback so
  the UI's incorrect-password prompt fires again.
- /plugins/installed returns the manifest's icon (string only).
- /starlark/apps/<id>/toggle coerces `enabled` and delegates to
  _toggle_starlark_app (disk before memory, no KeyError, "false" is false).
- /api/v3/ JSON GETs are sent Cache-Control: no-store; non-JSON keeps 5s.
- ScrollHelper.set_scrolling_image converts non-RGB input (alpha onto black);
  create/set_scrolling_image reset last_update_time like reset_scroll.
- LogoHelper backs off a failed download per path for
  MISSING_LOGO_RECHECK_SECONDS; cleared on invalidate/clear_cache.
- refresh_placeholder_timestamp saves atomically.
- FontManager.clear_cache / _clear_plugin_font_cache bump cache_generation.
- Odds manager: per-game logs to DEBUG; JSON decode error caught before
  RequestException (same cooldown).
- element_style mangled continuations; startup validator skips null plugin
  blocks and reuses the controller's discovery.
- src/common/README lists frame_timing, json_body, render_gate.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:26:05 -04:00
ChuckandClaude Opus 5.5 b518c51679 fix(plugin-system): load/enable failures, atomic state files, pip lock, test-double parity (#645)
- load_plugin: an on_enable() that raises unregisters the instance, so the
  next load retries instead of returning True "already loaded".
- get_plugin_info: guard plugin.get_info(); one plugin raising no longer
  breaks /api/v3/plugins/installed.
- plugin_state.json and the operation history are written with
  atomic_write_text under their lock.
- plugin_loader: module-level lock serialises pip installs across the
  parallel startup loaders.
- store_manager._install_via_download: extract dir cleanup moved to finally.
- Test doubles: draw_image() warns (DeprecationWarning; the real
  DisplayManager has none), MockDisplayManager.draw_text accepts the real
  signature's optional params, VisualTestDisplayManager logs draw errors at
  WARNING.
- Docs/comments: compatibility.py method name, PluginState.LOADED meaning,
  brittle schema count, why _report_skip_once uses setdefault.
- Remove unused PluginOperationQueue.get_active_operations().

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:25:45 -04:00
ChuckandClaude Opus 5.5 6bc13a8934 fix(display): Vegas resumes after live priority, and six smaller runtime fixes (#644)
- Vegas: a live-priority pause was only lifted from inside run_frame(),
  which returns before that check while paused, so the ticker never came
  back until a restart. run_iteration() now resumes it (the controller
  only calls it when nothing preempts Vegas); start()/stop() clear the
  pause state. Iteration length is timed with the monotonic clock.
- Dim schedule: a per-day disabled day now updates the minute-gate cache,
  so brightness no longer flips back to dim within each minute.
- On-demand: a second request no longer overwrites the rotation resume
  index with the first request's mode.
- Render pipeline: reset() drops the prepared group and deferred queue,
  and a prefetch in flight across a reset discards its result.
- Sync: stop() removes the status file (and the controller's cleanup now
  calls it), standalone removes a stale one at startup, and writes use a
  unique mkstemp temp file.
- render_gate.swap_releases_gil() delegates to frame_timing.
- Stale docstrings/comments corrected.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:25:26 -04:00
ChuckandClaude Opus 5.5 bcef1957a9 fix(security): refuse unsafe plugin ids, keep secrets private, validate request bodies (#643)
* fix(security): refuse unsafe plugin ids, keep secrets private, validate bodies

- install_from_url and the registry install's manifest rename refuse a
  plugin id that is not a single safe name (no ../ out of plugins_dir).
- Uninstall and config reset refuse core config sections and ids with
  path parts; uninstall of a plugin whose directory is gone still works.
- separate_secrets checks a field's own x-secret marker before recursing,
  so object/array secrets no longer land in config.json.
- Backup restore creates missing secrets/wifi/ytm files with mode 640;
  export skips non-object manifests and no longer collides on same-second
  exports.
- SYSTEM_FONTS includes every bundled font from BUNDLED_FONTS.
- Raw config/secrets saves and validate_request_json require a JSON object.
- A blank max_dynamic_duration_seconds keeps the stored value; other values
  are validated to 30-1800 instead of raising a 500.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(security): validate the id before install_plugin moves anything; claim backup names atomically

- install_plugin set aside plugins_dir / plugin_id before any id check, so
  "../x" moved a directory outside the plugins dir (the rollback moved it
  back, but only if the install path got that far)
- two exports finishing in the same second could both see a free name and
  the later os.replace destroyed the first archive; the name is now
  claimed with O_EXCL before the archive is swapped in

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 08:24:43 -04:00
ChuckandClaude Opus 5.5 da5937da3d fix: six bugs found testing main on a real Pi (#641)
* fix: six bugs found testing main on a real Pi (ledpi)

- Stopping the service now runs cleanup. systemd stops ledmatrix.service
  with SIGTERM, whose default action ended Python before run()'s finally
  block, so the update worker, Vegas and the panel were never torn down.
  main() now turns SIGTERM into KeyboardInterrupt, the Ctrl-C path.
- "Now showing" no longer turns into "unknown". display_current_state was
  only written on a mode change and the web UI reads it with max_age=120,
  so a live game or a single plugin on screen for longer read as unknown.
  It is republished every 30 s while unchanged.
- Switching Vegas on in the web UI works when it was off at startup. The
  coordinator was only created at startup; the config watcher now flags it
  and the render thread creates it.
- configure_web_sudo.sh finds reboot and poweroff in /usr/sbin. Run as the
  web user it could not, silently dropped their rules and still said it
  granted them, so the web UI's Reboot/Shutdown stopped working.
- check_system_compatibility.sh reports installed packages as installed.
  `dpkg -l | grep -q` under pipefail failed when grep exited early.
- A network failure fetching GitHub repo info logs a WARNING, not ERROR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): fixes found testing on a Pi

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: run the Linux-only script tests correctly

The sbin-lookup test set PATH=/nonexistent and then could not find bash
itself; call it by absolute path. The dpkg-query stub read $4, but the
package name is the third (last) argument. Both now pass on a Pi.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): cover the follower, long-render and startup cases

Review follow-ups on the ledpi fixes:
- The pending Vegas start is applied before the sync-follower branch too
  (_apply_pending_vegas_init), which skips _is_vegas_mode_active() while a
  follower is connected but needs the coordinator for the leader's image.
- _service_pending_changes(), which runs inside Vegas iterations and long
  screens, republishes a stale display_current_state as well; the main
  loop alone could be away for a 240 s Vegas iteration.
- The SIGTERM handler is installed after DisplayController() is built, so a
  stop during parallel plugin loading keeps the default immediate exit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 18:13:49 -04:00
ChuckandClaude Opus 5 c4927e82a3 fix(config): normalize nullable arrays and objects instead of refusing them (#642)
* fix(config): normalize nullable arrays and objects instead of refusing them

`element_style._nullable` widens every `customization.modes.<mode>` override
with 'null' so a blank means "inherit the base", which turns a colour declared
"array" into ["array", "null"]. `normalize_config_values` only knew how to
convert null/integer/number/boolean out of a union, so a valid [0, 249, 0]
matched nothing and logged

    Could not normalize field customization.modes.upcoming.odds_text.text_color:
    value=[0, 249, 0], type=<class 'list'>, schema_type=['array', 'null']

The warning was the harmless half. It `continue`d past the single-type handling
below, where `prop_type == 'array'` coerces items, so a nullable array never had
its items normalized while a plain one did. A form posts numbers as strings, so
["0", "249", "0"] survived to the validator and was rejected with "Expected type
integer, got str" -- setting a per-mode colour in the web UI failed outright.
Every per-mode override of a structural or string type was exposed, across all
eight scoreboard plugins, not only colours.

Re-enter the single-type handling with the matched member rather than bailing,
accept a string that matches, and warn only on a genuine mismatch so the
diagnostic still reaches the validator.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ET8e5weDrb5Ju5QTKLU7zh

* fix(config): convert only integral numbers for integer array items

Review catch on the previous commit. Routing nullable arrays into the shared
item handling made its integer coercion reachable for them, and that coercion
called int(v) on any number: a client sending [2.5, 249, 0] for an RGB array
got 2 stored and a 200 back, so a wrong value was silently corrected into a
valid-looking one rather than refused.

Convert only genuinely integral values, at both the union-item and the plain
'array' item branch so the two cannot drift. A whole float -- 2.0, which is all
JSON can express for an integer -- still converts. This also settles an
inconsistency that predates the change: int('2.5') raises, so the string form
was always preserved and rejected while the numeric form was truncated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ET8e5weDrb5Ju5QTKLU7zh

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-26 16:33:23 -04:00
ChuckandClaude Opus 5.5 9964dd2183 feat(vegas): render plugin content off the render thread, and keep it off the GIL when the panel needs it (#630)
DisplayManager.offscreen() gives a thread its own canvas, so Vegas renders every plugin's ticker content on its prefetch thread instead of pausing the scroll for canvas-bound plugins on the render thread. A render gate (src/common/render_gate.py, vegas_scroll.prefetch_gate, on by default with the GIL-releasing binding) lets the prefetch thread run Python only while the render thread waits in SwapOnVSync: on hdpi, frames 2+ refreshes late fell eightfold and late frames overall from 0.90% to 0.60%. See docs/OFFSCREEN_RENDERING.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 19:57:03 -04:00
ChuckandClaude Opus 5.5 865d62f67b feat(display): compensate for the panel's scan order while scrolling (#634)
A 1:N-scan HUB75 panel lights the two rows either side of its middle at opposite ends of each refresh, so a scroll at one pixel per refresh shows a 1px step across the middle of every panel. While something scrolls at one frame per refresh, DisplayManager now shows the half whose seam row lights first one refresh behind the other (src/scan_order.py), which lines the two up again. Only for layouts whose row order is known; display.scan_order_compensation "off" disables it. Confirmed on hdpi before and after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 19:47:44 -04:00
ChuckandClaude Opus 5.5 8ad9d191a7 feat(perf): frame timing for every presented frame, a soak tool, a render bench and a stall watchdog (#629)
src/common/frame_timing.py times every frame the display presents, whoever drew it, and writes cumulative counters to /dev/shm. scripts/frame_soak.py grades a running service (late frames, freezes, where the time goes) and scripts/render_bench.py the hardware and render path alone. A stall watchdog logs the stacks behind any scroll held up for 250 ms or more (LEDMATRIX_STALL_WATCHDOG_MS lowers that). See docs/SCROLL_PERFORMANCE.md, "Soaking a rig".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 19:38:49 -04:00
ChuckandClaude Opus 5.5 7f9c73e9aa fix(vegas): smooth Vegas scroll pacing -- whole pixels per refresh, measured refresh, off-thread preview writes (#628)
Vegas scrolls a whole number of pixels per panel refresh, locked to SwapOnVSync, against the refresh the panel really holds (measured from swap gaps), instead of blending sub-pixel positions against the refresh cap. The web preview PNG is encoded off the render thread while scrolling, with writes ordered and retried. On hdpi, late frames fell from 6.3% to 0.7%. See docs/SCROLL_PERFORMANCE.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 19:38:22 -04:00
ChuckandClaude Opus 5.5 b9416ef803 fix(display): Vegas teardown and 240s default; remove dead Vegas buffer code (#637)
* fix(display): tear down Vegas mode on controller cleanup

DisplayController.cleanup() never called VegasModeCoordinator.cleanup(),
so the Vegas teardown (stop, pipeline/stream reset, adapter cache drop)
was unreachable. Call it before the display manager is cleaned up, and
skip it when Vegas was never created.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): default max_cycle_duration to the documented 240s

The template, the web UI help, CONFIG_REFERENCE and the controller all
say 240, but the code defaulted to 600 in two places, so a config
without the key ran Vegas iterations 2.5x longer than documented.

from_config now falls back to the dataclass field defaults instead of
repeating each one, so the two copies can no longer drift, and the
controller's follower scroll-speed default reads VegasModeConfig's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): let run.py -d show display_manager's DEBUG output

display_manager pinned its logger to INFO at import, overriding the root
level, so debug mode never showed its DEBUG lines. Use get_logger() from
src.logging_config like the rest of the core and leave the level to the
logging setup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): run each startup validation check once

StartupValidator.validate_all() ran twice at boot, before and after the
plugin manager was created, so every config, cache, display and
systemd-unit warning was logged twice. The second pass now runs only the
plugin checks. Drop the commented-out raise_on_errors line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): one INFO line per plugin-list refresh

StreamManager logged "=" * 60 banners and a line per plugin (INCLUDED,
SKIPPED, FETCHING CONTENT, SEGMENT CREATED) at INFO on every refresh and
fetch, i.e. at each cycle start and every 30s. Log one INFO summary of
the rotation per refresh and move the per-plugin detail, the weighting
breakdown and "no content this cycle" to DEBUG (the adapter still warns
when every content path fails).

Also drop the check/cross marks from the controller's log messages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): drop the per-iteration static-mode plugin scan

run_iteration() rebuilt _static_mode_plugins on every iteration, asking
every plugin for its display mode and logging the set at INFO, but
nothing ever read it: static pauses are triggered by
_check_static_plugin_trigger() from the next segment. Delete it, the
coordinator's get_ordered_plugins() that only it used, and the
write-only _static_pause_plugin / _static_pause_start.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(vegas): remove the staging buffer that was never filled

StreamManager and RenderPipeline carried a double-buffer design that
nothing used: _staging_buffer was only ever cleared or swapped, so
swap_buffers() never did anything and should_recompose()'s
staging_count > 0 branch was dead, and _active_scroll_image,
_staging_scroll_image, _is_rendering, _last_frame_time and
_frame_interval were written but never read. Delete the machinery and
rewrite the docstrings around what actually carries updates:
_pending_updates, consumed by process_updates() in swap mode and
invalidate_pending_updates() in continuous mode.

should_recompose() no longer builds a buffer-status dict every frame.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(display): tidy the display controller without changing behaviour

- Import VegasModeCoordinator locally instead of through module globals
  (there is no circular import to avoid).
- Drop hasattr() checks on attributes PluginManager.__init__ always sets
  (plugin_executor, plugin_last_update, get_plugin_lock,
  run_scheduled_updates*, stop_update_worker) and the dead "older
  manager" fallbacks; keep the health_tracker None checks, now via
  _health_tracker().
- Extract _display_once() for the per-frame display call both render
  loops copied, _advance_on_demand() for the two on-demand rotations,
  _reset_on_demand_fields() for the error and clear paths, and
  _timezone() / _in_window() for the two schedule checks.
- Remove always-true conditions and the unreachable non-plugin else
  branch in run(), and read _was_display_active / _last_published_mode /
  vegas_coordinator directly now that __init__ declares them.
- Declare the follower render state in __init__, name its tuning
  constants, add _follower_sign(), and share the 90/s sync send
  interval with the render pipeline (SYNC_SEND_INTERVAL).
- Delete history narration and the "Opt #N" labels, fix the comment
  that called _scroll_speed constant (hot reload updates it), and drop
  a startup timing log that measured nothing.
- render_pipeline / plugin_adapter: read display_manager.width/height
  as the properties they are, drop an empty TYPE_CHECKING block, an
  aliased threading import and a duplicated `if result and
  self.sync_manager:`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(display): trim dead code from display_manager

- Add _new_canvas() for the image/draw/fontmode="1" setup that was
  copied six times.
- Call resolve_double_sided() and compose_pixel_mapper_config() directly
  instead of through a module alias and a passthrough method, and replace
  the comment that said the passthrough read class attributes.
- Delete the unused _initialized flag and _ORIENTATION_ROTATE_DEGREES
  alias (no core or monorepo reader; tests stop resetting the flag), the
  test pattern's unreachable no-matrix branch (it only runs once the
  matrix exists), `del old_image  # help GC` (a no-op on a local), a
  duplicated early return in process_deferred_updates, and stale
  comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(vegas): remove unread fields and test-only helpers, fix docstrings

- ContentSegment: drop total_width, fetched_at, is_stale, image_count and
  is_static, none of which is read.
- StreamManager: drop _current_index (never advanced) and the test-only
  get_all_content_for_composition() and has_pending_updates();
  VegasModeConfig: drop the test-only is_plugin_included().
- geometry.find_blank_cut() has had no production caller since the crop
  moved to item boundaries; delete it and its tests.
- PluginAdapter: the _finalize docstring described separator_width
  between every image, and _crop_to_budget's said cuts snap to the
  nearest blank column; both now describe what the code does.
- Coordinator: the static-pause interrupt log no longer blames follower
  mode for every interrupt, and set_update_callback names the callback
  the controller actually wires.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(scroll): correct ScrollHelper comments and drop dead branches

- Four comments said the strip always starts with display_width of
  blank; it does only when lead_gap is None (Vegas passes its own).
- Delete the "Width calculation mismatch" warning: the image is created
  at the calculated width, so the two can never differ.
- Remove the two scroll_delay <= 0 fallbacks (which disagreed with each
  other): set_scroll_delay clamps it to at least 0.001 and nothing in
  core or the plugin monorepo assigns it directly.
- Trim the scipy history from the blend docstring.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(run): drop a redundant comment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): log set_scrolling_state only when it changes

Vegas and scrolling plugins set the scrolling state every frame, so once
display_manager's DEBUG output became visible in debug mode it printed
"Scrolling state set to: True" about 120 times a second. Log only when
the value differs from the previous one; the state, activity timestamp
and frame hold still update on every call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): display-vegas

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:37:15 -04:00
ChuckandClaude Opus 5.5 f6afbdbb15 fix(web): Cache/Logs error mix-up, store errors, tab fallbacks; remove ~2.5k lines of dead JS (#639)
* fix(web): keep Cache and Logs helpers out of each other's way

Both partials declared top-level showError and escapeHtml. Their scripts
run at global scope after every HTMX swap, so whichever tab was opened
last owned window.showError, and a Cache failure after visiting Logs
rendered into the Logs panel (and the other way round). Each script is
now an IIFE; Cache still exports deleteCacheFile for its row buttons.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): make the HTMX-failure fallbacks for tab panels actually run

- The "HTMX never loaded" fallback read appElement.__x.$data, which is
  Alpine 2. The page ships Alpine 3, so the check was always false and
  the Overview never loaded without HTMX. It now reads Alpine.$data().
- The Overview and WiFi panels used hx-on::htmx:response-error, which
  htmx expands to "htmx:htmx:response-error", an event that never fires.
- loadTabContent sent requests with <body> as the source, so htmx fired
  its events on <body> and no panel's hx-on handler ran at all. The
  panel is now the source. htmx also resolves its promise on a 4xx/5xx,
  and the panel was stamped data-loaded anyway, leaving a skeleton that
  never retried; it is now stamped only when no responseError fired.

loadPluginsDirect, loadOverviewDirect and loadWifiDirect are merged into
one window.loadPartialDirect(id, url), which also runs the partial's
inline scripts before Alpine sees the markup (as htmx-config.js does on
htmx:afterSwap). The ~10 s "htmx never arrived" path in loadTabContent
uses it for every tab instead of four hard-coded ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): store and registry failures no longer wipe the Plugin Manager

showError replaced the whole #plugins-content with an error message, so
one failed store search, custom-registry install or saved-repository
call took the installed list, the store and every control with it, with
no way back short of reloading the tab. Those failures are now error
notifications. The full-panel message is kept only for a first load of
the installed list that failed (nothing to show yet); a failed refresh
of an already-rendered list is a notification too. showSuccess's
fallback branch, which wrote the message into innerHTML unescaped, is
gone: showNotification always exists.

The "Please try refreshing your browser" hint tested for the text
"Failed to Fetch", which no browser produces (Chrome says "Failed to
fetch", Firefox "NetworkError..."), so it never appeared. It now keys on
the failure itself: a TypeError from fetch(), or PluginAPI's
NETWORK_ERROR wrapper around one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): escape plugin action and install output on every path

executePluginAction escaped data.message and data.output when an action
failed but put data.message straight into innerHTML when it succeeded,
and set the OAuth step-2 button's innerHTML from the manifest's
step2_button_text. Plugin actions run plugin code, so that is plugin- or
server-controlled markup in the page. Both paths now escape, and the
button label is set with textContent.

The same pattern sat in the install-from-GitHub-URL status lines
(plugin_id, the server's message, and error.message, which can echo a
repository URL) and the custom-registry load error; those are escaped
too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): file-upload widget owns the image list and schedule editor

plugins_manager.js loads after the widget bundle, so its older copies of
deleteUploadedFile, updateImageList, hideUploadProgress, formatDate,
openImageSchedule, toggleImageScheduleEnabled, updateImageSchedule{Mode,
Time,Day} and updateCheckboxGroupData replaced the widget's. They are
deleted; the widget files are the only definitions.

Before switching over, the two sets were diffed and fixed so nothing
regresses:

- The old copy labelled the schedule/delete buttons for screen readers
  and lazy-loaded thumbnails; the widget now does both.
- The schedule button did nothing on a card rendered by
  plugin_config.html whenever the image id is a UUID (every upload): the
  template turns "-" into "_" in the editor's id, and neither JS copy
  did. Both now use the template's rule.
- The widget's "keep the open editor open" copied the editor's innerHTML
  into the new list. That dropped its event listeners and showed the old
  values, so after the first change the editor looked live but ignored
  input. A schedule edit now saves to the hidden input and updates the
  card's summary in place without re-rendering the list; a list re-render
  (upload, delete) rebuilds an open editor from the data. Editor controls
  are routed by one delegated change listener, so there are no
  per-element listeners to lose.
- The old deleteUploadedFile had a JSON branch that removed a
  #file_<id> element and skipped the re-render. No template or script
  renders such an element, and JSON uploads are listed through
  updateImageList like images, so re-rendering (the widget's behaviour) is
  the consistent one; the branch was not carried over.
- The template always renders the summary line (".image-schedule-summary",
  "Always shown" when unscheduled) so an edit has a line to update.

The inline-handler test evaluated plugins_manager.js's updateImageList;
test_file_upload_widget.js now covers the widget's list and editor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete the unused handleCredentialsUpload

Its last caller went when plugin_config.html switched credential uploads
to the file-upload widget's handleSingleFileSelect. Nothing in the web
UI, the tests or the plugin monorepo references it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete dead and shadowed front-end code

Nothing calls any of these (checked across web_interface/, test/ and the
ledmatrix-plugins monorepo, including hx-*/x-*/onclick attributes):

- app-shell.js: the Alpine methods refreshPlugins (it called a
  nonexistent this.searchPluginStore), loadPluginConfig,
  savePluginConfig, getSchemaPropertyType, escapeCssSelector,
  formatCommitInfo and formatDateInfo, and the top-level copies of
  savePluginConfig, getSchemaPropertyType, escapeCssSelector,
  formatCommitInfo, formatDateInfo and togglePluginFromTab. Plugin config
  forms save through hx-post in plugin_config.html.
- window.reconnectSSE (app-shell.js); window.updateArrayTableAddButtonState
  (array-table.js).
- toggleNestedSection, defined twice (app-shell.js and
  plugins_manager.js) and called from nowhere.
- plugins_manager.js: the window.initializePlugins wrapper around an
  IIFE-local origInit that was always undefined, and __pluginDomReady,
  which was written but never read.
- display.html's fixInvalidNumberInputs fallback: app-shell.js defines it
  before any partial loads.
- base.html's window.loadCodeMirror and the two CodeMirror stylesheet
  preloads, and the .CodeMirror rules in plugins.html. The raw JSON
  editor is a plain textarea.

Also deleted: app-shell.js definitions that a later script always
replaced, so they never ran: executePluginAction (plugins_manager.js
assigns its own), uninstallPlugin and its pollUninstallOperation
(plugins_manager.js), and updateAllPlugins (install_manager.js).

vendor/codemirror stays: test/test_web_smoke.py still requests
codemirror.min.js as a sample static asset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): call showNotification without checking it exists

app-shell.js defines window.showNotification (a stand-in that queues
until the notification widget loads) and base.html runs it, deferred,
before every other script that notifies: app.js, the utilities, the
widget bundle, plugins_manager.js, and all partials, which HTMX loads
after the page. The 81 `typeof showNotification === 'function'` /
`!== 'undefined'` checks, the `window.showNotification || console.log`
and `|| alert` fallbacks, and their else branches (alert(), console
output, and schedule.html's own hand-built toast) could never take the
fallback path. They are removed, as is fonts.html's second copy of the
queueing stand-in.

The stand-in in app-shell.js keeps its guard (it must not replace the
widget's implementation if load order ever changes), and BaseWidget's
public notify()/getNotificationFunction() keep their shape for widgets
that plugins ship.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one HTML escaper, window.LEDEscape

About 30 files each carried their own escapeHtml / escapeAttr / escHtml /
_esc / escapeJs. They disagreed: several (notification.js, display.html's
escapeHtml, operation_history.html, the app() stub) did not escape
quotes, google-calendar-picker.js and tools.html's escHtml left ' alone,
and some turned 0 into ''. Most were fine only because the quote-safe
widget copies were preferred at runtime.

window.LEDEscape now lives at the top of app-early.js, a blocking script
in <head>, so it exists before any other script runs:

  html(v)          & < > " ' as entities, null/undefined as ''
  attr(v)          the same, for call sites that want to say "attribute"
  jsStringAttr(v)  a JS string literal safe inside an inline handler

Every former copy is now a one-line name for it (kept so call sites do
not change), widgets included, with no fallback. plugins_manager.js
loses its four escapeJs wrappers (callers use jsStringAttr), the
duplicate escapeAttr and escapeHtml inside renderInstalledCards and
renderCustomRegistryPlugins, and the window.escapeHtml /
window.escapeAttribute exports, which nothing read.
addArrayObjectItem's fallback markup (with a sixth hand-written escape
chain) is gone too: window.renderArrayObjectItem is defined earlier in
the same file, so the fallback could not run. The unused escapeHtml
methods on the Alpine app (app-early.js stub and app-shell.js) are
deleted.

test_html_escaping.js now runs LEDEscape and every remaining name for it,
and fails if a hand-rolled escaper reappears anywhere in web_interface/.
Suites that evaluate slices of plugins_manager.js or widget files load
LEDEscape from app-early.js through test/js/led_escape.js.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): stop htmx re-running partial scripts after every tab load

htmx-config.js runs each swapped-in <script> itself on htmx:afterSwap
and meant to turn htmx's own script handling off with
htmx.config.allowScriptTags = false. It did that once, while setting up,
but base.html loads htmx with a dynamic <script>, so htmx was not defined
yet and the setting never applied. On every tab load htmx then tried to
run each script again in its settle phase, found it already replaced
(no parent node) and threw "Cannot read properties of null (reading
'insertBefore')" into the console, which also skipped the rest of that
swap's settle tasks.

The setting is now applied in the afterSwap handler, which always runs
after htmx exists and before htmx settles the same swap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): show "--" for a system stat the server could not read

The stats stream and /system/status now send null for a metric they
cannot read (cpu_temp off a Pi, for one) instead of 0. updateSystemStats
built the header and Overview text as value + unit, so a null showed as
"null°C". CPU, memory and temperature, in the header and on the
Overview, now render "--" plus the unit for null or a missing field --
the same placeholder the page starts with, and what tools.html already
shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one Alpine accessor and one plugin-list signal

window.getApp() (app-early.js) returns the root <body x-data="app()">
component through Alpine's public Alpine.$data, or null before Alpine
has initialised it. It replaces the private el._x_dataStack[0] reads in
app.js, app-early.js, app-shell.js, settings-search.js, overview.html and
plugins_manager.js, the three local getAppComponent/appData/getAppData
copies, and the Alpine 2 el.__x.$data fallbacks, which Alpine 3 never
provides.

Publishing the installed-plugin list: one load set window.installedPlugins
and dispatched pluginsUpdated twice (loadInstalledPlugins, then
renderInstalledPlugins), then wrote into the Alpine component through
_x_dataStack[0] and called its updatePluginTabs() directly, and
app-early.js's global listener set window.installedPlugins a third time
and called updatePluginTabs() again. Now renderInstalledPlugins is the
one publisher: it sets window.installedPlugins and dispatches
pluginsUpdated once, and the full app()'s listener (app-shell.js) is the
receiver. The app-early.js listener only builds the tab row while the app
is not the full implementation yet. The "grid not loaded yet" case is a
normal state (Plugin Manager tab not opened), so it logs through
pluginLog instead of console.warn.

updatePluginTabs had a "Debounce" comment and clearTimeout over a timer
nothing ever set, and two identical branches; it now just calls
_doUpdatePluginTabs (app-early.js detects the full implementation by
that name in its source, which the new comment says).

app()'s unused baseComponent lookup is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): reload the plugin list after installs and failed toggles

Several callers refreshed the installed list with

    if (typeof loadInstalledPlugins === 'function') loadInstalledPlugins();
    else if (typeof window.loadInstalledPlugins === 'function') ...

but loadInstalledPlugins is local to the plugin-manager IIFE and
window.loadInstalledPlugins is never defined, so from outside that IIFE
both tests were false and nothing reloaded:

- A failed plugin toggle left the switch drawn in the new state while
  the data said the old one. It now re-renders from the reverted data.
  The optimistic in-place edit also has to forget the grid's
  last-rendered markup, or setGridHtmlIfChanged sees identical HTML and
  skips the revert. A successful toggle still keeps the switch (and
  focus) as drawn.
- Installing from a GitHub URL (the early handleGitHubPluginInstall),
  installing or uploading a Starlark app, and toggling a Starlark app on
  its config tab never refreshed the list, so the new app had no tab or
  Installed badge until the page was reloaded. They now force a reload
  through window.pluginManager.loadInstalledPlugins(true), and the
  Starlark grid redraws when that finishes instead of after a fixed
  500 ms.
- The Starlark uninstall inside the IIFE reloaded from the 3 s cache,
  which could still hold the app; it now forces a reload.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): route debug output through debugLog

base.html defines window.debugLog, gated on localStorage.pluginDebug.
plugins_manager.js read the same key twice more into its own flags
(_PLUGIN_DEBUG_EARLY, and PLUGIN_DEBUG behind a pluginLog() wrapper), and
api_client.js's RequestThrottler had a separate `debug` property with a
setDebug() that nothing called. All of it now goes through debugLog. The
"functions defined" dumps with their ✓ lines, and two per-plugin
"enabled=" loops that ran on every render, are dropped; "[PLUGINS STUB]"
labels on code that has not been a stub for a long time read
"[PLUGINS]".

Ungated console.log calls that announced normal events on every page
load or action (settings search and tooltips registering, every toast
repeated to the console, the schedule pickers initialising, widget
registry unregister/clear) go through debugLog too. What remains on
console.log is the widget registry's on-demand LEDMatrixWidgets.debug()
dump and BaseWidget.notify's no-notifier fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop waits and guards that could never fire

- handlePluginAction polled up to 10 x 50 ms for window.togglePlugin,
  configurePlugin, updatePlugin and uninstallPlugin before calling them.
  All four are defined when the scripts load, before any card can be
  clicked, so the poll always succeeded at once; it now calls them.
  The long thinking-aloud comment over the toggle state is replaced by
  two lines on why the stored state, not the checkbox, decides.
- initializePlugins checked typeof on setupGitHubInstallHandlers and
  applyStoreFiltersAndSort, function declarations in the same IIFE, and
  wrapped window.checkGitHubAuthStatus(), which returns a promise with
  its own .catch, in try/catch.
- searchPluginStore wrapped each "#store-count" update (a getElementById
  and an innerHTML assignment) in try/catch four times; one
  setStoreCount() helper does it. The store's post-render re-attach of
  the GitHub token handler dropped its try/catch and existence checks
  for the same reason.
- The load-time fallback outside the IIFE tested typeof
  initializePluginPageWhenReady, which is IIFE-local and so always
  undefined there; it calls window.initPluginsPage directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete two unused plugin-manager helpers

stopOnDemand (IIFE-local; the page's stop button calls window.stopOnDemand
from app-shell.js) and debounce had no callers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): document the plugin-config handlers templates call

validatePluginConfigForm, handleConfigSave, handleToggleResponse,
handlePluginUpdate and refreshPluginConfig each get a JSDoc naming the
attribute in partials/plugin_config.html that calls it and what the
return value means (only validatePluginConfigForm's matters: false
cancels the submit).

- The `if (!window.__pluginConfigHandlersInitialized)` wrapper is gone:
  app-shell.js runs once per page, so it was never false. The block is
  dedented one level; `git diff -w` shows the real change.
- The three handlers read xhr.responseJSON first. XMLHttpRequest has no
  such property (it is jQuery's), so that branch never ran; one
  xhrJson(xhr) helper parses responseText for all of them, with the same
  fallbacks as before.
- runPluginOnDemand and stopOnDemand checked that plugins_manager.js's
  openOnDemandModal/requestOnDemandStop exist; plugins_manager.js is on
  every page, so they call them.
- fixInvalidNumberInputs had a stray "Notification helper function"
  comment on top of its own; a leftover "section toggle ... duplicate
  definition removed" note is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): one toast per save, and a failed durations save says so

app.js's global htmx:afterRequest listener showed the server's message
for every htmx request, and every form and button that posts through
htmx (plugin config save/toggle/update, Display, Durations, General,
Schedule, Dim schedule, the Overview actions) also reports its own result
from hx-on after-request. Each save showed two toasts. The global
listener now stays quiet for a request whose element, or its form, has
its own after-request handler.

That exposed the Rotation & Durations form's handler, which read
xhr.responseJSON: XMLHttpRequest has no such property, so it always said
"Durations saved" in green, even when the save failed (the global toast
had been the only place the error showed). display.html already had a
correct version (2xx only counts as saved; the server's message wins;
its status may refine success but never overturn failure). That is now
window.showSaveResult(xhr, savedText, failedText) in app.js, used by the
Display, Durations and General forms; General's inline copy of the same
logic is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(web): file headers and comments that say what the code does now

- plugins_manager.js, app-shell.js, app.js and app-early.js open with a
  header: what the file owns, how base.html loads it and in what order
  relative to the others, and the globals it defines. app-early.js's
  app() stub also says why it exists and that, with app-shell.js now
  loaded before Alpine, it does not run in practice.
- base.html's note on plugins_manager.js said it must load last to win
  over same-named functions in app.js/app-shell.js; there are none left,
  so it now gives the real reason (it uses everything loaded before it).
- Change-narration and "already defined at the top, no need to redefine"
  notes are gone or rewritten as present-tense reasons; comments that
  were wrong are fixed ("Toggle password visibility" over the function
  that opens the token panel, "Insert before the closing </nav>" over an
  appendChild, "(from v2)", the export note that still listed
  escapeHtml). About forty comments that restated the line below them
  are removed, and a second window.currentPluginConfig = null outside the
  IIFE is dropped (the IIFE sets it).
- The file-upload, checkbox-group and custom-feeds widgets' render()
  stubs say plainly that the widget is rendered server-side, instead of
  "for now" / "placeholder for future client-side rendering".

test_plugin_action_delegation.js sliced the source up to one of the
removed notes; it now ends the slice at the next section header.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): keep the escapeHtml/escapeAttribute globals for plugin pages

6da77363 removed window.escapeHtml and window.escapeAttribute because nothing in core or the plugin monorepo read them. Plugin web UIs served through serve_plugin_web_ui and third-party plugin pages may still call them, so they come back as aliases of window.LEDEscape.html and .attr, defined in app-early.js before any other script runs. test_html_escaping.js checks the aliases exist.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): web-frontend

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): encode the image thumbnail path; match script tags case-insensitively

CodeQL flagged the upload widget building an <img> src from a stored path,
and the escaper test extracting inline scripts with a case-sensitive regex.
Each path segment is now URL-encoded (still a same-origin path, and correct
for names with spaces or

* fix(web): clear Codacy findings in the escaper, app shell and upload widget

- LEDEscape looks entities up in a Map instead of indexing an object.
- showNotification is declared as a global for app-shell.js.
- openImageSchedule checks the index is a non-negative integer and reads
  the image with Array.prototype.at.
- The schedule editor calls escapeHtml directly and documents why its
  innerHTML template is safe: every value is escaped or constrained.
  The remaining rule hits are suppressed on that line with the reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): build the image schedule editor with DOM calls

Codacy does not honour inline suppressions, and the editor's innerHTML
template kept tripping its XSS rules even though every value was escaped.
The editor is now built with a small element helper (createElement and
setAttribute), so no value is ever parsed as HTML, and the file's own
escapeHtml goes away.

Also for Codacy:
- LEDEscape.attr is its own function rather than a second name for html.
- The tab loader records a failed load on the panel (data-load-failed)
  from a named handler, instead of a closure over a local flag.

The fake DOM in test_file_upload_widget.js gains append/replaceChildren,
its hostile-id check now asserts the id arrives as attribute data with no
innerHTML anywhere in the editor, and test_html_escaping.js drops the
file-upload.js escaper it no longer has.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): schedule editor helpers as plain functions

Codacy's lint flags arrow functions held in local constants and a forEach
callback that returns a value. The editor's pieces are now named function
declarations (displayStyle, scheduleModeOption, scheduleRangeTime,
scheduleDayTime, scheduleDayRow) taking what they need as arguments, and
the element helper loops with for...of. htmx is declared as a global in
app-shell.js. Output is unchanged; test_file_upload_widget.js passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:36:53 -04:00
ChuckandClaude Opus 5.5 3a81f38f09 fix(web): uniqueItems saves, /health count, Vegas order wipe; one list-repair helper (#638)
* fix(web): drop repeats from uniqueItems lists before validating a plugin save

dedup_unique_arrays lost its only caller in #330, so submitting a value a
uniqueItems list already holds (a stock symbol saved once and posted again)
failed the whole save with a validation error. _prepare_plugin_config_for_save
runs it again just before validation, which covers both POST /plugins/config
and plugin sections posted to /config/main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): /health counts the discovered plugins and logs the checks it fails

The plugin check counted plugin_manager.get_available_plugins(), which
PluginManager does not have, behind a hasattr guard that made plugin_count 0
on every device. It now counts the discovered manifests, discovering first
when nothing has been scanned yet.

The config, plugin and hardware checks answered "see logs for details"
without logging anything. Each now logs a warning with the traceback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): store refresh no longer claims a commit-metadata refresh

POST /plugins/store/refresh read fetch_commit_info (or fetch_latest_versions)
only to append "(with refreshed commit metadata from GitHub)" to its message.
It never fetched any: the route re-downloads the registry and nothing else.
search_plugins takes the flag, but it reads commit info through its cache,
so passing it on would not refresh anything either. The flag is ignored now
and the message says what happened.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): refuse a malformed Vegas plugin order instead of clearing it

A vegas_plugin_order or vegas_excluded_plugins value that was not JSON, or
not a list, was stored as [] and the save answered 200, so a bad value wiped
the saved order or exclusions. Both now answer 400 and save nothing, the way
plugin_rotation_order already did; the three share one parser. A list that
holds anything but plugin-id strings is refused as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): per-plugin health and metrics read the display service's latest

GET /plugins/health/<id> and /plugins/metrics/<id> called get_health_summary
and get_metrics_summary without force_reload, so they answered with whatever
the web process read first and kept in memory, while the display service kept
writing newer state. They now pass force_reload=True, as the list routes do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): plugin config reset saves through the shared atomic save

POST /plugins/config/reset called config_manager.save_config directly, so it
took no backup, and a failed write escaped as an unhandled exception. It then
handed on_config_change the raw stored section, not the prepared config a
loaded plugin runs with. It now saves through _save_config_atomic with a
backup, answers CONFIG_SAVE_FAILED when that fails, and notifies with
_prepared_plugin_config, as POST /plugins/config does.

POST /plugins/toggle carried its own copy of _save_config_atomic's
save_config_atomic-or-save_config fallback; it calls the shared helper now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): one reading and one "unavailable" for each system metric

system_metrics.collect_system_metrics() promised None for a metric it could
not read, but returned cpu_temp as 0 off a Pi, and the whole no-psutil
fallback as zeros. GET /system/status measured the same numbers a second time
with its own code, and answered None there. Now both come from
collect_system_metrics(), and "unavailable" is None everywhere.

/system/status keeps its 0.1s CPU sample and its 10s cache, and gains
nothing it did not already send. Two differences: without psutil it answers
200 with null metrics instead of 503, and a disk it cannot stat is null
instead of a 500.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): /display/current sends the snapshot as-is and logs a failed read

GET /display/current PIL-decoded the preview snapshot and re-encoded it before
base64-ing it, spending CPU on the Pi to send the same picture, and dropped
any failure with `except Exception: pass`. The /stream/display SSE stream
already passed the PNG's bytes straight through.

Both now read through web_interface/display_preview.py and answer with the
same payload. A missing snapshot is still a null image; any other read
failure is logged as a warning. /health reads the snapshot path from the same
module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one helper puts a submitted plugin config's lists back

The plugin-config save turned position-keyed dicts ({"0": ..., "1": ...})
back into lists in five copies: four in the form path's
fix_array_structures (whose prefix branches never ran, since no caller
passed one), and _fix_json_arrays on the JSON path. It then force-fixed
the news plugin's feeds.custom_feeds by name, in case the generic pass had
missed it. src/web_interface/config_arrays.coerce_array_shapes now does it
for both paths, custom_feeds included. ensure_array_defaults duplicated
_fix_none_arrays and is gone.

In the same function: the union-type re-checks that the null handling
above them made unreachable, the "(temporary)" random_seed debug log, and
a commented-out log line are removed. A failed validation is logged once
as a warning, not four ERROR lines and a WARNING.

Element types are left to normalize_config_values, which already converted
them for both paths. One difference: the form path no longer adds an empty
{} for a nested object the post left out that has no defaults.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): import at module top and log through the module logger

The web_interface.cache imports in config.py and fonts.py were wrapped in
`except ImportError` fallbacks. It is an in-repo module that imports nothing
from the project, so it cannot fail to import; it is imported once at module
top, as system.py now does. cache.py's docstring said blueprints import it
lazily "to avoid circular imports"; it now says why that is unnecessary.

Five logging.error calls in the dim-schedule GET and three logging.warning
calls in plugins.py went to the root logger; they use the module logger.
Function-local re-imports of json, os, shutil, logging and Path, all
already imported by the module, are gone. The `import os` inside two except
blocks of save_plugin_config also made os a local name for the whole function.

execute_plugin_action's step-1 handler gets a comment saying why it stays:
it looks like a copy of the blueprint handler, but without it a
TimeoutExpired from the plugin's script would reach the route's own
`except subprocess.TimeoutExpired` and be answered as a 408.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): app.py loses dead CSRF and reconciliation state, comments fixed

- csrf was always None, so `if csrf: csrf.exempt(...)` never ran, and its
  note that the api_v3 blueprint "is exempted above" named an exemption that
  does not exist. Both are gone; the reason there is no CSRF protection stays,
  shortened.
- The SSE rate-limit comment called the default "tight" at 20 per minute. The
  default is 1000 per minute and the streams' 200 is the tighter one; the
  comment now says so. The limits are unchanged.
- _reconciliation_done was written and never read. The docstring that
  explains why reconciliation runs once keeps its reason, in the present
  tense.
- Removed: a dangling "import cache functions" comment with no import under
  it, a "security check ... within project_root" label on an existence check,
  the "(simplified version)" narration, and the note that no redirect route is
  needed. The preview loop's sleep comment no longer mentions a PIL encode
  that the loop does not do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(web): api_v3 comments name the package __init__, not a _common module

Every route module's docstring said the shared blueprint comes "from
._common", a module the package split never created; they name the
package __init__. The PROJECT_ROOT comment described the path from
_common.py; it now describes this package and keeps the incident it
guards against. The "(corrected) in this commit" note in
resolve_pull_command and the /health comment the split's mechanical
time -> _pkg.time rewrite garbled ("Stamp the start _pkg.time") read
correctly again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop hasattr checks for attributes PluginManager always has

PluginManager.__init__ sets health_tracker and resource_monitor (to None
until they are configured), so the seven
hasattr(api_v3.plugin_manager, ...) guards in the health, metrics and limits
routes were always true. The falsy checks that do the work stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): pages_v3 dispatches partials from a dict with one error handler

load_partial chose a loader through a fourteen-branch if/elif, and thirteen
of the loaders then wrapped themselves in the same try/except, logging
"Error loading partial" without saying which. The route now looks the name up
in _PARTIAL_LOADERS and has the one handler, which logs the partial's name.
The loaders just render. _load_tools_partial keeps its own messages. The
search index's _partial_html already catches a loader that raises.

serve_plugin_web_ui repeated _plugin_dir_for inline (containment plus the
ledmatrix- prefix fallback); it calls it now. Also removed: the unused
markupsafe.escape import, function-local json/Path re-imports, and unused
exception bindings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): remove unused imports, locals and a try that cannot fail

- get_error_aggregator was imported by the api_v3 package and used by no
  one; seven names config.py imported, and Path in misc.py and logging in
  plugins.py, likewise.
- branch_info in install_plugin was built and never logged; test_config in
  /health was bound and never read (the load_config call is the check).
- An f-string with no placeholders in the asset upload route.
- _installed_plugin_ids wrapped list(manifests.keys()) in try/except;
  _discovered_plugin_manifests always returns a dict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): start.py logs its startup lines and drops unreachable branches

The startup banner went to stdout with print(); it goes through a logger
now, which the app import has already configured, so it reaches the journal
with a level and timestamp like every other line. The "no addresses" branch
is gone: get_local_ips() always returns at least "localhost".

The except around app.run re-raised "only if it's not a client
disconnection error" from inside the branch that had just established it
was one, so that raise could not run. It is one check now, on a named
tuple of the errnos, which the werkzeug log filter uses too. The comment
on threaded=True counts three SSE endpoints, which is how many there are.
Trailing whitespace is stripped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): save_main_config names its General fields once

The General tab's field names were listed twice, once to detect a General
form post and again, with four more, to keep the remaining-keys merge from
storing them as top-level keys. GENERAL_FIELDS and _MAPPED_TOP_LEVEL_FIELDS
hold them now, and the four per-section skip checks are one set.

The comment on that merge said plugin configs are handled "here too", and
"(including plugin keys)". Plugin sections are handled and removed from the
body before it runs; the comment says so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): plugin directories come from the plugin manager only

Six lookups fell back to PROJECT_ROOT/plugins/<id> when there was no plugin
manager: GET /plugins/config's of-the-day data, POST /plugins/action, the
plugin static-file route, the calendar credentials upload and the calendar
OAuth routes. The loader never scans plugins/ (PluginManager.discover_plugins
reads only the configured directory, plugin-repos by default), so what they
found there was a plugin that never runs. _plugin_directory() asks the
manager and answers None without one, which each route already reports as
"not found".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): web-backend

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:35:49 -04:00
ChuckandClaude Opus 5.5 7b90759252 fix: /errors stack traces, Wi-Fi disconnect and save, plugin fonts, API cache TTL (#636)
* fix(errors): record the exception's own stack trace

record_error() called traceback.format_exc(), which only sees an
exception while its except block is running. plugin_executor records
exceptions caught on a worker thread after that block has ended, so
every trace on /errors read "NoneType: None". The trace is now built
from the exception's __traceback__. The executor's log call had the
same problem with exc_info=True and now passes the exception.

record_error() also merged LEDMatrixError context into the caller's
dict in place; it now works on a copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(wifi): point at configure_wifi_permissions.sh instead of a sudoers list

The module docstring told users to grant NOPASSWD sudo on iptables and
ip. configure_wifi_permissions.sh refuses those grants on purpose: a
wildcard rule for either runs an arbitrary program as root. Point at
the script and say why it leaves them out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(wifi): disconnect finds the saved profile by SSID

disconnect_from_network() asked `nmcli -f NAME,802-11-wireless.ssid
connection show` for the profile to take down, but nmcli rejects that
column for `connection show`, so the lookup always failed and only the
device was disconnected. The per-profile lookup _connect_nmcli() already
used is now _find_profile_for_ssid(), and both callers share it. It
also splits terse output on the last colon and unescapes "\:", so a
profile name containing a colon is found.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(wifi): write wifi_config.json atomically and report a failed save

_save_config() opened the file for writing in place and swallowed any
error, so a wifi_config.json left owned by root made the web toggle for
auto-enabling AP mode report success while nothing was saved, and a
crash mid-write could truncate the file. It now uses atomic_write_json,
which also keeps the file's owner and shared group when root saves it,
and returns False on failure. POST /wifi/ap/auto-enable answers 500 in
that case.

The file is now written with indent=4, like the other config files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): resolve plugin:// fonts in the plugin's own directory

FontManager looked for a plugin's bundled fonts under Path("plugins") /
plugin_id: relative to the process cwd, and not the default install
directory (plugin-repos/), so a manifest's plugin:// fonts never loaded.

register_plugin_fonts() takes an optional plugin_dir, and PluginManager
passes the directory it loaded the plugin from. Callers that omit it get
a lookup in the configured plugin_system.plugins_directory, then plugins/,
resolved against the install root.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(api-helper): cache responses for the requested cache_ttl

APIHelper.get(cache_ttl=...) and set_cache(ttl=...) dropped the ttl on
the claim that CacheManager does not support one, but CacheManager.set()
takes a ttl, stores it with the entry, and both cache tiers honour it
over a reader's max_age. Without it every response expired after the
300-second default read age, whatever the plugin asked for. The ttl is
now passed through, and the cache read passes cache_ttl as max_age for
entries written without one. The class docstring describes what the
helper actually does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(style): one scale range for the schema, element_scale and LogoHelper

The generated Scale field allowed 0.1 to 10, element_style's reader
capped at 10 with no floor, and LogoHelper accepted 0.05 to 8 and reset
anything else to 1.0. A logo scale of 9, which the form accepts, drew at
the shipped size.

MIN_ELEMENT_SCALE / MAX_ELEMENT_SCALE (0.1, 10.0) in src.element_style
are now the schema bounds and the clamp every reader applies through
coerce_scale(): a positive number outside the range is clamped, and
anything that is not a finite positive number means the default. That
also stops element_scale() passing NaN through, since min(nan, 10.0)
is nan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(logos): placeholder lands at the requested path; empty logos list

download_missing_logo() wrote its fallback placeholder to
<normalize_abbreviation(abbr)>.png in the logo directory rather than to
the logo_path the caller passed, so it could return True while nothing
existed where the plugin looks (e.g. "TA&M.png" vs "TAANDM.png").
create_placeholder_logo() takes an optional filepath, and
download_missing_logo passes the requested one.

download_missing_logo_for_team() only caught KeyError, so a team whose
"logos" list is empty raised IndexError; it now treats KeyError,
IndexError and TypeError as "no logo URL".

The placeholder is drawn with PLACEHOLDER_SIZE / PLACEHOLDER_BG, the
constants is_placeholder_logo() recognises it by, instead of repeated
literals.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): resolve bundled font paths against the install root

TextHelper's default font_dir, the logo placeholder's font and
FontManager's font_overrides.json were all relative to the process cwd,
so a process started anywhere but the install root (the plugin safety
harness, a manual run, a unit without WorkingDirectory) drew with PIL's
default face and read no overrides. They now go through
font_layout.resolve_asset_path; the overrides file sits in the install
root's config/.

The resolver docstrings described an order the code does not follow:
resolve_asset_path never consults the cwd, and sports_shared's
_resolve_font_path tries the cwd first. Both docstrings now say what
the code does, and _resolve_font_path calls resolve_asset_path instead
of probing FontManager for it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(sync): the web UI reads the sync status file the display writes

sync_manager writes its status to tempfile.gettempdir(), but
GET /api/v3/sync/status read a hardcoded /tmp/led_matrix_sync_status.json
and defaulted the port to a literal 5765. Wherever TMPDIR is set (or on
any non-/tmp host) the page only ever showed "starting". The endpoint now
uses sync_manager.STATUS_FILE and SYNC_PORT.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(http): the rankings resolver sends the project's User-Agent

DynamicTeamResolver fetched ESPN rankings with a bare requests.get, so
it sent python-requests' default User-Agent, which ESPN rejects; the
AP_TOP_N favourites then resolved to nothing. It now sends
DEFAULT_HTTP_HEADERS. BaseOddsManager carried its own copy of the
User-Agent string and now uses the same shared headers (which also adds
Accept-Language).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(backup): record the core release and read the configured plugin dir

The manifest's ledmatrix_version came from a VERSION file that does not
exist, then from .git/HEAD: a 12-character sha, or "ref: refs/he" when
the branch's ref was packed. It is now src.__version__.

list_installed_plugins() scanned a hardcoded plugin-repos/, so on an
install whose plugin_system.plugins_directory points elsewhere, plugins
missing from plugin_state.json were left out of the backup. It now reads
the configured directory from config/config.json, defaulting to
plugin-repos.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(startup): report a missing display section once

A config without a display section produced three errors for the one
problem ("Missing required configuration key: display", "Display
configuration is missing or empty" and "Display configuration is
missing"), and an empty one produced two. _validate_config now reports
it once, as a missing key or an empty section, and
_validate_display_config leaves it to that.

The module docstring said the validator fails fast; nothing in the
display service calls raise_on_errors(), so it now says the errors are
reported and startup continues.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(wifi): share the copied blocks and name the AP constants

- _parse_nmcli_wifi_list() is the one parser behind _scan_nmcli and
  _scan_nmcli_cached.
- _verify_connected(), _wait_for_device_idle(), _failsafe_ap() and
  _mark_forced() replace blocks that were pasted two or three times in
  the connect and enable-AP paths. The device-idle wait now checks
  before its first one-second sleep instead of after it.
- _check_command() calls _find_command_path() instead of repeating it.
- AP_IP, PORTAL_PORT, AP_PROFILE_NAME and AP_PROFILE_NAMES name values
  that were spelled out 14, 12, 8 and 2 times; the two deletion loops
  now walk the same tuple. The iwconfig status path compares the AP
  address exactly: startswith() also skipped 192.168.4.10-19.
- Dropped a second WIFI.SIGNAL query that repeated the first, a no-op
  "if ssid: continue", the try/except around _connect_wpa_supplicant's
  constant return, and a second save of a scan scan_networks already
  saves.
- _ensure_wifi_radio_enabled's docstring says it returns True when the
  radio state cannot be read at all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(config): drop dead branches and history comments in ConfigManager

- The module docstring pointed plugin authors at update_plugin_config(),
  which does not exist; it now names save_config_atomic() and
  save_raw_file_content().
- load_config's FileNotFoundError handler tested the message for
  "config_secrets.json", but a missing secrets file is handled where it
  is read, so only config.json reaches it; the check is gone.
- save_raw_file_content's `file_type == "main" or "secrets"` guard was
  always true (anything else raised earlier).
- get_raw_file_content('secrets') already returns {} for a missing file,
  so the os.path.exists() in front of two calls to it is gone.
- Comments that narrated earlier behaviour are rewritten as what the
  code does now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(background-data): present-tense comments, drop unused API

- Comments that told the history of each fix (what "used to" happen,
  "the old per-delivery release") now state the invariant the code keeps.
- get_statistics() no longer reports a constant 'queue_size': 0, and the
  uncalled clear_completed_requests() is gone (_cleanup_completed_requests
  does that job on every completion). Neither is referenced in core, the
  web UI or the plugin monorepo.

shutdown_background_service() has no production caller either, but it
is the only way to tear down the get_background_service() singleton,
which the tests rely on, so it stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(odds): drop the unread cache_ttl and merge the odds_data branches

BaseOddsManager loaded base_odds_manager.cache_ttl from config and never
used it: cached odds live for the update interval (get_odds' ttl=interval).
No core or monorepo code reads the attribute, so it is gone along with
its log line. The two consecutive `if odds_data:` blocks are one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(backup): one table for the single-file sections

config, secrets, wifi and ytm_auth were each spelled out in create,
preview, validate and restore. _SINGLE_FILE_SECTIONS lists them once,
with the RestoreOptions flag that restores each, and all four walk it.
Restore error messages keep their wording ("Failed to restore
<file name>").

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(fonts): drop FontManager's write-only state and duplicate logs

- fonts_config, font_metadata and font_dependencies were written and
  never read; the performance_stats keys font_load_times, render_times,
  total_renders and the per-call "resolve" timings
  (_record_performance_metric) likewise. get_performance_stats() reads
  only the counters that remain. Nothing in core or the plugin monorepo
  references any of them.
- A failed BDF load was logged twice, by _load_bdf_font and again by
  get_font; get_font's line is the one kept.
- Removed "NEW:" and commented-out cozette entries, the "Copy font to
  assets/fonts" comment on code that copies nothing, and local imports
  of names the module already imports. The deprecated add_font() now
  resolves assets/fonts against the install root.

The @deprecated methods stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(text-helper): cache loaded fonts; drop the pre-textlength fallback

TextHelper declared _font_cache, cleared it and reported its size, but
never stored anything in it. load_fonts() now keeps each (file, size)
it loads there, so clear_font_cache() and get_font_cache_stats() mean
what they say and repeated load_fonts() calls reuse the fonts.

get_text_width() no longer catches AttributeError for Pillow releases
without ImageDraw.textlength; requirements.txt pins Pillow>=12.2.
The class docstring describes what the helper does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(common): fix wrong docstrings in api_helper, permission_utils, snapshot_policy

- permission_utils called 0o2775 "sticky bit"; the 2 is setgid, which is
  what makes new files take the directory's group.
- snapshot_policy pointed at web_interface/blueprints/api_v3.py, which
  is a package now; the health check is in api_v3/misc.py.
- APIHelper.clear_cache() lost a history note and a fallback to a
  clear() method that neither CacheManager nor the testing
  MockCacheManager has. The session headers are built from
  DEFAULT_HTTP_HEADERS instead of a copy of them, and the module
  docstring says what the module offers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(sports): present-tense comments in the shared scoreboard renderers

- sports_scroll and sports_game_renderer comments that referred to "this
  PR", "the old flat 128px card" or what the renderer "previously" did
  now describe the current behaviour and its reason.
- The block explaining why non-finite settings are rejected sat above
  _score_reserve_width; it describes _center_gap_width and now lives in
  it.
- unshare_element_fonts wrapped its import of font_layout.load_truetype
  in an `except ImportError` that cannot fire inside core; the import
  stays at call time so tests can spy on the pinned loader.
- sports_card docstrings that told the history of a fix say what the
  code does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sports-shared): drop dead code, name the ESPN limit

- _get_weeks_data asked for limit=1000, which fetch_espn_scoreboard
  clamps to ESPN_MAX_LIMIT anyway; it now names that constant. Its
  unused `immediate_events = []` is gone.
- _get_season_schedule_dates() returned ("", "") and has no caller in
  core or the plugin monorepo.
- _should_log keeps its warning_type parameter (part of the inherited
  signature, though nothing in core or the monorepo calls it) and its
  docstring says the cooldown is shared across types.
- An unused ImageFont import is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(sync): one follower-mode switch, shared panel defaults

- The class docstring said the leader sends PNG frames. Frames go over
  UDP as raw RGB; PNG is only the Vegas scroll image sent over TCP. It
  now describes both paths.
- _enter_follower_mode() replaces the two copies of "note the leader,
  switch from standalone to follower, log, write status" in the frame
  and scroll-position handlers.
- The rows/cols fallbacks use DEFAULT_ROWS / DEFAULT_COLS from
  src.display_geometry, as chain_length already did.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(style): drop _layout_axis, name the layout group title

- ElementStyleResolver._layout_axis() had no caller in core or the
  plugin monorepo.
- _element_block_from_spec checked spec['size'] was a dict again after
  size_spec already had; it reads size_spec.
- The "Layout Offsets" title written into three generated schema blocks
  is _LAYOUT_TITLE.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(logo-helper): say what the placeholder draws; name the 1.5 box factor

- _create_placeholder_logo's docstring said it draws the team
  abbreviation; it draws an outlined grey box and nothing else. The
  docstring says so, and the "in a real implementation you'd want text"
  comments are gone.
- The 1.5 x panel default logo box, written out six times, is
  DEFAULT_LOGO_BOX_FACTOR.
- ImageDraw is imported with Image at the top of the module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(logos): drop dead code and a duplicate regex in logo_downloader

- _SAFE_LEAGUE_CODE_RE was the same pattern as _SAFE_LEAGUE_RE; both
  checks use the one.
- get_logo_filename_variations reassigned the TA&M case to the list it
  already had; the function returns the two names directly.
- _get_team_name_variations() had no caller in core or the plugin
  monorepo.
- fetch_single_team's docstring was copied from fetch_teams_data; a log
  message read "for{team_id}".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: drop the Pillow<9.1 resample shim and a catch-and-reraise

- adaptive_images fell back to Image.LANCZOS/NEAREST for Pillow < 9.1;
  requirements.txt pins Pillow>=12.2. RESAMPLE_LANCZOS and
  RESAMPLE_NEAREST keep their names (src.common re-exports them).
- CacheManager.save_cache caught CacheError only to re-raise it; the
  disk write is now called directly, with the same result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(api-helper): stop the real CacheManager's cleanup thread

The cache-lifetime tests built a CacheManager and left its cleanup
thread's class-wide claim on the directory in place, which broke
test_cache_cleanup_thread_ownership when it ran later in the session.
The fixture now stops the thread on teardown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): core-common

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:32:29 -04:00
ChuckandClaude Opus 5.5 b11bcfa204 fix(plugins): store and plugin-manager bugs; tidy src/plugin_system (#635)
* fix(store): don't read a ZIP-installed plugin's remote from the LEDMatrix repo

update_plugin looked up remote.origin.url with `git -C <plugin> config
--local` for plugins that are not git checkouts. Under plugin-repos/ git
walks up to the enclosing LEDMatrix repository, so the lookup returned
LEDMatrix's own URL and a plugin missing from the registry was
"reinstalled" from the LEDMatrix repo. Only ask git when the plugin
directory has its own .git, the test _get_local_git_info already uses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(schema): report each missing required field once, by name

validate_config_against_schema ran its own required-fields loop after
Draft7Validator.iter_errors, which already yields one `required` error
per missing field, so every missing top-level field was listed twice.
The validator's copy also printed the schema's whole `required` list
("Missing required property '['api_key', 'city']'") instead of the field.
Drop the loop and take the field name from the error itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(store): stop mangling repository URLs that contain ".git"

install_from_url and fetch_registry_from_url cleaned URLs with
`rstrip('/').replace('.git', '')`, which removes ".git" anywhere:
https://github.com/user/my.github.io became .../myhub.io, so installing
or browsing that repository asked GitHub for one that does not exist.

Add src/plugin_system/repo_urls.py with one anchored normalize_repo_url(),
same_repo() for comparisons, github_owner_repo() and github_api_headers(),
and use them for the five copies of the owner/repo parsing and GitHub
headers in the store and for saved repositories. GitHub URLs are now
recognised by urlparse().hostname everywhere: _get_latest_commit_info
used a substring test, and _install_from_monorepo_api parsed any host.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(store): install a repository whose only branch is not main/master

_install_via_git returned None both when every clone failed and when the
last-resort clone of the repository's default branch succeeded.
_install_plugin_impl papered over it with `and not plugin_path.exists()`;
install_from_url did not, so a repository whose only branch is e.g.
`develop` was cloned, then treated as a failure, then "downloaded" from
main/master archives that do not exist.

After a default-branch clone, return the branch the clone checked out
(read from .git/HEAD), so None means failure and nothing else, and give
both callers the same `branch_used is None` fallback.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): judge the memory limit on each call's own growth

monitor_call stores `metrics.memory_mb = max(previous, growth)`, and
_check_limits compared that high-water mark with max_memory_mb. It never
decreases, so once one update() grew the process past the limit every
later call raised ResourceLimitExceeded and the circuit breaker kept
reopening. Pass the call's own RSS growth to _check_limits; keep the
high-water mark for reporting and document what it measures.

Remove ResourceMetrics.update_average_execution_time: nothing called it,
and it overwrote the running total with the average.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): reload_plugin re-reads the manifest from the discovered directory

reload_plugin read `plugins_dir / plugin_id / "manifest.json"`, ignoring
the discovery map and the plugin_dirs rules. For a plugin whose
directory name differs from its manifest id the path did not exist, the
re-read was skipped without a word, and the reload kept the stale
manifest. Resolve the directory with find_plugin_directory, as
load_plugin does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): drop the always-null last_display from plugin state info

PluginStateManager reported `last_display` from `_last_display`, which
nothing ever wrote, so it was null for every plugin. Recording it in
PluginExecutor.execute_display would not help: get_state_info's only
reader is the web process, whose PluginManager never calls display().
Remove the field, its dict and get_last_display() (no caller in core,
the web UI or the plugin monorepo).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(store): share the rollback and requirements helpers, drop dead code

- install_plugin and _reinstall_with_rollback set aside, discard and
  restore the old copy through _set_aside/_discard_backup/_restore_backup
  instead of two copies of the same blocks.
- The loader and the store run the same pre-pip checks through
  contained_plugin_dir() and requirements_to_install() in plugin_loader.
  They still invoke pip differently (sys.executable -m pip vs. the sudo
  wrapper). `except (BrokenPipeError, OSError)` + `isinstance(e, OSError)`
  becomes `except OSError` checking errno.EPIPE.
- load_module never returns None, so load_plugin's check is gone and the
  docstring says what it raises.
- Remove the always-true JSONSCHEMA_AVAILABLE, the inline re-imports of
  re and permission_utils, the fake status_result object nobody reads,
  hasattr(git_error, 'cmd'), a redundant "merge conflict" test and
  `import traceback` (exc_info=True does it).
- Correct comments: install_from_url names the directory for the
  caller's id when given (not always the manifest id), _get_local_git_info
  saves one git subprocess (not four), _enrich calls two helpers,
  search_plugins documents all its arguments, _find_plugin_path states
  its behaviour instead of a TODO, and history narration is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): tidy base_plugin, correct plugin_manager/state comments

- base_plugin: drop the unused `import logging`; get_display_duration
  runs the instance value and the config value through one
  _positive_seconds() helper instead of two copies of the coercion; the
  'static'/'none'/fallback branches of get_vegas_display_mode, which all
  returned FIXED_SEGMENT, are one; fix the mis-indented validate_config
  example; say that get_supported_vegas_modes/get_vegas_segment_width
  are not consulted by core (kept, plugins override them).
- schema_manager: import expand_style_elements normally rather than
  swallowing an ImportError of a core module.
- plugin_manager: the plugins directory is the configured one
  (plugin-repos/ by default), not plugins/; get_config() returns the live
  dict, not a copy, so the interval cache comments say what it saves.
- state_manager: config_version and the file version are not used to
  detect corruption; say what they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): stop writing data/plugin_operations.json

PluginOperationQueue wrote its finished-operation history to
data/plugin_operations.json after every operation, and read it back only
into its own in-memory list, which only get_operation_history() exposes
-- and nothing calls that. The operation-history endpoint reads
OperationHistory (data/operation_history.json). No code in src/,
web_interface/, scripts/ or test/ reads the file.

Drop the history_file/lazy_load parameters and the load/save code; the
bounded in-memory history stays. web_interface/app.py and the
integration test stop passing the removed arguments. An existing
data/plugin_operations.json is left in place (data/* is gitignored).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): plugin-system

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:32:02 -04:00
ChuckandClaude Opus 5.5 3967a6cffc fix(security): re-harden root sudo helpers; installer fixes; ARCHITECTURE and PERMISSIONS docs (#640)
* docs: add ARCHITECTURE and PERMISSIONS guides

ARCHITECTURE.md maps the processes, the state the display and web
services share through the cache, the display loop, the plugin system,
the web UI and the update path, with links into the code and a
where-to-start table.

PERMISSIONS.md lists who owns what after install, both sudoers files
(and why iptables is not granted), the polkit rule, and which
scripts/fix_perms script to run as which user.

Both are linked from the docs index, along with the MQTT bridge README
and src/common/README.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: correct stale setup, service and troubleshooting claims

- README: quick actions run systemctl on ledmatrix.service (run.py), not
  display_controller.py; use_short_date_format has no effect; the
  installer uses system pip with --break-system-packages, not a venv.
- CONFIG_DEBUGGING: LEDMATRIX_DEBUG must be "true"; logs are in journald.
- GETTING_STARTED, WEB_INTERFACE_GUIDE, TROUBLESHOOTING: enabling a
  plugin, plugin settings, brightness and Vegas settings apply without a
  restart; matrix hardware settings still need one.
- TROUBLESHOOTING: install dependencies with sudo so the root service
  sees them; point permission problems at PERMISSIONS.md instead of a
  project-wide chown.
- ADVANCED_FEATURES: real BackgroundDataService stats keys; Vegas hooks
  return VegasDisplayMode and None falls back to capture; cache files
  are 0660; fix_web_permissions.sh runs as the web user and does not
  touch sudoers.
- STARLARK_APPS_GUIDE: only the linux-arm64 pixlet binary is downloaded.
- HOW_TO_RUN_TESTS: test class examples that exist.
- CLAUDE.md: PluginStoreManager, plugin_dirs.py, monorepo installs via
  the Trees API with ZIP fallback, requirements.txt is optional.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: mark deprecated plugin APIs and state manifest fields once

Methods @deprecated("3.7.0") (the set pinned in test_deprecation.py)
were shown as current API in the quick reference, API reference,
advanced guide, development guide and FONT_MANAGER. Each is now marked
deprecated with its replacement. FONT_MANAGER is rewritten around the
current API; the override editor is gone and override methods are
deprecated.

Required manifest fields were stated three different ways. The API
reference now has one section: the 7 schema-required fields, the 4 the
store refuses without, class_name for the loader, and the 8 to set.
The other guides link to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document every src/common module and every widget

- src/common/README.md covered 7 of 17 modules. It now has a table of
  all of them (purpose, whether plugins import it, release to floor
  on), a short entry each, and logging advice that matches the code.
- SPORTS_UNIFICATION listed two shared modules and called
  sports_helpers the first; it now lists all six.
- The widgets README lists all 28 registered widgets plus the support
  files, and absorbs the parts that only docs/widget-guide.md had
  (x-options.labels, x-advanced, x-display hidden, plugin-file-manager).
  docs/widget-guide.md is now a pointer to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(security): fix_web_permissions.sh re-hardens the root sudo helpers

The script chowns the whole project to the web user. That included
scripts/fix_perms/safe_plugin_rm.sh and safe_pip_install.sh -- the two
helpers /etc/sudoers.d/ledmatrix_web lets the web user run as root -- so
running it turned both into a root shell for whoever can edit them. It
also re-grouped config_secrets.json away from ledmatrix.

After the chown it now does what first_time_install.sh's Steps 11 and
11.1 do: helpers back to root:root 755, and config_secrets.json back to
the web unit's User=:ledmatrix 640. Each step is non-fatal and prints the
manual command if it fails.

Also fixes what the script and its docs claimed: it never configured
sudoers, its closing hint pointed at ./configure_web_sudo.sh (wrong
path), and the README and ADVANCED_FEATURES.md said to run it with sudo,
which it refuses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(security): validate and harden every sudoers drop-in the scripts write

configure_wifi_permissions.sh copied its rules into
/etc/sudoers.d/ledmatrix_wifi without `visudo -c`. A malformed drop-in
makes sudo refuse every command for every user, which on a headless Pi
leaves no way back in. It now checks first and leaves the installed file
alone when the rules do not parse, as the other two writers do. (It
already used mktemp, so that part of the review did not apply.)

It also grants the two literal commands wifi_manager.py runs for
NetworkManager's shared-mode dnsmasq drop-in -- `cp
/tmp/ledmatrix-nm-dnsmasq.conf .../dnsmasq-shared.d/ledmatrix-captive.conf`
and `rm -f` of that file. The directory's mkdir was granted, the file was
not. Both are pinned in test_sudo_allowlist_covers_calls.py.

configure_web_sudo.sh wrote its rules to /tmp/ledmatrix_web_sudoers_$$,
a predictable name in a world-writable directory; it now uses mktemp with
an EXIT trap, as first_time_install.sh does. It sets mode 440 on the
installed file instead of leaving the temp file's mode, and finds visudo
in /usr/sbin when that is not on the user's PATH, which skipped the
check silently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): escape the project path in the DNS-fix and MQTT unit renderers

install_dns_fix.sh and install_mqtt_bridge.sh substituted
__PROJECT_ROOT_DIR__ with the raw path, while the other three renderers
go through sed_escape_replacement from lib_systemd_render.sh. A checkout
under a path containing `&`, `\` or `|` rendered a corrupted unit from
these two only. Both now source the helper and use it, and a test checks
that every placeholder substitution in scripts/install uses an escaped
value.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): stop the installer scripts reporting things that are not true

- first_time_install.sh printed "Password: ledmatrix123" for the setup
  access point. wifi_manager creates it as an open network ("No
  password" on the panel), so it now says so.
- Step 10.1 printed "✓ WiFi management permissions configured" straight
  after its own failure message; install_wifi_monitor.sh printed
  "✓ Package installation completed" after a failed apt install. The
  tick now only follows success.
- Step 7 printed "Web dependencies already installed ... in Step 5" in
  the one branch that runs because Step 5 did not install them, then
  created .web_deps_installed on that basis. It now warns and leaves the
  marker off so the next run retries, as the comment below it intends.
- check_system_compatibility.sh called Debian 12 Bookworm "full
  compatibility confirmed" while first_time_install.sh refuses anything
  but Debian 13. Bookworm, older Debian and non-Debian systems are now
  errors. Its counters used ((X++)), which under `set -e` exits the
  script at the first warning or error (the expression is 0), so the
  check never reached its summary on any system with one.
- configure_web_sudo.sh and configure_wifi_permissions.sh finished by
  testing `sudo -n test -f ...` and `sudo -n nmcli device status`,
  neither of which is granted, so they always reported a failure. They
  now ask `sudo -n -l` about commands the new rules do grant, which
  checks the rule without running anything.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(install): print the completion summary before rebooting

With -y -- and so for every one-shot `curl | bash` install, which always
passes -y -- first_time_install.sh ran `reboot` about 180 lines before
its "Installation Complete / Web UI Access" summary. reboot returns at
once, so the summary printed while the Pi was going down and the SSH
session usually dropped before the web UI address could be read.

The reboot block moves, unchanged, to the very end of the script. The
interactive prompt now also follows the summary. Because the summary now
runs before the -y reboot, its one command that could fail under
`set -Eeuo pipefail` (the SSID lookup, when nmcli reports a connected
device but no active network line) gets `|| true`; a missing SSID was
already handled as "SSID unknown".

one-shot-install.sh prints its "Next steps" after the installer returns,
by which time the reboot is under way, so it now says so, and README's
Quick Install mentions the automatic reboot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(scripts): correct wrong comments and messages, drop dead code

No behaviour change except the output text noted below.

- 2775 is setgid, not the sticky bit (first_time_install.sh Step 3.1,
  fix_plugin_permissions.sh), and root needs no "PWM hardware access"
  to plugin files.
- The 777 comments in first_time_install.sh Step 3's fallback and
  fix_assets_permissions.sh said root needs it to write. Root ignores
  mode bits; the comments now say what 777 actually opens. The 777
  itself is unchanged.
- apt_remove ends in `|| true`, so Step 12's "Some packages could not be
  removed" branch could never run; it is gone and the helper stays
  non-fatal.
- detect_web_service_user's comment named Step 8 for the web unit
  (install_service.sh installs it in Step 7.5) and now says which
  branch actually runs.
- Step 5 described an "already installed" check that does not exist;
  the ACTUAL_USER comment described the re-exec backwards.
- on_error printed a literal "\n" before "Common fixes:".
- Dead code: one-shot-install.sh's uncalled fix_tmp_permissions,
  LEDMATRIX_ELEVATED=1 (never read) on the sudo re-exec, and
  configure_web_sudo.sh's unused PYTHON_PATH, which also made a missing
  python3 fatal for rules that never mention it.
- start_display.sh / stop_display.sh said "for user: <you>"; the
  service runs as root.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(fix_perms): fix_cache_permissions.sh uses setup_cache.sh's model

There were two models for /var/cache/ledmatrix. setup_cache.sh (the
installer's Step 2) and install_web_service.sh share it through the
ledmatrix group: root:ledmatrix, 2775, files 660, which is also what
DiskCache relies on to give files the directory's group.
fix_cache_permissions.sh instead made it 777 and re-grouped it to the
invoking user's group, undoing that.

It now runs setup_cache.sh for /var/cache/ledmatrix and keeps its own
handling of ~/.ledmatrix_cache. Dropped: /var/cache/ledmatrix/
placeholder_logos (nothing reads it) and the checks against the
`daemon` user (no service runs as daemon).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: pin actions/checkout in the Claude workflows, drop template comments

claude.yml and claude-code-review.yml used actions/checkout@v4 while
test.yml and release-version-check.yml pin the v4.2.2 commit SHA; they
now pin the same SHA. The commented-out starter-template settings
(prompt, claude_args, paths, author filter) are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(scripts): index every script and list removal candidates

New scripts/README.md gives one line per top-level script and scripts
directory, marked keep, dev-only or diagnostic, and lists the eight
scripts nothing in the repo refers to as candidates for removal (kept
for now). The install, utils and dev READMEs now list the files they
were missing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: tighten two checks that mutation testing showed were too loose

- The wifi sudoers check matched `visudo -c -f "$TEMP_SUDOERS"` in the
  error report too, so replacing the check with `if false` still passed.
  It now requires the command as the condition.
- The summary test never had the setup access point up, so reinstating
  the bogus "Password: ledmatrix123" line went unnoticed. A case with
  hostapd active now checks the AP is described as open.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(permissions): describe the repaired fix_perms scripts and new WiFi grants

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): docs-scripts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 17:31:41 -04:00
ChuckandClaude Opus 5.5 4e61d7248a refactor(web): one error-response path for api_v3 (#624)
* refactor(web): answer unhandled api_v3 errors from one blueprint handler

Fifty-three api_v3 routes ended in a copy of the same catch-all: log the
traceback, return {status, "An error occurred; see logs for details",
details: describe_exception(e)} with a 500. They are replaced by one
errorhandler on the api_v3 blueprint that returns exactly that body.

It lives on the blueprint rather than falling through to app.py's global
handler because the two answers differ: the global one adds
error_code: UNKNOWN_ERROR, and api_client.js sends a body with an
error_code to the error modal and one without to a plain toast. A
blueprint handler also gives tests that mount api_v3 on a bare Flask app
the same answer the real app gives.

Only handlers that were byte-for-byte that shape were removed (matched on
the AST, and each rewritten function re-parsed and compared). Handlers
with their own message, extra keys, operation-history records or cleanup
stay, as does execute_plugin_action's step-1 handler, which sits inside
an `except subprocess.TimeoutExpired` arm that would otherwise turn a
plugin's timeout into a 408.

HTTPExceptions raised inside a route go back as themselves in the global
handler's 4xx shape. Where a removed catch-all used to swallow one (only
delete_plugin_asset's non-silent get_json() is reachable), a malformed
request now gets its 415/400 instead of a 500.

Most of the diff is re-indentation from unwrapping the try blocks;
`git diff -w` shows the real change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): plugin action errors name the real failure, not UnboundLocalError

execute_plugin_action bound a local `logger` in its JSON-parsing arm,
which made `logger` local to the whole function. Every other
`logger.error` in it then raised UnboundLocalError, so a failing OAuth
step-1 script was reported as "UnboundLocalError: cannot access local
variable 'logger'" -- from the step-1 handler, and before the previous
commit from the route's outer catch-all too. Use the module logger.

Found by comparing every api_v3 route's forced-failure response before
and after the catch-all consolidation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop the error category and exception-name code guessing

WebInterfaceError derived an ErrorCategory from every error code and put
it in each structured error body as `error_category`. Nothing reads it:
not the web UI (static/ and templates/), not the tests beyond the ones
pinning the mapping itself, and not any plugin in ledmatrix-plugins. The
enum, the inference table and the JSON key go.

from_exception() could also guess an error code from the exception's
class name ("Config" -> CONFIG_LOAD_FAILED, and so on). Every caller
passes a code, so the guess never ran; error_code is now required.

suggested_fixes stays: the error dialog in static/v3/js/utils/
error_handler.js lists them.

The REST reference loses error_category and says what an unanticipated
exception in an /api/v3 route answers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one call for the from_exception error responses

Nine plugin routes built a structured error by hand:

    from src.web_interface.errors import WebInterfaceError
    error = WebInterfaceError.from_exception(e, ErrorCode.X)
    return error_response(error.error_code, error.message,
                          details=error.details, context=error.context,
                          status_code=500)

That is now exception_error_response(e, ErrorCode.X) in api_helpers, so
error_response() is the only structured-error entry point the routes
use. The three operation-history routes never passed the context, and
with_context=False keeps their bodies exactly as they were; a test
compares the helper against the hand-written pair for both forms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): one api_v3 error-response path

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:53:19 -04:00
ChuckandClaude Opus 5.5 ece416c4e5 refactor(plugins): one plugin-directory resolver (#623)
* refactor(plugins): one resolver for plugin id -> directory

Five places mapped a plugin id to its directory, each with its own rules
and each re-reading manifests per lookup: PluginManager discovery and
get_plugin_directory, PluginLoader.find_plugin_directory,
PluginStoreManager._find_plugin_path / list_installed_plugins, and
state_reconciliation.disk_plugin_ids. They disagreed on backup dirs,
on whether the manifest id or the directory name is the id, on duplicate
ids and on path safety.

src/plugin_system/plugin_dirs.py now holds the rules once:
PluginDirectoryIndex scans one directory and reads each manifest once;
resolve_plugin_dir() searches directories in order. What legitimately
differs per caller is an explicit argument: search dirs (discovery and
the loader: configured dir only; the store: configured then sibling
plugins/), ledmatrix- prefix (not for the store), case folding (loader
only), manifest pass (not for get_plugin_directory, whose discovery map
already holds it).

Behaviour changes, all for layouts installs do not produce:
- a directory whose manifest declares the id beats one merely named for
  it (discovery already worked this way; the loader and store now agree)
- the store searches the configured dir completely before plugins/
- backup and hidden dirs are skipped everywhere (the loader's case and
  manifest scans and list_installed_plugins used to return them)
- duplicate ids resolve deterministically (exact name, then
  ledmatrix-<id>, then by name) with a one-time warning; discovery no
  longer lists the id twice
- disk_plugin_ids / list_installed_plugins report manifest ids, falling
  back to the directory name; auto-update looks the directory up
- ids that are not one plain path segment resolve to nothing in every
  caller (the loader used to truncate them, the store to join them)

The .standalone-backup- marker is one constant, BACKUP_MARKER, used by
store_manager's rename-aside names and every lookup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): one plugin-directory resolver

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:52:52 -04:00
ChuckandClaude Opus 5.5 13bbb537f3 refactor(web): one logging setup and one TTL cache for the web process (#621)
* refactor(web): use src.logging_config in the web process; routine requests to DEBUG

The web interface had its own logging setup (web_interface/logging_config.py)
that replaced the root handlers with a plain stdout formatter. The web
service's journal lines therefore never carried a syslog priority, so
`journalctl -p err -u ledmatrix-web` returned nothing while errors were
logged, and the line shape differed from the display's (the log viewer's
prefix stripping only matched the display format). It also ran after the
module-level managers were built, so their INFO lines at import (including
"Re-removed N uninstalled plugin(s)") were dropped.

app.py now calls src.logging_config.setup_logging() first thing, the same as
run.py: journald priorities under systemd, LEDMATRIX_DEBUG honoured,
LEDMATRIX_JSON_LOGGING still selects JSON.

Per-request logging moves to web_interface/request_logging.py. Every request
used to be logged at INFO, so the UI's polling filled the journal
("GET /api/v3/errors/summary - 200" every minute per tab). Now a successful
GET/HEAD/OPTIONS is DEBUG, a successful write is INFO, 4xx WARNING, 5xx
ERROR. Durations use perf_counter and print to 0.1ms.

The duplicate module is deleted; nothing else imported it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one thread-safe TTL cache for the web process

web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.

Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
  counted and get_cached defaulted to 60s. An entry now expires after the TTL
  it was stored with; a reader's ttl_seconds can only shorten that. Both
  current callers pass the same value on both sides (fonts_catalog 300s,
  system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
  same expired key could raise KeyError (reproduced), which the endpoints
  turned into a 500.

app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.

Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): web logging and TTL cache

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): only ask systemctl about known units

Codacy flagged the systemctl argv built from a variable. The unit now has
to be one of two literals, and anything else raises.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): response_time_ms reads the same clock request_logging stamps

request_logging now stamps request.start_time from perf_counter, but
success_response still subtracted it from time.time(), so metadata
reported ~1.8e12 ms. Found testing on ledpi.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:52:17 -04:00
ChuckandClaude Opus 5.5 afe9001aed refactor(fonts): one BDF loader and one BDF rasterizer (#627)
* refactor(fonts): one BDF loader and one BDF rasterizer

BDF faces were loaded three ways (FontManager._load_bdf_font,
element_style._load_bdf, DisplayManager._load_fonts) and drawn by two
copies of the same per-pixel loop (DisplayManager._draw_bdf_text and the
plugin test harness's "replicated" copy), which golden images and
check_plugin/dev_server previews rely on matching the panel.

src/common/bdf_font.py now owns both:
- load_bdf_face(path, size) -> (face, realised_px): native-strike fallback
  for sizes the file lacks, one bounded LRU cache keyed on path, size and
  mtime. FontManager, element_style and DisplayManager delegate to it;
  read_bdf_native_size moves here (the old names delegate).
- draw_bdf_text(draw, text, x, y, face, color, clip): builds each glyph as
  a 1-bit mask and fills it with ImageDraw.bitmap instead of a draw.point
  per pixel. A blending Draw (RGB image, "RGBA" mode) keeps the point path
  so translucent colours still blend.

Pixel-identical: 220,032 renders (every bundled BDF at native and
off-strike sizes, 14 strings, 4 colours, clipped on every edge, through
each old loader x rasterizer) match origin/main byte for byte.
test/test_bdf_font.py keeps a lightweight version against a frozen copy of
the old loop. DisplayManager._draw_bdf_text goes from 1.4-23 ms to about
0.1 ms per string (the old loop re-read FreeType's buffer as a Python list
for every pixel).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(testing): harness calendar_font is sized like the panel's

VisualTestDisplayManager built its 5x7 calendar_font / bdf_5x7_font as a
bare freetype.Face. With no size set its ascender reads 0, so BDF text
drawn with it landed 6px above where DisplayManager draws it -- entirely
off the canvas at y=0 -- and get_font_height() returned 0. Golden images
and check_plugin / dev_server previews showed text the panel does not.

Load it through load_bdf_face at the panel's 7px, so it is the very face
DisplayManager uses. Across the differential run this changes only the
cases drawn with the harness's own calendar_font (968 of 220,032), which
now match the panel's output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): one BDF face per thread

The shared face cache now hands every loader (FontManager, element_style,
DisplayManager, the harness) the same freetype.Face. FreeType does not allow
two threads to use one face at once, since load_char rewrites its glyph
slot, so key the cache by thread as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:51:53 -04:00
ChuckandClaude Opus 5.5 abedc46104 refactor(sports): merge the sports_shared/sports_card twins that behave identically (#626)
* refactor(sports): wrap the sports_card twins that behave identically

SportsCoreSharedMixin (switch mode, via each scoreboard's sports.py) and
sports_card (scroll/Vegas mode, via game_renderer.py) carried the same
helpers twice. test/test_sports_twins.py now calls every pair with the
same inputs -- the eight scoreboards' harness fixture games in flat,
flat+nested and nested-only shapes, plus edge cases (favourites by id and
abbreviation, NRL's colliding abbreviations, missing and non-numeric
scores, bad zones, out-of-range dates, shared font faces).

Identical pairs become thin wrappers over the sports_card function:
_card_option, _vs_text, _format_game_time, _coerce_rgb, _crisp_size (with
the class's own tables), _unshare_element_fonts (with the class's own
element map, via a new optional argument), and the colour/month/weekday/
font-grid tables (dicts copied, not aliased). _format_game_date shares the
card's formatting body but keeps its own setting, weekday zone and month
table; _schema_font_size shares the parser but keeps its per-class cache,
because a reloaded plugin gets new classes and a shared path cache would
stop it seeing an edited schema. _resolve_font_size agrees but keeps its
body so it still dispatches through the overridable hooks.

No behaviour change: old and new mixin/card agree on all 22,994
comparisons over the test corpus, and the pairs that do differ
(favourite-result colours on nested payloads and by favourites source,
the weekday's timezone, the element-name map, per-mode colours) are left
alone and pinned in TestPinnedDivergence for an owner decision.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(sports): pin that an ambiguous NRL abbreviation tints in both modes

NRL's resolver passes a shared abbreviation ("NEW") through with an error
and its _is_favorite_game matches ids only, but both favourite-colour
helpers match on abbreviation as well, so both display modes tint a
Knights or Warriors result for a user who typed "NEW". The twins agree;
neither consults the _favorite_key seam. Pinned so a fix is deliberate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:51:35 -04:00
ChuckandClaude Opus 5.5 1fe7237799 refactor(install): generate the web sudoers rules in one place (#622)
* refactor(install): generate the web sudoers rules in one place

/etc/sudoers.d/ledmatrix_web was written by two copies of the same
allow-list: a heredoc in first_time_install.sh Step 10 and a block of
echo lines in scripts/install/configure_web_sudo.sh. They drifted before
(safe_pip_install.sh was granted by one only), and a test existed just
to catch that.

Both now call web_sudoers_rules() from the new
scripts/install/lib_sudoers.sh and keep their own validate (visudo -c),
install and confirm flows.

- first_time_install.sh output is byte-for-byte unchanged, so a device
  re-running the installer gets "already up to date". If the library is
  missing, Step 10 keeps the installed file and carries on, the same way
  it handles rules that fail visudo (an empty file would pass visudo).
- configure_web_sudo.sh now writes the installer's layout: same 18 rules,
  different comments and order. It still leaves out reboot, poweroff and
  journalctl when they are missing; the library does that for both.

The drift test now pins the generator's grants, checks that neither
installer writes rules of its own, and runs each installer's call line
to check the argument order. Tests that read the rule text now read the
library.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(install): detect the web service user in one function

first_time_install.sh pasted the same WEB_SERVICE_USER detection block
three times (Step 3.1's fallback, the plugin-repos setup and Step 11).
The copies were identical apart from comments; they now call
detect_web_service_user(), whose body is that block unchanged.

Behaviour is the same: the function sets the same global and always
returns 0, as the inline if-chain did. Checked on Linux against all
three original copies across 13 layouts (installed unit with and without
User=, the repo as shipped, each grep branch, template placeholders).

The comment notes that the install_web_service.sh / install_service.sh
greps no longer match anything, so until Step 8 installs the unit the
result is "root". That behaviour is left as it was.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:51:14 -04:00
ChuckandClaude Opus 5.5 ddf5f085a5 perf(cache): tell a stale record from its header instead of parsing it (#633)
The sports plugins cache whole season schedules: 53MB for MLB, 18MB for
NHL, 17MB for NCAA baseball. On a Pi 4, orjson.loads of the MLB file
takes ~1.8s with the GIL held, and every thread in the display service
waits -- the stall watchdog caught the render thread frozen 0.5-1.3s with
the interpreter itself blocked, right on these reads. When a season record
expired, DiskCache.get paid that whole parse only to find the timestamp
too old and throw the result away.

CacheManager.set now writes timestamp and ttl ahead of the data, and
DiskCache.get reads them from the first 256 bytes of the file, applying
the same rule as before (a per-entry ttl wins over max_age; no limit
means never stale). A record that is stale is refused without being
parsed. Files in the old layout, and records from other writers, don't
match the header and are parsed in full as before.

Also: ESPN responses in the background data service and espn_dates are
parsed with orjson when it is installed (src/common/json_body.py). The
stdlib parser behind response.json() takes 3.1s on the MLB season
against orjson's 1.8s, both with the GIL held. espn_dates imports it with
a fallback, since plugins bundle copies of that module for older cores.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:50:50 -04:00
ChuckandClaude Opus 5.5 5baf983fe0 docs(scroll): explain the tear across the middle on fast scrolls (#620)
* docs(scroll): explain the tear across the middle on fast scrolls

A 1:32-multiplexed 64-row panel lights row 31 almost a whole refresh after
row 32, so fast scrolls show a sideways offset at mid-height of about
speed x refresh period. Documents the cause, how to read the real refresh
rate (show_refresh_rate prints with a carriage return), what was measured on
a single-chain 2x128x64 Pi 4 (pwm_bits, gpio_slowdown and an uncapped
refresh barely help; gpio_slowdown 2 glitches), and the fix that does help:
fewer pixels per output via parallel chains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(scroll): limit the 1:32 row-pair explanation to panels that scan that way

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:50:33 -04:00
ChuckandClaude Opus 5.5 82f3a3a3e4 fix(redaction): make credential redaction linear, not quadratic (#631)
* fix(redaction): make URL-userinfo redaction linear, not quadratic

_REDACT_URL_USERINFO could start a match at every letter of a run of
scheme characters, and each attempt read to the end of the run looking
for `://`. On a long unbroken run of letters or digits (a hex digest, an
ID, part of a response body) that is quadratic: 1.6s for 20k characters.

The display service redacts every message, stack trace and context value
it publishes in the error snapshot, holding the aggregator lock, and
re.sub holds the GIL for the whole call, so one such exception stalled
every thread, render loop included (~0.5s measured for 20k chars of hex).
It also made test_snapshot_stays_small the slowest test in the suite by
far: 142s of a 383s run, 139s of it in this one regex.

A match may now only start where a run of scheme characters starts
(negative lookbehind). Leading digits and `+.-` are captured in group 1
so the substitution restores them, and the scheme still has to start
with a letter, so what gets redacted is unchanged: old and new output
were identical on 300k fuzzed inputs. 20k chars now take ~0.5ms, 200k
~6ms, and test_snapshot_stays_small takes 0.8s.

test/test_redaction.py pins the exact output for schemes that begin after
digits or `+.-`, and bounds 50k-character runs at 1s; against the old
pattern those timing tests fail at 3-11s each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK

* fix(redaction): make Authorization-header redaction linear too

_REDACT_AUTH_HEADER matched the value's opening as `\s*["\']?\s*`: two
`\s*` separated only by an optional quote. With no quote, a whitespace
run could be split between them in every possible way, and when no
credential followed (end of text, or `,` `"` `<` ...) the engine tried
them all before giving up: quadratic, 8s for `authorization:` and 20k
spaces, 17s with `Proxy-Authorization:` (tried again at the inner
`authorization`). Same stall as the URL pattern: re.sub holds the GIL,
and the display service redacts everything it publishes.

The quote and the whitespace after it are now one optional unit,
`\s*(?:["\']\s*)?`, which matches the same strings with only one way to
split them. Output is identical to the old pattern on 300k fuzzed
inputs; 20k spaces now take ~1.6ms. A scan of all three redaction
patterns over prefix/run/suffix shapes finds none left that scales
superlinearly.

test/test_redaction.py pins exact output for quoted, tabbed, multi-line
and credential-less headers, and bounds header + 20k whitespace at 1s;
against the previous pattern those fail at 8-17s each.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 15:49:51 -04:00
ChuckandClaude Opus 5.5 c1ce0b7b04 fix(web): two api_v3 paths called names that no longer exist (#625)
The Pixlet editor stop route restarts the display after a SIGKILL with
_run_systemctl_command, which starlark.py never imported (since #554). The
Starlark device-location resolver fell back to _ensure_cache_manager, which
#609 deleted; the resolver already accepts no cache manager. Both raised
NameError on the rare path that reaches them. pyflakes finds no other
undefined names in src/ or web_interface/.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:49:29 -04:00
ChuckandClaude Opus 5.5 f3894916a9 feat(web): show which plugins use each font; warn before deleting one (#619)
* feat(web): show which plugins use each font, warn before deleting one

The Fonts tab lists font files from the web process's own scan, and the
plugins that register fonts run in the display process, so the tab had no
way to say whether a font was in use before deleting it.

The display service now publishes {catalog key: [plugin ids]} to the
shared cache (font_usage_snapshot, src/font_usage.py), built from the
loaded plugins' FontManager.register_manager_font() registrations. A
daemon thread checks every 10 s and writes only when the usage changed
(plus a daily refresh so cache cleanup cannot expire it); it never raises.
Families, aliases (press_start, four_by_six, ...) and paths are resolved
through FontManager's catalog to the file stem the Fonts tab keys rows by;
fonts outside assets/fonts are left out. Unloading a plugin drops its
registrations (new FontManager.forget_manager_fonts).

GET /api/v3/fonts/catalog merges used_by into each row per request (the
5-minute scan cache is copied, never edited): a list of plugin ids, or
null when the display service has not reported. The tab shows a Used by
column ("unknown" / "-" / ids, rendered as text) and deleting an in-use
font names the plugins in the confirmation, from a fresh read. The server
still refuses only system fonts. Catalog fetches bypass the browser's
5-second API cache, which otherwise served the pre-delete list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: call forget_manager_fonts through a hasattr check pylint can follow

getattr(..., None) then callable() is fine at runtime, but pylint's E1102
("not callable") can't see through it, and Codacy fails the check on it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:50:42 -04:00
ChuckandClaude Opus 5.5 9a1f94f793 fix(starlark): blank app locations use the device location, not San Francisco (#617)
* fix(starlark): blank app locations use the device location, not San Francisco

A Starlark (Tidbyt) app whose Location field is blank rendered at its
author's hard-coded DEFAULT_LOCATION -- usually San Francisco -- even with
the device city set under General settings. A user in Charlotte, NC got San
Francisco weather and radar with nothing in config.json to explain it.

src/device_location.py fills unset location fields at render time (display
plugin and the web standalone render): the device city is geocoded once via
Open-Meteo, preferring a match in the configured state/country, and cached
permanently. A saved location always wins; if the lookup fails the field is
dropped so the app uses its own default, and the failure is not retried for
30 minutes.

Also fixes the config form: clearing a location omitted the key, and the
save merges, so the old value could never be removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(starlark): say what happens when the device location can't be used

A blank app Location only renders at the device's city when one is set and
the Open-Meteo lookup finds it. With no city, no match, or the geocoder
unreachable (retried after 30 minutes), the app gets no location and keeps
its author's default. The guide, the config page hint, CONFIG_REFERENCE and
the CHANGELOG entry now say so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:48:02 -04:00
ChuckandClaude Opus 5.5 61e462c635 refactor: remove the skin system and the unused src/base_classes package (#615)
* refactor: remove the skin system

Skins never rendered with the current scoreboard plugins: the only hook was
SportsCore._render_game in src/base_classes, which no plugin builds on, so
the UI and store already treated them as unsupported. The owner decided on
2026-09-23 to remove them outright.

Removed src/skin_system/ (runtime, base class, fixtures), skins/,
scripts/validate_skin.py and their tests; the store's "type": "skin"
installer, uninstaller and hide/refuse filters (the official registry lists
no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins.

Stored skin/skin_options config values are handled in the next commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(config): drop retired skin/skin_options keys instead of validating them

A config.json written while the skin system existed can carry skin and
skin_options in any plugin section, and most plugin schemas set
additionalProperties: false. They are no longer core plugin properties;
RETIRED_PLUGIN_KEYS in schema_manager lists them and
drop_retired_plugin_keys removes them (unless the plugin's own schema
declares the name) in prepare_plugin_config, which loading, hot reload,
GET /plugins/config and both web saves already share, and in
validate_config_against_schema for callers that validate a raw section.
POST /plugins/config and /config/main also drop them from the stored
section they merge into, so they leave config.json on the next save.

Tests cover the load path (real PluginManager.load_plugin: no schema
warning, not degraded), raw and prepared validation,
validate_all_plugin_configs, and the JSON, form and /config/main saves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: remove the unused src/base_classes package

No scoreboard plugin builds on src.base_classes: the nine monorepo
scoreboards ship their own sports.py and share code through src/common
(docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins
imports it. The one import anywhere, baseball-scoreboard's
rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing
instantiates.

Removed the package and the eight test files that only tested it
(test_api_extractors, test_data_sources, test_sports_base_characterization,
test_sports_capabilities, test_sports_core_promotions,
test_sports_logo_cache_bounded, test_sports_modes_promotions,
test_sports_odds_fanout). test_common_is_hardware_free no longer lists
src.base_classes as a forbidden import, and comments in sports_helpers.py
and base_odds_manager.py stop pointing at it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: drop the skin system and src/base_classes from the docs

Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to
them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins
section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the
term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was
removed and shared code lives in src/common, in the Layering section and
the view-model-contract rule. Other docs stop pointing at the removed
package. CHANGELOG records both removals under Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(store): hide and refuse registry entries that aren't plugins

The skin filters went with the skin system, but a custom registry can still
list "type": "skin" entries, and installing one as a plugin would unpack it
into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing
type means plugin) now hides non-plugin entries from the store and
custom-registry listings, and install refuses them, in the route with a
clear 400 and in _install_plugin_impl for any other caller.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:33:24 -04:00
ChuckandClaude Opus 5 4fe3cdd906 fix(starlark): stop the root display service locking the web UI out of starlark-apps (#604)
* fix(starlark): stop the root display service locking the web UI out

Reported after a fresh install: installing an app from the Starlark tab
failed with "install failed: Failed to install from repository", and so did
uploading a .star file and installing from a GitHub directory. The reporter
found the cause only by reading service logs, and fixed it with

    sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/starlark-apps

starlark-apps is gitignored, so it is never checked out -- it is created
lazily by whichever process reaches it first. Those processes run as
different users. systemd/ledmatrix.service is User=root and constructs this
plugin at startup, which is where _get_apps_directory() is called from;
systemd/ledmatrix-web.service runs as the login user and is what actually
installs apps.

The documented first step is to install pixlet and reboot, so on a fresh
machine the display service usually wins that race and mkdir() leaves the
directory root-owned. The web process then fails in _install_star_file() on
app_dir.mkdir(), which catches nothing, so PermissionError reaches the
route's outer `except Exception` and becomes the generic message the user
saw. All three install paths write to the same directory, which is why all
three failed.

The web user cannot repair this -- chown needs root. So root does it, on
every startup, which also heals machines already broken by this without the
owner having to find the chown themselves. It is a no-op when not root, when
the platform has no POSIX ownership, and when the checkout genuinely belongs
to root; a chown that fails warns rather than killing startup.

Also made the failure legible if the handover is ever prevented: a
PermissionError now names the directory, the automatic repair, and the
manual chown, instead of a message that names neither path nor cause.

Verified by mutation: dropping the handover call, chowning a genuinely
root-owned checkout, and letting a non-root process chown each fail their
own test. 121 starlark tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): address the review on the ownership repair

Findings from the automated review of #604.

Symlinks (CWE-59, the serious one). A root chown that follows links is a
privilege-escalation primitive: anyone able to write in starlark-apps could
point a link at a root-owned file and have the repair hand it over. Entries
are now read with os.lstat, symlinks are skipped outright, and the chown
passes follow_symlinks=False. Descendants are processed before the directory
itself, so the container does not change hands while its contents are still
being walked.

install_app() caught PermissionError in its broad handler and returned
False, which both routes report as a generic install failure -- the exact
shape of the bug this PR exists to fix, since the caller could not tell
"this app is broken" from "this process cannot write here". PermissionError
is now re-raised; every other failure still returns False.

The test fixtures skipped on bare Exception, which would have turned a
syntax error or NameError in the plugin into a green run. They now skip only
for a named absent dependency and re-raise anything else.

Also fixed the _Stat stub that failed in CI but passed locally: it carried
only st_uid/st_gid, and pathlib reads st_mode while walking. It now wraps
the real stat result and overrides ownership alone.

NOT taken: the CodeQL "information exposure through an exception" finding on
the hint response. Dropping `details` would contradict this package's
documented rule -- "if it returns 5xx, it says why" -- which
test_no_api_v3_handler_discards_its_exception enforces with an allowance
that may shrink and never grow. The Starlark routes are the ones that policy
was written for: they answered 500 with no detail for three releases.
describe_exception already redacts credentials and truncates. Keeping the
detail is the deliberate trade-off, so the finding is declined rather than
silently worked around.

Verified on hdpi with the updated code: a symlink to /etc/shadow planted in
starlark-apps was skipped while the directory was handed back, and
/etc/shadow stayed root:shadow.

Mutation-checked all three behaviours. The symlink test was vacuous on the
first attempt -- the link already had the target owner, so it was skipped
for the wrong reason and the mutation passed. It now forces the link to look
like it needs handing over, and fails when the check is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 16:33:12 -04:00
ChuckandClaude Opus 5.5 4a1fd7464a fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)
* fix(errors): serve /api/v3/errors/* from the display service's aggregator

The error aggregator is a per-process singleton and only the display
service runs plugins, so only its aggregator records anything. The routes
read the web process's own, empty one and always reported no errors.

The display service now publishes a bounded snapshot of its aggregator to
the shared cache (plugin_error_snapshot) from a daemon thread: at most once
every 10 s and only when something changed, never raising into the caller.
The routes read it and keep their response shapes, adding
snapshot_available, generated_at and clear_pending; exception text has
credentials redacted.

POST /errors/clear writes a clear request (plugin_error_clear_request) that
the display applies on its next 5 s tick via the new clear_before(), which
keeps errors recorded after the cutoff and rebuilds the counts. Until the
snapshot acknowledges the request, reads hide everything before the cutoff,
so a snapshot written just before the click cannot bring errors back. Adds
"all": true; cleared_count is null when only the display can know it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): show plugin errors in the Logs tab

A compact panel under the log viewer: per-plugin error counts, repeating
errors (type, count, affected plugins, a sample message, last seen) and a
Clear button, with empty states for "no errors" and "display service
hasn't reported yet". Polls every 15 s while the tab is active; all text
goes through escapeHtml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe where plugin error reports come from and how clear works

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(errors): redact the published snapshot before clipping it

Keeping only a traceback's tail (or clipping a message) could cut an
`api_key=` marker off while keeping the secret after it, and the web side's
redaction would then have nothing to match. The display now redacts every
free-text field of the snapshot first. The patterns move to a Flask-free
src/redaction.py so the display service can use them; redact_text in the web
error handler uses the same function, unchanged in behaviour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:32:02 -04:00
ChuckandClaude Opus 5.5 cd5a4e2251 fix(display): apply on-demand, brightness and schedule changes mid-screen (#618)
* fix(display): apply on-demand, brightness and schedule changes mid-screen

The main loop read the on-demand mailbox, the on/off schedule and the
brightness target once per pass -- once per screen. A dwell can be a minute
and a Vegas iteration runs for max_cycle_duration (240s), so on a Pi an
on-demand request posted at 10:54:27 was activated at 10:57:24, and two
brightness saves 12s apart inside one 30s screen never reached the panel.
During Vegas nothing read the mailbox at all: _check_vegas_interrupt only
checked on_demand_active, which only the main-loop read sets.

_service_pending_changes does the main loop's on-demand poll, expiry,
schedule and brightness steps, throttled to PENDING_CHANGES_INTERVAL (the
existing 0.25s mailbox floor), on the display thread. It runs from the Vegas
interrupt checker, the high-FPS and once-a-second render loops (replacing
their direct on-demand poll) and _sleep_with_plugin_updates; between passes
it costs one monotonic compare. A brightness change re-pushes the current
frame, since the panel only shows it from the next push.

Callers act on what it leaves behind: Vegas yields on an on-demand start or
the display being scheduled off (and the main loop then blanks instead of
rendering a screen), the render loops break on a schedule-off as they
already did on a mode change, and the dwell sleep returns early on an
on-demand start/stop or a schedule flip -- so the 60s scheduled-off sleep
now wakes for an on-demand request. The main loop no longer rotates after
a dwell that ended that way, which advanced a new on-demand session past
the mode that was asked for.

A brightness set_brightness() refuses is not retried until the target
changes, so the 4Hz pass doesn't log the same failure (fallback mode)
four times a second.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(display): a screen scheduled off midway stops rendering

Covers the schedule-off break added to the high-FPS and once-a-second
render loops: with the display scheduled off halfway through a 120s screen,
neither loop renders for more than one redraw plus one service interval
past the boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:58:39 -04:00
ChuckandClaude Opus 5.5 3f8edf5113 fix(logos): one hardened logo download path; shared, real HTTP headers (#612)
* fix(logos): harden the plugin logo download and share core HTTP headers

download_missing_logo / LogoDownloader.download_logo, the path the
scoreboard plugins use, read response.content with no size cap and wrote
straight to the final path, so a failed or corrupt download could be left
in place and cached as the logo. It now goes through fetch_logo: streamed
with a 10 MB cap, image/* only, decoded by Pillow, converted to RGBA once,
and moved into place atomically. A failure leaves no partial or temp file
and keeps any logo already on disk. LogoHelper._download_logo delegates to
the same code. Public signatures and return values are unchanged; saved
files are pixel-identical to before (RGBA, palette+tRNS, L+tRNS, LA, JPEG).

download_missing_logo reuses one downloader per thread instead of a new
Session per logo. Per thread rather than behind a lock: Session is not
documented thread-safe, and a lock would serialise every plugin's
downloads behind the slowest one.

Placeholders are written atomically, without the test_write.tmp probe.

The logo downloader and background data service now send the real
ChuckBuilds User-Agent from src.common.api_helper (USER_AGENT,
DEFAULT_HTTP_HEADERS) instead of a yourusername/contact@example.com
placeholder, and no longer hand-set Accept-Encoding: br (brotli is not
installed).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(http): drop APIHelper's hand-set brotli encoding; LogoHelper sends the real UA

APIHelper advertised `br` though brotli isn't installed, so a server that
honoured it would send a body requests can't decode. LogoHelper sent a bare
`LEDMatrix-Common/1.0`, the kind of User-Agent ESPN has been rejecting.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:58:11 -04:00
ChuckandClaude Opus 5.5 f813ea2117 fix(web): define project_root when plugins_directory is absolute (#616)
project_root was only assigned in the relative-path branch, so an absolute
plugin_system.plugins_directory made web_interface/app.py raise NameError
at import (first use: the SchemaManager construction). Define it before the
if/else; plugins_dir resolution is unchanged.

Adds a regression test that imports the real module in a fresh interpreter
with an absolute and a relative plugins_directory.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:54:29 -04:00
ChuckandClaude Opus 5.5 9d024f24ef refactor(cache): remove the cache layer's duplicate cleanup and dead lookups (#613)
* refactor(cache): collapse CacheStrategy's all-60 defaults table and twin soccer branch

get_sport_live_interval() without a config manager looked the sport up in
a table where every value was 60, with 60 as the fallback; it now returns
60. get_data_type_from_key() had an `if 'soccer'` branch returning the
same 'sports_live' as its else.

test_cache_strategy_intervals pins the returned strategy for every data
type x sport key x config-manager shape; it passes unchanged on the old
code. A 2,544-entry dump of every CacheStrategy method over a wider grid
is identical before and after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(cache): drop CacheStrategy's `<sport>_scoreboard` config lookup

get_sport_live_interval() and get_cache_strategy() read live/recent/
upcoming intervals from config[f"{sport}_scoreboard"]. Those sections
belonged to the built-in scoreboards the plugin system replaced; plugin
config is keyed by plugin id ("football-scoreboard"), so on a current
config the lookup always fell through to the defaults (60 live, 1800
recent, 10800 upcoming), which are now returned directly.

The one input where this differs: a config.json upgraded from the
pre-plugin era that still carries e.g. an "nfl_scoreboard" section (no
code removes them), queried with an explicit sport key. No caller in core
or the plugin monorepo passes a sport key here -- get_with_auto_strategy
only derives one for keys classed sports_live/live_scores, and its callers
(odds managers, odds-ticker) use odds keys -- so the stale section was
unreachable in practice. A dump of every CacheStrategy method over 2,544
inputs differs from the previous commit only in those 45 legacy-config
entries; the test grid now includes that shape.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(cache): list cache files without holding the memory-tier lock

CacheManager.list_cache_files() held the in-memory cache's lock while it
listed and stat'd the whole cache directory -- 8,864 files on a real rig
-- so every get()/set() from the display loop and plugins waited out the
scan. The lock never protected the disk: DiskCache writes and deletes
under their own lock, and a file vanishing between listdir and stat was
already handled (logged and skipped). The body is unchanged apart from
the dedent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(cache): delegate memory-tier cleanup and stats to MemoryCache

CacheManager._cleanup_memory_cache() was a line-for-line copy of
MemoryCache.cleanup(), and get_memory_cache_stats() a copy of
MemoryCache.get_stats(), both reaching into the component's private
_cache/_timestamps/_lock through "backward compatibility" aliases bound
in __init__. So the component's own cleanup and stats only ever ran in
tests, and the aliases went stale whenever the component was swapped
(test_cache_ttl_honoured does). Both now delegate, and the aliases are
gone: nothing in core, the tests, or the ledmatrix-plugins monorepo reads
them.

Behaviour is the same. Compared line by line, the two cleanups differ
only in the sort key's fallback (0 vs 0.0, which orders identically),
range+bounds check vs slice for the eviction, and the logger name on the
DEBUG summary line (src.cache_manager -> src.cache.memory_cache). A
differential run over 20,000 random memory states (str/None/garbage/
future timestamps, orphan keys, sizes 0-12, forced and throttled runs)
gives identical removed counts, resulting dicts and last-cleanup times;
the same harness catches each of three seeded mutations of
MemoryCache.cleanup. The throttle clock also moves with it:
CacheManager kept its own copy of last-cleanup, the component's is used
now, and they started equal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(background): inline the sport cache key and drop the unused request queue

get_sport_cache_key() constructed a whole CacheManager -- ConfigManager,
config parse, cache-dir probing with test-file writes -- to return
f"{sport}_{date}". It now builds the key itself in the same format as
CacheManager.generate_sport_cache_key() (UTC date, %Y%m%d); tests check
the two agree for explicit dates and, with a frozen clock at 03:30 UTC,
for the default date. Median per call on Windows: ~0.6 ms -> ~2 us
(alternating runs); on a Pi the old path also wrote a probe file per call.

request_queue was a PriorityQueue nothing ever put into: requests go
straight to the executor, so `priority` never did anything. The queue is
gone; the `priority` parameter and FetchRequest field stay (every
monorepo scoreboard passes priority=) and are documented as ignored, and
get_statistics() keeps reporting queue_size, now a literal 0 as it
always was in practice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:54:07 -04:00
ChuckandClaude Opus 5.5 269385c97c fix(config): write config.json through one durable atomic writer (#611)
save_config() opened config.json with 'w' and streamed json.dump into it,
so a power cut or an unencodable value left the file truncated.
save_config_atomic() renamed a temp file into place but never fsynced it,
rewrote the unchanged secrets file on every save, and re-parsed every
backup to rotate them. save_raw_file_content() had its own third copy.

All of them, plus rollback and config creation from the template, now go
through atomic_write_text(): temp file in the same directory, fsync,
final mode set before the rename, rename (retried on Windows while a
reader holds the file), directory fsync. A root save copies the previous
owner onto the new file so a rename by the display service no longer
hands config.json to root; the shared-group fix-up is unchanged. The
mode is chosen from the file name, so a "secrets" directory in the
install path no longer makes config.json 0640.

The secrets file is rewritten only when its content changes, and backup
rotation works from filenames alone. Backups keep their names
(config/backups/config.json.backup.<version>, paired secrets backup) and
the five newest are kept, as before.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:52:50 -04:00
ChuckandClaude Opus 5.5 604f58ff07 feat: deprecate unused plugin-facing methods for removal in 3.7.0 (#610)
35 methods on CacheManager, DisplayManager, FontManager and PluginManager
have no caller in core, the ledmatrix-plugins monorepo or the registry's
third-party plugins, but plugins live elsewhere, so they stay for one
release. src.deprecation.deprecated logs a warning (and emits a
DeprecationWarning) the first time each is called in a process, naming the
release that removes it. The list and replacements are in CHANGELOG and
PLUGIN_API_REFERENCE's new Deprecated APIs section; a test pins the set.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:44:55 -04:00
ChuckandClaude Opus 5.5 a8b3e86775 refactor(web): delete dead routes, JS files and duplicate definitions (#609)
* refactor(web): drop validators nothing calls

escape_html, validate_image_url, validate_font_awesome_class,
validate_mime_type, validate_numeric_range, validate_string_length and
sanitize_plugin_config had no callers outside their own tests. Only
validate_file_upload (fonts upload) is imported by the web interface.

dedup_unique_arrays is kept: its one caller in save_plugin_config was
removed by the unrelated sync PR (#330), which looks accidental.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(api): remove the music-auth and of-the-day JSON routes

POST /plugins/authenticate/spotify and /plugins/authenticate/ytm had no
caller but their tests: the music plugin authenticates through its
web_ui_actions (authenticate_spotify.py / authenticate_ytm.py) via
/plugins/action.

POST /plugins/of-the-day/json/upload and /json/delete looked the plugin
up by the id ledmatrix-of-the-day (its manifest id is of-the-day), were
reachable only from a file_type "json" upload field that no schema
declares, and put the plugin directory on sys.path per request to
import scripts.update_config. of-the-day manages its files through
plugin-file-manager and its own web_ui_actions.

The of-the-day branch of GET /plugins/config stays: it matches the real
manifest id and still merges the on-disk category files into the form.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(api): read managers only from the blueprints

api_v3/__init__.py and pages_v3.py declared module globals
(plugin_store_manager, saved_repositories_manager, schema_manager,
operation_queue, plugin_state_manager, operation_history, sync_manager,
config_manager, plugin_manager) that nothing assigns: app.py sets the
managers as attributes on the Blueprint objects, and every route reads
them there. The one reader, backup restore's fallback to the module
plugin_store_manager, could only ever fall back to None.

_ensure_cache_manager() built a second CacheManager in the web process
instead of using the one app.py puts on api_v3. The display routes now
read api_v3.cache_manager, creating it on the blueprint only when
nothing set it (the same None handling as the /cache routes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(web): drop run.sh and the unused log_config_change

web_interface/run.sh was referenced only by web_interface/README.md;
the service starts the UI through scripts/utils/start_web_conditionally.py
and the README already documents `python3 web_interface/start.py`.
log_config_change() in web_interface/logging_config.py was never called.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete unreferenced store_manager.js, diff_viewer.js, htmx-sse.js

- js/plugins/store_manager.js (window.PluginStoreManager) and
  js/config/diff_viewer.js (window.ConfigDiffViewer) were loaded on every
  page but nothing reads either global.
- htmx-sse.js (plus its CDN fallback) was loaded after HTMX, but no
  template or plugin page uses sse-connect / hx-ext="sse": the live
  streams run through LEDStreams in app-shell.js.

js/plugins/state_manager.js stays: install_manager.js's updateAll()
reads and refreshes window.PluginStateManager.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): remove app.js helpers nothing calls

- hexToRgb, rgbToHex, validateForm, uploadFont and switchTab (whose
  'switch-tab' event had no listener) have no caller in the templates,
  static JS or the plugin monorepo.
- installPlugin: plugins_manager.js (loaded last) assigns
  window.installPlugin, and its own store cards are the only callers.
- The showNotification fallback could never install: app-shell.js is
  deferred ahead of app.js and defines the same fallback at top level.
- performanceMonitor only logged with ?debug=perf and read an unset
  this.measures; the marks it took on every load had no reader.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop app-shell.js refreshPlugin

A top-level function in app-shell.js, so a window global, but nothing
calls it (no inline handler, no window lookup, no string-built name).

The other plugin actions in that block stay. updatePlugin is the live
window.updatePlugin: plugins_manager.js only installs its own copy when
none exists. uninstallPlugin/pollUninstallOperation, updateAllPlugins,
executePluginAction and toggleNestedSection are replaced by later
deferred scripts, but a click that lands while those scripts are still
downloading reaches the app-shell copies, so removing them is not a
pure no-op.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): remove definitions plugins_manager.js always overrides

All of these are replaced before anything can call them, checked
against the load order in base.html and the live window.* values:

- openOnDemandModal/requestOnDemandStop stubs: the IIFE later in the
  same script assigns the real functions synchronously.
- updatePlugin and uninstallPlugin stubs (`window.X || stub`): app-shell.js
  already defined both, so the fallback never installed. Same for the
  later updatePlugin override, gated on the live function containing
  '[UPDATE]', which app-shell.js's never does.
- The first addArrayObjectItem/removeArrayObjectItem: reassigned by the
  top-level copies after the IIFE.
- The first `function formatDate` in the IIFE: a later declaration of
  the same name in the same scope wins.
- deleteUploadedImage, getCurrentImages, showUploadProgress,
  formatFileSize and getScheduleSummary: character-for-character
  copies of js/widgets/file-upload.js, which stays the owner.
- `typeof X === 'undefined'` fallbacks and `typeof X !== 'undefined'`
  re-exports after the IIFE: always false, or a self-assignment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): render the shell directly and delete index.html

index.html extended base.html with {% block content %}, but base.html
defines no blocks, so none of index.html ever rendered: rendering both
with jinja2 gives byte-identical output. index() still loaded the config,
read config.json and config_secrets.json raw and json.dumps'd them on
every page load for variables base.html never reads, and flashed errors
that base.html never shows. It now renders base.html with no context.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): stop htmx-config.js replacing console.error and console.warn

It swapped both globals for filters that dropped any error mentioning
insertBefore / "Cannot read properties of null" when "htmx" appeared in
the message or stack, and a list of Permissions-Policy warnings. That
hid real errors from every script on the page, and made every logged
error and warning report htmx-config.js as its source. The beforeSwap
target validation above it, which prevents the insertBefore errors in
the first place, stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(web): quiet the widget load announcements and debug logs

About 30 lines hit the console on every page load: one "... widget
registered" per widget file, one "[WidgetRegistry] Registered widget: X"
per registration, plus the registry, base widget and plugin loader
announcing themselves. The load-time announcements are removed; the
per-call ones (registry register, plugin widget loads, "Render called")
now go through the page's debugLog switch (localStorage.pluginDebug),
guarded because the widgets also load in node tests without it.
fonts.html and wifi.html debug logging goes through debugLog as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(api): drop the removed music-auth and of-the-day JSON routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:36:53 -04:00
ChuckandClaude Opus 5.5 a231d4dbc7 chore: delete unreferenced scripts and archived docs; fix stale doc claims (#607)
* chore(scripts): delete unreferenced helper scripts

None of these is referenced by an installer, systemd unit, CI workflow,
test, the web UI or src/:

- utils/cleanup_venv.sh removes venv_web_v2, which nothing creates
- utils/clear_python_cache.sh hardcodes ~/LEDMatrix and a .webassets-cache
  nothing uses
- install/migrate_config.sh only copies the template, which the installer
  and ConfigManager already do
- install/debug_install.sh, debug/debug_web_manual.py
- diagnose_web_ui.sh and verify_web_ui.sh overlap diagnose_web_interface.sh,
  which the docs point to
- fix_internet_connectivity.sh is iptables-only (stale on nftables)
- diagnose_plugin_permissions.sh, dev/validate_python.py
- download_nba_logos.py + README_NBA_LOGOS.md: logo_downloader fetches
  logos on demand
- setup_plugin_repos.py linked into the production plugin-repos/ dir; the
  dev workflow is scripts/dev/dev_plugin_setup.sh, and
  MULTI_ROOT_WORKSPACE_SETUP.md now uses it

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(config): drop unused plugin_system flags and a dead unit comment

- config.template.json: remove plugin_system.auto_discover,
  auto_load_enabled and development_mode. Nothing reads them; the web UI
  only stores them when a client sends them. ConfigManager's migration
  only adds template keys, so existing configs keep theirs unchanged.
- config.template.json: re-indent vegas_scroll's live_* keys.
- systemd/ledmatrix.service: remove the comment documenting
  LEDMATRIX_ON_DEMAND_PLUGIN / on_demand_env.conf; nothing reads either.
- CONFIG_REFERENCE.md: say the legacy keys are no longer in the template.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: delete docs/archive and PLUGIN_IMPLEMENTATION_SUMMARY.md

- docs/archive/: superseded guides; the repository history keeps them
  and no live doc links into the directory. The one open document in it,
  WEB_UI_AUDIT_2026-09.md, moves to docs/audits/ and is linked from the
  docs index.
- PLUGIN_IMPLEMENTATION_SUMMARY.md invented usage statistics, called
  v2.0.0 current, listed shipped auto-updates as future work and
  documented a BasePlugin.get_config() that does not exist.
- docs/README.md: drop both, and stop telling contributors to archive
  obsolete pages instead of deleting them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(plugin-api): fix extra_small_font size, cache metric key and scroll pacing example

- PLUGIN_API_REFERENCE: extra_small_font loads at 7, not 6 (crisp_size
  snaps it, src/display_manager.py); get_cache_metrics() returns
  cache_hit_rate, not hit_rate (src/cache/cache_metrics.py).
- ADVANCED_PLUGIN_DEVELOPMENT: the basic scrolling example slept in a loop
  and never passed frame_hold; use ScrollHelper + scroll_config.configure()
  and set_scrolling_state(True, frame_hold=...) as PLUGIN_API_REFERENCE does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(plugin-config): match the config tab, icon and web-action docs to the code

- PLUGIN_CONFIG_QUICK_START / PLUGIN_CONFIGURATION_TABS /
  PLUGIN_CONFIGURATION_GUIDE: there is no "Reset to Defaults" button (the
  tab has Refresh, Update, Uninstall, Save Configuration); plugin config
  hot-reloads (ConfigService + on_config_change), so no restart; the
  schema is found by the fixed name config_schema.json, not a manifest
  config_schema field; the tab row is "Plugin Manager", not "Plugins";
  forms are server-rendered from /v3/partials/plugin-config/<id>; the
  duration hook is get_display_duration()/display_duration; a class_name
  mismatch raises PluginError; the store requires id, name, class_name and
  display_modes (not version); plugin_system.debug/log_level do not exist
  (use run.py -d / LEDMATRIX_DEBUG). Drop "future" features that shipped.
- PLUGIN_CONFIG_CORE_PROPERTIES: list all of CORE_PLUGIN_PROPERTIES,
  including skin, skin_options and the vegas_* tuning keys.
- PLUGIN_CUSTOM_ICONS: icon is only a Font Awesome class (fallback
  fa-puzzle-piece); emoji/URL icons and getPluginIcon() never existed in
  v3. Note that /api/v3/plugins/installed currently omits icon.
- PLUGIN_WEB_UI_ACTIONS (+ example JSON): success_message, error_message
  and step1_message are never read.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(store): describe the monorepo registry and the store UI as they are

- PLUGIN_STORE_GUIDE: the Plugin Store is a section of the Plugin
  Manager tab; URL installs are "Install from GitHub" -> "Install Single
  Plugin"; bulk update exists (Check & Update All) plus opt-in weekly
  auto-update; PluginStoreManager() defaults to plugins/, so the Python
  examples pass plugin-repos; registry plugins are downloaded (GitHub API,
  ZIP fallback), not cloned; updates compare version with latest_version.
- PLUGIN_REGISTRY_SETUP_GUIDE: replace the per-plugin-repo + tag
  walkthrough with a short page on the monorepo registry (plugin_path,
  latest_version, update_registry.py) that points at the monorepo's own
  SUBMISSION.md. Drops the reference to the deleted
  PLUGIN_IMPLEMENTATION_SUMMARY.md and setup_plugin_repos.py.
- plugin_registry_template.json: use the real entry shape.
- PLUGIN_QUICK_REFERENCE: automatic background updates exist (opt-in);
  registry example and publishing steps use the monorepo, not tags.
- PLUGIN_DEVELOPMENT_GUIDE: tags/releases are not read by the store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(readme): fix the Triple Bonnet mapping, install prerequisites and backup names

- README: the Adafruit Triple Bonnet uses `regular` (3 outputs), not
  `regular-pi1` (1 output) -- src/matrix_support.py MAPPING_OUTPUTS, and
  the README's own hardware_mapping section; the template default mapping
  is adafruit-hat, the PWM mod switches it to adafruit-hat-pwm; manual
  install only needs git up front (first_time_install.sh installs
  python-dev-is-python3, cmake, ninja-build etc.; cython3/scons are not
  used); the Pi Zero 2 W is a supported low-memory board, consistent with
  PRODUCT.md, LOW_MEMORY_BOARDS.md and the installer's low-memory build;
  fix the "First_time_install.sh" spelling, an orphan "2." list item and
  the hello-world starter link (it lives in the plugins monorepo).
- CONFIG_DEBUGGING: automatic backups are
  config/backups/config.json.backup.<YYYYMMDD_HHMMSS_ffffff> (five kept),
  not config_YYYYMMDD_HHMMSS.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(dev): correct the test-running and rgbmatrix build instructions

- HOW_TO_RUN_TESTS: coverage is not collected by a plain pytest run and
  pytest.ini has no threshold; the only one is --cov-fail-under=52 in the
  core unit-test job of .github/workflows/test.yml, which runs the whole
  test/ tree (not an allowlist). Almost no tests carry markers, so
  -m integration / -m slow select nothing; drop them and -m unit as the
  quick check. Replace the hardcoded /home/chuck path.
- DEVELOPMENT: the rgbmatrix package is built with pip install . from
  the submodule root (scikit-build-core + CMake + Ninja), as
  first_time_install.sh does; there is no make build-python /
  bindings/python step, and the build deps are python-dev-is-python3,
  cmake and ninja-build, not cython3/scons.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(wifi): the setup AP is open; auto-enable can be turned off without code changes

- WIFI_NETWORK_SETUP / SSH_UNAVAILABLE_AFTER_INSTALL: both AP paths in
  src/wifi_manager.py create an open network and nothing reads
  ap_password, so drop the "ledmatrix123" password and the ap_password
  key/advice.
- SSH_UNAVAILABLE_AFTER_INSTALL: disabling automatic AP mode does not
  need code changes -- auto_enable_ap_mode is a WiFi-tab toggle and
  POST /api/v3/wifi/ap/auto-enable; note the monitor daemon reads
  wifi_config.json at start, so restart it after changing the setting.
  Use the ledpi username and a relative install path like the other docs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(reference): add auto_update, drop drifted line numbers, fix UI and service details

- CONFIG_REFERENCE: document the top-level auto_update.enabled key (read
  by web_interface/auto_update.py and src/auto_update_setup.py); replace
  drifted file:line references with function names; the template's
  dim_schedule mode is "global".
- ADVANCED_FEATURES: core does not read a per-plugin background_service
  block (the sports plugins read their own), and priority is "higher
  number = higher priority" on FetchRequest but not used for ordering.
- WEB_INTERFACE_GUIDE: the General tab toggle is "Web Display Autostart"
  (web interface service), brightness is 1-100, and config paths are
  relative to the LEDMatrix folder, not /config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: drop references to code removed in #608

get_installed_plugin_info, WiFiManager's saved_networks and the six
always-skipping plugin test files are deleted there. NetworkManager already
remembers joined networks; LEDMatrix no longer stores WiFi passwords.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: don't link SKIN_SYSTEM.md from the core-properties page

#615 deletes SKIN_SYSTEM.md; with this link, whichever of the two merged
second would break test_doc_links. The skin/skin_options entries go when
#615 removes the keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:36:38 -04:00
ChuckandClaude Opus 5.5 84afa9d64f refactor: delete dead Python code in the core (and stop storing Wi-Fi passwords) (#608)
* refactor(plugins): remove the no-op PluginHealthMonitor

Its monitor loop did nothing (`if callbacks: pass`), register_health_check
had no callers and api_v3.health_monitor was never read by any route. The
live health data comes from PluginHealthTracker, which is untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(store): drop the never-set uninstall tombstones

Nothing in production called mark_recently_uninstalled, so the
reconciler's was_recently_uninstalled check was always False. The
persistent uninstall registry is what actually stops resurrection; the
reconciler test now exercises that gate instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(common): delete unused config/display/game helpers, utils and error_handler

Nothing in core, the web UI, scripts or the plugin monorepo imports
config_helper, display_helper, game_helper, utils or error_handler; only
their own tests did. The error_handler re-exports leave src.common's
__all__; APIHelper, TextHelper, ScrollHelper, LogoHelper and the adaptive
layout exports are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(config): drop ConfigService's unused versioning and save API

ConfigVersion, get_version/get_version_history/get_version_config,
rollback, save_config, reload, get_plugin_config and the backward-compat
load_config/get_config_path/get_secrets_path had no callers. The display
controller only uses get_config, subscribe, unsubscribe and shutdown,
plus the file watcher. Change detection now compares against the
current checksum instead of the last history entry.

The subscriber tests asserted `callback.called or True`; they now
reload the way the watcher does and assert the notification.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): drop unread plugin state history and callbacks

plugin_state.PluginStateManager kept a bounded per-plugin transition
history that only get_state_history (tests only) read; get_state_info
reports a separate lifetime count, which stays. set_error_info and
record_display had no callers, and set_state_with_error's `error`
argument only fed the history.

The web-side state_manager.PluginStateManager loses
subscribe_to_state_changes, _notify_callbacks, set_plugin_error and
get_state_version, none of which had callers; with no subscribers the
old-state copy in update_plugin_state went with them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): remove unused PluginManager methods and attribute guards

update_all_plugins was only called by a test (the display loop uses
run_scheduled_updates); get_plugin_health_metrics,
get_plugin_resource_metrics and get_plugin_state had no callers; and
plugin_modules was written but never read. plugin_directories is now
initialised in __init__, so the hasattr() guards around it go.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): remove unused executor, loader, store and package helpers

- PluginExecutor.execute_safe: no callers.
- PluginLoader._parse_semver: only its own tests; compatibility.parse_semver
  is the live copy and test_compatibility.py already covers it.
- PluginStoreManager.get_installed_plugin_info: no callers.
- PluginResourceMonitor._local: never read.
- src.plugin_system.get_store_manager and __api_version__: no importers in
  core, scripts or the plugin monorepo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(wifi): stop storing Wi-Fi passwords in wifi_config.json

WiFiManager appended every joined network's SSID and password, in
plaintext, to saved_networks in config/wifi_config.json, and nothing
(web UI, backup restore, scripts) ever read them back: NetworkManager
keeps its own credentials. The writes are gone, and loading the config
now drops any saved_networks key and rewrites the file, so passwords
already on disk are scrubbed.

Also removes _check_dnsmasq_conflict (never called) and _detect_trixie,
whose result only reached one log line, along with the
NM_CONNECTIONS_PATHS constant only it used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(display): remove unreachable and unused DisplayController code

- _follower_rebuild_scroll_image: never called.
- mode_duration (never read) and last_mode_change (write-only).
- The `chosen_cap <= 0` branch: chosen_cap is either the minimum of
  caps already filtered to > 0 or DEFAULT_DYNAMIC_DURATION_CAP (180).
- The `max_duration < min_duration` branch directly after
  `max_duration = max(min_duration, max_duration)`.
- The circuit-breaker branch's `display_result = False` and
  `manager_to_display = None`: the first is overwritten a few lines
  later, the second is already None there.
- The bool-to-bool conversion of execute_display's result, which is
  always a bool.
- The `loaded_plugins` lookup in _update_modules: PluginManager has no
  such attribute, so it always fell through to `plugins`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(vegas): remove unused config update, boundary finder and refresh

VegasModeConfig.update had no callers outside its own tests (the
coordinator rebuilds the config with from_config on a change);
geometry.find_item_boundary and StreamManager._refresh_plugin_content
had no callers at all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(run): drop the debug block that pretended to import the plugin system

In debug mode run.py put src/plugin_system itself on sys.path and printed
"Plugin system import successful" without importing anything. Nothing
imports plugin_system modules by bare name, so the path entry did
nothing either.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: delete tests that test nothing

- test/plugins/test_{basketball_scoreboard,calendar,clock_simple,
  odds_ticker,soccer_scoreboard,text_display}.py skip everywhere the named
  plugins are not installed, including CI (LEDMATRIX_PLUGINS_DIR holds only
  the fixture plugin); test_plugin_matrix.py already covers every
  discovered plugin. Their PluginTestBase and the fixtures only it used
  (plugins_dir, mock_display_manager, mock_cache_manager,
  mock_plugin_manager, base_plugin_config in test/plugins/conftest.py) go
  with them.
- test_plugin_system.py: test_discover_plugins (body was `pass`) and
  test_dependency_check (a comment), plus the test_plugin_manager fixture
  only the former requested.
- test_display_manager.py: test_draw_image asserted that an image it had
  just assigned was not None.
- test_display_controller.py: the rotation and schedule-override tests
  re-implemented the run-loop arithmetic inline and asserted on their own
  result without calling the controller.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: expect one plugin_last_update success stamp after update_all_plugins

EveryStampRecordsACompletion required at least two success-path stamps;
the second was update_all_plugins, removed as test-only. The worker and
synchronous paths share the remaining stamp in _execute_update_now, and
the check that every stamp calls _note_update_completed is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:36:26 -04:00
ChuckandClaude Opus 5.5 e1ce7189f1 fix(install): make one-shot retry() retry, and drop root grants on user files (#606)
retry() in one-shot-install.sh used `if ! "$@"; then status=$?`, where $? is
the status of the negation -- always 0. A failed command was never retried
and retry() reported success, so a failed `git clone` carried on until a
later check noticed the missing checkout. It now retries (3 attempts) and
returns the command's status. The two apt steps stay non-fatal: warning and
continuing is what they effectively did before, and making them fatal would
stop installs that work today. A clone that keeps failing stops the install,
as it already did, just sooner and with the one-shot's own error message.

Both installers granted the web user NOPASSWD root on display_controller.py,
start_display.sh and stop_display.sh. Those files are owned by the user after
Step 11's chown, so the grant let the web user rewrite them and run them as
root, and nothing ever ran them through sudo. Removed from both installers,
with a test that every project file granted as root is a root-owned
fix_perms helper.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:34:20 -04:00
ChuckandClaude Opus 5.5 342e9164b8 fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound

Each repeat of an error pattern appended every plugin in the time window to
the pattern's list again, so a plugin failing in a loop grew the display
process's memory without limit: 3,000 errors from three plugins reached 2.5
million entries. Keep the list unique.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): load a BDF font at its native size instead of PIL's default

FreeType rejects any size but a BDF strike's own, and FontManager answered
that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf
requested at 8 or 10px rendered as PIL's default font. Retry at the native
strike, as element_style already does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): plugin toggle failures no longer claim "operation in progress"

Every exception in POST /plugins/toggle was mapped to
PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation
is already in progress". Report the failure as what it is, and record the
plugin id in the operation history for form posts too (it read a `data`
variable that only the JSON path set).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): route plugin card clicks through handlePluginAction

The document-level delegation checked `typeof handlePluginAction`, which is
scoped inside the plugin-manager IIFE and so never visible to it. Every card
click took a copied fallback that stopped propagation (the grid's own
listener never ran), confirmed an uninstall twice, and sent Starlark app
uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>.
Expose the handler on window and delegate to it.

Also run every test/js/unit suite under pytest: they need only node, but CI
ran one of the eight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): apply Rotation durations, WiFi messages and Vegas settings

Three settings the web UI saves never reached the display:

- Rotation & Durations: display.display_durations was never read. Every
  plugin inherits get_display_duration() and the plugin was asked first. A
  saved value now wins. The page shows unsaved screens blank with the
  plugin's own duration as a placeholder, and saving a blank removes the
  override, so one save no longer pins every screen.
- WiFi status overlay: the controller looked for wifi_status.json one
  directory above the repo. Both sides now use
  wifi_manager.get_wifi_status_path(). The message is written by rename so
  the display never reads it half-written, and the resumed plugin redraws the
  whole panel afterwards.
- Vegas: nothing called coordinator.update_config(), so saved Vegas settings
  never reached a running scroll. They are now queued when
  display.vegas_scroll changes, and applied while Vegas is stopped too, so a
  disable then re-enable works. The follower's scroll-speed default (75) now
  matches VegasModeConfig's (50).

Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us
per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows
with each plugin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: keep affected_plugins order when serialized; guard non-Element targets

ErrorPattern.to_dict() ran the now-ordered list through set(), so
get_error_summary() listed plugins in an unstable order. The document-level
card-action listener called event.target.closest() without checking the
target is an Element.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:29:16 -04:00
claude[bot]andClaude Opus 5 0903f9055f docs(changelog): record #600 and #602 in the 3.5.0 section (#603)
* docs(changelog): record #600 in the 3.5.0 section

#600 merged into main while the release PR was open, so the 3.5.0 section
went in without it. Nothing in that PR touched the CHANGELOG, and no check
covers "everything merged since the last tag is written down", so tagging
v3.5.0 as main stands would ship the standings-endpoint fix undocumented.

The entry goes under Sports data, next to the other ESPN fetch changes, and
is written from the commit: what the old order did, why a college league's
200 defeated the 404 fallback, and what is now treated as routine.

No version change: 3.5.0 is not tagged yet, so this belongs in that section
rather than a new one. `scripts/check_release_version.py v3.5.0` still passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

* docs(changelog): record #602 in the 3.5.0 section

#602 merged into main after #601, the same way #600 merged during it, and
also touched no CHANGELOG. So the section was still a commit short of what
v3.5.0 will actually ship.

It gets its own "Installers" subsection rather than a line under "Small
fixes": a malformed drop-in in /etc/sudoers.d makes sudo refuse every command
for every user, which on a headless Pi is unrecoverable over SSH. That is not
a small fix, and someone reading the release notes to decide whether to update
should see it.

Written from the commit: what both installers did, what `visudo -c` now gates,
and the fixed /tmp path that mktemp replaced.

`scripts/check_release_version.py v3.5.0` still passes, and this branch is
rebased onto 967f3a05 so the section now covers every commit since v3.4.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-23 10:28:31 -04:00
claude[bot]andClaude 967f3a0567 fix(install): parse the sudoers rules before installing them (#602)
Both installers generated the ledmatrix_web rules and copied them straight
into /etc/sudoers.d without ever parsing them. Every rule is built from
`which` lookups, so an empty or surprising path produces a malformed
drop-in -- and a malformed file in /etc/sudoers.d makes sudo refuse every
command for every user. On a headless Pi that is unrecoverable over SSH.

first_time_install.sh now runs `visudo -c` on the generated file and, if it
does not parse, prints what visudo said and leaves the installed file
untouched rather than replacing it with a broken one. configure_web_sudo.sh
does the same before it offers the rules for confirmation.

first_time_install.sh also built the file at a fixed /tmp path as root;
mktemp now picks the name.

test/test_sudoers_is_validated.py renders the installer's own sudoers
heredoc and checks the result with visudo -- the check neither installer
had -- and asserts the install stays gated on it.


Claude-Session: https://claude.ai/code/session_01Dby94z9PV3zVM25fqGNXTt

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-21 16:08:29 -04:00
claude[bot]andClaude Opus 5 21c8a54f68 chore: prepare the 3.5.0 release (#601)
* chore: prepare the 3.5.0 release

Turns the CHANGELOG's Unreleased section into `## 3.5.0` and bumps
`src.__version__`, the value plugin `ledmatrix_min_version` floors compare
against. No behaviour change; nothing outside the CHANGELOG, `src/__init__.py`
and one docs line is touched.

The staged entries are reshaped into the `### ` subsections every released
section already uses, and the "new modules a plugin may import via `src.*`"
block moves to the top as the plugin-facing summary, the same shape as 3.4.0.
Its floor, written as "the release that ships this" while it was staged, is now
3.5.0, and `docs/SPORTS_UNIFICATION.md` says 3.5.0 for `sports_helpers.py`
instead of "(unreleased)".

Four merged changes had never been written down. They are added under the
subsection each belongs to, from the commits and their measurements:

- the idle back-off clamped to the next kickoff (#599)
- concurrent ESPN date chunks (#596)
- the three web routes that consulted plugin manifests before anything had
  discovered plugins, one of which wrote a plugin API key to config.json in
  plain text (#594)
- the cache permission fix and its systemd unit changes (#593), which get
  their own subsection

No tag and no release: `scripts/check_release_version.py v3.5.0` passes, so
tagging is a separate, deliberate step.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

* ci: let Claude Code Review run on PRs the Claude app opens

The review action refuses a workflow whose actor is a GitHub App unless the
app is named in `allowed_bots`, which this workflow never set:

  Actor is a GitHub App: claude[bot]
  Actor type: Bot
  Action failed with error: Workflow initiated by non-human actor: claude
  (type: Bot). Add bot to allowed_bots list or use '*' to allow all bots.

It aborts about two seconds in, before the diff is read, so the check is red
on every such PR and re-running cannot help: the actor does not change. Until
now no PR here had a bot author, so nothing tripped it.

`'claude'` rather than `'*'`: the action lowercases each entry and strips a
trailing `[bot]` before comparing it to the actor
(`isAllowedBot` in `src/github/validation/actor.ts`), so this admits
`claude[bot]` and no other app. `'*'` would admit any app that can trigger a
workflow here, with a prompt it controls — the action's own docs warn about
that on public repositories, and this one is public.

The write-permission check already allowed the app; `checkHumanActor` was the
only gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rqzd6Nz2bQJp5K7DD5dS4X

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-21 16:08:05 -04:00
ChuckandClaude Opus 5 cf02538d2e fix(sports): ask the endpoint the league actually publishes for standings (#600)
* fix(sports): ask the endpoint the league actually publishes for standings

ESPNDataSource.fetch_standings tried /standings first regardless of league
and fell back to /rankings only on a 404. College leagues answer /standings
with a 200 that carries no poll, so the fallback never fired and the poll
came back empty every time. Nothing failed; the rank badge simply never
appeared, and anything keyed off rankings quietly did nothing.

Endpoints are now ordered by whether the league publishes a poll, a 200
that lacks the key counts as a miss so a league answering both still ends
up with whichever one carries the poll, and only a 404 is treated as
routine -- it is how a league says it has none. A connection error, a
timeout or an unparseable body is logged as an error again.

This is the implementation the football, baseball and hockey boards already
ship; core was the last copy still on the old one. Verified against live
ESPN: mens-college-basketball returns a populated rankings key where it
previously returned nothing, and nba still resolves from /standings alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(standings): stop the endpoint handler from swallowing its own bugs

Addresses both CodeRabbit findings on #600.

The handler caught `Exception`, so an AttributeError or TypeError raised
while *inspecting* the payload was indistinguishable from an endpoint that
failed. The loop would move on and, if the other endpoint had nothing
either, return {} -- silently dropping rankings for a league that has them.
That is the precise failure this function was written to fix, so the
handler was able to reintroduce it.

Only the request is guarded now. `requests.RequestException` covers the
transport failures and `ValueError` covers a body that will not parse;
payload inspection happens after the handler, where a bug surfaces instead
of being logged as a missing poll. A non-dict payload is treated as a miss
explicitly rather than by tripping over `.get`.

Tests: the fallback paths had no coverage -- the old single-endpoint code
would have passed the suite unchanged. Added order assertions for both
league kinds, a 200-without-a-poll fall-through, 404 and non-404 recovery,
a non-object payload, and a guard proving a bug is no longer swallowed.

`test_fetch_standings_returns_empty_on_error` faked a transport failure
with a bare `Exception`, which only passed because the handler caught
everything. It now raises ConnectionError, which is what actually happens.

Verified by mutation: restoring standings-first fails 5 tests, restoring
the catch-all fails the bug-not-swallowed guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 16:07:51 -04:00
ChuckandClaude Opus 5 81e1bc596f fix(sports): stop the idle back-off sleeping through a kickoff (#599)
* fix(sports): stop the idle back-off sleeping through a kickoff

A league with no live games backs its poll off as empty checks mount,
capped by live_idle_max_interval. The escalation counts empty looks and
nothing else, so a league three hours before kickoff is indistinguishable
from one three months out of season. Both reach the ceiling -- and the
ceiling then *is* the blind spot.

Measured on two rigs on 2026-09-19: gaps of up to 928s between looks, ten
of them at or above 900s. Reproduced in the wild on 2026-09-20, where an
unpatched rig sat for fifteen minutes with eight NFL games in progress and
had not noticed any of them. That is the "it doesn't pick up new live
games until I restart it" report -- restarting being the one thing that
forces an immediate look.

The clamp costs no extra request: the live fetch already downloads the
whole day's scoreboard, upcoming games included, so the earliest start
still ahead of us falls out of the payload the manager already has.
Before a kickoff the wait is shortened so it cannot run past it; just
after one, the live cadence is held for _KICKOFF_GRACE_SECONDS, because a
provider that has not yet flipped the status would otherwise look like
another empty check and escalate the back-off again, right when the game
is starting.

The grace window needed a second pass. A soak caught it as dead code: the
just-passed kickoff was replaced by the next fixture on the card the
instant it passed, `now < start` went true again, and the back-off
returned to its ceiling. Observed live -- the rig polled at 13:00:45,
found nothing because ESPN had not flipped the status, then went quiet for
a quarter of an hour. A kickoff inside the grace window is now kept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(sports): pin absolute tolerances and correct a wrong grace expectation

pytest.approx defaults to a relative tolerance. On a unix timestamp that is
roughly 1790 seconds, so every kickoff assertion here was effectively
vacuous -- it called a kickoff half an hour away "equal". All seven now
pin abs=1.

That hid a wrong expectation. test_an_earlier_kickoff_still_wins_during_the_grace
asserted a game ten minutes out should displace one that kicked off moments
ago. It should not, and the code does not: while the grace holds, the wait
is the live cadence (30s), which is strictly tighter than clamping to the
nearer kickoff would give (~600s). Letting the candidate win would set a
ten-minute wait at the exact moment games are starting -- the dead grace
window this branch exists to fix.

The test now pins the real behaviour plus the safety property that makes it
correct, and is renamed to say what it checks.

Reported by CodeRabbit on the PR. The finding was right that code and test
disagreed; the suggested fix was the wrong way to resolve it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 18:25:33 -04:00
ChuckandClaude Opus 5 19686ab761 fix(web): plugin settings form shows schema defaults for unsaved keys (#597)
The server-rendered plugin settings partial rendered straight from the
saved config, so an option added in a plugin update (geochron 1.2.0's
show_date / show_date_line, default true) drew as an unchecked box, and
the save route's missing-checkbox handling then stored it as false.
Enum dropdowns likewise showed their first option instead of the default.

- _load_plugin_config_partial runs the stored section through
  prepare_plugin_config (as GET /plugins/config does) before masking
  secrets, so a secret's schema default is masked too.
- render_field falls back to the field's own default, covering children
  of objects that declare a default of their own (where the defaults
  extraction stops).
- The legacy-boolean parity test now compares against the config the
  plugin actually runs with (defaults included).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 11:09:25 -04:00
ChuckandClaude Opus 5 92f1960d00 perf(sports): fetch ESPN date chunks concurrently (#596)
* perf(sports): fetch ESPN date chunks concurrently

Since ESPN started rejecting `dates=YYYYMMDD-YYYYMMDD` on 2026-09-15, one
season request became a chunk per month -- and a month over the 500-event
cap becomes a request per day. A cold college-baseball season is about 130
requests, and they went out one at a time.

That is slower than the 20s budget `_update_plugins()` shares across every
plugin at startup, so scoreboards were logging `update() timed out` on
first run and being deferred to the scheduled tick with nothing on the
panel. Measured on a Pi 4 against live ESPN, March+April college baseball
(63 requests, 3101 events): 11.2s sequential, 1.6s concurrent. Over a whole
boot that moved football-scoreboard, ledmatrix-flights and birdnet-go
inside the budget -- 13 plugins deferred before, 10 after.

Chunks now go out six at a time, in two passes: months and edge days first,
then the days of any month that came back capped. Six keeps the shared
Session under requests' default pool_maxsize of 10, so no connection is
discarded. Merged events still follow `espn_date_chunks` order -- a capped
month's days are spliced back into its own slot -- so the payload does not
depend on which request won the race.

Request order is no longer significant, so the three tests that pinned it
compare the chunks as a set and keep asserting the merged event order,
which is the part callers actually see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): drop capped month payloads before fetching their days

Review of the concurrent chunk fetch found it raised the worst-case peak
memory more than the concurrency explains. The old loop discarded a month
that came back at the 500-event cap the moment it saw it; the rewrite kept
every capped month alive in `results`/`slots` until all of their day
requests had finished.

Measured on a Pi 4 fetching 20260201-20260531 college baseball (four capped
months, 5462 events), peak RSS growth over the call:

  sequential (main)               83 MB
  concurrent, months retained    121 MB  (+43)
  concurrent, one worker         108 MB  -- the retention alone was +25
  concurrent, months dropped      98-100 MB (+16)

docs/LOW_MEMORY_BOARDS.md puts a 1 GB Pi 3B+ at under 200 MB of headroom,
where running out makes the board unreachable until a power cycle, so the
difference matters. The remaining +16 MB is six responses parsing at once;
three workers saved about 6 MB more, within run-to-run noise, so the worker
count stays at six.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(sports): state what ESPN_CHUNK_WORKERS was measured to do, not more

The comment claimed the sequential fetch made scoreboards blow the 20s
startup update() timeout. A boot on this branch still deferred 12 plugins
and timed out baseball-scoreboard while its season fetches took 0.74s and
1.12s: the startup budget is spent on other per-plugin work. Say what was
measured -- 17.7s sequential, 2.6-3.3s concurrent -- and nothing else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 09:56:50 -04:00
ChuckandClaude Opus 5 116abb0daa fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service

BackgroundDataService always sent a season range first and, on a 400,
fell back to chunks without recording the rejection, so every background
season fetch spent a doomed request and live scoreboards learned nothing
from it (or it from them). The worker now consults and sets the same
6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight
to month/day chunks, and if every chunk fails the range is asked once for
a real error without re-spending the chunks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep plugin asset and action routes inside their directories

POST /plugins/assets/upload, GET /plugins/assets/list and POST
/plugins/assets/delete joined the request's plugin_id onto assets/plugins
unchecked, so '../../config' created, wrote, listed and deleted outside
it. #561 guarded only the route that serves the files. All three now go
through path_safety.resolve_under and answer 400 for anything but a
plain name, and delete only unlinks a metadata path that resolves into
that plugin's uploads directory.

PluginManager.get_plugin_directory refuses ids that are not one plain
path segment, so /plugins/action (which runs a manifest script from the
returned directory) and every other caller get the guard; the action
route also rejects such ids up front, covering its no-manager fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): report a no-op plugin update as already up to date

update_plugin() returns True both for a real update and for "nothing to
do" (a ZIP-installed monorepo plugin already at the registry version, a
bundled plugin). With no git commit to compare, POST /plugins/update
called every such success "updated successfully", so Check & Update All
counted most official plugins as updated on every run.

The route now reads what changed off the plugin itself (commit, else
manifest version, else last_updated) and returns data.update_status
(updated / up_to_date / local_only). The update-all toast is summarised
by PluginInstallManager.summarizeUpdateResults from that status, falling
back to the message for older servers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): scoreboard scroll speed no longer follows target_fps

sports_scroll computed the crisp speed ladder against the global
target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since
frame-locked presentation (#545) the helper steps a fixed number of whole
pixels per presented frame and the panel presents at its real refresh, so
the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a
50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s.

The ladder now uses the display manager's refresh_hz, then
display.hardware.limit_refresh_rate_hz, then the default. target_fps is
not consulted. Docstrings now say scroll_delay is ignored for pacing (no
behaviour change there) and describe the fixed-step model.

Tests: replace the tests that pinned target_fps as the ladder refresh and
described time-based stepping; assert speed independence from target_fps
(unit and end-to-end presented px/s against the real helper), that the
fixed per-frame step is applied, and that scroll_delay does not change
speed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): escape registry and upload values in plugin manager inline handlers

The store, saved-repository and custom-registry buttons built
onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone,
so a custom registry entry whose id contained ' closed the attribute and
added its own handler. One helper, jsStringAttr(), now HTML-escapes the
JSON literal for every one of those handlers, and the store View button
opens only http(s) repo links.

The live window.updateImageList (plugins_manager.js loads last, so its
copy wins over the file-upload widget's) wrote the uploaded file's
original name, path and ids into markup raw; they are escaped now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note plugin asset, action and inline handler guards

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): let the root pip wrapper install web_interface/requirements.txt

Update Code, the automatic update's health check and Install Base
Requirements install web_interface/requirements.txt through
safe_pip_install.sh, which only allowed the root requirements.txt. The
first commit changing that file would fail its dependency install, and
the automatic updater rolls back any update whose dependencies did not
install -- on every device, for every newer commit.

The wrapper now lists both core requirement files. Only their folders
are resolved, so a requirements.txt symlinked out of the project is
compared by its target and refused (previously the root file's own
symlink target was what got allowed). The updater's file list is a
named constant, and a test runs the real wrapper (pip stubbed) on
every file Update Code and the rollback install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): do not retry plugin requests that got an HTTP answer

PluginAPI.request wrapped everything that was not a structured error as
NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a
JSON error without error_code included. Check & Update All retries
NETWORK_ERROR, so those updates were re-sent five more times with
backoff, contrary to the #587 contract that an HTTP error response is
the server's answer.

NETWORK_ERROR now means only that fetch() rejected. Any HTTP response
without an error_code, or with a body that is not JSON, is API_ERROR
with the HTTP status attached. Tested against the shipped api_client.js.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): restart the stats window when an idle gap is dropped by size

#582 dropped an idle gap from the frame stats two ways: the reset_scroll()
sentinel, which also restarts the 5s window timer, and a size guard for
scrollers that never call reset_scroll(), which did not. On that path the
first real frame after the gap found the boundary overdue and logged a
stats line for a one-frame window. Both paths now share one seeding helper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): leave plugins alone when update_core's own rollback fails

update_core returns rollback_failed directly when a partial pull or an
update whose health check never started cannot be rolled back. run()
only held plugins back for 'verifying', so those devices still got new
plugin versions and a display restart on top of a core in an unknown
state -- the opposite of what the health-check path does, and of the
3.4.0 changelog (plugins are left alone if the rollback fails).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(api): make the REST reference match the api_v3 package

Every documented request body, query parameter and response shape was
re-checked against the handlers in web_interface/blueprints/api_v3/.
Fixes calls that failed as documented (repo_url, action_id/params,
files/image_id, font_file+font_family, ?font=, cache key,
auto_enable_ap_mode, plugin limit keys), removes the font-override
endpoints dropped in #566, corrects response shapes (plugins/config,
plugins/schema, health, metrics, operation history, github-status,
fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds
the 26 routes it omitted (backup, system auto-update/git, wifi radio,
starlark editor, MQTT bridge, status endpoints, skins).

Documents the merge semantics of partial JSON saves to /config/main and
/plugins/config and the dim-schedule POST accepting GET's days shape,
which land in the same change set. Replaces app.py line numbers and the
removed api_v3.py path with file and function names.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): remove the General-tab plugin system toggles that did nothing

plugin_system.auto_discover, auto_load_enabled and development_mode had
General-tab toggles whose help tips promised dormant plugins and verbose
logging, but nothing reads them: every enabled plugin is discovered and
loaded regardless. Remove the three toggles.

The keys stay tolerated in stored configs. The save handler now stores
a flag only when a client sends it; treating a missing key as an
unchecked box would otherwise rewrite all three to false on every
General-tab save, which still posts plugins_directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(scroll): remove dead code left by #523/#570

- Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read
  them since the numpy blend replaced the scipy path.
- Drop ScrollHelper._last_integer_position and frame_time_target, which
  were written but never read.
- Keep target_fps and set_target_fps() but document them as
  informational: nothing paces off them, yet ledmatrix-elections'
  test_scroll_pacing.py reads helper.target_fps back and third-party
  plugins may call the setter.
- Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to
  throttle", set_sub_pixel_scrolling's "default: True", and
  set_frame_based_scrolling's claim that it steps.

The plugins monorepo was grepped for every removed name; none is used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI

FONT_MANAGER.md told plugins to read display_manager.font_manager, which
does not exist, so a plugin following it failed to load with
AttributeError. The shared FontManager lives on the PluginManager and
BasePlugin._get_font_manager() returns it (with a fallback for harnesses).

Also removes the Fonts-tab override workflow and element-override panels
that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls

/plugins/store/search does not exist (404) and the list endpoint reads
query, not q. The registry guide's curl examples omitted the JSON
Content-Type, so the handlers saw an empty body and answered 400.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): use the shared core-key list in the last three private copies

StartupValidator warned "Plugin 'auto_update' is enabled but not found" on
every display start with auto-update or a dim schedule on; the reserved
plugin-id check missed auto_update, sync, location and the rest; and
ConfigManager's (uncalled) orphan cleanup would have deleted display,
schedule and auto_update. All three now read src/core_config_keys.py, which
also gains CORE_SECRETS_KEYS for the github/youtube secrets sections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): partial JSON saves to /config/main change only what they send

A JSON body with one field reset every checkbox in the sections it touched:
the MQTT bridge's brightness slider turned off disable_hardware_pulsing,
inverse_colors, show_refresh_rate and use_short_date_format, and a
timezone-only save turned off web-UI autostart and weekly auto-updates.
Missing-means-unchecked now applies only to form posts: form-encoded bodies
and the v3 forms, which mark themselves with a hidden __form_section input.

Also on the config routes:
- vegas_min/max_cycle_duration no longer match the generic *_duration rule,
  so they stop landing in display_durations and a blank one no longer
  rejects the whole Display save;
- saving from the Raw JSON editor calls start_setup_if_needed like the
  General form, so enabling auto-update there finishes its setup;
- the schedule and dim-schedule POSTs accept the per-day days.<day> shape
  their GETs return, as well as the flat form keys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): install plugin dependencies from the configured plugins directory

install_plugin_dependencies.sh scanned only plugins/, but the Plugin
Store installs into plugin_system.plugins_directory (default
plugin-repos), so the documented "Recommended" fix found 0 plugins on
every store install. It now reads plugins_directory from
config/config.json (relative to the project root or absolute, default
plugin-repos) and also scans plugins/ for dev symlinks, installing a
plugin reached through both only once.

With set -e alone, `pip ... | tee` took tee's exit status, so a failed
pip install was reported as success; set -o pipefail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: replace stale API names, line numbers and the api_v3.py path

- ADVANCED_FEATURES: StreamManager methods that exist
  (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the
  real on-demand status envelope ({status, data: {state, service}})
- app.py:199 / :144 / :607-619 line citations and
  web_interface/blueprints/api_v3.py (now a package) replaced with file and
  function names in ADVANCED_FEATURES, CONFIG_DEBUGGING,
  PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE,
  PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README
- CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use
  /config/raw/main to replace the file; describe where validation runs
- TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints
  usage)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): verify the web interface that actually ships, on port 5000

verify_installation.sh failed every healthy install: it required the
long-removed web_interface_v2.py and looked for a listener on port 5001,
while the web interface binds 5000 (web_interface/start.py). It now
checks the files ledmatrix-web.service runs (start_web_conditionally.py,
web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the
same 5001 port in its listen check, HTTP probe and printed URLs.

Port matches are anchored so :50001 no longer counts as :5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): one display-size contract: display_manager.width/height

CLAUDE.md (#580) says to read display_manager.width/height because
matrix is None when hardware init fails; the development guide, the
safety-harness doc and two DisplayManager docstrings still recommended
matrix.width/height. The bundled starlark-apps plugin read matrix.width
unguarded, so its magnify recommendation and frame scaling raised in
fallback mode (e.g. after the Pi 5 hardware refusal).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): make install_service.sh --help print usage instead of installing

install_service.sh parsed no arguments, so `sudo ./scripts/install/
install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md)
rewrote ledmatrix.service, ledmatrix-web.service and both update-verify
units and enabled/started them. It now handles -h/--help (usage, exit 0,
no changes) and rejects any other argument with exit 2 before doing
anything. Running it with no arguments, as first_time_install.sh does,
is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scroll): describe the fixed-step model and document frame_hold

Since #545 a crisp speed from scroll_config.configure() makes the helper
advance a fixed whole-pixel step per presented frame with no clock, and the
display manager's frame hold is part of the speed. The docs still described
the removed wall-clock model:

- scroll_config's module and configure() docstrings said speed is applied
  in time-based mode and that omitting the hold "falls back to fractional
  pixels"; omitting it actually runs the scroll frame_hold times too fast.
- SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both
  modes, and read a 20 ms stats median as missed refreshes although that
  is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step,
  the hold-dependent healthy median, that target_fps plays no part, and
  that a hand-added scroll_pixels_per_second loses to a schema-default pair.
- PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling)
  without frame_hold; it now documents the parameter (core 3.4.0) with a
  configure() + set_scrolling_state example.
- update_scroll_position/set_scroll_speed and set_scrolling_state
  docstrings say the same.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark target_fps legacy; describe what Vegas scroll_delay does

- General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after
  the sports_scroll fix nothing in core scrolling reads it. The field and
  its API validation stay so saved configs and plugins that read
  global_config['target_fps'] keep working. CONFIG_REFERENCE says the same.
- Vegas frame_based_scrolling/scroll_delay were described as frame-count
  stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based
  mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and
  still advances by elapsed time, so the applied speed is
  clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The
  config comments, render_pipeline comment and CONFIG_REFERENCE rows now
  say so. No behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(deps): describe how plugin dependencies are really installed

The guides said the web service runs as root, that installs pick --user
from os.geteuid(), and quoted a warning and a
PluginManager._install_plugin_dependencies() method that don't exist. The
web unit runs as the installing user; store installs go through
install_requirements_file() and sudo safe_pip_install.sh (root), with a
user-level fallback that says so, and load-time installs run in the
display service's own (root) interpreter.

Manual paths now use the configured plugins directory (plugin-repos/ by
default) instead of plugins/, which store installs no longer use, and
install_plugin_dependencies.sh is described as scanning that directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): count local changes one way for the preflight and the pull

The automatic update's preflight ignored mode-only changes and anything
whose status line contained plugins/ or plugin-repos/, then promised
"Automatic updates will not stash your changes". perform_core_update
used plain git status (modes count) and ignored only 'plugins/', then
ran 'git stash push -- :!plugins', which nothing ever pops. So an edit
to a bundled plugin under plugin-repos/, or the installer's chmods on
tracked scripts, passed the preflight and was stashed away for good.

- auto_update.local_changes() is the one predicate both use:
  core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/
  excluded by leading folder rather than substring (a core file under
  web_interface/static/v3/js/plugins/ now counts).
- Update Code's explicit stash leaves out both plugin folders; the
  pull's --autostash carries their edits and mode changes across and
  reapplies them.
- The automatic updater calls perform_core_update(stash_local_changes=
  False), which refuses instead of stashing edits that appeared after
  the preflight; update_core reports that as 'blocked'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): diagnostics follow the web autostart default and api_v3 package

#556 made a missing web_display_autostart mean "start" (only an explicit
false/off keeps the web interface down), but the diagnostics still said
otherwise: diagnose_web_ui.sh reported a missing key as "defaults to
false", diagnose_web_interface.sh said the web interface "will not start
unless this is set to true" and recommended enabling it, and
debug_web_manual.py printed False. Troubleshooting a down web UI pointed
users at a non-cause.

Both shell scripts now evaluate the setting with the launcher's own
autostart_enabled() (inline fallback if it cannot be imported) and report
on / off / not set (on) / unparseable config; debug_web_manual.py uses
the same function. They also check web_interface/blueprints/api_v3/
__init__.py: api_v3.py became a package in #553, so every healthy
checkout was reported as missing a file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(install): what install_service.sh installs; verify script port; no sudo for --help

install_service.sh installs and starts ledmatrix, ledmatrix-web and the
update-verify units, not only ledmatrix.service (systemd/README.md,
README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help'
as a harmless check; it now shows --help without sudo and warns what a
real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh
checks the web interface on port 5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note update-all, plugin system settings and script fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): size the preview after orientation and pixel mappers

display_geometry.physical_size claimed to give DisplayManager's answer but
only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured
after the library's pixel mappers, so a Rotate:90 / orientation 90 chain
previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for
128x64.

Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc
does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown
names leave it alone), and move the orientation composition here so
DisplayManager and the preview share it. The module docstring no longer
claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH.

Audit finding F18.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): refuse settings the rgbmatrix library aborts on, on every board

The library answers several settings with a NULL matrix or abort() rather
than an error, so the display service crash-looped (Restart=on-failure)
instead of reaching fallback mode: rows above 64, chain_length above 255
(uint8_t binding setter, documented as "no upper limit"), a misspelled
hardware_mapping, and parallel 2-3 on a single-output mapping, reachable
from the Display form on the default adafruit-hat(-pwm) mapping. #586 only
guarded the Pi 5 subset.

- src/matrix_support.py holds the rules for every board (Options::Validate
  ranges, binding integer types, mapping names and outputs from
  lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the
  API's numeric ranges.
- DisplayManager checks them before building options and raises
  MatrixSettingsRefused, so a hand-edited config falls back with a logged,
  reported reason. Emulator mode only warns.
- The config API refuses them with a 400 naming the setting; combinations
  are checked against stored values but reported only when the request
  sets a field involved.
- The hardware status file gains "cause" (settings/library/forced). The
  fallback log and Display banner give the Pi 5 rebuild hint only for a
  library failure instead of rebuild + gpio_slowdown advice for every
  failure; one Pi 5 slowdown recommendation (1-3, start at 1).
- The Display form offers classic/classic-pi1 and orientation 90/270 and
  renders any other stored mapping selected with a warning, so an
  unrelated save no longer rewrites them; the API accepts 90/270.

Audit findings F03, F16, F19, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(display): library limits, template defaults and Pi 5 slowdown

- rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs,
  classic/classic-pi1 mappings and orientation 90/270 documented.
- Defaults are the config.template.json values: config migration adds
  missing keys from the template, so the listed "code defaults" never
  applied.
- One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode,
  starting at 1.
- Troubleshooting describes the refused-settings fallback, and CHANGELOG
  corrects the Unreleased "no upper limit" entry.

Audit findings F19, F20, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py opens the panel with the service's options

--measure and --demo built RGBMatrixOptions from a private copy of the
display service's builder that had drifted: gpio_slowdown came from
display.hardware (default 2) instead of display.runtime (default 3), and
rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors,
pixel_mapper_config and orientation were skipped, with different defaults
(hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown
was measured -- or garbled -- in a setup the service never drives.

The option filling in DisplayManager._setup_matrix moves, unchanged, into
DisplayManager.apply_matrix_options(options, config), which _setup_matrix
calls and the script reuses (overriding only limit_refresh_rate_hz for
--measure). The script now loads the whole config rather than the hardware
block. Tests pin the script's options to the service's attribute for
attribute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py recommends keys the resolver honours

The ladder ended by telling users to set
display_options.scroll_pixels_per_second. scroll_config ranks that key
below the scroll_speed + scroll_delay pair, deliberately, and several
plugin schemas default the pair into config, so the advised key was
silently ignored (a schema-default 1/0.02 pair plus an advised 66 still
resolved to 50 px/s).

The advice is now the pair that selects the crisp speed exactly
(pixels_per_frame every frame_hold/refresh seconds), explains that the
pair outranks scroll_pixels_per_second, and gives the scoreboards'
per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the
printed pair over a schema-default pair and check it lands on the
advertised speed and hold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula

- SPORTS_UNIFICATION.md still presented honouring global target_fps as
  sports_scroll's added behaviour and its one user-visible gain; note that
  it was withdrawn because it had become a speed multiplier.
- ADVANCED_FEATURES.md gave Vegas scrolling as
  (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s
  by elapsed time, through a 0.1-5 px per scroll_delay clamp when
  frame_based_scrolling is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): scroll model fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(dev): link-github links plugins from the ledmatrix-plugins monorepo

link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git,
and those per-plugin repositories no longer exist: official plugins are
directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the
monorepo once into the dev directory, finds plugins/<name>,
plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and
links it under its manifest id. With an explicit repo URL it still links a
single-repository plugin as before.

dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a
fork), plus plugins_repo and plugins_branch; github_pattern, which was
documented but never read, is dropped and warned about. Ships
dev_plugins.json.example and git-ignores dev_plugins.json, both of which
the guide promised. Reading JSON falls back to python3 when jq is missing
(get_plugin_id silently returned nothing without jq).

update/status/list find the git checkout above a monorepo plugin
directory (its .git is not in the plugin dir), and update pulls a shared
checkout once. status no longer exits 1 when nothing is broken.

Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration,
workflow, store integration, hello-world link, submission), and the
nonexistent scripts/git-hooks/pre-push-plugin-version and
scripts/bump_plugin_version.py replaced with the real rule: bump the
manifest version and run update_registry.py. scripts/dev/README.md and
CLAUDE.md updated to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scripts): monorepo workspace layout; fix_perms and install READMEs

MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin;
setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into
plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the
workspace file opens LEDMatrix plus ../ledmatrix-plugins.

scripts/fix_perms/README.md listed cache directories
fix_cache_permissions.sh never touches and a 'ledmatrix' service user
that doesn't exist (also in scripts/install/README.md); adds
safe_pip_install.sh. install/README: install_service.sh installs the web
and update-verify units too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): keep the rollback's pip retries inside the unit time limit

The health check reinstalled the previous requirements by trying the
next bash path after any failure, including a 600 s pip timeout. Two
files, two paths: up to 40 minutes of pip alone, while systemd stops
ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the
rollback half-way and leaving the update 'verifying' until the web UI
calls it lost.

- Like permission_utils.install_requirements_file, only a sudo refusal
  moves on to the next bash; a pip that ran and failed or timed out is
  not repeated. The refusal wording is one list
  (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only
  verifier and pinned equal by a test.
- All reinstalls in one rollback share a 600 s budget.
- WORST_CASE_SECONDS adds up every timeout on the longest path (27.5
  min); a test holds it under the unit's TimeoutStartSec and that under
  the web UI's VERIFY_LOST_SECONDS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools

Plugin config was prepared differently depending on how it arrived:

- JSON POST /plugins/config built a partial body on schema defaults, so
  {"enabled": true} reset every other setting of the plugin. It now merges
  onto the stored section first, as the form path already did.
- Legacy-boolean normalization (#588) ran only at load: GET /plugins/config
  returned the raw boolean, posting it back failed validation, and hot
  reload handed plugins the raw section (a legacy dynamic_duration: true
  came back as a boolean). schema_manager.prepare_plugin_config (normalize,
  then defaults) is now used by PluginManager.load_plugin, both save paths,
  GET, the save notifications and DisplayController's hot-reload callback.
- The JSON save's filter kept only enabled/display_duration/live_priority
  and dropped a submitted skin, skin_options or vegas_* tuning key. There
  is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES,
  used by validation and by the save filter; PluginManager's
  CORE_OWNED_CONFIG_KEYS is its vegas subset.
- Plugin sections posted to /config/main were stored verbatim, including
  values /plugins/config rejects. They now go through the same preparation
  (_prepare_plugin_config_for_save, extracted from save_plugin_config), and
  a failing section rejects the whole save before anything is written.
- dev_server read only top-level defaults and let a schema enabled:false
  win; build_full_config shallow-merged overrides, dropping sibling
  defaults; the harness extracted defaults differently from the device.
  loading.build_config now uses the device's extraction and preparation,
  and dev_server, check_plugin, render_plugin and the harness all use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt-bridge): brightness changes apply live and touch nothing else

The display service's hot reload applies a saved brightness within a few
seconds, and /config/main no longer resets other display settings on a
brightness-only JSON body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): automatic update hardening

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI

It described web_interface_v2.py and index_v2.html (both gone), client-side
form generation, one POST per field with {key, value}, and 'no nested
objects'. The v3 UI renders plugin forms server-side from the schema
(pages_v3 partial + plugin_config.html macros, nested sections and
x-widgets), posts the whole form once, and save_plugin_config() merges onto
the stored section, validates, splits x-secret fields and notifies the
plugin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt): brightness saves apply via hot reload and leave other settings alone

The bridge README said brightness is applied on the display's next
restart; the display controller's config hot reload applies it within
seconds. It also now states that the bridge's partial JSON save changes
only brightness (the /config/main merge fix in this change set).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): don't log pip's output from the health check's reinstall

pip can echo a private index URL with embedded credentials;
permission_utils redacts it, the stdlib-only verifier cannot, so it
logs the exit code only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark the plugin_system toggles as unused legacy keys

auto_discover, auto_load_enabled and development_mode are read by
nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE
said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and
the REST reference listed them as live settings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): docs and developer tools group

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): legacy plugin-system toggles no longer count as a General save

auto_discover, auto_load_enabled and development_mode have left the General
form, so a post carrying only one of them is not a general-settings save and
must not treat web_display_autostart and auto_update as unchecked. The
plugin_system block itself is left as on main for the branch that reworks it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): config-save and plugin-config preparation fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(claude): re-check matrix_support.py rules when the library submodule is bumped

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address Codacy findings on the core audit PR

- plugin_manager.prepare_plugin_config: when the fallback legacy-boolean
  pass also fails, log a warning instead of a bare except/pass.
- api_client.js: request() refuses any endpoint that is not a plain path
  under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace,
  control characters) with INVALID_ENDPOINT before calling fetch(), and
  plugin ids are URL-encoded wherever they are put into a URL (also in the
  app-shell batch load).
- test_update_all.js: pins both against the shipped client.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): check endpoint control characters without a control-character regex

Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags
the \x00-\x1f range in checkEndpoint's regex. Test the char codes
instead; the endpoints refused are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(auto-update): make the seed script executable on disk, not only in the index

On Linux Repo.publish() commits with -a, which recorded scripts/run.sh
as 100644 upstream because the seed file was never chmod +x. The pull
then brought in the same mode the installer chmod had made locally, so
installer_chmod saw no mode change left to check. The updater was fine:
with the upstream commit at 100755 the --autostash carries the device's
chmod across. Verified under Linux (WSL, git 2.43): the old helper fails
exactly as CI did, the fixed one passes all 63 tests in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 16:37:29 -04:00
ChuckandClaude Opus 5 7e5967e160 fix(web): find installed plugins before anything has discovered them (#594)
The web process discovers plugins lazily: plugin_manifests is empty until
some endpoint calls discover_plugins(). Three routes consulted it without
discovering, so they misbehaved for as long as nothing else had run --
which, after every ledmatrix-web restart, is until someone opens the
dashboard:

- POST /display/on-demand/start answered 404 "Plugin <id> not found"
  (or "Mode <mode> not found"). Measured on a rig: 404 for over three
  minutes after a web restart, until GET /plugins/installed ran. The
  browser UI loads the plugin list first, so API-only callers (the Home
  Assistant MQTT bridge, scripts) are the ones who hit it.
- POST /plugins/toggle answered 404 "Plugin not found".
- POST /config/main did not recognise a plugin section, so it skipped
  secret separation and merged the section as-is: the plugin's API key
  was written to config.json in plain text instead of config_secrets.json.

Add _discovered_plugin_manifests(), which discovers when nothing has been
yet, and rescans once when a specific plugin id (or, for on-demand by
mode, a mode) is not found, so a plugin installed since the last scan is
found too. _installed_plugin_ids() now uses it.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 12:38:55 -04:00
ChuckandClaude Opus 5 f475038895 fix(cache): web UI can read what the display service caches again (#593)
* fix(cache): web UI can read what the display service caches again

ledmatrix-web.service carried CacheDirectory=ledmatrix. With User= set to
the installing user, systemd re-owns /var/cache/ledmatrix and everything
in it to that user and its primary group whenever the directory's owner
differs -- for a directory root created, on the first start. That erased
the root:ledmatrix setgid layout the installers set up, so every file the
display service (root) wrote afterwards was root:root 0660 and unreadable
by the web interface:

  WARNING - Permission denied loading cache for display_current_state ...

Since #547 install_service.sh renders the web unit from the template, so
every fresh install hit this. Measured on one rig: 392 unreadable files,
and the web UI's display status, on-demand state and plugin health empty.

Existing installs only receive `git pull`, never a reinstalled unit, so
the fix for them is in the code the root display service runs:

- DiskCache.set gives each file the directory's group (when the directory
  is group-writable) and 0660 on the open descriptor before the rename,
  independent of setgid. This also closes a window where a fresh file was
  visible as mkstemp's 0600.
- DiskCache.share_existing_files repairs files an older version left
  behind, once per process from the cleanup thread. It works through
  O_NOFOLLOW descriptors and skips hard links and other users' files: the
  directory is writable by the web user, and root must not be steered
  into changing a file outside it.

For new installs, the web unit drops CacheDirectory=/CacheDirectoryMode=,
and install_web_service.sh stops replacing an existing directory's
ledmatrix group with the user's group.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): on-demand and current-display status read the display's latest state

Found testing the cache-permission fix on a rig: once the web interface
could read display_on_demand_state at all, /display/on-demand/status kept
answering "active" for over 100 seconds while the file on disk said
"idle". Both status routes read the display service's keys through the
web process's memory tier, which serves the first copy it read for the
full max_age (120s). Read them with memory_ttl=0, as every other
cross-process reader (plugin health/metrics, the on-demand mailbox)
already does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): re-group the cache dir whenever the web user is outside its group

install_web_service.sh replaced an existing cache directory's group only
when it was root's. A directory in any other group the web user is not a
member of -- root:ledmatrix, for a user who is not in ledmatrix -- was left
alone, and every file root wrote there stayed unreadable to the web
interface. Replace the group whenever the installing user is not in it.

A directory whose group the user is already in (ledmatrix, or the user's
own group where CacheDirectory= left it) is still left as it is: re-grouping
a working directory strands the files already in it on the old group.

When the group does change and root-owned JSON files carrying the old group
are present, try-restart ledmatrix.service so DiskCache.share_existing_files
re-groups them through its symlink- and hard-link-safe path, rather than a
recursive chgrp.

Verified under WSL's systemd for seven directory states (user group,
ledmatrix member, ledmatrix non-member with and without root files,
root:root, missing, unnamed gid); the previous version left the non-member
case unchanged.

Addresses CodeRabbit review on #593.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 12:38:39 -04:00
ChuckandClaude Opus 5 1d51efe4c7 fix(backup): restore over existing files on hosts without os.chown (#592)
_copy_file() replaces each restored file and then carries the previous
owner across with os.chown. On Windows os.chown does not exist and
st_uid/st_gid are 0 rather than absent, so the ownership branch always
ran and raised AttributeError. That is not an OSError, so it escaped
every per-section handler in restore_backup(): a restore over any
existing config aborted at config.json and restored nothing.

Skip the ownership step where os.chown is missing, as
auto_update_setup.py already does. No change on POSIX.

test_restore_over_a_file_the_user_cannot_write simulates root-owned
files with chmod 0o444; on Windows that sets the read-only attribute,
which blocks any rename over the file, so it is skipped there. The
modes the app writes (0o644/0o640/0o600) replace fine on Windows.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 09:26:21 -04:00
ChuckandClaude Opus 5 7ae614aa35 fix(sports): recover from ESPN rejecting scoreboard date ranges (#591)
* fix(sports): recover from ESPN rejecting scoreboard date ranges

Since 2026-09-15 ESPN's site API answers `dates=YYYYMMDD-YYYYMMDD` with
400 "Failed to get events endpoint." for every sport. Single days, months
(`YYYYMM`) and season years still work. Every season and weeks-window fetch
in core failed, including the background service the scoreboards submit
their season schedules to.

src/common/espn_dates.py re-asks a rejected range as whole-month chunks
plus the leftover edge days, which tile the window exactly (a season is
8 requests, not 213). A month that comes back with exactly 500 events is
truncated (college baseball's March) and is re-asked day by day.

It also clamps `limit` to 500: above that ESPN truncates silently, e.g.
college football returns 25 of 68 games for one Saturday at limit=1000.

BackgroundDataService recovers rejected ranges on the worker thread and
advertises `handles_espn_date_ranges` so plugins can tell whether to hand
it a range. SportsCore, sports_shared, ESPNDataSource and APIHelper route
through the helper or the clamped limit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): ESPN date-range fallback and limit clamp

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): stop re-sending ESPN date ranges once one is rejected

Live scoreboards refresh every 30 seconds, and each refresh sent the range
first, got the 400, then fetched the chunks: three requests where one used
to do. After a rejection, ranges now go straight to chunks for six hours,
then the range is tried again so the workaround retires itself if ESPN
reverts. A single-day 400 does not set the memo, and when every chunk fails
the range request supplies the error without the chunks being fetched a
second time. Per-fetch chunk logging drops to debug; the rejection itself
stays a warning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): clamp limit only on ESPN scoreboard submissions

The background service is generic, and limit above 500 only truncates
scoreboards. /teams needs limit=1000 (college football has 762 teams and
limit=500 returns 500), so a teams submission must keep its limit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 17:12:28 -04:00
ChuckandClaude Opus 5 2082665252 fix(config): load config_secrets.json on hosts without os.geteuid (#590)
ensure_shared_group_ownership() - the chgrp self-heal ConfigManager runs
before reading config_secrets.json (#416) - looked up os.geteuid
unguarded. That name does not exist on Windows, and the AttributeError
is not an OSError, so it escaped the helper's best-effort handling and
every except clause in load_config(). Any Windows checkout with a
config/config_secrets.json got a ConfigError from every config load and
could not import web_interface.app.

That is what made test_update_all_plugins.py error at setup: its client
fixture imports web_interface.app. It was not state leaked between test
files - the trigger is whether the checkout has a secrets file.

Return early when os.geteuid or os.chown is missing. No change on POSIX.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 17:12:18 -04:00
ChuckandClaude Opus 5 f9b3d6ae52 fix(web): accept every panel size and row address type the rgbmatrix library does (#586)
* fix(web): accept every panel size and row address type the rgbmatrix library does

The Display form capped columns at 128 and chain length at 24, and its
submit handler (fixInvalidNumberInputs) rewrote anything larger to the cap,
so wide panels and long chains silently saved as the wrong size. The config
API checked none of the hardware numbers, so values the library rejects (odd
rows, parallel 4, PWM dither bits 3) saved and the matrix then refused to
start.

- Form limits now match the pinned library: rows even 8-64, cols >= 16 and
  chain_length >= 1 with no upper bound, parallel 1-3, PWM dither bits 0-2,
  PWM LSB nanoseconds 50-3000.
- save_main_config rejects out-of-range rows, cols, chain_length, parallel,
  brightness, scan_mode, pwm_bits, pwm_dither_bits, pwm_lsb_nanoseconds and
  gpio_slowdown with a 400.
- A stored gpio_slowdown or pwm_dither_bits of 0 renders as 0 instead of the
  default, so saving the tab no longer overwrites it.
- Row Address Type offers 5 (SM5368 / B707 row shift register). Verified on a
  Waveshare 96x48 V2 (24S-A1) on a Pi 4 with the Adafruit Triple LED Matrix
  Bonnet: rows 48, cols 96, row address type 5, BGR, GPIO slowdown 8.
- Help text and docs: FM6124-family panels use Panel Type Standard; on a Pi 5
  the library supports only row address types 0 and 2.

No change to the rpi-rgb-led-matrix submodule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): drop the rows cap and document every display setting accurately

Rows: no upper limit in the form or the API. Still even and at least 8. The
current rgbmatrix library rejects more than 64 per panel, so a larger value
saves but the matrix won't start; the help tip, README, config reference and
troubleshooting section all say so, and nothing here needs changing if the
library lifts the limit.

limit_refresh_rate_hz: the form accepts 0 (the library's "no cap"), a stored
0 no longer renders and re-saves as 120, and the API rejects negatives.

pwm_dither_bits stays 0-2: the library rejects 3 and 4, so the old form's
0-4 only ever let users save a config the display couldn't start with.

Docs and help tips, checked against the pinned library and its README:
- panel_type and rp1_rio get README entries
- show_refresh_rate prints to stdout; it never drew on the panel
- dither bits raise the refresh rate; the tip said they lowered it
- scan_mode is about interlacing at low refresh, not wrong colours
- disable_hardware_pulsing: hardware pulsing needs OE on GPIO 18 and the
  onboard sound driver off; software timing makes rows flash brighter
- gpio_slowdown guidance agrees between the README and the UI
- all 22 multiplexing values listed; every numeric setting states its range
- troubleshooting for a blank panel after a settings change, jumping rows
  and brightness flashes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): reject true and 5.5 for row_address_type and multiplexing

Both still went straight through int(), so a JSON true saved as 1 and 5.5
as 5. They now use the shared hardware range check like the other panel
fields. Review feedback on #586.

Also: the RP1 Backend tooltip said it is ignored on Pi 3/4 (it is ignored
on every model but the Pi 5), and the README gave the dynamic-duration
default cap as 90s; the code default is 180s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: refuse matrix settings a Raspberry Pi 5 can't drive

On a Pi 5 the pinned rgbmatrix library drives the panel through the RP1
chip, and that path supports only row address types 0 and 2, parallel 1-3
and the regular / regular-pi1 / classic / adafruit-hat(-pwm) mappings
(Rp1PioConfigSupported in lib/rp1/rp1_pio_backend.cc). For anything else
CreateFromOptions returns NULL; the Python binding doesn't check, so the
display process crashed on its first call into the matrix and systemd
restarted it into the same crash every 10 seconds.

- src/pi5_matrix_support.py: the rule and Pi 5 detection, matching the
  library's /proc/device-tree/model check
- DisplayManager raises before creating the matrix, so it is a logged init
  failure (reported by /api/v3/hardware/status) and fallback mode
- the config API rejects those settings on a Pi 5 when a request sets
  row_address_type, parallel or hardware_mapping
- the Display form offers only row address types 0 and 2 on a Pi 5, and
  warns when a stored value can't be used
- CLAUDE.md: re-check the rule whenever the submodule is bumped

Review feedback on #586.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 19:17:48 -04:00
ChuckandClaude Opus 5 9616a5a054 fix(web): auto_update and other core settings are not orphaned plugins (#589)
v3.4.0 shows "Plugin Config Warning - In config but not installed:
auto_update. Reinstall via the Plugin Store, or remove these entries from
config.json." auto_update is the core weekly-update setting from #581.
Reconciliation treated every top-level dict not in its private
_SYSTEM_CONFIG_KEYS list as a plugin id, and #581 could not know to extend
that list.

- Move core top-level keys into src/core_config_keys.py (CORE_CONFIG_KEYS)
  and use it in reconciliation. Tests fail if a config.template.json key or
  a key written by the general-settings save is missing from it.
- A secrets-file key only counts as a non-plugin when no installed plugin
  has that id. Plugin secrets are namespaced by id, so installed plugins
  with secrets were reported as missing from config on every run.
- still_unresolved() drops "not on disk" findings whose id is no longer a
  plugin entry in config, so a stored verdict clears without a restart.
- A plugin whose id is a core key is skipped with a warning, and the fix
  never writes a plugin stub over or in place of a core setting.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:19:39 -04:00
ChuckandClaude Opus 5 c200b5837d fix(plugins): normalize legacy boolean settings before schema validation (#588)
The news plugin's schema turned global.dynamic_duration from a boolean
into an {enabled, min_duration_seconds, ...} object. Installs that have
not saved the news settings since still hold `true`, so every start
logged "Plugin news config does not match its schema (loading anyway):
Field 'global.dynamic_duration': Expected type object, got bool" and
flagged news degraded.

The settings form already reads such a boolean as {"enabled": <bool>}
(render_nested_section in plugin_config.html) and the next save writes
the object. The loader did not. It now applies the same rule before
merging schema defaults and validating, so the defaults fill in the rest
of the object and the plugin receives it in the new shape.

The rule lives in schema_manager.legacy_bool_as_object /
normalize_legacy_booleans. It applies at any depth of nested objects
but not inside arrays, matching the form, and only to a real bool under
an object-typed property with an `enabled` child. Every other mismatch
still warns. A parity test renders the template macro against the helper
so the two cannot drift.

Nothing is written to config.json at load: the normalization is in
memory, and the next save of the plugin's settings persists the object.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:09:23 -04:00
ChuckandClaude Opus 5 9f2743471c fix(web): update-all skips Starlark apps and no longer misses plugins (#587)
* fix(web): update-all skips Starlark apps and no longer misses plugins

Check & Update All posted every entry from /plugins/installed to
POST /plugins/update, including the virtual starlark:<app_id> entries
that list installed Starlark apps. The store manager cannot find those,
so each answered 500 "plugin not found". Update-all now sends only
plugin ids (install_manager.js, and the older app-shell.js copy), and the
route answers a starlark: id with a 400 saying it is a Starlark app.

A request that got no HTTP answer was recorded as failed and never sent
again. On a device, a web-service restart mid-run killed the in-flight
request and refused the next one, stock-news, which was left on 2.6.2
with 2.8.0 available. Such requests are now re-sent with backoff
(about 30s) before being reported as failed. HTTP error answers are not
retried.

Tests: test/js/unit/test_update_all.js (run from pytest via
test/web_interface/test_update_all_plugins.py so CI covers it) and the
route contract for starlark: ids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): walk update-all retry delays without indexed lookup

Codacy's ESLint security/detect-object-injection rule flagged
retryDelays[attempt] as a High issue. The index was a bounded loop
counter over a fixed array, but shifting a per-plugin copy of the
schedule gives the same backoff without the pattern. No behaviour
change: test/js/unit/test_update_all.js (21) and
test/web_interface/test_update_all_plugins.py (7) pass unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:09:11 -04:00
ChuckandClaude Opus 5 fddb0e06db feat(web): honour x-display: hidden in plugin settings (#585)
* feat(web): honour x-display: hidden in plugin settings

Plugins keep deprecated and internal keys declared so stored configs keep
validating (weather api_key/radar_zoom, countdown's auto-generated row id),
but the settings form drew them as live controls.

A property marked "x-display": "hidden" -- or an object whose children are
all hidden -- now gets no control at any depth: top level, nested sections,
Advanced Settings (not counted either), array-table columns and the row
editor. A hidden top-level key is not reported in __rendered_section.

Saving never changes a hidden value. Plain and nested fields aren't posted,
so the save's deep merge keeps them; _set_missing_booleans_to_false skips
hidden booleans at every depth. A posted array row replaces the stored item,
so hidden row properties are carried as JSON-encoded hidden inputs and
decoded exactly on save (an id "1" stays a string). New rows get none.
JSON API saves are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): read hidden row keys without dynamic property access

Build the set of x-display: hidden item properties once and look values up
through Object.entries, instead of indexing objects by a variable key on
the lines this branch added (Codacy: object injection sink, 6 warnings).
Behaviour is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 15:47:06 -04:00
ChuckandClaude Opus 5 8360220809 feat(common): sports_helpers — the helpers all nine scoreboards carry identical copies of (#583)
* feat(common): sports_helpers, the helpers all nine scoreboards copy verbatim

Add src/common/sports_helpers.py: the helpers the scoreboard plugins'
sports.py carry byte-identical copies of (docstring-stripped AST, checked at
ledmatrix-plugins f09bff2), so a later plugins PR can delete its copies once
it floors on the core release that ships this.

- Free functions: clamp_window, clamp_seconds, logo_needs_refresh (lazy
  src.logo_downloader import, as in the plugins), spread_weighted_order,
  MIN_WINDOW_DAYS / MAX_WINDOW_DAYS. All nine plugins.
- SportsHelpersMixin (no __init__, stateless): _mode_customization,
  _setting_int, _reset_dwell_on_reentry, _next_switch_index,
  _spread_weighted_order (all nine), _odds_color and
  _upcoming_date_and_time_text (all but ufc), plus the _favorite_key seam
  from base_classes core.py for later phases.

A new module rather than more methods on sports_shared: a plugin that
deletes a copy and relies on an existing module having grown the method
fails at runtime with AttributeError on an older core, which neither the
loader nor check_min_core_version.py can see; a missing module fails at load.

Tests: behaviour for every helper, a derived host contract, and a parity
test that AST-compares every body against every plugin copy when
LEDMATRIX_PLUGINS points at a checkout (skipped otherwise).
test_common_is_hardware_free.py imports src.common and every sports_* module
with rgbmatrix blocked and scans src/common for module-level imports of
src.base_classes, src.display_manager and src.plugin_system (no existing
violations).

Nothing in core imports the new module; no behaviour change. CHANGELOG
Unreleased entry and a converging note in docs/SPORTS_UNIFICATION.md.
__version__ is not bumped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(common): address review on sports_helpers and the hardware-free test

- SportsHelpersMixin docstring and CHANGELOG: constructor-free, but it keeps
  lazy state on its host (_reset_dwell_on_reentry, _next_switch_index).
- test_common_is_hardware_free: the runtime check now filters every
  FORBIDDEN package, src.plugin_system included; the AST scan resolves
  relative imports against src.common, so `from .. import plugin_system`
  and `from ..plugin_system import x` are caught. Guard tests for both.
- Parity skip reason names the CI guard that runs the same comparison:
  ledmatrix-plugins scripts/check_sports_helpers_parity.py (#495).
- _odds_color: line-level pylint disable for a not-callable false positive
  (getter is None-checked); the AST is unchanged, parity still passes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 15:32:21 -04:00
ChuckandClaude Opus 5 9e3f184d81 docs(changelog): complete 3.4.0 with weekly auto-updates and scroll timing fix (#584)
* docs(changelog): complete 3.4.0 with weekly auto-updates and scroll timing fix

#581 (weekly automatic updates) and #582 (scroll frame-stats idle gap)
merged after the 3.4.0 section was written in #580. v3.4.0 will be tagged
on main including both, so they belong in 3.4.0.

Also record src.common.font_layout (#539, #565), a src.* module plugins
may import that shipped in 3.4.0 but was never listed, and mark
display_geometry and auto_update_setup as core-internal.

Correct the 3.3.0 historical note: remote tags v3.3.0 (bc2dbf38) and
v3.3.1 (32d637a4) both report "3.3.0" and both ship sports_shared.py.
The "3.2.0" claim came from a stale local tag.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): cover the whole 3.4.0 release since v3.3.1

The 3.4.0 section listed plugin-facing API and per-element customization
but not the rest of what merged since v3.3.1. Group it under subheadings:
Install and updates, Scrolling, Plugins, Web interface, Tools and
security, Fixes, and put the existing customization block under its own
heading.

Omitted on purpose: #569 (fixes a regression and an editor race in the
unreleased per-element framework), #570 (no runtime change), and
test-only, refactor and dev-tooling PRs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 15:24:38 -04:00
ChuckandClaude Opus 5 869e36fb2f feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback

A General-tab toggle (off by default) checks for and installs LEDMatrix and
plugin updates once a week, overnight in the configured timezone.

- Pre-update checks skip (and report) instead of forcing: local edits or
  commits, merge/live rebase, no upstream, low disk, missing health check, or
  a version that was already rolled back. An abandoned rebase (HEAD back on a
  branch) is cleared, since it would otherwise block every pull.
- The pull reuses the Update Code path (now perform_core_update(), which
  reports dependency install failures as data).
- ledmatrix-update-verify.service, started via a .path unit from a request
  file, restarts the services from its own cgroup, requires them to come up
  and stay up, and otherwise resets to the previous commit and reinstalls the
  previous requirements. It runs a copy of the checker taken before the pull.
- No SSH needed: switching the toggle on restarts the display service, which
  (as root) installs the two units from the repo templates for the web user.
  first_time_install.sh installs them too and takes --enable-auto-update /
  LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh).
- Plugins update after the code passes its check; failures, blocks and
  rollbacks raise an Overview banner and show under the toggle.

Tested end to end on a Pi: web-UI setup, a good update, a broken web service
and a broken display (both rolled back), a blocked local edit, and an
abandoned rebase found on the device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(auto-update): address static-analysis findings

- Replace the subprocess.CompletedProcess the verifier fabricated for a
  command that could not start with a plain namedtuple; nothing is executed
  there, but the scanner flags any CompletedProcess built from variables.
- Mark the subprocess imports with the repo's standard B404 annotation (all
  calls are list-form argv, no shell).
- Mark the rollback-failed message as not SQL (B608 matched its wording).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): CI failures on Linux

- Keep the setup result when chown fails. CI runs as a non-root user, where
  chown to the web user raises; that discarded the result file, so the
  General tab would never learn whether setup worked. Regression test added.
- Register the two new /api/v3/system/auto-update routes in the URL map
  snapshot.
- Use utility classes app.css defines (space-y-1, hover:text-red-600).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): address review feedback

- Health check: a failed restart command no longer lets the check run
  against the still-running old process; it counts as a failure (and after a
  rollback, as a failed rollback). An unreadable restart count is never
  treated as stable, since a crash loop looks healthy between attempts.
- Installer writes the auto_update setting to a temp file and swaps it in,
  keeping mode and owner, so a running config watcher never reads a
  truncated config.json.
- Verify unit quotes its command-line paths (install folders with spaces);
  setup refuses folder names systemd would reinterpret (%, quotes,
  backslashes, control characters) and says so on the General tab.
- The auto-update status route no longer returns exception text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): keep error detail in the status route's 500

test_web_error_detail requires every 5xx handler to log the traceback and
return describe_exception(e), which redacts credentials, so failures are
diagnosable from the web UI. Dropping it for CodeQL broke that policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): dismiss route rejects non-object JSON with 400

A JSON array or scalar body made `.get('alert_id')` raise, returning 500.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): let the app-wide handler answer status-route errors

CodeQL (py/stack-trace-exposure, #709) flagged the route's own except,
which returned describe_exception(e). web_interface/app.py's error handler
already logs the traceback and returns the same redacted detail for any
unhandled exception, so the local copy is removed: same response, no new
exception-to-response flow, and test_web_error_detail's policy still holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:58:57 -04:00
ChuckandClaude Opus 5 d01da3bd9f fix(scroll): stop timing the idle gap between scrolls as a frame (#582)
ScrollHelper.last_frame_time was set once in __init__ and thereafter only
at the end of log_frame_rate(). Nothing re-armed it when a scroll began, so
the first frame of every scroll was timed against the last frame of the
*previous* one and the whole idle period between them was recorded as a
single frame.

Measured over 3 hours on a 256x64 Pi 4, that produced 31 windows reading

    Scroll frame stats - 0.0 fps over 1 frames | median 136776.02ms
    p95 136776.02ms max 136776.02ms min 136776.02ms | stalls 0 (0.0%)

and -- worse, because it is not obviously wrong -- put the same gap in the
max field of otherwise healthy windows, where the worst values were 537s
and 604s. It also counted as one stall per scroll start: at ~500 frames to
a window that is ~0.2%, against measured stall rates of 0.07-0.16%. The
stall rate is the number used to judge whether a scroll change worked, and
it was the same order of magnitude as its own artefact.

The first frame of a scroll has no predecessor, so it has no frame time.
last_frame_time is now None until one is rendered, and reset_scroll() puts
it back -- the same treatment last_update_time already gets three lines
above, for the same reason. reset_scroll() alone is not enough, because the
scrollers actually emitting these lines never call it, so a sample at or
past the 5s log interval is dropped as well: nothing that renders a scroll
takes that long over one frame. Seeding also restarts the window timer, or
the boundary is already overdue when the second frame arrives and every
scroll opens by reporting a window of exactly one frame. A window whose
samples were all dropped now logs nothing rather than reporting the gap.

docs/SCROLL_PERFORMANCE.md documented the diagnostic in terms of a
"Frame time: N ms" line that 6031e705 replaced with the aggregate, so its
grep matched nothing on any rig. The section now describes the line that is
actually emitted, reads duplicate frames off skips and a below-median
result rather than a 2ms mode, and adds a command that ranks every scroller
by p95 -- verified against 3 hours of journal, where it reproduces
src.base_odds_manager at p95 44.08ms against 10.19ms for the two scrollers
already on src/common/scroll_config.py.

requirements.txt still offered scipy for the sub-pixel interpolation path
deleted in #570. Installing it has no effect; the entry says so.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 18:44:52 -04:00
ChuckandClaude Opus 5 814c21de1c chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0 (#580)
* chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0

Skins: no current scoreboard plugin builds on src.base_classes, so the only
skin hook (SportsCore._render_game) never runs. The plugin schema endpoint no
longer injects the Visual Skin dropdown, the store hides and refuses
"type": "skin" registry entries, and GET /api/v3/skins reports
supported: false with a message. Stored skin config still loads and saves.
src/skin_system/ and its tests are unchanged apart from the support flag.

Docs: check_plugin.py/render_plugin.py examples use --plugin; document
BasePlugin.get_update_interval() and its interaction with the manifest
update_interval; CLAUDE.md drops the stale template line number and
recommends display_manager.width/height.

Preview size: new src/display_geometry.py holds the size computation and
defaults DisplayManager uses (double-sided applied, chain_length default 2).
The web preview, /display/current, Starlark magnify default, sync handshake
and two dev scripts use it.

Release: __version__ 3.4.0, CHANGELOG 3.4.0 section plus a 3.3.0 tag note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address CodeRabbit review on #580

- Preview fallbacks (SSE stream and /display/current) use logical_size({})
  (128x32, the shared default) instead of a hard-coded 128x64.
- display_geometry treats a non-mapping display/hardware block as missing,
  so a malformed config.json falls back to defaults instead of raising
  AttributeError (which turned the Starlark render into an HTTP 500).
- Docs: the static update interval falls back manifest -> plugin config
  -> 60s, in both the API reference and the architecture spec.

Not taken: validating double_sided copies against chain_length/parallel.
An orientation Rotate: or U-mapper pixel mapper decides which axis panels
lie on, so the counts would reject working setups (the existing
vertical-split test is one).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display_geometry): a non-finite hardware size raises ValueError, not OverflowError

CodeRabbit flagged the Starlark magnify default in
_standalone_render_starlark_app for truthy non-mapping display values. That
case was already handled by a9e1bd0b (_display/_hardware treat a non-mapping
block as missing, covered by test_non_mapping_display_config_uses_the_defaults),
and the magnify it produces from the 128x32 defaults is the same as from 64x32.

Checking the same path found one input that still escaped: Python's JSON
parser accepts Infinity, and int(inf) raises OverflowError, which neither the
Starlark path (TypeError, ValueError) nor the preview stream in app.py caught,
so a hand-edited "rows": Infinity returned HTTP 500. physical_size now raises
ValueError for it, matching its documented contract, so every caller's
existing fallback applies. DisplayManager already caught Exception.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:36:58 -04:00
ChuckandClaude Opus 5 914bf2002f fix(install): grant and harden safe_pip_install.sh in first_time_install.sh (#579)
first_time_install.sh granted the web user safe_plugin_rm.sh but not
safe_pip_install.sh, unlike scripts/install/configure_web_sudo.sh. On devices
set up only by the first-time installer, install_requirements_file could not
use the root wrapper and fell back to a user-level install that root-run
ledmatrix.service may not see.

Also harden both sudo-granted helpers to root:root 755. first_time_install.sh
never did this, and Step 11's project-wide chown to the user would undo it if
placed in Step 10, so it runs at the end of Step 11.1.

Add a test that parses the ledmatrix_web sudoers rules from both installers
and asserts they grant the same commands, and that every granted helper is
hardened (after the chown, in first_time_install.sh).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 17:45:03 -04:00
ChuckandClaude Opus 5 47afaaac2b fix(install): build rgbmatrix on ARMv6 Pi Zero / Pi 1, and stop locking users out of the submodule (#577)
The pinned rpi-rgb-led-matrix commit emits the ARMv7-only `dmb ishst`
instruction in lib/rp1/rp1_rio_backend.cc, guarded only by __arm__, so the
build fails on every ARMv6 board ("selected processor does not support
`dmb ishst' in ARM mode"). Bump the pin to upstream 1ee4f76, which merges
12d839f (guard on __ARM_ARCH >= 7) plus docs only.

The installer also needed two changes for that bump to reach anyone:

- git pull never moves an existing submodule checkout, so a device that
  already failed would keep building the broken commit. The build step now
  moves the checkout forward to the pin — never backward or sideways (a
  `git submodule update --remote` checkout is left alone), and never fatal.
- The submodule git commands ran as root on the user's clone (git's SUDO_UID
  exemption allows it), leaving .git/modules/rpi-rgb-led-matrix-master
  root-owned and the user unable to run git in it. They now run as the
  project directory's owner, and root-owned leftovers are handed back.
  Root-owned installs keep running as root.

test/test_install_rgb_checkout.py covers the non-root sync scenarios under
the installer's strict mode, checks that every called _helper is defined
before use, and pins the one-shot-install.sh -> first_time_install.sh
contract. Verified the tests fail on five deliberate mutations. Root/owner
scenarios were exercised manually under WSL Ubuntu, and the library was
cross-compiled for arm1176jzf-s at both pins (old: rp1_rio_backend.cc fails
at line 120; new: 16/16 sources compile).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 17:38:17 -04:00
ChuckandClaude Opus 5 6d1cbfb70b fix(web): plugin config page survives stored values the schema outgrew (#578)
Two stored shapes broke the config form:

* A scalar under a field that is now an object. News' dynamic_duration
  was a boolean and is becoming an object; render_nested_section did
  `key in true` and the whole page failed to render. Look into dicts
  only, and carry a legacy boolean over as the object's `enabled`, so
  the next save upgrades it without switching the feature off.
* A custom feed logo with a path but no id. The template always emitted
  an empty `logo.id` input, which the save route parsed to null, failing
  the id's string type on every save. Emit it only when there is an id,
  as custom-feeds.js already does.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 14:03:15 -04:00
ChuckandClaude Opus 5 8e6d7c280f test(element-style): cover the stateless element_color clamp path (#576)
#569 fixed _normalize_color and #572 covered the resolver path. That test's own
docstring notes the resolver "normalizes colour separately from element_color",
and the other path had no test: the stateless element_color(), which
src.common.sports_card delegates to and which every one of the nine scoreboard
plugins takes for each per-element colour it draws.

That is the path that regressed. element_color() moved here with the per-element
customization framework, the coercion rejected out-of-range components where the
reader it replaced clamped them, and a rejection reads as "not configured" -- so
one component over 255 painted the element white while the user's colour sat in
their config. Every scoreboard's test_element_text_colors.py failed on it, and
it took two plugin PRs red on CI to surface.

Six cases: clamping, in-range untouched, hex, unparseable fallback, missing
element, and agreement with sports_card.coerce_rgb. The last is the point --
the two shared readers disagreed about the same value, so this asserts against
coerce_rgb directly rather than restating the arithmetic, and any future move
of element_color has to keep them consistent.

Verified by mutation: restoring the rejecting coercion fails two of the six,
alongside the resolver test from #572.

Tests only; no source change.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:27:10 -04:00
ChuckandClaude Opus 5 dac71afedc fix(web): full-height plugin config form and full-width plugin card descriptions (#573)
* fix(web): let the plugin config form use the full page height

The form wrapper has carried `max-h-96 overflow-y-auto` since #145, but
the class was a no-op until #568 defined `.max-h-96` in app.css. That
silently capped the whole config form at 24rem with a nested scrollbar.
Drop the cap so the form flows naturally and the page scrolls.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): give installed plugin card descriptions the full card width

The enable/disable toggle was a flex sibling of the whole text column
(name, metadata, description), so it reserved its width for the full
height of the card body. Descriptions wrapped into a narrow strip,
leaving blank space under the toggle and making cards very tall.

Move the toggle into a header row with just the name and badges, and
render the metadata and description below at full width.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:27:00 -04:00
ChuckandClaude Opus 5 a6b3384032 fix(web): show the action script's error in the file-manager widgets (#574)
A failing plugin action returns a 400 whose JSON body carries the
script's own message, but both file-manager widgets threw it away:
plugin-file-manager's toggle always said "Toggle failed", and
json-file-manager's request helper threw "Server error 400" before
reading the body. That hid of-the-day's "Category ... not found in
config", which is why its toggles looked broken for no reason.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:26:46 -04:00
ChuckandClaude Opus 5 11bf39cd66 fix(web): two plugin config saves that always returned 400 (geochron, news) (#575)
* fix(web): render widget-less arrays of objects as a table, not comma text

An array of objects with no x-widget (geochron's `cities`) fell through to
the comma-separated text input. Jinja joined each item as a Python dict
repr, the save route read them back as a list of strings, and the schema
rejected them -- so every save of the plugin returned 400 "Configuration
validation failed", whatever setting was changed.

Default such arrays to the existing array-table widget, which already
edits arrays of objects and posts `field.N.key` inputs the save route
rebuilds into a list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): don't leave an empty object stub in array items on save

The unchecked-checkbox pass walked into every nested object of an array
item looking for booleans, creating it when absent. A news custom feed
with no logo came out with `logo: {}`, which fails the logo's
`required: [id, path]`, so every save of the news plugin returned 400.

Recurse into a scratch dict instead and attach it only if a boolean was
actually set in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:04:11 -04:00
ChuckandClaude Sonnet 5 bc60b41445 test(element-style): cover ElementStyleResolver's colour clamp path (#572)
* fix(colour): clamp out-of-range text_color components instead of dropping them

_normalize_color returned None for a triple with a component outside 0..255,
and None means "not configured" to element_color -- so configuring
[300, 0, 20] silently handed the element its *default* colour rather than red.
Every scoreboard reads its per-element colours through this path, so the bug
reached all eight.

It is also the odd one out: sports_card.coerce_rgb and
SportsShared._coerce_rgb both clamp, and core's own test is named
test_coerce_rgb_clamps_rather_than_rejecting. The rejecting normaliser arrived
with the shared readers in 82a65ad2 (#425) while the eight plugins' colour
tests kept asserting the clamping behaviour they had before, so the two sides
have disagreed ever since.

Clamped inline rather than delegating to coerce_rgb: sports_card already
imports element_style, so importing back would be circular.

Adds the core assertion whose absence let this drift -- element_color had no
test covering an out-of-range component.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(element-style): cover ElementStyleResolver's own colour clamp path

CodeRabbit flagged that the new sports_card clamp regression test only
exercises element_color(); ElementStyleResolver._resolve() normalizes
configured colours through a separate call to the same _normalize_color,
comparing against a schema/classic reference to decide user_forced_color.
Add a resolver-level case so a future regression in that path (e.g. going
back to rejecting out-of-range components instead of clamping) is caught
too.

Mutation-checked: fails if _normalize_color rejects instead of clamps.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KP6kWxjUtJi72c56GaMmC8

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 13:03:23 -04:00
ChuckandClaude Opus 5 f9b1f87e8d fix: clamp colour components, and let the style editor actually take over (#569)
* fix(element-style): clamp out-of-range colour components instead of rejecting

A regression this framework shipped. The eight scoreboards used to read their
colours through sports_card.coerce_rgb, which clamps; routing them through the
shared element_color sent them through _normalize_color, which rejected any
component outside 0..255 and fell back to the default. So a configured
[999, -5, 20] -- a typo'd bright red -- rendered white instead of (255, 0, 20).

Their own test_element_text_colors.py caught it: one case of nineteen, in all
eight plugins, failing only once the core change reached main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): the style editor takes over its own blocks -- and gets to at all

Two defects, both found by rendering the real partial in a browser rather than
by reading the code.

It was losing a race to its own fields. The hand-off guard asked "do any
fallback controls differ from their server-rendered defaults?" as a proxy for
"is someone editing this?". But the fallback holds this block's own font
fields, and the font-selector widget populates them on the same 50ms timer --
so a plain page load, with nobody touching anything, raced into "dirty" and the
editor removed itself, leaving the 701-line accordion form it exists to
replace. Measured: seven customization.*.font selects dirty ~60ms after
injection, clean again by 400ms. The question is whether a *person* typed, and
event.isTrusted answers exactly that; the listeners now go on synchronously,
because the edit worth protecting can happen before initWidget runs.

It took over too much. Taking over removed the whole fallback section, but a
customization block can hold more than styling -- football keeps
favorite_result_colors there -- so that removed the only UI those fields had,
and the editor also rendered them as an element, giving every row an "enabled"
and three colour columns. Core now marks the blocks it recognises as styling
(the compact declaration already did; hand-written adoption did not), the
widget renders only those, and the template drops only the children the widget
reports owning.

Verified on football's real schema: 28 rows across four mode tabs, columns
Element/Font/Size/Colour/X/Y, favorite_result_colors still editable with its
ten inputs, no duplicated field names, no console errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): style editor no longer drops layout-only fields it never rendered

CodeRabbit flagged elementKeys() in style-editor.js: render() claims the
whole customization.layout child as the widget's own (removing it from the
generic fallback renderer, since posting the same offset twice is worse),
but elementKeys() only listed keys that also have their own top-level style
block. A hand-written schema can put a key under layout that never got one
-- a logo, a timeout indicator, a possession arrow with a position but no
font or colour -- and that key's only control silently disappeared: no row
in the style editor's table (elementKeys never listed it) and no fallback
section either (layout was removed wholesale).

elementKeys() now appends any layout-declared key not already covered by a
style element, so table() renders a row for it (layout columns only, no
style columns) and the wholesale layout ownership claim stays truthful.

Verified against current code before fixing. New regression test
(test/js/unit/test_style_editor_element_keys.js, following this repo's
existing eval-extraction pattern for testing widget JS without a browser)
fails against the reverted function and passes with the fix; added to
run_all.js and the suite table in test/js/README.md.

Full pytest suite: 4887 passed, 62 skipped, 2 failed -- both the
pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias gap on this sandbox,
identical on origin/main, unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dpg3HLWohdCUdzz2QNHanm

* fix(web): style editor no longer strands leaf-valued layout fields

A prior fix on this PR made elementKeys() append any layout-only key with
no style block of its own (a logo, a timeout indicator, a possession
arrow), so table() draws a row for it instead of losing it when the
wholesale `layout` claim removes the generic fallback. That covers a
layout-only key shaped like an object (x_offset/y_offset, ...), because
columnsFor() only ever produced columns from a key's *sub-fields*.

It missed the case where the layout-only key's own value is itself a
leaf -- a plain "show_logo" boolean directly under layout, no x/y object
underneath. elementKeys() still lists it (any row: no matching column),
so it renders as an uneditable blank row and its only control -- the
generic fallback checkbox -- is still gone. Confirmed by executing the
real widget's render() against a synthetic schema in Node (a DOM-stub
harness, not committed): the field's name never appeared as an <input>.

columnsFor() now gives such a leaf key a column keyed to itself
('layout-leaf'), and elementRow() binds it to the leaf's own path
(customization.layout.<key>, matching the name the fallback would have
used) instead of leaving every cell blank.

New regression test (test/js/unit/test_style_editor_layout_leaf_columns.js,
following this PR's existing eval-extraction pattern) checks the leaf
column is produced, is self-keyed, doesn't duplicate, and that a schema
with no leaf-valued layout key is unaffected; wired into run_all.js and
the suite table in test/js/README.md.

test/js/run_all.js: 84 + 6 + 6 = all suites passed (jsdom unavailable
here, DOM suites skip as before). Python suite untouched by this change;
test_style_editor_extra_fields.py, test_style_editor_save_roundtrip.py
and the one PIL-dependent style_editor_takeover.py case fail identically
before this commit -- missing flask/PIL in this sandbox, not this PR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep layout-leaf style-editor columns distinct from name collisions

columnsFor() keyed a layout-only leaf field's column by its bare field
name. If an unrelated element's style block or another element's layout
axis block happened to declare a sub-field with that same name, the
`!seen.has(key)` guard skipped creating the leaf's column, silently
dropping its only control again -- the same failure the leaf-column fix
was meant to close, just reached through a name collision (CodeRabbit
review on 324a7ea).

Key layout-leaf columns under a namespaced id so they can never be
shadowed by an unrelated column sharing their name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): CSS.escape() the owned key before it becomes a selector

container.dataset.ownedKeys round-trips schema property keys through a
DOM dataset attribute, and the takeover handoff spliced each one
straight into '[data-child-key="' + k + '"]' with no escaping --
inconsistent with this codebase's own convention elsewhere
(plugin-file-manager.js, app-shell.js's escapeCssSelector) for building
a selector from a dynamic value. A key containing a quote or backslash
would break the selector or be steerable; Codacy's static analysis
flagged this pattern (1 high ErrorProne finding on PR #569, current
head at the time) as a new issue, though its dashboard is unreachable
from this sandbox (egress to app.codacy.com is blocked) and the
check-run API returned no detail text -- verified and fixed by reading
the diff directly rather than the tool's own description.

Added a source-assertion regression test alongside this file's
existing ones (this behavior lives in an inline script no Python test
executes).

Full suite: 4888 passed, 62 skipped, 2 failed -- both the pre-existing
Europe/Kiev/Asia/Calcutta tzdata-alias gap in this sandbox, identical
on origin/main, unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): every advertised layout offset gets a control in the style editor

The editor took the whole layout section over but matched offsets to style
rows by exact key. A hand-written schema's two blocks were never named alike --
football styles score_text but positions score -- so of football's eleven
positionable things only status_text had a control. Score, odds, both logos,
timeouts, possession, down-and-distance, date, time and records were options
the schema advertised and the renderer reads, reachable nowhere in the UI.

Core now resolves each style element's layout key through alias_keys, the map
the resolver already reads offsets with, and records it as x-layout-key. The
widget reads that rather than carrying a second copy of the rules, and posts
under the key the schema declares: football's own offset reader looks up
layout.score, so a value saved as layout.score_text would be kept and never
drawn. Layout entries no style element claims get an "Other positions" table
with its own columns, in every mode panel as well as the base one, in the order
the plugin declared them.

Verified in a browser against football's real schema: 92 of 92 layout fields
(23 base, 23 per mode) rendered exactly once under their declared names, none
posted under a style key, no duplicated field names, favorite_result_colors
still editable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:48:36 -04:00
ChuckandClaude Opus 5 7e580dc005 fix(wifi): make Connect work from the setup AP (#571)
* fix(wifi): make Connect work from the setup AP

Joining a network from LEDMatrix-Setup has to take the AP down first, which
drops the phone that sent the request. The connect endpoint answered only
after the attempt finished, so the browser never got a reply and the WiFi
tab's Connect button appeared to do nothing.

- /wifi/connect answers 202 immediately while the AP is active and connects
  in a background thread; the result (never the password) is reported via
  /wifi/status as last_connect_attempt. A second connect while one is
  pending gets 409.
- connect_to_network holds a /tmp flag for the attempt; the monitor daemon
  skips AP management while it is fresh. Previously the daemon's
  disconnected counter, accumulated over the whole AP session, re-enabled
  the AP on its next tick in the middle of the connect.
- The WiFi tab and captive setup page explain the handoff up front, and on
  reopening show why the last attempt failed. The wrong-password message
  now works: the route sets the error_type the captive page checks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(wifi): serialize connect attempts on both paths

Addresses CodeRabbit review on #571:

- Check for a pending attempt before branching on AP state. A background
  attempt takes the AP down long before it finishes, so a second click
  used to bypass the 409 and start a competing synchronous connect.
- Record pending for the synchronous (non-AP) path too, so two requests
  can't overlap and have the first clear the daemon's in-progress flag
  while the second is still connecting.
- Clear the pending state if the background thread fails to start, rather
  than refusing every later request until restart.
- Say the setup network returns "within a few minutes": a stale flag plus
  the daemon's grace period can take longer than one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:40:41 -04:00
ChuckandClaude Opus 5 d51f7ada14 chore(scroll): drop the dead sub-pixel path, and two dev-tooling papercuts (#570)
Three independent changes, none of which alter runtime behaviour.

1. Remove ScrollHelper._get_visible_portion_subpixel and
   _interpolate_subpixel (162 lines). get_visible_portion dispatches only to
   _blend_visible_portion, so _get_visible_portion_subpixel had no caller, and
   _interpolate_subpixel was reachable only from inside it -- a closed island.
   _blend_visible_portion's own docstring already records that the scipy path
   it replaced was dead; the replacement landed but the corpse stayed.

2. scripts/check_plugin.py: also search ../ledmatrix-plugins/plugins. The
   scoreboards live in the sibling checkout, so --all silently skipped every
   one of them and only --plugin-dir reached them.

3. .gitignore: ignore config/.config_secrets.json.tmp.*, which the suite
   leaves behind several of per run.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 09:42:50 -04:00
ChuckandClaude Opus 5 d1e821c625 fix(web): harden, polish and optimize the web UI per the Sept 2026 audit (#568)
* fix(web): harden, polish and optimize the web UI per the September 2026 audit

Works through docs/archive/WEB_UI_AUDIT_2026-09.md (health 8/20).

Implementation integrity (P0)
- app.css now defines every utility class the templates and JS use,
  including .hidden, so the ~145 JS show/hide toggles work. Button reset,
  and base component rules (.btn, .form-control) wrapped in :where() so
  utility classes on the same element win. New static-audit test fails
  when a used utility class has no rule.

Accessibility
- Focus rings render (the old ring rule referenced undefined variables);
  one :focus-visible outline everywhere; skip link; labelled nav landmarks.
- Shared dialog helper (js/utils/dialog.js): role/aria-modal, focus trap,
  Escape, focus return, applied to every modal.
- Named icon-only buttons and labelled ~70 form fields.
- Toasts announced once; errors persist >= 10s; one showNotification.
- Captive WiFi page: live region, timeouts, dark mode, 16px inputs.

Performance (Pi Zero 2 W)
- SSE streams and tab timers pause when hidden or off-tab; the display
  stream only runs while a preview is visible. app-shell.js deferred.
- Widget scripts served as one versioned bundle (/assets/widgets.js):
  52 -> 21 script tags, 66 -> 35 requests on first load.
- Stdlib gzip fallback when flask-compress is missing: first-load JS/CSS
  1358 KB -> 291 KB on the wire. SSE untouched.

Theming and responsive
- File managers, form fields and Fonts upload on theme tokens; bare
  inputs themed in dark mode; no more white surfaces.
- No horizontal overflow at 375px on any tab; 44px touch targets on
  coarse pointers; reduced-motion respected; header title truncates.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings on #568

- json-file-manager: focus-trap releases kept in a Map (no dynamic
  property access or delete; no value-returning forEach callback)
- notification / schedule-picker: style and day-label lookups via Map
- app.js: move the pending-queue assignment out of the expression
- diff_viewer / error_handler: named function declarations instead of
  arrow consts

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: check the OAuth widget ships in the widget bundle

base.html no longer tags widget scripts one by one; they load through
/assets/widgets.js. Assert the page requests the bundle and the bundle
contains google-oauth.js, which is what the test was protecting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): address review feedback on #568

- widget bundle version fingerprints every file (name, mtime_ns, size)
- gzip fallback appends Accept-Encoding to an existing Vary header
- dialog helper: releasing a non-top dialog no longer moves focus out of
  the dialog the user is in
- labels: file-upload targets its file input; fallback config fields get
  label for/id pairs; native color input has a fallback name
- utility audit also reads class names inside bound :class expressions

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): give the native color-picker input an accessible name

CodeRabbit flagged this on PR #568 as an outside-diff finding (never
posted inline, so it was missed in the round of fixes that addressed
the other 6 review comments). The <input type="color"> only carried a
title attribute; screen readers don't reliably announce title, and
there's no other label naming the control when showHexInput is false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings in app-shell.js

- drop the unused catch binding on the SSE JSON parse
- move the pending-notification queue assignment out of the expression

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): contain plugin widgets/ dir and bound style-editor retries

From CodeRabbit review on #568 (code that arrived with the main merge):
- serve_plugin_widget resolves widgets/ with resolve_under before
  resolving the manifest script under it, so a symlinked widgets
  directory can't become the containment base (CWE-22). New test.
- style-editor init stops polling after ~10s when the widget never
  registers and leaves the plain fallback fields in place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 09:42:24 -04:00
ChuckandClaude Opus 5 69d408b321 feat(core): one per-element display-customization framework, wired into the web UI (#566)
* fix(sports): rebuild un-shared faces through the pinned layout engine

unshare_element_fonts re-instantiates a duplicate font face so two
elements can be told apart by id(). It did so through bare
ImageFont.truetype, which takes PIL's default layout engine rather than
the one src/common/font_layout.py pins. Raqm and Basic disagree on
fractional advances -- that disagreement is the reason the pin exists,
having broken golden images across machines -- so a rebuilt face could
measure differently from the shared face it replaced, on any host where
Raqm is installed.

These were the only two call sites in src/ bypassing the pin.

The guard asserts that the rebuild goes through the pinned loader rather
than comparing engine values: where Raqm is absent, bare truetype returns
BASIC anyway, so an engine comparison passes whether or not the pin is
honoured. The first draft of this test did exactly that and passed with
the bug reintroduced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): drop the two dead client-side config-form renderers

generateConfigForm and generateSimpleConfigForm (580 lines) were defined
on the Alpine component and never called: server-side Jinja replaced them,
as pages_v3.py:641 records. Nothing in any template invokes them -- there
is no x-html in the templates and no bracket access on the component.

They carried their own x-widget dispatch, which made them an active trap:
the next person adding a widget would reasonably think both renderers
needed updating.

plugins/config_manager.js (PluginConfigManager, 133 lines) goes for the
same reason -- loaded on every page from base.html, referenced only by
itself and by an archived doc.

Kept, having checked them: widgets/example-color-picker.js is the worked
example docs/widget-guide.md points plugin authors at, and
widgets/plugin-loader.js is the client half of a documented feature
(manifest-declared plugin widgets) whose server route is missing --
soccer-scoreboard already ships a widgets/custom-leagues.js that this
loader is meant to fetch. That is an unfinished feature to complete, not
dead code to delete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): serve plugin-declared widgets, and actually ask for them

LEDMatrixWidgets.loadPluginWidget has always fetched
/static/plugin-widgets/<plugin>/<widget>.js, and docs/widget-guide.md has
always documented that path, but nothing served it. soccer-scoreboard has
shipped a 17KB widgets/custom-leagues.js since August that could never
load. Both halves were missing, not just the route:

- serve_plugin_widget serves the script from the plugin's widgets/
  directory as text/javascript. The manifest is the allowlist -- only a
  widget the plugin declares is reachable -- so installing a plugin does
  not publish everything it ships. Path handling mirrors the sibling
  serve_plugin_web_ui: allowlist regexes, os.path.basename, resolve() +
  relative_to() containment, and the ledmatrix- prefix fallback. The
  declared script name is guarded too, since it comes from the plugin
  rather than the request.

- The config form never requested one. Its x-widget dispatch is a
  hardcoded list of core widget names, so a plugin's own widget fell
  through to a plain text input. An unrecognised x-widget on a string
  field now asks ensureWidget() for it. The text input stays as the
  fallback and is removed only once the widget has actually rendered, so
  a missing or broken widget costs the user an editor rather than their
  configured value on the next save.

- manifest_schema.json gains "widgets", so the declaration is validated
  rather than merely tolerated by additionalProperties.

Verified in a browser against the real partial: a declared widget loads,
registers and renders, and its field posts exactly one value; a field
whose widget 404s keeps its text input and still posts its value.

Not addressed: loadPluginWidgetsFromManifest still has no caller. The
per-field ensureWidget path is lazier and is what the form now uses, so
that bulk helper is dead weight -- worth removing, but left alone here
rather than inventing a call site for it.

Known limitation, documented: only string-typed fields take this path.
object/array/boolean/number fields and enums are dispatched by the
template's own branches, which still only know core widgets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(element-style): a wrong-size BDF now keeps its font, not its size

BDF fonts are fixed-size bitmap strikes: FreeType accepts only the pixel
size baked into the file and raises for anything else. 32 of the 35
shipped fonts are BDF, so a size picked in the web UI usually is not a
valid strike -- and load_font caught that failure with its generic
"unloadable font" handler, which substitutes PressStart2P. Asking for
5x7.bdf at size 10 therefore rendered a completely different typeface,
silently.

It now falls back to the file's own native size instead, which is what
SportsCore._load_custom_font_from_element_config has always done. The
native size is read via FontManager._read_bdf_native_size rather than a
fourth copy of that parser, matching how core.py already delegates.

Also here, because they are the same code path:

- native_bdf_size() is exposed for the web UI, which needs to know when a
  size field can take effect at all. None means "free choice".
- ElementStyle.font_size now reports the size actually realised rather
  than the one requested. Callers lay out from it, and reserving space
  for a size nothing was drawn at is how this surfaces.
- The module font cache is a bounded LRU (256) instead of an unbounded
  dict. The display process runs for weeks and every config save can add
  a (font, size) pair; every other hot cache in the codebase is bounded
  this way.

Untouched configs are unaffected: the shipped classic fonts are the three
TTFs, so nothing was hitting the substitution path by default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): per-mode style and offset overrides

Lets one element be styled differently per situation -- a scoreboard's
live / upcoming / recent cards, weather's current / hourly / daily
screens -- under customization.modes.<mode>.

The mode is bound at construction rather than passed per call. That is
what makes this cheap to adopt: SportsUpcoming and SportsRecent are
already separate instances with distinct SKIN_MODE values, so binding
once makes every existing style()/offset_value() call site mode-aware
without editing any of them. A per-call mode argument exists for the rare
host that renders more than one mode.

The two layers answer different questions, deliberately:

- The base layer keeps the existing "differs from the schema default"
  rule, because the save flow writes the full default object into
  config.json whether or not the user touched it.
- A mode layer is pure override -- its fields default to None, so
  presence is intent. Nothing writes into it unasked, so there is nothing
  for the stricter rule to protect against.

None therefore means inherit, and has to stay distinct from 0: a mode
y_offset of 0 means "sit at the base position", not "no preference".
This is the same distinction scroll_card.switch_* draws with "inherit".

A malformed mode value falls back to the resolved base value rather than
to the caller's default -- caught by the degradation tests, which is what
they are for: resolving the mode first let one bad string in a mode block
silently discard a good base offset.

With no modes block, and for every existing caller, resolution is
unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): declare per-mode overrides in config_schema.json

A plugin adds "x-style-modes": ["live", "upcoming", "recent"] alongside
its x-style-elements declaration and gets a customization.modes.<mode>
group per mode, with every field of every declared element repeated as an
override.

Those override fields are typed nullable and default to null, which is
the whole trick. The save flow writes schema defaults into config.json
wholesale, so giving a mode field the base element's default would make
every mode a frozen copy of the base the first time a user pressed Save,
and the base would stop reaching them. Null means inherit. The mutation
test for this is explicit: with concrete defaults, a base font_size of 14
resolves as 10 with user_forced set.

min/max from the declaration carry into the mode blocks, so an
out-of-range override is rejected by validation rather than clamped
silently at render time.

Also: the emitted font field now carries "x-widget": "font-selector". The
widget already shipped and the config form already allowlisted it -- the
hint was simply never emitted, so the field rendered as a bare text box
that the user had to type a font filename into.

Verified through the real SchemaManager path -- load_schema, defaults
extraction, merge_with_defaults, validation, then resolution -- rather
than against a hand-built dict, since the thing at risk is what that
pipeline does to a null.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): render the config form from the schema the save route validates

The form read config_schema.json with a raw json.load while
api_v3.save_plugin_config went through SchemaManager. Those are not the
same schema: SchemaManager applies expand_style_elements, which turns a
compact customization.x-style-elements declaration into the per-element
blocks the form knows how to render.

Without it, that customization object has an x-style-elements key and no
"properties", so the template's object branch matched nothing and the
section rendered as empty space -- while saving still validated against
the expanded shape. of-the-day ships the compact form, so its
customization section has been invisible in the web UI.

pages_v3 gains a schema_manager the way it already has config_manager and
plugin_manager. use_cache=False matches the save route, so an edited
schema is not served stale during plugin development. The raw read stays
as a fallback for callers that register this blueprint without one.

Checked before making the change: load_schema does nothing here except
read, validate and expand -- inject_skin_selector is a separate method it
does not call -- so this is not a behaviour change for schemas without
the declaration.

The test pair renders the same compact schema with and without a
SchemaManager, so it documents exactly what was broken as well as what is
fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): style-editor widget -- a row per element instead of 65 accordions

Rendered element by element, a realistic scoreboard's customization block
is 65 nested sections, and reaching one per-mode font size takes five
levels of expanding. The widget collapses that to one compact row per
element -- font, size, colour, X, Y -- with a tab per declared mode.

It emits ordinary inputs under the same dotted names the generic renderer
would produce, so the save/validate/merge pipeline is untouched: no hidden
JSON blob and no new server-side parsing. It is driven entirely by the
schema block it is handed, so fields added to the schema later appear
without editing the widget. If it fails to load or throws, the generic
nested rendering it replaces is left in place.

Fixing two things the save path got wrong for nullable fields, found by
posting what the widget actually emits:

- The indexed-array recombiner (text_color.0/.1/.2 -> one list) compared
  the declared type to the string 'array', so a per-mode colour, typed
  ["array", "null"], was never reassembled and failed validation on save.
  _parse_form_value_with_schema had the same comparison.
- A blank nullable field became [] rather than None, which then failed the
  minItems the colour array declares. Null is the inherit sentinel, so it
  has to survive.

And two things the widget itself got wrong, found by looking at it:

- An unset base control fell back to the select's first option, so an
  untouched scoreboard claimed every element used 10x20.bdf -- and the
  size box then locked itself to that bitmap font's fixed size. Base
  controls now show the schema default; mode controls stay blank, because
  blank there means inherit.
- Elements arrived alphabetised (Detail and Odds above Score). Flask's
  JSON provider sorts keys, so declaration order has to be stated
  explicitly; expand_style_elements now emits x-propertyOrder, which the
  generic renderer already honoured too.

Size is disabled and shown as fixed for a bitmap font, using the
scalable/native_size the font catalog now reports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): visibility, alignment and scale per element

Completes the customization vocabulary: hide an element, align it, and
resize a logo, alongside the font/size/colour/offset that already existed.
All three per mode.

They resolve to "change nothing" until the user asks for something -- True,
None and 1.0 -- rather than to whatever the schema declares. That is the
same invariant the font fields keep: a caller that honours them still
renders an untouched config exactly as it did before they existed. A
schema default therefore does not count as a choice, which matters because
the save flow writes that default into config either way.

scale sits in the layout block with the offsets rather than in the element
block, because it is geometry: a logo has a scale and no font. The widget's
columns come from the schema, so a logo row shows visibility, offsets and
scale and no empty font cell.

Two bugs found by the tests rather than by reading:

- A nullable enum needs null in its enum list, not just in its type. The
  mode copy of `align` defaulted to null and then failed its own schema, so
  a plugin declaring any enum field with modes could not save at all. Six
  tests failed on this before any of them reached what they were testing.
- defaults_from_schema only ever extracted font/font_size/text_color, so
  the schema defaults for the new fields were invisible to the resolver and
  a declared default read as a user choice.

Widget: the table scrolls horizontally and pins the element-name column.
Nine columns do not fit the config panel, and clipping them hid the offsets
entirely while scrolling them made every row anonymous.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): resolve elements under the names plugins actually use

Two naming conventions collided as the scoreboards grew. Counted across
the published schemas: the style block names elements with a _text suffix
(score_text, status_text, detail_text), while the layout block mostly uses
the bare noun (score, date, time, odds) -- except status_text, which kept
the suffix in seven plugins and lost it in two. records vs record splits
seven to two the same way.

A lookup now tries the exact name first and then the spellings that mean
the same thing. Exact-first is what makes this inert for any config that
already matches; the aliases only decide cases that resolved to nothing
before.

This is also what makes migrating to the compact declaration form safe.
That form uses one key for both blocks, so a scoreboard adopting it asks
for layout.score_text while its users have layout.score saved -- without
the aliases, every offset they had dialled in would silently become 0.

Applies to the style block, the layout block, the schema defaults and the
per-mode overrides, since the drift shows up in all four.

Not attempting to canonicalise on write: renaming keys in config.json
would break the plugins still reading the old spelling from their own
bundled code, and the drift costs a dict miss rather than correctness.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(plugins): BasePlugin.styles -- per-element styling every plugin inherits

Adopting the element-style system meant repeating three things in every
plugin: a guarded import, finding its own config_schema.json, and
rebuilding the resolver when on_config_change swapped the config dict.
This is those three things once, on the class all 45 plugins already
inherit from.

    title = self.styles.style('title_text',
                              classic_font='PressStart2P-Regular.ttf',
                              classic_size=8, classic_color=(255, 255, 255))

The classic_* arguments are the adoption contract: with nothing configured
they come back verbatim, so a plugin that switches to this renders exactly
as before until a user changes something.

A plugin with one instance per display mode sets STYLE_MODE on the class
and every existing lookup becomes mode-aware without a call site changing
-- which is the point of binding the mode to the resolver rather than
passing it per call. styles_for() covers a plugin that renders several
modes from one instance.

Schema discovery reads the concrete class's own module rather than this
file, because this file lives in src/plugin_system where no plugin schema
exists -- the same trap SportsCore._config_schema_path documents. The
first mutation test for that passed anyway: an installed plugin's module
directory and its entry under plugins_dir are the same path, so the test
could not tell the two apart. The case where they diverge is a plugin
symlinked in for development, and the test now forces that shape.

Getting discovery wrong is silent rather than loud: with no schema the
resolver has no defaults to compare against, so every configured value
reads as a deliberate override and the plugin quietly stops honouring its
own shipped styling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): adopt hand-written customization blocks, and widen the font list

Nineteen plugins spell their style elements out longhand instead of
declaring them -- football's block is 701 lines for seven elements -- and
predate this system entirely. Core now recognises that shape, so they pick
up the row-per-element editor and the real font picker on a core update
rather than on a plugin release. Checked against every published schema:
21 plugins adopt, and the defaults of each still validate against the
schema generated for it.

Detection requires *every* field in a block to be one this system
understands. A looser "has at least one style field" rule sweeps in
baseball's `count`, which carries a text_color beside geometry that means
nothing here. That distinction took three attempts to test: the first two
assertions passed under both rules, because an over-eager rule leaves a
fontless block looking untouched and only surfaces as an extra row in the
editor.

The hardcoded font enum is replaced rather than extended. Football lists
five of the thirty-five installed fonts, which is why a font a user
uploads can never appear in one. It is not a curated safe set -- it omits
some twenty other faces that fit the declared size cap just as well -- it
is the fonts that happened to exist when it was written.

Widening it does need a guard, though, and not the one the schema already
has: a bitmap font ignores font_size and renders at its size baked into
the file, so `maximum: 16` cannot stop a 27px face. The picker now filters
out fixed-size fonts taller than the element's own declared ceiling, which
drops exactly the four that would overflow a 32px panel and keeps the
other thirty.

Per-mode overrides stay opt-in: core cannot invent a plugin's display
modes, so `x-style-modes` remains the one line that unlocks them. Their
layout half covers every positionable element rather than only those with
a style block -- the two namespaces do not line up in a hand-written
schema, and football positions six things (logos, timeouts, possession)
that have no style block at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): remove the two Fonts-tab panels that reported invented data

"Element Font Overrides" let a user configure an override, showed a
success toast, and changed nothing. All three endpoints behind it were
stubs -- GET returned a hardcoded {}, POST and DELETE returned success
without calling anything -- each marked "This would integrate with the
actual font system".

Wiring them to FontManager would not have fixed it. The machinery there is
real (_load_overrides/_save_overrides persist config/font_overrides.json,
resolve_font applies them, and the countdown plugin genuinely consumes
it), but the panel's element dropdown offered eleven invented keys --
nfl.live.score, clock.time, weather.current -- that no plugin has ever
read. An override saved against one of those would have persisted
correctly and still done nothing.

"Detected Manager Fonts" goes for the same reason. It claimed to show
"fonts currently in use by managers (auto-detected)"; its own comment said
"we'll simulate this", and it listed every font in the catalog with a
hardcoded usage_count of 1 -- the panel beside it, with fabricated
numbers attached.

Per-element font choice now lives in each plugin's own config editor,
against the elements that plugin actually has, and covers size, colour,
offsets, visibility, alignment and scale rather than family and size.

Kept: the font library (upload, preview, delete), which works, and
/fonts/tokens, which is a stub but genuinely feeds the preview's size
dropdown. FontManager's override methods are untouched -- countdown uses
them.

Verified in a browser with the tab's JS running: no console errors, 35
fonts listed, upload and preview intact. Removing the panel meant unwiring
it from populateFontSelects too, which would otherwise have bailed out
early on the missing select and left the preview dropdown empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(sports): one reader for element colours and layout offsets

There were two copies of the per-element colour read and three of the
layout-offset read. They had already drifted -- the scroll-card renderer
carries a comment about having ignored offsets its own schema advertised
-- and each new capability had to be added to all of them or silently work
in some places and not others.

All of them now go through src.element_style, which is what carries the
alias handling and the per-mode lookup. That lands immediately for the
nine plugins importing these modules: a scoreboard asking for `score_text`
offsets finds the `layout.score` its users configured, and a Live instance
resolves its own colours through SKIN_MODE without any call site passing a
mode.

_normalize_color learned "#RRGGBB" in the process. The scoreboards' own
readers have always accepted it, so the shared one had to, or consolidating
would have quietly dropped a form users' configs may hold. _coerce_offset
picked up the non-finite guard the scroll-card reader had and the other two
did not.

_get_layout_offset is promoted onto SportsCoreSharedMixin. Each plugin
still carries its own copy in its bundled sports.py, which wins by MRO --
so adopting this is a deletion in the plugin, and until that deletion
nothing changes for it.

Note for whoever runs the suite next: test_display_dirty_tracking.py is
order-dependent. Fifteen of its tests failed in one full run and passed in
the next with no change in between, and pass in isolation. Pre-existing,
unrelated to this, but it makes a full-run diff untrustworthy until it is
fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): record the element-style work under Unreleased

This file's own preamble asks for it: a plugin may delete its bundled
fallback copy of a core module only when its manifest floors on the first
release that shipped that module, which requires the additions to be
recorded here against a version.

Names a plugin can now import and floor on -- the stateless layout_offset
and element_color readers, alias_keys, native_bdf_size, the resolver's mode
binding, BasePlugin.styles, and the promoted
SportsCoreSharedMixin._get_layout_offset -- plus the schema and web-UI
changes, the four fixes and the three removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(fonts): log the BDF native-size read failure instead of swallowing it

The bdf-native-size lookup in get_fonts_catalog() caught any exception
and silently discarded it. Every other guarded read added in this PR
(the manifest parse in _declared_widget_script, the SchemaManager
fallback in _load_plugin_config_partial) logs before falling through
to the same degraded behavior. This one didn't, which is the shape a
silent-exception-swallow lint rule flags. Behavior is unchanged --
native_size still comes back None -- but a corrupt or unreadable BDF
file now leaves a trace.

Verified: font-related tests (140) and the full suite still pass,
with only the 2 pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias
failures already present on origin/main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address CodeRabbit findings on the style-editor/font-selector PR

- Fix _load_font_sized double-wrapping the (font, size) tuple on the
  missing-font path, which handed callers a tuple instead of a font.
- Fix _set_nested_value skipping an explicit None when the key already
  existed, which silently kept stale overrides when a user cleared a
  nullable per-mode field or blanked all channels of an indexed color.
- Preserve BDF scalable/native_size metadata through fetchFontCatalog's
  catalog-format mapping so maxFixedSize filtering actually applies.
- Stop caching an empty array on a failed font-catalog fetch so a later
  call can retry instead of being stuck with the failed result.
- Keep a saved font selected in the style editor even when it no longer
  fits a newly declared maxFixedSize, instead of silently deselecting it.
- Don't drop in-progress user edits to fallback fields when a plugin
  widget finishes loading asynchronously and takes over the form.
- Tighten the removed font-override endpoint test to assert 405, not
  just != 200.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): a partial save no longer switches off checkboxes it never showed

An HTML checkbox posts nothing when unchecked, so the save route walked the
schema and forced every boolean missing from the form to False. That is right
for the rendered form and wrong for every other caller: a script, the MQTT
bridge or a curl against the documented endpoint never rendered a checkbox, and
reading its silence as "all off" turns a one-field save into a mass disable.

Found on hardware. Posting four customization.* keys to a live device switched
off nfl.enabled, ncaa_fb.enabled and every display-mode toggle in one request.

The form now reports the top-level sections it drew (__rendered_section), and
inside those an absent checkbox still means unchecked -- including a section
whose only fields are checkboxes that are all off, which no heuristic could
recover. A post with no marker only touches objects it actually posted a field
from. Meta fields are dropped before form keys are treated as config paths,
because unknown keys are otherwise written straight into config.json.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(sports): resolve element colour by name, and honour visible/align/scale

Two of the three gaps this framework shipped with.

Colour by name. A draw resolved its colour by comparing the *identity* of the
font object it was handed, which cannot tell two elements apart when they share
a face -- so those draws went out white. Every bitmap font is in that case,
because a freetype.Face cannot be re-instantiated to un-share it, which is how
an element rendered in any of the 32 shipped BDF fonts silently lost a colour
its picker had offered all along. _draw_text_with_outline now takes
element="score_text" and reads the colour by name; the identity path remains
for un-annotated callers, but narrows before giving up -- one configured colour
among the sharers is the only thing the user can have meant.

Visible, align and scale. The resolver has understood these since the
framework landed and nothing consumed them: an element could be marked hidden
in the web UI and still render. Adds the stateless readers, the mixin
accessors, and a scale parameter on the one shared logo-sizing seam (keyed into
the cache, so two elements scaled differently cannot be served each other's
image). Naming an element in a draw also honours its visibility.

Untouched configs are unaffected: every new parameter defaults to today's
behaviour, and all ten affected plugins render pixel-identically to main across
every harness size.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): how to declare styleable elements; harden the widget's lookups

The plugin-author guide for the compact x-style-elements declaration -- what
each key does, how to read values back without breaking the "user-forced only
when it differs from the default" rule, and why a hand-written block needs no
changes to be adopted.

Also clears the static-analysis findings on style-editor.js. Every lookup in
that file is keyed by something out of a schema or a saved config, so a key of
__proto__ or constructor would walk the prototype chain and hand back a
function instead of a schema; reads now go through an own-property helper. The
panel registry became a list, and the flagged vars moved to their function
roots.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear the remaining static-analysis findings

Five, all on lines this branch touched.

The Python one is not a new defect: _set_missing_booleans_to_false's first
parameter was always named `config`, which shadows the `config` submodule
imported for its side effects at the bottom of this module. Editing the
signature simply put the existing warning on a changed line. The parameter is
the plugin's config dict, so `plugin_config` is what it should have been called
anyway; callers pass it positionally and are unaffected.

The JavaScript ones are the object-injection rule firing on reads keyed by
data. own() now goes through a property descriptor, so the one unavoidable
data-keyed read is no longer a computed member access; at() consumes its path
instead of indexing it; and the column set is a Map, which has no prototype to
pollute and needs no guarded reads at all.

Verified the widget still renders identically against football's real schema:
29 element rows, all four mode tabs, values populated, no console errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): drop the hasOwnProperty alias the descriptor read made redundant

own() now reads through Object.getOwnPropertyDescriptor, so the alias it used to call has no remaining reference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 11:53:50 -04:00
ChuckandClaude Opus 5 92ac231138 fix(fonts): load 4x6 on its pixel grid, from any working directory (#565)
* fix(fonts): load 4x6 on its pixel grid, from any working directory

`extra_small_font` loaded 4x6-font.ttf at 6, off the face's 7px grid.
Under `draw.fontmode = "1"` the mono rasteriser thresholds each glyph at
50% coverage, so every glyph lost its fourth column and deformed:
christmas-countdown rendered "UNTIL" as "VM1JL". The advance is 5px at
both sizes, so snapping to 7 reflows nothing.

- Sizes in DisplayManager._load_fonts go through crisp_size() instead of
  literals. crisp_size / FONT_PIXEL_GRID / FONT_NAME_ALIASES move to
  src/common/font_layout.py; sports_card re-exports them.
- Mirror the fix in VisualTestDisplayManager, the harness's fork of
  _load_fonts. Without it every golden is blessed at the old size.
- Resolve bundled font paths against the install root, not the cwd.
  FontManager._resolve_asset_path now delegates to
  font_layout.resolve_asset_path (kept by name; plugins probe for it).
- The startup banner's middle rung snaps to 7; the 5 rung stays off-grid
  on purpose (the only size that fits a dotted quad on 64px).
- loading.py reads all plugin JSON as UTF-8 (cp1252 on Windows aborted
  check_plugin.py on a 0x9d byte).
- check_plugin.py reports in ASCII and never dies on an unencodable char.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(fonts): resolve relative asset paths from the install root, not the cwd

resolve_asset_path checked os.path.exists(relative_path) unconditionally,
so a relative asset path was still resolved against the process cwd first
-- exactly the dependency this module exists to remove. An unrelated
working directory that happens to contain assets/fonts/4x6-font.ttf (a
stale checkout, a copied assets folder, another project) would shadow the
real bundled font instead of the install root ever being consulted.

Only an absolute path is now returned as-is; a relative path always
resolves against _INSTALL_ROOT first, matching the docstring's stated
contract. FontManager._resolve_asset_path delegates to this function, so
it's covered by the same fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 10:47:01 -04:00
ChuckandClaude Opus 5 772258f73e docs: add PRODUCT.md and September 2026 web UI audit (#567)
* docs: add PRODUCT.md product context for web UI design work

Captures durable product truth (users, positioning, operating context,
constraints, principles) so design passes on the web control panel share
one source. Open decisions (offline-only, CSS build step, WCAG target)
are recorded as undecided rather than adopted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: add PRODUCT.md and September 2026 web UI audit

PRODUCT.md captures durable product context (users, positioning,
operating context, constraints, principles) for web UI design work.
Open decisions (offline-only, CSS build step, WCAG target) are recorded
as undecided rather than adopted.

docs/archive/WEB_UI_AUDIT_2026-09.md records the technical audit of
web_interface/ (8/20): the hand-rolled Tailwind subset in app.css leaves
333 used utility classes undefined (including .hidden), focus rings never
render, modals lack dialog semantics, and SSE/polling never pause. Includes
a verified-and-rejected section so the cache-busting false positive is not
re-raised.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 16:53:45 -04:00
ChuckandClaude Opus 5 9ad7528c9b fix(config): stop same-second backups overwriting each other (#564)
* fix(config): stop same-second backups overwriting each other

A backup's version is its identity. save_config_atomic() hands the path
back, rollback_config(backup_version=...) looks that version up, and the
paired secrets backup is found by reusing the same string.

The version was stamped at second granularity, so two saves inside the
same second produced the same filename and the second shutil.copy2()
silently overwrote the first backup. The path a caller was still holding
then pointed at different content, and rolling back to it restored the
wrong config. A user saving twice in quick succession lost a restore
point with no error.

list_backups() made it worse. It parsed the version off Path.stem, which
drops only the last dot-component, so for config.json.backup.20240101_120000
parts was ['config', 'json', 'backup'] and parts[-2] was 'json' -- never
'backup'. The filename branch was unreachable: every backup fell through
to the mtime fallback and reported a second-granularity restamp of its
mtime rather than the name on disk, so a unique filename alone would not
have been enough for rollback to find the right version.

Stamp microseconds, and never overwrite an existing backup -- on a
collision bump a -N suffix rather than lose a restore point. Parse the
version off the exact glob prefix so it round-trips with the filename,
still reading the legacy second-granularity format so restore points that
predate this keep working.

Two tests had encoded the bug:

  - test_multiple_config_changes asserted a rollback produced plugin1=45
    with plugin2=15, a state no single backup ever held -- 45 was only in
    the second backup, 15 only in the first. It passed because the two
    saves collided onto one file, so the first version resolved to the
    second's content. Corrected to the state that backup actually holds.

  - test_backup_rotation asserted against a hardcoded max of 3 while
    setUp configured 5, and still passed: every save in its loop collapsed
    onto a single filename, so there was only ever one backup to count and
    rotation was never exercised. It now asks the manager for its limit
    and overshoots it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): fold collision suffix into ordering, close backup-path race

_parse_backup_version() stripped any trailing "-segment" unconditionally,
so a collision-suffixed backup parsed to the exact same timestamp as its
sibling and list_backups() had no deterministic way to order them. Only
strip the suffix when it's numeric, and fold it back in as extra
microseconds so same-tick collisions sort newest-first reliably.

_create_backup() also checked backup_path.exists() before shutil.copy2(),
which two concurrent callers can both pass for the same path -- the second
copy2() then silently destroys the first call's restore point. Reserve
each path (config and, when configured, secrets) with exclusive file
creation instead of a check-then-copy, retrying on a real conflict.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 16:13:50 -04:00
ChuckandClaude Opus 5 6b3028ad58 test: isolate DisplayManager globals across modules, and name the failure (#563)
Follow-up to #562. That commit fixed the actual cause of the intermittent
15-test failure in test_display_dirty_tracking.py -- the emulator's fixed TCP
port 8888, a machine-wide singleton that a concurrent pytest process takes
away. This adds the two things that would have made it a five-minute
diagnosis instead of a long one, and closes the other door into the same
failure.

Confirmed the module is order-independent as it stands, on this checkout:

  pytest test/ -q, three times          115 failed / 4464 passed / 63 skipped,
                                        byte-identical failure sets, the
                                        module 21/21 passed each time
  module forced last (197 files first)  identical failure set
  module forced first                   identical failure set
  module after each of test_display_manager, test_display_controller,
    test_display_controller_vegas_tick, test_skin_system, test_sports_scroll,
    test_initial_update_budget, test_display_double_parity,
    test_initializing_screen                                     all pass
  four concurrent processes on the file                          21/21 each

And reproduced the original, to be sure the diagnosis in #562 is the whole
story. Holding 0.0.0.0:8888 from a separate process:

  HEAD's test/conftest.py        21 passed
  pre-#562 test/conftest.py      15 failed, 6 passed

The 15/6 split is not arbitrary: the six survivors are the only tests in the
file that never touch dm.matrix.

conftest.py: DisplayManager is a process-wide singleton and the RGBMatrix /
RGBMatrixOptions names it constructs through are module globals, bound once at
import. All three are shared by every test module in the run, so a module that
leaves an instance in _instance -- or leaves patch('src.display_manager.
RGBMatrix') standing -- changes what the NEXT module builds, invisibly, and
only in a full run. A module-scoped autouse fixture now resets the singleton
and restores either binding if a patch outlived its module. Module-scoped
rather than per-test so that files sharing one manager across their own tests
keep doing so; only the leak across the module boundary is cut. Autouse
fixtures are set up ahead of requested ones, so this is finalised after a
module's own DisplayManager fixture. Verified with a throwaway pair of probe
modules -- one leaks a patch and a singleton, the next asserts both are clean
-- which passed and were then removed.

test_display_dirty_tracking.py: _setup_matrix() swallows every construction
failure and falls back to matrix=None, so a broken environment arrived as
fifteen identical "'NoneType' object has no attribute 'SwapOnVSync'" errors
naming neither the fixture nor the cause. The fixture now fails once, and
says where to look; under a held port it reads

    DisplayManager fell back to matrix=None: RGBMatrix construction raised...
    Known causes: the emulator adapter losing a fixed TCP port to another
    process -- see pytest_configure in test/conftest.py -- or a
    patch('src.display_manager.RGBMatrix') leaked from an earlier test module.

with WinError 10048 in the captured log directly above it.

No regressions: full suite with both changes is 115 failed / 4464 passed /
63 skipped, failure set identical to the pre-change baseline. The 115 is the
pre-existing Windows-environment baseline (os.geteuid, POSIX modes, fcntl);
CI on Linux remains authoritative.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 16:13:38 -04:00
ChuckandClaude Opus 5 59997594ac test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths

Two shared-state problems, both of which show up as a permanently dirty
checkout or an unreproducible test failure.

test_display_dirty_tracking.py builds a real DisplayManager, whose
_snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web
UI reads. Every pytest process on the machine shares that one file, so two
concurrent runs -- CI shards, a second worktree, an agent running the suite
alongside -- overwrite each other's snapshot and the mtime assertions stop
meaning anything. The module fixture now points it at a session-unique temp
path; the individual tests that care still override it further.

To be clear about what this does and does not fix: this is a real shared-path
hazard, but it is NOT the cause of the intermittent 15-test failure in that
module. That turned out to be the emulator's fixed TCP port, fixed in the
follow-up commit. This change stands on its own merits.

web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json
and data/operation_history.json as the web interface runs, into a directory
that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig
that ever opened the web UI -- and every test run that constructs the app --
left three untracked files behind and a permanently dirty `git status`. Only
data/.gitkeep is tracked under data/, so the negation keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop the emulator binding a fixed port, so concurrent runs can't collide

This is the cause of the intermittent full-suite failures we have been chasing:
runs of identical code landing anywhere between 100 and 130 failures, while
every implicated test passed in isolation.

Six test modules set EMULATOR=true and build a real DisplayManager. The repo's
emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to
serve the dev preview. That port is a machine-wide singleton, so a second pytest
process -- a CI shard, another worktree, an agent running the suite alongside --
loses the bind. RGBMatrix construction then raises, DisplayManager catches it and
falls back to `self.matrix = None`, and every test that subsequently touches the
matrix dies with

    AttributeError: 'NoneType' object has no attribute 'SwapOnVSync'

which names neither a port nor a socket, and points at the wrong file entirely.
Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its
matrix-touching tests fail together or not at all -- the 15-test swing that made
the totals look random.

Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process
and running test_display_dirty_tracking.py:

    without this change    15 failed, 6 passed
    with this change       21 passed

The "raw" adapter renders in memory and binds nothing. Only display_adapter is
overridden, in a throwaway config written per pytest process; the repo's
emulator_config.json is untouched and `run.py -e` still opens the browser
preview on 8888. Nothing in the suite referenced the adapter, and the tests
wrap SwapOnVSync on the matrix object itself, so they are indifferent to what
sits underneath. allow_adapter_fallback is forced off -- falling back would
land us on the browser adapter and its fixed port, which is the whole problem.

CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set
to an absolute path: the previous behaviour depended on where pytest was invoked
from, and silently wrote a default config into whatever directory that was.

Verified no regressions: full suite on this branch and with origin/main's
versions of the touched files, same machine, back to back -- 115 failed /
4347 passed on both sides, zero failures unique to either. That 115 is the
pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell
scripts, Linux-only binaries); CI on Linux remains authoritative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: mark the shell entry points executable

Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh`
fails with "Permission denied" and only works if you know to prefix `bash`.
That one matters most: the web UI's own error hint, added in #560, tells users
to run exactly that path when a system action fails for want of passwordless
sudo, and following that instruction verbatim did not work.

All eleven carry a shebang and are invoked directly, never sourced. The two
sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately
left non-executable, which is what distinguishes a library from an entry point.

Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied
with `git update-index --chmod=+x` because this checkout is on Windows, where
core.fileMode is off and the working-tree bit is not tracked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:45:13 -04:00
ChuckandClaude Opus 5 f6367d63ae security: triage the CodeQL backlog — 129 alerts, three of them live (#561)
* fix(web): escape quotes in every HTML escaper, not just & < >

The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:

    value="${escapeHtml(v)}"   title="${escapeHtml(v)}"

so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.

It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.

app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.

cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `&#39;`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.

url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).

test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): stop request-supplied names from reaching paths outside their base

Three of the py/path-injection alerts were live, not lint:

* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
  whose resolved path *string-prefixed* the plugin directory. Flask's
  default converter forbids a slash but not dots, and
  get_plugin_directory('..') returned the parent of the plugins directory
  because it exists -- so every file under the project root then prefixed
  that directory, config/config_secrets.json included. The prefix check
  was also wrong on its own terms: with plugin dir "plugin-repos/foo",
  "../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
  start with "plugin-repos/foo".

* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
  body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
  file_id of "../../../../etc/something" deleted that file. This is the
  one finding in the batch that destroyed data rather than exposing it.

* POST /api/v3/cache/delete passed the body's key through
  CacheManager.clear_cache to DiskCache, which joined it as a filename and
  called os.remove. Same shape, same result. The guard goes in
  DiskCache.get_cache_path, the single choke point get/set/clear share, so
  every caller is covered rather than just this route. Real keys are the
  stems of files already flat in the cache directory -- that is how
  list_cache_files derives them -- so nothing legitimate is turned away.

The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.

Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.

test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): refuse a plugin id that is not a plain name, don't truncate it

pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.

Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(web): say what the plugin web_ui iframe actually is

The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.

Also drops the now-unused os/os.path imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): inline url-input's scheme guard at the previewLink.href sink

CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.

Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.

Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): address CodeRabbit findings on the CodeQL triage PR

- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
  credential-saving or connect flow runs. NetworkManager only accepts
  printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
  was previously saved/attempted before nmcli itself rejected it.

- web_interface/blueprints/api_v3/config.py: fail closed when the
  plugin config schema path can't be resolved under the plugins
  directory (e.g. a symlinked plugin dir). Previously this fell
  through with secret_fields left empty, so submitted credentials for
  that plugin were saved as ordinary, unencrypted configuration.

- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
  splicing the JSON day/column key into an inline oninput="..." handler
  string. escHtml() escapes quotes for a normal HTML attribute, but the
  browser HTML-decodes the attribute before running it as script, which
  undoes that escaping and lets a crafted column name (e.g. from an
  uploaded JSON file) break out of the JS string and execute. Cell
  edits now travel through data-day/data-col attributes read by one
  delegated 'input' listener instead.

  While in this file: fixed 6 pre-existing missing-')' typos on
  multi-line safeSetHTML(...) calls (already flagged by Biome in this
  PR's own CodeRabbit run as syntax errors blocking its lint pass).
  These predate this PR (present on main too) but made the whole file
  fail to parse in any JS engine, which is a bigger problem than the
  XSS finding itself and directly touches the same lines.

Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:32:03 -04:00
ChuckandClaude Opus 5 5137e86d16 feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab

PR #544's change, ported onto the api_v3 package split (#553). Identical
behaviour; only the placement of the new code differs.

The original added 508 lines to web_interface/blueprints/api_v3.py, which #553
deletes, so every hunk of it would conflict irreconcilably. Ported by AST:
26 new top-level items sorted to where the split puts each kind --

  __init__.py   2 imports, 11 constants, 7 helpers
  starlark.py   4 routes  (/starlark/editor/{apps,status,start,stop})
  misc.py       2 routes  (/integrations/mqtt-bridge{,/config})

Everything outside api_v3.py -- the Tools partial, the installer scripts, the
JS tests -- applied unchanged.

Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is
regenerated to match, which is exactly what test_api_v3_url_map.py is designed
to make you do when routes are added.

Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR
ships could not be run here -- node is not installed on this machine -- so
test/js/dom/test_tools_sections.js is unverified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST

The AST-based port of #544 onto the api_v3 package split dropped `time`
from starlark.py's import list. start_pixlet_editor() and
stop_pixlet_editor() both call time.time()/time.sleep() directly, so
every start (NameError building `state['started_at']`) and every stop
that has to wait out the EXIT trap crashed with a 500. No test caught
it because the route's own tests mock subprocess.Popen but never
actually invoked it before now.

Also carries over #544's later fix that this port branched before:
env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an
operator who had already pinned PIXLET_EDITOR_HOST to loopback,
forcing the unauthenticated `pixlet serve` process onto the LAN
regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as
api_v3.starlark.py's siblings already do for _pkg-owned names.

Both fixes route the shared _pkg.time reference the rest of the
package's route modules already use for anything a test might need to
patch, rather than a bare `import time` local to this file.

Ported the existing regression test from #544
(TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's
module layout (web_interface.blueprints.api_v3.starlark instead of the
old monolithic api_v3 module), which is what caught the NameError.

Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this
branch and on origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to
this change).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): clear the six lint errors this rebase introduced

All six were introduced by rebasing this branch onto the merged blueprint
split, not by the split itself. Confirmed by diffing pyflakes output against
main with line numbers normalised -- everything else it reports is present on
main too and is the package's deliberate re-export pattern.

starlark.py used _STARLARK_APPS_DIR three times without importing it (F821).
The rebase resolved an import-list conflict as a union of both sides, and that
symbol was on neither side of the conflict hunk, so it was silently lost. It is
defined in __init__.py and is now imported like its neighbours. This was the
only one of the six that would fail at runtime rather than merely lint.

__init__.py imported contextlib twice (F811): the cherry-pick added one next to
the existing import. Removed the duplicate; the original at line 19 is used.

__init__.py imported signal purely to re-export it to starlark.py, so pyflakes
saw it as unused (F401). signal is stdlib and does not need routing through the
blueprint package, so starlark.py imports it directly and __init__.py no longer
does. contextlib stays re-exported because this module genuinely uses it.

_read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule
this module imports at the bottom for its route side effects (F811). Renamed to
`settings`, with a comment saying why, since the name is otherwise the obvious
one to reach for.

Verified: pyflakes now reports nothing on this branch that main does not, the
package imports, all nine route modules load, and 117 routes register, matching
the pinned URL-map snapshot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply

Two CodeRabbit findings on the bridge settings endpoint, both of which returned
200 while doing something other than what the caller asked.

`request.get_json(silent=True) or {}` turned a missing or unparseable body --
and the JSON literals null, [] and false -- into an empty dict, which then
satisfied the isinstance(data, dict) guard on the very next line. The guard was
there to reject exactly those bodies. Dropping the `or {}` lets None fail it.

The same `or {}` on /errors/clear is left alone: its docstring documents the
body as optional, so an absent body legitimately means "use the defaults". The
difference is that saving settings has nothing sensible to do with no body.

`if data.get('clear_password'):` accepted any truthy value, and the string
"false" is truthy in Python -- so a client echoing the field back as a string
wiped a password it meant to keep. Now coerced through the package's existing
_coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else.

test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus
a missing body, and clear_password across truthy and falsy spellings. Verified
against the unfixed code -- reverting the body guard fails 5, reverting the
coercion fails 3.

Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials
when TLS is off (CWE-319). That is a policy decision about the feature rather
than a defect -- unencrypted MQTT on a trusted LAN is common and often
deliberate -- so it is raised on the PR for a maintainer call instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: work through the remaining review findings on the editor and bridge

allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses
the network in cleartext. Refused now rather than merely warned about -- but
refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal
deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults
false, and is coerced like the other booleans so the string "false" cannot
switch the guard off.

starlark.py:796 -- the supported service runs Flask threaded, so two start
requests could each see running=False, each launch an editor, and the second
state write replace the first PID, orphaning a process that holds the display
down with nothing recording it. The check-launch-write sequence now takes a
module-level lock.

starlark.py:848 -- if the state write failed the route returned success with an
editor running and no PID recorded: status and stop both reported no session
while the display stayed down until the timeout expired. It now terminates the
process group and returns an error.

starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so
nothing hands the display back, yet the response said "the display is
restarting". After an escalation the display is now restarted explicitly, and a
failure to do so returns an error naming the manual step instead of a success.

pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no
timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the
failure lands before the display is stopped rather than after.

pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface
address are three cases, not two. Collapsing the last two printed a URL saying
"localhost" whenever PIXLET_EDITOR_HOST named a LAN address.

tools.html:1254 -- escHtml does not encode single quotes, and the app id was
interpolated into an inline onclick="startPixletEditor('...')", so a directory
containing an apostrophe could break out of the JS string and run script. The
handler binds with addEventListener and reads the id from dataset, where it is
only ever parsed as an HTML attribute.

Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in
in both directions. The tools DOM suite gains three guards asserting the edit
buttons carry no inline onclick and pass the id via dataset -- those need jsdom
and did not run here, so CI verifies them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): log the traceback on the editor state-write failure

The 848 fix answers 500 when the session state cannot be written, and logged
that at error level -- but without exc_info, so the traceback never reached the
log. test_web_error_detail.py guards exactly this: a handler returning 5xx must
write an error-level record *with* the traceback and return the sanitized
detail, because checking that merely something was logged is too weak.

Caught by Core unit tests on the previous commit, not locally: the guard parses
every module under web_interface/blueprints/api_v3 as one source, so it only
fires once the whole package is read together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 11:32:12 -04:00
ant456 a3d505384d Add render_width/render_height support to Starlark Apps (#552) 2026-09-11 11:22:52 -04:00
ChuckandClaude Opus 5 bdb9a94033 refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package

web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181
functions in one module -- 9% of the core by line count and three times the
next largest file. It becomes a package of nine route modules grouped by path
segment, plus __init__.py for the shared imports, constants, Blueprint and
helpers.

Every route module decorates the SAME api_v3 Blueprint object, so endpoint
names stay api_v3.<function>, the URL map is unchanged and app.py is untouched.
Verified: 111 routes before, 111 after, byte-identical rules, endpoints and
methods, and every endpoint still on the one blueprint.

  plugins   3,867   config    1,178   starlark  692   system  619
  fonts       452   misc        398   wifi      361   display 326   backup 212
  __init__  1,787 (imports, constants, Blueprint, 56 helpers)

Two things the URL-map check could not catch, both found by running the suite:

1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one
   directory deeper made that resolve to web_interface/ instead of the project
   root. Nothing failed at import; it surfaced as ~110 tests failing with 404s
   and "installation script not found", because every path built from it was
   one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts
   PROJECT_ROOT/run.py exists so the next move cannot repeat it.

2. Module-attribute patching. Tests do
   monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route
   module that binds such a name by value never sees the patch. The shared code
   therefore stays in __init__.py rather than moving to a _common submodule --
   it has to live on the module the tests patch -- and the eleven names tests
   patch are read back through the package (_pkg.X) instead of bound by value.
   Those eleven were found by AST-scanning every setattr in the test tree, not
   by guessing; "time" is among them, used to drive a fake clock through the
   second-resolution credential-backup filenames.

Test changes are confined to what genuinely moved: patch targets that now name
the owning route module, imports of helpers, and six tests that scan the api_v3
source as a file and now read the package directory.

Full suite: 4,278 passed, 68 skipped, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(api-v3): address CodeRabbit findings from the blueprint-split review

Fixes to the api_v3 package split (PR #553), one per finding verified
against the actual code:

- __init__.py: _redact_credentials only blanked scalar values under a
  credential-named key; a bare list of secrets under such a key (e.g.
  tokens: ["a", "b"]) passed through untouched, since the list branch
  recursed with no memory that its key looked like a credential. Nested
  dicts still walk normally (a documented, tested behaviour -- a container
  like secrets: {api_key: ..., note: ...} is a section name, not a value to
  blank outright), but any value reached under a credential-shaped key is
  now actually blanked.

- __init__.py: the OAuth helper script's raw stderr/stdout went to
  logger.error unredacted (CWE-532) right next to a comment claiming this
  was deliberate; the HTTP response already used the existing redact_text
  helper. Routed the log line through the same helper.

- __init__.py / starlark.py: the standalone Starlark manifest fallback
  (used when the plugin instance isn't loaded) read-modified-wrote
  manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe
  (plugin-repos/starlark-apps/manager.py), which already holds an flock for
  the same file when the plugin is loaded. Added _starlark_manifest_lock,
  mirroring that pattern, and wrapped every standalone read-modify-write
  call site in it. The app-config update route also wrote config.json and
  the manifest as two separate, non-transactional writes (a second,
  distinct finding at the same call site); config.json is now rolled back
  if the manifest write that follows it fails.

- backup.py: restore options used bare bool() on values from the request,
  so {"restore_secrets": "false"} restored secrets anyway (bool("false") is
  True). Switched to the existing _coerce_to_bool helper already used for
  this exact purpose elsewhere in the package.

- config.py: an automated import-rewrite mangled four user-facing
  validation strings and their neighbouring comments -- "Invalid start
  time" had become "Invalid start _pkg.time" (and likewise for "end time")
  in both the schedule and dim-schedule per-day validation paths.

- display.py: `import _pkg.time as time_module` -- _pkg is a local alias
  for the package, not a real importable module, so this raised
  ModuleNotFoundError whenever a caller restarted an already-running
  display service via /display/on-demand/start, after the on-demand
  request was already written to cache. Fixed to `import time`. Audited
  the rest of the package for the same `_pkg.<module>` import mistake;
  every other `_pkg.` reference is a legitimate attribute read-through
  (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken
  import statement.

- fonts.py: validate_file_upload's max_size_mb parameter is silently
  unused by that helper (it only checks filename/extension) -- the font
  upload route saved arbitrarily large files as a result. Added the same
  seek-and-check pattern already used for the sibling .star upload.

- wifi.py: two ad hoc, inconsistent bool coercions. POST
  /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled
  it. POST /wifi/radio's enabled/force parsing recognized real bool and
  some strings but not int 1/0 (1 is True is False in Python). Factored one
  small _parse_bool_ish helper local to this file and used it at all three
  sites.

Not changed: the "unknown/misspelled restore option keys default to True"
half of the backup.py finding -- the file's own comment documents that a
missing key deliberately means "restore everything," matching the
already-existing JSON-parse-failure guard a few lines above it; only the
bool-coercion defect was a real bug.

Added or extended regression tests for every fix, following each area's
existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed
on both this branch and origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to this
change) -- no new failures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5

* fix(api-v3): reject unknown restore option keys

CodeRabbit's review of the blueprint split (#553) asked that
POST /backup/restore reject option keys outside RestoreOptions'
known set. The follow-up commit fixed the bool("false")-is-True
bug with _coerce_to_bool but never added the key check: a typo'd
or renamed key (e.g. "restoreSecrets") is silently ignored by
opts_dict.get(key, True), so the flag stays at its True default
and secrets get restored despite the caller's request saying
otherwise -- with no indication anything was wrong.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb

* fix(api-v3): address CodeRabbit findings on the blueprint split

- _redact_credentials: blank scalar descendants of objects reached
  through a credential-owned list (e.g. tokens: [{"value": "secret"}])
  regardless of field name -- the existing name-based walk only
  protected direct dict values under a credential key, not list items.
- wifi.py: reject enabled/force/auto_enable_ap_mode values
  _parse_bool_ish can't recognize (400) instead of silently treating
  them as False, which could disable Wi-Fi or the radio itself.
- Starlark manifest locking: lock a stable manifest.json.lock sidecar
  instead of manifest.json itself, in both the standalone route path
  (_starlark_manifest_lock) and the plugin path
  (StarlarkAppsPlugin._save_manifest / _update_manifest_safe).
  manifest.json is replaced by an atomic rename on every write, which
  swaps in a fresh inode; a lock held on the old inode does not
  exclude a second locker that opens the path afresh right after the
  rename and gets the new inode, so two writers could race despite
  each holding "a lock". A sidecar that no write ever touches always
  resolves to the same inode for every locker.

Skipped as stale: the "serialize the complete manifest
read-modify-write" finding at api_v3/__init__.py -- every standalone
handler that calls _write_starlark_manifest is already wrapped in
_starlark_manifest_lock() on this branch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): re-check reconciliation findings by the reconciler's own rules

Both CodeRabbit findings on the merge commit, verified against the code first.

Major, plugins.py: the stale-findings filter derived its own notion of "in
config" and "on disk", and both were looser than the reconciliation module's.
set(load_config()) also contains system keys, the secrets-file keys load_config()
merges in, and non-dict values; and any directory holding a manifest.json
counted as installed even when that manifest does not parse. Either looseness
clears a finding that is still true -- and a secrets key read as a plugin is the
precise bug the filter exists to stop reporting, so reintroducing that asymmetry
while re-checking was the wrong way round.

The two extractions now live in state_reconciliation.py as config_plugin_ids()
and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys()
alongside. _get_config_state() and _get_disk_state() use them too, so there is
one definition rather than two that can drift. _get_disk_state() re-reads each
manifest for version/name after taking membership from the shared extractor;
that costs one extra small read per plugin on a path that runs once per boot.

Minor, the new test: the fixture assigned api_v3.config_manager and
api_v3.plugin_manager directly. Those live on a module-level blueprint
singleton, so the mocks leaked into every later test that imports api_v3 --
pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr,
which restores them. This is the same pollution class that made an earlier test
in this session break seven unrelated ones, so it is worth getting right.

Five cases added for the parity itself: a secrets key, a system key and a
non-dict value must not clear an "installed but missing from config" finding,
and neither an unparseable manifest nor a .standalone-backup- directory may
count as installed. All five fail against the looser version.

Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and
CodeRabbit all pass. Codacy reads action_required on every commit of this
branch including the first, so it is pre-existing and not from this work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 10:07:38 -04:00
ChuckandClaude Opus 5 0ab95586fb fix(web): say when a system action failed for want of passwordless sudo (#560)
* fix(web): say when a system action failed for want of passwordless sudo

POSTing reboot_system to a Pi returns, in full:

  {"message": "Action failed; see logs for details", "status": "error"}

The cause is that the web interface runs unprivileged, and its
systemctl/reboot/journalctl calls only work once
scripts/install/configure_web_sudo.sh has granted NOPASSWD. first_time_install.sh
never invokes that script and no user-facing doc mentions it, so on a fresh
device every privileged action fails -- start_display, stop_display, the
autostart toggles, reboot, and the log viewer.

That last one closes the loop: "see logs for details" is unreachable advice
when journalctl is refused for the same reason. This is exactly the failure
src/web_interface/error_handler.py's describe_exception() was written to break,
and /system/action's exception handler was still discarding the cause instead
of using the helper the module already imports.

Two changes, no behaviour change when things work:

- The exception path now returns 'details': describe_exception(e), matching how
  the other handlers in this blueprint already report.
- A failure whose stderr or exception text is sudo refusing to prompt ("a
  password is required", "no tty present", "a terminal is required") reports
  what to do about it, naming configure_web_sudo.sh. Unrelated failures keep
  the generic message and their stderr, so a missing unit is not blamed on
  sudo.

Granting the sudo rights is left alone deliberately: auto-running a script that
hands out NOPASSWD is a security decision for the maintainer, not something to
slip into an installer. Making the refusal legible is the part that is
unambiguously an improvement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): apply the sudo hint on the on-demand start_display path too

start_display with a mode builds its own response and returns before the shared
nonzero-result path, so a recognized sudo refusal there reported only "Failed to
start display" and said nothing about the passwordless sudo that refused it --
the exact gap the rest of this PR closes everywhere else.

Raised by CodeRabbit on #560 and verified against the code before fixing: the
branch at api_v3.py:2058 does return early past the shared handler.

Three regression cases: the on-demand branch reports the sudo cause, keeps its
"Display started" message on success, and does not blame an unrelated failure on
sudo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:45:40 -04:00
ChuckandClaude Opus 5 39f27d285d fix(plugins): stop reconciliation inventing plugins and telling users to delete real config (#557)
On a device running four installed, configured, working plugins, the overview
banner read:

  Stale plugin config entries found: football-scoreboard, odds-ticker, data,
  ledmatrix-weather, starlark-apps. Remove them from config.json or reinstall
  via the Plugin Store.

Every claim in that sentence was wrong, and following its advice would have
deleted 4.9KB of working league settings. Four separate defects combined.

1. Secrets keys became phantom plugins. load_config() merges
   config_secrets.json into the config it returns, and the ignore list named
   only 'github' and 'youtube'. A 'data' key in that file therefore read as a
   plugin id and was reported as "in config but not on disk" forever. Read the
   secrets file's own top-level keys instead of hardcoding two of them.

2. The auto-fix clobbered real config. The handler for "on disk but not in
   config" assigned `config[plugin_id] = {'enabled': False}` unconditionally,
   so whenever detection was wrong it replaced a plugin's entire configuration
   with a stub. On the reported device it only failed to do so because the
   write hit EACCES. Now it refuses to overwrite an entry that already exists.

3. The banner gave backwards advice. plugin_missing_in_config ("on disk, not in
   config") and plugin_missing_on_disk ("in config, not on disk") are opposite
   problems, and both were rendered as "stale config entries ... remove them
   from config.json" -- which is correct for the second and destructive for the
   first. They are now reported separately, each with the advice that fits.

4. A stale verdict was served indefinitely. The result is a snapshot written
   once per run to a status file, and a run that fails to apply a fix also
   declares it will not retry. A condition that had since resolved kept being
   reported for hours. The status endpoint now re-checks stored findings
   against current state, dropping only what it can prove stale and keeping
   any kind it cannot re-verify.

The secrets-key lookup is deliberately fail-safe: an unreadable, absent,
malformed or non-path secrets location narrows the ignore set rather than
raising. An earlier revision let TypeError escape, which the broad handler in
_get_config_state() swallowed as "Error reading config state" -- emptying the
config state and making every downstream detection wrong. The existing
reconciliation tests caught it; there is now a regression test for it too.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:45:21 -04:00
ChuckandClaude Opus 5 aba96e25b3 chore: delete three functions nothing calls (#550)
src/base_classes/baseball.py       _get_baseball_display_text   45 lines
  src/web_interface/api_helpers.py   validate_request_params      22
  web_interface/blueprints/api_v3.py _validate_time_range         14

Each has exactly one occurrence across both repositories -- its own
definition. No decorator, no __all__, no getattr dispatch, nothing in
templates or JavaScript.

A fourth candidate was dropped after checking: _unshare_element_fonts in
src/common/sports_shared.py looked unreferenced, but eight scoreboard plugins
call SportsCore._unshare_element_fonts directly from their
test_element_text_colors.py, plus their own copies at runtime. It is live API.
The earlier reading came from a plugins checkout 84 commits behind main, which
is a good argument for re-verifying this kind of claim against a fresh tree
rather than trusting an earlier scan.

Full suite: 4,265 passed, 68 skipped.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:45:07 -04:00
ChuckandClaude Opus 5 577f5501a6 perf(plugins): stop re-deriving a display() signature the caller already cached (#549)
display_controller resolves once, and caches, whether a plugin's display()
takes a display_mode keyword -- self._plugin_accepts_display_mode, populated
right before the dispatch. It then handed the executor a
types.SimpleNamespace wrapping a closure, and execute_display() ran
inspect.signature() on that to work out the same thing.

Because the SimpleNamespace is rebuilt per call, the callable was new every
time, so nothing inside the executor could ever cache it either. Measured at
~39us per dispatch on a Pi 4, for a value the caller had a line earlier.

execute_display() now takes accepts_display_mode, falling back to inspecting
only when a caller does not pass it, so existing callers are unaffected.

Also documents two things that read as bugs and are not:

- execute_with_timeout()'s timeout is advisory. Nothing cancels the thread --
  Python cannot -- so on expiry the operation runs to completion in the
  background and only the caller gives up. A permanently hung plugin leaks a
  daemon thread per attempt. This is why callers holding a lock across the
  call must release it from inside the wrapped callable, as run()'s
  _release_display_lock already does.

- Only the first display() of each mode goes through the executor; the
  per-frame loops call display() directly. That is deliberate: a thread per
  frame would cost more than an advisory timeout buys. Both loops now say so,
  so the asymmetry does not read as an oversight.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:42:49 -04:00
ChuckandClaude Opus 5 dcd6e39c96 fix(web): report real disk usage and MemAvailable on the live status stream (#558)
The SSE status stream sent 'disk_used_percent': 0 as a literal, so every
consumer of the live view showed 0% disk no matter how full the card was.
/api/v3/system/status computed it correctly; the stream that the dashboard
actually watches did not. On a Pi with a modest SD card that is the warning a
user most needs, and it was guaranteed to never appear.

The stream also omitted memory_available_mb. /api/v3/system/status carries it
with a comment spelling out why it matters: MemAvailable accounts for
reclaimable page cache, so it is what separates a board reading 70% "used" that
is fine from one reading 70% that is about to fail fork(). A 1GB Pi 3B+ can sit
at either. The number that predicts the failure was missing from the live view.

An unreadable disk now reports None rather than 0. The UI already renders null
as '--'; a confident 0 reads as "plenty of room", which is worse than a blank.

Metric collection moves to web_interface/system_metrics.py, with no Flask or app
imports. That is not cosmetic: importing web_interface.app constructs the Flask
application and a CacheManager, and the latter claims the cache directory with a
cleanup thread. The first version of these tests imported the generator directly
and broke test_cache_cleanup_thread_ownership ("one thread per directory") plus
four starlark route tests through that side effect. Reading a CPU percentage
should not boot a web application, and testing it should not either.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:42:30 -04:00
ChuckandClaude Opus 5 ad5bc4b819 perf(sports): LRU-bound the decoded logo cache (#559)
SportsCore._logo_cache was a plain dict keyed by team abbreviation with no
eviction. Its entries are not file bytes but decoded RGBA thumbnails sized to
display*1.5 -- roughly 36KB on a 256x64 panel, more for wide wordmarks -- and
assets/sports/ncaa_logos ships 307 of them. A plugin that walked a full league
held the whole league resident: about 11-18MB per manager instance, and a league
runs three (live/recent/upcoming) that each keep their own cache, so the same
logos were duplicated across them.

On the 1GB Pi 3B+ this was measured on, one board was sitting at 439MB resident
with ~290MB available, so tens of megabytes of duplicated league logos is real
money. Bounded to 64 entries, which holds a full "other games" cycle (on the
order of 20 games, 40 teams) without thrashing while capping the cache well
below a 307-team league.

Eviction is LRU rather than clear-when-full, using the OrderedDict/popitem
pattern the neighbouring caches in this codebase already use (_IMAGE_CACHE_MAX,
_FIT_CACHE_MAX, _TEXT_WIDTH_CACHE_MAX). That ordering matters: the logos on
screen right now are precisely the ones that must not be discarded, so a cache
hit moves the entry to the end.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:42:18 -04:00
ChuckandClaude Sonnet 5 fb3b293ace fix(plugins): let a plugin ask to be polled faster while it has live content (#555)
* fix(plugins): let a plugin ask to be polled faster while it has live content

Reported: "the football plugin with live games only updates the live game in
progress if I restart the display."

The data path was never the problem. NFLLiveManager fetches ESPN with no cache,
SportsLive.update() refreshes current_game in place when the game IDs are
unchanged, and the scorebug redraws from the game dict every frame -- which is
why the reporter's logs look healthy.

The problem is cadence. _get_plugin_update_interval() read only the manifest's
static update_interval, football's manifest pins that to 60, and the plugin's
own live_update_interval (15s) was invisible to the scheduler. Measured on a rig
during the fourth quarter of the game in the report:

    23:21:49  23:22:50  23:23:50  23:24:50  23:25:50   <- exactly 60s apart

A clock and score up to a minute stale during a two-minute drill reads as a
frozen panel, and a restart is the one moment it is ever current.

A single static number cannot say "every 15 seconds while a game is on, every 15
minutes in July", and only the plugin knows which is true. get_update_interval()
lets it say so per tick; returning None means "no opinion" and the existing
manifest/config resolution applies, so every plugin that predates this is
unaffected.

Requests are clamped to MIN_DYNAMIC_UPDATE_INTERVAL (5s): a plugin returning 0
would otherwise be re-entered on every tick of the render loop, busy-waiting
against its own API. A hook that raises or returns a non-number is ignored
rather than propagated -- a scheduler that fails on one plugin's bug stops
updating all the others.

Deliberately NOT changed: the manifest still beats config in the static path.
That looked like the obvious fix -- user config being silently ignored -- until
checking a real rig, where football and baseball both carry update_interval 3600
in config against a manifest 60, and weather 1800 against 60. Those values are
stale precisely because nothing has been honouring them; making config win would
have slowed three plugins by 60x, turning a one-minute lag into an hour. The
dynamic hook makes the flip unnecessary. There is a test pinning the current
precedence with that reasoning attached.

Full suite: 4,283 passed, 68 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(plugins): drive the real scheduler, not just the interval resolver

test_plugin_dynamic_update_interval.py asserts that
_get_plugin_update_interval() returns the number the plugin asked for. That is
not the same claim as "the plugin gets updated more often", and the gap between
those two is exactly where the original bug lived: the plugin knew it wanted
15s, said so in live_update_interval, and nothing downstream acted on it.

So this ticks the real run_scheduled_updates() through a simulated hour and
counts dispatches. Against pre-fix core it reports "10 updates in 10 minutes of
a live game" -- the 60s manifest cadence, matching what was measured on a rig
during the reported game. Against the fix it reports ~40.

Also pins the regression that would be worse than the bug: an idle hour must
still be ~60 updates, not 240. Asking for the live interval year-round would
poll ESPN four times a minute all summer.

Scope note, since it is easy to over-read this fix: the *switch* display path
already refreshed the manager immediately before drawing, via
_try_manager_display() -> _ensure_manager_updated(), which honours the manager's
own 15s interval. So a switch-mode card was already <=15s stale at draw time
before this change. What this fixes is the background cadence, which is what
live-priority detection, Vegas content and scroll preparation all read.

Full suite: 4,288 passed, 68 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): reject bool and -inf hook results in dynamic interval

get_update_interval() ran bool through float() (bool is an int subclass,
so True/False became 1.0/0.0) and only checked for +inf, not -inf. Both
cases landed on the MIN_DYNAMIC_UPDATE_INTERVAL floor by coincidence
instead of falling back to the static/manifest interval as invalid
input should.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Tst9cied2ri9bH4QRWa6H

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:58 -04:00
ChuckandClaude Opus 5 8da13f02f8 chore: ignore team logos fetched at runtime (#551)
logo_downloader.py and LogoHelper write into assets/sports/<league>_logos/
whenever a plugin meets a team whose logo is not on disk. Those directories are
also tracked -- 209 NCAA logos and 153 soccer ones ship with the repo -- so
every rig accumulates untracked files nobody intended to commit. This checkout
had 62; hdpi shows the same.

The cost is not the files, it is that a permanently dirty `git status` trains
everyone to ignore the one signal that says a checkout is not what you think it
is. That is how a stale tree sat unnoticed on a rig for hours until a restart
surfaced four sports plugins that could no longer import.

Ignoring a directory does not untrack what is already in it, so the logos that
ship keep shipping -- verified: 209 and 153 still tracked, no deletions in the
diff. Only new downloads are hidden.

Adding a logo on purpose stays possible and is what the escape hatch in the
comment documents. It is also rare: the last deliberate addition was #415, four
named NCAA logos a plugin needed, and `git log` finds no other in a year. So the
common case is noise and the rare case is explicit, which is the right way round.

Untracked files: 62 -> 0.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:44 -04:00
ChuckandClaude Opus 5 28bc79566f fix(logo): remember a missing logo instead of re-warning every rotation (#548)
* fix(logo): remember a missing logo instead of re-warning every rotation

load_logo() stat'd the path and logged a WARNING on every call, and the
positive cache never covered it because a miss returns None and caches
nothing. A file that is simply not there therefore produced one warning per
rotation for as long as the process ran -- measured on a live rig at 114 lines
in 24 hours for a single missing ticker icon, for a file nobody was going to
add.

Misses are now remembered for 10 minutes: warn once, then return None without
touching the disk. Bounded rather than permanent because logo_downloader
writes logos at runtime, so a file that appears later must still be picked up
without a restart. Downloads through load_logo_with_download() clear the entry
outright -- load_logo() consults the miss record before it stats the disk, so
without that a freshly downloaded logo would stay invisible for the whole
window.

This is in the core rather than in ledmatrix-stocks, where it was found, so
every plugin that goes through LogoHelper gets it.

_cache_order stays a list. Swapping the pair for an OrderedDict would shave an
O(n) scan per cache hit, but n is capped at cache_size (100 by default) and
test_logo_helper.py pins the current structure; not worth the churn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(logo): make the miss TTL longer than the rotation it is meant to outlast

Deployed the previous commit to a live rig and measured it: no change at all.
"Logo not found for VOO" stayed at ~6 lines an hour, exactly the baseline.

The TTL was 600s and the display rotation is ~618s, so every recheck expired
just as the plugin came round again and the negative cache never once got to
suppress a warning. The fix was correct in shape and useless in practice,
which only measuring on the rig would show.

An hour instead. That is safe because the TTL is not the main way an entry
clears: load_logo_with_download() drops it the moment a download succeeds and
clear_cache() drops all of them. The TTL only covers a file that appeared some
other way -- someone copying one in by hand -- and waiting up to an hour for
that, or restarting, is a fair trade for not re-warning about a file nobody is
going to add.

The general lesson is in the comment: a TTL has to be long relative to the loop
that does the asking, not merely "a while".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:32 -04:00
ChuckandClaude Opus 5 12f3790994 fix(install): render the systemd units from their templates, not from heredocs (#547)
* fix(install): render the systemd units from their templates, not from heredocs

The installers carried their own inline copies of units that also exist as
templates under systemd/, and the copies drifted.

install_service.sh renders ledmatrix.service from the template correctly, then
wrote ledmatrix-web.service from a heredoc that predated it -- missing
Wants=network-online.target, RestartSec=10, SyslogIdentifier, CacheDirectory,
CacheDirectoryMode and Environment=USE_THREADING=1. install_web_service.sh had
a third copy, and install_wifi_monitor.sh a fourth, that one already differing
from its template (syslog where the template says journal).

startup_validator.py compares the installed unit against the template, so a
rig installed this way warned on every boot -- and the remedy the warning
names, "re-run scripts/install/install_service.sh", reinstalled the same stale
copy. The warning could never clear. Reproduced on a live rig running exactly
that unit.

All three installers now render systemd/*.service through the same placeholder
substitution. The template gains a __USER__ placeholder rather than hardcoding
User=root, because the web interface runs as whoever installed it.

That last point was a second, independent cause of a permanent warning: the
validator substituted a fixed "root", so any non-root install reported drift
forever. It now reads User= from the installed unit -- an install-time
decision, not something the template dictates -- and compares everything else
strictly. first_time_install.sh already reads the installed User= the same way.

Tests cover a non-root web unit not warning, a genuinely changed directive in
that unit still warning, the User= fallback, and a grep-based guard that no
installer under scripts/install/ contains an inline unit body. That guard is
what found the install_wifi_monitor.sh copy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(install): escape sed replacements, use mktemp, and make render failures fatal

Address CodeRabbit findings on install_service.sh, install_web_service.sh and
install_wifi_monitor.sh:

- Values interpolated into each script's sed expression (project root path,
  username) were not escaped, so a value containing &, \ or the | delimiter
  would corrupt the rendered systemd unit. Add a shared
  sed_escape_replacement() helper in the new scripts/install/lib_systemd_render.sh
  (sourced by all three scripts) and apply it to every sed replacement.
- install_service.sh rendered the main and web units to the predictable path
  /tmp/ledmatrix.service.tmp before installing them -- a symlink/TOCTOU race
  (CWE-377). Use mktemp for both, with a trap to clean up on exit.
- install_service.sh treated a missing template as a mere warning and then
  checked only whether a unit already existed at the destination before
  enabling/starting it, so a render failure could silently fall back to
  enabling a stale, previously-installed unit. Both unit blocks now exit
  non-zero on a missing template or a failed render.

Also rename the ambiguous loop variable `l` to `line` in
test/test_systemd_unit_drift.py (Ruff E741); ruff isn't wired into any CI
workflow in this repo today, so this isn't currently CI-blocking, but the
rename is trivial and correct regardless.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5

* test(install): cover sed_escape_replacement against sed-special characters

CodeRabbit asked for regression coverage using a project path containing an
ampersand; the earlier commits on this branch already fixed the escaping,
mktemp usage, and enable/start-on-fatal-render-failure findings, and the
l->line rename was already applied -- this closes the one remaining gap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:41:19 -04:00
ChuckandClaude Opus 5 2df273ecfc fix(web): default web_display_autostart to true, as the installer already does (#556)
start_web_conditionally.py read the flag with
`config_data.get("web_display_autostart", False)`, so a config that simply
lacked the key got no web interface. Both config/config.template.json and
first_time_install.sh ship the key as true, so the code default contradicted
the shipped default in two places: absence means an older or hand-edited
config, not a request to stay down.

The failure mode was silent in the worst way. The "not starting" path exits 0,
so `systemctl status ledmatrix-web` reported the unit as successfully started
while nothing was listening on the port, and the only trace was one journal
line saying the flag was "false or not set" -- which reads as a deliberate
setting rather than a missing key.

Also start the web interface when config.json is missing or unparseable,
instead of exiting. The web interface is how a config gets created and
repaired, so a broken config is exactly when the user needs it most; leaving
it down means there is no way back in. Only an explicit false/off disables
autostart now, and the disabled message says "explicitly disabled" so the
journal distinguishes a real setting from a default.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 17:42:00 -04:00
ChuckandClaude Opus 5 a29c84208e fix(scroll): advance whole pixels per frame, not per wall-clock second (#545)
* fix(scroll): advance whole pixels per frame, not per wall-clock second

Smooth motion is not a frame-rate property, and measuring it as one is why
this survived three rounds of fixes. odds-ticker's frame timing is excellent
-- 100.0 fps, 10.00ms median, 0% stalls, worst in-scroll frame 19.95ms -- and
it still visibly stuttered.

What the eye judges is whether the strip advances the same number of whole
pixels on every presented frame. update_scroll_position derived position from
scroll_speed * delta_time and get_visible_portion truncated it with int(), so
jitter in delta_time decided which side of a pixel boundary the position
landed on. The live windows show why that matters: a rock-steady 100.0 fps
whose individual frames still range 5.6ms to 15.2ms, which at 100 px/s is
0.57px to 1.44px of movement.

Run the measured frame times through the real helper and 5.8% of frames
advance 0 or 2 pixels instead of 1 -- about six hitches a second. A frame that
moves nothing followed by one that jumps two is exactly what micro-stutter
looks like.

It is worst at a crisp speed, which is the part that stings: at 100 px/s on a
100Hz panel the accumulator sits exactly on integer boundaries, so
sub-millisecond jitter flips it either way and the motion beats at around
50Hz. Snapping to the crisp ladder fixes the average and the wall clock then
throws away the per-frame uniformity the ladder was bought for.

So when scroll_config snaps to a crisp speed it now also puts the helper in
fixed-step mode: each presented frame advances exactly pixels_per_frame and no
clock is consulted. 100% of frames move by the same amount, whatever the
jitter.

This is only correct because SwapOnVSync blocks until the panel has taken the
frame, which makes the frame count a truer clock than time.time(). Before the
swap was locked to vsync it would have run at whatever speed the loop spun at.
Related: frame-based mode used to step discretely and was converted to
elapsed-time accumulation earlier in this series, because its threshold
comparison flipped on jitter. That was right for the code as it stood -- but
it treated the symptom, replacing a broken discrete step with a smooth-looking
accumulator instead of asking why a wall clock was involved at all.

Non-crisp speeds keep pacing off time, and set_scroll_speed() clears the fixed
step so a legacy caller changing speed is not silently ignored.

Trade-off worth naming: speed is now tied to the presentation rate rather than
to real time. If the loop cannot keep up with the panel the scroll runs slow
rather than jumping to catch up. That is the better failure -- uniform motion
at a slightly wrong speed beats correct average speed with a hitch six times a
second -- and a loop that cannot hit the resolved rate is a measurement
problem for the crisp ladder, not something to paper over with uneven steps.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(scroll): make the time-based pin actually pin something

Review caught that test_time_based_stepping_is_what_it_replaces could pass
against perfectly uniform motion, and it was right.

update_scroll_position sets last_update_time on its way through, so the very
first call sees a delta_time of zero and moves nothing in time-based mode.
_advances counted that synthetic frame, which put a guaranteed zero in every
histogram -- enough on its own to satisfy "uneven > 0". The test asserting the
defect exists would have passed after the defect was gone.

The first call is now primed and discarded, and the assertion is a proportion
rather than "more than zero": against these frame times the old path misses
roughly one frame in twenty, so 1% is well below the real rate and far above
anything a stray frame could produce.

Re-measured with the artefact removed, the numbers in the PR description are
unchanged: 5.85% of frames uneven before (114 zero-advance and 120 double
frames in 4000), 0.00% after.

Also fills in the docstrings the review flagged: everything in the new test
file, plus three pre-existing one-liners in scroll_config that the diff
touched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:25:45 -04:00
ChuckandClaude Opus 5 1198615d19 test: install PyYAML so the starlark route tests can load the plugin (#546)
Tests has been red on main since #535. All 13 failures in
test/web_interface/test_starlark_pixlet_routes.py are the same
ModuleNotFoundError: No module named 'yaml'.

The test loads plugin-repos/starlark-apps/tronbyte_repository.py by path --
deliberately, "the way the blueprint does", since the core web blueprint
really does exec that plugin module -- and the plugin imports yaml.

Nothing is undeclared. The plugin's own requirements.txt already pins
PyYAML>=6.0.2, and on a real rig the plugin store installs it. CI installs
only requirements.txt and requirements-test.txt, so a core test that reaches
into a plugin gets none of the plugin's dependencies.

PyYAML goes in the test requirements rather than the core ones because it is
not a core dependency: nothing in src/ or web_interface/ imports yaml. This is
the same shape as the psutil entry directly above it -- a package the core does
not require, installed so a test can exercise a real path instead of a stub.

Verified locally: with yaml available the file goes from 13 failures to 64
passing. (One unrelated failure remains on Windows only, where os.geteuid does
not exist.)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:25:30 -04:00
ChuckandClaude Opus 5 f8e2e89edc refactor(sports): put the scoreboards on the shared scroll resolver (#542)
* refactor(sports): put the scoreboards on the shared scroll resolver

Eight sports scoreboards -- afl, baseball, basketball, football, hockey,
lacrosse, nrl, soccer -- scrolled through this module's own pacing while the
other eleven scrolling plugins went through src/common/scroll_config. Two
implementations of the same job, and this one was on the losing side of every
difference.

It never called set_scrolling_state. Two consequences, both of which this
release's work was about:

- The frame hold is applied through that call, so a speed the crisp ladder
  could render in whole pixels still presented a new frame every refresh.
- Core only runs deferred updates while nothing is scrolling. Believing
  nothing was, it ran blocking work in the middle of these scrolls.

The default is non-crisp today: scroll_speed 50.0 with scroll_delay 0.01 is
50 px/s, which on a 100Hz panel is half a pixel per refresh. That cannot
render as motion -- it alternates 0px and 1px steps and judders at a 50Hz
beat, on every scoreboard, out of the box. Resolved through the ladder it
stays 50 px/s and holds each frame for two refreshes: same speed, whole-pixel
motion.

The stepping disagreement that used to justify a separate module is gone.
scroll_config avoided frame-based mode because it stepped on a wall clock at
1/scroll_delay with scroll_delay set to the frame period, so the decision sat
on its own threshold and flipped on sub-millisecond jitter. That branch now
accumulates elapsed time, identical arithmetic to the time-based one, so the
two differ only in the units the speed arrives in.

What is NOT shared, and must not be: the two modules read identically-named
keys with different meanings. Here scroll_speed is px/SECOND and scroll_delay
only converts to px/frame; in scroll_config scroll_speed is px per STEP, so
px/s is speed/delay. Passing this module's settings dict to the resolver turns
50 px/s into 5000, clamped to 500 -- a tenfold speed-up everywhere. So
_get_scroll_settings keeps sole ownership of reading sports config, including
the league merging, and hands the resolver a plain px/s. A test pins that
specific number, because it is the mistake the refactor invites.

MIN/MAX_PIXELS_PER_FRAME are gone; the resolver bounds speed and the helper
clamps FPS. _resolve_target_fps stays, re-purposed: under the old model that
key was the rate frames were presented at, so it is the faithful translation
into the refresh the ladder is computed against, used when no hardware
refresh is configured.

Speed changes for panels that are not 100Hz: 50 px/s becomes 60 at 60Hz
(+20%) and 48 at 120Hz (-4%). At 100Hz it is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): drop the frame hold when a scroll times out, not just when it says so

set_scrolling_state(False) clears the hold. The other way a scroll ends is
is_currently_scrolling() deciding, after scroll_inactivity_threshold of
silence, that it is over -- which is what happens when the rotation moves on
mid-scroll or a plugin is torn down. That path cleared the flag and kept the
hold, so every later plugin, scrolling or static, was presented at refresh/N
by whoever scrolled last, until something called the explicit stop.

The method's own docstring already states the rule this breaks: the hold "must
not outlive the scroll that asked for it". The timeout was the exception it
did not cover.

Pre-existing, but reachable by three plugins before and eleven after the
sports scoreboards moved onto the shared resolver, so it belongs with that
change. The test ages the activity timestamp past the threshold rather than
sleeping.

Also adds scripts/sports_scroll_check.py. The sports scroll path is per-league
opt-in, so a rig showing static game cards never constructs a
SportsScrollDisplay and none of its pacing can be observed from a normal run
-- which is exactly what happened when this change was first put on hardware:
26 minutes, zero sports scroll lines. The script drives the path directly with
synthetic games and asserts the three things the resolver is meant to buy: the
speed lands on whole pixels, the hold is published, and it is released after.
It never starts or stops the display service, matching scroll_speeds.py, so a
crash here cannot leave the panel dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): refuse to grab the panel while the display service has it

The module docstring already said to stop ledmatrix first. Nothing enforced
it, and running the script against a live service is not a harmless mistake:
rpi-rgb-led-matrix configures GPIO directions and the hardware PWM inside
RGBMatrix(), and when the root check fails it calls exit() from C with no
cleanup. The service keeps rendering and swapping onto pins that have been
reconfigured underneath it, so the panel goes black while every diagnostic
says the display is healthy -- fresh framebuffer, every pixel lit, "RGB Matrix
initialized successfully", nothing in the log. A restart fixes it, once you
work out that is what happened.

Found the hard way: this is what took the panel down on the test rig, not the
change the script was written to verify.

--fallback skips the check, since it never opens the matrix. --force is there
for anyone who means it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(scripts): annotate the subprocess call the way this repo already does

Codacy fails a PR on one new issue, and bandit B404 fires on any subprocess
import. scripts/run_plugin_tests.py carries the same suppression with the same
justification -- list-form argv, no shell -- so this follows it rather than
inventing a second convention.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-09 16:13:04 -04:00
ChuckandClaude Opus 5 4423ec33d5 feat(plugins): search, filter and sort for Installed Plugins, on a shared ListFilter helper (#540)
* feat(plugins): add search, filter and sort to Installed Plugins, on a shared helper

The Installed Plugins grid had no way to narrow it down: no search, no way to
see only what's enabled, disabled, or out of date. On a rig with a couple dozen
plugins that means scrolling the whole grid to find one.

The two sections below it already solved this, twice, independently — the
Plugin Store and Starlark Apps carried a copy-paste fork of the same ~600 lines
(filter state, apply-filters-and-sort, page renderer, pagination strip,
active-filter badge, listener wiring). Rather than add a third copy, this
extracts the shared machinery and builds the new toolbar on it.

New: web_interface/static/v3/js/plugins/list_filter.js — ListFilter.create()
owns debounced search, filter axes, sort, the active-filter count, Clear, and
optional pagination/persistence. Callers keep their own card markup via a
`render` callback. Three control types cover every axis the page uses: pills
(new), select (store category, starlark author) and cycle (the tri-state
All -> Installed -> Not Installed button).

Installed Plugins gets a compact toolbar: search box, one-click All / Enabled /
Disabled / Updates pills, and a sort dropdown (A-Z, Z-A, updates first,
recently updated, category). Filters reset on load, so you never come back to a
mysteriously short list. No new CSS — this is the first consumer of the
.filter-pill rules already sitting unused in app.css.

renderInstalledPlugins() is split so it still publishes canonical state while
renderInstalledCards() draws only the visible subset; the filtered list is
never assigned to window.installedPlugins, which the toggle handler,
isStorePluginInstalled(), runUpdateAllPlugins() and the Alpine config tabs all
read as their source of truth. Toggling a plugin while filtered pins its card
so it doesn't vanish from under the cursor.

The Store and Starlark migrations are behaviour-preserving: same element ids,
same localStorage keys (storeSort/storePerPage, starlarkSort/starlarkPerPage),
same tri-state button markup, same pagination. Verified by differential tests
that run the old and new implementations side by side against identical
fixtures and compare every observable after each interaction. The only visible
change is the pagination attribute (data-store-page/data-starlark-page ->
data-list-page), which nothing outside its own click handler referenced.

Net -156 lines in plugins_manager.js while adding a feature.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): keep raw search text, and stop the store search refetching

Two review findings from CodeRabbit on #540.

Do not write the trimmed search value back into the input. setSearch() trimmed
before storing, and syncControls() then copied that trimmed value back over what
the user had typed. Pausing longer than the debounce after typing a space
deleted the space (and reset the caret), making multi-word terms effectively
untypable. The raw text is now kept alongside the trimmed one: filtering and
activeCount() still use the trimmed value, while the input keeps exactly what
was typed.

Remove the legacy #plugin-search / #plugin-category listeners in
initializePlugins(). They bound searchPluginStore as the event handler, so the
DOM event arrived as its `fetchCommitInfo` argument — always truthy, which
skipped the cached-filter fast path and refetched /api/v3/plugins/store/list
with commit info on every keystroke burst and category change. The store's
ListFilter controller already filters the cached list, which is what those two
controls should do. This double-binding predates this PR (the old code guarded
with _listenerSetup and _storeFilterInit, two different flags, so both sets
stayed live); it is fixed here because the refactor owns that wiring now.

Both fixes are covered by tests that fail without them: the trailing-space
regressions in the installed-plugins DOM suite, and a new whole-file jsdom test
that counts fetches while typing (1 request at init, 0 thereafter; previously
1 -> 2 -> 3 -> 5).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): build pagination via DOM APIs, drop computed member access

Addresses the five Codacy security findings, all in list_filter.js.

Pagination no longer assembles an HTML string (3 findings: 2 critical + 1 high,
"unsafe assignment to innerHTML"). The interpolated values were only page
integers and local class constants, so there was no injection path, but
concatenating markup into innerHTML is the pattern the scanners flag and
createElement is no less clear. Each button now also owns its click listener
directly instead of the container being re-queried afterwards, and the strip is
cleared with textContent = '' rather than by assigning empty markup. No
innerHTML assignment remains in the file.

haystack() now walks Object.entries(item) and keeps the configured fields,
instead of reading item[field] per field ("generic object injection sink").
Field order no longer drives the haystack order, which is irrelevant to the
substring test. matches() iterates controls with for...of instead of an index
("variable assigned to object injection sink").

The rendered pagination is unchanged: same buttons, labels, page numbers,
disabled states and classes. The old-vs-new differential tests now compare
pagination structurally (tag, text, page, disabled, sorted class list) rather
than as an HTML string, since building nodes legitimately serialises
differently — «/» as characters rather than &laquo;/&raquo;, disabled="" rather
than a bare attribute. That comparison is stronger than the string one it
replaces, and the real-DOM suite still drives the actual page buttons.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): keep configured field order when building the search haystack

The previous commit swapped item[field] for Object.entries(item) to clear a
static-analysis object-injection warning, and in doing so changed the order of
the haystack: entries follow the object's own key insertion order, not the
configured `fields` order. Since the values are concatenated, that order decides
which values end up adjacent, so a multi-word query spanning a field boundary
matched differently. For store fields [name, description, author, id, ...] and
API objects keyed {id, name, description, author, ...}, "bob plugin-01" matched
before and stopped matching after.

That contradicted the behaviour-preservation claim for the store and starlark
migrations, and the differential tests missed it because every fixture query was
a single word.

Values now come out of a Map built from Object.entries, iterated in `fields`
order: the original haystack is restored, and there is still no computed member
access for the analyser to flag.

Regression coverage for the ordering itself, at both levels:
  - unit: phrases spanning name->id and category->tags, plus the reverse
    (object-key) order asserted NOT to match
  - differential: the same class of query compared old-vs-new, with a guard that
    the phrase actually matches something so a mutual zero-result cannot pass
    vacuously

Verified both fail without this fix (3 unit, 2 differential) and pass with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(web): add JS suites for ListFilter and the plugin-manager grids

No JS toolchain exists in this repo, so these are plain node scripts with no
framework: each prints ok/FAIL lines and exits non-zero. `node test/js/run_all.js`
runs everything, skipping the DOM suites (rather than failing) when jsdom is
absent or nothing is listening, so it stays useful in a bare checkout.

  unit/test_list_filter.js    ListFilter search/filter/sort/count/sticky, driven
                              through the installed-plugins config eval'd
                              verbatim out of plugins_manager.js so the test
                              cannot drift from the real configuration
  unit/test_render_cards.js   renderInstalledCards markup, both empty states,
                              and escaping of hostile plugin metadata
  dom/test_installed_dom.js   the toolbar in a real DOM, including the HTMX
                              partial re-swap and a getComputedStyle check that
                              .filter-pill[data-active] matches what we emit
  dom/test_store_dom.js       store pagination, per-page, category, tri-state
                              Installed button, persistence across a re-boot
  dom/test_no_double_fetch.js loads the whole plugins_manager.js and counts
                              requests, so a keystroke cannot refetch the store

The DOM suites deliberately fetch the partial and the plugin data from a running
web interface instead of using fixtures, so a renamed element id or a changed
payload shape fails them loudly. Point them at a rig with a full plugin set when
it matters (BASE=http://host:5000); a dev box with two plugins installed passes
while exercising very little.

Several assertions exist to stop specific bugs recurring: trailing spaces
surviving the search debounce, a query spanning two adjacent search fields
(haystack field order is load-bearing), and window.installedPlugins staying at
full length while the grid is filtered. Others guard against passing vacuously —
counting only non-skeleton cards, and checking a search phrase matches something
before comparing two result sets.

The old-vs-new differential suites that verified the store and starlark
migrations are not included: they compared against the pre-refactor code, which
now exists only in git history.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 20:04:47 -04:00
ChuckandClaude Opus 5 26769ee37f fix(starlark): the store authenticated with a key nothing writes (#541)
#535 restored the thirteen routes, so the store stopped answering 404 --
and still would not load. Confirmed against a running device before
anything was changed: /repository/browse answers 200 with 1000 apps in
27s, so the routes are fine. Two things underneath them are not.

**The store never used the token the user configured.** The three
repository routes read `github_token` off config.json. Nothing writes
that key -- it is not in config.template.json, no setting offers it, and
it appears nowhere else in the codebase. The configured token goes to
config_secrets.json as `github.api_token`, which PluginStoreManager
loads and every other GitHub caller uses. So the store could never be
authenticated: 60 requests/hour, on the same per-IP budget 48 installed
plugins spend on update checks, while the 5000 the user had already
configured sat unused. On the device, /plugins/store/github-status
reported authenticated with a limit of 5000 at the same moment
/starlark/repository/browse reported 60, with 18 left. The store going
blank was that 60 running out.

**Every failure looked identical.** list_all_apps_cached turned any
listing failure -- rate limit, DNS, timeout, non-200 -- into an empty
app list, and the route sent that out as `status: success`, so a rate
limit and an empty repository drew the same blank grid with no error
anywhere. It now returns the reason, the route answers 502 with it, and
a failure is no longer cached as an empty repository for two hours.

The guard for a bad response was itself a crash: _make_request catches
`(json.JSONDecodeError, ValueError)` but `json` was never imported, so
evaluating the tuple raises NameError and the guard written for exactly
this case never ran. Reachable whenever something on the path answers
with HTML -- a captive portal, a proxy page, a DNS-hijacking router.

Seventeen handlers answered 5xx with no detail at all.
test_no_api_v3_handler_discards_its_exception is meant to prevent that
across api_v3, but it matched one exact message string, and all thirteen
Starlark routes wrote their own wording. The guard now keys on the shape
that matters: if it returns 5xx, it says why. The 15 pre-existing
non-Starlark functions are listed as a set that may shrink, never grow.

**The listing was capped at 1000 and did not say so.** The contents API
truncates a directory silently; tronbyt/apps has 1075 app directories,
so the store showed a truncated repository and looked complete doing it.
Now listed via the git trees API, which reports `truncated`, with the
contents API kept as a fallback.

Not addressed: the 27-second cold load -- 1075 manifests fetched five at
a time behind skeleton placeholders -- which is probably the largest part
of what "does not load" feels like, and wants its own change.

25 new tests.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:04:27 -04:00
ChuckandClaude Opus 5 968b953a51 fix(display): pin one text layout engine, and give the 5x7 BDF face a size (#539)
* fix(display): pin one text layout engine, and give the 5x7 face a size

Two ways a font could render differently on two machines running the same
code, both found while diagnosing four plugins whose golden images passed on
the machine that generated them and failed everywhere else.

**Layout engine.** `ImageFont.truetype` picks its engine at load time: Raqm
where the host Pillow was built with libraqm, Basic otherwise. The two round
fractional glyph advances differently. `PressStart2P-Regular.ttf` at 8px has
whole-pixel advances, so they agree — which is why most of the fleet matched
everywhere and hid this. `4x6-font.ttf` at 6px does not: glyph positions drift
cumulatively along a run, and the four plugins that draw body text in it
(geochron, of-the-day, christmas-countdown, ledmatrix-weather's almanac) are
exactly the four whose goldens travelled badly.

Every core font load now goes through `src/common/font_layout.load_truetype`,
which pins the Basic engine, so a render depends on the font file and the size
and nothing else. Basic gives up complex-script shaping and kerning pairs;
neither applies to bitmap-grid faces on an LED panel. Output is unchanged on a
host without libraqm.

**Zero font height.** `DisplayManager` built the 5x7 BDF face with
`freetype.Face(path)` and never called `set_char_size`, so `face.size.height`
stayed 0 and `get_font_height()` returned 0 for it — callers stacking rows by
`prev_y + prev_height + gap` drew two lines on top of each other. The
start-up line `Calendar font size: 0 pixels` has been printing the symptom all
along. `font_manager._load_bdf_font` already called `set_char_size`, so
whether measurement worked depended on which path loaded the face.

`DisplayManager` now sets it too, and `get_font_height()` falls back to the
strike the file declares rather than returning a zero line height.

Fixes ChuckBuilds/ledmatrix-plugins#397
Refs ChuckBuilds/ledmatrix-plugins#371, #375, #378, #391

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): give the startup banner a rung that fits a full address at 64px

CI caught what pinning the layout engine exposed rather than caused.
`_fitting_font` walks PressStart2P then 4x6 at 6px, and "255.255.255.255" --
the widest thing the startup banner ever shows -- measures 66px at 4x6/6px
against the 62 a 64x32 panel has to give. It used to squeak in only because
the measurement depended on which layout engine the host Pillow happened to
have; with the engine pinned it does not, so the rung the worst case actually
needs is now in the ladder instead of implied: 4x6 at 5px, which measures 51.

The fallback was wrong in the same place. When nothing in the ladder fit, it
returned `self.font` -- the *widest* option, and precisely how "Initializing"
came to run off the side of a 64px panel to begin with. It returns the
narrowest face that loaded now.

test/test_initializing_screen.py: 34 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): name the exceptions the BDF strike read can raise

Codacy flagged the try/except/pass. It was already narrow in intent -- a
malformed strike table on the measurement path must degrade to "size unknown"
rather than take the display down -- but a bare `except Exception: pass` says
neither of those things and hides a genuinely broken font behind a silent 8px
fallback. It now catches what reading `available_sizes` can actually raise and
logs which face failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: drop logo PNGs the render harness downloaded into the worktree

These are fetched at runtime by the logo cache; they are not source, and they
rode in on a `git add -A` while I was running check_plugin.py against this
branch. Nothing in the change needs them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 09:40:00 -04:00
ChuckandClaude Opus 5 e23f1f45d3 feat(starlark,on-demand): the third-party fixes worth taking, plus a Home Assistant MQTT bridge (#538)
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge

Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of
patches and services built while running this project on Starlark apps
under MQTT control. Its patches are whole-file copies taken against an
older tree, so applying them as written would revert #523's frame
pacing, #534's display() bool returns and the GitHub token masking in
plugins_manager.js. Three of its claimed fixes are already in main, and
its api_v3 Starlark routes are #535's. What follows is the rest --
verified against current code, and reimplemented where the patch's
approach did not hold up.

**On-demand display.** `pinned` reached the controller from the API, was
stored on it and republished in the status payload, but never narrowed
the rotation -- a pinned request still cycled every mode its plugin
owns. Right for a sports plugin, whose modes are views of one subject;
wrong for a plugin whose modes are unrelated, which is every Starlark
app. Now honoured, and it survives a restart.

Restarting while on-demand was active loaded *only* the on-demand
plugin, so normal rotation had nothing to return to for the life of the
process -- and a restart mid-session is routine, since that is how an
update is applied. The panel came back cycling one plugin's modes with
no way out but clearing the cache by hand. Every enabled plugin loads
now; on-demand still resumes on its saved mode.

Stop requests are exempt from the duplicate guards on purpose, so that a
second click stops a mode a race left running -- which means consuming
the mailbox is the only thing that ends one. It was never consumed, so
the same stop was re-read and re-processed on every poll, forever. Both
paths now share one compare-before-delete helper.

**Starlark rendering.** `extract_schema` parsed the source with a regex,
which can only see option lists written out literally: an app whose
dropdown is filled from a live API call inside `get_schema()` came back
empty, and the config form offered nothing to pick. Now runs `pixlet
schema`, which executes the app, and falls back to the parser when
Pixlet is absent, too old for the subcommand, or the app fails to run.
The third-party patch replaced the parser outright and hardcoded
/usr/local/bin/pixlet; this keeps the fallback and the binary search.

A `|` in a config value was dropped by a shell-metacharacter filter,
though the command is a list with no shell involved -- and apps do use
it as a separator inside one value. The key went missing silently and
the app rendered its own "not configured" screen with nothing to say
why. And a 0-byte render was reported as success: Pixlet exits 0 and
writes nothing when an app has no content, which read downstream as a
working app drawing a black panel.

**Starlark display.** `display()` ignored the mode it was called with,
so a specific app could not be addressed. It now accepts `display_mode`
-- which is the whole mechanism, since the controller inspects the
signature before passing it. Found while there: `_select_next_app` ran
only while `current_app` was unset, so with several apps installed the
first was picked once and shown forever while the rest were rendered on
schedule and never displayed. And `enable_scrolling` was missing, so
multi-frame apps were called once per rotation slot and never advanced
past frame one.

**GET /api/v3/display/modes.** Every mode that can be requested
on-demand, with the plugin that owns it. Nothing exposed this, so
anything driving the display from outside the web UI read each plugin's
manifest.json off disk and reimplemented PluginManager's fallbacks. It
also triggers discovery, which is otherwise lazy and normally happens
because a person opened the dashboard.

**integrations/mqtt_bridge.** Home Assistant control over MQTT
Discovery: a mode select, a stop button, power, brightness. Rewritten
against the API rather than the filesystem, so it needs no read access
to config.json and cannot drift from the web UI. paho-mqtt 2.x
VERSION2, TLS, an availability topic that is also the last will, and
secrets from the environment.

**Two opt-in extras.** A DNS single-request unit, for glibc's parallel
A/AAAA lookup stalling ~5s per name on routers that answer only the A
query -- which makes any plugin calling an external API slow and
Starlark apps, which have a render timeout, fail outright. And a Pixlet
config editor: a script you run and Ctrl+C rather than the third-party
version's always-on unauthenticated Flask service, since it stops the
display for the length of a session. Neither is installed by default.

Long Starlark app names now wrap instead of overflowing their card.

115 new tests across 5 files. Also unblocked
test_starlark_display_contract.py, which was silently skipping wherever
fcntl is absent. Whole suite: no new failures against main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(mqtt_bridge): the five issues Codacy flagged on this branch

All in the new bridge, all real:

  * requests floor was 2.31.0, which carries CVE-2024-35195,
    CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which
    is what the project's own requirements.txt already pins.
  * `import time` was never used.
  * `"mqtt_password": None` in DEFAULTS read as a hardcoded credential.
    It is the "no password configured" default; marked nosec B105, the
    convention used elsewhere in the repo.

Also dropped an unused `build_app` from the display-modes test imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: the review findings on this PR

Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and
is answered below.

**One bad config section blanked the whole mode list.**
`/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`,
so a non-dict under a plugin id -- a shape DisplayController already
guards, so it happens -- raised AttributeError mid-loop and answered 500
with no modes at all. Every MQTT bridge entity is built from that list.
Now skipped with a warning.

**The DNS scripts reported success they had not earned.** Three separate
paths: `resolvconf -u` failing was swallowed by `|| true`; the
systemd-resolved branch exited 0 without applying anything, so the
oneshot unit recorded success while the workaround was inactive; and the
installer's `|| echo` turned a failed start into "installation
complete." with exit 0. All three now fail loudly. `single-request` is a
glibc resolv.conf option with no resolved.conf equivalent, so on those
hosts the honest answer is that it cannot be applied.

A NetworkManager-generated resolv.conf is regenerated on connection
changes, not only at boot, and the unit is oneshot with RemainAfterExit
-- so the option can vanish mid-boot with nothing to put it back. Now
detected and stated plainly rather than implied to be permanent.

**`Before=` does not order a manual restart.** It only orders units
already in the same transaction, so `systemctl restart ledmatrix` could
bypass the fix. install_dns_fix.sh now writes a ledmatrix.service
drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround
failing should not stop the display.

**The Pixlet editor's `--lan` is gone.** `pixlet serve` has no
authentication, and a printed warning is not access control. Loopback
only, with the SSH port-forward in the header where the flag used to be
documented -- SSH does the authenticating and nothing is left listening.

**The MQTT example config now defaults to TLS** on 8883. The installer
copies it verbatim, and without TLS the broker password and every
command cross the network in cleartext. A plaintext broker is still
supported and documented, and the bridge warns once at startup when a
password is configured without TLS.

**Not taken: "the upstream Pixlet CLI has no `schema` subcommand."**
Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh`
installs `tronbyt/pixlet`, whose `cmd/schema.go` is
`schema [PATH]` -> JSON on stdout, built on
`runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is
exactly what extract_schema_via_pixlet calls. A binary without the
subcommand exits non-zero and falls back to the source parser, which is
already covered by a test.

**CodeQL stack-trace exposure: not taken either.** I removed `details`
first and that broke
test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception,
which enforces `describe_exception` across all ~75 handlers -- written
because a device with failing storage answered "see logs for details"
from the log viewer itself. describe_exception redacts credentials; the
trade-off is the project's and is already made. Restored, with the
reasoning in a comment.

11 new tests. Whole suite: no new failures against main, 4127 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 08:34:52 -04:00
ChuckandClaude Opus 5 793b988d33 fix(starlark): toggle the app id the list published, and store a relocatable star_file (#537)
The two review nitpicks left over from #535. Both are still on main after
that merge; the five findings alongside them landed with it.

**The toggle could not find what the list had just shown.**
`_starlark_virtual_plugins` publishes the raw manifest key as
`starlark:<key>`, and `_toggle_starlark_app` passed it back through
`_validate_and_sanitize_app_id`, which lowercases and rewrites every
character outside `[a-z0-9_]`. An app stored as `My-App` was listed as
`starlark:My-App` and looked up as `my_app`, so toggling an app the page
had drawn a moment earlier answered 404. Keys written by
`_install_star_file` are already sanitised, so this only shows up for
manifests written by the starlark-apps plugin itself or edited by hand.

`_validate_starlark_app_path` rejects traversal without rewriting, so it
is the check to use here -- listing and toggling now agree on one key.
The updater also uses `setdefault` rather than indexing: the app is
loaded but its on-disk entry need not exist, and `_update_manifest_safe`
does not catch `KeyError`, so that escaped as a 500 rather than writing
the entry.

**`star_file` was stored absolute.** Readers join it to the app's own
directory -- `_standalone_render_starlark_app` does `app_dir /
app_data.get('star_file', f'{app_id}.star')` -- so the key's default is a
bare filename and an absolute value gave it a second meaning. Since
`Path.__truediv__` discards the left side when the right is absolute,
the manifest was pinned to whatever PROJECT_ROOT installed it, and a
moved or redeployed install could not find its own file. Storing
`dest.name` matches the default and stays relocatable. Read paths are
unchanged, so manifests already holding an absolute path keep working.

7 new tests. Whole suite: no new failures against main, 4013 passed
against 4007.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 16:35:56 -04:00
Chuck 50258635a8 fix(starlark): restore the API routes #330 dropped (#535)
The Pixlet install button reported "Pixlet install failed: Resource not found" -- Flask's 404 handler, because the route did not exist. #253 added thirteen Starlark routes; #330 rewrote api_v3.py and dropped all of them, along with the `starlark:<app_id>` entries that surface installed apps in the plugins list and the toggle branch that enables them.

Restores all thirteen routes, the plugin-list entries and the toggle path, so Pixlet installs, the app store browses and installs, and an installed app can be managed like any other plugin.

Not a straight revert. Three error paths stopped returning exception text to the caller; the manifest write moved off a shared temp filename that two concurrent writers could interleave; both dynamic importers stopped leaving half-initialised modules in sys.modules; the config update rolls back when the save fails; the toggle checks that persistence succeeded; and the path check returns the validated path instead of a boolean so callers stop re-joining the raw value. New tests no longer reach GitHub.

Verified on a 256x64 Pi: Pixlet installs and runs (v0.53.1), the store lists 1000 apps, install/toggle/uninstall round-trip, and traversal and command-injection probes are rejected at every entry.

25 CodeQL alerts dismissed as verified false positives -- path-injection where traversal is blocked, and one list-form subprocess with no shell. Both classes already present on main.

Full core suite: 3981 passed.
2026-09-07 16:06:09 -04:00
ChuckandClaude Opus 5 c9289e3a1d fix(store): update_plugin silently did nothing for four installed plugins (#536)
install_plugin() deliberately renames a plugin's directory to the MANIFEST id
when it differs from the REGISTRY id, so registry `stocks` lands in
`ledmatrix-stocks/`. Every lookup in _find_plugin_path() is by directory name,
so update_plugin("stocks") found nothing, logged "Plugin not installed", and
returned False.

Nothing surfaced that to the user. Clicking update in the web UI was a no-op
with no error, and the plugin stayed on a stale version indefinitely. Four
installed plugins hit this on a real device -- leaderboard, music, stocks and
weather -- found because a scripted update of eleven plugins failed on exactly
those four.

Adds a manifest-id scan as the LAST step of the resolution chain, so the two
documented lookups above it (configured dir, then the sibling plugins/
fallback) keep their exact meaning and ordering. That ordering is pinned by
test_discovery_path_contract.py, which characterises the divergence between
the three resolvers on purpose; this extends the chain rather than reordering
it. Directories renamed aside with '.standalone-backup-' during an install or
rollback are skipped, since matching one would report a half-finished install
as a live plugin.

Also adds scripts/audit_render_path.py, which walks the call graph from
display() and reports blocking calls reachable from it. display() runs on the
render thread, so anything slow there stalls the panel; on a vsync-paced loop
a single 15ms call drops a frame and a network round trip freezes the marquee.
Two instances were already found the slow way, by reading frame-time
histograms -- odds-ticker reading the scoreboard cache per frame, and
soccer-scoreboard timing out inside update(). The audit finds that shape in
the source instead. It is a heuristic and says so: a hit behind an interval
check may be fine.

It currently flags 23 calls across six plugins. The clearest is
ledmatrix-music, whose display() falls back to an inline
requests.get(timeout=5) when album art has not been prefetched -- a deliberate
"show the art rather than go blank" tradeoff by its author, but up to five
seconds of frozen panel. Reported, not changed; that is its owner's call.

185 store tests pass. Three of the six new tests fail without the fix.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 15:55:46 -04:00
ChuckandClaude Opus 5 d12323e7f1 perf(scroll): pace frames to the panel — 44→100 fps, stalls 14% → 0.02% (#523)
* perf(scroll): pace frames to the panel, not to a fixed sleep

Scrolling ran at 44-46 fps on a 2x128x64 chain and 14-17% of frames took
41-53ms, which reads as judder. Four independent causes, each measured on
the hardware; details and the diagnostic recipe are in
docs/SCROLL_PERFORMANCE.md.

The high-FPS loop slept a flat 8ms after every render. display() has
already blocked on the panel's vsync by then, so that sleep was added to a
wait that had happened: ~4ms of render plus 8ms put each iteration at ~12ms
against a 10ms refresh grid, so every swap missed a refresh and the loop
settled at 50fps while asking for 125 -- with no headroom, so a further
14% of frames slipped again. It now sleeps only the remainder, with a 1ms
floor so plugin threads still get the GIL.

ScrollHelper stepped position on a wall clock at 1/scroll_delay steps per
second. Plugins set scroll_delay to the frame period, so that comparison
sat exactly on its own threshold: a frame arriving a hair early moved zero
pixels and rendered an identical frame, dirty-tracking skipped the swap, it
returned in ~2ms, and the beat repeated. No scroll_delay value tunes that
out -- a shorter delay trades stalled frames for periodic double-steps.
Both modes now accumulate elapsed time at the same configured speed, so
position stays proportional to real time.

Sub-pixel blending goes back to off by default. It renders a half-step by
mixing two adjacent columns, which on a coarse panel showing pixel-font
text alternates crisp and smeared frames and reads as shimmer -- visibly
worse than integer stepping on the hardware. Vegas mode still opts in.

disk_cache uses orjson when importable, falling back to the stdlib. Encoding
a ~1MB record drops from 14.8ms to 5.4ms end-to-end, and that work holds the
GIL while a marquee is on screen. display_manager also checksummed the whole
framebuffer twice per frame (dirty tracking, then the preview snapshot); the
snapshot now takes the checksum the caller already computed.

New src/common/scroll_config.py resolves scroll settings in one place. Five
ticker plugins each hand-rolled this and disagreed: odds-ticker ranked the
deprecated scroll_pixels_per_second above the documented scroll_speed/delay
pair, and because that key carries a schema default the documented settings
were dead for every user (ChuckBuilds/ledmatrix-plugins#408), while
ledmatrix-leaderboard read the same key only as a fallback. The resolver also
warns when a speed will not advance a whole number of pixels per refresh,
which is the property that actually determines whether a scroll looks smooth.

scripts/build_rgbmatrix_nogil.sh rebuilds the rgbmatrix binding so it
releases the GIL. Upstream declares SwapOnVSync without nogil, unlike
SetPixel/Clear/Fill beside it, so the render thread held the GIL for the
whole vsync wait and starved background threads into long uninterruptible
bursts. The script patches, builds and self-verifies into a scratch tree;
--install backs up the original and rolls back if the service does not come
back healthy.

Measured after: 100 fps locked, no stalls observed, render thread down from
51% to 19% of one core.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): keep the panel swap locked to vsync while scrolling

Dirty tracking skipped SwapOnVSync for byte-identical frames. That is the
right call for static content, but SwapOnVSync is also what paces the render
loop, so skipping it skips the wait for the panel: a duplicate frame returns
in ~8ms instead of ~10ms on a 100Hz panel, advances the strip only 0.8px
instead of 1.0px, and so makes the next frame more likely to repeat as well.
The effect sustains itself once it starts.

Measured over 20 minutes on a 2x128x64 chain, both scrollers configured
identically at 100 px/s:

    leaderboard   10ms x35, 11ms x3            (clean)
    odds-ticker   10ms x26, 8ms x7, 15ms x5    (~20% duplicates mid-scroll)

The duplicates were not end-of-cycle idling -- 38% of fast frames fell within
90s of a scroll completion against 35% of normal frames, a null result. The
trigger is per-frame work: odds does more of it, and more variably, so it is
first to land a frame that advances less than a whole pixel.

Pushing an identical frame costs one canvas copy. Falling out of vsync lock
costs smooth motion. Static content is untouched, because
is_currently_scrolling() expires on its own inactivity threshold -- covered
by test_stale_scrolling_state_stops_forcing_pushes so a plugin that stops
scrolling without saying so cannot pin the panel into always-push.

Also de-flakes test_snapshot_still_written_on_skip, which asserted a strict
mtime increase between two writes that can land in the same filesystem tick;
it failed about two runs in three on Windows regardless of the code under
test. The file is now backdated before the check.

156 tests pass on the Pi. Not yet confirmed by eye on the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): report the frame-time tail, and stop the row-major blit

Two problems, both found by looking at the panel rather than the metric.

The frame-stats line reported ONE instantaneous frame every 5 seconds --
about 1 frame in 500 -- printed beside a 100-frame average. Both hide exactly
the fault they are used to chase: a 2ms duplicate and a 21ms double-wait
average to precisely 10ms, so a ticker stalling on half its frames still
reports a healthy "Avg FPS: 100.0". That reading cost several rounds of
chasing the wrong layer. The line now aggregates every frame since the last
log and reports median, p95, max, min, and explicit stall and skip rates
(past 1.5x the median missed a refresh; under half never reached the panel,
because dirty tracking skipped the swap so the frame never waited on vsync).

On the hardware this now reads:

    leaderboard  100.0 fps over 501 frames | median 10.00ms p95 10.05ms
                 max 10.34ms | stalls 0 (0.0%) skips 0 (0.0%)

The binding rebuild's blit patch becomes opt-in (RGB_PATCH_BLIT=1, default
off). Reordering that loop to row-major changes what a torn frame looks like:
column-major tearing shows as a vertical seam, row-major as a horizontal split
between the panel's upper and lower halves. On a 1/32 scan panel that reads as
a one-pixel fold across the middle of every panel, which is what was reported
on hardware and what went away when the blit was reverted. All of the measured
gain comes from the SwapOnVSync change, so the risky half is simply not worth
taking; the header says so.

Also fixes --install resolving its paths against $HOME, which is /root under
sudo, so it looked in /root/rgbmatrix-nogil-build and died with "no built
module found" on a machine where the build had just succeeded. It now resolves
SUDO_USER's home. Both build paths are verified on the Pi: default yields one
GIL-release site, RGB_PATCH_BLIT=1 yields two.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(scroll): let users pick a crisp speed for their own panel

Whole-pixel motion was previously only available at multiples of the refresh
rate -- 100, 200, 300 px/s on a 100Hz panel. 100 px/s crosses a 256px panel in
2.6s, which is brisk for reading, and everything slower had to blend (blur) or
repeat frames unevenly (judder). There was no way to ask for 50 px/s and get
clean motion.

SwapOnVSync takes a framerate_fraction the display manager never passed. It
holds each frame for N panel refreshes; the panel keeps refreshing at its full
rate throughout, so holding costs nothing in flicker and only changes how often
a NEW image is presented. That turns 50 px/s into one whole pixel every second
refresh instead of half a pixel every refresh.

The crisp speeds are therefore refresh_hz / hold * pixels_per_frame, and that
ladder depends on the panel: a Pi Zero on a long chain has a different set of
good speeds from a Pi 4 on a short one. crisp_ladder() enumerates them and
solve_crisp() picks the best match for a requested speed.

solve_crisp weights motion quality rather than picking the numerically nearest
entry, which matters more than it sounds. Asked for 30 px/s, nearest-by-value
answers 28.6 -- 2px jumps at 14fps -- over 33.3, which is single-pixel motion
at 33fps and obviously better on the panel. The target is also clamped into the
ladder's range first, because relative error saturates near 1.0 for a target
far outside it and the quality penalty would otherwise answer "10000 px/s" with
the slowest entry.

configure() snaps to the ladder and applies the hold when given a display
manager. Without one the hold silently cannot happen and motion falls back to
fractional pixels, so it warns rather than failing quietly. set_frame_hold()
resets to 1 when scrolling stops, so one plugin's pacing cannot leak into
whatever is on screen next.

scripts/scroll_speeds.py is the user-facing part: it prints the ladder for the
configured rate, measures what the panel ACTUALLY manages (--measure, for
hardware that cannot reach its configured limit), highlights the nearest option
to a wanted speed, and demos one live. It never starts or stops the display
service itself -- doing that inside a script stranded the panel twice today.

Speeds below ~20 px/s remain stepped regardless. That is the pixel pitch, not a
software limit.

Also fixes the dirty-tracking test spy, which stubbed SwapOnVSync with a
single-argument function and would have masked the new call as a failed push,
and rewrites a configure() test that had started passing for the wrong reason:
it asserted a judder warning, which snapping now prevents, and was matching the
unrelated "hold could not be applied" warning instead.

183 tests pass on the Pi.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): tie the frame hold to the scroll, not the plugin

The hold applied in configure() never reached the panel. Plugins share one
display manager, and set_scrolling_state(False) -- fired whenever ANY other
plugin finishes its scroll -- reset the hold to 1. A hold set once at plugin
construction was therefore always gone by the time that plugin rendered.

The symptom was a log line that lied. ledmatrix-stocks reported

    Scroll configured: 50.0 px/s (1px every 2 refreshes = 50.0 fps, smooth)

while the panel measured 100.0 fps, median 10.00ms. Config, resolution and
snapping were all correct; only the pacing silently was not applied.

set_scrolling_state(is_scrolling, frame_hold=1) now carries it, so the hold
lives exactly as long as the scroll that asked for it. configure() reports the
value as ScrollSettings.frame_hold instead of applying it -- applying it behind
the caller's back could never have been right on a shared display manager.
Existing callers are unaffected; the default keeps one frame per refresh.

Verified on hardware: stocks at 50 px/s now measures

    50.0 fps over 251 frames | median 20.00ms p95 20.09ms | stalls 0 skips 0

20.00ms being exactly two refreshes, with the panel still refreshing at 100Hz
underneath so flicker is unchanged.

test_another_plugin_stopping_does_not_strand_a_hold pins the interaction that
broke this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll,cache): resolve CodeRabbit review on #523

Eight findings, all reproduced before fixing.

scroll_config.configure() read the refresh rate *after* resolve() had
already used it. resolve() fills in target_fps, pixels_per_frame and the
judder warning from that rate, so on a 60Hz panel every one of them
described 100Hz -- and with snap_to_crisp=False nothing downstream
corrected it, so set_target_fps() paced the helper to 100 FPS. The rate
is now settled first, and falls back to the global config rather than
straight to the default.

refresh_hz_from_config() used `(cfg.get("display") or {}).get(...)`,
which raises AttributeError when either level is truthy but not a
mapping -- out of a function whose whole contract is a rate or a default.

The frame-stats line reported the upper-middle sample as the median and
the 96th sorted sample as p95 of 100. Both are also thresholds (stalls
at 1.5x the median, skips at 0.5x), so the counts were biased too. The
arithmetic is now in frame_stats()/format_frame_stats(), testable
without a clock.

configure()'s docstring and docs/SCROLL_PERFORMANCE.md still said it
applies the frame hold and warns when it cannot. It deliberately does
neither since "tie the frame hold to the scroll, not the plugin"; a
caller following the old text would omit set_scrolling_state() and slow
snapped speeds would still present every refresh.

disk_cache had no policy for non-finite floats: orjson writes null,
the stdlib writes NaN/Infinity, and orjson then rejects those legacy
files so DiskCache.get deleted them as corrupt. One behaviour on both
paths now -- write null, keep legacy records readable. allow_nan=False
detects the values; the replacement walk runs only when there is one,
so the ordinary write path is byte-identical and pays nothing.

build_rgbmatrix_nogil.sh picked the build artifact with a glob piped to
`head -1`, which sorts cpython-311 ahead of cpython-313, so a stale .so
staged in from the source tree was installed as core.so while the GIL
check -- which reads the generated core.cpp, not the .so -- still passed.
It now requires the current interpreter's exact ABI name and fails
closed. Its systemctl calls were also unchecked under `set -uo pipefail`:
a failed stop left the old service running, the following start
succeeded as a no-op, and the health check reported SUCCESS for a
binding that was never loaded.

orjson floor raised to 3.11.6 for CVE-2025-67221 (unbounded recursion
in dumps); it covers the project's Python 3.10-3.13 range.

Adds test/test_cache_nonfinite_floats.py (14) plus regression tests in
test_scroll_config.py and test_scroll_helper.py. 9 of the cache tests
and 9 of the scroll_config tests fail against the pre-fix code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(harness): keep the visual double's signature tied to production

Moves set_scrolling_state's frame_hold into the test double here, where
DisplayManager gains it, rather than in #534 where it arrived a PR early.
CodeRabbit flagged the #534 version correctly: a double that accepts an
argument production does not lets the call pass every harness run and
raise TypeError on the panel, which is the one failure a safety harness
exists to prevent.

The drift has now gone both ways across two branches -- double behind
production on this branch, double ahead of it on #534 -- so it is pinned
instead of remembered. test_display_double_parity.py compares the two
signatures and fails with the direction of the drift named. It reads the
files with ast rather than importing them, because display_manager
imports rgbmatrix at module scope and this check should hold on a laptop
and in CI as well as on a Pi.

Plugins begin passing frame_hold in ledmatrix-plugins#462, which is why
production and the double both need it before that lands.

Full suite: 3889 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 13:37:54 -04:00
ChuckandClaude Opus 5 a0d3e64099 fix: ten defects found validating the whole plugin fleet on hardware (#534)
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers

Three independent fixes found while validating every plugin on a 256x64 rig.

FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf
ships with the core, so every plugin offering it logged "Font family
'tom_thumb' not found" (16 warnings per countdown render) and had to carry a
private loader to use a bundled font. Closes #524.

VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter
that DisplayManager gained, so any plugin passing it died with TypeError at
render time and failed every size. Nine plugins now make that call;
ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other
seven only passed because their scroll path was unreachable without data.
Closes #525.

api_v3 declared module-level config_manager/plugin_manager = None that nothing
ever assigned -- app.py sets the blueprint attributes, which the other 150+
call sites use. Three sites read the decoys, so /health reported the config
unreadable and the plugin system uninitialised (making "degraded" permanent and
unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on
every rig. The decoys are removed rather than assigned, so a bare name is now a
NameError at test time instead of a silent None. The same function's first-call
uptime was computed from two separate clock reads and came out negative.
Closes #529.

Verified on the rig: both previously-failing plugins render, the tom_thumb
warnings are gone, /health reports "healthy" with all three checks passing, and
/display/current reports the real 256x64.

* fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark

The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is
world-writable and sticky, and the display service runs as a different user
from the tooling, so a leftover temp owned by anyone else became unopenable
even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a
sticky directory. The preview and the health check's liveness proxy then froze
until someone deleted the file by hand; on the test rig that meant 23 hours of
a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with
cleanup on failure, matching the hardware-status write a few hundred lines
above. Closes #528.

_poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults
the in-memory TTL to max_age -- so the first request was pinned in memory for an
hour and every later poll returned that stale copy. No second on-demand request
was honoured until the service restarted, while the API kept returning 200.
get() already documents memory_ttl=0 for exactly this cross-process case.
The consumed request is also now deleted: leaving it on disk meant a restart
replayed the previous request, activated it, and ignored the one the caller had
just made. Closes #530.

starlark-apps returned None from display() when it has no app to show, which is
the state of every install without Pixlet and of a fresh one before any app is
added. The controller only skips on a boolean False, so that held a black panel
for the full display_duration instead of rotating on. Closes #456 (core side).

Verified on the rig: two consecutive on-demand requests with no restart between
them are both activated, where the second was previously dropped in silence.

* perf(harness): share one cache across a plugin's renders

_instantiate built a fresh MockCacheManager for every (size, mode), and that
mock is a per-instance in-memory dict, so each render was a cold start. A plugin
that fetches per game or per player re-fetched everything N times over --
baseball-scoreboard at one size took 840s for nine renders where the arithmetic
said ~72s, and at eight sizes it exceeded a 900s timeout.

The second and later renders also never exercised the cache-hit path, which is
what a running rig executes almost all of the time, so a caching regression
could not be caught here.

The cache is now built once per render_plugin_matrix call and threaded down.
The display manager stays per-render -- the bounds checking depends on that --
so only fetched data is shared.

Measured on the rig, same render counts and same goldens:
  tide-display         2s -> 1s   (32 renders)
  cricket-scoreboard  10s -> 3s   (24 renders)
No pass/fail change across tide-display, cricket-scoreboard, clock-simple,
geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages.

Closes #533.

* fix(scripts): run standalone plugin tests instead of collecting nothing

run_plugin_tests.py discovered every plugin test file and handed the lot to
pytest. Most plugin tests are standalone scripts -- module-level main() plus an
`if __name__ == "__main__"` guard, signalling through an exit code -- and pytest
collects zero items from those. The run printed how many files it had *found*,
then "no tests ran", and exited without executing any of them. On a rig with all
44 first-party plugins that is 151 of 248 files.

Files are now classified and each kind runs under the right runner: pytest for
real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1
fail convention ledmatrix-plugins' own runner established (a script that wants a
tty or an LED matrix is a skip, not a regression).

Before:
    $ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos
    Found 1 test file(s)
    collected 0 items
    no tests ran in 0.31s                      rc=0

After:
    Found 1 test file(s) -- 0 collectable, 1 standalone script(s)
    1 passed, 0 skipped, 0 failed (scripts)    rc=0

Verified across three shapes: countdown (1 script), jellyfin-now-playing and
pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files
split 4 collectable / 7 scripts, all seven of which had never run).

Closes #532.

Running the flights scripts for the first time also surfaced four genuinely
failing tests there, hidden by the mirror-image bug in the plugins repo's own
runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465.

* fix(harness): give an empty-looking mode a few frames before warning about it

check_plugin's "drew nothing but display() returned X" warning fired on a single
frame, rendered with force_clear=True, under a frozen clock. All three defeat a
scrolling plugin, whose first frame is legitimately its blank scroll-in buffer.
Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which
people stop reading a warning, which matters because the true positives are
real: a mode that draws nothing and does not return False holds a blank panel
for its whole display duration.

An apparently-empty frame is now re-driven for up to 48 more frames with
force_clear=False (force_clear means "reset the scroll", so repeating it would
redraw frame 1 for ever) and with the clock advancing -- freezegun's factory
where time is frozen, a real sleep where it is not, since scroll position is
usually a function of elapsed time. The first frame that draws content replaces
the result.

The clock is moved back afterwards. It is shared by every render in the matrix,
so time borrowed by the probe leaked into later modes and drifted their goldens
-- f1_upcoming picked up 5 spurious drifts before this was restored.

Measured on the rig:

                        empty warns          check
                        before  after
  f1-scoreboard            42      0    48 PASS / 0 FAIL, goldens intact
  ledmatrix-elections      16      0    16 PASS / 0 FAIL
  on-air                    8      8    true positive, kept
  nfl-draft                 8      8    true positive, kept
  clock-simple/geochron/    0      0    unchanged
  christmas-countdown

58 false positives gone, both true positives kept, no golden regressions. Cost
is confined to modes that really are blank: plugins that draw immediately are
unchanged (clock-simple and tide-display still 2s), while on-air -- eight
deliberately blank modes -- goes to 21s.

Closes #527.

* fix(harness): load nested schema defaults, and merge caller config at leaf level

load_config_defaults read only top-level properties. An object property carries
its defaults on its children, not on itself, so everything nested was dropped --
2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of
565. render_plugin_matrix's comment says the plugin then "behaves like a real
install", which for most of the fleet it did not.

_defaults_from_properties now recurses. merge_config deep-merges the caller's
config onto the result so an override lands at the leaf: a shallow merge would
let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard
every other nhl default, which is the same class of bug being fixed here.

Measured before/after across all 49 installed plugins on the rig: **no render
changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact.
Plugins already fall back to the same values internally via config.get(key,
default), so supplying them explicitly agrees with what they were doing. The
defaults really are arriving now:

  ufc-scoreboard        9 -> 87 defaults
  ledmatrix-flights    51 -> 95
  masters-tournament   10 -> 51
  cricket-scoreboard   22 -> 50
  tide-display         12 -> 18

and hockey-scoreboard, which used to load nhl.enabled=None, now gets
nhl.enabled=True with its full display_modes block.

Caveat worth carrying: the eight plugins with the most nested config
(soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of
the 2,386 dropped defaults, 68%) could not be measured. They import
src.common.sports_shared, which the test rig's core branch predates, so they
fail to load there identically before and after. Re-run this comparison against
a core that has that module before trusting the "nothing changed" result for
them; those are exactly the plugins whose renders should change most.

Closes #531.

* refactor: narrow the exception handlers this branch introduced

Codacy flagged the new code; it passes on other recent PRs, so the finding is
mine. Four of the five broad `except Exception` clauses I added were catching
far more than they needed to, which is the same shape as several bugs this
branch fixes -- hello-world's TypeError sat invisible for exactly this reason.

  freezer() / move_to() / tick()   -> (AttributeError, TypeError, ValueError)
  cache_manager.delete()           -> (OSError, AttributeError, KeyError)

The fifth stays broad and now says why: it wraps a call into a plugin's own
display(), which can raise anything, and the first frame has already rendered --
so a failure there must not turn a good result into an error.

Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty
warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows
7 golden drifts both before and after this branch, so it is not from these
changes -- its committed goldens predate #521's 1-bit text rendering.

* fix: resolve CodeRabbit review and Codacy findings on #534

CodeRabbit raised six; all six were real.

The test double had drifted ahead of production. VisualTestDisplayManager
accepted set_scrolling_state(frame_hold=...) while DisplayManager did not,
so such a call passed every harness run and would raise TypeError on the
panel -- the one failure a safety harness exists to prevent. frame_hold
belongs to the change that adds it to DisplayManager (#523), so it moves
there and the double matches main again.

The harness swallowed exceptions from re-rendered frames. _settle_loop
re-renders a mode that came back blank, to give a scroll time to draw;
returning silently on a crash meant a mode that renders one good frame
and then explodes was reported as passing. Recorded on result.error now,
keeping the captured frame so the failure stays inspectable.

starlark-apps display() returned True after _display_frame() failed, so
the controller held a dead frame for the whole display_duration instead
of rotating on. _display_frame now returns bool on all three paths.

run_plugin_tests.py used env.setdefault for PYTHONPATH and
LEDMATRIX_CORE, so an inherited value won and the subprocess imported a
different core than the one under test -- ledmatrix-plugins#467 exactly.
Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally.

The on-demand mailbox is polled after every frame, ~125x/second on a
scrolling mode, and the read is deliberately uncached, so it was that
many disk reads per second to find nothing. Floored at 250ms, which is
imperceptible for a web-UI click. Consuming it also deleted whatever was
present rather than what had just been processed, so a request posted
while the previous one was in flight was thrown away and never ran; the
delete is now keyed by request_id. That narrows the window rather than
closing it -- a true atomic claim needs a primitive the cache layer does
not offer, and the code says so rather than implying otherwise.

Codacy's 2 criticals were bandit B404/B603 on the subprocess call added
to run_plugin_tests.py. Fixed interpreter, argument list, no shell;
annotated with the repo's existing nosec convention. Bandit is clean on
the file.

Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py
(4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of
those fail against the pre-fix code. Full suite: 3961 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: satisfy Codacy's subprocess checks on the new test runner

Codacy runs Bandit and Opengrep (its Semgrep fork). The new
subprocess.run in scripts/run_plugin_tests.py trips three patterns, on
two different lines:

  Bandit   B404 on the import, B603 on the call
  Opengrep dangerous-subprocess-use-audit          on the run( line
           dangerous-subprocess-use-tainted-env-args on the argv line

A nosemgrep applies only to its own line, so the call line and the argv
line each need one; a single comment on the call covered neither rule
fully. Suppression is the right answer here rather than a rewrite: the
interpreter is sys.executable, the arguments are a list, and no shell is
involved, so there is nothing to word-split or expand.

Matches the pair the rest of the repo already uses for this shape --
permission_utils.py, plugin_loader.py, install_dependencies_apt.py.

Codacy: 0 new issues, up to standards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: leave visual_display_manager untouched so #523 can merge

The only change this branch made to that file was a docstring, and it
collided with #523's rewrite of the same method -- so #534 and #523 each
merged cleanly against main but conflicted with each other. Reverted to
main's text; #523 owns this method and adds frame_hold to it.

The note the docstring carried ('frame_hold arrives in #523') would have
been stale the moment #523 landed anyway. The parity test in #523 is
what actually keeps the two signatures honest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: add the Ruff suppression nosec/nosemgrep do not cover

Ruff reports S603 on the same call Bandit and Opengrep do, and none of
the three suppressions covers the others. Confirmed the precondition
first: path comes from discover_plugin_tests(), which globs test files
inside the repo, and the call is a fixed interpreter with a list argv
and no shell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 13:37:38 -04:00
ChuckClaude Opus 5coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
696acdbc7b feat(render_plugin): add --display-mode so multi-mode plugins can be rendered (#522)
* feat(render_plugin): add --display-mode so multi-mode plugins can be rendered

render_plugin.py always called plugin.display(force_clear=True) with no mode.
A plugin that declares one display mode is fine, but the sports scoreboards
declare three or more and keep their per-mode state on sub-managers; their
no-argument path selects nothing and returns False, so the render came out
blank with nothing to say why. Measured on nrl-scoreboard with identical
seeded state:

  live.display() directly                 True,  1892 lit pixels
  plugin.display(display_mode="nrl_live") True,  1892 lit pixels
  plugin.display()                        False,    0 lit pixels

--display-mode passes the requested mode through. It is only passed when
asked for, so the many plugins whose display() takes no display_mode keep
working untouched, and a plugin that declares modes but does not accept the
argument degrades to its default screen with a warning rather than a
TypeError.

This is what lets the plugin READMEs show a scoreboard at all, and it also
unblocks screens like birdnet_stats and the weather plugin's hourly, daily and
almanac modes, which could previously only be described in prose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Only fall back when plugin display rejects display_mode

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-09-06 17:28:22 -04:00
ChuckandClaude Opus 5 91d15a8943 fix(display): draw text 1-bit, so glyphs stay crisp on the LED grid (#521)
An LED panel has no partial brightness. PIL defaults ImageDraw's fontmode to
"L", which anti-aliases TrueType glyphs into a grey fringe the panel can only
round off -- a 4px glyph arrives smeared into 3px.

DisplayManager creates its shared `draw` in six places and set fontmode at
none of them, while _load_fonts loads extra_small_font as 4x6-font.ttf at
size 6. Measured at draw time, that face at that size puts 74% of its lit
pixels at partial coverage. Every plugin drawing small text through the
shared draw inherited the blur; geochron was the case that surfaced it.

The harness's VisualDisplayManager had the same gap, which mattered more than
it looks: goldens were recording anti-aliased text that production would not
produce, so the harness could not have caught this. Fixing only production
left geochron still blurry under the harness -- that is how the second site
was found.

Both are set to "1" so the harness renders what the panel renders.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 16:02:04 -04:00
Chuck 6bea1a7c21 fix(sports): say when the schema cannot be read, instead of failing silently (#520)
_schema_font_size swallowed every exception and cached an empty dict. That is
not cosmetic. With no schema, a configured font size can no longer be compared
against the schema default, so every size is treated as a deliberate user
choice and skips the snap to the font's pixel grid -- which renders
4x6-font.ttf at 6 instead of 7: a 3px-wide glyph instead of 4px.

That shipped. On a 256x64 panel it made the odds, the team records and the date
row hard to read, and it was found by a user counting pixels on a photo of the
panel rather than by anything here. The cause (_plugin_dir returning None under
the real plugin loader) is fixed in #519; this makes the same class of failure
audible next time:

    Orphan: could not read config_schema.json (FileNotFoundError: ...); every
    font size will be treated as user-chosen and will skip its pixel grid
    snap. Font sizes may render a pixel narrow.

The message names the consequence, not just the error, because the error alone
does not suggest "your fonts are a pixel narrow".

Logged rather than raised: an unreadable schema must not stop a plugin
rendering. The cache is built once per class (per schema path in sports_card),
so this cannot repeat per frame.

Scope deliberately small. An audit of the three shared modules found 23 handlers
that swallow and return a default, but all 23 catch specific types -- TypeError,
ValueError, ImportError -- turning bad config values into defaults, which is
what they are for. Of 77 broad handlers across the font and odds paths, 74
already log. Only these two were both broad and silent.
2026-09-04 16:01:50 -04:00
Chuck 0730d95200 fix(sports): let the plugin declare its own directory, don't deduce it (#519)
_plugin_dir() returned None on every device. The consequence was silent and
reached the panel:

    _plugin_dir()       -> None
    _schema_font_size() -> None for every element
    -> a configured size equal to the schema default stops looking like a
       default and is treated as a deliberate user choice
    -> the snap to the font's pixel grid is skipped
    -> 4x6-font.ttf renders at 6 instead of 7: 3px-wide glyphs, not 4px

On a 256x64 panel that made the odds, the team records and the date row hard to
read. Both `odds` and `detail` were affected -- anything resolving a
grid-snapped schema default was a pixel narrow.

Why it was invisible here. PluginLoader._namespace_plugin_modules renames every
bare module a plugin brought in (sports, game_renderer, ...) to
"_plg_<plugin_id>_<module>" and REMOVES the bare sys.modules entry, so two
plugins owning a module of the same name cannot collide. A class defined in
sports.py still reports __module__ == "sports", but sys.modules["sports"] is
gone, so walking the MRO for a module with a __file__ finds nothing.

Every test here imported plugins directly, which leaves the bare entry in
place, so the walk succeeded. The safety harness loads plugins its own way and
never reproduced it either. It was found by a user counting pixels on the
panel.

The directory is now declared by the plugin (_PLUGIN_DIR) and only deduced as
a fallback, for hosts that declare nothing -- the plugins' own probe harnesses
build classes with type().

Verified on hardware, which is the only place the original failure appeared:
before, the live service logged plugin_dir=None and 4x6-font.ttf@6 for all six
football managers; after, plugin_dir resolves and both odds and detail are @7.

Five regression tests, including the production shape: a class whose __module__
is absent from sys.modules still resolves via its declared directory, and the
precondition that the MRO walk alone returns None is pinned so the test keeps
meaning something if the fallback changes.
2026-09-03 17:01:30 -04:00
Chuck 32d637a446 fix(store): read the core version from disk, not from a stale import (#518)
* fix(store): read the core version from disk, not from a stale import

Updating the core to 3.3.0 and then updating plugins refused all eight sports
scoreboards:

    Refusing to install nrl-scoreboard: NRL Scoreboard supports LEDMatrix
    >=3.3.0, but this system is running 3.2.0.

while src/__init__.py on that machine read 3.3.0. Observed on hardware, not
theorised.

The gate ran `from src import __version__ as core_version`, which binds
whatever the process loaded at start. The plugin store's gate lives in the web
UI, a long-lived service of its own, and the update route deliberately restarts
nothing -- it replaces files on disk and asks the user to restart. Its prompt
named only the *display* service, so a user who followed it left the web
process holding the previous number.

Stale by exactly one release is the case that bites: every plugin flooring on
the release you just installed is refused, blaming a core version that is
already correct on disk. It reads as a broken plugin store. 3.3.0 is the first
release where this hits a whole family at once, since all eight scoreboards
floor there.

compatibility.current_core_version() reads the version from the file instead,
falling back to the imported value on any failure -- so it can only ever be as
correct as before, never worse. All four gate call sites use it: three in
store_manager (install, the git-pull update path, install_from_url) and one in
plugin_loader's advisory warning.

The restart prompt now names both services.

Twelve tests, including the hardware failure itself: a process holding 3.2.0
while disk says 3.3.0 refuses hockey, and reading fresh allows it. The inverse
is asserted too -- a genuinely old core still refuses, so the gate has not
become permissive. One test greps both modules for the old import-bound read;
reintroducing that line fails it, which is what stops this coming back.

Not changed: web_interface/__init__.py also imports __version__, but for
display rather than gating, and the API endpoint already reports a fresh
git describe.

* fix: drop the unused os import

Left over from a first draft that joined paths by hand before this used
pathlib. Flagged by CodeRabbit on #518; confirmed dead -- no os. reference
remains in the module.
2026-09-03 16:06:40 -04:00
ChuckandClaude Opus 5 bc2dbf3824 feat(sports): share the sports.py surface that is identical in all eight scoreboards (#515)
* feat(sports): share the sports.py surface that is identical in all eight

Nine scoreboards ship their own sports.py -- 41,326 lines. Comparing executable
ASTs across the eight that share a lineage, 48 method bodies are byte-identical
in every one: 1,007 lines carried eight times, so 8,056 lines that must be
edited eight times to fix once.

They are the parts with no sport in them: the selection and rotation engine
(_round_robin_favorites, _favorites_first, _compose_selection,
_check_ranking_coverage, _game_divisions, _normalise_quality), the font/colour/
date subsystem (_scale_headline_fonts, _scorebug_font, _resolve_font_size,
_format_game_date, _font_color), and the switch-mode upcoming card
(_draw_upcoming_center_switch). Nothing here knows what an inning is.

Mixins rather than free functions: every one of these reads host state, so
rewriting 48 bodies into free functions would be a rewrite rather than a move,
and it is the move that keeps the renders identical.

Three of the 48 are deliberately left in the plugins, because a byte-identical
body is not automatically safe to move:

- _get_timezone calls resolve_timezone, imported from a per-plugin module
  (hockey_timezone, soccer_timezone, ...). All eight of those differ -- each
  carries its own _WRITEBACK_FIXED_IN -- so hoisting the caller would silently
  bind every scoreboard to one plugin's copy.
- _extract_game_details and _fetch_data are @abstractmethod stubs. They are the
  sport contract; satisfying them from a mixin would let a plugin instantiate
  without implementing its own sport.

_resolve_font_path went the other way: a module-level function, identical in all
eight, that _scale_headline_fonts needs -- so it is inlined here.

_schema_font_size needed a real change rather than a move. It located the
plugin's config_schema.json with __file__, which here is src/common/, so the
load failed silently, the cache stayed empty and every element fell back to an
unsnapped size -- measured at 81% anti-aliased edges on a panel that should be
pixel-crisp. It now recovers the plugin directory from the instance. Note that
type(self).__module__ alone is not enough: SportsCore is an ABC, so a subclass
built with type(name, bases, ns) -- which the plugins' own tests do -- reports
its module as "abc". _plugin_dir walks the MRO past those synthetic classes to
the first module sitting beside a config_schema.json.

Worth recording: the 176 harness renders did NOT catch that regression. The
plugin's own test_fonts_are_crisp.py did. Renders alone were not a sufficient
gate here.

Not merged with src/common/sports_card.py despite fourteen same-named twins.
Only five are provably equivalent by source comparison; the other nine differ in
ways inspection cannot settle, and a wrong guess silently changes what every
scoreboard draws. That merge needs differential testing and is its own change.

* fix(sports): declare the constants the mixins read, and test the contract

CodeRabbit found _QUALITY_CHOICES and _RANKING_COVERAGE_SECONDS read by
_normalise_quality and _check_ranking_coverage but never defined on a mixin.
Confirmed: both are declared by all eight scoreboards, so nothing fails today --
it would only have bitten the ninth plugin to adopt this, at runtime, mid-render.
Both are identical everywhere, so they get defaults here; each plugin's own copy
still shadows them.

Auditing for others showed those two were the only ones, but also that the
host-contract docstring was substantially incomplete: it listed 21 attributes
where the mixins actually read about 40, and omitted five hooks
(_is_favorite_game, _is_game_really_over, _is_ranked_game,
_passes_other_filters, _get_timezone). The section is now derived from that
audit rather than remembered.

test_sports_shared.py covers what is genuinely new, not the moved bodies:

- The contract itself. It parses the module for every ALL-CAPS `self.X` the
  mixins read and asserts each is defined, so the next omission fails here
  rather than in the field.
- _plugin_dir, the only new logic in the move. Including the case that made it
  necessary: SportsCore is an ABC, so a subclass built with type(name, bases,
  ns) -- which the plugins' own tests build -- reports __module__ as "abc". The
  test asserts that precondition before asserting the walk steps past it.
- The three SportsLive bodies. Hockey and lacrosse disable live mode in their
  harness fixtures, so the 176 renders never reach this path; testing the mixin
  directly means coverage no longer depends on which plugin happens to have a
  unit test.

Two of those tests pin things that would otherwise be silently undone.
SportsRecentSharedMixin does carry an __init__ -- SportsRecent.__init__ was one
of the 48 byte-identical bodies. Its bare super() binds to where it is defined,
now the mixin, so it only reaches the host because the mixin is listed first in
the bases. One test proves the chain runs; the next proves that reversing the
order silently skips the host constructor.

* fix(sports): drop three unused imports and let the matcher narrow

Codacy flagged five issues on this file.

Three are unused imports: math, abc.abstractmethod and zoneinfo.ZoneInfo.
Nothing in the module references any of them -- the timezone work goes through
pytz, and the @abstractmethod mention in the module docstring describes the
two stubs that deliberately stayed behind in each plugin, not anything
declared here. pyflakes agrees; all three are removed.

The other two are "team_in is not callable" on the round-robin favourite
matcher. That call is already guarded by callable(), so it cannot raise at
runtime, but callable() is not a narrowing construct a static analyser
follows: the name still carries the None from getattr's default. Normalising a
non-callable to None and branching on `is None` gives the analyser a test it
does understand, and keeps the guard.

Behaviour is unchanged. _round_robin_favorites has no test coverage, so I
exercised it directly on both paths -- a host with _team_in (id matching, the
NRL case) and one without (abbreviation matching) -- across limits 1 to 4, and
the selections are identical before and after. A host whose _team_in is
present but not callable still falls back to abbreviation matching rather than
raising.

test/test_sports_shared.py: 27 passed. The 9 collection errors under
`pytest test/ -k sport` reproduce identically on the unmodified branch and are
not from this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-03 13:29:37 -04:00
Chuck cb0545ecb3 release: report 3.3.0, so the version the gate reads matches the tag (#516)
v3.3.0 is tagged, but src/__init__.py still says "3.2.0" -- and that string,
not the git tag, is what the compatibility gate compares
(store_manager.py: `from src import __version__ as core_version`).

The effect is that every plugin flooring at 3.3.0 is refused on a device
running 3.3.0. Checked against the real gate and the real manifest:

    core __version__ reported to the gate : 3.2.0
    hockey floor                          : 3.3.0
    verdict                               : REFUSE
      "supports LEDMatrix >=3.3.0, but this system is running 3.2.0"

That is all eight sports scoreboards plus calendar 1.2.3, and it would read as
a broken plugin store rather than a stale constant.

The TRUSTWORTHY_FLOOR escape hatch does not cover this: it exempts cores
reporting below 2.0.0 as "unknown rather than old", and 3.2.0 is above it, so
the number is trusted and compared.

This is the same slip as v3.1.0, which was tagged six weeks before its version
string was bumped and shipped __version__ = "1.0.0" -- the reason that escape
hatch exists at all.

With the bump, the same gate call returns ALLOW for hockey at both the
sports_card and sports_shared floors, and for calendar 1.2.3.

CHANGELOG.md gains the 3.3.0 section. That file is what plugin authors read to
decide which release to floor on, so it records the three new modules --
src/common/sports_card.py, sports_game_renderer.py and sports_shared.py --
against this version, along with the three traps in adopting the mixins.

No test pinned the old literal; the ones that care monkeypatch __version__.
135 compatibility tests pass.
2026-09-03 10:32:08 -04:00
Chuck 300cdaa250 feat(sports): share the scroll-card geometry the scoreboards all duplicate (#514)
* feat(sports): share the scroll-card geometry the scoreboards all duplicate

The card helpers moved to src/common/sports_card.py, which shared the eight
scoreboards' settings lookups. Their *geometry* stayed duplicated: nine
methods deciding how wide the centre strip is, how much room each logo gets,
and where an upcoming card's date and time land. Five were byte-identical in
all eight plugins; the other four were identical in seven, each with a
different single outlier.

That shape is why this is a mixin and not free functions. Comparing executable
ASTs against the eight plugins, 67 of the 70 method bodies are inherited
unchanged and 3 become ordinary overrides -- baseball keeps its own
_logo_slot_width and _draw_upcoming_game_status, hockey its own
_upcoming_date_and_time. No per-sport branching goes inside the base.

It deliberately has no __init__ and no state. The plugins' constructors differ
six ways and none of it is worth unifying, so adoption is one line on the
class statement plus deleting what now comes from here.

Placed in src/common/ rather than src/base_classes/sports/ on purpose:
importing that package pulls core.py -> DisplayManager -> rgbmatrix, and this
is pure geometry that must not drag a hardware import into every plugin that
uses it. It sits next to sports_card.py, which the same plugins already use.

Only _SCORE_PROBE varies between plugins, so leagues that reach three digits a
side override that one ClassVar; the four gap constants are identical
everywhere.

The tests drive the mixin through a host that provides exactly the surface the
module docstring names and nothing else, so the mixin growing a new self.*
dependency the plugins do not have fails the contract test rather than
shipping.

* fix(sports): reject non-finite card settings before they abort the render

A center_gap of inf passes `isinstance(x, (int, float)) and x >= 0` unharmed
and then raises OverflowError out of int(). The surrounding guards caught only
(TypeError, ValueError), so it escaped and took the whole card render with it.
The same holds for center_gap_ratio, the two clamp bounds, and layout offsets,
where "inf" arrives as a string and float() is happy to produce it.

Four of the five paths crashed; only a NaN ratio happened to survive, by
accident of min/max rather than by design.

This is pre-existing behaviour -- the bodies moved here verbatim from the eight
plugins and every one of them has it today. Fixing it in the mixin fixes it in
all eight at once, which is the argument for the mixin.

Guarded with math.isfinite() before any int()/round(), falling back to the same
defaults the finite paths already use, plus OverflowError added to the except
clauses as a backstop. Ordinary settings are untouched: all 192 scroll-card
renders (8 plugins x 8 panel sizes x 3 game types) stay byte-identical to
pristine main.

Found by CodeRabbit on #514 and confirmed by running it before fixing.
2026-09-02 16:22:25 -04:00
ChuckandClaude Opus 5 10da2f97fd feat(sports): share the card helpers the eight scoreboards each carried (#513)
Twenty methods were byte-identical in all eight scoreboards' game_renderer.py:
the colour pickers, the scroll_card settings lookup, the date and time
formatting, the favourite-team rules and the font-size grid snapping. Every fix
to any of it had to be made eight times, and a new scoreboard began by copying
them a ninth.

src/common/sports_card.py holds them once. 245 lines leave each plugin.

**Free functions, not a base class.** Every helper takes config/logger/fonts as
arguments rather than reading them off an instance, so a plugin keeps its
method and delegates the body -- call sites, signatures and override points are
all untouched. Adoption is therefore per-function and reversible, which is what
let all eight move with byte-identical renders.

The bodies are the plugins' code moved, not rewritten. Two deliberate
differences, both verified:

- crisp_size takes the seven-plugin guard (`not desired`) rather than
  football's. They agree on every real input; the extra guard only stops a
  None size raising TypeError, so adopting it is a no-op for seven plugins and
  removes a crash path for the eighth.
- schema_font_size caches per schema PATH. The plugins cached on their own
  class, which is the same distinction expressed without a class to hang it
  on; two plugins never share an entry. The path has to be passed in because
  the plugins derived it from __file__, and __file__ here is the core's.

Verified before any plugin was touched: 534 differential comparisons of the
helpers against afl's originals and 704 more of the font-sizing chain against
all eight plugins' originals -- 1,238 comparisons, zero differences. Writing
the constant tables by hand introduced two errors that check caught: the tie
colour was (255,255,0) instead of the plugins' (255,200,0), and a
"five_by_seven" alias that does not exist. Both are now taken from the plugins
verbatim.

43 tests pin the contract, including the cases the plugins' own comments record
as having bitten: a three-character string must not iterate into a colour, a
shared font face must give up rather than guess an element, a font_size equal
to the schema default carries no intent, and a bad timezone falls back to UTC
rather than blanking the card.

Full suite 3768 passed, 6 skipped.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 16:22:09 -04:00
ChuckandClaude Opus 5 92f9d06af9 fix(logos): stop a failed download pinning a team to a grey box forever (#512)
* fix(logos): stop a failed download pinning a team to a grey box forever

When a logo download fails, create_placeholder_logo writes a 64x64 grey PNG
under the *real* logo's filename. Every later call then hits
`if filepath.exists(): return True` and reports success, so the real logo is
never attempted again. One transient failure -- no network at boot, ESPN
blipping -- permanently costs that team its logo.

This is not hypothetical. Five of the eleven cached AFL logos in my checkout
were 384-byte stubs written in a single bad minute, and they had stayed that
way ever since; the scoreboard rendered COLL, FRE, NMFC, PORT and SYD as grey
text boxes on every card.

Placeholders are now stamped with a `ledmatrix_placeholder` PNG text chunk
carrying their creation time, and `is_placeholder_logo` recognises them. It
also matches on the placeholder's exact geometry and background colour, so the
stubs already sitting on users' disks are picked up too -- without that, this
fix would only help teams whose logos break in future. Verified against the
real stubs: all five detected, all six real logos untouched.

`download_missing_logo` now treats an existing placeholder as the failed
download it is and retries, rather than as a satisfied request. The retry is
rate-limited to PLACEHOLDER_RETRY_SECONDS (6h) so this does not trade a
permanent grey box for an ESPN request every frame; a failed retry rewrites the
placeholder, restarting the clock. The age comes from the stamp rather than
mtime, so a backup restore, an rsync, or a permissions script cannot silently
reset it.

`download_missing_logos_for_league` gets the same treatment -- a bulk pass is
exactly where a previously failed logo should get another chance -- and
`LogoHelper.load_logo_with_download` no longer accepts a stale placeholder as a
cache hit. That import is lazy and guarded so the module still works against a
core build predating the marker.

`LogoHelper._create_placeholder_logo` needs no change: it returns an in-memory
image and never writes it to disk, which is the behaviour this bug argues for.

Tests cover marked and legacy-unmarked detection, the two false-positive cases
(a real 500x500 logo, and a 64x64 image that is merely the same size), the
retry, the rate limit, and that the age survives an mtime touch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(logos): address review — unify eligibility, invalidate cache, restart back-off

Three findings from the review on #512, all confirmed against the code:

1. The three download sites each had their own idea of "already have it".
   download_missing_logos_for_league() retried *any* placeholder, ignoring the
   back-off entirely, while download_all_ncaa_football_logos() was never
   updated and still skipped placeholders forever. They now share one
   should_attempt_download(), which also covers force_download, so the sites
   cannot drift apart again. download_missing_logo() reads through the same
   helper.

2. LogoHelper.load_logo_with_download() answered from the in-memory cache
   before touching the disk, so after a stale placeholder was successfully
   replaced the *cached placeholder image* was still returned -- the real logo
   would not have appeared until the process restarted. The cache entry for
   that file (every size of it) is now dropped after a successful download.

3. A failed retry left the stale placeholder on disk with its old timestamp,
   so the next call saw it as stale again and retried immediately: a download
   attempt per call, which is precisely what the back-off exists to prevent.
   refresh_placeholder_timestamp() restamps it, and the helper calls that on
   the failure path. It refuses to touch anything that is not a placeholder.

Tests cover both bulk loops in both directions (fresh placeholder skipped,
stale one retried), the eligibility rule including force_download, the
timestamp refresh, and the two LogoHelper paths -- including that a
freshly-downloaded logo is actually what comes back rather than the cached
placeholder.

Two of the new bulk-loop tests initially passed for the wrong reason: the
fetch_teams_data stub returned {}, which is falsy, so the loops bailed before
reaching the eligibility check at all. Fixed to return a truthy payload.

Re-verified end to end: with both halves in place, rendering the AFL scoreboard
took FRE.png from a 362-byte stub to a 12,928-byte logo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 13:21:07 -04:00
ChuckClaude Opus 5coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>CodeRabbit
6e361e05cc docs(sports): record B6 as done, and what running it found (#435)
* docs(sports): record where B6 stands, and why it is waiting

The phase table had B4 as "next" and B5 as "after B4" while both had shipped,
and described B6 as blocked on B4's gate — which is now merged and released. A
plan that misreports which phase it is in is worse than no plan: the next
person reads it and repeats finished work.

Corrected, and three things that were only ever decided in conversation are now
written down:

  * **B6 is deliberately held.** 3.2.0 published 2026-08-03; 3.1.0 ran nine
    months before it. B6's premise is that cores without the module are gone,
    and there is no release-asset count or install telemetry to show that.
    Running it now strands users on their current plugin versions. The gate
    that makes it safe is already built and tested — it is the calendar that is
    missing, and no amount of further code changes that.
  * **Stop adopting further shared modules** (data_sources, game_renderer,
    base_odds_manager) until B6 closes. Each adoption adds a copy to keep in
    step against a payoff contingent on B6.
  * **A B5 retrospective**, because "the adoption went fine" is not what
    happened: four of eight shipped with scroll mode broken on a 3.2.0 core.
    The bundled fallback did not protect against it — the break was on the
    modern path — which is an argument for the sunset, not against it. Records
    the ledger too: net negative on disk until B6 runs.

Also replaces the "what's next" list, whose first five items were all done,
with what actually remains.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix: apply CodeRabbit auto-fixes

Fixed 1 file(s) based on 2 unresolved review comments.

Co-authored-by: CodeRabbit <noreply@coderabbit.ai>

* docs(sports): stop a wrapped PR reference reading as a heading

A line wrapped onto "#433), the newest manifest entry ...", which
markdownlint reads as a malformed ATX heading (MD018). Reflowed so the
line starts with "(#431, #433)" instead.

Not the suggested fix: adding a space after the hash would have turned
the PR reference into "# 433". The B5 safety claim raised alongside this
was already corrected in ac44b5a.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* docs(sports): scope the B5 safety claim to the fallback

The heading read "B5 — adoption is safe by construction", which this same
document disproves two sections later: four of the eight adopted plugins
shipped with scroll mode broken on a 3.2.0 core and were repaired in
plugins #251.

The body was already careful -- it says fallback compatibility is what is
guaranteed, and that correctness on a core which *does* ship the module
needs object-level and scroll-mode validation. The heading was not, and a
heading is what a reader scanning the plan actually takes away.

Retitled to name both halves, with a sentence up front saying why the
unqualified claim is false and pointing at the retrospective that shows
it. The phase intro said "one of them is safe by construction and the
other is not"; that now says what it actually means -- one cannot break a
user on an old core, the other can.

The second review point, MD018 on the ATX heading at line 409, does not
reproduce: that line now begins "(#431, #433)" rather than "#433)", so
there is no bare-hash heading. `grep -cE '^#+[^ #]'` returns 0 for the
whole file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* docs(sports): re-check the hold, and close two items that are already done

The remaining-work list had two entries that finished without the doc noticing,
which is the failure mode this file exists to prevent.

- The stale plugin-test tranche is gone. run_plugin_tests.py --all now reports
  174 passed, 2 skipped, 0 failed across the whole fleet. Recorded how to
  re-check it too: these are standalone scripts, not a pytest suite, and one
  calls sys.exit(1) at import, so pointing pytest at a plugin directory
  collapses into an INTERNALERROR that looks nothing like the real state.
- CLAUDE.md already says eight panel sizes.

That leaves the hardware soaks as the only open item needing work rather than
calendar time.

B6's prerequisite is now built -- core test/test_sports_sunset_matrix.py
(#505) -- so the phase table and the regression-test section say so, and the
two modelling traps it had to work through are recorded for whoever touches it
next: the copy-removed shape must be an unguarded import or the failure names
scroll_display_legacy instead of the core module, and only the leaf module may
be hidden because a pre-3.2.0 core still ships src/common/.

The hold itself is re-checked and unchanged: v3.2.0 is still latest,
__version__ is still 3.2.0, no 3.3.0, 23 days rather than the few months the
gate asks for. Also worth stating plainly -- the core updates by git pull, not
by downloading a release, so release-asset counts would not measure uptake even
if we had them. Whatever unblocks this has to come from the store side.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
2026-09-02 11:34:18 -04:00
ChuckandClaude Opus 5 a686932c7e fix(store): land the install_from_url gate, which never reached main (#511)
#510 shows as merged, but into fix/gate-git-pull-updates -- #508's branch --
rather than main. #508 reached main first, so the sideload gate was left behind
on a branch. Same failure as plugins #350/#351, which merged into each other's
bases; worth knowing the pattern, because GitHub reports these as MERGED and
`gh pr list` shows nothing outstanding.

main today has two of the three routes gated: install_plugin (#431/#433) and
update_plugin's git branch (#508). install_from_url validates required manifest
fields and then installs whatever it found, never comparing the core version.

Cherry-picked unchanged from the orphaned branch -- it applies to main with no
conflict. TestSideloadGate pins the three cases the other routes pin: refuses a
floor above this core leaving nothing behind, still allows a compatible plugin
(the guard against a gate that refuses everything), and does not block a 2.0.0
floor on a core reporting an untrustworthy version.

Full suite 3725 passed, 6 skipped.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 09:51:27 -04:00
ChuckandClaude Opus 5 154525beb8 fix(background): release the payload after every callback, not after each one (#509)
Callers that join an in-flight fetch share one FetchResult. #499 released the
payload inside the delivery loop, so the first callback got the data and every
joiner got `result.data is None`.

That is not a quiet degradation. Consumers read `result.data.get('events')`, so
they raise AttributeError -- which the delivery loop catches and logs. The
entire failure surfaced as one line:

    ERROR - src.background_data_service - Error in callback for request
    nhl_2026_...: 'NoneType' object has no attribute 'get'

and a manager that silently never received its schedule. Seen on hardware:
NHLRecentManager logs "Background fetch completed for 2026: 1000 events" and
the very next line is the error, from NHLUpcomingManager's callback on the same
request -- which had already logged "No events found in shared data."

Deduplication is the normal case, not a corner. A sport's recent, upcoming and
live managers all want the same season schedule, so the second and third are
joiners on almost every cycle. _release_payload's own docstring said "once A
callback has been handed the data", singular, which is the assumption that
broke: the loop above it was written for many, and says so.

Moved after the loop, and guarded on `callbacks` being non-empty. The guard
matters: a request submitted without a callback must keep its payload, because
polling get_result() is then the only way to collect it. The per-delivery
release got that right by accident -- an empty list never entered the loop body
-- and the existing test for it caught the omission.

test_background_payload_release.py gains TestJoinersAllGetTheData: two
submitters on one in-flight cache_key, asserting both are handed a populated
payload, plus that the memory fix still happens once they have all had it.
test_background_fetch_dedupe.py already proved the joiner's callback FIRES; it
never checked what the callback received, which is the gap that let this
through.

Verified the new test bites: restoring the release inside the loop fails it
with "'second' was handed a released payload".

Full suite 3716 passed, 6 skipped.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 09:54:04 -04:00
ChuckandClaude Opus 5 9b522d412c fix(store): gate the git-pull update path (#508)
* fix(store): gate the git-pull update path

`install_plugin` gates every route that re-downloads, `_reinstall_with_rollback`
included. `update_plugin` has one branch that re-downloads nothing: a git
checkout pulls in place, installs dependencies, and returns True. A pull could
therefore deliver a manifest flooring above this core and nothing would notice
until the plugin failed to load — which surfaces as one line in the journal and
a display that silently stopped appearing.

Checked after the pull rather than before it, for the same reason
`_install_plugin_impl` checks after the download: the registry carries no
compatibility field, so the incoming floor is only knowable once the new commit
is on disk.

Undone with `git reset --hard` to the pre-pull commit rather than by removing
the directory. This is a live checkout, the old commit is still in the object
store, and the reset leaves the user on the exact version they were already
running — the same promise `_reinstall_with_rollback` makes, reached by the
means this path actually has, with no window where the plugin directory does
not exist. An unreadable manifest allows: it is not evidence of a floor.

Scope, stated plainly: monorepo plugins install as archives and update through
`_reinstall_with_rollback`, so they were already gated. Only registry entries
with no `plugin_path` reach this branch. It is closed anyway because the sunset
rule in the plugins repo's `08-shared-sports-code.md` names, as condition 3,
that the core enforces the floor "at install/update time" — and B6 rests on
that being true rather than merely written down. `install_from_url` is still
ungated; the tests say so rather than letting the next reader assume otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(store): do not pull what the gate cannot un-pull

Review of the gate found a data-loss path it had introduced, plus two smaller
scope errors. All three from CodeRabbit on #508.

**The stash failure was load-bearing and was not treated as one.** update_plugin
stashes local changes before pulling; when that stash failed or timed out it
logged a warning and pulled anyway. That was harmless while nothing ever undid
a pull. It is not harmless now: the gate's rollback is `git reset --hard`, which
discards uncommitted tracked edits -- exactly the edits the stash existed to
protect. A pull does not refuse on a dirty tree as long as the incoming commit
touches other files, so the sequence completed silently: pull succeeds, gate
refuses, reset takes the user's work with it.

update_plugin now returns before pulling unless the tree was already clean or
was successfully stashed. Refusing costs an update in a case that had already
gone wrong; the alternative costs data. That also makes `--hard` safe by
construction in _gate_pulled_commit, and its comment now says so rather than
observing it in passing.

Pinned by test_a_failed_stash_stops_the_update_before_pulling, which writes a
local edit, forces the stash to fail, and asserts both that HEAD did not move
and that the edit is still on disk. Verified it bites: with the new guard
removed the file comes back as `class P: pass`, the edit gone.

**_HAS_GIT could take the module down instead of skipping it.** With no git on
PATH, subprocess.run raises FileNotFoundError, and this runs at import time --
before skipif can act, so the whole file errors rather than skipping. Now
catches OSError.

**The doc overclaimed the gate's reach.** It said the floor is enforced on
"every route that installs or updates" while the same passage notes
install_from_url is ungated. Both spots now scope the claim to registry-managed
installs and the two supported update paths, and name the sideload exception.

Full suite 3720 passed, 6 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-31 17:35:00 -04:00
Chuck cbc540a679 fix(assets): drop case-colliding duplicate league logos (#506)
Four league fallback logos were each tracked at two paths differing
only in case:

  assets/sports/mlb_logos/MLB.png  +  mlb.png
  assets/sports/nba_logos/NBA.png  +  nba.png
  assets/sports/nfl_logos/NFL.png  +  nfl.png
  assets/sports/nhl_logos/NHL.png  +  nhl.png

On Windows and default macOS the filesystem is case-insensitive, so
both index entries map to one physical file. Whichever git writes last
wins and the other entry reports as permanently modified, so `git
status` can never be clean and any `git pull` flips which one is dirty.

For MLB/NBA/NFL the two entries pointed at the same blob, so the
collision was only cosmetic. NHL was not: NHL.png is 1439x1621
(672455 B) and nhl.png is 768x768 (107184 B), so which resolution the
league fallback logo loaded depended on checkout order rather than on
the code.

Keep the uppercase path in each pair. Every logo lookup uppercases the
abbreviation before building a filename -- LogoDownloader
.normalize_abbreviation and .get_logo_filename_variations
(src/logo_downloader.py) and LogoHelper.normalize_abbreviation
(src/common/logo_helper.py) all do -- and nothing in the tree requests
a lowercase league logo, so the uppercase name is what the code
actually asks for. For NHL that is also the higher-resolution asset.

Removed with `git update-index --force-remove` so the literal index
entry is dropped without the case-insensitive working tree deleting the
survivor.
2026-08-30 09:08:42 -04:00
Chuck 4aeb0033e0 fix(install): don't abort when journald settings are absent (#507) 2026-08-29 18:32:50 -04:00
ChuckandClaude Opus 5 eae063700f test(sports): pin the four-case sunset matrix before B6 runs (#505)
B5 and B6 make different promises about an adopted scoreboard, and only the
second is obvious. The table:

                     | bundled copy present   | bundled copy removed
  pinned old core    | loads (the fallback)   | ERROR, naming the module
  current core       | loads, using core code | loads, using core code

The top-left cell is the one worth having. Nothing in the suite proved that an
adopted plugin still runs on a core predating src.common.sports_scroll, and
that claim is the entire basis for having shipped B5 ahead of B6's gate.

The bottom-left cell is asserted through PluginManager.load_plugin rather than
a bare import, deliberately. The manager catches the ModuleNotFoundError, so a
test written around pytest.raises would pass against a core where the module is
merely broken rather than absent, and would say nothing about what the user
meets: a plugin parked in ERROR and one log line. The assertion is the ERROR
state plus an error naming the module, which is what makes the log actionable.

Modelling notes, both of which were wrong first time and matter:

- The copy-removed shape is an UNGUARDED import, not the guarded one with the
  legacy file deleted. Keeping the guard while removing its fallback only
  mislabels the failure -- the plugin reports a missing scroll_display_legacy
  and never mentions the core module that is actually absent.
- Only the leaf module is hidden. A pre-3.2.0 core still ships src/common/;
  hiding the package would be a harsher core than any that shipped, and would
  make the failure name the package instead of the module.

Ablation: disabling the old-core simulation fails three cells, and dropping the
error object from load_plugin's set_state fails the fourth, so none of it
passes vacuously. The install gate -- the other half of the guarantee -- stays
in test_plugin_compatibility_gate.py rather than being duplicated here.

Refs docs/SPORTS_UNIFICATION.md, "B6 -- why the sunset needs more than a
version floor".


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 11:25:10 -04:00
ChuckandClaude Opus 5 af96bd5cb0 fix(memory): join an in-flight fetch instead of starting a duplicate (#503)
* fix(memory): join an in-flight fetch instead of starting a duplicate

submit_fetch_request() had no notion of "already fetching this". request_id
embeds a millisecond timestamp, so every submit looked new, and
active_requests is keyed by that id rather than by what is being fetched.
Two submits for the same cache_key therefore started two identical fetches.

It is not a rare race. _fetch_data in the sports managers branches: the Live
manager fetches only today's games, but Recent and Upcoming both pull the
full season schedule under the SAME cache_key. On a cache miss both miss,
both submit, and nothing stops the second. On a running 512x64 board:

    138 background fetches in 24 hours, arriving in pairs at identical
    millisecond timestamps, roughly hourly:

      2  2026-08-25 11:47:26.612
      2  2026-08-25 10:46:28.064
      2  2026-08-25 08:01:55.962

Half of them redundant. Each duplicate costs a second download, a second
JSON parse -- the expensive part on a Pi -- and a second parsed copy
resident at the same time. Schedules on that board run 256KB to 20MB, 106MB
across all sports. Because the pairs land in the same millisecond they also
occupy two of the three executor slots with identical work, which is what
makes two large parses peak simultaneously.

A submit for a cache_key already in flight now joins that request: its
callback is added to the existing one and the existing request_id is
returned, so get_result() works for both. Different keys are untouched, and
dedupe applies only while a fetch is in flight -- a submit after completion
fetches again, because this is not a second cache layer.

Three details:

  - The in-flight entry is dropped and the callback list snapshotted in the
    SAME critical section as filing the result. Otherwise a submitter could
    join a fetch whose callbacks had already run and never be called back.

  - Cancellation is the other way a request leaves active_requests, so it
    releases the key too. And the join path looks the request up rather than
    trusting the id, so an entry stranded any other way cannot wedge a key
    permanently -- it is dropped and a fresh fetch starts.

  - One callback raising no longer prevents the others being delivered.
    Previously there was only ever one.

Interaction with #499, whichever merges second: that PR releases the payload
after the callback runs. With several callbacks the release must happen
after ALL of them, and must not happen at all if a joined submitter passed
no callback, since polling get_result() would then be its only delivery
path. The callback list built here is the hook for that.

test_background_fetch_dedupe.py -- 8 tests, covering the join, callback
delivery to both submitters, one callback raising, distinct keys not being
coalesced, a post-completion submit fetching again, cancellation releasing
the key, a stranded entry not wedging one, and the reported count. Verified
non-vacuous by removing only the join branch: 3 fail. The callback tests
assert the ids coalesced, without which they would pass trivially on two
independent requests.

Full suite: 3698 passed, 60 skipped, 1 failure that reproduces identically
on unmodified main (test_install_lowmem, environment-dependent: /var/tmp is
disk-backed on this machine).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(memory): discard a cancelled fetch instead of letting it commit

Review follow-up on the dedupe.

Cancelling releases the cache_key, so a replacement fetch for that key can
start immediately. But _fetch_data_worker() had no cancellation check: the
cancelled worker still wrote its response to the cache, flipped its own
status from CANCELLED to COMPLETED, and ran its callbacks. The stale
response could therefore land on top of the replacement's fresher data.

The worker cannot abort an HTTP call in flight, so the response is discarded
on return instead: no cache write, no callbacks, status left CANCELLED. The
check sits immediately before the cache write, which is the first
side effect.

Also fixed, found by the new test rather than by reading:

    request_id was f"{sport}_{year}_{milliseconds}", which is not unique.
    Two submits inside the same millisecond produced the SAME id -- the
    test's two sequential fetches collided on a fast mocked response, and
    one request silently replaced the other in active_requests and
    completed_requests. Rare before this PR; load-bearing now, because
    dedupe hands that id back to every joiner as their handle for
    get_result(). A per-service counter is appended.

Two test problems of my own, both fixed here rather than left to flake:

  - The cancellation test synchronised with time.sleep(0.4). A slow worker
    would have made it pass for the wrong reason. It now waits for the
    request to be filed in completed_requests.
  - The id-uniqueness test patched session.get, but submits are async: the
    50 workers outlived the patch and made real DNS calls to the dummy host.
    It stubs the executor instead, which is what a submit-time test should
    exercise.

20 consecutive runs of the dedupe file: 0 failures. Full suite: 3700 passed,
60 skipped, 1 failure that reproduces identically on unmodified main
(test_install_lowmem, environment-dependent).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(memory): make cancellation terminal, not advisory

Three paths wrote request.status without checking whether the request had
already been cancelled, so a cancel could be silently undone and the work it
was meant to stop went ahead anyway.

- A request cancelled while queued had CANCELLED overwritten with IN_PROGRESS
  the moment its worker started, defeating the discard check entirely: it
  downloaded, cached and called back for work the caller had withdrawn. It
  now skips the fetch outright, which is also the cheapest possible cancel.
- The cancelled-check and the cache write were separate critical sections, so
  a cancel landing between them left the payload in the cache with the
  callbacks suppressed -- every submitter that joined the fetch waited for a
  call that never came. The worker now claims the commit in the same critical
  section that reads the status, and cancel_request refuses once claimed. The
  write stays outside the lock: it serialises a multi-megabyte payload to the
  SD card, and holding the service lock across that would stall every submit,
  status query and cancel behind it.
- A cancelled request that then failed was relabelled FAILED, which slipped
  past the CANCELLED-only callback gate and delivered a spurious error
  callback. The except path now leaves CANCELLED alone.

get_request_status() also reported a cancelled request as FAILED, since it
inferred status from result.success; the final status is now recorded on the
result. Both early returns assign to `result` so completed_requests files the
outcome that was reported rather than the untouched placeholder.

Tests cover cancellation before worker start, during the commit, and during an
HTTP failure; all three fail against the unfixed code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* test(memory): reach the exception path as a cancelled request

test_a_failure_after_cancelling_stays_cancelled cancelled the request while
its worker was still queued, so once the pre-start branch landed the worker
returned there and never reached the exception handler the test is named for.
It passed against the unfixed code only because that branch did not exist yet;
with it, the test passed for the wrong reason and reverting the except-path
guard did not fail it.

Cancel while the worker is parked inside the HTTP call instead, and assert the
fetch actually started so the test cannot silently degrade into the pre-start
case again. Reverting each of the three guards now fails exactly one test.

Also read the payload inside the callback rather than off the FetchResult
afterwards: #499 releases result.data once delivery is done, so the later read
saw the released object and not what the caller was handed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 10:03:43 -04:00
Chuck 5e5979973e Update README.md (#504)
Signed-off-by: Chuck <33324927+ChuckBuilds@users.noreply.github.com>
2026-08-26 09:13:23 -04:00
ChuckandClaude Opus 5 bdced206dc fix(plugins): retain state history by age, with the count as a ceiling (#502)
* fix(plugins): retain state history by age, with the count as a ceiling

Follow-up to the cap in this PR. A flat entry count answers the wrong
question: what a reader wants from this history is "the last couple of
hours", and how many transitions that is depends entirely on the plugin's
update interval. On a real board those span 2s to 3600s, so 200 entries is

    interval   200 entries covers
        2s          3.3 minutes     (flights, live)
       10s         16.7 minutes     (jellyfin)
       60s          1.7 hours       (default)
      300s          8.3 hours       (news)
     3600s          4.2 days

-- the plugin churning hardest, the one worth looking at, keeps the least.

So transitions are now trimmed by AGE first
(STATE_HISTORY_MAX_AGE_SECONDS, two hours), which makes the retained window
comparable whatever the cadence, and the count cap becomes purely a memory
ceiling for pollers fast enough to exceed it inside that window. The
ceiling rises 200 -> 2000: at ~230 bytes an entry that is ~0.5MB per plugin
worst case, and only plugins updating faster than roughly every 4s can
reach it. Steady-state memory is unchanged for everything slower, since the
age trim binds first.

Two details worth stating:

  - The trim reads time.monotonic(), stored alongside each transition,
    rather than the datetime already inside it. A DST shift or an NTP step
    would otherwise make every entry look ancient and flush the history in
    one go. The human-readable timestamp is untouched and still what
    get_state_history() returns.

  - Trimming happens on append, so a plugin that goes quiet keeps its last
    window until it writes again. That is deliberate: it is bounded either
    way, and a lazy trim costs nothing on the hot scheduling path. The
    guarantee is therefore about the SPAN of retained history, not its age
    against the current clock, and the test asserts it that way.

The public shape is unchanged: get_state_history() still returns the same
list of transition dicts, and state_history_count is still the lifetime
total.

test_plugin_state_history_retention.py adds 7 tests. Verified against this
branch with only the age trim removed: 4 fail, 3 pass -- the three that
survive are testing the count ceiling and the monotonic clock, which this
commit does not change. Full suite 3753 passed, 60 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(plugins): build get_state_info() as one locked snapshot

Every field was read under its own lock, so an unload running concurrently
could be observed half-done: 'state' read before clear_state() removed it
and 'state_history_count' read after, handing PluginManager.get_plugin_info()
a plugin that is ENABLED with zero transitions.

The whole payload is now built in one critical section. _lock is an RLock,
so the helpers called inside it can still take it.

The regression test runs a reader against a thread that repeatedly fills and
clears the same plugin, and fails on the first torn snapshot. Verified by
removing only the lock: fails on 3 of 3 runs, passes on 3 of 3 with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 08:56:45 -04:00
Ron PierceandClaude Opus 5 39e7f8cbe0 fix(plugins): cap the per-plugin state transition history (#501)
* fix(plugins): cap the per-plugin state transition history

PluginStateManager recorded every state transition in a per-plugin list
and never trimmed it. The only code that removed entries was
clear_state(), called solely from PluginManager.unload_plugin(), so a
plugin that stays loaded -- normal operation -- never released one.

The list is written on the hot scheduling path. Every update cycle
appends twice: _reserve_for_update() sets RUNNING and _finish() sets
ENABLED back again. At the default 60s update interval that is 2,880
entries per plugin per day, and nothing reads them -- get_state_info()
only takes their len(). Pure dead weight.

Measured against the unpatched class, ten plugins on a 60s interval:

    sim uptime   history entries   heap growth
          1 day           28,810        7.7 MB
          7 days         201,610       53.9 MB
         30 days         864,010      230.9 MB   (still climbing)

With the cap it is flat at 2,000 entries / 0.5 MB from day one.

On a 1 GB board 231 MB of garbage is fatal on its own, and the failure
is not a clean OOM: once MemAvailable falls far enough fork() starts
returning ENOMEM, so sshd accepts connections and closes them before its
banner while the kernel still answers pings. The board looks like a
hardware fault and needs a power cycle. Same family as the ceilings
added in #464.

Retain the most recent 200 transitions per plugin in a deque and let the
rest age out. state_history_count is surfaced through the web API, so
the lifetime total is tracked separately rather than plateauing at the
cap. get_state_history() now returns a copy under the lock; it was
handing out the manager's own list, which a caller could mutate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(plugins): copy history entries out, lock clear_state

Review follow-ups on the transition history.

get_state_history() copied only the outer list, so a caller holding a
returned transition could rewrite the manager's record of what happened
-- which contradicted the defensive-copy guarantee in its own docstring.
Copy each entry too. Every value in a transition is immutable, so a
shallow copy per entry is enough. test_get_state_history_entries_are_copies
pins it; without the change it fails with 'tampered' == 'enabled'.

clear_state() mutated five shared dicts without holding _lock, while
every other mutator takes it. A concurrent set_state() could interleave
and leave a plugin with history but no state. Drop the five as one unit.

This does not close the wider unload-vs-worker race, which lives in
PluginManager.unload_plugin() and predates this change: an update worker
still in flight can call set_state() after clear_state() returns and
recreate the entry. Serialising that needs the per-plugin lock held
across worker join in unload_plugin(), which is a separate change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 09:45:17 -04:00
ChuckandClaude Opus 5 f90638a9ec feat(web): show available memory in Tools diagnostics (#500)
* feat(web): show available memory in Tools diagnostics

System Diagnostics reported memory as used-percent plus used/total GB.
Neither distinguishes a healthy board from one about to fail, because
page cache counts as used and is reclaimable on demand -- a Pi can read
70% used and be fine, or read the same and be minutes from trouble.

MemAvailable is the kernel's own estimate of what a new allocation can
actually obtain, and it is the number that tracked the failure on a 1GB
Pi 3B+: healthy running sat above 500MB, and the crash came at 73MB. By
that point fork() was failing, so sshd could not spawn a session and
systemd could not respawn the display, while the kernel carried on
answering pings at 0% loss. Used-percent gave no warning at any point on
the way there; available memory fell steadily for hours.

/api/v3/system/status now returns memory_available_mb from
psutil.virtual_memory().available, and Tools renders it as its own tile,
coloured against the thresholds that failure implies: red under 150MB,
amber under 300MB, green above.

The existing memory tile is left alone -- used/total is still what you
want when sizing a workload; this answers the different question of how
much room is left right now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): round available memory once, so the tile agrees with itself

The colour was classified from the raw value while the label was rounded,
and the API sends one decimal place. At the boundaries the two disagreed:
149.6 rendered as "150 MB" in red, and 299.6 as "300 MB" in amber -- each
contradicting the threshold its own colour claims to apply ("red under
150MB"). A reader checking the tile against the documented thresholds would
conclude the readout was broken.

Rounding once and using that number for both restores agreement. It moves
those two boundary cases up a band, which does not matter: the thresholds
come from a measured failure at 73MB, so which side of the line a spare
0.4MB falls on carries no information. The tile agreeing with itself does.

Null handling and the thresholds themselves are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 08:40:28 -04:00
ChuckandClaude Opus 5 333fd17d28 fix(memory): release fetched payloads once they have been delivered (#499)
* fix(memory): release fetched payloads once they have been delivered

BackgroundDataService kept the fetched body on the FetchResult it filed
in completed_requests, which is swept hourly and capped at 500 entries
by count. For a status record that costs nothing; for a season schedule
it costs a tenth of the board.

Measured on a 1GB Pi 3B+ with a 1-second RSS profile: the display
process sat at 404MB after plugin load, then stepped +21MB when NFL
fetched its season and +90MB when NCAA football fetched 946 games for
2026 -- and stayed at 494MB. Not a leak; a staircase that never came
down. When a later fetch landed while headroom was low, available memory
reached ~70MB, fork() began failing, and the board stopped being able to
start a process at all: sshd accepted connections and closed them before
its banner, systemd could not respawn the display, and the panel went
dark while the kernel carried on answering pings.

The cache-hit path was the worse of the two. It runs once per update
interval per sport, mints a fresh request_id each time, and files
whatever the cache returned. The memory tier is capped at 150 entries on
a 1GB board, so a miss re-parses the payload from disk into a genuinely
new object -- separate copies accumulating toward the 500-entry cap,
not shared references.

Releasing is safe: the payload is written to the cache under the
request's cache_key before the result is built, the callback is handed
the object directly, and consumers read it back from the cache
afterwards (the plugins' callbacks use it only in passing, to log a
count, before reading the cache). Nothing is lost -- it moves from RAM
to the disk cache that was already holding it.

Requests submitted without a callback keep their payload, since polling
get_result() is then the only way to collect it. That keeps the existing
contract, and the existing tests covering it, intact.

Not addressed here: max_workers=3 allows three concurrent fetches, so
three large parses can peak at once, and there is no in-flight dedupe by
cache_key -- a second submit for a key already being fetched starts a
second fetch. Both bound the transient peak rather than what stays
resident, and both are behaviour changes worth their own review.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(memory): file the cache-hit result before running its callback

Restores the original ordering. Releasing the payload after the callback
meant filing the result after it too, so a callback that queried
get_result() or is_request_complete() for its own request would not have
found it -- a behaviour change unrelated to the memory fix.

The dict holds a reference to the same object, so releasing after filing
still clears the payload.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: wait for the payload release, not just the filing

The callback test waited on is_request_complete(), which goes true as soon
as the worker files the result in completed_requests. The worker then runs
the cleanup pass, then the callback, then releases the payload. Both of the
test's assertions therefore raced the worker: `seen` is populated by the
callback, and `data is None` only after the release that follows it.

It passes today because a one-line callback usually finishes inside the
20ms poll interval. Confirmed by making the callback sleep 0.4s: _wait()
returns with seen == {} and the payload still resident.

_wait_for_release() polls for the released payload instead. Release happens
strictly after the callback returns, so a released payload also means the
callback has finished and one wait covers both assertions. Verified against
the same 0.4s callback.

_wait() stays for the other three fetch-path tests, which assert only what
is already true when the result is filed -- the success flag, the error,
and the cache write that happened during the fetch itself. Its docstring
now says so, so the next reader picks the right one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 08:32:08 -04:00
ChuckandClaude Opus 5 a4a55a23fc fix(web): reject non-finite JSON numbers instead of raising (#497)
* fix(web): reject non-finite JSON numbers instead of raising

POST /api/v3/config/dim-schedule with {"dim_brightness": Infinity} answered
500. So did /api/v3/errors/clear with max_age_hours, and /api/v3/config/main
with multiplexing or row_address_type.

json.loads accepts Infinity/-Infinity/NaN by default -- they are not valid
JSON, but Python's parser emits them -- and Flask's get_json passes them
straight through. int(float('inf')) raises OverflowError, which is neither
ValueError nor TypeError, so validation blocks that carefully caught those let
it past and Flask turned it into a 500.

The status code was not the real damage. dim-schedule answered with
CONFIG_SAVE_FAILED and suggested "Check file permissions on config directory"
and "Check available disk space" for what was an invalid number. Every one of
these sites already had a correct 400 response written; they just never
reached it.

NaN already returned 400, because int(nan) raises ValueError. That is why this
only ever showed up for the infinities, and why it survived: the obvious test
case passes.

OverflowError is now caught alongside ValueError/TypeError at the 27 sites in
this file whose try block performs a numeric coercion. An AST sweep confirms
no int()/float() of request-derived data is left outside a block that catches
it.

Verified end to end through Flask's test client rather than by reasoning about
the parser: all four routes returned 500 before and 400 after.

Tests: five Infinity cases (which fail against the previous except tuples),
two NaN cases pinned so narrowing the tuple cannot quietly break them, and a
check that ordinary input is not rejected by the widened guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* test(web): assert 400 exactly, and prove valid input is accepted

Both review points were right, and the first is the failure mode this file
exists to catch.

Accepting any 4xx meant a 404 would have passed. Renaming one of these routes
would have left the test green while it tested nothing -- the same "looks like
coverage, points somewhere safe" shape that hid the composer injections. Now
asserts exactly 400.

Both infinity signs are exercised for every route. int() raises OverflowError
either way, but only +Infinity was in the original report, and a guard that
special-cased the sign would have passed a one-sided test.

The valid-input test previously asserted "not a 400", which did not show what
it claimed: the mocked save path fails for any input, so that assertion held
whether or not validation had accepted the value. It now gives load_config a
real dict and stubs _save_config_atomic, so the endpoint reaches its success
response and the test can assert 200 -- which only happens if the value passed
validation.

8 of the 11 checks fail with OverflowError removed from the except tuples.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:44:23 -04:00
ChuckandClaude Opus 5 085fb93a87 test(logging): stop the location assertion matching the clock (#496)
test_location_toggle asserted that ":42" -- a bare colon plus the record's
hardcoded lineno -- is absent from a line formatted with include_location=False.
But every formatted line starts with an HH:MM:SS.mmm timestamp, so ":42" also
matches the clock whenever the minute or the second is 42. The test fails for
roughly 3% of runs with nothing wrong:

  2026-08-22 08:05:42.274 - INFO - test.logger - hello
                     ^^^ matches ":42"

Assert on the whole "module.funcName:lineno" token the format string actually
emits ('%(module)s.%(funcName)s:%(lineno)d') instead of a fragment of it. That
cannot collide with a timestamp, and it checks the thing the test is named for.

Confirmed by formatting a record stamped 08:42:42 -- both minute and second
colliding: the old assertion fails, the new one passes.


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:44:00 -04:00
ChuckandClaude Opus 5 c321b94085 fix(display): retry a plugin that is enabled but failed to load (#495)
* fix(display): retry a plugin that is enabled but failed to load

A plugin whose validate_config() returns False is treated as a hard load
failure. The API then reports enabled=true, loaded=false, error=null: the
plugin is simply absent, with nothing saying why. hockey-scoreboard sat in
that state on a live rig for four days.

The recovery path existed but could not be reached. _reconcile_enabled_plugins
computes to_add = desired - current, and a plugin that failed to load is never
in current, so it stays in to_add and would be retried. But the reconcile is
queued by _enabled_set_changed(), which compares only top-level `enabled`
flags -- and the edit that actually fixes such a plugin (enabling a league,
filling in an API key) is nested inside the plugin's own config section. No
top-level flag changes, so no reconcile is queued, and the save that should
have fixed it does nothing. Only toggling some unrelated plugin -- which does
change a top-level flag -- queues the global reconcile that recovers it.

Add a second gate: queue a reconcile when a discovered plugin is enabled in
config but absent from the running set.

It is deliberately narrow rather than "reconcile on any config change".
Reconcile calls discover_plugins(), a ~39-manifest filesystem scan, and it
runs on the render thread; doing that on every config save would trade this
bug for a frame hitch. Gating on plugin_manifests also keeps non-plugin
sections that carry their own `enabled` flag (schedule, display) from
queueing a reconcile they can never satisfy. In the steady state -- every
enabled plugin loaded -- the new check is False and costs nothing.

The same valid-but-unconfigured => hard-fail shape still exists in
text-display, youtube-stats, birdnet-go, ledmatrix-flights and
mqtt-notifications; this makes all of them recoverable without a restart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(display): snapshot the plugin mappings under their locks

Addresses the review finding on the cross-thread reads.

_enabled_plugin_not_running runs on the config-watcher thread and read two
mappings the render thread mutates. Catching RuntimeError was not a fix: it
turned a torn read into a coin flip between an unnecessary discovery scan and
a missed retry, which is the bug this PR exists to remove.

Both reads are now snapshots taken under the lock that guards their writes:

- plugin_manifests via a new PluginManager.discovered_plugin_ids(), which
  copies the ids while holding the existing _discovery_lock. Discovery
  rebuilds that mapping entry by entry, so an unsynchronised reader can see
  it half-populated.
- plugin_display_modes under a new controller lock, taken at the only two
  sites that mutate it (_register_loaded_plugin / _unregister_plugin).

The locks are never nested -- each snapshot is taken and released before the
next -- so this cannot deadlock against discovery, which holds _discovery_lock
while it rebuilds.

No cost on the per-frame path. Both mutation sites run during reconcile, which
is rare, and every hot-path read of plugin_display_modes is on the render
thread itself, same thread as the writes, so those stay lock-free.

Tests: the accessor returns a snapshot rather than a live view, and actually
takes the discovery lock (proved from a second thread, since an RLock is
reentrant on the owning one) so a later refactor cannot quietly drop it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(display): consume the reconcile request before serving it

Addresses the second review finding: a lost update on
_pending_plugin_reconcile.

The flag was cleared after a successful reconcile. Reconcile has already read
its config by that point, so a config change arriving mid-flight set a flag
that the trailing clear then erased -- a request that was never served, and
the newest config never reconciled. That is the same "my save did nothing"
symptom this PR exists to remove, so leaving it would have undercut the fix.

Consume the request before running it instead, and re-arm only on a retryable
failure. A change that lands during reconcile now stays set and is picked up
on the next pass.

The per-frame read stays lock-free. It is a fast path that can only produce a
false negative -- the watcher setting the flag just after it is read is seen
on the next iteration -- never a false positive that loses a request. The lock
is taken only when a reconcile is actually pending or a config change arrives.

Extracted _service_pending_reconcile() so the sequence is testable rather than
buried in run()'s loop; the review asked for a regression test that invokes
the subscriber during reconciliation, which is not reachable otherwise.

Tests: 4 new, covering a request racing in mid-reconcile, the quiet success,
the retryable-failure re-arm, and not reconciling when nothing is pending.
Two of them fail against the previous clear-after-success semantics.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:43:39 -04:00
ChuckandClaude Opus 5 5a1f121e6b Stop array-item secrets being wiped, and logging them (#493)
Three review findings from #485 that I missed when addressing that PR;
it has since merged, so they land here.

1. Array-item secrets destroyed by any unrelated save (data loss).

remove_empty_secrets recursed into dicts but let a list fall through to
the scalar branch and kept it verbatim. Lists merge by *replacement*, so
the blanks the masked form posts back went straight over the stored
array:

    stored   [{"name":"a","token":"REAL-A"}, {"name":"b","token":"REAL-B"}]
    posted   [{"name":"a","token":""},       {"name":"b","token":""}]
    merged   [{"name":"a","token":""},       {"name":"b","token":""}]
             -> both credentials gone

Same failure as the scalar api_key case fixed earlier, one container
deeper. Lists now prune element-wise, and a list with nothing real in it
is dropped so the stored one is left alone. Where one entry does change,
the new merge_secrets merges by index instead of replacing.

Two details the first attempt got wrong, both caught by existing tests:

- An emptied dict item must stay {}, not None. ConfigManager's
  _strip_secrets_recursive treats a secrets list as *parallel* to the
  regular one ({} = "item i has no secrets"); a None makes it stop
  looking parallel, and it then drops the whole key from the main config
  -- silently deleting the items' non-secret fields too.
- The incoming list's length wins. The regular config's list is
  authoritative about how many items exist, so preserving surplus stored
  entries would let the two fall out of step and make deleting an entry
  impossible.

2. Submitted credentials written to the journal (security).

save_plugin_config logged `Full config: {plugin_config}` at INFO and
`Config that failed: {plugin_config}` at ERROR. Both run before
separate_secrets, so plugin_config still held the values just typed into
the form. Now keys only. Swept the rest of web_interface/ and src/ for
the same shape -- these were the only two.

3. Restart banner kept stale wording.

showRestartPending() cleared the stored custom text but left the DOM
element alone, so a config save could show the previous update's
message. The default is read back from the server-rendered copy rather
than duplicated in JS, so the template stays the one owner of the string.

Verified: 556 passed, 1 skipped across the web suite. Mutation-checked --
reverting api_v3 fails the logging guard and the array-merge test;
reverting either half of the secret_helpers change fails the unit tests.
New end-to-end coverage drives the real endpoint, not just the helpers.


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:43:08 -04:00
ChuckandClaude Opus 5 6138a3cbef test(install): stop assuming pytest's tmp_path is on disk (#492)
test_returns_nothing_when_tmpdir_is_already_disk_backed asserted that
lm_disk_backed_tmpdir prints nothing when TMPDIR is already disk-backed,
and used pytest's tmp_path as the "disk-backed" directory:

    # tmp_path is on the regular filesystem, so the default must be kept.
    assert call("lm_disk_backed_tmpdir", env={"TMPDIR": str(tmp_path)}) == ""

That premise is false on the platform the helper was written for. Debian
13 mounts /tmp as tmpfs -- which is the entire reason lm_disk_backed_tmpdir
exists -- and pytest puts tmp_path under /tmp. So on the target platform
TMPDIR is memory-backed, the helper correctly answers /var/tmp, and the
test fails:

    E  AssertionError: assert '/var/tmp' == ''

The helper is right; the test was wrong. Reproduced on a box where
/tmp is tmpfs and / is ext4.

The test now looks for a directory whose backing store is actually disk
-- tmp_path, else a scratch dir under /var/tmp, else beside the library
-- using the same findmnt lookup the helper itself uses, and skips only
if no disk-backed directory exists anywhere. An earlier version of this
fix skipped whenever tmp_path was tmpfs, which made it skip on every
machine with a tmpfs /tmp; that is barely better than asserting the
wrong thing, so it now searches instead of giving up.

Verified: 31 passed, 0 skipped. Mutation-checked -- deleting the
"is the current TMPDIR memory-backed?" guard from lm_disk_backed_tmpdir
fails this test, so it still catches the regression it is there for.


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:42:51 -04:00
ChuckandClaude Opus 5 568cb6d77f perf(vegas): report the frame rate when it is worth reporting (#487)
Vegas logged an FPS line at INFO every five seconds for the whole of
every run. Measured over two hours on a rig: 1410 samples, 98.5% of them
within 10% of target. The 1.5% that were not included a reading of
8.6fps against a target of 60 -- a real stall, completely invisible
inside 1389 lines reading "59.6". INFO is now reserved for a shortfall,
the recovery from one, and a slow heartbeat so a healthy marquee still
shows a pulse. Scroll-progress tracing drops to DEBUG for the same
reason: it runs for the whole of every scroll and is what you turn debug
on to watch.

Three review findings, all fixed here.

1. Per-frame timing used the wall clock (critical). The loop sleeps the
   remainder of each frame budget:

       frame_elapsed = <now> - frame_started
       time.sleep(max(0.0, frame_interval - frame_elapsed))

   These devices have no RTC, so the clock jumps by however wrong boot
   time was when NTP first syncs. A backward step makes frame_elapsed
   negative, `frame_interval - frame_elapsed` then exceeds the whole
   budget, and the render loop stalls for the size of the correction. A
   forward step instead inflates the p99 and worst-frame figures this
   telemetry exists to report. Both per-frame timestamps are monotonic
   now. start_time stays wall-clock: it is only used for the iteration
   duration report, where a human-readable clock is the point.

2. FPS health state reset every iteration. last_fps_health_log and
   was_degraded were locals of run_iteration(), which is called once per
   cycle. Starting at 0.0 against a monotonic clock, `due` was true on
   the first sample of every iteration, so the 300s heartbeat degenerated
   into one report per cycle -- reintroducing the noise this change is
   about. A recovery that crossed an iteration boundary was never
   reported either, since was_degraded had already gone back to False.
   Both now live on the coordinator and reset in start().

3. The degraded threshold read as an off-by-one. 90% of target is
   deliberate -- a marquee jitters constantly, so "anything below target"
   would report forever and mean nothing -- but nothing said so, leaving
   55fps-against-60 looking like a missed case. The constant now states
   the band and gives that exact example.

Also drops two soccer logo PNGs that a `git add -A` had swept into the
first commit. They are unreferenced, unrelated to frame-rate telemetry,
and 210KB.

Verified: each fix mutation-checked -- restoring the wall clock on either
per-frame timestamp, or making the health state local again, fails the
new tests. 566 passed across the vegas, coordinator and scroll suites.


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:42:33 -04:00
ChuckandClaude Opus 5 6b74506695 fix(sports): fetch odds for the games shown, not the whole schedule window (#494)
* fix(sports): fetch odds for the games shown, not the whole window

SportsUpcoming.update() walked every upcoming game in the schedule
window and called _fetch_odds() on each one inside that collection
loop, narrowing to upcoming_games_to_show only afterwards. Each call is
a separate sequential ESPN request.

The comment sitting above it said odds were fetched "only for games that
will be displayed". The only narrowing it actually applied was
show_favorite_teams_only, which is not the default, so in the usual
configuration nothing narrowed it at all.

Measured on devpi, where the football plugin has the same shape:

  467 odds requests in one 35s burst, 467 distinct events
  315 NFL + 152 college-football -- roughly a whole season
  plugin football-scoreboard operation timed out after 30.0s

The burst repeats each time the 1h odds TTL expires: 67 -> 327 -> 957 ->
1261 requests/hour across four consecutive hours. Between expiries the
cache works and the rate is zero, so this is a thundering herd on
expiry, not a caching failure.

The fetch now runs after selection, over team_games -- the list already
cut to upcoming_games_to_show. This mirrors the fix the football plugin
already carries; the shared base class never got it.

SportsLive is deliberately left as it is: it walks the raw event list
because it has to find which games are live, but only fetches odds for a
game that has already passed the is_live/is_halftime test, so its
fan-out is bounded by how many games are actually in progress. The test
pins that distinction rather than assuming it.

The test reads the AST rather than the source text, and asserts the full
set of call sites, so a new one has to be classified deliberately
instead of inheriting whichever behaviour it happens to land in. Writing
it that way is what turned up the SportsLive site, which I had missed.

Verified: reverting the fix fails the test with the offending iterable
named ("iterates over 'events'"). 525 passed, 9 skipped across the sports
and odds suites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* test(sports): check the odds guard structurally, not by its text

Review caught that _guards_above() collected an `if` test even when the
call sat in that if's `else`, so moving _fetch_odds() into the else of
the is_live/is_halftime test would still pass -- while fetching odds for
exactly the non-live games the guard exists to exclude.

Verifying that turned up a wider hole in the same assertion. It matched
substrings of the *unparsed source*, so a negated condition satisfied it
too:

    if not (details["is_live"] or details["is_halftime"]):
        self._fetch_odds(details)      # every non-live game

Both names still appear in that text, so `"is_live" in guards` held and
the test passed on code doing the opposite of what it claims to check.

The guard test is now structural. It walks the AST for an enclosing `if`
whose *body* (never its `else`) contains the call, and whose test
references both names without either sitting under a `not`.

Verified by mutation: fetching odds for non-live games now fails with
"does not sit in the true branch of a test requiring the game to be in
progress". Moving the call into the else of the *favourites* test still
passes, which is correct -- the game there is still live, so the
in-progress contract holds and the fan-out stays bounded by how many
games are actually in play.

525 passed, 9 skipped across the sports and odds suites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 16:24:15 -04:00
ChuckandClaude Opus 5 5f29243e87 fix(config): make the device location the default for plugin location fields (#490)
* fix(config): make the device location the default for plugin location fields

A user in Kansas City reported their radar centred on Dallas, TX with
nothing in config.json to explain it.

The radar is the `ledmatrix-weather` plugin's `radar` mode, and it centres
on the same coordinates as every other weather mode: `forecast_data`
lat/lon, geocoded from the plugin's own `location_city` /
`location_state` / `location_country`. Those ship with schema defaults of
Dallas / Texas / US. A user who never opened the weather plugin's config
form therefore has no `location_city` on disk, and `PluginManager` merges
the schema default in at load time — so the whole plugin (not just the
radar) silently runs on Dallas. Radar is just the only mode that draws a
recognisable map and gives the mismatch away.

Meanwhile the device-wide `location` block that General settings writes
was read by nothing at all, despite its own help text promising it was
"used for weather, sunrise/sunset, and other location-based content".

`SchemaManager.generate_default_config()` now substitutes the device
`location` into the three fully-namespaced `location_*` keys before
handing defaults back, so the promise holds:

- Only `location_city` / `location_state` / `location_country` are
  substituted. A bare `state` key is left alone — `ledmatrix-elections`
  uses it for a two-letter code, and rewriting it would break that plugin.
- A value the user saved on the plugin still wins: this replaces the
  schema default, and `merge_with_defaults` puts user config on top.
- The substitution is applied on the way out of the defaults cache rather
  than into it, so changing the device location takes effect immediately.
- No config manager, no `location` block, or an unreadable config all
  fall back to the plugin's own schema defaults.

Every caller benefits: the plugin loader, the config form (which now
pre-fills the user's real city), config save, and reset-to-defaults.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GNLrSZ32FNKpHRaduKEJsg

* docs(web): name the exact plugin keys the device location seeds

Review follow-up. The General settings help text said the device location
was "the default for every plugin that asks for a city", which overstates
what the code does: only the fully-namespaced `location_city` /
`location_state` / `location_country` keys are substituted. A plugin with
a bare `city` key gets nothing — deliberately, since `ledmatrix-elections`
uses `state` for a two-letter code. The tips now name the exact keys.

Worth noting for anyone editing these: `ui.help_tip(...)` takes a
single-quoted Jinja string, so an apostrophe in the tip text has to be
escaped or written around. The wording here avoids them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GNLrSZ32FNKpHRaduKEJsg

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-21 16:22:48 -04:00
ChuckandClaude Opus 5 1fbe244e49 fix(plugins): say when discovery skips a directory (#489)
* fix(plugins): say when discovery skips a directory

A plugin can be enabled in config, enabled in plugin state, present on disk
with a valid manifest and an importable entry point -- and simply absent from
the running process, with nothing anywhere to say why.

That is not hypothetical. hockey-scoreboard on a live rig is enabled in both
places, imports cleanly when loaded by hand, and is listed in the Vegas plugin
order, but is not among the 22 plugins the process actually holds. Establishing
even that much meant comparing cache-file mtimes to find it had last run three
days earlier. The journal had nothing, because discovery does not report what
it declines to load.

Two paths were silent. A directory with no manifest.json was skipped without
comment, which is defensible until it is the thing you are trying to explain.
Quieter still, a manifest that parsed but carried no "id" was read
successfully and then dropped on the floor -- no warning, no trace, and the
plugin simply does not exist as far as the rest of the system is concerned.

Both now log a warning naming the directory and the reason.

This does not explain the rig above; its manifest has an id. It makes the next
occurrence diagnosable from the journal instead of from file timestamps.

Reverting the change fails both tests. 65 plugin-system tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(plugins): warn once per directory, not once per scan

Self-review catch. Discovery runs on every web UI page load and every config
reconcile, so warning unconditionally about an unloadable directory would put
a line in the journal each time someone opened a page -- the same log-volume
problem this change exists to help diagnose.

The skip is now reported once per directory per process. The diagnostic value
is unchanged: the reason a plugin is missing still appears in the journal,
once, where before it appeared nowhere.

Test added covering five consecutive scans producing one warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(plugins): one unusable manifest no longer aborts the whole scan

json.load accepts any JSON value, so a manifest.json holding null, [],
"text" or 42 parses without complaint and then raises AttributeError on
manifest.get('id'). Nothing catches that: the outer handler around the
scan takes OSError and PermissionError only.

So a single malformed manifest did not skip that one directory -- it
aborted _scan_directory_for_plugins outright, and every other plugin on
disk, however healthy, silently failed to register. Reproduced with
three directories, the middle one holding `null`:

    SCAN ABORTED -> AttributeError: 'NoneType' object has no attribute 'get'
      the two valid plugins never registered

That is the same failure this PR set out to fix, in its most severe
form: a plugin enabled in config, enabled in plugin state, present on
disk, and absent from the running process with nothing to say why --
except here it takes every other plugin with it.

A manifest that is not a JSON object is now skipped like any other
unusable directory, named once, with what it actually was:

    Skipping bad-null: its manifest.json is NoneType, not a JSON object
    Skipping bad-list: its manifest.json is list, not a JSON object
    scan returned: ['aaa-good', 'zzz-good']

Verified: removing the guard fails 6 of the 10 tests. Covers null, list,
string, int and bool, and asserts the healthy plugins either side of the
bad one still register.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 16:19:33 -04:00
ChuckandClaude Opus 5 863e4a1ecd consolidate(perf): cut SD writes, log volume, and metrics churn (#486)
* fix(plugins): one bad metrics cache entry should not stop every plugin

Caught live on a rig: every plugin failing, once each, continuously.

    ERROR - src.plugin_system.plugin_manager - plugin geochron operation failed:
    ResourceMetrics.__init__() got an unexpected keyword argument
    'consecutive_failures'

    ERROR - ... plugin text-display operation failed: ...
    ERROR - ... plugin news operation failed: ...
    ERROR - ... plugin odds-ticker operation failed: ...

with /api/v3/health reporting plugin_system: not_initialized while the display
process itself kept running and updating the panel.

`consecutive_failures` is a plugin_health field, not a metrics one.
get_metrics() does ResourceMetrics(**cached), which raises TypeError on a
single unrecognised key, and that exception escapes into plugin_manager and is
reported per plugin. One malformed cache entry takes the whole plugin system
down.

How a health-shaped record came to sit under a plugin_metrics key on that
machine is not established, and I could not finish the diagnosis: the rig went
back into its EIO failure mode partway through -- SSH resetting pre-banner,
systemctl unexecutable -- while the web API kept answering from RAM. Checked
before that: the cache files on disk are correctly shaped and separate, and
CacheManager.get() returns the right record for each key, so it is not a live
key collision. A restored backup mixing two machines' caches is the likeliest
explanation, and that rig had one restored onto it.

Either way the loader should not be brittle enough for the answer to matter.
plugin_health already repairs its records field by field rather than trusting
what is on disk; this does the same. Known fields are kept, unknown ones are
dropped and named once in the log so a genuine schema change stays visible
rather than being silently discarded, and a non-mapping entry no longer raises.

Keeping the known fields matters: discarding the record wholesale would throw
away real call counts and timings because of an unrelated stray key.

Mutation-checked: restoring ResourceMetrics(**cached) fails 6 checks, dropping
the whole record fails the field-preservation check, and dropping unknown
fields silently fails the logging check. 28 tests pass across the resource
monitor and plugin health suites.

* perf(health): stop rewriting a health record on every healthy cycle

Every successful plugin update called record_success(), which persisted the
record unconditionally. In steady state the only fields that had changed were
total_successes and last_success_time -- a counter and a timestamp that
health_monitor surfaces for display and that nothing reads back after a
restart. Nothing alerts on the age of last_successful_update; it is carried in
the metrics dataclass and shown.

Measured on a rig running 24 plugins, all steady-state (0 consecutive
failures, circuit closed): a five-minute sample caught 22 health-file
rewrites, about 4.4 a minute or 6,300 a day. Each write is ~400 bytes through
cache_manager.set(), which writes a file per call, so each one costs a
filesystem block plus an ext4 journal write.

That lands on an SD card, where the unit of cost is an erase-block cycle
rather than the bytes involved, and where wear is what eventually kills the
card. Two cards have already failed on the other rig with the same
signature -- unreadable block device, EIO on exec, sshd unable to read its
host keys.

The circuit breaker still has to survive a restart, so the write is kept for
exactly the fields it is rebuilt from: consecutive_failures, circuit_state,
circuit_opened_time, half_open_start_time. A failure, a circuit opening and a
recovery are all still written the moment they happen. In-memory state is
updated every time either way, so the health API and web UI show what they
always did.

Tested: 100 healthy cycles now perform zero writes after the first, the
counters remain accurate in memory, and a failure, a recovery and a
half-open-to-closed transition each still reach disk. One test kills and
rebuilds the tracker from the cache to prove the breaker's state genuinely
survives what is no longer written.

Mutation-checked both ways: persisting unconditionally again fails the
steady-state test, and widening _DURABLE_FIELDS to include last_success_time
fails it too. The 46 existing health tests pass.

(cherry picked from commit 14abea2d24)
(cherry picked from commit 0f77bd2345)

* perf(vegas): trace the content path at DEBUG instead of INFO

plugin_adapter narrates every step of acquiring content from every plugin --
"Has get_vegas_content", "Native: calling get_vegas_content()", "Native
content returned None", "Has scroll_helper", per-item sizes -- once per plugin
per cycle, all at INFO.

Measured on a live rig: 13,408 log lines an hour, of which 13,366 were INFO
and 35 were WARNING. Roughly 223 lines a minute of string formatting on a Pi
that is also driving the panel, written through journald to the SD card, with
the 35 lines that actually indicate a problem buried among them.

Top repeated messages in that hour:

    717  Scroll progress: elapsed=... total_scrolled=.../... px
    399  [plugin] --> INCLUDED in Vegas scroll
    323  [plugin] content_type=static, display_mode=fixed
    195  [plugin] Has get_vegas_content: True
    195  [plugin] Native: calling get_vegas_content()
    168  [plugin] Native: get_vegas_content() returned None
    168  [plugin] Native content returned None        <- the same fact, twice

54 logger.info calls in plugin_adapter become logger.debug, along with the
per-frame scroll-progress line in scroll_helper. Together those are 3,174 of
the 13,408 lines an hour, a 23% cut, and the ~3,600 odds-manager lines are
addressed separately by ledmatrix-plugins#300.

Nothing is lost: the 19 warning/error/exception calls in the module are
untouched, so real failures still surface at their own level. This is a
logging-level change only -- no control flow, no behaviour.

One INFO call is deliberate and stays. The padding-strip message picks its
level at runtime (`logger.warning if (left and right) else logger.info`) and
test_vegas_plugin_adapter.py pins that choice; it survives because it is not a
direct logger.info call site. That test still passes.

Mutation-checked both ways: reintroducing a single INFO trace fails the guard,
and demoting the warning/error calls along with the trace fails a second guard
written for exactly that mistake. 537 vegas and scroll tests pass.

(cherry picked from commit e496d95dfe)
(cherry picked from commit 8d1e43c15a)

* fix(logging): give the journal the real severity of each line

Everything this process writes to stdout reaches the journal as PRIORITY=6,
whatever the Python level was, because journald has nothing else to go on.
Measured on a live rig over 24 hours:

    lines containing " - ERROR - "      55
    lines containing " - WARNING - "    13
    journald PRIORITY recorded          6, for every one of them

So `journalctl -p err -u ledmatrix` returns nothing while errors are being
logged, and `-p warning` likewise. Triage falls back to grepping message text,
which is slower and unreliable: during this audit a search for "oom" matched
the radar logging "zoom=9" twenty-four times and briefly looked like the OOM
killer had been firing.

systemd reads a leading "<N>" on each stdout line and takes it as the priority
(sd-daemon(3)), so a formatter that prefixes one costs no dependency. Every
line of a multi-line record is tagged, not just the first -- the journal splits
them, and an untagged continuation reverts to the default, which would leave
the body of a traceback filed as informational while its first line was an
error.

Applied only when JOURNAL_STREAM is set, which systemd sets for services whose
output it captures. Run from a terminal, in the emulator or under pytest the
prefixes would be literal noise, and the file handler keeps the plain
formatter for the same reason.

Mutation-checked three ways: prefixing unconditionally fails the
outside-systemd test, prefixing only the first line fails the multi-line test,
and mapping ERROR to 6 fails the level mapping. 39 tests pass across the
logging suites.

(cherry picked from commit 780fca6365)

* fix(logging): let callers see through the journald formatter wrapper

CI caught what local testing could not: two existing tests in
test_logging_config.py assert that setup_logging() selected a
StructuredFormatter or a ContextualFormatter, by checking the console
handler's formatter directly. Wrapping that formatter to tag each line with
its syslog priority makes those assertions false.

They passed locally and failed on the runner because the wrapper is applied
only when JOURNAL_STREAM is set -- absent in a terminal, present in CI. An
environment-dependent break, which is the kind that gets shipped.

The wrapper now exposes the formatter it delegates to, and those two tests
look through it. They are about which formatter format_type selects, and that
behaviour is unchanged; only the object they have to reach for moved.

Verified both ways this time: 39 tests pass with JOURNAL_STREAM set and with
it unset.

* perf(plugins): stop rewriting a plugin's metrics file on every call

Plugin metrics were persisted to the cache inside monitor_call, so every
call by every plugin rewrote a small JSON file. Measured on a running rig:
one plugin's plugin_metrics file changed nine times a minute, with fourteen
such files active. Each is around 350 bytes, which on ext4 costs a 4KB block
plus a journal entry, so the cost is dominated by the write itself rather
than the payload. Cache writes accounted for essentially all of that device's
2.4 MB/min of SD traffic, on a card that wears out and has already failed
twice on the other rig.

Metrics cannot be de-duplicated the way health state can, because call_count
changes on every call and the timings usually do too. So they are rate-limited
instead: at most one write per plugin per 30 seconds.

The in-memory copy stays authoritative and exact -- a plugin's call_count is
still precise the instant after it runs. Only the cross-process snapshot the
web UI reads is delayed, and telemetry up to half a minute old is still a fair
description of a long-running plugin.

reset_metrics clears the throttle timestamp, so a reset is not left showing a
deleted key for the rest of the interval.

Extrapolating the sampled rate, this takes metric writes from roughly 126 a
minute to 28. Health persistence, the other half of the churn, is handled
separately in #475.

Verified by reverting the throttle: the churn test then reports 50 writes for
50 calls. 88 tests pass across resource monitor, plugin system and web API.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix: use a monotonic clock and only mark metrics persisted once written

Two review findings on the throttle, both right.

The interval compared wall-clock timestamps. These devices have no RTC, so
the clock jumps by however far off boot-time was the moment NTP first syncs
-- a forward jump would allow an early write, a backward one would stall the
snapshot well past the interval. time.monotonic() is not subject to either.

The timestamp was also recorded before cache_manager.set(). A set() that
raised would buy the next interval's silence without leaving a snapshot
behind, which is the one case where skipping the write is least affordable.
Recorded after the write lands instead, so a failure is retried on the next
call.

Verified by restoring the original ordering: the new test then reports one
write where two are expected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* Address all six review findings on the perf consolidation

CodeRabbit reported six; this is all six, checked against its own
"Actionable comments posted: 6" rather than against what I happened to
scroll past.

Three are real defects in the code:

1. _under_systemd() trusted the presence of JOURNAL_STREAM.

systemd publishes JOURNAL_STREAM as "dev:ino", and every child process
inherits it -- including one whose stdout has been redirected to a pipe
or a file. The variable outlives the descriptor it describes, so a
subprocess would decide it was talking to the journal and emit the "<N>"
priority prefixes as literal noise into that captured output. That is
exactly the noise the function exists to prevent. It now parses the pair
and fstats stdout, per systemd's own guidance, and returns False for
missing, malformed, mismatched, or unusable descriptors.

2. Cached metrics were not type-checked.

A dataclass does not enforce its annotations, so
ResourceMetrics(call_count="not a number") builds happily and only
fails later, deep inside monitor_call:

    TypeError: can only concatenate str (not "int") to str

Values are now coerced to their declared type at load, where there is
still a cache key to name in the warning, and a value that cannot be
coerced starts the plugin fresh instead of arming a delayed failure.
A numeric string is accepted rather than discarded -- a JSON round-trip
can widen an int, and that is recoverable.

3. The first metrics snapshot was skipped for the first 30s of uptime.

_persist_metrics used 0.0 as the "never written" default. monotonic() is
time since boot on Linux and systemd starts this service at boot, so
`now - 0.0 < 30` was true for the first half-minute of every run: the
throttle swallowed the very first write, the one that matters most after
a restart. The sentinel is now None and the interval is only applied when
a previous write exists.

Three are tests that could pass without testing anything:

4. test_health_write_churn's fake cache stored by reference, so the
   tracker kept mutating the object already in the store -- a record
   could look persisted when no write had happened, which is precisely
   what test_durable_state_survives_a_restart exists to detect. Both
   directions now deep-copy, like a cache that serialises to a file.
   Verified: disabling the one real cache write now fails three tests.

5. test_values_of_the_wrong_type_do_not_raise asserted only that a
   dataclass had been constructed, which was true with the bad value
   still in it. It now asserts the loaded metrics are usable -- the
   field is numeric, and arithmetic on it does not raise -- across four
   kinds of bad value.

6. test_vegas_log_volume counted "logger.error(" in the source text,
   which also matches comments, docstrings and string literals --
   including that module's own docstring, which names those levels. A
   real error call could be demoted with the tally unmoved. It now walks
   the AST, reusing the helper already in the file. Verified: demoting
   all 18 warning/error/exception calls now fails the test.

Verified: every fix mutation-checked by reverting it and confirming the
matching test fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 15:50:55 -04:00
ChuckandClaude Opus 5 9c0c0dc851 consolidate(web): credential exposure, secret loss, and the update path (#485)
* fix(web): stop /config/main handing out every credential it holds

The endpoint returned the raw config to anyone who could reach the port, and
this web interface has no authentication of any kind. An unauthenticated
request against a live rig returned:

    github.api_token                40 chars
    incoming-packages.ha_token     183 chars
    jellyfin-now-playing.api_key    32 chars
    ledmatrix-weather.api_key       32 chars
    on-air.mqtt_password             8 chars
    youtube.api_key                 20 chars
    youtube-stats.api_key           39 chars

A GitHub token and a Home Assistant long-lived token among them. Anything on
that LAN could read them.

The x-secret masking the plugin config endpoints use does not reach here: this
route never consults a schema, and core keys such as github.api_token have no
schema to carry the marker. Several of the fields above *are* tagged x-secret
in their plugin's schema and were still returned in full, which is what rules
out the schema route as the fix for this endpoint.

Credential-named fields are now blanked. Matching on the name is blunt, and
for a whole-config dump that is the right default: anything named like a
credential should not leave the process, and a new plugin adding a
differently-shaped secret is covered without anyone remembering to tag it.

Blanked rather than removed, and safe to blank: POST /config/main merges into
the freshly loaded config and writes only the keys it was given, so a client
that round-trips this response cannot erase a secret it never saw. The web API
suites confirm it -- 81 passing, unchanged.

On the test that matters: the first version of this suite exercised the two
helpers and nothing else, and reverting the single line that wires the
redactor into the route passed all thirty of them. A property asserted on a
helper is not a property asserted on the endpoint, and it is the endpoint that
is exposed to the network. The added test goes through the view function, and
it does fail on that revert.

This also corrects an earlier claim of mine. I reported that GET /api/v3/config
did not expose these values; that path 404s, so the check proved nothing. The
real route is /config/main and it exposed all of them.

* fix(web): stop an unrelated config edit from erasing a plugin's secret

Saving any field on a plugin's config form destroyed that plugin's stored
credential. On a rig with a weather API key, changing the city silently
emptied the key, and the plugin stopped working at the next fetch with no
indication why.

The path had no guard at any step. The config partial masks secrets before
rendering (pages_v3.py:740), so the browser posts them back blank; _parse_value
deliberately preserves "" for optional string fields; separate_secrets routes
that "" into secrets_config, which is a truthy dict; deep_merge writes it over
the stored value; save_raw_file_content persists it.

The blank does not even need the round-trip. merge_with_defaults injects the
schema's api_key default ("") into every save, so a client that never sends
the field at all still erases it. test_secret_count_message_counts_top_level_keys
was counting exactly that injected blank as a saved secret field -- the visible
edge of the bug, pinned as expected behaviour.

remove_empty_secrets() already existed for this, with seven unit tests and a
docstring describing this precise scenario ("clients will send those empty
strings back ... so that existing stored secrets are not overwritten with
blanks"). It was never wired into a call site. This wires it into both save
paths that merge into the secrets file.

A blank now means "unchanged" rather than "delete", which is the same contract
the helper's tests already describe. The cost is that a secret can no longer be
cleared by emptying the field; clearing needs its own affordance, since a
control that erases credentials as a side effect of ordinary edits is not one.

Verified by reverting the guard: the new round-trip test then fails with the
stored key read back as ''. 262 web tests pass with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop dumping the config and request headers to the journal

save_main_config logged its entire POST body and the full request headers at
ERROR on every save. The body is the configuration itself, and the headers
carry the session cookie, so a routine settings change wrote both to the
journal -- at a level that guarantees they survive any sane log filter.

The lines are leftover debug output: they say "DEBUG:" in the message while
calling logging.error, and they went through the root logger rather than the
module logger, bypassing the level configured for this blueprint.

Replaced with a debug-level line recording the shape of the request, which is
the part with diagnostic value. The local `import logging` went with them; it
shadowed a module-level import that was already there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop /config/secrets handing out every credential it holds

GET /api/v3/config/secrets returned config_secrets.json in full to anyone who
could reach the port, and this interface has no authentication. Probed against
a real rig it produced six populated credential fields: a 40-character GitHub
token, a 183-character Home Assistant token, and Jellyfin and weather API keys.
This is the second door onto the same credentials; #477 closes the first.

Masking the response alone would have been worse than the leak. The only
client fetches every secret, edits one field and posts all of them back, and
save_raw_file_content replaces the file wholesale -- so a masked GET followed
by the client's own save would write the mask over every credential the user
had not touched. That is why this was left open when the leak was found; it
needs both halves.

Read side: mask_all_secret_values(), which already existed for exactly this
endpoint -- its docstring names it -- and had never been wired to a call site.
It leaves empty values and YOUR_* placeholders alone, so a client can still
tell "set" from "not set" without being told the secret.

Write side: strip the echoed mask and blanks from the submission, then merge
onto what is stored, so "unchanged" means unchanged. The cost is that a secret
can no longer be cleared by blanking it; that wants its own affordance, since
a control that erases credentials as a side effect of saving an unrelated one
is not one.

Browser side: the token field is now left empty rather than filled from the
response. Filling it with the mask would have stored eight bullet characters
as the token the next time the user pressed Save, and filling it with the real
value is the thing being fixed. It reports whether a token is saved instead.

Verified end to end through the Flask endpoints, not the helpers. Reverting
the masking fails the leak tests; reverting the merge fails the preservation
tests; both halves are independently guarded. 278 web tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop reporting "no update" when the update check could not run

check-update returned update_available=False whenever git failed. The banner
is the only route to the update button, so a checkout git refuses to touch
looked exactly like a current one -- permanently, with nothing on screen to
act on and only a log line recording why.

The common cause is an install performed as root. scripts/install/one-shot-install.sh
clones into ${HOME}/LEDMatrix, never consults SUDO_USER, and contains no chown
at all, while its own error text suggests running the whole thing under sudo.
The result is a root-owned checkout, and on a rig this is what every git
command in it does:

    fatal: detected dubious ownership in repository at '...'

including the fetch this endpoint runs. Verified on real hardware rather than
assumed.

A failed check now reports check_failed with a message the user can act on --
for dubious ownership, the chown that fixes it. The banner shows that message
instead of hiding itself, with the update button suppressed since updating
cannot work until the cause is fixed. The success path is untouched.

This does not fix the installer, which is the real cause; it stops the symptom
being invisible. The installer needs SUDO_USER handling and a chown, and its
suggestion to run as root should go.

Reverting the endpoint change fails four of the five new tests; the fifth
guards the success path and correctly does not move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): stop the installer chmod stripping exec bits on every update

git tracks five scripts as mode 644 that first_time_install.sh then chmods to
755 (start_display.sh, stop_display.sh, the two install_*_service.sh, and
one-shot-install.sh does the same to first_time_install.sh). With
core.fileMode true, the default on Linux, git reports all five as modified
from then on, in files the user never touched.

The update button stashes local changes before pulling, so it is not blocked
by this. But it never pops that stash -- stash pop and stash apply appear
nowhere in the update flow -- so the mode change is stashed away and left
there, and the files revert:

    === file modes after the update button's stash ===
      664  first_time_install.sh      <- installer had made these 755
      664  start_display.sh
      664  stop_display.sh
      664  scripts/install/install_service.sh

So every web-UI update silently strips the executable bit from the installer's
own scripts, and leaves a stash entry holding the difference. start_display.sh
and stop_display.sh stop working from the shell afterwards.

A manual `git pull --rebase` over SSH fails outright, since nothing stashes for
it: "cannot pull with rebase: You have unstaged changes". That is the likely
source of the reports, since plenty of people update that way.

Tracking the five as 755 -- what they should always have been, as the
installer chmodding them attests -- removes the spurious mode change
entirely: nothing to stash, nothing stripped, no stash entry, and manual
pulls work.

The pull also passes --autostash, for the case the code explicitly tolerates:
when the stash fails it logs a warning and pulls anyway, and that pull is what
then fails. Autostash also pops what it stashes, which the manual stash does
not.

Note that `git add -A` after `git update-index --chmod=+x` silently reverts
the index to the on-disk mode, so the modes here were set by chmodding the
files themselves.

Regression test asserts the five stay tracked executable; reverting any one
of them fails it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): ask for the restart that makes an update take effect

The update button pulls new code and restarts nothing. There is no systemctl,
restart, reload or reboot anywhere in the 172-line git_pull handler -- it
stashes, pulls, installs changed requirements, re-removes plugins the user had
uninstalled, and returns "Code updated successfully."

Meanwhile both services go on running the code they loaded at boot. So the
display keeps rendering the old build, the web interface keeps serving the old
build, and the user is told the update worked. Nothing on screen suggests
otherwise, and the next reboot is what actually applies it -- whenever that is.

The affordance for this already exists: the restart-pending banner, raised
after main-config saves, with a Restart Now button wired to the display
service. A code update is a stronger reason to show it than a config save is.

The response now reports restart_required, and applyUpdate raises the banner
with wording for a code update rather than a config save. The banner's message
became a parameter and is persisted next to the flag, since it outlives the
page that raised it.

restart_required is only true when the pull actually moved HEAD. "Already up
to date" is a success too, and prompting after a no-op would train users to
dismiss the prompt unread.

This covers the display service, which is what the Restart Now button drives
and what users notice. The web interface still picks up its own new code on
its next restart; restarting it from inside a request it is serving is a
larger change than this one.

Reverting the flag fails the test that a pull which moved HEAD asks for a
restart. 290 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* Mask list-shaped secrets element-wise, close two vacuous tests

mask_all_secret_values treated any non-empty list as a scalar, so a
secrets file holding

  "accounts": [{"name": "a", "token": "tok-a"}, {...}]

came back as a single "••••••••". The caller could not see how many
entries existed, and the raw editor was handed a string where the file
holds an array. Recurse into lists in both _mask_value and _contains_mask.

Lists merge by replacement, not key-wise, so strip_masked_values now
drops a list outright if any element still carries the mask -- storing a
half-masked list would discard the untouched entries.

Two tests could pass without exercising what they claim to check:

- test_git_pull_resolution asserted modes only for paths git ls-files
  returned. A renamed or deleted installer target is simply absent from
  that output, so its mode was never checked. Assert every CHMODDED path
  is tracked first.
- test_config_secrets_masking never checked the POST status. A 500
  leaves the old file in place, which satisfies every assertion that
  follows. Assert 200 before reading the file back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 14:30:29 -04:00
ChuckandClaude Opus 5 71739d85d1 harden: cap malloc arenas, warn on unit drift, and grant the portal's sudo commands (#476)
* perf(systemd): cap glibc malloc arenas on the display service

Measured on a live rig 2.5 hours after start:

    RSS                          1030 MB
    Private_Dirty                 988 MB
    anonymous mappings > 10 MB       23
    largest        104, 79, 66, 63, 63 MB, on 64 MB-aligned addresses
    threads                           9
    cores                             3   -> glibc ceiling = 8 x 3 = 24 arenas

23 against a ceiling of 24, all 64 MB-aligned: these are glibc's per-thread
malloc arenas, not live objects. The data the process was actually holding
accounts for perhaps 15 MB -- the widest scroll strip observed was 35,746 x 64,
about 7 MB as RGB and the same again for its numpy mirror.

It is bloat rather than a leak: sampled four times over 135 seconds, RSS sat
between 990 and 1030 MB rather than climbing. glibc gives each allocating
thread its own arena, grows them to hold peak demand, and never gives them
back. A process that builds and drops large images across several threads is
exactly the shape that produces this.

The device had 59 MB free at the time, on 1845 MB total.

MALLOC_ARENA_MAX=2 trades a little allocator concurrency for that resident
memory. It is a tuning knob rather than a fix for a defect, so the rationale
and the measurements sit next to it in the unit file, and a test asserts they
stay there -- a bare environment variable invites removal by whoever meets it
next.

Two things this is NOT, both checked rather than assumed:

- Not an OOM problem today. A grep for "oom" in the service journal returned
  24 matches, all of which were the radar logging zoom=9 and zoom=7. The kernel
  OOM killer has not fired: dmesg has zero matches.
- Not currently capped by the unit's MemoryMax=85% either. That directive is in
  this file but absent from the unit actually installed on the rig, which
  reports MemoryMax=infinity, so nothing is enforcing a ceiling there.

The saving is unmeasured on hardware: applying it needs a service restart,
which blanks the panel, so that is the user's call rather than something to do
mid-audit. If p99 frame time regresses -- it sits at 18.4 ms against a 16.7 ms
budget for 60 FPS, so there is not much headroom -- raise the value rather than
remove it.

(cherry picked from commit 446207ffbc)

* test(systemd): pin the arena value instead of accepting a range

Review follow-up. The range check accepted 1, 3 and 4, so a change to 4 --
which hands most of the resident saving back -- passed a test whose whole
purpose is to notice that.

Pinned to the value the unit ships, in one named constant. Raising it is still
a legitimate response to a frame-time regression, but it should be a visible
edit here rather than silent drift, and the failure message says so.

Mutation-checked: changing the unit to 4 now fails.
(cherry picked from commit 73fff8d2d5)

* fix(startup): warn when an installed systemd unit has drifted from the repo's

Nothing re-applies systemd units after the first install. `git pull` -- which
is what the web UI's update button runs -- brings a new template into the
checkout, but no code in web_interface/ or src/ copies it to
/etc/systemd/system, and nothing anywhere runs `systemctl daemon-reload`. The
unit that actually runs is whatever first_time_install.sh wrote on day one.

So every hardening added to a unit is inert on existing installs, silently.
Measured on a live rig:

    installed  /etc/systemd/system/ledmatrix.service   2026-08-06
    template   systemd/ledmatrix.service               2026-08-19
    contents                                           differ

with the practical result that the MemoryMax=85% the repo's template specifies
was not being enforced at all -- `systemctl show` reported
MemoryMax=infinity. Anyone reading the template would reasonably believe the
service was capped.

Startup now compares each installed unit against its substituted template and
warns when they differ, naming install_service.sh as the remedy.

A warning, not an error, and deliberately not a silent rewrite: editing files
under /etc and restarting services is the installer's job, not something a
display process should do to a machine while it is booting. Making it fatal
would also brick every development checkout whose unit is legitimately absent
or hand-edited.

Comparison ignores comments, blank lines and ordering. The template carries
explanatory comments the installed copy will not have, and systemd does not
care about order within a section, so a literal comparison would warn on every
boot and be ignored within a week.

Mutation-checked three ways: never reporting drift fails, making it fatal
fails, and -- after the first attempt missed it -- comparing raw text now fails
too. That last gap is worth noting: the comment-insensitivity tests originally
exercised the helper directly, so a comparison that stopped calling the helper
passed them all. The test that catches it goes through _validate_systemd_units.

29 startup-validator tests pass.

(cherry picked from commit cf521bdfd8)

* fix(install): grant the sudo commands the captive portal actually runs

The installers write two allow-lists, /etc/sudoers.d/ledmatrix_web and
ledmatrix_wifi. Anything the code runs under sudo that is not in one of them
needs a password, which a service cannot supply, so the call fails.

Five commands were being run and none of them granted:

    sysctl -w net.ipv4.ip_forward=0|1     wifi_manager.py:788, 883
    nft add|delete table ip ledmatrix     wifi_manager.py:835, 895
    rfkill unblock wifi                   wifi_manager.py:1811
    iptables ...                          wifi_manager.py:796, 813, 818, 871
    mkdir -p .../dnsmasq-shared.d         wifi_manager.py:922

Together these are the captive portal: unblock the radio, bring up the AP,
add the redirect, turn on forwarding, and undo all of it afterwards. Without
the grants a hardened install would associate clients to the access point and
then fail to route them.

Why it has gone unnoticed: a stock Raspberry Pi image ships
/etc/sudoers.d/010_pi-nopasswd granting the default user

    <user> ALL=(ALL) NOPASSWD: ALL

which satisfies every one of these regardless of what the allow-lists say.
Confirmed on a live rig -- `sudo -n -l` permits sysctl there, and the blanket
rule is why. The allow-lists are effectively decorative on a default image and
only start mattering once that rule is removed or the service runs as another
user.

test_sudo_allowlist_covers_calls.py extracts every argv-style sudo call in
src/ and web_interface/ and asserts an installer grants it, so the next command
added without a rule fails here rather than on someone's hardened box.

Getting that test honest took three passes, each worth recording:

- Matching the literal "systemctl" against rules written as
  `$SYSTEMCTL_PATH enable ...` reported six gaps that did not exist. Binary
  path variables are now normalised before comparing.
- Scanning the whole installer let `NFT_PATH=$(command -v nft)` -- a variable
  definition, not a grant -- satisfy the check on its own, so deleting the
  actual nft rules still passed. Only NOPASSWD lines are considered now.
- `sudo -n <tool>` reported "-n" as the binary. sudo's own flags are skipped.

Each of the five grants is individually mutation-checked: removing any one
fails the suite.

(cherry picked from commit a372b43cd1)

* fix(install): drop the iptables wildcard, and pin each grant properly

Review follow-up. Two findings, both right, and the first is a hole I opened
myself.

`NOPASSWD: iptables *` is a root shell for the web user by another name.
`iptables --modprobe=/path/to/anything` runs that path as root, so a wildcard
grant on iptables escalates rather than restricts. I added that rule while
fixing a permissions gap, which is a worse outcome than the gap. It is gone,
and a test now fails on any trailing-wildcard grant to a tool that can execute
another program -- iptables, nft, tcpdump, find, awk, sed, perl, python, env.

The other finding: checking only the binary made the coverage test far weaker
than it looked. With `sysctl` present anywhere in the allow-list, deleting the
`net.ipv4.ip_forward=0` grant still passed -- and the portal would then be
unable to restore forwarding on teardown. Each required command is now matched
in full, and each is mutation-checked individually, including that exact
single-line case.

Scope pulled in deliberately. The first version of this test tried to assert
that *every* sudo call in the codebase is granted. Run honestly, it showed the
portal also runs iptables, nft, `ip addr`, `ip link` and `cp` with arguments
built at runtime -- an interface name, a port. Those cannot be granted safely
in a sudoers file: the rule needs a trailing wildcard, and that is the
escalation above. Closing that half needs a privileged helper that builds the
rules itself and takes only an interface and a port, granted the way
safe_plugin_rm.sh already is. That is a design decision, not a one-line grant,
so the test now pins the four commands this change actually grants and the
docstring says plainly what it does not cover.

Better a narrow test that is true than a broad one that is not.

(cherry picked from commit 500cfbc9f4)

* fix(install): pin PATH, keep unit order, and tighten the sudoers assertions

Three review findings, all correct.

The installer resolved binaries through an inherited PATH and wrote whatever
it found into sudoers as NOPASSWD grants. first_time_install.sh re-execs
itself with `sudo -E`, which preserves the caller's environment, so a writable
directory early in PATH turned a compromise of the low-privilege web user into
permanent root -- via a file the installer itself wrote. PATH is now pinned to
the system directories before anything is resolved, and every resolved binary
must be root-owned and unwritable by anyone else before it reaches the
sudoers file.

_unit_body() sorted a unit's lines before comparing. Order is not noise in a
systemd unit: repeated ExecStartPre=/ExecStartPost= run in the order they
appear, and a directive that moves between [Unit], [Service] and [Install]
means something different where it lands. The drift check reported no drift
for units that had genuinely changed. Order is preserved now.

Two of that check's own tests asserted the wrong thing --
test_reordered_directives_are_not_drift said so in its name -- and are
inverted, with a second covering a directive moved between sections. The
cosmetic-difference test now varies comments, blank lines and indentation,
which is what the installer actually drops, rather than reversing the file.

The sudoers assertions matched command prefixes, so
`sysctl -w net.ipv4.ip_forward=0 *` satisfied the requirement while granting
the caller arbitrary trailing arguments as root. They are exact now. The
wildcard check also normalises ${NFT_PATH} the same way as $NFT_PATH; the
brace is not a word boundary, so that spelling was skipped entirely.

Verified by reintroducing each: a widened required grant fails the exact
match, `${NFT_PATH} *` fails the wildcard check, and require_trusted_binary
refuses a non-root-owned, world-writable, or missing binary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 14:29:45 -04:00
ChuckandClaude Opus 5 cc258aaffd fix(install): tag the secondary installer's journalctl grants too (#491)
Review was right on all three counts, and the first is the one that matters:
scripts/install/configure_web_sudo.sh writes the same three wildcard
journalctl rules as first_time_install.sh and none of them carried NOEXEC. So
this PR closed the pager escape on one installer path and left it open on the
other, which is close to no fix at all -- a rig configured through that script
still hands out a root shell via less's "!command".

The test could not have caught it, for two independent reasons. INSTALLERS
did not list the file. And even listed, _grant_lines() kept the raw source
line: that installer echoes its rules, so each one ends in a quote rather
than the wildcard, and the trailing-* check skipped every one of them. Either
alone would have hidden it.

Both fixed: the file is covered, and an echoed rule is unwrapped to the
sudoers line it actually emits.

The selector test now covers -t ledmatrix as well. It asserted only the two
-u forms, so deleting the -t rule would have passed.

Verified by removing NOEXEC again from the secondary installer: four of the
six tests fail, where before the suite passed with the vulnerability present.


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-21 14:28:23 -04:00
Chuck fe5a3aa99d harden(install): tag the journalctl sudo grants NOEXEC (#472)
journalctl starts a pager when its output is a terminal, and from less a "!sh"
is a shell with whatever privileges journalctl was given. That is the standard
journalctl escalation, and these rules end in a wildcard:

    <user> ALL=(ALL) NOPASSWD: /usr/bin/journalctl -u ledmatrix *

Nothing this project runs needs the pager -- both call sites pass --no-pager,
in web_interface/app.py and api_v3.py. But a sudoers rule cannot require a flag
that sits in the middle of a command line, and reasoning about what a trailing
wildcard does and does not admit is exactly the kind of subtlety that produces
a hole. sudo's NOEXEC tag stops the command executing another program at all,
which closes it without depending on that reasoning.

NOEXEC works by LD_PRELOAD, so it applies to dynamically linked binaries.
Checked on the target hardware: journalctl there is dynamically linked. The
generated rules were run through `visudo -c` -- parsed OK.

Found while auditing the pre-existing wildcard grants, prompted by review
catching a far worse one I had added myself in the same area: `iptables *`,
where --modprobe runs an arbitrary path as root.

Reachability, stated plainly: on a stock Raspberry Pi image none of this
matters, because 010_pi-nopasswd already grants the default user
`ALL=(ALL) NOPASSWD: ALL`. It matters on a hardened install, or where the
service runs as a user without that blanket rule.

Two mutation checks: dropping NOEXEC from a rule fails, and deleting the rules
rather than tagging them fails too -- that second one matters, since "make the
test pass" and "remove the feature" would otherwise look the same.
2026-08-21 13:03:04 -04:00
ChuckandClaude Opus 5 10e75b977f Cover the next tier of untested modules and endpoints, and fix the 43 bugs that surfaced (#459)
* test(sync): cover the display sync protocol, and fix what that surfaced

DisplaySyncManager had no tests at all — it appeared in the suite only as
a MagicMock() stand-in, so none of its framing, handshake, or socket
handling was ever exercised. Writing that coverage surfaced three bugs.

Both receive loops caught the generic Exception and immediately retried.
A socket left in a bad state raises on every call, so the thread spun at
100% CPU logging the same line; the reverted-code run of the new
regression test takes 24 seconds where the fixed one takes 0.2. Both now
back off briefly before retrying.

The follower dispatched on `data[:8] == _RAW_MAGIC or len(data) > 512`.
That size threshold is not part of either wire format: a control message
over 512 bytes — a hello_ack carrying a long incompatibility error, for
instance — went to the image decoder and was dropped, and a raw frame
under 512 bytes went to the JSON parser. Both formats are already
self-describing, so dispatch on the magic prefix and treat a JSON parse
failure as the legacy unmarked PNG, with the shared frame bookkeeping
factored into _handle_received_frame().

_oversized_frame_warned was created on first use through
getattr(self, ..., False) rather than in __init__, alone among the
instance attributes.

75 tests: role parsing, the hello compatibility matrix, watchdog
timeouts, both receive loops, the TCP image server's length and
dimension caps and decompression-bomb guard, status shape per role, and
one end-to-end loopback handshake so the wire format is exercised for
real and not only against mocks.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(logos): cover LogoHelper, and stop bad downloads poisoning the cache

Nothing in test/ referenced logo_helper.py, so its caching, resizing and
download-fallback logic was entirely unexercised. Two bugs surfaced.

_download_logo wrote response.content to disk with no size limit and no
check that the bytes were an image. A logo URL is remote input, so the
response chose how much went into the assets directory; worse, an
undecodable one stayed there, and because load_logo() only reports the
decode failure and returns None, every later call re-read the same
corrupt file. The download path never retried, so a single bad response
made a logo permanently blank rather than falling back to the
placeholder. Cap the response, verify it decodes, and delete it if not,
which lets the existing fallback in load_logo_with_download do its job.

get_cache_stats() divided by self.cache_size with no guard, so a helper
built with cache_size=0 raised ZeroDivisionError from what is only a
stats call.

37 tests: size-qualified cache keys, LRU eviction and refresh, the four
load_logo_with_download paths, download permissions and timeout,
placeholder generation, and the abbreviation normalizer — including a
test pinning its deliberate divergence from
LogoDownloader.normalize_abbreviation, since logo filenames on existing
installs depend on both behaviors staying put.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(web): cover the error and response builders, and stop dropping empty values

errors.py and error_handler.py's response builders had no direct tests,
though every API response passes through them. Two bugs surfaced.

WebInterfaceError set suggested_fixes with `or`, so a caller passing []
to mean "I have no suggestions for this one" got the default list
instead. Only None should fall back.

create_success_response gated `data` on `is not None` but `message` and
`metadata` on truthiness, so an explicitly-passed "" or {} vanished from
the response while 0 and False survived — the response shape depended on
the value. api_helpers.success_response() then re-gated metadata the same
way, which is the path every api_v3 endpoint actually calls, so fixing
only the inner function would have changed nothing observable. Both now
use `is not None`.

That wrapper also merged request timing into the caller's own metadata
dict in place. A caller reusing a dict across requests would accumulate
previous responses' timings; it now copies before adding.

79 tests: category inference for every error code, mapped vs fallback
suggestions, the JSON shape including which keys are omitted when empty,
exception-to-code inference, and the success/error builders end to end.
Two behaviours are pinned as deliberate rather than fixed: an empty
context stays out of the response body, and from_exception's `message`
is the fixed per-code string, never the raw exception text.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(web): cover the input validators, and close three holes in them

validators.py had tests for dedup_unique_arrays only; the other eight
functions were untested. Three bugs surfaced.

validate_image_url checked for '..' only inside its relative-path
branch, so http://host/../secret passed validation while /../secret was
rejected — the traversal check now runs before the branch split, which
is where a safety check on the whole URL belongs.

validate_file_upload lowercased the uploaded filename's extension but
compared it against the caller's list verbatim, so allowed_extensions of
['.TTF'] rejected every valid .ttf file. Both sides are lowercased now.
The one in-tree caller passes lowercase already, so this only widens what
future callers can hand it.

validate_numeric_range accepted True and False, because bool subclasses
int; a boolean then compared as 1 or 0 against the range and validated
cleanly. Excluded explicitly, matching how base_plugin.py already handles
the same trap for display_duration.

84 tests. Two behaviours are pinned rather than changed:
sanitize_plugin_config deliberately does not HTML-escape strings, since
escaping at this layer would store the escaped form in config.json — the
docstring said "prevent injection", which read as a promise it does not
keep, and now says what it actually does. validate_font_awesome_class's
second 'fa-' check is unreachable behind its own regex; harmless, so
characterized rather than removed.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(api): cover wifi and registry endpoints, and fix bodyless POSTs

The /wifi/* routes drive the host's real networking and the registry
routes reach GitHub, and neither had endpoint-level tests. Covering them
surfaced a bug affecting six endpoints.

Six handlers read their body as `request.get_json() or {}`. The `or {}`
says every field is optional and a missing body should fall back to
defaults — but get_json() without silent=True raises UnsupportedMediaType
when there is no JSON Content-Type, and it raises before `or {}` is ever
evaluated. Each handler's catch-all then reported that as a 500. So
POSTing with no body — what curl sends by default, and what a fetch()
without options sends — failed on /plugins/store/refresh,
/display/on-demand/start, /plugins/config/reset,
/plugins/of-the-day/json/delete, /plugins/{id}/limits and
/plugins/authenticate/spotify. The shipped UI always sends a JSON object,
which is why this stayed hidden.

All six now use silent=True. test_api_v3_optional_body.py covers the
affected endpoints and adds a source check, since the combination of
`or <default>` with a non-silent read is self-contradictory wherever it
appears and is easier to catch by inspection than by exercising each
endpoint by hand.

Also adds test/_api_v3_test_helpers.py: the blueprint holds its managers
on a module-level singleton rather than in Flask app state, so a test
that mocks them leaks into every later test unless the originals are
restored. The existing _make_client() does this for unittest classes;
this is the pytest-fixture equivalent, for the five suites still to come.

69 endpoint tests: connect/disconnect/AP/radio including the string-aware
boolean coercion these endpoints deliberately use, the radio's
lockout-refusal path, registry refresh and fetch-from-URL, and a guard
that WiFiManager is never constructed for real.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(api): cover the music auth endpoints, and always clean up the wrapper

The Spotify step-2 handler writes a Python wrapper script to a temp file
with the user's redirect URL embedded in its source, then executes it.
That is the most dangerous shape in the blueprint and had no tests.

The wrapper was deleted in the success/failure branch and again in the
TimeoutExpired handler. Any other failure from subprocess.run — no
interpreter, a fork failure, an interrupted call — reached neither, and
left a world-readable temp file containing the user's redirect URL on
disk. Cleanup moves to a finally block, which is what "delete this
whatever happens" should have been from the start.

The injection tests are the point of this file. Eight adversarial
redirect URLs (embedded quotes, backslashes, newlines, triple quotes, a
full `"; import os; os.system("id"); "`) are each pushed through the
endpoint and the generated wrapper is parsed with ast: it must still be
valid Python, the URL must still be a single string literal bound to
redirect_url, and no os.system call may appear anywhere in the tree.
json.dumps holds up, but nothing was checking that it does.

40 tests. Also pins that the two endpoints are not symmetrical despite
the matching names — only Spotify has a two-step flow and a wrapper; YTM
runs its script directly — so a later change does not "restore" a parity
that was never there.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(api): cover the credentials upload, and stop it hoarding secrets

The endpoint that receives the user's Google OAuth credentials file had
no tests. Two bugs surfaced.

The OAuth-shape check ran inside `except Exception: pass`. A JSON
document that parses but is not an object — a bare 42, true, null, a
list — makes `'installed' not in creds_data` raise TypeError, which the
bare except swallowed, and the file was then written out as
credentials.json regardless. The check now decides the outcome instead
of being advisory, so anything not credentials-shaped is refused up
front rather than failing later inside the calendar plugin.

Every overwrite copies the old file to credentials.json.backup.<ts> and
nothing removed them, so a user who re-uploaded ten times had ten
complete sets of OAuth client credentials sitting in the plugin
directory, indefinitely. Keep the newest five. Pruning is housekeeping,
so a backup that cannot be removed logs and leaves the upload alone.

27 tests: size and extension limits, malformed JSON, the shape check,
0600 permissions on the written file, backup-on-overwrite, and pruning
including the repeated-upload case that stays bounded.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(api): cover the install endpoints, and make 14 dead guards reachable

/plugins/install and /plugins/install-from-url were tested only at the
PluginStoreManager layer, so the route logic — the queue-versus-direct
branch, schema invalidation, discovery, state and history recording — was
unexercised.

Covering them surfaced the wider form of the body-parsing bug fixed for
the `or {}` handlers in the previous commit. Fourteen handlers read
`data = request.get_json()` and immediately guard with `if not data:
return 400, 'No data provided'`. That guard cannot run: get_json()
without silent=True raises UnsupportedMediaType for a request with no
JSON body, so the catch-all answered 500 "an error occurred; see logs
for details" where the handler plainly meant to answer 400 and say
which field was missing. Every one of these endpoints told a caller who
simply forgot the body to go read the server logs.

All fourteen now use silent=True, so the guard each author already wrote
is the one that runs. This covers /config/raw/main and /config/raw/secrets
among them, whose own bodyless case had the same shape.

The two remaining bare reads are left alone: neither declares what a
missing body should do, so there is no stated intent to honour.

31 install tests plus 17 body tests. The install pair is checked against
each other rather than only individually — the same install logic is
written twice, once in the queue callback and once in the fallback, so
the tests assert both produce identical schema, discovery, state and
history effects. They agree today; the one difference is the success
message wording, which is characterized rather than changed.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(api): cover the raw config write endpoints

/config/raw/main and /config/raw/secrets write whatever JSON they are
given straight to config.json and config_secrets.json, bypassing the
secret-separation path the rest of the config surface goes through. Given
how carefully that surface keeps secrets out of config.json, the pair
that skips it was worth pinning precisely. Backed by a real
ConfigManager over tmp_path, so the assertions are against files on disk.

20 tests covering both routes: what lands in which file, that a raw
secrets write never touches config.json and vice versa, the GitHub token
reload, the uninitialized-manager and empty-body branches, and the
ConfigError path that carries config_path through to the response.

The bypass itself is pinned as intentional rather than changed — these
back the raw JSON editor, so writing the body verbatim is the feature.
The test says so explicitly, because the failure mode is someone later
routing plugin config through here as a convenience and silently losing
secret separation.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(api): cover backup restore and path containment, and fix restore scope

Restore is the most destructive thing the web interface can do — it
overwrites config, secrets, WiFi settings and fonts, then reinstalls
plugins — and neither it nor the file routes beside it had tests.

A malformed `options` field fell back to {}. Every RestoreOptions flag
defaults to True, so a caller who asked for a narrow restore and
mis-serialized the request got a full one instead, secrets included, and
was told it succeeded. Valid JSON that is not an object was worse:
`"null"` or `"[1,2]"` reached .get() on a non-dict and raised, so the
request died as a generic 500. Both are now refused with a 400 that says
what was wrong, and restore_backup is never reached.

The other file routes take a filename straight out of the URL and turn it
into a path — one to read, one to unlink. _safe_backup_path is the only
thing keeping those inside the export directory, and it was untested. No
bypass was found; the thirteen traversal shapes are pinned so a later
loosening of that pattern has to argue with something. The delete route's
by-name enumeration is covered too, including that a directory sharing a
backup's name is not removed.

84 tests. Two behaviours are pinned as intentional: a failed plugin
reinstall turns the whole restore into an error even though file
restoration succeeded, and omitting `options` entirely still means
restore everything — that is the documented default, and it is only the
mis-serialized case that was wrong.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* ci: raise coverage floor to 52%

Measured 54.45% after the Tier 1 and Tier 2 suites, up from 50%. Keeping
the same two points of headroom the 45 -> 48 ratchet used.

The modules this branch set out to cover: sync_manager 0 -> 97%,
logo_helper 0 -> 98%, errors and error_handler 0 -> 100%, validators
0 -> 97%. api_v3 moved less in percentage terms because it is 4,341
statements, but the endpoints covered are the destructive and
credential-handling ones.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(sync): probe for a free port on loopback, not every interface

CodeQL flagged the ephemeral-port probe in the handshake test for
binding to all interfaces. The probe only needs a free port number, so
loopback is both sufficient and correct — a test should not open a port
to the network to discover one.

The manager under test still binds to all interfaces, which is
deliberate and already marked nosec: a follower has to receive the
leader's UDP broadcast.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: bound the logo download, and stop malformed input reading as a fault

Review findings on the coverage branch.

The download size cap I added checked len(response.content), which has
already buffered the whole body -- it stopped the bytes reaching disk but
not memory, which was the point. A server that omits Content-Length and
never stops sending would still exhaust the process. Stream it instead,
counting as it arrives, into a sibling .part file that is replaced over
the target only once it decodes. A transfer that dies midway now leaves
nothing behind rather than a truncated logo for load_logo() to cache.

The follower's control-message handler caught three exception types, but
two reachable UDP payloads raise others: a bare JSON scalar makes
msg.get() raise AttributeError, and an "sx" carrying a non-numeric x
raises ValueError or TypeError from float(). Those escaped to the outer
handler, skipping the legacy-PNG fallback and -- since this branch added
a backoff there -- charging one malformed packet a 0.1s stall on the
receive path. The legacy-PNG path also decoded without the dimension cap
its TCP counterpart applies, so a crafted 65KB frame could force a large
allocation on the render thread; both paths now share one constant.

Three repo_url handlers called .strip() on client input without checking
it was a string, so {"repo_url": 12345} answered 500. The credentials
upload parsed the same file twice, the second time inside a bare except
that a preceding parse had already made unreachable. And both raw-config
handlers kept a json.JSONDecodeError arm that get_json(silent=True) had
turned into dead code, collapsing "sent something unparseable" into "sent
nothing" -- they now say which.

Two of the new tests were not testing what they claimed. The pruning
round-trip wrote ten backups inside one second, so all ten landed on the
same int(time.time()) filename and overwrote each other; it never reached
the limit it asserted. And the sync clock helper patched attributes on the
stdlib time module, freezing time process-wide for every daemon thread
earlier tests had left running.

Full suite: 3352 passed, coverage 54%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(sync): probe broadcast by sending, not by listening

The broadcast check added in the previous commit bound INADDR_ANY to
receive its own probe datagram, and the free-port probe did the same to
pick a port. CodeQL flagged both, correctly: a test suite has no reason
to open a socket the whole network can reach.

Sending is enough for what the probe is actually for. An environment
that refuses broadcast raises on sendto, which is the case that occurs
in sandboxes and is the one worth skipping over; confirming delivery
would have required the listening socket. A network that accepts the
send and silently drops it still reaches the assertion, exactly as it
did before either commit. The port probe binds loopback -- it only needs
a number, and the manager's own bind is the one that has to succeed, with
the retry loop already covering a port taken elsewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: keep callback faults out of the frame-decode fallback

Review follow-up on the previous two commits.

Widening the control-message except tuple put the callback dispatch
inside it, so an _on_new_cycle() that raised ValueError, TypeError or
AttributeError sent a perfectly good control packet to the legacy PNG
decoder -- which reported it as an image decode error and buried the
real fault. Split the two: whether the payload parses as JSON decides
frame vs control message, a second guard covers reading the fields of an
attacker-shaped body, and the callback fires outside both. It still
cannot kill the receive thread; the loop's own handler catches it, and
now says what actually went wrong.

The logo download's temp file was a fixed "<name>.part". Two plugins
asking for the same logo at once would interleave writes into it,
publish the mixture, or delete each other's partial. mkstemp gives each
download its own name in the same directory, so os.replace stays atomic.
Its descriptor is adopted by fdopen before the request runs, since a
request that raises before the write would otherwise leak the fd --
quietly, because load_logo_with_download swallows that.

Two test fixes: the oversized-frame test replaced PIL.Image.open
process-wide, the same hazard the clock helper documents, and Ruff B007
on an unused loop variable.

Full suite: 3355 passed, coverage 54%.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test(sync): cover the announce loop, and reject non-finite scroll positions

Three review findings from the follower receive path.

Non-finite scroll x reached follower rendering. json.loads accepts the
bare NaN/Infinity literals and float() accepts them as strings, so
"x": NaN arrived as a real float and was stored verbatim. NaN loses
every comparison the scroll code makes, so a follower given one sits on
a position it can never advance past. It now raises through the existing
malformed-control-message guard, which logs and drops the packet and
leaves the last good position in place.

_broadcast_available() only proves the host accepts sendto() for a
broadcast; a network that accepts the send and drops the packet would
let TestRealSocketHandshake run to its five-second deadline and fail on
assertions the code did not break. The deadline now distinguishes the
two: if not one packet crossed in either direction, that is the
environment, and the test skips rather than reporting a protocol
failure.

That skip could hide a real regression in the announcing side, so
TestFollowerAnnounceLoop covers it on mock sockets, where no network is
involved and nothing can skip: hello carries this display's hardware
config and goes to the broadcast address, heartbeats follow, an empty
hardware config falls back to 32x64x1, hello is not resent before its
interval, and a send failure is swallowed rather than killing the loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-21 11:16:52 -04:00
Chuck cf0a551f7b fix(web): stop checkbox groups posting back options they cannot show (#465)
The enum that lets a checkbox group draw its options is also what validates
the saved value. When an option goes away -- a league retires a team code, a
schema drops a choice -- a config still holding the old value has no checkbox
to render for it, but the value stayed in the hidden _data input anyway:
that input is seeded from the stored array and only rebuilt by
updateCheckboxGroupData() on change.

So the stale value was posted back on every save the user did not happen to
touch that widget for. The schema rejected it and the save endpoint returned
400 CONFIG_VALIDATION_FAILED, which blocks editing *any* field on that
plugin until the user works out which invisible entry is at fault -- with
nothing on screen naming it, because the offending value is precisely the one
with no checkbox.

Runtime was never affected: load_plugin() treats schema violations as
warn/degrade, and a retired code already matched nothing. Only the web UI
blocked.

Values not in the enum are now dropped before the hidden input is seeded, and
listed above the group so the selection is not lost silently. Only when the
widget has options -- an empty enum means there is nothing to check against,
and filtering on it would wipe the field.

This is not hypothetical. ledmatrix-plugins #212 ("correct team abbreviations
so config save no longer 400s") and #234 (removed the retired NHL code UTA
from a picker across four plugins) are both this failure mode, fixed one
league at a time. Nine shipped plugins use checkbox-group today; all of them
get the fix.

Tested by rendering the checkbox-group block lifted out of the shipped
template, following test_enum_option_labels.py, so the tests exercise the
production expression rather than a copy. Mutation-checked: removing the
filter fails 2 tests, filtering unconditionally fails the empty-enum test,
and dropping the notice fails the one asserting the value is named.
2026-08-19 17:40:52 -04:00
ChuckandClaude Opus 5 9018fa23cd fix(web): verify the onboarding timezone step, don't compare it to the default (#462)
* fix(web): verify the onboarding timezone step, don't compare it to the default

The Getting Started card's timezone step ticked when the saved timezone
differed from the value config.template.json ships (America/New_York),
OR-ed with the saved city differing from Tampa. Both halves were wrong.

"Differs from the default" answers "did somebody edit this?", but what the
checklist needs to know is whether the value is right. A user genuinely in
America/New_York could never satisfy it, so the card nagged forever with
four of five steps done -- the case that prompted this, on a panel whose
timezone was correct all along.

The city half was worse than useless: the saved city says nothing about
whether the timezone is set, and because the two were OR-ed, saving a city
ticked the step off with the timezone still wrong. That is the direction
that actually breaks displays, since event times then render in the wrong
zone.

The browser already knows its own zone, so compare against that. No new
persisted state, no network, and it catches the reverse case the old test
got backwards: a panel still set to the old zone after a move now stays
unticked, where before it ticked the moment the value stopped being the
default. Zones are compared by the wall-clock time they produce for one
instant rather than by identifier, so aliases (Asia/Calcutta vs
Asia/Kolkata, Europe/Kiev vs Europe/Kyiv) don't read as a mismatch. When
they genuinely differ the step names the browser's zone, so an unticked box
says why. Configs with no timezone, an unparseable zone, or a browser
without Intl leave the step open for the existing manual tick.

The step still deep-links to the General tab, and the location value stays
visible in its label -- it just no longer votes on whether the timezone is
configured.

Tests render the partial across configured zones and both cities: the step
never pre-ticks server-side, carries the configured zone for the client to
check, is unmoved by the city, and the panel-size step still resolves
server-side. Reverting the template fails 9 of the 11.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): compare zones on fields Intl has always had

dateStyle/timeStyle are late additions to Intl -- Firefox shipped them in
91 -- and an implementation that does not know them ignores them and
formats the date alone. The comparison would then read New York, Chicago
and Madrid as the same zone and tick the step for a timezone that is
plainly wrong, which is the failure the check exists to catch. Silent, and
only on older browsers.

Explicit numeric fields (year/month/day/hour/minute) have been in Intl
since ECMA-402 v1, so there is nothing left to degrade to.

The options look like a stylistic choice, so a test pins them: it reads the
comparison with comments stripped -- the comment names dateStyle to explain
why it is not used -- and fails if either style option comes back or a
time field is dropped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix(web): sample both sides of DST when comparing zones

CodeRabbit caught this and it is right: comparing the wall clock at one
instant treats zones that merely coincide right now as the same one.
America/New_York and America/Lima hold the same offset all winter, so a
panel set to the wrong one of the two ticked the step in January and then
ran an hour off from March -- a silent false pass, which is the failure the
whole check exists to prevent. Same shape as the dateStyle problem in the
previous commit: a comparison coarser than it looks.

Three instants now, all of which must agree: now, and mid-January and
mid-July of the current year. Those sit either side of DST in both
hemispheres, so only zones that agree year-round match. Toronto still
matches New York, which is correct -- either renders the same times.

Two tests. A static one asserts the comparison samples more than the
current instant, since reverting to `[now]` looks like a simplification.
And a table pinning which pairs must count as the same zone: aliases and
same-rule zones equal, seasonal coincidences (New York/Lima,
Phoenix/Los_Angeles, Sydney/Guadalcanal) not. That table mirrors the
algorithm rather than executing the shipped JS -- there is no JS runtime
here and the repo has no JS test infra -- so it records the verdicts the
browser code has to reach, and the static guard keeps the two aligned.

Mutation-checked: reverting to a single instant fails the static guard.

Also documented what the city test compares. CodeRabbit read it as always
failing, on the grounds that the label differs between Tampa and Seattle.
It does, but timezone_step() returns the opening tag only, so the
comparison is over data-done and data-tz and the label is not in it. The
assertion is left as an equality over the whole tag, which is stronger than
checking the two attributes by name; the docstring now says so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 12:49:46 -04:00
ChuckandClaude Opus 5 0c5b9c57d3 fix: keep low-memory boards reachable under load (#464)
* fix(service): survive corrupt health cache and clean exits

Three independent failure modes that each end with a dark panel and no
automatic recovery.

1. PluginHealthTracker._load_health_state returned the cached value
   verbatim. If that value is not a dict, every caller raises
   AttributeError: 'list' object has no attribute 'get' — during
   DisplayController.__init__, so the process dies before the display
   loop starts. systemd restarts it, the same bad entry is read back
   from disk, and it dies again: an unattended restart loop that
   survives reboots because the cause is persisted. Observed in the
   field with plugin_health:<id> holding an unrelated plugin's list
   payload. Now non-dict entries are discarded with a warning and the
   defaults are rebuilt.

2. ledmatrix.service used Restart=on-failure, so any exit with status 0
   left the unit stopped and the panel dark indefinitely — systemd
   treats it as success and never brings it back. Restart=always.

3. ledmatrix-wifi-monitor.service used StandardOutput=syslog, which
   systemd has marked obsolete; it warns and rewrites it to journal on
   every load.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* perf(memory): size the cache to the board and stop reinstalling deps

On a 1GB Pi 3B+ the display process settles around 600MB RSS of 905MB
total. When the remaining headroom runs out the failure is not a clean
crash: fork() starts returning ENOMEM, so sshd accepts connections and
closes them before its banner, timer jobs stop running, and the panel
goes dark, while already-resident processes keep serving normally. The
board looks healthy from outside and cannot be logged into. Only a power
cycle clears it.

Three contributing causes:

- MemoryCache had a fixed 1000-entry ceiling. Entries are parsed API
  payloads of tens of KB, so one ceiling cannot serve both a 512MB Zero
  2 W and an 8GB Pi 5. Now scaled from MemTotal (150 entries at <=1GB,
  1500 at >=8GB), overridable with LEDMATRIX_CACHE_MAX_ENTRIES.

- requirements_are_satisfied() returned False for any requirement with
  extras, so a plugin depending on python-socketio[client] re-ran pip on
  every single start: ~8s, a network dependency, and a 100-200MB spike
  at the least convenient moment. During a restart loop it repeats for
  each restart. Extras are now resolved one level deep against installed
  metadata, keeping the conservative "anything unverifiable falls
  through to pip" contract.

- ledmatrix.service had no memory ceiling. MemoryMax=85% expressed as a
  percentage so one unit file suits every board. Note this needs the
  memory cgroup controller, which Pi firmware disables by default;
  first_time_install.sh now adds cgroup_enable=memory to cmdline.txt,
  and the unit file documents how to verify it took effect.

first_time_install.sh also enables persistent journald storage (capped
at 64M). Default storage is volatile, so every reboot destroys the logs
that would explain why the board rebooted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: guidance for 512MB and 1GB boards

Documents the memory ceiling on small boards and, more usefully, what
running into it actually looks like: sshd accepting connections and
closing them before the banner, the web UI still responding normally,
clean ping, a dark panel, and a wrong clock after the next boot. None of
those read as "out of memory", which makes the failure hard to identify
from the symptoms.

Cross-referenced from SSH_UNAVAILABLE_AFTER_INSTALL.md, since "I can't
SSH in any more" is how most people will first meet this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address review findings on the low-memory work

Nine CodeRabbit findings, five in code.

**Health state (the one that matters).** The non-dict guard did not cover a
dict missing fields the callers index directly, which is the shape actually
seen in the wild: a record carrying only circuit_state produced
`plugin clock-simple operation failed: 'circuit_state'` about fifty times a
minute with the panel frozen. The record is now completed against the
defaults per field rather than trusted or discarded wholesale. Per field
matters: a first pass rejected any incomplete record outright, which reset a
tripped breaker and real failure counts to healthy because one optional
field was absent -- an existing test caught it. Values of the wrong type
(a counter persisted as a string, an unknown circuit_state) fall back
individually, valid neighbours survive, and newer fields the schema has
grown since (degraded, degraded_reason) are carried through untouched.

**Cache ceiling.** MemoryCache.set() accepted entries without bound between
cleanup sweeps, which run every 300s by default, so a burst could take the
cache far past max_size -- the unbounded growth the limit exists to stop.
Eviction now runs under the same lock on every write, sharing one helper
with the periodic sweep so the two cannot drift.

**Installer, cgroups.** Only cgroup_enable=memory was checked, so a board
carrying that without cgroup_memory=1 reported success and got no change,
leaving MemoryMax= inert. Each parameter is now checked and appended
independently; verified against all four combinations, single line preserved.

**Installer, journald.** Persistence was inferred from /var/log/journal being
non-empty, which proves neither Storage=persistent nor a size cap -- the
directory survives a switch back to volatile. The effective configuration is
read instead (systemd-analyze cat-config, falling back to the conf files),
and an explicitly configured SystemMaxUse is preserved rather than
overwritten. Verified across volatile, persistent-without-cap,
persistent-with-user-cap, cap-without-storage, and commented-only configs.

**Dependency extras.** _extras_are_satisfied stopped at one level, so a
gated dependency that itself requests an extra (requests[socks]) passed on
the base distribution's version while the extra's own dependency was
missing, and pip was skipped. It now recurses, with a visited
(distribution, extras) set so a cycle terminates.

Docs: both kernel command-line paths documented (the installer falls back to
/boot/cmdline.txt), daemon-reload and restart added after the systemd
override example, memory exhaustion added to the SSH summary with its
power-cycle-only recovery, and a language on the fenced block for MD040.

Tests: five for the health-state repair including the exact wild shape and
that record_failure/record_success no longer raise against it, and one for
the cache ceiling. Both mutation-checked. Full suite 2927 passed, with the
one pre-existing tmpfs failure that also fails on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* fix: harden the health-state repair and confirm journald took effect

Second review round; all three findings were valid and two were bugs in the
repair added last commit.

The repair could raise out of itself. An unhashable circuit_state (a list or
dict on disk) hit `value in {...}` and raised TypeError -- from the code
whose whole job is to stop a malformed record crashing the caller. It now
requires a str before the membership test.

bool is a subclass of int, so True passed the timestamp check and then
compared as 1.0: enough to expire a cooldown the instant the breaker opened,
while False would stop the elapsed check firing at all. Timestamps now
exclude bool explicitly.

The regression test for the original crash was seeded with a record that
*contained* circuit_state, so it passed against the old raw-return behaviour
too -- the counters are read with .get(), so circuit_state is the only field
whose absence used to raise. Reseeded to omit it, and it now fails against
raw-return as intended.

journald: drop-ins apply in lexical order, so a local file sorting after
ledmatrix-persistent.conf still wins and writing ours proves nothing. The
effective Storage is re-read afterwards and a warning naming the diagnostic
command is printed if persistence is still not active, rather than reporting
a success that was not verified.

Full suite 2934 passed, same single pre-existing tmpfs failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 12:28:22 -04:00
ChuckandClaude Opus 5 9083df9f5c fix(pixlet): resolve the release tag correctly when downloading (#461)
* fix(pixlet): resolve the release tag correctly when downloading

Starlark apps render through the pixlet binary, and the installer that
fetches it silently produced nothing, so every app failed with "Pixlet
not available - Starlark apps will not work".

Two compounding defects:

The version lookup parsed the wrong token. GitHub returns the release
JSON on a single line, so `grep '"tag_name"'` matches the whole document
and the greedy `sed 's/.*"([^"]+)".*/\1/'` captures the LAST quoted
string in it. That resolved to "mentions_count", giving a download URL
for a release that does not exist. The `[ -z "$PIXLET_VERSION" ]`
fallback never fired, because the value was not empty -- just wrong.

And `curl -L -o` without `-f` writes a 404 body to the file and exits 0,
so the download was reported as successful and the first sign of trouble
was tar complaining "not in gzip format" about a page of HTML:

    → Downloading linux-arm64...
      Extracting...
    gzip: stdin: not in gzip format
    ✗ Failed to extract archive: .../pixlet_mentions_count_linux-arm64.tar.gz
    Download complete: 0/1 succeeded

Now the tag field is matched directly and the value taken from it, and
the result is checked for a version shape rather than merely being
non-empty -- a wrong-but-non-empty value is exactly what made this
silent. curl gets -f so an HTTP error is a failure, and the archive is
gzip-tested before extraction, since a proxy can return 200 with an
error page.

Verified on an arm64 rig: v0.53.1 resolved, 1/1 downloaded, the binary
runs, and the plugin's own detection finds it at
bin/pixlet/pixlet-linux-arm64.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(pixlet): anchor the version check, and don't echo response bytes raw

Both CodeRabbit findings were valid.

The shape check accepted partial matches, so "v0.53garbage", "0.53" and
"v0.5" passed it and built a download URL for a release that cannot exist
-- the failure the check was added to stop, just one step later. Anchored
at both ends now. Every tronbyt/pixlet release to date is vX.Y.Z (all 38
verified against the API), with an optional suffix left for a future -rc.1
or +build tag.

The invalid-response diagnostic printed bytes straight from whatever
answered the request. NUL and newline were filtered but escape, carriage
return and backspace were not, so an error page could rewrite the output
or bury it in a CI log. Non-printable bytes are stripped and it goes
through printf. CodeRabbit suggested hex-encoding the lot; printable
characters are kept instead, because "<!DOCTYPE html>" is the diagnostic
-- hex would make the line safe and useless.

Also corrected the comment above the parse. It asserted GitHub returns
this JSON on a single line; the API is pretty-printed by default, and I
could not get a single-line response from two machines across five header
variants. The single-line case is real (it is what produces
"mentions_count", and the failing device's error named
pixlet_mentions_count_linux-arm64.tar.gz), but it is a shape to be robust
against, not a constant. As written the comment invites the next reader to
check by hand, see pretty JSON, and conclude the fix was unnecessary.

Tests drive the real script with a stubbed curl: the tag resolves from
both response shapes, non-release values fall back, an HTTP error is
reported as a download failure rather than surfacing later as a tar error,
a non-archive body is rejected before extraction, and the diagnostic
cannot carry control bytes. The stub honours -f the way real curl does --
without that, the HTTP-error test passed against the old script too, since
both end at 0/1 and only the reporting layer differs.

Mutation-checked: 10 of the 16 fail against the pre-fix script, and the 5
covering these two findings fail against this branch's previous state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:35:02 -04:00
ChuckandClaude Opus 5 0901d044d3 fix(plugins): report updates that completed, not ones that were queued (#460)
run_scheduled_updates_with_changes() snapshotted plugin_last_update,
called run_scheduled_updates(), and diffed the two to answer "whose data
just changed".

But run_scheduled_updates() only enqueues. The work runs on the update
worker and stamps plugin_last_update there, after this method has already
returned, so the two snapshots were always identical and the result was
always an empty list. The only path that ever worked was the synchronous
kill-switch, where update() runs inline.

Vegas is the caller. That empty list is what feeds mark_plugin_updated(),
which drops the cached content for a plugin whose data moved -- so a
segment kept scrolling whatever it was first built from. It is the
failure the coordinator's own comments describe: last night's live game
still drawn as live the next morning. On a live rig: zero update ticks in
twenty minutes, with weather, stocks and news all updating on schedule.

The worker now records each completed update in a ledger and the call
drains it, reporting what has finished since the previous poll rather
than what this call enqueued. That costs one tick of latency -- Vegas
polls every ~4s -- and is correct whichever side of the queue the work
lands on. Failure paths are excluded: they stamp the timestamp too, to
space out retries, but no fresh data exists.

Verified on the rig it was found on: 0 update ticks before, 208 in
twenty-five minutes after, naming real plugins.

The behavioural tests here would pass with both production call sites
deleted, which mutation testing caught -- they drive the ledger directly.
So there is also a structural test asserting the invariant at the source:
wherever a successful update stamps plugin_last_update, it must record
the completion. Writing it immediately caught that _record_update_failure
stamps the same field and must not be included.

Mutation-checked: removing either call site, removing both, and dropping
the drain's clear are all caught.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:28:40 -04:00
ChuckandClaude Opus 5 08265c1135 feat(vegas): let live content keep its place in the ticker (#457)
* feat(vegas): let live content keep its place in the ticker

Live content used to preempt Vegas outright: while any plugin reported
live priority the display controller refused to run the ticker at all and
showed a full-screen scoreboard instead. Keeping the marquee meant not
seeing live scores; seeing live scores meant losing the marquee.

Two changes, both off by default.

vegas_scroll.live_in_ticker keeps the ticker running through a live game.
Three places assumed the takeover and all three now honour it: the
controller's gate, the coordinator's per-frame pause, and the rotation
switch that would otherwise move current_mode_index underneath a ticker
that never yields.

And the rotation is no longer a strict round robin. It was one slot per
plugin per cycle, so with a dozen plugins enabled a live score came round
once a lap and could be minutes old on screen. A plugin can now hold
several slots, placed by Smooth Weighted Round-Robin -- the same
scheduler the sports plugins already use to rotate their own games. The
property that matters is that repeats are spread through the cycle
rather than clumped: three in a row and then silence would be worse than
no boost at all.

Weight comes from the plugin first, via a new optional
get_vegas_priority_weight(), then from the core: live content earns
live_weight, everything else 1. So existing plugins gain the behaviour
without changes, and the hook exists for the one thing the core cannot
work out -- the core can see that a game is live but not whose, so only
the plugin can say a favorite is playing.

Documented in ADVANCED_FEATURES (worked example, why weights are per
plugin not per game, and that frequency is not freshness),
CONFIG_REFERENCE, PLUGIN_API_REFERENCE, and the config template.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): carry the new keys through config, and correct two docs

Three findings from CodeRabbit, all valid.

to_dict() and update() enumerate keys explicitly and had not learned the
three new ones, so get_status() never reported them and a live config
change never applied -- turning live_in_ticker on in the web UI would
have done nothing until a restart. update() clamps the weights exactly
as from_config does.

The vegas_scroll key count in ADVANCED_FEATURES said 29; the template
has 30. My arithmetic, not the reviewer's.

The third was a documentation error rather than a code one, and I have
fixed it the other way round. The docs claimed a raising
get_vegas_priority_weight() is treated as weight 1. The code instead
falls through to the core's own live-content check, and that is the
better behaviour: the hook is only how a plugin asks for *more* than
live_weight, and has_live_priority/has_live_content are separate methods
guarded separately, so a plugin with a broken weight calculation should
lose the favorite distinction and keep the live boost. Said so in the
code, the base-plugin docstring and the API reference.

The test fake now fails in each place independently, because the two
failures mean different things: a broken hook still earns live_weight, a
plugin that cannot say whether it is live has nothing to fall back on
and weighs 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ui/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): stop the heaviest plugin doubling across the cycle seam

Smooth Weighted Round-Robin spaces repeats well within a pass, but it
schedules the heaviest item first and usually last as well. The strip
loops, so those two are neighbours: the marquee showed the same plugin
twice running at exactly the one join a within-cycle check cannot see.
Observed on a live rig at 28 slots -- gaps of 6, 7, 7, 7 and then 1.

Rotating the list does not fix it. Rotation preserves the cyclic order
exactly, so it moves where the seam is drawn rather than the adjacency
itself; the trailing entry has to be swapped with one from the middle.

The first version swapped with the first slot that merely fitted, which
undid the spacing this exists to protect -- it moved a repeat from a gap
of 7 into a gap of 2, more clumped than the seam had ever been. It now
picks the candidate furthest from any other appearance, so the repeat
lands in the widest gap.

Left alone when no candidate exists. A plugin holding most of the slots
has to neighbour itself, and scheduling it is better than refusing to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): stop the seam repair creating the duplicate it removes

Swapping the trailing repeat with a middle slot moves two elements, and
the candidate filter only guarded one of them. It checked the neighbours
`repeated` would acquire at j, but not what the displaced element would
sit beside at the end -- so ['a','b','c','d','x','y','x','a'] came back
as [...,'x','x'], the seam duplicate traded for a fresh one. Reported by
CodeRabbit with that exact case.

Adding the missing condition fixed it and immediately broke something
else: schedule[j] is schedule[-2] when j is the second-to-last slot, so
that candidate was always excluded, and ['a','b','c','a'] lost the only
repair it has. The same class of mistake twice, from reasoning about
which neighbours two moved elements end up with.

So it no longer reasons. It performs each candidate swap, counts the
cyclic duplicates in the result, and keeps the best one that has none --
preferring whichever leaves the boosted plugin most evenly spread. When
no such swap exists the schedule is returned untouched, which is the
unavoidable case: a plugin holding most of the slots has to neighbour
itself.

Fuzzed across 6,956 seam schedules: none made worse, none lost an entry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:02:36 -04:00
ChuckandClaude Opus 5 5713fd20a7 fix(calendar): implement the OAuth and calendar-listing endpoints (#458)
* fix(calendar): implement the OAuth and calendar-listing endpoints

The plugin's config advertised a three-step setup and only step 1
existed. Step 3's picker fetched
/api/v3/plugins/calendar/list-calendars, which was never registered, so
Flask fell through to the global 404 handler and the user saw "Resource
not found" -- a message that names nothing and points nowhere. Step 2
had no endpoint at all, so even a working picker would have found no
token to list with.

Two routes, following the pattern the spotify and ytm plugins already
use for their own auth scripts:

  POST /plugins/calendar/authenticate    two-step Google OAuth
  GET  /plugins/calendar/list-calendars  calendars for the picker

The authenticate route drives calendar_registration.py, which the plugin
already ships and which was written expressly for this -- it reads a
redirect URL on stdin and prints one JSON object. It takes two calls
because a human has to visit Google in between; the script persists the
PKCE verifier from the first call for the second, without which the
exchange fails with "Missing code verifier".

The listing route reads the token directly rather than shelling out
again: the picker is interactive and a subprocess per click is slower
than the API call it would wrap. It refreshes an expired token in place,
sorts the primary calendar first, and drops entries with no id, which
could not be selected anyway.

Both name the plugin when it is not installed, rather than reproducing
the anonymous 404 that started this.

Verified against the live Google API on the dev rig: HTTP 200 with the
account's real calendars.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(calendar): one input for the auth code, and a louder warning

Two things reported after testing the flow.

There were two boxes and no way to tell which to use. The config
template's string branch dispatches widgets from an allow-list of names,
and anything missing from it falls through to a plain input type=text --
so the field rendered both the widget's own box and a stray one for the
same key. google-oauth is now on that list, which is all the widget ever
needed to render in place of the fallback rather than beside it.

And the warning that the redirect page fails to load was small grey text
under a link, which is where it is least likely to be read. It is now an
amber callout that leads with "The next page will fail to load. That is
expected." The failure lands at exactly the moment the user has to act
on it, and it looks precisely like the flow breaking rather than
working. The paste box is labelled too, rather than relying on a
placeholder that vanishes on focus.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(calendar): redact script diagnostics, and page the calendar list

Three findings from CodeRabbit, all valid.

Raw subprocess output was being returned to the client -- the script's
stderr on one path, and its own error payload on another. CodeQL flagged
the same line. That script handles OAuth client secrets and interpolates
exceptions into its messages, so either could carry a secret or a path.
Both now go to the log unredacted, where they are worth having in full,
and reach the client through a redactor.

That redactor already existed inside describe_exception, which only
takes exceptions. Split out as redact_text: an exception is not the only
thing worth returning, and a subprocess's stderr is just as capable of
quoting a token.

calendarList.list returns 100 entries per page by default, caps at 250,
and hands back a nextPageToken when there are more. Reading one page
would have hidden calendars from the picker with nothing to say the list
was cut short. It now pages, asking for 250 at a time, bounded at ten
pages so a malformed token cannot spin.

And a test helper was a lambda where ruff wants a def.

The five new tests fail against the previous commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(calendar): redact the last two raw exception interpolations

CodeQL flagged four exposure paths. Two were mine and genuinely raw: the
OSError from failing to spawn calendar_registration.py, which carries the
interpreter path and whatever the OS chose to say, and the ImportError
for the Google libraries, whose message named the missing module by
interpolating the exception directly. Both now go through
describe_exception, and the unredacted text goes to the log.

The other two are the repo-wide pattern from PR #448 -- 67 handlers on
main already return details=describe_exception(e), and these two new
handlers follow it. That function is the sanitizer: it strips URL
userinfo, auth headers and credential-shaped key=value pairs, collapses
to one line and caps the length. CodeQL's taint tracking cannot see a
sanitizer it has no model for, so it reports the flow regardless.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(calendar): announce status changes, name the paste box, drop a no-op

Three more from the review, all valid.

The status line is written after every async call -- the consent link is
ready, the exchange failed -- and was a plain paragraph, so a screen
reader was told none of it. It is a live region now.

The paste box had a visible label that was never associated with it, so
its only accessible name was the placeholder, which disappears on focus:
precisely when the value is being pasted. The label now points at the
input by id.

And a conditional in the test helper returned the same value from both
branches, which Ruff flags as RUF034. It was left over from making the
fake page; one page is all those cases need, and TestPagination builds
its own sequences.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* test(calendar): assert the accessibility relationships, not their parts

The previous assertions searched for role="status", aria-live, a label
`for` and an input `id` independently, so they passed whether or not
those belonged together. Two attributes on different elements announce
nothing, and a `for` that names something other than the input leaves it
just as anonymous.

Both attributes are now asserted on the status element itself, and the
label and input are checked to go through the same identifier rather
than merely both existing. Verified by mutation: a mismatched pair and a
displaced aria-live are both caught.

Reported by CodeRabbit, against tests I had written two commits earlier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 15:25:43 -04:00
ChuckandClaude Opus 5 9cf30bbbef fix(startup): bound the initial plugin update so the panel lights sooner (#456)
* fix(startup): bound the initial plugin update so the panel lights sooner

DisplayController.__init__ calls _update_modules() once to populate
plugin data before the first frame. It walks every loaded plugin in
turn, and each update blocks the calling thread for up to the executor's
30s timeout, so the uncapped total is the sum of every slow plugin on
the system. The rig's own log:

    Initial plugin update completed in 82.255 seconds
    Initial plugin update completed in 55.123 seconds
    Initial plugin update completed in 25.975 seconds

The panel shows nothing for all of it.

Nothing is lost by stopping early. A plugin that has never updated is
immediately due, so run_scheduled_updates() collects it seconds later --
with the display already running rather than blank.

A deadline alone was not enough: it is checked before each plugin, so
the last one to start could still block for the full 30s, and a 20s
budget produced a 31.8s pass on the rig. The remaining budget is now
passed down as that update's timeout too, with a floor so a plugin
starting on the last sliver is not handed ~0s and recorded as having
timed out for a slot it never had. Measured after: 20.006s.

Found while profiling a scroll freeze with py-spy, which caught the main
thread 9.34s inside execute_with_timeout's join. Worth being clear that
this is startup latency, not the recurring stutter -- _update_modules
has exactly one caller and runtime updates already run off the display
thread.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* feat(display): show the device address on the startup screen

That screen is what the panel holds for the whole initial plugin update,
and on a headless Pi it is the only place the address appears without
going looking for it -- so it now carries the address under
"Initializing".

The lookup connects a UDP socket, which sends no packets: it only asks
the kernel which source address it would route from. That costs 0.03ms
and works with the network down so long as a route exists. Deliberately
not `hostname -I` plus a systemctl probe for AP mode, which is how the
web launcher does it -- two subprocesses with multi-second timeouts, on
the startup path this branch exists to shorten.

Two things had to change for the address to be worth putting there.

The text is now sized to fit rather than fixed at 8px: "Initializing" is
96px in PressStart2P, drawn at x=10, so it already ran off the side of a
64px panel before an address was added. It falls back to 4x6 where that
does not fit, and both lines are centred.

And the test pattern is punched out from behind the block, with the text
drawn white rather than blue. The diagonal runs through the middle of
the panel, which is exactly where this sits, and blue on black reads
fine on a monitor but is marginal on a dim panel. An address that cannot
be read off the wall is not worth showing.

The rendering tests assert against pixels -- no green left behind the
text at any supported size, enough lit pixels to be visible -- rather
than against the geometry that produced them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(display): keep the startup text blue -- it is a channel reference

The test pattern lights one pure channel per element: red border, green
diagonal, blue text. That is how a glance at the panel tells you whether
led_rgb_sequence is right -- wire it BGR and the border comes up blue
and the text red. Drawing the text white, as the previous commit did for
contrast, lights all three channels and destroys the only blue reference
on the screen.

Reverted to blue, with the reason written down so it is not treated as a
style preference again, and with tests that pin it: the text must be
pure blue, nothing on the screen may be white, and all three primaries
must be present.

The punched-out backdrop stays. It only removes the diagonal from behind
the glyphs, which costs nothing diagnostically -- the diagonal is still
plainly visible across the rest of the panel -- and it is what makes the
address readable at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(startup): defer a plugin with too little budget, rather than clamp it

The per-plugin timeout was clamped up to a floor, so a plugin that began
with a sliver of budget left was granted the full floor and ran on past
the deadline: a 20s budget could take 22. The floor existed to stop a
plugin being handed a slot too short to use and then recorded as having
timed out, which is a real concern, but clamping solved it by breaking
the bound.

Deferring solves both. Below the floor the plugin is left to the update
tick, which was already the fate of everything after the deadline, so
nothing new is lost -- a plugin that has never updated is immediately
due. Above it, the timeout is the exact remainder, and the pass cannot
outlast its deadline.

Measured on the rig after the change: 20.002s, 5 plugins deferred.

Also names an unused binding in the initializing-screen test.

Both reported by CodeRabbit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 14:58:05 -04:00
ChuckandClaude Opus 5 a51fb7ce11 feat(vegas): make scroll stutter visible, and catch it in the act (#454)
* feat(vegas): make scroll stutter visible, and catch it in the act

The loop reported only a mean FPS over a five-second window. At 120fps
that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee --
moves the average from 120.0 to 115.4 and reads as healthy. Stutter was
literally unmeasurable.

The FPS line now carries p99, the worst frame, and a hitch count. On the
dev rig that immediately turned "it sometimes stutters" into a number:
two freezes of 3.2s and 0.7s in twenty minutes, with every other frame
under 81ms.

Statistics say a stall happened but not what caused it, and by the time
they are logged the stack is gone. So there is also a watchdog that dumps
every thread's stack while the loop is still wedged. It is off unless
LEDMATRIX_STALL_WATCHDOG is set to a threshold in seconds, since it
prints a lot. Pointed at the 3.2s freeze it named the culprit on the
first try: a plugin generating a 17,000px scroll image, logo PNG decode
and all, synchronously on the render thread.

The hitch threshold is relative to what frames actually cost, not to the
configured target. The target is routinely set above what the panel can
hold so vsync does the pacing; measured against that budget every
ordinary frame counts as a hitch, and the first version of this counter
duly reported 250 per window on a display running perfectly smoothly.

The watchdog is owned by the coordinator, not created per iteration --
run_iteration is called repeatedly, so building one there would leak a
thread each time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): let the watchdog see stalls that hold the GIL

The watchdog only noticed a late heartbeat, which a whole class of
freeze can never produce: if the loop is inside one long C call that
holds the GIL, this thread cannot run during the stall, and by the time
it does the loop has already checked in. On the dev rig that hid a
recurring 3.2s freeze completely -- twenty minutes of watching produced
one dump, for an unrelated 0.4s stall.

What it can still observe is that its own sleep ran long. A badly
overshot wait is now reported as a stall in its own right. The stacks
are stale by then and the message says so, but knowing the freeze is
GIL-holding is most of the diagnosis: it rules out lock contention and
scheduling, and points at a single long C call.

This also explains why lowering sys.setswitchinterval changed nothing --
the switch interval cannot preempt a C call that never releases the GIL.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* feat(vegas): report the worst frame, not just the mean

The loop logged only a mean FPS over a five-second window. At 120fps
that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee
-- moves the average from 120.0 to 115.4 and reads as perfectly healthy.
Stutter was unmeasurable, which is why "it sometimes freezes" went
unpinned for so long.

Adding p99 and the worst frame turned that into a number immediately: on
the dev rig, two freezes of 3.2s and 0.7s in twenty minutes with every
other frame under 81ms. Not general slowness -- two rare, total stalls,
which is a different problem with a different fix.

Costs 0.96us per frame, about 0.012% of an 8.3ms frame.

This replaces an earlier version that also shipped a stall watchdog and
a hitch counter. The watchdog never found anything -- one dump in
forty-five minutes, for an unrelated stall -- because it can only notice
a late heartbeat, and the freeze happens in coordinator.start() before
the frame loop begins beating. py-spy found the cause in one recording
by sampling the process externally, which needs no code here. The hitch
counter went with it: it needed a rolling median every frame, which was
most of the cost, to produce a number the worst frame already tells you.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): use the nearest-rank index for p99

int(n * 0.99) is off by one, and at exactly 100 samples it selects the
maximum -- which is the number logged immediately beside it as the worst
frame. The two columns exist to say different things, p99 the
bad-but-ordinary frame and worst the outlier, so they agreed precisely
when the sample was smallest and least informative.

Nearest rank is ceil(n * fraction) - 1. Extracted so it can be tested
directly rather than only through a five-second logging interval.

Reported by CodeRabbit on the PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 14:06:32 -04:00
ChuckandClaude fce1fdac57 Add panel orientation setting for upside-down mounting (#455)
Adds a display.hardware.orientation config field ("normal" / "180")
so panels mounted upside down (e.g. to put the Pi/wiring on a more
convenient side) render correctly without custom pixel_mapper_config
edits. Composes onto the existing pixel_mapper_config as a trailing
"Rotate:180" mapper, so it stays independent of any custom mapper
string (e.g. U-mapper chain layouts) already in use.

Exposed as a "Panel Orientation" dropdown in the web UI's Display
settings, validated server-side, and documented in README and
CONFIG_REFERENCE.


Claude-Session: https://claude.ai/code/session_01FakipqMDHQLpsFjTuBdSFQ

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-12 09:40:48 -04:00
ChuckandClaude Opus 5 7171e6c022 fix(cache): one cleanup thread per cache directory, not per manager (#453)
The display process ran three cleanup threads over one directory:

    14:22:59.954  display_controller        (the real manager)
    14:22:59.973  startup validation, run 1 (discarded)
    14:23:01.055  startup validation, run 2 (discarded)

Two of those managers existed only to read a directory path.
StartupValidator._validate_cache_directory built a whole CacheManager to
call get_cache_dir(), and validation runs twice -- once before the
plugin manager exists and again after. Each construction also probes
writability by writing and deleting .writetest on the card.

The discarded ones never went away. cleanup_loop closes over `self`, so
the thread keeps its manager alive: two objects that could never be
collected, waking every 24 hours to re-scan the same 9,000-file
directory. Nothing stopped them either -- stop_cleanup_thread had no
callers anywhere in the tree.

Two changes. The validator now takes the CacheManager the application
actually uses, which is also the more correct thing to validate; when
no caller supplies one it still builds its own, but stops the thread
afterwards. And CacheManager now tracks which directory it is sweeping,
so the second manager over a directory skips starting a thread at all.
That is the right granularity regardless of call sites: the sweep lists
a directory and deletes from it, so a second thread only duplicates the
scan. Ownership is released on stop, so a survivor can take over rather
than leaving the directory permanently unclaimed by a dead owner.

Measured directly, three managers over one directory: 3 threads before,
1 after.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 08:42:56 -04:00
ChuckandClaude Opus 5 9fbdd71941 fix(cache): collect the temp files abandoned writes leave behind (#452)
* fix(cache): collect the temp files abandoned writes leave behind

DiskCache.set() writes through mkstemp then os.replace, and removes its
own temp file in a finally. That covers a write that fails, but not a
process that dies between the two -- a SIGKILL, a lost restart race, a
power cut, all ordinary on a Pi.

Nothing ever collected what was left. The temp names are
".<key>.json.<random>", and cleanup_expired_files listed only names
ending in .json, so every one of them was invisible to the sweep for as
long as the card had been in service. On the dev rig: 76 files,
1,050 MB, 81% of the whole cache directory, the oldest six months old.
The startup sweep reported "18/8864 files deleted, 0.01 MB freed" while
sitting on top of a gigabyte it could not see.

They are removed after an hour. A real write holds its temp file for
milliseconds, so that is far outside any in-flight write while still
clearing the same day's debris, and it is deliberately not tied to the
retention policies: those say how long data stays useful, and a
half-written file never was.

The predicate is tested harder than the sweep, because a false positive
deletes real data. It matches the shape set() creates rather than just a
leading dot, so a completed ".json", a stray .gitignore, and a
"weather.json.bak" are all left alone -- and one test drives set()
itself and asserts the names it produces are matched, so the writer and
the predicate cannot drift apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(cache): count swept temp files as scanned

files_scanned only counted completed .json files, so a sweep that
removed orphans reported more deleted than it had looked at -- the
summary line renders "<deleted>/<scanned>", which came out as "76/1".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 08:40:07 -04:00
ChuckandClaude Opus 5 2add759f40 fix(odds): identify the odds requests to ESPN (#451)
The odds fetch used a bare requests.get, so it went out as
python-requests/x.y -- the one agent ESPN is known to reject. Around
2026-08-04 it began 403ing browser strings and bare custom tokens alike;
what it accepts is a token carrying a URL that says who is calling.
Every other ESPN caller in the tree already sends that header
(src/common/api_helper.py, src/base_classes/data_sources.py); this path
was simply missed.

It is the worst one to miss. Odds are fetched per live game from inside
the live update loop, so its failures are the ones that cost the caller
its whole update budget -- the same path the 5s timeout and the cooldown
were added to protect.

Sent via a session rather than per-call, which also reuses the
connection across a slate. Deliberately no retry adapter, unlike
api_helper: retries multiply request_timeout, which is 5s precisely to
stay inside the 30s operation budget.

The existing tests patched the module's requests.get, which this change
bypasses -- test_base_odds_manager was consequently reaching the real
ESPN and taking 404s. Both files now patch the session, and the new
tests pin the agent against api_helper's live value so the two cannot
drift apart the next time ESPN moves the goalposts.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 16:10:56 -04:00
ChuckandClaude Opus 5 bb1a1671ec fix(cache): make the ttl parameter actually control expiry (#450)
CacheManager.set(key, data, ttl=...) stored the number and no read path
ever consulted it. Expiry came from a max_age inferred from substrings
in the key -- "live", "odds", "stock" -- so all 52 callers passing a ttl
were writing a value that did nothing. The docstring said so outright:
"stored for compatibility but expiration is still controlled via max_age
when reading". It is easier to read that as a note than as a defect,
which is presumably how it survived.

Both cache layers already hold the record when they decide, so each now
prefers an explicit ttl and falls back to max_age when there is none.
The caller that wrote the record knows what its data is; a substring
guess is a reasonable default for records that never said, and a poor
override for records that did.

Measured against a device's real cache of 8,875 entries carrying a ttl,
the inferred and intended values disagreed nearly everywhere:

    stocks    max_age  600  vs ttl    1800   4903 entries
    news      max_age 3600  vs ttl     600   1770 entries
    odds      max_age 1800  vs ttl    3600   1301 entries
    images    max_age  300  vs ttl 2592000     20 entries

In every case the ttl matches what the plugin plainly intended: stock
quotes cached for half an hour rather than ten minutes, headlines
refreshed every ten minutes rather than hourly, bird photographs that
never change kept for a month rather than five minutes.

Two things make this safe to land now. No sports_live entry carries a
ttl at all -- the live-score path does not use set(ttl=) -- so live
freshness is untouched, which matters with a season two weeks out. And
replaying the change against that real cache, 997 currently-expired
entries become live while not one live entry becomes expired, so there
is no invalidation spike on deploy.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 14:14:25 -04:00
ChuckandClaude Opus 5 8159afca43 fix(odds): stop a stalled ESPN taking the whole plugin update with it (#449)
Odds are fetched per live game from inside SportsLive.update(), with
show_odds defaulting on, and the plugin executor kills an operation at
30s. The odds request timeout was also 30s, so a single stalled request
consumed the entire budget and the update carrying every game's score
was killed.

Out of season that is invisible: preseason week 1 returns one game. A
Sunday slate is around sixteen, so the odds of at least one slow request
rise sharply just as the cost of losing the update does.

Shorten the request timeout to 5s, and after a network failure skip the
network for 60s. The timeout alone is not enough -- sixteen consecutive
5s timeouts still blow through -- and when ESPN is unreachable it is
unreachable for the whole slate, so the first failure already answers
the question for the rest of the pass.

    before: one stalled request = 30s = the entire budget
    after : 5s, the rest of the slate skipped, retry after 60s

The stale-cache fallback is unchanged: the cache is consulted before any
of this, and the failing request still falls back to it.

An earlier version of this branch also jittered the cache TTL to stagger
expiry across a slate. That has been dropped: CacheManager.set() stores
ttl for compatibility but the read path expires entries by a per-type
max_age (1800s for odds), so the jitter was inert. Making the read path
honour a per-entry ttl is a real fix but changes a contract 48 plugin
call sites already rely on, which is not a change to make two weeks
before the season.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 13:56:53 -04:00
ChuckandClaude Opus 5 44f59ede07 fix(web): say what actually went wrong instead of "unknown" (#448)
* fix(web): say what actually went wrong instead of "unknown"

Every failing endpoint returned "An error occurred; see logs for
details" and nothing else. That is survivable until the logs are the
thing you cannot reach: a device whose SD card was failing answered the
restart action, /system/status and /logs with that same sentence -- the
log viewer included, because journalctl could not be executed -- while
the exception underneath said

    [Errno 5] Input/output error: 'systemctl'

which names the fault outright. The only endpoint that helped was
/health, and only because it happens to pass a subprocess's stderr
through. Diagnosis came down to guessing which endpoint leaked something.

Add describe_exception(), returning "TypeName: message" on one line, and
populate the `details` field that the response schema has always had and
nothing ever filled. The type alone carries information -- a bare
PermissionError says more than any generic sentence.

Exception text is not automatically safe to echo: a requests error
quotes the URL it failed on, and plugins that authenticate by query
string put their key there. Credential values are redacted while the
parameter name is kept, since knowing which credential was involved is
part of the diagnosis. Length is capped and newlines collapsed so a
parser's context cannot flood a JSON field.

Nine handlers in api_v3 bound the exception and never used it, so the
promised log entry was never written either -- "see logs for details"
was false, not merely unhelpful. Those now log with a traceback and
carry the detail. The other 60 already logged and are unchanged; they
can adopt the helper as they are touched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(web): redact auth headers and URL userinfo, and cover every handler

Three review findings.

The sanitizer missed two credential shapes that requests puts in its
exception text verbatim: `Authorization: Bearer <token>` and
`https://user:password@host`. Both would have gone straight into a
response. The auth-scheme name and the username are kept -- they say
which credential and whose without being the secret.

The AST test only asked whether *something* had been logged, so a
`logger.info("failed")` satisfied it while discarding the exception just
as completely. It now requires an error-level record carrying exc_info
and `describe_exception()` called on the handler's own bound exception.

Enforcing that revealed the first cut had scoped itself wrongly. I had
converted the nine handlers that logged nothing and left the sixty that
logged, reasoning their detail was at least in the journal. But
/system/status is one of the sixty, and on the failing device it told me
nothing -- the journal was exactly what could not be read. Splitting
them left most of the diagnostic surface unhelpful for the case this
change exists for, so all sixty-nine now carry the detail.

Two handlers had no bound exception name, and three passed the message
through a variable rather than a literal; both shapes needed doing by
hand. Full suite: 2383 passed, one pre-existing unrelated failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(web): stop reporting client errors as server faults

Werkzeug's HTTPExceptions subclass Exception, so the catch-all handler
saw them too and turned every 405, 400, 413 and 415 into a 500
UNKNOWN_ERROR. A GET on a POST-only route answered "an error occurred;
see logs for details", which tells the caller nothing and blames the
wrong side -- found while probing a device whose POST-only config
endpoints did exactly that.

Hand HTTPExceptions back as themselves, with their own status and
description. A genuine server fault still reports as one, with the
detail this branch adds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(web): redact any auth scheme, and require the detail in the response

Two review findings.

The auth-header pattern listed Bearer, Basic, Digest and Token, so
`Authorization: ApiKey SECRET` or `Negotiate SECRET` went to the client
intact. A fixed list silently leaks whatever it does not name, and
plugin APIs invent their own schemes, so match any scheme name and keep
it while redacting the credential.

The AST test accepted a describe_exception(e) call anywhere in the
handler, which a handler could satisfy by computing the detail and
dropping it before returning the generic message. It now requires the
call inside every return expression, which is where it has to be to
reach the caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 13:56:16 -04:00
ChuckandClaude Opus 5 ca26c1b83b fix(vegas): stop the width cap emitting fragments and stale windows (#446)
* fix(vegas): stop the width cap emitting fragments and stale windows

Two defects in the rotation that narrows an oversized plugin to its
width budget. Both were found while investigating "cut off early /
starts in the middle" reports and are the reason the cap is no longer
on by default; they still bite anyone who sets one.

A rotation's last window was whatever happened to be left over. Windows
are placed by walking forward from the previous one, with nothing
looking at the remainder, so a 1,840px stocks ticker against a 1,536px
budget split 1,492 + 348 -- every other appearance showed seven seconds
and cut. Absorb a remainder below half a budget into the window before
it. That overruns the budget by at most half, which is the better trade:
the budget guards against one plugin holding the panel for minutes, not
against a 20% overshoot. The floor is measured against the budget rather
than the panel because snapping to item boundaries already lands an
ordinary window short of it -- a 512px budget over 182px-pitch items
yields 348px windows, so an absolute floor merges windows that were
never fragments.

The stored offset also outlived the content it was recorded against. It
was a pixel column, reused verbatim after the plugin re-rendered, so
once anything ahead of it changed width the window pointed at unrelated
items -- observed as news refreshing 9,793px -> 9,505px mid-rotation.
Track the rotation as an index into the strip's item boundaries instead,
since the Nth boundary survives a digit appearing in a price, and record
alongside it what the offset indexes into: a row list, a boundary list,
or a column in a gapless image. A mismatch restarts the rotation rather
than reinterpreting the number, which also closes the case where one
plugin's row index was read back as a pixel column after its content
changed from several rows to one wide strip.

Replaying the four plugins that actually hit the cap on a live 512px
panel: no window is now a fragment, none exceeds 1.5 budgets, and every
rotation still covers the whole strip.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): apply the runt floor to the multi-row rotation too

The floor only guarded the single-image path. I had reasoned the row
path could not produce a runt because it wraps, which is wrong: wrapping
only helps when the row wrapped to actually fits. Rows of 450, 450 and
100 against a 512px budget give the 100 a pass of its own -- two seconds
against nine, which is the symptom this branch exists to remove.

Reproduced before changing anything:

    pass 1: 450px    pass 2: 450px    pass 3: 100px

A window may now overrun the budget while it is still shorter than the
floor, bounded at the same 1.5 budgets the single-image path allows, so
the short row is carried with its neighbour instead of standing alone.

    pass 1: 450px    pass 2: 450px    pass 3: 550px

A next row too wide to absorb within that cap still leaves a short
window standing -- rows of 900 and 100 keep alternating. Merging them
would mean a window of nearly two budgets, and the rule that always
shows an oversized first row already makes the same trade.

Three regression tests: the reported shape, that the overrun stays
bounded when a row cannot be absorbed, and that absorbing never drops a
row from the rotation. The single-image path is untouched -- the four
plugins that actually hit the cap on a live panel replay identically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 09:41:25 -04:00
ChuckandClaude Opus 5 f887063434 test(harness): flag a mode that draws nothing without reporting it (#447)
The controller skips a mode whose display() returns False and treats
anything else -- including None -- as "content was shown". A mode that
draws nothing and does not return False is therefore never skipped, and
because a mode switch clears the panel first, it sits on a blank screen
for its whole display duration. Two sports plugins shipped exactly that.

The harness rendered those modes and passed them, because it called
display() and discarded the result. Capture it, and warn when a render
produced no lit pixels while claiming content.

Warn-only by default, and deliberately so: a scroll mode's first frame
is legitimately its blank scroll-in buffer, which is 42 of these on the
F1 scoreboard alone. Plugins whose modes are known to draw on their
fixture data can opt into failing via harness.json {"empty_check":
"strict"}, matching how the fill check is staged.

Worth being clear about the limit: this only sees what the fixtures
render. It would not have caught the sports bug, whose fixture seeds
games so the empty path never renders -- that needs the source-level
gate in the plugins repo. What it does catch is the same mistake in any
plugin whose empty state the harness does happen to reach, which is
coverage there was none of before.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 08:17:57 -04:00
ChuckandClaude Opus 5 6287acd591 fix(vegas): stop capping plugin width by default (#445)
* fix(vegas): stop capping plugin width by default

Vegas plugins read as "cut off early" or "starting in the middle". That
was the per-plugin width budget, not the scroll engine:
overflow_mode=rotate is designed to resume mid-content on each
appearance, so the symptom was the feature working as specified.

Measured over a 17-plugin fleet on a 512px panel, the cap was a bad
trade. Only four plugins were ever wide enough to hit the 3.0 default --
leaderboard 11,518px, news 10,021px, odds-ticker 4,643px, hockey
1,508px. Weather is 650px and flights 512px; the cap never touched them
or the other eleven. So it bought nothing on thirteen plugins while
costing two visible faults on four: content entering mid-item (a news
ticker started at column 6027 of its own strip), and a final rotation
window of whatever happened to be left -- 348px of an 1,840px stocks
ticker, seven seconds of panel time.

Default max_plugin_width_ratio to 0 (uncapped), so every plugin
contributes all of its content and is always entered at its beginning.
The cap remains available, and vegas_max_width_screens still caps an
individual plugin -- which is where the knob belongs, since a genuinely
long ticker is a property of that plugin rather than of the fleet.

Verified on a live 512px device: 327 budget crops in the preceding six
hours, none after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* test(vegas): cover the width cap still working when asked for

Defaulting the cap off must not quietly remove it. Asserts that an
explicit ratio is honoured and validates, alongside the existing check
that omitting it means uncapped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:17:28 -04:00
ChuckandClaude Fable 5 ee59caa577 Follow-ups from #441: secret-helper migration, ten more bug fixes, and coverage for every remaining untested module (#444)
* refactor(web): use canonical secret helpers in api_v3; make ConfigManager secret strip/merge array-aware

api_v3.py carried three inline nested copies of find_secret_fields/
separate_secrets (main-config save, plugin-config save, plugin-config
reset). They drifted from each other (one lacked isinstance guards) and
none supported the canonical module's array-item secrets
(accounts[].token). All three endpoints now import from
src/web_interface/secret_helpers.

Adopting the canonical behavior makes array-item secrets reachable, and
their parallel-placeholder shape ([{'token': ...}, {}] alongside the
regular list) was not survivable by ConfigManager's round-trip:
_strip_secrets_recursive dropped the whole key (losing the regular
fields from config.json) and _deep_merge replaced the regular list
wholesale on load. Both are now array-aware:

- strip removes the secret fields from each item and ALWAYS keeps the
  list so indices survive for merge-on-load; whole-key secrets (scalar
  lists, shape mismatches) still drop the key entirely — never leak.
- merge folds each secrets item into the config item at the same index,
  skipping {} placeholders. The regular list's length is authoritative
  in both directions: a user deleting an array item never has it
  resurrected from a stale secrets entry (extras warn and are ignored).

api_v3's own deep_merge intentionally still replaces lists wholesale —
form posts carry complete arrays and index-merging would resurrect
deleted items; a comment now documents that.

Tests: the parity guard flips from 'exactly 3 inline copies' to 'zero,
and the canonical import must exist'; TestArraySecretStripAndMerge
covers the new strip/merge semantics incl. length-mismatch contracts;
new test_api_v3_secret_roundtrip.py drives all three endpoints through
a Flask client with a REAL ConfigManager+SchemaManager over tmp_path,
proving secrets land in config_secrets.json, config.json stays clean,
and a fresh load merges them back into the right array items.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: repair broken helper paths across display, cache, odds, logging, resolver, repos, config, validator

Nine fixes for bugs surfaced while writing coverage for previously
untested modules (plus the bool-duration quirk pinned in PR #441):

- base_plugin.get_display_duration: exclude bools from both numeric
  branches — display_duration=True no longer reads as a 1-second slot;
  it falls through to config, then the 15.0 default.
- display_helper: draw_error_message/draw_no_data_message called
  _draw_centered_text with the wrong arguments and crashed with
  AttributeError — both now delegate to draw_centered_text.
  draw_scorebug_layout drew status and clock at the same y, overprinting
  each other — they now share one combined top line.
  draw_ticker_layout drew its text starting at x=display_width (fully
  off-canvas), returning a blank frame every time — now draws at x=0;
  scroll_speed stays accepted-but-unused and is documented as such.
- api_helper.clear_cache guarded on a nonexistent CacheManager.clear()
  method, silently never clearing anything; it now uses the real surface
  (clear_cache/delete/list_cache_files) and no-ops safely otherwise.
- base_odds_manager._extract_espn_data raised AttributeError when ESPN
  sent explicit JSON nulls ("homeTeamOdds": null) — every level now
  null-safes with 'or {}'. format_odds_summary gated on
  is_odds_available, which deliberately ignores money lines, so
  ML-only odds formatted as "No odds available" — it now gates only on
  empty/no_odds data and formats money lines.
- logging_config.ContextualFormatter mutated record.msg in place, so a
  second handler prepended the context prefix twice; it now formats a
  copy. log_error hardcoded exc_info=True and raised TypeError when the
  caller passed exc_info — now kwargs.setdefault.
- dynamic_team_resolver wrote its "shared" class cache through self,
  creating instance shadows — the cache was per-instance and every
  scoreboard refetched rankings. Writes now go through the class.
- saved_repositories cleaned URLs with an unanchored .replace('.git','')
  that mangled URLs merely containing '.git' (my.github.io -> myhub.io);
  now strips only a trailing suffix. add/remove also roll back the
  in-memory list when the save fails, so memory always matches disk.
- config_helper.merge_configs shallow-copied the base, aliasing every
  un-overridden nested dict into the result — now deep-copies.
- startup_validator.validate_all accumulated errors/warnings across
  calls — now resets both lists per run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: cover the previously untested modules

Nine new suites plus an extension, asserting the Phase-1b fixed behavior
and pinning the quirks deliberately left alone:

- test_logging_config.py: formatters (JSON shape, no record mutation,
  single prefix through two handlers), PluginLoggerAdapter precedence,
  setup_logging handler hygiene and LEDMATRIX_DEBUG, log_error exc_info.
- test_startup_validator.py: exact messages, error-vs-warning split,
  accessor split (load_config vs get_config), cache-dir branches with
  os.access monkeypatched (root can write anything in CI), idempotence,
  raise_on_errors classification precedence.
- test_config_helper.py (full): load/save round trips, dot-notation
  get/set incl. silent-failure contract, post-fix no-aliasing merge,
  schema validation branches, the '{id}_config' key pin, default-enabled
  pin.
- test_saved_repositories.py: three load shapes, bare-list rewrite pin,
  trailing-only .git strip (my.github.io regression), save-failure
  rollback, type-classification case-sensitivity pin.
- test_api_helper.py: rate-limit math, cache-hit short circuit, ESPN
  URL/key formats, exact User-Agent guard, retry adapter, post-fix
  clear_cache against the real CacheManager surface, ttl-dropped pin.
- test_base_odds_manager.py: cache-key/URL construction, no_odds
  sentinel round trip, stale-cache fallback, null-safe extraction,
  ML-only formatting, is_odds_available truth table (ML-blind by
  contract), config key/attr mismatch pin.
- test_dynamic_team_resolver.py: expansion/dedup/slicing, dropped
  unknown-dynamic names (TOP_ substring hazard pinned), genuinely
  shared class cache (second instance: zero HTTP), TTL expiry,
  failure degradation without raising.
- test_display_helper.py (full): the fixed error/no-data renders,
  combined scorebug top line, non-blank ticker with scroll_speed
  no-op pin, composite upconversion, logo bleed positions, square
  orientation pin.
- test_skin_runtime_cache.py: discovery-cache hit/invalidation
  semantics (manifest mtime, .py edits pinned as non-invalidating),
  sys.modules namespacing contract incl. bare-name restore and stdlib
  shadowing, entry-module execute-once, API minor-version tolerance,
  skin_matches_target table.
- test_sports_capabilities.py (extended): _draw_celebration_layout
  executed for real (flash window, matrix-dims fallback, highlight
  alternation, logo-failure isolation), _should_celebrate_for direct,
  strict duration boundary, score_to_int edges, both-teams-score
  precedence, expired-coalesce refire, disabled-win baseline
  preservation, id-less prune.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: real schedule/dim coverage for DisplayController; fix two vacuous schedule tests

New test_display_controller_schedule.py drives _check_schedule and
_check_dim_schedule on a bare controller stub: same-day and
midnight-crossing windows with inclusive boundaries, global vs per-day vs
legacy-inferred modes (and dim's global-only default — no legacy
inference), per-day disabled days, invalid %H:%M fallbacks, unknown
timezone -> UTC, dim_brightness default 30, inactive-display short
circuit, and the _was_display_active/_was_dimmed transition flags.

test_display_controller.py's test_schedule_disabled and
test_active_hours patched config_service.get_config — which
_check_schedule never reads — so both asserted the init-default value
and could not fail. Rewritten on the test_inactive_hours pattern
(inject controller.config['schedule'], reset the minute gate, flip the
flag to the opposite state first so the assertion has teeth).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* ci: raise coverage floor to 48%

Measured 50% with the new suites in place (was 47% baseline when the
gate was introduced at 45); floor stays two points under measured.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: address CodeQL alert and review findings

- config_manager: the "secrets list longer than config list" warning now
  interpolates only config-side data (no key name or secrets-derived
  values), resolving the CodeQL clear-text-logging alert.
- base_plugin: validate_config rejects bool display_duration, matching
  get_display_duration (bool is an int subclass and would otherwise pass
  as a positive number).
- config_helper: merge_configs deep-copies override values in the
  non-recursive branch so mutating the merged result cannot reach back
  into override_config.
- saved_repositories: saves are atomic (temp file + fsync + os.replace),
  so a failed write can no longer truncate saved_repositories.json.
- tests: regression cases for each fix, plus a pin that whole-item
  array secrets (key[] + key[].field both marked) strip to empty {}
  skeletons — no secret values can reach config.json.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 16:17:11 -04:00
Chuck fc25a70d75 fix(web): make the update button work on branches without tracking (#443)
Reported from a pi whose checkout sat on a local branch:

    git pull failed (returncode=1): There is no tracking information for
    the current branch. Please specify which branch you want to rebase
    against.

The Tools tab reported that as "Update failed; check logs for details",
which tells the user nothing they can act on, and the underlying git
message never reached the UI at all.

A branch with no upstream is easy to end up on — checking one out by
name, restoring a backup, or following a guide that names a branch — and
until now it left the update button permanently broken with no way out
except SSH.

resolve_pull_command() now decides how to pull:
  - upstream set                  -> git pull --rebase, as before
  - no upstream, origin/<branch>  -> git pull --rebase origin <branch>,
    then attach tracking so the next update is a plain pull
  - no upstream, no remote branch -> an error naming the branch and
    pointing at Switch branch
  - detached HEAD                 -> says so, rather than failing obscurely

That resolution happens BEFORE the stash. Previously the handler stashed
local changes and then discovered it could not pull, putting the user's
work away for an update that was never going to run.

Failures now surface git's own message instead of "check logs".

Adds a branch picker to the Tools tab, backed by GET
/system/git-branches (local + remote-only) and a checkout_branch action.
Switching attaches tracking, so Pull Latest works afterwards. Branch
names are validated against a strict pattern before reaching a subprocess
argument list.

Local edits block a checkout, as they should. Rather than a truncated
one-line error, the response carries git's full list of blocking files
and a can_retry_with_stash flag; the UI then offers "Stash and switch" as
an explicit choice. Stashing is never done unasked — putting someone's
edits away without consent is worse than refusing the switch.

Verified on the pi that produced the report: on its untracked 'audit'
branch the update now returns the actionable message, git-info reports
upstream='' and can_pull=false, and an injected branch name is rejected.
27 tests build real git repositories and cover each path, including the
stash route that could not be exercised safely on the device.
2026-08-07 13:27:43 -04:00
ChuckandClaude 003312f4ff feat(web): let schemas label enum dropdown options (#442)
* feat(web): let schemas label enum dropdown options

An enum property renders as a dropdown whose option text is derived from
the value — underscores replaced, title case applied. That works when the
value reads as its own label and fails when it does not: "vs" renders as
"Vs", and "abbrev" tells the user nothing about the "Sep 19" it produces.
Schemas had no way to say otherwise, so the label was whatever the config
key happened to look like.

Enum dropdowns now take their option text from x-options.labels when the
schema supplies it. This is not a new convention: the checkbox-group
widget has read x-options.labels since it was written, with the same
humanised fallback. This extends it to plain enums and to array-table
columns.

Display only — the option value, and so the saved config, is unchanged.
The map may be partial; unlabelled values keep the humanised fallback, so
every existing schema renders exactly as before. Older cores ignore
x-options entirely, which means a plugin can ship labels without
requiring users to upgrade first.

Array-table columns get the same lookup but keep the raw value as their
fallback rather than the humanised one. Those columns hold values such as
ticker symbols, where "aapl" -> "Aapl" would be wrong, and they were not
being humanised before this change.

Verified against the running web service: with labels the hockey plugin's
date dropdown reads "Sep 19 / 9/19 / 19 Sep / 19/9 / Fri Sep 19"; with
the pre-change template and the same schema it falls back to
"Abbrev / Numeric / Day First / ...", confirming the degradation path.

* fix(web): label enum options in dynamically added table rows, and test the
shipped template

Both points from the CodeRabbit review on #442.

array-table.js built enum <option> elements with o.textContent = opt, so a
row added with "Add row" showed the raw value while the server-rendered
rows above it showed the schema's label — the same column reading two
different ways until the page was reloaded. Both option-building sites now
go through a shared enumOptionLabel(), which mirrors the template exactly:
x-options or x_options, labels map, raw value as the fallback.

The tests rendered a copy of the template expression, so they could pass
while production drifted. They now extract the live enum <select> block
out of plugin_config.html and render that, and assert on the full
value -> label map rather than substring presence.

Mutation-checked, since a guard that cannot fail is not a guard:
  - remove the labels lookup      -> 5 of 9 fail
  - change only the fallback to   -> 2 of 9 fail
    option|upper (keeping the
    enum_labels.get call intact)
  - revert the JS to raw values   -> 1 of 9 fails
The middle case is the one the review called out as able to slip through.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 13:27:20 -04:00
ChuckandClaude 31d607f6b3 fix(backup): make a restored device match the one that was backed up (#439)
* fix(backup): make a restored device match the one that was backed up

Found by wiping a working device and reinstalling from scratch. Every
problem here is invisible until you actually do that, which is why a
green test suite and eleven hours of uptime had not surfaced any of them.

**Four enabled plugins vanished on restore.** Weather, stocks, music and
leaderboard have a registry `id` that differs from the `id` in their own
manifest: the registry calls them `weather`, everything else calls them
`ledmatrix-weather`. Installation already prefers the manifest id for the
directory name and warns when the two disagree, so on disk, in
config.json and in a backup they are `ledmatrix-weather` -- but nothing
resolved that in reverse. Restore asked the store for `ledmatrix-weather`
and got "Plugin not found in registry", four times, and the device came
back missing four plugins the user had enabled.

Registry lookup now falls back to matching `plugin_path`, which already
records `plugins/ledmatrix-weather`. Renaming the published ids would
have orphaned `plugin_state.json` entries keyed on the old ones. Exact id
still wins, so a path that collides with another entry's id cannot
shadow it. Against the live registry and a real 28-plugin install this
takes unresolvable directories from five to one -- the one being
starlark-apps, which is genuinely not in the registry.

**Secrets could not be restored at all.** A fresh install left
config_secrets.json group-readable but not group-writable, and the web
interface -- which is what performs a restore -- does not necessarily run
as the owner. Every other file in the backup restored; secrets failed
with EACCES. Now group-writable, so the account running the web UI can
put them back.

**A partial restore reported "Restore had errors" and nothing else.**
That is the same message whether the whole thing failed or it quietly
dropped your API keys. It now names what was restored, what failed, and
which plugins were not reinstalled.

**ytm_auth.json was never in the backup.** It sits in config/ beside the
three files that are, and is pure device-local auth: losing it silently
signs the user out of YouTube Music. Backed up and restored with the
wifi config, which it resembles.

**Backups were written inside the directory a reinstall deletes.**
config/backups/exports is destroyed by the reinstall the user was told to
make it before. Exports now go beside the install, falling back to the
old path when that is not writable.

**The installer reboots without asking in non-interactive mode**, which
the README did not mention -- easy to hit when piping the install, and
alarming when a device you are installing onto disappears. Documented,
with --no-reboot-prompt. Its log also claimed root:ledmatrix while
printing a hardcoded group name rather than the one it used.

Tests: registry resolution gets its own suite, including the collision
case and third-party entries with an empty plugin_path. The existing
round-trip test passed throughout this because its fixture plugin has a
directory name equal to its id -- the one shape that cannot fail -- so it
now carries ytm_auth too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* ci: run the backup suites

test_backup_manager.py existed but was never enrolled, so the tests that
should have guarded backup and restore have not run on a pull request.
That is part of why the restore bugs in the previous commit reached a
device: the suite was there, it just was not watching. Adds it alongside
the new registry-resolution tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(backup): restore must not depend on owning the file it replaces

Follow-up from testing the previous commit on real hardware, where the
secrets fix turned out to be both too narrow and slightly wrong.

Too narrow: config.json, wifi_config.json and ytm_auth.json are installed
root-owned and group-readable exactly like the secrets file, so all four
were unrestorable by the web service, not just one. `shutil.copy2` opens
the destination for writing, which needs permission on the *existing
file*; the web user could create files in that directory all day and
still not replace them.

Slightly wrong: the previous commit loosened the secrets file to
group-writable. That was treating the symptom. The real error was
deciding ownership from `ledmatrix.service` -- the display service, which
runs as root and only ever *reads* secrets -- when the account that
*writes* them is the web interface, which deliberately does not run as
root. Ownership now follows the web service's user and the mode stays
640.

`_copy_file` writes a temporary file alongside the target and renames
over it. That needs only directory permission, so a restore no longer
cares who owns the destination, and it is atomic: a crash mid-restore can
no longer leave a half-written config. The destination's mode is carried
across so restoring secrets does not widen them to the umask, and its
owner is carried across too when the OS allows it -- only root can hand a
file to another user, so a restore run by the web service keeps its own
ownership rather than pretending to preserve root's.

Verified on a device with all four config files set root-owned 640 and
unwritable by the web user: before, every one failed with EACCES; after,
the restore reports success with no errors and all four sections
restored, mode still 640, root still able to read them and the web
service still able to write them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(backup): address CodeRabbit and CodeQL findings on PR #439

- first_time_install.sh: verify chown/chmod succeed and the final
  owner/group/mode on config_secrets.json before reporting success;
  exit with a clear error otherwise instead of swallowing failures.
- api_v3.py: replace the predictable .writetest probe with an
  exclusive NamedTemporaryFile to avoid a race with concurrent
  resolvers; log the preferred/fallback export path and OSError when
  falling back to the reinstall-deleted directory.
- api_v3.py: mark a restore as failed when plugin reinstalls fail,
  even if file restoration itself succeeded, so the endpoint no longer
  reports HTTP 200 success on a partial restore.
- api_v3.py: stringify plugin IDs before joining them into the error
  message so a malformed backup's non-string plugin_id can't raise a
  TypeError and mask the detailed response.
- backup_manager.py / api_v3.py: stop putting raw exception text (originating
  from a user-controlled backup file) into restore results returned to
  the client; log full details server-side instead. Addresses the
  CodeQL "stack trace information exposure" alert.
- test coverage: add a test for get_plugin_info() resolving a
  manifest id, and assert the disabled restore_wifi path also skips
  and omits ytm_auth.json.

Co-authored-by: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 13:02:59 -04:00
ChuckandClaude Fable 5 d6c5f97c13 Test suite overhaul + fixes for the three bugs it uncovered (#441)
* ci: run the whole test tree and make the plugin-safety job assert something real

The unit-tests CI job ran an explicit 24-file allowlist that had rotted:
63 of 90 test files (display, vegas, store manager, web API, web_interface)
never ran on a PR. The job now runs all of test/ (minus test/plugins, which
the plugin-safety job owns) so new test files are enrolled by default and
any exclusion needs a visible, commented --ignore.

The plugin-safety job was a green no-op: plugins/ is empty in CI, so every
test skipped with 'Manifest not found'. It now renders a bundled
deterministic fixture plugin (test/fixtures/plugins/ci-fixture-plugin,
golden images included for all 8 default sizes) via LEDMATRIX_PLUGINS_DIR,
and sets LEDMATRIX_REQUIRE_PLUGINS=1 so discovering zero plugins fails
loudly instead of skipping green. The per-plugin suites document that they
target dev machines with real plugins installed.

Coverage is now measured and enforced in exactly one place — the CI
unit-tests step (--cov=src --cov=web_interface --cov-fail-under=45, from a
measured 47% baseline). pytest.ini previously declared --cov-fail-under=30
but CI always passed --no-cov, so the gate had never run anywhere; local
pytest is now coverage-free and fast.

Enabling the 63 unenrolled files surfaced three cases of test rot, fixed
here: test_display_controller_vegas_tick.py could not collect without the
hardware rgbmatrix module (now uses the emulator convention), the
state-reconciliation unrecoverable-cache tests broke when production added
the is_plugin_uninstalled tombstone check (bare Mock returned truthy),
and test_get_system_status assumed the optional psutil dependency
(now installed via requirements-test.txt and guarded by importorskip).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: replace can't-fail tests with real assertions

test_font_manager.py was 5 of 6 tests shaped as 'try: call(); assert True /
except: assert True' — running in CI while unable to fail on any
regression. Rewritten against the real FontManager API and the bundled
assets/fonts: returned font types, cache-hit identity, distinct entries per
size, default-font fallback for unknown families and corrupt files
(recorded in failed_loads), BDF native-size reading, text measurement, and
cache lifecycle.

test_display_manager.py's test_draw_text ended in 'assert True'; it now
renders onto a known-black canvas and asserts pixels were actually lit —
which required un-breaking the fixture's freetype MagicMock so draw_text's
isinstance check doesn't silently swallow the draw.

test_display_controller.py carried a permanently-skipped test whose skip
reason already declared it redundant; deleted.

Both display test files now set EMULATOR=true before importing
display_manager (the same convention as test_display_dirty_tracking.py) so
they collect standalone instead of depending on which test module imports
display_manager first.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: cover the untested fragile logic (compatibility gate, secrets, config merges, durations, skin cards)

New unit tests for pure or filesystem-only logic that previously had zero
direct coverage:

- test_compatibility.py: the semver install gate (parse_semver suffix
  handling, every range operator, TRUSTWORTHY_FLOOR behavior for cores
  reporting untrustworthy versions, 'more restrictive wins', and the
  malformed-manifest shapes that used to raise).
- test/web_interface/test_secret_helpers.py: the canonical x-secret
  helpers — find/separate/mask/remove, array-item secrets, no input
  mutation, and a separate->recombine round-trip.
- test/web_interface/test_api_v3_helpers.py: the module-level helpers
  behind the plugin config save endpoint (_is_plugin_update_available,
  _coerce_to_bool including the int==1 quirk, deep_merge including its
  shared-subtree shallowness, _parse_form_value, dotted-key-aware
  _get_schema_property/_set_nested_value).
- test_base_plugin_duration.py: get_display_duration's full coercion
  ladder (instance attr -> config -> 15.0), including the bool-is-int
  quirk where display_duration=True means one second.
- test_config_manager_secrets.py: the secrets round-trip — deep-merge on
  load, strip on save, group pruning, the load fast path — and two
  characterized sharp edges marked SUSPECTED BUG: an unreadable secrets
  file at save time writes secrets into config.json in plaintext, and a
  same-mtime-same-size content swap is served stale.
- test_schema_manager_merge.py: merge_with_defaults branch behavior (None
  replacement vs falsey preservation, dict-vs-scalar mismatches, arrays
  replaced wholesale, defaults never mutated).
- test_skin_system.py (extended): render_skin_card shares _render_game's
  3-strike counter but never resets it on success — the asymmetry is
  pinned in both directions, along with card fallthrough and the disable
  interaction between the two paths.

Suspected bugs are characterized, not fixed — each carries a comment so a
future behavior change is deliberate rather than accidental.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: add drift guards for cross-file contracts

Three guard suites that pin contracts spanning multiple files, where one
side changing unilaterally breaks the other silently:

- test_version_comparison_consistency.py: the repo's four version
  comparators (compatibility.parse_semver, api_v3's packaging-based
  _is_plugin_update_available, store_manager update_plugin's raw string
  equality, skin_runtime._major) answer differently on the same inputs.
  A table pins each one's verdict; update_plugin is driven through its
  real code path to show the SUSPECTED BUGs: 'v1.2.0' vs '1.2.0'
  triggers a full reinstall the UI calls unnecessary, and a locally-ahead
  plugin gets downgraded. A pairwise-ordering check keeps parse_semver
  agreeing with packaging on plain X.Y.Z.
- test/web_interface/test_secret_separation_parity.py: api_v3.py carries
  three inline copies of find_secret_fields/separate_secrets that lack
  the canonical module's array-item support. The copy count is asserted
  exact (it may only go down; new copies must import
  src/web_interface/secret_helpers), the missing-array-support gap is
  asserted so it can't grow silently, and the canonical behavior that
  migration will adopt is documented executably.
- test_discovery_path_contract.py: the three 'where is plugin X'
  resolvers (PluginManager discovery, StoreManager._find_plugin_path,
  SchemaManager.get_schema_path) agree on the configured directory, and
  their divergent fallback chains are characterized. Also pins the
  .standalone-backup- naming contract shared by store rollback and
  discovery, and _resolve_skin_target's path-traversal rejection.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* test: address review feedback — fixture lifecycle, test names, ClassVar

- ci-fixture-plugin: call display_manager.clear() before rendering (per
  plugin guidelines — the fixture should model a well-behaved plugin),
  add a class docstring, and document why Pillow is deliberately not
  pinned in its requirements.txt (core dependency; harness installs
  nothing).
- Rename two tests whose names contradicted their assertions:
  test_unparseable_core_version_is_compatible ->
  test_unparseable_core_with_high_floor_is_blocked, and
  test_unreadable_secrets_file... -> test_corrupt_secrets_file...
- Annotate TestGetSchemaProperty.SCHEMA as ClassVar (RUF012).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* ci: allow manual test.yml runs via workflow_dispatch

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: unify version comparison, refuse secret-leaking saves, reset skin strikes on card success

Fixes the three suspected bugs this PR's characterization tests pinned,
flipping those tests to assert the corrected behavior:

- plugins/store: ONE shared update comparator. New
  compatibility.is_update_available() (PEP 440 via packaging) is now used
  by both the web UI's update badge (api_v3._is_plugin_update_available
  is a thin alias) and store_manager.update_plugin's reinstall decision.
  Previously update_plugin used raw string equality: 'v1.2.0' vs '1.2.0'
  triggered a full reinstall the UI called unnecessary, and a locally-
  ahead plugin (2.0.0 installed, registry 1.9.0) was silently DOWNGRADED.
  Now equivalent spellings skip the reinstall and locally-ahead versions
  are never downgraded; unparseable versions still reconcile by
  reinstalling from the registry.

- config: save_config and save_config_atomic now refuse (ConfigError)
  when config_secrets.json exists but cannot be loaded. Both previously
  proceeded without stripping, writing the merged secrets into
  config.json in plaintext. The shared _load_secrets_for_save() helper
  raises with an actionable message instead; a missing secrets file is
  still fine (nothing to strip), and _migrate_config's catch-all keeps
  boot resilient.

- skins: render_skin_card resets _skin_failures on both success paths
  (vegas card returned, or mode renderer handled), mirroring
  _render_game. Transient card failures no longer accumulate across a
  session until they permanently disable a working skin.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

* fix: harden shared comparator edges from review

- is_update_available: reject truthy non-string versions (a malformed
  manifest can carry a number; packaging raises TypeError on those) by
  surfacing the mismatch instead of raising.
- store_manager.update_plugin: drop the truthiness gate around the
  comparator so a missing version on either side follows the shared
  'no update' verdict, keeping the store consistent with the UI badge;
  a missing manifest still uses the reinstall recovery path.
- config_manager._load_secrets_for_save: catch only expected read/parse
  failures (OSError/ValueError/RecursionError) so implementation bugs
  propagate as themselves, and log with traceback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 10:17:30 -04:00
ChuckandClaude Fable 5 d9683e28be Codebase audit: fix shipping bugs, remove verified-dead code, repair doc drift, add regression guards (#438)
* fix(web): implement delete_cached so the font catalog cache actually invalidates

api_v3.py's font upload/delete handlers import delete_cached from
web_interface.cache, but the function was never defined. The surrounding
except ImportError silently swallowed the failure, so the fonts_catalog
cache entry survived uploads/deletes and newly uploaded fonts did not
appear until the TTL expired or the service restarted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(web): remove dead weather/stocks partial routes that returned 500

The partial dispatcher still routed 'weather' and 'stocks' to loaders
rendering v3/partials/weather.html and stocks.html — templates that no
longer exist since weather and stocks became store plugins. Requesting
either partial raised TemplateNotFound, which the catch-all turned into
a 500. No template or JS references these partials (the only 'weather'
hit in the front end is a plugin-store category filter option), so the
branches and both loader functions are removed; unknown partials now
fall through to the existing 404 handler.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(deps): align contradictory psutil/Flask-Limiter/freetype-py pins

requirements.txt's optional-install comment recommended psutil>=5.9,<6.0
while web_interface/requirements.txt hard-requires >=6.0,<7.0 — anyone
following the comment ends up with an unsatisfiable pair. The comment now
recommends the same range the web interface requires (all psutil APIs
used — Process, boot_time, cpu_percent, disk_usage, virtual_memory — are
stable in 6.x). Flask-Limiter gains the same <4.0 cap in both files and
freetype-py the same >=2.5.1 floor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(config): add template keys the code already reads

display.hardware gains pixel_mapper_config, row_address_type,
multiplexing and panel_type (read at display_manager.py with these exact
fallbacks — users on non-standard panels previously had no way to
discover them from the template). vegas_scroll gains
frame_based_scrolling and scroll_delay, the only two of its 27 keys the
template omitted (read in src/vegas_mode/config.py). plugin_system gains
development_mode, which the web UI reads and writes but the template
never declared.

Every added value is byte-identical to the code-side .get() fallback, so
ConfigManager._migrate_config() merging these keys into existing user
configs cannot change behavior on any installed device.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(scripts): repair broken sys.path setup in utility scripts

clear_cache.py and download_nba_logos.py pointed sys.path at a 'src'
directory relative to the script's own folder (scripts/utils/src and
scripts/src — neither exists), so both crashed on import; they now insert
the project root and import via the src package like the other scripts.
debug_web_manual.py resolved 'project root' to scripts/debug/ instead of
two levels up. fix_nhl_cache.sh is removed: it used Python docstring
syntax in a bash script and invoked clear_nhl_cache.py, which does not
exist anywhere in the repo — it cannot ever have worked in its current
location.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: correct stale file:line references and the loader-fallback contradiction

CLAUDE.md and .cursorrules disagreed about plugin-directory fallback
behavior; the code (SchemaManager.get_schema_path) probes plugins/
BEFORE plugin-repos/, and the main discovery path has no fallback at
all — both files now describe the real behavior, preferring symbol names
over line numbers so the references rot slower. REST_API_REFERENCE.md
pointed at app.py:144/:607 for mounts that live at :199/:799 and counted
92 routes where there are 94. PLUGIN_ARCHITECTURE_SPEC.md's historical
banner gains a note that its example imports
(src/plugin_system/base_classes/*_plugin.py) never shipped — the real
base classes are src.base_classes.sports.SportsCore and
src.base_classes.hockey.Hockey.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: fix broken links, phantom script references, and stale CI description

Repairs every broken relative link in active docs (targets renamed or
archived long ago: PLUGIN_DEVELOPMENT.md -> PLUGIN_DEVELOPMENT_GUIDE.md,
API_REFERENCE.md -> REST_API_REFERENCE.md, PLUGIN_STORE_USER_GUIDE.md ->
PLUGIN_STORE_GUIDE.md, plugin_docs/ dir, TROUBLESHOOTING_QUICK_START.md,
and MIGRATION_GUIDE's README link that silently resolved to the docs
index instead of the project README). Replaces commands invoking scripts
that do not exist (scripts/update_stats.py, validate_registry.py,
check_updates.py, fix_permissions.sh) with the real tooling, and
rewrites HOW_TO_RUN_TESTS.md's CI section, which described a
security-audit workflow that was never committed and a pytest workflow
'queued to land' that landed long ago as test.yml.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: complete the docs index and refresh the web interface file tree

docs/README.md's own policy says every page must be linked from the
index, yet five weren't — including the entire skin system
(SKIN_SYSTEM.md, CREATING_SKINS.md), ADAPTIVE_LAYOUT.md,
plugin-safety-harness.md and SPORTS_UNIFICATION.md. Each is now listed
in the section it belongs to, and PLUGIN_ARCHITECTURE_SPEC.md is marked
historical in the index (the doc itself already carries the banner).
web_interface/README.md's static/v3 tree showed only app.css/app.js;
it now reflects the actual contents.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove dead modules confirmed unused in-repo and across all store plugins

- src/common/cli.py: imports a 'ledmatrix_common' package that exists
  nowhere (not in this repo, any requirements file, or the plugin
  monorepo), so it cannot ever have run; its README section claimed
  scripts/dev/* used it, which was also untrue.
- src/web_interface/logging_config.py: zero callers — the web app uses
  web_interface/logging_config.py (a different module), and nothing
  imports the src copy.
- handle_errors decorator in src/web_interface/error_handler.py: zero
  call sites (the module's response helpers stay — they are used).
- ConfigManager.get_clock_config(): reads a 'clock' config key that no
  longer exists anywhere; only caller was its own unit test.

Deliberately kept despite zero in-repo callers: DisplayError,
src/common/config_helper.py and display_helper.py — all documented as
plugin-facing API (docs/PLUGIN_ERROR_HANDLING.md, src/common/README.md),
and third-party plugins outside the official monorepo cannot be
enumerated. Verified against a fresh clone of ledmatrix-plugins (43
plugins): zero references to any removed symbol.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove manager-era NBA test files and one-off debug scripts

The four test_nba_*.py files imported nba_managers, leaderboard_manager
and odds_manager — top-level modules deleted when sports displays became
plugins — inside try/except blocks that swallowed the ImportError, so
they passed while exercising nothing. test_nba_data_structure.py and
debug_nba_api.py (a diagnostic script living in test/) made live ESPN
API calls rather than testing repo code. None were enrolled in CI.

scripts/debug/direct_fix_imports.py and check_imports.py were one-shot
artifacts that edited/inspected a hardcoded ~/LEDMatrix/web_interface/
app.py to fix an import problem solved long ago; nothing references
them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove generate_report.py, which aggregates artifacts of CI jobs that do not exist

The script's only function is to merge JSON artifacts
(bandit/semgrep/pip-audit/safety/gitleaks results) produced by a
security-audit workflow that was never committed —
.github/workflows/ has no such jobs, so there is nothing for it to
aggregate and no way to run it usefully. Its siblings stay:
prove_security.py and audit_plugins.py both run standalone (verified),
and .codacy.yml stays because the Codacy service (README badge) reads it
server-side without a workflow file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(web): load the three widget scripts store plugins already declare

time-picker.js, file-upload-single.js and plugin-file-manager.js
register widgets that installed store plugins reference in their config
schemas (countdown uses x-widget: time-picker and file-upload-single;
of-the-day uses plugin-file-manager), but base.html never included the
scripts. plugin_config.html renders such fields as an empty container
that polls LEDMatrixWidgets.get(...) on a 50ms loop forever, so those
plugin config fields appeared permanently blank. The audit initially
flagged these files as dead code; the monorepo cross-check proved the
opposite — they were unreachable, not unused.

example-color-picker.js (the documented custom-widget example) gains an
explicit warning that including it in base.html would shadow the
built-in color-picker widget.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: drop the legacy youtube block from the secrets template

No code reads a top-level youtube secrets key: the youtube-stats plugin
receives its API key namespaced under its own plugin id (declared via
x-secret in its config schema), like every other store plugin. The key
survives only in state_reconciliation.py's non-plugin-key exclusion set,
which stays — existing installs still carry the key in their generated
config_secrets.json, and the exclusion prevents it from being
misclassified as a plugin config. New installs simply stop being asked
for a YouTube API key they have nowhere to use.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore(deps): remove packages nothing imports, declare direct imports, move mypy to test deps

Removed from requirements.txt: python-socketio, python-engineio,
websockets, websocket-client — zero imports anywhere in this repo, and
the one store plugin that needs Socket.IO (ledmatrix-music) declares it
in its own requirements.txt, which the plugin store installs. Removed
the same quartet plus timezonefinder, geopy, google-auth-oauthlib,
google-auth-httplib2, google-api-python-client, unidecode, icalevents,
python-dateutil, flask-wtf and the werkzeug pin from
web_interface/requirements.txt — all leftovers from the deleted built-in
weather/calendar/music displays (flask-wtf was doubly dead: app.py
explicitly disables CSRF and sets csrf=None). scripts/
install_dependencies_apt.py, which mirrors these lists for the
first-time installer, drops the same packages.

Added: urllib3 (imported directly in four core modules), jinja2 and
markupsafe (imported directly in pages_v3.py) — previously reachable
only as transitives. mypy moves from runtime requirements to
requirements-test.txt.

Verified in a fresh venv: all four requirements files co-install, pip
check is clean, the full CI-enrolled suite (907 tests) and a Flask boot
smoke pass with the trimmed dependency set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* refactor: single canonical DateTimeEncoder

src/cache_manager.py and src/cache/disk_cache.py each defined an
identical DateTimeEncoder (datetime -> ISO-8601). The disk_cache copy is
the only one actually used for serialization; cache_manager now
re-exports it instead of defining a twin, so the two can never silently
diverge. Import compatibility is preserved — from src.cache_manager
import DateTimeEncoder still works and is the same class object.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs(code): document deliberate duplicates instead of merging them

The audit surfaced several near-duplicate implementations that turned
out to be either deliberate forks or behaviorally different — merging
any of them would risk changing behavior on installed devices, so each
now carries an explicit comment stating the relationship:

- VisualDisplayManager: headless fork of DisplayManager; header now
  lists the ~15 mirrored methods and warns that DisplayManager changes
  must be mirrored.
- normalize_abbreviation: LogoDownloader's version (called directly by
  nine scoreboard plugins) replaces filesystem-unsafe characters;
  LogoHelper's strips spaces. Logo filenames on existing installs
  depend on both behaviors staying put.
- The two PluginTestBase classes: the shipped one is plugin-author
  API, the repo's own richer harness lives in test/plugins/ — now
  cross-referenced.

Also verified (no change needed): ConfigManager's backup/rollback
methods genuinely delegate to AtomicConfigManager, and SportsCore
already delegates _read_bdf_native_size to FontManager.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: add a unified configuration reference

There was no single place documenting what lives in config.json —
display.* keys were scattered across README sections, vegas_scroll lived
in ADVANCED_FEATURES.md, and dim_schedule, display.double_sided,
sync.follower_position, plugin_system.development_mode and the four
newly-templated hardware keys were documented nowhere. CONFIG_REFERENCE.md
now lists every template key plus the code-read-only keys, each with
type, default, and the code location that reads it, and explains the
secrets file's plugin-id namespacing. Linked from the docs index.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: bring the README's feature tour into the plugin era

The Core Features section still presented clock/weather/sports/stocks/
music displays as built into the project, when all of them are store
plugins installed from the ledmatrix-plugins monorepo — only
starlark-apps and web-ui-info ship in this repo. The intro now says so
(the showcase itself is unchanged; those are real displays available in
the store). The display_durations reference drops its built-in-calendar
example in favor of plugin-id keys, and the Configuration section links
the new CONFIG_REFERENCE.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: archive the custom-icons status report, cross-link config docs, document assets/

PLUGIN_CUSTOM_ICONS_FEATURE.md was a 'What Was Implemented' status
report duplicating the actual guide (PLUGIN_CUSTOM_ICONS.md) — moved to
docs/archive/ per the docs index's own policy. The overlapping
plugin-config docs keep their content but PLUGIN_CONFIG_ARCHITECTURE.md
now states up front which doc is canonical for which purpose.

assets/README.md is new and load-bearing: assets/stocks, weather,
news_logos and broadcast_logos have zero references in this repo's code,
which makes them look deletable — but store plugins (ledmatrix-stocks,
ledmatrix-weather, news, odds-ticker) resolve those exact paths at
runtime against the install directory. The README records that evidence
so a future cleanup doesn't break installed plugins.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* test: add regression guards for the bug classes fixed in this PR

Three lightweight static checks, all enrolled in CI's unit-test
allowlist along with the new web-cache test:

- test_template_targets.py: every literal render_template() target must
  exist (would have caught the weather/stocks partial 500s at commit
  time).
- test_widget_scripts.py: every widget JS file must be script-included
  in base.html or explicitly allowlisted with a reason (would have
  caught the unloaded time-picker/file-upload-single/plugin-file-manager
  widgets), and allowlisted files must NOT be included (prevents the
  example widget from shadowing the real color-picker).
- test_doc_links.py: relative markdown links in active docs must
  resolve (docs/archive/ exempt).

Each guard was verified to fail against the pre-PR tree and pass now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(deps): restore the werkzeug version floor

Commit 1ec22db removed the werkzeug>=3.1.6,<4.0.0 pin along with the
genuinely-unused packages, but this one was a version floor on Flask's
transitive dependency, not a phantom: Flask 3.1.3 itself only requires
werkzeug>=3.1.0, so dropping the pin let fresh installs resolve
3.1.0-3.1.5. Restored with a comment explaining why it exists.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix(web): route pixel_mapper_config into display.hardware; guard time-picker registration

pixel_mapper_config was the only display.hardware key absent from both
the display_fields detection allowlist and the hardware write loop in
the settings save path. No form posts it today, but if one ever did the
key would fall through to the generic handler and land at the TOP level
of config.json — where state_reconciliation would mistake it for a
missing plugin id and loop auto-repair attempts (the failure class the
'github'/'youtube' exclusion comment documents). It now round-trips
into display.hardware like its siblings.

time-picker.js gains the same LEDMatrixWidgets-undefined guard its two
sibling widgets already have; correct today only via defer ordering.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: align installer leftovers with the dependency cleanup

first_time_install.sh's fallback secrets heredoc (used only when the
template is missing) still wrote the legacy youtube block — now matches
the template (github only). install_dependencies_apt.py drops the
IMPORT_NAME_MAP entries for packages no longer in its install lists and
a stale google-api reference in a docstring.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix: address CodeRabbit review findings

Verified each finding against the code; fixes for the valid ones:

- install_dependencies_apt.py: the installer listed 'freetype', but the
  declared dependency is freetype-py — an apt miss would pip-install the
  wrong PyPI package. Now installs freetype-py with an import-name
  mapping (pre-existing bug, surfaced by the review).
- api_v3.py: pixel_mapper_config is validated as a string before being
  saved to display.hardware (JSON callers could previously store an
  object/list the matrix library can't use).
- .cursorrules: the Plugin Loading Process and File Organization
  sections still said discovery scans plugins/ — now consistent with the
  corrected overview (configured directory, default plugin-repos/).
- README.md: removed the stale '(except the core calendar)' claim — no
  core calendar exists in src/ — and qualified the plugin inventory
  (official plugins in the monorepo; third-party from their own repos).
- CONFIG_REFERENCE.md: hardware_mapping now shows the code fallback
  (adafruit-hat-pwm) alongside the template value.
- PLUGIN_REGISTRY_SETUP_GUIDE.md: check_plugin.py takes --plugin, not a
  positional id.
- scripts/fix_perms/fix_*.sh: exec bits set so the documented
  'sudo ./...' invocations work.
- Guard tests hardened: template guard now catches multi-line
  render_template() calls; widget guard parses actual <script> src
  values and fails if the widgets dir goes missing; type hints and
  docstrings added per repo coding guidelines.

Skipped with reasons (noted on the PR): limit_refresh_rate_hz 100-vs-90
is documented as intentional in CONFIG_REFERENCE.md; the psutil comment
already names the enforcing manifest; docs/archive/ findings are out of
scope per the docs policy (archive may rot).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: remove Cursor IDE tooling, consolidate its guidance into CLAUDE.md

The maintainer no longer uses Cursor. .cursorrules, .cursorignore and the
.cursor/ tree (rules, plugin templates, a parallel 751-line plugins
guide) are removed; measurement showed near-zero literal overlap risk —
the canonical content already lives in docs/. Unique guidance worth
keeping moved before deletion:

- CLAUDE.md gains the dev workflow (dev_plugin_setup.sh, dev_server.py,
  run.py -e, check_plugin.py), the plugin-secrets namespacing contract,
  and the no-draw_image()/paste-onto-PIL pitfall.
- PLUGIN_DEVELOPMENT_GUIDE.md absorbs the plugin version-management
  rules (pre-push hook install, SKIP_TAG, version resolution order) that
  its own text previously linked out to .cursorrules for.
- The one completed plan doc (.cursor/plans/) is archived to
  docs/archive/ per the docs policy rather than deleted.

One of the deleted rule files (sports-managers.mdc) targeted
src/*_managers.py globs that have matched nothing since the plugin
migration.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* chore: second-pass cleanup — dead installer branch, broken script, orphaned JS, misfiled test deps

- first_time_install.sh: removed the pip fallback branch that installed
  from requirements_web_v2.txt — a file that has not existed since the
  v2 web interface was removed (the branch always printed its own
  'not found; skipping' warning).
- scripts/remove_plugin_backups.sh deleted: its PROJECT_ROOT resolved to
  the repo's PARENT directory, and its verify_submodules() checks for
  plugin submodules from an era before plugins moved to the store — it
  could never have worked from its current location.
- plugins_manager.js: removed three functions with zero call sites
  anywhere (addKeyValuePair, formatCommit, togglePasswordVisibility) —
  verified against all templates, all JS, and the dynamic window[name]
  dispatch sites, which resolve widget-registry keys only. Also replaced
  base.html's misleading 'Legacy ... during migration' label: the file
  is deliberately loaded last and provides the LIVE implementations of
  seven window.* plugin actions that shadow same-named definitions in
  app.js/app-shell.js.
- pytest/pytest-cov/pytest-mock moved from runtime requirements.txt to
  requirements-test.txt (CI already installs both files; the installer's
  line-by-line loop simply installs three fewer packages on devices; no
  store plugin declares pytest). HOW_TO_RUN_TESTS.md updated.
- scripts/add_defaults_to_schemas.py and analyze_plugin_schemas.py
  scanned the empty legacy plugins/ dir — now scan plugin-repos/.

Verified: fresh venv installs all four requirements files with pip check
clean and pytest available; bash -n on the installer; node --check on
the JS; widget/cache guard tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* docs: correct semantically stale content across the user and developer guides

A second-pass content audit checked the guides' substantive claims
against the code (the first pass only fixed mechanical drift). Fixes:

- GETTING_STARTED: described booting a prebuilt SD image and seeing
  default clock/weather plugins — neither exists. Now documents the real
  install (Pi OS Lite + one-shot installer / first_time_install.sh) and
  that displays come from the Plugin Store. Duration and ordering
  instructions moved to the Rotation tab where the controls actually
  live.
- WEB_INTERFACE_GUIDE: three whole tabs were undocumented (Rotation,
  Backup & Restore, Tools) and the Display tab's Vegas Scroll section
  was unmentioned. Fonts overrides are per display element (not per
  plugin); Logs has an Auto-scroll checkbox (not a Pause button); the
  aspirational keyboard-shortcut list and no-JS claim removed.
- TROUBLESHOOTING: the hand-written service-file template (wrong user,
  wrong ExecStart, dropped the autostart gate) replaced with the real
  systemd/ units + install scripts; recovery steps no longer copy
  placeholder units verbatim; WiFi curl endpoint corrected to /api/v3/;
  cache-clearing advice now targets the real cache locations.
- ADVANCED_FEATURES: removed a false claim that CacheManager has no
  delete(); fixed two example snippets that raise TypeError
  (BackgroundDataService and get_config_file_mode signatures); fixed
  cache paths, a 5-minute TTL that is actually 1 hour, and the vegas
  table now links the complete 26-key reference.
- EMULATOR_SETUP_GUIDE: documented run.py flags that don't exist
  (--plugin/--test-plugins) removed in favor of dev_server.py and
  check_plugin.py; shipped emulator config values corrected (browser
  adapter default on :8888, not pygame).
- PLUGIN_QUICK_REFERENCE: drag-and-drop reordering is shipped, not
  'not yet supported'; discovery-fallback and registry-repo claims
  corrected. PLUGIN_API_REFERENCE: get_vegas_segment_width returns
  panels, not pixels. CONTRIBUTING: the repo uses flake8/mypy/bandit
  pre-commit hooks, not black/ruff, and tests need requirements-test.txt.
- SKIN_SYSTEM/DEVELOPER_QUICK_REFERENCE: stale module paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

* fix: raise the two new dependency floors past their CVEs, silence a deliberate re-export

All three of these were introduced by this PR, which is what makes them
worth fixing here rather than deferring.

`urllib3` and `jinja2` were added to the requirements so that direct
imports stop relying on transitives — right call, but both floors were
set to the version that introduced the API rather than a version that is
safe to install. `urllib3>=1.26.0` sits below roughly ten CVEs including
a decompression-bomb safeguard bypass, and `jinja2>=3.1.0` below five
including two sandbox breakouts. Raised to 2.7.0 and 3.1.6, which is what
a working device already runs, so no install is disturbed. The comments
now say the floor is a security floor, since the next person to read
"imported directly" would otherwise reasonably lower it again.

This is the same reasoning the PR already applied to werkzeug; these two
just missed it.

The `DateTimeEncoder` import in cache_manager is unused on purpose — the
canonical class moved to src.cache.disk_cache and this re-export keeps
the documented import path working. flake8 cannot see intent, so it gets
an explicit `# noqa: F401` rather than being removed and quietly breaking
anything importing it from here. Verified the re-export still resolves to
the same object and still serialises datetimes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix: address CodeRabbit re-review — installer robustness and doc lint

- install_dependencies_apt.py: an import-only check let Debian
  Bookworm's python3-freetype 2.3.0 satisfy the freetype-py>=2.5.1 pin.
  check_package_installed() now verifies the installed freetype-py
  version, and an apt install that lands below the minimum falls through
  to pip instead of counting as success.
- first_time_install.sh: the .web_deps_installed marker was created even
  when the smart installer failed, so re-runs skipped installation with
  dependencies missing. The marker is now created only on success.
- CONTRIBUTING.md: document installing the pre-commit CLI before
  'pre-commit install' (the requirements files don't provide it).
- Doc lint: fence language on the on-demand cache example (MD040),
  blockquote continuation in GETTING_STARTED (MD028), and the
  suppress_adapter_load_errors key removed from the emulator debug
  example to match the options table.

Skipped one finding with reason (noted on the PR): the per-plugin
display_duration field in PLUGIN_QUICK_REFERENCE's example is not
obsolete — BasePlugin.get_display_duration() reads it and
PLUGIN_CONFIG_CORE_PROPERTIES.md documents it as a core property;
display.display_durations is a per-mode override, not a replacement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 14:04:23 -04:00
ChuckandClaude Opus 5 2af41c561b fix(plugins): make update scheduling atomic so update() cannot run twice at once (#437)
* fix(plugins): make update scheduling atomic so update() cannot run twice at once

Closes #401.

`run_scheduled_updates()` decided whether to update a plugin with a
check-then-act sequence: `can_execute()` and `set_state(RUNNING)` were
separate calls with nothing between them, so two scheduler threads could
both observe ENABLED and both go on to call the same plugin's `update()`.
`update_all_plugins()` had the identical pattern.

Two schedulers really do run at once. The render loop calls
`_tick_plugin_updates()`, and Vegas mode fires its own `vegas-plugin-tick`
daemon thread that is never joined when `VegasModeCoordinator.play()`
returns — a slow `update()` still in flight overlaps the next tick from
the main loop. A plugin running `update()` twice concurrently is unsafe
unless it happens to be reentrant; shared mutable state, a non-thread-safe
HTTP session or cache all break.

The async path was already covered by the `_pending_lock` dedup in
`_enqueue_update`, so the live exposure was the synchronous kill-switch
path and `update_all_plugins()`. Both now claim the plugin through
`_reserve_for_update()`, which holds one lock across the eligibility
check, the due-time check and the RUNNING transition — and nothing more.
Holding it across `execute_update()` would serialize slow plugins behind
each other and reintroduce the render stall the async worker exists to
avoid.

The due-time check moved inside the lock deliberately. Left outside, a
thread that had already decided "due" could claim the plugin the instant
the winner finished, running `update()` twice within one interval.

Two supporting changes fall out of it:

- `_enqueue_update()` no longer sets RUNNING (the reservation did), and
  hands the reservation back if the pending-dedup ever fires. Otherwise a
  reserved-but-unqueued plugin would sit in RUNNING with nothing left to
  release it, and `can_execute()` would refuse it forever.
- `_finish()` now clears the pending entry *before* flipping the state
  back to ENABLED. The old order left a window where a scheduler saw
  ENABLED, reserved the plugin, then had its enqueue silently dropped by
  the dedup — harmless as a missed tick before, a stuck plugin once a
  reservation is involved.

Regression suite added and enrolled in CI, along with
test_async_plugin_updates.py which was not previously run there. The
overlap tests delay `can_execute()` to hold every thread inside the
check-then-act gap: the real window is a couple of bytecodes wide, so a
plain hammering test passes against the unfixed scheduler and proves
nothing. With that delay the suite reports `update() ran 8x concurrently`
on both affected paths before the fix, and passes after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* test: fix a race in the reservation suite's own wait loops

`test_async_path_never_overlaps` failed in CI with "update() ran 0x
concurrently" — the test's bug, not the scheduler's. It polled
`plugin._active` to wait for the update to finish, but before the worker
picks the item up nothing is active yet, so the loop fell straight
through and asserted on a plugin that had never run.

Both async waits now key on `update_calls >= 1` as well, so they wait for
an update to have started *and* finished. The stranded-state test gets
the same guard for a second reason: ENABLED is also the starting state,
so without it that assertion passes vacuously on a plugin that was never
scheduled.

Verified over 12 consecutive local runs, 12 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(plugins): roll back the claim when dispatch fails

_enqueue_update() reserved the plugin and added it to the pending set,
then started the worker and queued the item. Thread.start() raises
RuntimeError when the OS refuses a new thread — not hypothetical on a Pi
under memory or thread pressure — and nothing is queued at that point to
release the plugin. It stayed RUNNING with a stale pending entry, so
can_execute() refused it for the rest of the process, and the exception
escaped run_scheduled_updates() and skipped every remaining plugin in
that tick.

That is the same stranded-RUNNING failure the reservation was introduced
to prevent, just reached through the dispatch rather than the dedup, so
it is handled the same way: discard the pending entry, hand the
reservation back, log the cause. Swallowed rather than raised so one
plugin failing to queue cannot abort the others' turn.

Both new tests fail against the un-rolled-back version — the second on
the escaping RuntimeError itself — and pass with it. 74 tests across the
reservation, async-update, plugin-system, health, Vegas-adapter and
controller-toggle suites still pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WexvwNDtWLVymGVqKD7BGk

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 19:38:11 -04:00
ChuckandClaude Opus 5 5b81cca684 fix(espn): send an identifying User-Agent so ESPN stops returning 403 (#436)
Around 11:00 EDT on 2026-08-04 ESPN's site.api began rejecting the agents
this repo sends. Every scoreboard that goes through the shared data
sources returned `403 Client Error: Forbidden` — standings, game
summaries and scoreboards alike. A device that had been running fine
logged 287 ESPN errors in a day.

The filter is not the familiar one. Probing site.api across agents and
libraries, using requests as the plugins do:

    bare 'LEDMatrix/1.0'                     403   (with or without Accept)
    browser string                           403   (no header rescues it)
    'LEDMatrix/1.0 (+https://github.com/...)' 200
    requests / urllib / curl defaults        200

So it rejects browser-style strings outright and bare custom tokens, and
accepts honest client tokens or an agent that identifies the client and
links to it. The instinct to "just send a browser User-Agent" is now
exactly backwards — that is the one thing guaranteed to stay blocked.

Both call sites here sent a bare token: `LEDMatrix/1.0` in the sports
data sources and `LEDMatrix-Common/1.0` in the API helper. The Accept
header the data sources already sent does not save it. Both now send an
agent carrying the project URL, which also gives ESPN someone to contact
rather than an anonymous token to rate-limit.

Verified on a live device: ESPN errors went from a steady stream to zero
across a restart, with live MLB games fetching again.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 16:55:45 -04:00
ChuckandClaude Opus 5 d305be6089 fix(plugins): core's own config keys no longer flag plugins as degraded (#434)
Found while sweeping devpi for issues. Nine of 27 installed plugins were
reported degraded in the web UI -- including baseball-scoreboard and
f1-scoreboard -- for using a documented core feature.

The core reads three tuning keys out of each plugin's own config block:
vegas_width_pct and vegas_overflow (vegas_mode/plugin_adapter.py) and
vegas_max_width_screens (base_plugin.py). No plugin declares them, and 37 of
the 42 published config schemas set "additionalProperties": false -- so schema
validation reported them as violations.

That is not just log noise. _validate_config_schema_soft sets `degraded` in
the health tracker, which the web UI surfaces, so a user who tuned a core
Vegas setting saw the plugin marked broken.

The keys are stripped before validation. Fixing it plugin-side would mean 42
schema edits and 42 version bumps -- 42 store updates for a contract the core
owns.

Listed explicitly rather than matched on a `vegas_` prefix: vegas_mode is the
opposite case, plugin-owned and declared in schemas, and a prefix rule would
silently stop validating it.

Verified on devpi: degraded went 9 of 27 -> 0 of 27, schema-mismatch warnings
9 -> 0, 22 plugins still load, no tracebacks. 800 core unit tests pass,
8 of them new -- including that a genuine violation is still reported, so the
check has not been turned into a no-op, and that the caller's live config dict
is never mutated.


Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 13:46:28 -04:00
ChuckandClaude Opus 5 53af53b4a1 feat(store): evaluate compatible_versions, not just the floor (#433)
* feat(store): evaluate compatible_versions, not just the floor

Closes the gap CodeRabbit surfaced on #427. `compatible_versions` is the
canonical compatibility contract -- schema/manifest_schema.json marks it
required, all 42 published manifests carry it -- and it is the only field that
can express an *upper* bound. `ledmatrix_min_version` is a floor and cannot
say "not compatible with 4.x".

The gate read only the floor, so a plugin declaring ["2.0.0 - 2.9.9"] would be
installed on 3.2.0 regardless of having said it stops at 2.x.

check() now evaluates both and the more restrictive wins. The array is a set of
alternatives (satisfying any one entry suffices), supporting every form the
schema permits: >=, <=, >, <, ~, ^, a bare exact version, and an inclusive
"A - B" range, with prerelease/build suffixes tolerated.

Refusal still requires evidence. Anything unparseable, absent, or below
TRUSTWORTHY_FLOOR resolves to compatible.

That last point needed a new strict parser. parse_semver is deliberately
lenient -- it strips non-digits and yields (0, 0, 0) for a string with no
numbers at all. Harmless for a floor (0.0.0 never blocks) but wrong for a
range, where the same leniency turned an unreadable spec into a *refusal*: a
manifest whose only entry was garbage got compared against 0.0.0 and refused.
Range specs are now shape-checked first, so garbage reads as "no evidence".
parse_semver itself is unchanged, since the loader depends on its behaviour.

Verified: 815 core unit tests pass, 18 of them new. Swept the real registry --
all 42 published manifests, at cores 1.0.0 / 2.0.0 / 3.1.0 / 3.2.0 / 4.0.0 --
and nothing is refused at any of them. The gate stays inert for shipped
plugins, which is the property that makes it safe to land ahead of B5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* feat(store): protect the one population the sunset would break

The B6 sunset deletes each plugin's guarded-import fallback, so a plugin that
floors at 3.2.0 must never reach a core that lacks the 3.2.0 modules. The gate
could not stop that for the population most at risk.

A device installed from the v3.1.0 release reports __version__ = "1.0.0". The
gate treated anything below TRUSTWORTHY_FLOOR as "unknown, do not block" --
correct while every manifest floors at 2.0.0, because blocking would have
emptied the plugin store for those users. But after the sunset it hands them a
3.2.0-floored plugin with no fallback, which fails to load with one log line.
Nothing else in the system protects them: they cannot be told apart from a
genuine 1.0.0 install.

On an untrustworthy core the gate now refuses a floor ABOVE 2.0.0 and still
allows anything at or below it. A floor above the ecosystem baseline says the
plugin needs modules that arrived after 2.0.0, and a core reporting below that
-- whether it is the v3.1.0 release or something genuinely ancient -- will not
have them. Refusing leaves the user on the version they already run instead of
one that cannot load.

Measured against all 42 published manifests:

  today (every manifest floors at 2.0.0)
    core 1.0.0 / 2.0.0 / 3.1.0 / 3.2.0 / unparseable -> 0 of 42 refused
  after B6 (same manifests floored at 3.2.0)
    core 1.0.0 -> 38 refused, core 3.1.0 -> 38 refused, core 3.2.0 -> 0

So nobody loses the store today, and the sunset cannot reach a core that
cannot run it.

Two older tests asserted the previous "allow everything" behaviour; they now
express the new rule with a 2.0.0 floor, which is what their no-lockout intent
was actually about.

The 38-of-42 in that measurement surfaced a separate B6 trap, recorded here
because it will bite whoever raises the floors: four plugins (flights,
leaderboard, music, stocks) declare the floor as a TOP-LEVEL
`min_ledmatrix_version`, a third spelling, which declared_min_version checks
before the versions[] array. For those, editing versions[0] is a silent no-op
and the floor stays at 2.0.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(store): address review — malformed manifests, and suffixed versions

Both CodeRabbit findings on #433 verified against the code and fixed.

1. Malformed manifest sections raised instead of degrading. `requires` as a
   list hit AttributeError ('list' object has no attribute 'get') and
   `versions` as a mapping hit KeyError: 0. Both reproduced.

   This got worse with the sunset rule in the previous commit: that branch
   resolves the floor for *every* manifest on an untrustworthy core, where the
   old code returned early. One hand-edited or third-party file with the wrong
   shape would have taken down the whole install path rather than just itself.
   Container types are now validated and an unrecognised shape reads as "no
   declared floor".

2. Prerelease and build metadata leaked into the version numbers. The digit
   scrape parsed "3.2.0+build42" as (3, 2, 42) and "3.2.0-rc1" as (3, 2, 1) --
   a release candidate ranking above its own release. Both fed reject
   decisions, and the consequence was demonstrable: a plugin pinned to exactly
   "3.2.0" refused a core running 3.2.0+build42, which is that same version.

   The suggested remedy -- use the strict token parser -- would not have
   fixed it. _parse_strict validates the shape but delegates the numbers to
   parse_semver, so it returned the same (3, 2, 42). The bug is in the scrape,
   so suffixes are now dropped before it. Prereleases compare equal to their
   release rather than below it; full prerelease ordering is more than any
   caller needs and equal is far closer to right than what it did before.

parse_semver is shared with PluginLoader, so its suite was re-run: unchanged,
and it only ever gets more correct here.

Verified: 839 core unit tests pass, 21 of them new -- six malformed shapes,
five suffixed forms, and the two demonstrated regressions. The real-registry
sweep is unchanged at 0 of 42 refused across cores 1.0.0, 3.1.0, 3.2.0,
3.2.0+build42 and an unparseable string.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 13:46:05 -04:00
ChuckBuildsandClaude Opus 5 183e23edb3 docs(changelog): record the compatibility gate in 3.2.0
The 3.2.0 section described the unified sports library but none of the
install-path work that landed in #428 and #431 -- which matters more than a
normal changelog omission, because the sunset rule keys on this section to
tell plugin authors what a given floor buys them.

The headline addition: 3.2.0 is the first release that *enforces*
ledmatrix_min_version. Before it the floor was advisory, so a plugin could
declare one and still be delivered to a core that could not run it. That is
the property B6 waits on, and it is now stated where a plugin author will
look for it -- along with the caveat that a core reporting below 2.0.0 is
treated as unknown rather than old and is never blocked.

Also records compatibility.py (and that it does not yet read
compatible_versions), check_release_version.py and its workflow, the
install-preservation fix, the reentrant-lock deadlock fix, and the
web_interface version re-export.

No version bump: 3.2.0 is unreleased, so this describes the release being
cut rather than a new one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
2026-08-03 16:54:53 -04:00
ChuckandClaude Opus 5 970ca2d04f feat(store): refuse to install a plugin that needs a newer core (re-target of #429) (#431)
* feat(store): refuse to install a plugin that needs a newer core

`ledmatrix_min_version` was decoration. The loader logged an advisory warning
and continued; the store never compared the core version at all, so a routine
"update" delivered a plugin that could not run. That is the gap phase B6 (the
sports-unification sunset) cannot be done over: deleting a plugin's bundled
fallback while nothing enforces the floor hands un-updated users a scoreboard
that raises ModuleNotFoundError at load and is reported only as one line in
the journal.

The gate lives in install_plugin, after the manifest is on disk and before
dependencies are installed. That is the earliest knowable point -- the
registry carries no compatibility field, so the floor is not visible until
the files are down -- and it is also the chokepoint: _reinstall_with_rollback
calls install_plugin, so a refused *update* restores the version the user
already had, for free.

Floor resolution and the comparison move to src/plugin_system/compatibility.py,
shared with the loader so the two cannot drift. Both read all four spellings
published manifests use, including the deprecated `ledmatrix_min`.

Refusal requires evidence. An undeclared floor, an unparseable version on
either side, or a core below TRUSTWORTHY_FLOOR (2.0.0) all allow the install.
That last one is deliberate and load-bearing: the v3.1.0 release reports
__version__ = "1.0.0" while nearly every published manifest floors at 2.0.0,
so a strict gate would lock those users out of the plugin store entirely --
much worse than the problem being solved. They stay unprotected until they
update the core, which is also what fixes their version string.

Verified: 782 core unit tests pass, including 25 new ones and the existing
loader-warning suite unchanged (the refactor is behavior-preserving). The
install tests drive the real install_plugin path with the download stubbed --
the allow and refuse cases differ only in the declared floor, so the refusal
is demonstrably the gate and not an earlier bail-out.

Follow-ups, deliberately not in this PR: surfacing the reason in the store UI
rather than only the log, and publishing the floor in plugins.json so the
store can refuse before downloading.

Phase B4 in docs/SPORTS_UNIFICATION.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(store): a failed install must not destroy the plugin it replaced

Found while validating the compatibility gate. `_install_plugin_impl` deletes
the existing plugin directory *before* downloading, so any failure after that
point leaves the user with nothing. `_reinstall_with_rollback` protects the
update path exactly this way; a direct `install_plugin` had no equivalent.

The gate made this reachable in a new way: a plugin whose declared floor
exceeds the running core is now refused *after* the old copy is already gone.
Floors are hand-written and can be over-declared, so the refusal could remove
a plugin that had been working fine on that core.

install_plugin is now a thin wrapper that renames any existing install aside,
delegates to _install_plugin_impl, and restores it on failure -- including
when the implementation raises, which is re-raised after the restore. It is a
pass-through when nothing is installed and when called from
_reinstall_with_rollback, which has already moved the old copy aside; a test
pins that so the two mechanisms cannot start nesting.

The aside name embeds '.standalone-backup-' because
plugin_manager._scan_directory_for_plugins keys on exactly that substring to
skip backups. A different name would have made the backup discoverable as a
duplicate plugin; a test pins that too.

789 core unit tests pass, including 7 new ones.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(store): serialize concurrent installs, and make the lock reentrant

Second bug found while validating the previous commit on hardware.

install_plugin's new set-aside/restore had no lock. The web UI runs Flask
threaded, so a double-clicked Install button gives two threads the same
plugin_id; interleaved, one thread's restore deletes the other's freshly
installed copy. _reinstall_with_rollback already guards exactly this with a
per-plugin lock, and install_plugin needs the same one.

Taking that lock naively deadlocks. _reinstall_with_rollback holds it across
its call to install_plugin, and threading.Lock is not reentrant -- so the
request thread hangs forever on the standard monorepo update path
(update_plugin -> _reinstall_with_rollback -> install_plugin), which is to say
on every plugin update. Verified by reverting to a plain Lock: the regression
test times out after 10s instead of passing.

The per-plugin locks are now RLocks, and install_plugin holds one for its
whole set-aside/install/restore sequence.

Verified on devpi (Pi, Python 3.13.5, real registry and network):
- update_plugin on an up-to-date plugin: True in 5.4s
- update_plugin forced through the full reinstall-with-rollback path:
  True in 13.1s, correct version restored, old copy replaced, no backup
  directories left behind
- install -> reinstall-over-existing -> failed-reinstall-restores: all pass
  against real downloads
- 22 plugins load, no tracebacks, web API and UI 200, steady-state journal
  50 lines/min

791 core unit tests pass, including 2 new concurrency tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* ci: run the new suites, and check tag/version agreement at release time

These were split out of #428/#429 because the token pushing them lacked the
`workflow` scope. Folding them in here rather than opening a stacked PR --
#429 was merged into its stacked base after that base had already been
squash-merged, so its content never reached main, and one such near-miss is
enough.

All three enrolled suites exist on this branch: test_version_consistency.py
came with #428 and is on main; the other two arrive with the commits above.
Enrolling them in a separate PR would have either raced with this one on
test.yml or briefly pointed CI at files main did not have.

- test.yml: enroll test_version_consistency, test_plugin_compatibility_gate
  and test_install_preserves_existing in the core unit job. Until now these
  32 tests existed but nothing ran them automatically.

- release-version-check.yml: run scripts/check_release_version.py on pushed
  v* tags and published releases, plus workflow_dispatch so a tag can be
  checked *before* it is created. No dependencies -- it reads src/__init__.py
  and CHANGELOG.md only.

Verified: both workflow files parse, and the release check still passes for
v3.2.0 against this tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 16:52:42 -04:00
ChuckandClaude Opus 5 f2b246ef03 fix(version): make the core version have exactly one answer (#428)
* fix(version): make the core version have exactly one answer

Plugin compatibility floors compare against src.__version__, so that string
has to be trustworthy. It has not been. v3.1.0 was tagged 2026-05-31 while
src/__init__.py still said "1.0.0"; the bump did not land until 2026-07-12.
Every device installed from that release reports 1.0.0, which is below the
(2, 0, 0) floor in PluginLoader._warn_if_incompatible -- so those users are
silently exempt from every plugin compatibility warning.

web_interface carried a third answer, a hardcoded "3.0.0" that nothing read
and that had drifted two majors from the core. It now re-exports the
canonical value, so it cannot disagree again.

Adds:

- test/test_version_consistency.py (enrolled in the core unit CI job):
  src.__version__ is parseable semver, matches the newest CHANGELOG heading,
  the CHANGELOG's headings are unique and descending, and web_interface
  tracks the core. src.plugin_system.__version__ is deliberately excluded --
  it versions the plugin API and moves independently.

- scripts/check_release_version.py + a release-version-check workflow that
  asserts the tag, the CHANGELOG and src.__version__ agree. Runs on pushed
  v* tags and published releases, and via workflow_dispatch so a tag can be
  checked *before* it is created:

      python scripts/check_release_version.py v3.2.0

Verified: 757 core unit tests pass including the four new ones; the script
exits 0 for v3.2.0 and non-zero for both a mismatched tag (v3.1.0) and a
non-semver one (v2.5); web_interface and web_interface.app still import.

Prerequisite for cutting v3.2.0 -- phase B4 in docs/SPORTS_UNIFICATION.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(version): address review — regex strictness, OSError, stale doc claim

From CodeRabbit on #428, all three valid:

- The module docstring claimed the tag check "runs at release time in
  .github/workflows/release-version-check.yml". That workflow is held back to
  a follow-up PR (the pushing token lacks the `workflow` scope), so the claim
  was false as written. Both files now describe the script as a manual
  pre-flight and say the CI wiring is still to come.

- `\d` also matches non-ASCII decimal digits, which int() happily parses, and
  `\s` matches newlines -- so "##\n3.2.0" read as a version heading. Patterns
  now use [0-9] and [ \t], kept in step across the test and the script, with a
  regression test pinning both behaviours.

- A missing or unreadable CHANGELOG.md raised OSError out of read_text() and
  printed a traceback. In a release gate that reads as "the tooling is
  broken"; it now reports the path and a recovery action and exits 1.

Verified: v3.2.0 passes, a mismatched tag exits 1, and a missing CHANGELOG
exits 1 with the new message instead of a traceback. 5 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:21:47 -04:00
ChuckandClaude Opus 5 963ab8292a docs(sports): split adoption from sunset, and say what makes the sunset safe (#427)
* docs(sports): split adoption from sunset, and say what makes the sunset safe

The phase table folded two steps with very different risk profiles into one
B5: adopting core imports (safe by construction -- the guarded import keeps
the bundled fallback) and deleting the bundled copies (removes the fallback,
so the import becomes a hard dependency). They are now B5 and B6.

Reading the enforcement path showed the declared floor protects nobody today:

- PluginLoader._warn_if_incompatible is advisory, and skips entirely when the
  parsed core version is below 2.0.0.
- The v3.1.0 release ships __version__ = "1.0.0" -- the tag was cut
  2026-05-31 and the string was not bumped until 2026-07-12 -- so the skip
  matches exactly the users most likely to be behind.
- Neither StoreManager.install_plugin nor .update_plugin compares the core
  version, so a store update delivers a plugin that floors above the core.

Verified against a v3.1.0 worktree: sports_scroll, element_style and the
sports package are absent there, and the import fails with
exc.name == 'src.common.sports_scroll' -- a guard set of {"src"} does not
match it.

B4 therefore grows to include making the version number trustworthy and
adding the install/update gate; B6 waits on that gate having shipped and
reached users. Also adds an ordered "What's next" and the durable lessons
this migration paid for.

Documentation only; no code changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* docs(sports): address review — test matrix, and compatible_versions

Both CodeRabbit findings on #427 hold up against the code; one sub-point was
already moot.

1. The compatibility regression test was described as one case (bundled copy
   removed on an old core) when it needs four. The case that actually matters
   is the one that was missing: an adopted plugin loading *with* its bundled
   copy on an old core, which is the entire basis for claiming B5 is safe to
   run ahead of the gate. Now a 2x2 table.

   The assertion was also wrong. "Fails loudly and specifically" is
   aspirational -- PluginManager.load_plugin catches ModuleNotFoundError, so
   nothing propagates and it fails into PluginState.ERROR with one log line.
   A test expecting a raise would pass for the wrong reason. Specified as
   PluginState.ERROR plus the exact missing module path, which is also what
   the rest of this document already says the failure looks like.

2. `compatible_versions` -- not `ledmatrix_min_version` -- is the canonical
   contract: schema/manifest_schema.json requires it, all 42 published
   manifests carry it, and it holds semver ranges ([">=2.0.0"] in 41,
   [">=1.0.0"] in 7-segment-clock). The gate as merged reads only the floor.

   Harmless today: no manifest uses an upper bound, and the two fields agree
   everywhere except 7-segment-clock. But the schema's range syntax permits
   upper bounds, so a plugin declaring ["2.0.0 - 2.9.9"] would be installed on
   3.2.0 regardless. Recorded as a named gap the gate must close before B6,
   and the migration step now has to reconcile both fields across every
   manifest the gate can refuse.

   Skipped, with reason: the finding also asked to migrate the deprecated
   top-level `ledmatrix_version`. No manifest carries it -- verified across
   all 42 -- so there is nothing to migrate.

Documentation only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:21:34 -04:00
ChuckandClaude Opus 5 16fbb7ebeb fix(install): survive the rgbmatrix build on low-memory Pis (#430)
* fix(install): survive the rgbmatrix build on low-memory Pis

The one-shot installer failed at Step 6 on a 1GB Pi with "Failed building
wheel for rgbmatrix", and told the user to install build tools they already
had. The real cause was the kernel OOM killer.

Upstream's pyproject.toml declares no [tool.scikit-build] options, so
scikit-build-core drives Ninja at its default of nproc+2 jobs -- six
concurrent compiles on a 4-core Pi. CMakeLists.txt compiles the same 14
sources three times (~45 translation units), two of them Cython-generated
C++ where a single cc1plus peaks near 800MB. That does not fit in 512MB-1GB
of RAM.

Add scripts/install/lib_lowmem.sh and wire it into the installer:

- Cap build parallelism at max(1, min(cores, RAM/768)) via
  CMAKE_BUILD_PARALLEL_LEVEL, which is what cmake --build actually reads.
  MAKEFLAGS is ignored by Ninja and is set only as a Makefile-generator
  fallback. A 4GB Pi 4 still gets 4 jobs; 512MB and 1GB boards get 1.
- Add a temporary swapfile sized to bring RAM+swap to 3GB (capped at 2GB),
  removed once the build finishes. An EXIT trap is the backstop for the
  error path. Nothing is written to /etc/fstab or /etc/dphys-swapfile.
  Existing swap is measured excluding zram, which is compressed RAM and so
  does not help a build OOM.
- Keep pip's build tree off tmpfs. Debian 13 mounts /tmp as tmpfs, so the
  default held the whole C++ build tree in RAM alongside the compiler.
- Diagnose OOM failures from the build log and the kernel ring buffer,
  instead of always blaming missing build tools. The OOM killer writes
  nothing to pip's output, which is why this was misreported.
- Report RAM and the chosen job count in the Step 1 preflight, and emit a
  heartbeat during the compile so a deliberately serial 15-25 minute build
  does not look like a hang.

New flags --skip-swap and --build-jobs N, with LEDMATRIX_SKIP_SWAP and
LEDMATRIX_BUILD_JOBS equivalents.

Also skip the duplicate apt-get update that the one-shot installer and
first_time_install.sh each ran a minute apart, and complete the
dphys-swapfile advice in diagnose_dependencies.sh with the CONF_MAXSWAP
line, without which raising CONF_SWAPSIZE above 2048 is silently clamped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VdsJs65WnUo8BtKHMAGA1q

* fix(install): address review findings on the low-memory build path

Three fixes from PR review:

- Validate --build-jobs / LEDMATRIX_BUILD_JOBS before check_memory's
  fallback return. When lib_lowmem.sh is absent that return also honoured
  the override, so a non-numeric value skipped validation and instead blew
  up later in an arithmetic test in Step 6 with a generic error.
- Fall back to the default TMPDIR when the disk-backed build directory
  cannot be created, rather than pointing the build at a path that does
  not exist. A nearly-full disk is the likely cause on exactly the devices
  this targets.
- Pass LEDMATRIX_APT_UPDATED explicitly to the sudo child instead of
  relying on -E, which a sudoers env_reset/env_keep policy can strip.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VdsJs65WnUo8BtKHMAGA1q

* fix(install): don't add 30s to every rgbmatrix build

The build progress heartbeat slept for the full 30s report interval
before re-checking whether the compile had finished, so every build paid
up to 30 seconds of dead wall time -- including fast ones on a Pi 4/5 and
every --force-rebuild run.

Poll every 2s and report every 30s instead. Measured: 30s of overhead on
an instant build drops to 2s, with heartbeats still emitted on the same
schedule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VdsJs65WnUo8BtKHMAGA1q

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-03 15:21:18 -04:00
ChuckandClaude Sonnet 5 21825cbfbc Sports unification phases 1–2: package split, promoted methods, opt-in capabilities (#426)
* fix(fonts): resolve asset paths against the install root, not the cwd

FontManager built its catalog from cwd-relative paths ('assets/fonts'),
so any process started outside the install root — the plugin safety
harness on CI being the recurring case — found no fonts and silently
degraded every plugin to PIL's default face. Several plugins grew
per-plugin workarounds for exactly this (countdown, text-display,
tide-display in the plugins monorepo).

Catalog population now falls back to the install root derived from this
module's location when the cwd-relative path is missing; behavior when
running from the install root is unchanged. Verified: resolve_font
returns the real FreeType face from a foreign cwd, and the full unit
suites (266 tests) pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* docs: seed CHANGELOG.md with the module-availability release discipline

The plugins monorepo's sunset rule ('delete a bundled fallback copy only
when the manifest floors on the first core release shipping the module')
needs core module additions recorded against version numbers. Seeds the
changelog at 3.1.0 and documents the discipline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* ci: enroll the core unit suites in a dedicated job

The existing workflow ran only the three plugin-harness suites; the
skin-system, font-manager, data-source, extractor, scroll-helper,
adaptive-layout, and loader-compat suites (266 tests) existed but never
ran in CI, so a refactor of src/base_classes or src/common could regress
them silently. Also enrolls the new sports characterization and
element-style suites landing in this branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* feat: ship src/element_style — the per-element style resolver plugins already expect

Three plugins (of-the-day, ledmatrix-music, football-scoreboard) import
src.element_style behind guarded try/except with classic fallbacks, but
the module never existed in core, so the richer per-element styling UI
those code paths implement has been dormant. This lands it:

- ElementStyleResolver.style() resolves per-element font/size/color with
  the key semantic the consumers encode: a config value counts as
  user-forced only when it differs from the schema default (the web UI
  bakes defaults into config.json on save), and untouched configs
  resolve to exactly the caller's classic values — byte-identical
  rendering, proven by of-the-day's committed goldens passing unchanged.
- defaults_from_schema_file parses both declaration forms (the compact
  x-style-elements map and hand-written customization blocks).
- expand_style_elements() expands x-style-elements into full config
  blocks; schema_manager.load_schema() applies it (guarded, no-op for
  schemas without the declaration) so the config form and defaults
  merging see the expanded UI.
- Fonts resolve cwd-independently with (path, size) caching; .bdf loads
  via freetype like FontManager; nothing in the module raises out of
  style().

Verified: 31 new unit tests; of-the-day's previously-skipped 9-test
spec suite now runs and passes; football's resolver tests pass (27);
music's 38 plugin tests pass; schema-manager suites pass (43).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* test: characterization suite for src/base_classes/sports.py ahead of unification

Pins current behavior before the planned merge of the nine drifted
plugin copies back into this ancestor: the _extract_game_details_common
key contract per sport (reusing GUARANTEED_KEYS from the skin tests),
update() flows for upcoming/recent/live against cache-seeded fixtures
under frozen time, rendering smoke per mode class, and guard rails on
the skin-system seam.

Five surprising behaviors are pinned AS-IS and flagged in comments so
the merge changes them knowingly or not at all: is_upcoming also
matching status.type.name; hockey dropping events whose competitors
lack 'statistics'; baseball reading the event-level status for innings;
no past-date filter in upcoming; and favorites-only mode with an empty
favorites list showing nothing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* ci: restrict the test workflow's GITHUB_TOKEN to contents:read

CodeQL flagged the new unit-tests job for running with the default
unrestricted token; the pre-existing job had the same exposure. Both
jobs only check out the repo and run pytest, so a workflow-level
contents:read is sufficient.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* refactor(sports): convert sports.py into a package (pure move)

Phase B1a of docs/SPORTS_UNIFICATION.md. src/base_classes/sports.py
becomes a package so the upcoming capability modules have a home and
diffs show their blast radius:

  sports/__init__.py   re-exports the public API
  sports/core.py       SportsCore
  sports/modes.py      SportsUpcoming / SportsRecent / SportsLive

No logic change: the 1515 class-body lines are byte-identical to the
original (verified by concatenating the two modules and diffing against
HEAD). Only module docstrings and the redistributed import blocks are
new. MRO and __abstractmethods__ are unchanged, and every existing
import site — including 'from src.base_classes.sports import SportsCore'
in the sport subclasses, the skin tests, and the characterization
suite — resolves through the package __init__.

One test edit was required: the characterization suite monkeypatched
'src.base_classes.sports.get_background_service', which is no longer a
module attribute on a package. Retargeted to
'src.base_classes.sports.core.get_background_service' — the module whose
globals SportsCore.__init__ actually resolves, so the patch is effective
exactly as before. No test logic or assertion changed.

Also adds docs/SPORTS_UNIFICATION.md: the architecture for the whole
B1-B5 sequence — how upgradability (guarded imports, capability probing,
frozen view-model keys, the sunset rule), reusability (promote only what
all nine copies share), and modularity (capabilities as opt-in mixins
rather than config branches, variants as named strategies, sport-unique
code as declared override points) are kept as three separate mechanisms.

Verified: characterization + skin 94 passed; the 10-file unit suite 338
passed; test/plugins 60 passed — all identical to pre-change counts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* feat(sports): promote the nine universal methods into the base classes

Phase B1b of docs/SPORTS_UNIFICATION.md. Every method here is present in
all nine bundled plugin sports.py copies and absent from core, so this is
reuse of code the fleet already agreed on — not new behavior. The
promotions are inert until B5: the plugins' own overrides still run.

SportsCore: cleanup, _get_layout_offset, _load_custom_font_from_element_config
SportsUpcoming: _select_games_for_display
SportsRecent: _get_zero_clock_duration, _clear_zero_clock_tracking,
              _select_recent_games_for_display
SportsLive: _is_game_really_over, _detect_stale_games

Where the copies disagreed, the canonical form was chosen on evidence and
the genuine per-sport differences became seams rather than branches:

- _favorite_key(game, side) -- NRL matches favorites on team id because its
  abbreviations are ambiguous (NEW is both Newcastle Knights and New
  Zealand Warriors). Default is the abbreviation; NRL overrides. Core never
  learns the string nrl.
- FINAL_PERIOD / CLOCK_COUNTS_DOWN -- hockey ends in P3, and soccer/afl/nrl
  clocks count UP, so 0:00 means kickoff, not expiry.
- _config_schema_path() / _font_root() -- plugin-supplied locations, never
  derived from this module's __file__.

BEHAVIOR CHANGE (baseball, ufc): the rejected variant coerced a missing or
non-str clock to the literal 0:00 and then declared the game over at
period >= 4. MLB has no game clock and period is the inning, so live games
were being evicted from the 5th inning onward; UFC likewise. The promoted
variant skips the clock check when the clock is unusable -- it fails safe
(keeps showing the game) instead of failing destructive.

Also fixes a regression from the package move in e591cec: the bodies were
byte-identical but __file__ gained a directory, so _resolve_project_path's
parents[2] silently began resolving to <root>/src instead of the repo root.
Both it and _font_root now derive from a single _INSTALL_ROOT constant, so
a future move needs one line changed rather than two hand-counted depths.
Tests assert the resolved values, not the index.

The font loader takes baseball's body (BDF memo cache + native-strike
retry) under hockey's Optional signature -- the older lineage is the
correct one here, and basketball's positional str default breaks on an
explicit None. It resolves through _font_root rather than the cwd, so it
does not reintroduce the bug just fixed for FontManager, and delegates to
FontManager for the alias table and BDF header parse instead of shipping
second copies. cleanup gained the two new font caches and still leaves
background_service alone -- it is a process-wide singleton.

Verified: 111 new tests (48 core + 59 modes + 4 install-root regression);
characterization + skin suites still exactly 94, unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* fix(sports): stop dropping hockey and baseball events on optional feed keys

Both bugs were pinned AS-IS by the B0 characterization suite so this
phase could change them knowingly. Both fixes are adoptions of code the
corresponding plugins already ship, not new inventions.

Hockey: the extractor read competitor["statistics"] unguarded, so a
competitor arriving without that array raised KeyError inside the
generator and the WHOLE event was discarded -- valid scores and status
included. Shot/save counts now default to 0, which is already what the
suite expects for an empty statistics array.

Baseball: for live games the extractor read game_event["status"], the
event TOP-LEVEL status, to get the inning. Real ESPN events duplicate
status there, but MiLB events (synthesized from the MLB Stats API into
an ESPN-like shape) populate only the competition-level one, so the
lookup raised a bare KeyError and dropped the event. It now reads the
competition-level status that _extract_game_details_common has already
validated, so it cannot be missing at that point.

The two characterization tests that pinned the old behaviour are
rewritten to assert the fix rather than deleted, so the suite still
documents the edge case -- and still totals 94.

CHANGELOG records these plus the live-clock change from aaabc61 under
Changed/Fixed, since all three are user-visible. The two new promotion
suites join the CI unit job (449 tests).

Verified: unit job 449 passed, plugin-safety job 60 passed, and the
hockey (16) and baseball (24) plugin harnesses render clean at every
panel size.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* fix(sports): harden the live game-over check and font/log init

Follow-up review findings on the promoted base-class methods.

_is_game_really_over:
- `period` present-but-None raised TypeError on `None >= FINAL_PERIOD`,
  taking down the whole live-update pass (_detect_stale_games has no
  try/except). Same failure shape as the null `period_text` already fixed.
- An expired clock spelled "00:00" normalizes to "0000", which matched
  none of the hand-listed literals, so a finished game with a two-digit
  minute clock stayed on the scoreboard forever. Compare numerically.

SportsCore:
- _load_fonts kept the cwd-relative "assets/fonts/..." literals the
  _font_root() seam exists to remove, so every scoreboard font degraded
  to PIL's default face outside the install root.
- _should_log read self._last_warning_time unguarded while only an
  unrelated method initialized it lazily; the first warning of a run
  raised AttributeError. Initialize it in __init__.

Also documents that game_update_timestamps is written by subclasses, not
by the base class, so the staleness branch is inert until B5 adoption.

14 new tests. Gates: 463 core unit, 60 plugin safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* feat(sports): opt-in celebration and rotation capabilities

Phase B2 of the sports unification. Both features exist in only some of
the nine scoreboards, so they ship as capabilities the plugin composes,
never as `if self.<feature>_enabled` branches inside the base classes: a
sport that does not opt in has none of this code in its MRO.

CelebrationMixin (afl, nrl, soccer, football)
The two lineages spelled this differently -- _check_for_goal /
celebrate_opponent_goals vs _check_for_score / celebrate_opponent_scores
-- but the bodies were identical apart from three things, each now a
seam rather than a branch:
  - wording -> score_phrase() / win_phrase() hooks
  - follow-up suppression -> COALESCE_SCORING_SEQUENCE, on for football
    where a touchdown lands as +6 then +1, off where two increments are
    two real goals
  - team identity -> _favorite_key, so nrl matches on team id without
    core learning why its abbreviations are ambiguous
Both config spellings are read, so a plugin adopting the mixin keeps
working with the keys already in its published schema.

Rotation strategies
The three "dialects" turned out to be one algorithm (SWRR) in two
shapes: an incremental picker holding state across calls, and a
precomputed per-cycle list. They agree within a cycle and differ only at
the boundary, so core ships both behind a name registry rather than
declaring a winner. weight_for is supplied by the host, so rotation.py
never learns what a favorite is; an unknown name degrades to "simple"
because it arrives from user config.

Each strategy is checked against a verbatim transcription of the plugin
code it replaces, over every live-game shape up to four games -- the
differential B5 will delete the bundled copies on the strength of.

185 new tests. Gates: 648 core unit, 60 plugin safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* feat(scroll): upstream the scroll orchestration layer; release 3.2.0

Phases B3 and B4.

B3 -- src/common/sports_scroll.py is deliberately NOT a superset of the
ten plugin scroll_display.py copies. A method-level comparison of the
eight that share a shape (f1 and ufc are genuine forks) found a sharp
split, and the module is drawn along it:

  promoted   orchestration -- get_all_vegas_content_items is identical
             in all eight; clear_all, get_scroll_info,
             get_dynamic_duration, is_complete and display_frame are
             96-100% similar
  promoted   settings -- one algorithm; the copies differ only in which
             league keys they walk, so the ladder is data
             (SCROLL_LEAGUE_KEYS) rather than a body per sport
  NOT        content -- prepare_scroll_content has 8 distinct bodies
             across 8 plugins (145 lines, 53% similar at worst) and
             _load_separator_icons 7 (6% at worst)

Same name, different job: prepare_scroll_content draws *this sport's*
game card. Merging those eight bodies would be exactly the mistake the
promotion rule exists to prevent, so the base raises NotImplementedError
rather than rendering something plausible -- a base that rendered
something would let a plugin ship a silently blank scroll.

The one behavior added over the plugin copies is native
global_config['target_fps'] support. The bundled copies hardcode ~100
FPS via scroll_delay and never consult the global target; Part A
threaded it through each copy by hand, and this makes that threading
legacy compatibility rather than the mechanism.

66 tests, including three against the real ScrollHelper rather than a
double -- a suite built entirely on MagicMock would sail straight past a
rename in the helper.

B4 -- bump src/__init__.py to 3.2.0 and close the CHANGELOG's Unreleased
section against it. This is the number the sunset rule keys on: the
first core release shipping the unified sports library, and therefore
the floor a plugin sets ledmatrix_min_version to before deleting its
bundled copies. The version bump and the changelog release heading move
together on purpose -- separating them would leave a commit whose
changelog announces 3.2.0 while the code still reports 3.1.0.

Nothing here changes what an existing plugin loads; adoption is B5.

Gates: 714 core unit, 66 plugin safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* fix(sports): per-type warning cooldowns and font-load logging

Follow-ups from the second review pass, both on already-fixed findings:

_should_log accepted a warning_type and ignored it, sharing one
timestamp across every kind of warning -- so an API-error warning
silenced an unrelated cache warning for the next minute, and whichever
fired first won. Cooldowns are now keyed by type. Nothing in core calls
this method, so no behavior regressed; _last_warning_time is kept in
step for subclasses that read it directly.

_load_fonts logged through the module-level logger, dropping the manager
context, and had no return type hint. It now uses self.logger (set well
before _load_fonts runs) and names the directory it searched -- the bare
"Fonts not found" sent people hunting for a font-format problem when the
actual cause is an install missing assets/fonts.

Gates: 717 core unit, 66 plugin safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* docs: record the validated hockey scroll-display pilot for B5

B5 cannot ship until this PR merges and 3.2.0 exists -- a plugin cannot
floor ledmatrix_min_version at a release that does not exist, and an
unguarded src.common.sports_scroll import would break every user on
3.1.0.

The pilot has been validated ahead of that gate: hockey's
scroll_display.py adopted against a core carrying 3.2.0 goes from 691 to
289 lines with all 16 harness renders byte-for-byte identical.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* fix(sports): harden the B2/B3 capabilities against bad config and subclasses

Review pass on the phase B2-B4 changes. Every fix here is the same shape
as the crashes this PR already fixed in the hockey and baseball
extractors: a config or feed value that is present-but-wrong reaching
arithmetic or a comparison on a path with no guard.

celebrations:
- celebration_duration is coerced and floored at init. It is compared
  numerically in display() *outside* any try block, so a string from a
  hand-edited config propagated a TypeError straight out; zero or
  negative armed a celebration that could never render.
- A render failure now disarms instead of staying armed. It previously
  retried the same broken render on every frame for the rest of the
  window -- a traceback per frame, and no scorebug either.
- prune_score_baselines() for the live set. Only _check_for_win removed
  entries, so a game that left the live list any other way leaked its
  baseline and the dict grew all season.
- display() reuses has_active_celebration() rather than repeating its
  window comparison, and log lines carry a [Celebrations] prefix.

rotation:
- MAX_WEIGHT ceiling. A cycle is sum(weights) long and each step scans
  every game, so an unbounded weight from a misread config spins the
  display thread -- on a Pi that stalls rendering outright.
- register_rotation_strategy rejects a non-subclass factory at
  registration instead of failing frames later inside schedule().
- schedule() previews through type(self), so a subclass overriding
  next_game is previewed with its own ordering -- which is what the
  method promises.

sports_scroll:
- scroll_speed / scroll_delay coerced. dict.get(key, default) only helps
  when the key is absent; present-but-null reached the multiplication
  inside __init__ and the display failed to construct at all.
- update_scroll_position and get_visible_portion moved inside the try.
  They ran outside it, so a raise there reached the plugin's frame loop
  despite the comment promising none can.
- prepare_and_display guards the subclass call, so one sport's bad
  payload cannot take down the shared orchestration for the others.
- _current_game_type spells "nothing active" as "" in both classes; the
  manager said None while the display said "".

Not taken: the report that baseball's favorite-team debug path still
reads event-level status. Verified against current code -- there are no
remaining game_event["status"] reads in that file; it was fixed in
2486bdb and the finding is stale.

Gates: 747 core unit, 66 plugin safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* fix(baseball): don't drop favourite MiLB games on the diagnostic path

The competition-level status fallback fixed the inning lookup, but the
favourite-team debug block a few lines above still read the event top-level
game_event["status"]. MiLB events (synthesized from the MLB Stats API into an
ESPN-like shape) populate only the competition-level status, so the identical
event that extracted fine for a non-favourite raised KeyError and returned
None once the team was a favourite.

Worst possible shape for the bug: it only hit the games the user cared most
about, and only on the path meant to help diagnose them. The existing
regression test missed it because it never passes favourites, so
is_favorite_game was False and the block never ran.

Uses the validated competition-level `status`, which
_extract_game_details_common guarantees is present by that point. Adds a
favourites-passing companion test; confirmed it reproduces the KeyError
without the fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* docs(sports): type-hint game_update_timestamps to match its sibling

Addresses the last remaining sub-point on the modes.py review thread. The
design finding itself is already handled: the base class documents that it
only reads game_update_timestamps and that a subclass's update() owns writing
"last_seen" (and afl/etc. do, so stale-game eviction works in practice). The
one concrete gap was the missing annotation -- _zero_clock_timestamps is typed
Dict[str, float] while this nested map had none. Now Dict[str, Dict[str, float]].

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Address CodeRabbit review: font-name traversal + offline test guard

Two Minor findings from CodeRabbit's first review of this PR.

- resolve_font_path: reject relative font names carrying path components.
  font_name comes from plugin config, which the web UI writes; a value like
  "../../config/config.json" escaped assets/fonts/ after os.path.join and let
  a config probe arbitrary paths for existence (disclosure unlikely, since
  Pillow/freetype reject non-font files, but the probe is real). Relative
  names must now be bare filenames (os.path.basename(name) == name); absolute
  paths keep their existing isfile() gate. Test confirms the traversal
  resolved the real config.json before the guard.

- build_manager fixture: patch requests.Session.get BEFORE constructing the
  manager. Construction creates both SportsCore.session and the
  ESPNDataSource.session; the old code only replaced manager.session after
  the fact, leaving data_source.session real and able to reach the network on
  an accidental fetch. Patching the class makes every session built in the
  fixture offline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* test(celebrations): make the expiry tests actually test expiry

CodeRabbit (Major) on the merge re-review: celebration_duration is clamped to
a 1.0s floor, so the two expiry tests that configured 0 and expected instant
expiration never actually hit the expiry branch. They passed only because
_draw_celebration_layout raises in the harness (no real fonts) and its
exception branch clears the celebration the same way -- so they were really
re-testing the render-failure path, not expiry.

Now use a valid 1s duration, backdate started_at past the window, and mock
_draw_celebration_layout with assert_not_called() so an expired celebration
provably does NOT render. Verified discriminating: both fail if
has_active_celebration is forced to never expire.

Production code unchanged -- the expiry logic was already correct; only the
tests were mismodelling it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-02 12:39:33 -04:00
ChuckandClaude Fable 5 82a65ad2a2 Sports unification phase 0: safety net, cwd-independent fonts, element_style (#425)
* fix(fonts): resolve asset paths against the install root, not the cwd

FontManager built its catalog from cwd-relative paths ('assets/fonts'),
so any process started outside the install root — the plugin safety
harness on CI being the recurring case — found no fonts and silently
degraded every plugin to PIL's default face. Several plugins grew
per-plugin workarounds for exactly this (countdown, text-display,
tide-display in the plugins monorepo).

Catalog population now falls back to the install root derived from this
module's location when the cwd-relative path is missing; behavior when
running from the install root is unchanged. Verified: resolve_font
returns the real FreeType face from a foreign cwd, and the full unit
suites (266 tests) pass.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* docs: seed CHANGELOG.md with the module-availability release discipline

The plugins monorepo's sunset rule ('delete a bundled fallback copy only
when the manifest floors on the first core release shipping the module')
needs core module additions recorded against version numbers. Seeds the
changelog at 3.1.0 and documents the discipline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* ci: enroll the core unit suites in a dedicated job

The existing workflow ran only the three plugin-harness suites; the
skin-system, font-manager, data-source, extractor, scroll-helper,
adaptive-layout, and loader-compat suites (266 tests) existed but never
ran in CI, so a refactor of src/base_classes or src/common could regress
them silently. Also enrolls the new sports characterization and
element-style suites landing in this branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* feat: ship src/element_style — the per-element style resolver plugins already expect

Three plugins (of-the-day, ledmatrix-music, football-scoreboard) import
src.element_style behind guarded try/except with classic fallbacks, but
the module never existed in core, so the richer per-element styling UI
those code paths implement has been dormant. This lands it:

- ElementStyleResolver.style() resolves per-element font/size/color with
  the key semantic the consumers encode: a config value counts as
  user-forced only when it differs from the schema default (the web UI
  bakes defaults into config.json on save), and untouched configs
  resolve to exactly the caller's classic values — byte-identical
  rendering, proven by of-the-day's committed goldens passing unchanged.
- defaults_from_schema_file parses both declaration forms (the compact
  x-style-elements map and hand-written customization blocks).
- expand_style_elements() expands x-style-elements into full config
  blocks; schema_manager.load_schema() applies it (guarded, no-op for
  schemas without the declaration) so the config form and defaults
  merging see the expanded UI.
- Fonts resolve cwd-independently with (path, size) caching; .bdf loads
  via freetype like FontManager; nothing in the module raises out of
  style().

Verified: 31 new unit tests; of-the-day's previously-skipped 9-test
spec suite now runs and passes; football's resolver tests pass (27);
music's 38 plugin tests pass; schema-manager suites pass (43).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* test: characterization suite for src/base_classes/sports.py ahead of unification

Pins current behavior before the planned merge of the nine drifted
plugin copies back into this ancestor: the _extract_game_details_common
key contract per sport (reusing GUARANTEED_KEYS from the skin tests),
update() flows for upcoming/recent/live against cache-seeded fixtures
under frozen time, rendering smoke per mode class, and guard rails on
the skin-system seam.

Five surprising behaviors are pinned AS-IS and flagged in comments so
the merge changes them knowingly or not at all: is_upcoming also
matching status.type.name; hockey dropping events whose competitors
lack 'statistics'; baseball reading the event-level status for innings;
no past-date filter in upcoming; and favorites-only mode with an empty
favorites list showing nothing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* ci: restrict the test workflow's GITHUB_TOKEN to contents:read

CodeQL flagged the new unit-tests job for running with the default
unrestricted token; the pre-existing job had the same exposure. Both
jobs only check out the repo and run pytest, so a workflow-level
contents:read is sufficient.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FgbA8SMutQQpXkMG8LMmC4

* Address CodeRabbit review: font-name traversal + offline test guard

Two Minor findings from CodeRabbit's first review of this PR.

- resolve_font_path: reject relative font names carrying path components.
  font_name comes from plugin config, which the web UI writes; a value like
  "../../config/config.json" escaped assets/fonts/ after os.path.join and let
  a config probe arbitrary paths for existence (disclosure unlikely, since
  Pillow/freetype reject non-font files, but the probe is real). Relative
  names must now be bare filenames (os.path.basename(name) == name); absolute
  paths keep their existing isfile() gate. Test confirms the traversal
  resolved the real config.json before the guard.

- build_manager fixture: patch requests.Session.get BEFORE constructing the
  manager. Construction creates both SportsCore.session and the
  ESPNDataSource.session; the old code only replaced manager.session after
  the fact, leaving data_source.session real and able to reach the network on
  an accidental fetch. Patching the class makes every session built in the
  fixture offline.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-02 12:01:56 -04:00
ChuckandClaude Sonnet 5 83f20b64fe Give plugins access to device-wide config, and a global scroll frame rate (#424)
* Give plugins access to device-wide config, and a global scroll frame rate

The sports scoreboards read `getattr(self, 'global_config', {})` to find a
shared scroll frame rate, but nothing ever set that attribute: the loader
constructs plugins with only plugin_id/config/display_manager/cache_manager/
plugin_manager (plugin_loader.py:671), `global_config` appears nowhere in
src/, no plugin manager assigns it, and BasePlugin has no __getattr__ to
synthesize it. The lookup always returned {}, so the ten scroll_display.py
copies that thread target_fps through to ScrollHelper could never fire on any
core. There was also no global target_fps to find -- the only one in the
template is display.vegas_scroll.target_fps, which is Vegas-scoped.

Adds the missing half:

- `BasePlugin.global_config` resolves the full config via
  plugin_manager.config_manager, then cache_manager.config_manager, then {}.
  Same order the sports timezone helpers already use. Exceptions are swallowed
  to debug so an unreadable config can never stop a plugin loading, and a
  non-dict result is rejected rather than handed to callers that will .get()
  it and feed the result to numeric code.
- A top-level `target_fps` (default 100), exposed on the General tab and
  validated 30-200 on save to match ScrollHelper.set_target_fps -- which
  clamps silently, so a rejected save reports a value that would otherwise
  appear to save and then behave differently.

The property has a setter deliberately. news, stock-news, ledmatrix-stocks,
ledmatrix-elections, ledmatrix-leaderboard and nfl-draft all assign
`self.global_config = config.get('global', {})`; without a setter that raises
"property has no setter" and those six plugins stop loading. Reproduced, then
pinned with a test.

target_fps is also kept out of the `is_general_update` key list: that branch
treats a missing web_display_autostart as an unchecked box, so counting a
target_fps-only POST as a General save would silently switch autostart off.

Verified end to end: config.json -> BasePlugin.global_config ->
scroll_display's existing block -> ScrollHelper.target_fps 120 -> 100, with no
plugin-side change needed. Suite 1441 passed; the 4 failures
(test_display_dirty_tracking, test_web_api::test_get_system_status, two in
test_state_reconciliation) are pre-existing and reproduce identically on a
clean tree.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Address review findings on target_fps validation and config resolution

- Reject floats and bools before int() in the target_fps save path. A JSON
  body can carry them, where int(90.5) silently stored 90 and true stored 1.
  Form posts send strings, so '90.5' already failed in int().
- Assert the template's target_fps is 100, not merely an int, so the
  documented default is actually pinned.
- Empty-config precedence: keeping the `and config` check deliberately, now
  spelled out in the comment and covered by a test. Both managers default to
  the same config/config.json, so falling through cannot pick up a different
  file's settings; treating {} as an answer would instead return {} when the
  first manager simply hasn't loaded yet, silently disabling every setting
  read through the property -- the failure this property exists to fix.

Suite 1446 passed. The float-rejection test was checked to fail without the
guard.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-02 10:40:12 -04:00
ChuckandClaude Sonnet 5 5b45f35888 Vegas mode: reclaim dead space and pace the rotation (#423)
* Vegas mode: reclaim dead space and pace the rotation

On a wide panel Vegas mode spent much of its time showing black. At 50px/s
on a 512px display, one display width of blank is 10.2 seconds, which makes
several long-standing behaviours expensive:

- ScrollHelper prepended a full display width of black as an "initial gap",
  charged once per cycle — 10.2s of black at the start of every rotation.
- Plugins without get_vegas_content() are captured off a full-display canvas,
  so their blank margins entered the ticker too. Measured: of-the-day drew
  35px of "No Data" on a 512px canvas (92% blank), youtube-stats 142px of
  content with 185px of black either side. Only the scroll_helper path had
  any trimming.
- Cycle transitions deliberately pushed a blank frame and then recomposed
  synchronously: 84ms at best, 4.8s at worst, every millisecond of it black.
- buffer_ahead doubled as the cycle size, so a 21-plugin install showed 3
  plugins per cycle and took ~7 cycles to come around.
- separator_width was applied between every image rather than at plugin
  boundaries, so a per-row ticker like the F1 scoreboard (116 images, which
  it renders 4px apart internally) got a 32px chasm between each row — and
  the width budget didn't count those gaps, so the plugin quietly occupied
  far more of the panel than intended.

Changes:

- src/vegas_mode/geometry.py: numpy column-ink primitives shared by the
  trimmer and the audit tool, so the number reported is the number acted on.
  A Python per-column loop over a 17,000px strip is far too slow for the
  render path.
- PluginAdapter trims every content path, not just scroll_helper. Only outer
  edges are cropped: interior blank columns are the plugin's own layout
  (logo left, score right) and closing them would corrupt the design. A
  plugin on a non-black background is inherently unaffected.
- ScrollHelper.create_scrolling_image takes an explicit lead_gap, still
  defaulting to display_width so the many standalone-ticker callers are
  unchanged. Vegas passes lead_in_width (default 0).
- Cycle end holds the last rendered frame instead of blanking, turning the
  recompose into a brief freeze rather than the panel switching off.
- plugins_per_cycle (default 6) is split from buffer_ahead, which goes back
  to being only a prefetch low-water mark.
- max_plugin_width_ratio (default 3x display width) caps one plugin's share
  of a cycle. Overflow is deferred, not discarded: a rotation offset advances
  each fetch so later rows appear on subsequent cycles. Single oversized
  images are cropped at a blank column so the cut misses glyphs.
- Composition groups images by plugin: rows are joined by intra_plugin_gap
  (default 8) and separator_width applies only between plugins. The width
  budget now counts those gaps.
- Plugin data updates no longer run on the Vegas render path.

All new settings are user-configurable in Display -> Vegas Scroll, including
min/max cycle duration and dynamic duration, which previously existed in code
but were reachable only by hand-editing config.json.

Measured with scripts/dev/vegas_audit.py on a 512x64 panel:

  mean ink coverage    42.7% -> 69.4%
  fully blank           5.9% -> 0%
  reads as empty        13.6% -> 0%
  worst blank stretch    4.8s -> 0s
  full rotation          414s -> 123s
  plugins per cycle         3 -> 6

Note the metric choice: a "fully blank" scan (>=95% black viewport) reported
only 0.4% and badly understated the problem, because two full-width segments
with mid-canvas content never fully blank the viewport — they hold it at ~28%.
window_coverage_stats grades every viewport position by how much ink it
carries, which is what tracks perceived dead time.

Known remaining: cycle transitions still freeze ~3.5s while the next cycle is
fetched. Fixing that needs background prefetch, which is deferred because the
fallback-capture path mutates the shared display_manager.image and racing it
against the render loop risks torn frames.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Drop unused Optional import from the vegas audit script

Flagged by Codacy (F401). Any, Dict and List are all still used.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Align Vegas API bounds with validate(), fix audit config plumbing

Both from review feedback on #423.

The web API's accepted ranges disagreed with VegasModeConfig.validate(),
which is what actually gates Vegas starting:

  scroll_speed      1-100  -> 1-200   (a slider value of 150 returned 400)
  separator_width   0-500  -> 0-128
  target_fps        1-200  -> 30-200
  buffer_ahead      1-20   -> 1-5

The three loose ones were the dangerous direction: the value saved with a
200, then VegasModeCoordinator.start() failed validation with only a log
line, so the ticker silently never ran. The UI already matched validate() in
all four cases, so the API was the odd one out.

test_vegas_api_bounds_match_validate parses the numeric_fields map out of
api_v3 and asserts every bound against validate(), plus that validate()
accepts both endpoints and rejects just outside them, so these cannot drift
apart again. That test immediately caught a missing upper bound on
min_plugin_width, now added — unbounded it would drop every segment and
leave a blank ticker.

Separately, vegas_audit.py constructed PluginAdapter without the config, so
it fell back to VegasModeConfig() defaults and would report trimming and
width-budget behaviour that differed from the user's config.json. It now
passes the loaded config exactly as the coordinator does. This is the same
class of drift the explicit lead_gap and grouping arguments already guard
against. Output is unchanged on a rig whose config matches the defaults.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Vegas mode: render plugins narrower, space rows by measured separation

Trimming reclaims blank margins but cannot compact a layout that genuinely
spans the display — a five-column forecast, a progress bar drawn at 100%
width, a stat block with the panel's whole width between its elements. Those
need the plugin to make different layout decisions, which means telling it the
screen is narrower while it renders.

DisplayManager.render_size() presents a smaller logical canvas for the
duration of a Vegas content fetch, reusing the same _LogicalMatrix
indirection double-sided mode already relies on so plugins see a consistent
size from every accessor. Plugins that size themselves from matrix.width need
no changes at all; one that wants to be explicit can read the new
BasePlugin.get_vegas_render_width().

Width is a percentage so a single setting travels across panel sizes:
vegas_scroll.render_width_pct globally, or vegas_width_pct in an individual
plugin's config. Measured on a 512x64 panel with real data:

  ledmatrix-weather   1536px -> 576px   (forecast becomes narrow cards)
  youtube-stats        353px -> 199px   (2% blank left, so genuinely compact)
  geochron             453px -> 153px   (ink density rises to 100%)
  ledmatrix-flights    950px -> 740px

The youtube-stats figure is the clearest evidence the layout itself changed
rather than being cropped: at full width the content had to be trimmed from
512px to 353px, whereas at 40% it arrives with almost no blank to reclaim.

Row spacing is now measured rather than added. A flat gap gets it wrong in
both directions at once — content drawn flush to its own edges ends up nearly
touching (reported for recent sports scores, which sat 8px apart), while
content already carrying wide margins gets pushed even further out.
separation_gap() measures the blank each pair already has and adds only the
shortfall, up to min_content_separation (default 24). intra_plugin_gap stays
as a floor applied regardless.

Two tests shipped in the previous commit encoded the old flat-gap arithmetic
and are updated to the measured semantics, including one renamed to reflect
that zero intra_plugin_gap alone no longer butts rows together.

Also fixes a real bug found while testing: the harness display manager had no
render_size(), and because the adapter catches broadly that surfaced as "no
content" rather than an error, silently dropping five plugins. Added the
context to VisualTestDisplayManager for parity, and _render_at() now degrades
to a no-op on any display manager lacking it, so a third-party or older
harness loses the narrowing rather than the content.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Vegas mode: end cycles before the wrap, keep the width budget honest

Three fixes, the first a regression from lead_in_width defaulting to 0.

get_visible_portion wraps: once scroll_position + display_width passes the end
of the strip it fills the right of the frame from the *head* of the same strip.
So the final display_width of travel showed the cycle's first plugin re-entering
on the right while its last plugin exited on the left, and the recompose that
followed replaced both at once. On a 512px panel at 50px/s that was 10.2s of
two plugins on screen at once, ending in a hard cut — reported as the ticker
"switching mid-scroll" from F1 to news.

That used to be invisible because the strip began with a full display_width of
blank, so the wrapped-in region was black. Removing that blank (it was 10s of
dead panel per cycle) exposed the wrap. Cycles now end one display width
earlier, before any wrapped content appears, clamped for strips no wider than
the display so they don't complete instantly and spin the recompose loop.

Verified on hardware: a 3936px strip now completes at 68.5s, exactly
(3936 - 512) / 50.

Second, auto_trim=False also skipped the width budget, which is an unrelated
concern — turning off margin cropping should not let one plugin hold the panel
for minutes. Seen in the field: the F1 scoreboard contributed 116 images and
14,848px untouched, giving a 33,821px cycle (11 minutes of content). The budget
now applies regardless of trimming; with it restored that cycle is 6,362px.

Third, the budget accounted for row gaps using the flat intra_plugin_gap while
the compositor had moved to measured separation, so it under-counted by up to
(min_content_separation - intra_plugin_gap) per row and a many-row plugin
overran its cap. Both now use the same separation_gap() rule, and a test
asserts the composed block fits the budget end to end rather than trusting the
two paths to agree.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Fix IndexError in find_blank_cut when the cut lands on the image edge

A cut position after the last column is legitimate — _crop_to_budget asks for
min(start + budget, img.width), which equals the width whenever the remaining
strip is shorter than the budget. find_blank_cut clamped target to width but
then walked leftwards starting at target itself, so ink[width] raised
IndexError.

Caught on hardware: it killed the ledmatrix-stocks fetch, and because
_fetch_plugin_content catches broadly that surfaced as the plugin silently
contributing nothing for the cycle.

Only reachable on the second or later pass of the rotating window over a single
oversized image, which is why the existing tests missed it — they all exercised
the first pass, where start is 0 and start + budget is comfortably inside the
image. Added TestRotationAcrossMultipleCycles, which walks the window round
several times and asserts content is never lost, plus direct coverage of
find_blank_cut at and beyond the image edge.

Both bounds now stop at width - 1 so neither direction can index past the end.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Only cut oversized segments at real gaps between items

The width-budget crop snapped to the nearest blank column, and in rendered text
the gap between two characters is a single column. So a cut routinely landed
inside a word: the cycle showed "Wednesda" and the orphaned "y" turned up as a
lone floating letter in the next cycle, positioned after whatever plugin
happened to precede it.

Measured on the clock-simple segment to confirm: its blank runs are
[1, 1, 1, 1, 1, 8, 8] — five single-column letter gaps, every one of which
find_blank_cut would happily have chosen.

Cuts now only land in a run of at least min_cut_gap blank columns (default 6),
which excludes letter spacing while still finding the gaps plugins put between
items (the stocks ticker uses 32px, baseball 48px). Where no boundary falls
inside the budget the cut waits for the next one and overruns, because
splitting an item is worse than a slightly long segment.

Continuous content is treated differently on purpose: an image with no internal
gaps is a map or a chart, where any column is as good as another, so it is still
cut to the budget exactly. The gap rule protects discrete items; letting a solid
image escape the cap in its name would be wrong.

blank_runs() is vectorised — 48ms for a 17,000px strip, against seconds for a
per-column Python loop.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Hold capture_mode for every plugin render, not just narrowed ones

The native content path only entered capture_mode when it was also narrowing
the canvas, so at full width — which is every plugin without a vegas_width_pct
override, i.e. most of them — a plugin calling update_display() while building
its Vegas content wrote straight to the hardware. That is a visible flash
mid-scroll, and it lines up with the flash reported at cycle transitions, when
several plugins are fetched back to back.

Suppression is now unconditional; the narrowing context stays separate because
it is already a no-op at full width.

Both contexts are reached through helpers that degrade to nullcontext when the
display manager lacks them. That matters more than it looks: the adapter's
handlers are deliberately broad, so an AttributeError from a missing context
does not surface as an error — it surfaces as the plugin contributing nothing.
Making the call unconditional without this turned 44 tests red for exactly that
reason, all of them reporting lost content rather than the real cause.

The test double now provides capture_mode and render_size too, so tests
exercise the real contexts instead of silently taking the degraded path.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Vegas mode: one continuous strip instead of swapping cycles

A cycle used to be a discrete strip that got replaced: motion stopped, every
pixel was substituted at once, and the next group started with the viewport
already full. That is the freeze, the flash and the jump.

The strip is now extended rather than replaced. ScrollHelper gains
append_content(), which adds items on the right without touching
scroll_position or total_distance_scrolled, so motion continues and the next
group simply arrives from the right. Because completion is measured against
total_scroll_width, extending also defers completion — there is no longer a
cycle boundary to see.

drop_scrolled_prefix() reclaims what has gone past, keeping the strip bounded
however long Vegas runs (observed 5,000-11,000px against an unbounded strip
otherwise). It shifts total_distance_scrolled and total_scroll_width together so
the completion arithmetic is unchanged, and refuses to run while the viewport is
wrapping: wrapping reads the head of the strip into the right of the frame, so
trimming the head there would visibly change the picture. A test caught that.

Groups are prepared off the render thread. The constraint is that the canvas and
the matrix proxy are process-wide mutable state, so narrowing or capturing
through them from another thread would corrupt the frame the render loop is
pushing. get_content() therefore takes offscreen_only: the background thread uses
only paths that avoid the canvas, and anything needing it is marked and picked up
on the render thread. That puts the expensive work (native renders of leaderboard
and baseball cards, seconds each) in the background and leaves the cheap work
(display capture, 40-600ms) in the foreground.

DisplayManager's capture flag is now thread-local. As a shared flag, a background
capture would have suppressed the render loop's own frame pushes for its
duration, freezing the panel precisely when the point was to avoid a freeze.

Canvas-bound plugins are drained one at a time rather than as a batch: six at
once held the render thread for 1.75s. Drains are also spaced by two seconds
while the lookahead is healthy, since taking them back to back turns one long
stall into a run of short ones. When the strip is genuinely running short the
throttle is ignored, because content matters more than smoothness there.

Measured on hardware: zero cycle-complete swaps, drains landing 2-4s apart,
lookahead holding at 1,200-3,500px, no errors.

Set continuous_scroll false to restore the swap behaviour; the old path is intact.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Pace the Vegas frame loop adaptively: 31.5 -> 78.7 fps

The loop slept a fixed frame_interval on top of however long the frame took, so
at a measured 31.6ms per frame a flat 8ms of that was pure idle — a quarter of
the budget spent not rendering. It now sleeps only the remainder of the budget.

Measured on hardware: 31.5 fps to 78.7 fps sustained, with CPU going *down* from
150% to 127%. Scroll speed is unchanged at 49.9px/s against a configured 50,
because motion is derived from elapsed time rather than frame count — this buys
smoothness, not speed.

Worth recording what the bottleneck was not: the per-frame render path measures
0.34ms in total (0.18ms for the numpy slice, 0.17ms for the dirty-tracking
digest), which is a theoretical 2900 fps. Optimising any of that would have been
wasted effort. The frame was idle, not busy.

Also nices the prefetch thread. Its work is PIL and numpy that releases the GIL,
so the scheduler can act on the priority, and without it the prefetch competes
for the same cores as the render loop.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Sub-pixel scrolling: motion at the frame rate, not the pixel rate

With integer positioning the number of distinct frames per second equals the
scroll speed in px/s, however fast the loop renders. Measured at 50px/s and
78.7fps, 36% of frames were byte-identical: the extra frames cost work and
bought no motion, and what was left was 50 discrete 1px steps a second.

Two things were wrong with the pre-existing sub-pixel support. get_visible_portion
never consulted sub_pixel_scrolling — it always took the integer path, so the flag
and _get_visible_portion_subpixel were dead code. And that implementation needed
scipy.ndimage.shift, which is not installed on the target devices (HAS_SCIPY is
False there), so it would not have interpolated even if reached. Verified both:
positions 1000.0 and 1000.5 produced identical frames either way.

Blending is now wired up and implemented with numpy. Two details make it
affordable: slice cached_array directly instead of building two PIL images only
to convert them straight back (the naive version measured 15x the integer path),
and use fixed-point uint16 multiply-add rather than float32, which suits the Pi's
cores and gives finer weighting than the panel can resolve. Result 0.939ms
against 0.237ms — 0.70ms added per frame, a 1065fps ceiling.

Measured on hardware: 81.2 fps with blending on, against 78.7 with it off, so no
cost within noise — and every frame is now a distinct position rather than one in
three being a repeat.

The trade is a slight horizontal softening of text, since each frame blends two
positions. Set smooth_scroll false for maximum crispness.

Also benchmarked and cleared as non-issues: extending the strip costs 9.4ms on an
11,000px strip and trimming 2.5ms, both under one frame at this rate.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Add overflow handling: keep ordered content whole instead of rotating a window

The width budget split any oversized plugin by advancing a window each cycle.
That is right for interchangeable items — news headlines, odds, stock prices —
but wrong for ordered content: a league table showed ranks 1-6, then resumed at
7 two rotations later, which reads as out of order and out of context. Nobody
needs rank 23 in a ticker; they need the top of the table, every time.

overflow_mode chooses between them:

  rotate   — advance a window each cycle so everything is seen eventually
             (unchanged default)
  truncate — always show the start and drop the rest, keeping ordered content
             coherent. Records no window state, so every pass starts at the top.

Per-plugin vegas_overflow overrides the global setting, since one install has
both kinds of plugin. Also adds per-plugin vegas_max_width_screens, so content
that must stay whole can be given more room — or uncapped with 0 — without
lifting the cap on every ticker.

Applied on the test rig: f1-scoreboard and ledmatrix-leaderboard set to
truncate, and baseball given 4.5 screens because it was showing 8 of 9 games
when the whole slate needed only a little more room. Verified: F1 now reports
"the first 10 of 116 ... the rest are not shown", baseball has dropped out of
the budget log entirely, and stocks, odds-ticker and stock-news still rotate.

Also corrects the crop log, which claimed "window advances next cycle"
unconditionally and so misreported truncated crops. A test now pins the
behaviour behind the message: truncate must leave no offset recorded.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Stop Vegas mode showing last night's games as if they were live

A game that was live in the evening was still being drawn as live the next
morning. Two faults combined to freeze plugin visuals indefinitely.

PR #291 added a call to plugin_adapter.invalidate_plugin_scroll_cache() so
a plugin's own cached scroll image would be rebuilt from fresh data. That
method was never implemented. hot_swap_content() wraps the call in a broad
except, so every hot swap has raised AttributeError and been swallowed
silently ever since — which is why the visuals it was meant to keep fresh
never were.

Continuous scrolling then removed the only path that reached it at all:
should_recompose() and hot_swap_content() are called from the
non-continuous branch of run_frame(), and continuous_scroll defaults to
True. So on a default install the pending-update flags were set by the
update tick, never consumed, and grew without bound.

Together these froze content completely, because refetching is not enough
on its own: the sports plugins' get_vegas_content() regenerates only "if
the cache is empty", so take_next_group() kept receiving the same picture
however often it asked.

Fixed by:

- Implementing invalidate_plugin_scroll_cache(). It covers both layouts —
  a helper directly on the plugin (stocks, news, odds-ticker) and one
  owned by a scroll-display manager (the sports scoreboards, which is the
  shape that produced this bug) — and clears cached_image and
  cached_array together, since the array is the image's numpy mirror.

- Adding StreamManager.invalidate_pending_updates() and calling it from
  the continuous branch. It only drops the caches; the plugin recomposes
  when it next comes round in the rotation. process_updates() is wrong
  here: it refetches synchronously and merges into the active buffer that
  continuous mode bypasses, and hot_swap_content() rebuilds and
  repositions the whole strip, which is the freeze-and-jump this mode
  exists to avoid.

Tests assert the fix rather than the implementation: 14 of the 17 new
tests fail without it. Includes the wiring itself, since the regression
was a call that was simply absent, and a check that the scroll position is
untouched so this cannot regress into the swap's visible jump.

All Vegas suites pass (355 tests). test_display_controller_vegas_tick.py
still cannot be collected off-device for want of rgbmatrix, identically
with and without this change.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* Fix two CodeRabbit-flagged test assertions in vegas density tests

test_prepared_group_is_used_without_refetching had a tautological final
assertion; now checks stream.calls directly. test_no_partial_letter_at_either_edge
required both crop edges to be blank, but the left edge here is always the
crop's start position with no lead-in gap in word_strip, so it legitimately
carries ink — only the right edge is an actual cut and needs the check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 09:40:38 -04:00
ChuckandClaude Opus 5 e2acbfb566 fix(web): don't validate double-sided settings when the feature is disabled (#422)
* fix(web): don't validate double-sided settings when the feature is disabled

Saving anything on the Display tab failed with a 400 when double-sided
mode was off:

    Double-sided copies (2) must divide chain length (3) evenly

The Display form posts every field in one request, including
double_sided_copies (default 2) and double_sided_axis, whether or not
the Enabled checkbox is ticked. The server block was gated only on "is
any double-sided field present in the payload" — it wrote
ds_config['enabled'] but never read it. A user with chain_length: 3 and
the untouched default copies: 2 was locked out of saving any display
setting at all: brightness, GPIO slowdown, Vegas, sync.

Gate the checks on the enabled flag:

- Divisibility against chain_length/parallel is hardware-relational and
  only runs when the feature is on.
- Structural checks (copies parses as an int in 2..8, axis in the
  whitelist) still 400 when enabled; when disabled they drop the value
  and leave the stored one untouched rather than rejecting the save.

The runtime already gated correctly (_resolve_double_sided returns None
when disabled), so nothing there changes.

Also in the Display tab:

- Hide Copies / Split Axis until Enabled is ticked, mirroring the Vegas
  Scroll pattern. Hidden rather than disabled, so the fields keep
  submitting and the server still sees an 'off' state to persist.
- Fix the save toast: the form's handler read xhr.responseJSON, a jQuery
  property that doesn't exist on a native XMLHttpRequest, so it was
  always undefined and every save reported a green "Display settings
  saved" — even the 400s. Parse responseText and use the real status.

Tests: two existing double-sided tests asserted 200 on payloads that the
divisibility check (added later, in #373) turns into 400s; the first now
supplies matching hardware values and the second passes as written now
that a disabled save skips the check. Added coverage for the reported
regression, for bad values while disabled, and for the check still
firing when enabled.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016aUvzTXpEYhNEWYQvFqQU9

* fix(web): treat only 2xx as a successful display save

Review feedback on #422.

- showDisplaySaveResult tested `xhr.status >= 400` for failure, so a
  network error — which reports status 0 — was waved through as
  "Display settings saved". That's the same class of false-success bug
  this branch set out to fix. Test the 2xx range instead, and let a
  response body refine a successful verdict without overturning a
  failed one.
- Annotate the _copies_fits_hardware helper, matching the annotated
  helpers already in api_v3.py.
- Cover the vertical divisibility branch: chain_length 2 would divide
  evenly, so only parallel 3 can produce the rejection, which pins the
  branch to the right hardware dimension.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016aUvzTXpEYhNEWYQvFqQU9

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-28 16:54:42 -04:00
ChuckandClaude Opus 4.8 3872a68ff7 Surface plugin update availability on the plugins page (#421)
* Surface plugin update availability on the plugins page

The plugin manager page already showed each installed plugin's version and
had an Update button, but nothing told users an update actually existed —
they had to guess. The store manager already compares the installed
manifest version against the registry's latest_version for its reinstall
decision; this surfaces that same signal in the UI.

- api_v3 /plugins/installed now returns `latest_version` (from the registry
  cache, no extra network call) and an `update_available` flag computed by a
  new semver-aware helper `_is_plugin_update_available`. A locally modified
  plugin whose version is ahead of the registry is not flagged.
- The installed-plugin card shows "vX.Y.Z available" next to the installed
  version, and the Update button becomes emphasized ("Update to vX.Y.Z" with
  a gentle pulse) when a newer version is published — mirroring the app's own
  update banner styling.
- Added tests for the helper and the endpoint fields.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012EW8yDbsk8EceDpq6Hjgqj

* Catch InvalidVersion specifically in plugin update comparison

Address review feedback: the version comparison caught a blind `except
Exception` (ruff BLE001). Split the two failure modes and catch each
specifically — ImportError for a missing `packaging` (a core dependency)
and InvalidVersion for an unparseable version string — while preserving
the existing "surface the mismatch" fallback behavior. Added a test for
the unparseable-version path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012EW8yDbsk8EceDpq6Hjgqj

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 08:06:36 -04:00
ChuckandClaude Sonnet 5 989162d28f fix(web): custom-feed logo upload uses the wrong request/response contract (#420)
Every custom-feed logo upload has been failing: handleCustomFeedLogoUpload
posts the file under field name "file" and reads the response from
data.data.files, but the backend endpoint it calls
(api_v3.upload_plugin_asset, /api/v3/plugins/assets/upload) requires the
field name "files" (checks 'files' not in request.files, 400s "No files
provided" otherwise) and returns the result in a top-level "uploaded_files"
key - there is no nested "data" wrapper in the response at all. Confirmed
by reading the endpoint directly, and cross-checked against
file-upload-single.js, a sibling widget that uses the correct contract
against the same endpoint.

- formData.append('file', file) -> formData.append('files', file)
- data.data.files / data.data.files[0] -> data.uploaded_files /
  data.uploaded_files[0]

No other call sites in this file used the stale contract (grepped for both
patterns after the fix - zero remaining). The response entries' 'path' and
'id' fields (both read further down in the same handler) are unaffected -
only the wrapper shape was wrong.

Found incidentally while re-verifying a CodeRabbit review on an unrelated
PR (#417) that had deleted a differently-named dead file
(custom-feeds-helpers.js) with the same bug; this widget (custom-feeds.js)
is the live code path and was never touched by that PR.

Validation: brace/paren balance check; explicit assertions that the old
field name and response shape no longer appear anywhere in the file. No
Python changed, no existing tests cover this endpoint's client flow.


Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

Signed-off-by: Chuck <33324927+ChuckBuilds@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-18 11:07:36 -04:00
ChuckandClaude Fable 5 cdf03fb107 Add skin system: user-installable visual overlays for sports scoreboards (#419)
* Add skin system: user-installable visual overlays for sports scoreboards

Skins restyle a scoreboard's live/recent/upcoming rendering while the
host plugin keeps doing data fetching, scheduling, caching, live
priority, and vegas mode — the anti-fork alternative for users who only
want a different layout.

- src/skin_system/: ScoreboardSkin API, SkinContext (canvas + adaptive
  layout + logo/font helpers), discovery/loading runtime with API major
  version gating and per-skin module namespacing
- src/base_classes/sports.py: _render_game() seam at the three
  _draw_scorebug_layout call sites; skin-first with built-in fallback,
  3-strikes session disable, slow-render warning; per-mode skin config
- scripts/validate_skin.py: headless multi-mode/multi-size validator
  with bundled per-sport fixtures (no hardware or network needed)
- skins/example-classic-baseball/: working reference skin
- Web UI: served-schema Visual Skin dropdown (validation never
  enum-restricted, so uninstalled skins can't invalidate configs) and
  GET /api/v3/skins
- Store: registry entries with type "skin" install to skins/
- docs/SKIN_SYSTEM.md (architecture), docs/CREATING_SKINS.md (author
  guide incl. Claude Code prompt), view-model contract locked by tests

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LrCusPasy1qeUN5anK3aA1

* Address review feedback on skin system

- skin_runtime: cache the entry module so the 2nd/3rd load of the same
  skin (live/recent/upcoming hosts) doesn't re-execute it with unbound
  sibling aliases; rebind cached sibling modules to their bare names
  around entry import and restore prior bindings after; include per-
  manifest mtimes in the discovery cache fingerprint so in-place skin
  updates are picked up
- sports.py: count render_skin_card exceptions toward the 3-strike
  session disable
- store_manager: validate skin ids (pattern + resolved-path containment
  in skins/), reject registry/manifest id mismatches, and stage+validate
  downloads in a temp sibling before replacing an existing skin
- schema_manager: leave the schema untouched when the configured skin
  value is a per-mode mapping (a string dropdown could overwrite it)
- validate_skin.py: reject non-positive sizes and non-object --options
  at parse time; support --output-dir outside the repo; type annotations
- example skin: validate accent_color once at load with logged fallback
- fixtures: pregame 0-0 scores in football/hockey upcoming fixtures
- api /skins: rely on the self-invalidating discovery cache instead of
  force_refresh
- docs: valid JSON manifest example, load_logo caching semantics spelled
  out, language ids on fenced blocks
- tests: view-model contract test now exercises the real extractor;
  regression test for repeated same-skin loads with sibling modules

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LrCusPasy1qeUN5anK3aA1

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-18 11:07:08 -04:00
ChuckandClaude 6a9d8014e5 Tag plugin logs structurally and surface the active plugin in System Logs (#418)
- get_logger() now returns a PluginLoggerAdapter when given a plugin_id,
  so every plugin log call is stamped with plugin_id automatically instead
  of only calls that explicitly passed extra={'plugin_id': ...}. This makes
  the "[Plugin: x]" prefix reliable in the journalctl-backed log stream.
- display_controller publishes the currently active mode/plugin to the
  shared cache whenever it changes, exposed via a new
  GET /api/v3/display/current-status endpoint.
- System Logs page: adds a "Now showing" banner backed by that endpoint, a
  plugin filter dropdown (populated from parsed log lines), a plugin badge
  per log entry, and fixes log parsing to handle the short-iso timestamp
  format journalctl actually returns (the old regex only matched syslog
  timestamps, so level/plugin extraction silently never ran).

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-18 10:59:36 -04:00
ChuckandClaude Sonnet 5 c90129285c Web UI: mobile navigation, guided onboarding, basic/advanced config tiering + performance (#417)
* chore(web): remove dead legacy client-side plugin-config generator (~2,300 lines)

Plugin config forms have been rendered server-side (plugin_config.html via
GET /partials/plugin-config/<id>) since the HTMX migration; the old
client-side generator survived as unreachable code. Verified dead by call
graph, not by naming: showPluginConfigModal and showGithubTokenInstructions
have zero callers anywhere in templates or JS, and everything removed here
is reachable only from those two roots.

Removed:
- plugins_manager.js: showPluginConfigModal, generatePluginConfigForm,
  generateFormFromSchema, generateFieldHtml, generateSimpleConfigForm,
  handlePluginConfigSubmit, the modal's JSON-editor view (initJsonEditor,
  switchPluginConfigView, syncFormToJson/JsonToForm, saveConfigFromJsonEditor,
  resetPluginConfigToDefaults, displayValidationErrors, closePluginConfigModal,
  savePluginConfiguration, currentPluginConfigState), their exclusive helpers
  (getSchemaPropertyType, escapeCssSelector, dotToNested, collectBooleanFields,
  normalizeFormDataForConfig, flattenConfig, loadCustomHtmlWidget), the
  orphaned-modal cleanup block, the modal's listener wiring, and the
  never-invoked showGithubTokenInstructions/closeInstructionsModal pair.
- plugins.html: the #plugin-config-modal markup those functions drove.
- base.html: the deprecated pluginConfigData() component and the
  window.PluginConfigHelpers shim (only ever called by pluginConfigData).

Deliberately kept, verified still live:
- renderArrayObjectItem, getSchemaProperty, escapeHtml/escapeAttribute
  (window-exposed for the top-level array-of-objects handlers the
  server-rendered form uses), toggleNestedSection, addKeyValuePair/
  addArrayObjectItem families, executePluginAction, and
  window.currentPluginConfig = null init (file-upload.js and
  executePluginAction read it, optional-chained).
- app()'s internal generateConfigForm/generateSimpleConfigForm methods in
  base.html: unreachable now but embedded in the live Alpine component;
  excising methods from a live object is deferred to keep this change
  zero-risk.

Validation: every deletion seam inspected line-by-line; Jinja parse of both
templates passes; repo-wide sweep confirms zero remaining references to any
deleted function or element id (deleted ranges contained no Jinja tags).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): mobile navigation drawer + responsive CSS gap fixes

Phones previously got the desktop layout squeezed: ~12 system tabs plus one
tab per installed plugin wrapped into many rows of small pill buttons, and
the header's settings-search and system-stats widgets were dropped entirely
(hidden below their breakpoints, never relocated).

- Off-canvas nav drawer below md: the existing nav markup (system tab row +
  #plugin-tabs-row, including dynamically injected plugin tabs) is wrapped in
  a #site-nav container that CSS repositions into a slide-in drawer on small
  screens. Same DOM nodes, same @click handlers, nothing duplicated. Tabs
  become full-width rows with 44px+ touch targets. A hamburger button
  (md:hidden) in the header and a backdrop toggle the new mobileNavOpen
  Alpine state (added to both app() definitions, mirroring activeTab).
  Clicking any tab, a search result, or the backdrop closes the drawer.
  At md+ hard CSS guards make all drawer styles inert - desktop renders
  exactly as before.
- Header widgets relocated, not hidden: placeHeaderWidgets() in app.js moves
  the #settings-search-wrap and #system-stats nodes (same elements, listeners
  intact - both are looked up by id from SSE/search code, so they must never
  be duplicated) into the drawer below md and back into the header above it,
  via a matchMedia listener.
- Fixed 13 breakpoint utility classes that templates referenced but app.css
  never defined (sm:block, sm:grid-cols-2, sm:text-sm, md:block, md:w-auto,
  lg:block, lg:flex, lg:w-64, xl:grid-cols-2/3, 2xl:grid-cols-2/3/4). This
  was a live bug: 'hidden sm:block' on the search box and 'hidden lg:flex'
  on the stats meant BOTH were invisible at every screen width. Audit method
  (repeatable): diff classes used in templates vs defined in app.css.
- Mobile modal sizing: one global rule caps .modal-content at 95vw/90vh with
  internal scroll below 640px - covers every modal without per-template
  changes.
- Horizontal-scroll affordance: pure-CSS edge-fade shadows on
  .overflow-x-auto containers (scrolling-shadows technique), plus larger
  in-table touch targets below md.

Validation: breakpoint used-vs-defined audit now returns zero gaps; Jinja
parse of base.html passes; all changes to desktop behavior are additive
(new utilities) or scoped inside max-width media queries.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): x-advanced schema flag groups plugin config fields under a collapsed Advanced Settings section

Plugin config pages show every schema property at equal visual priority,
which overwhelms first-time users. Plugin authors can now add
"x-advanced": true to any flat (non-object) property in config_schema.json
to move it into one collapsed "Advanced Settings (N)" section rendered after
the basic fields - progressive disclosure with zero loss of control.

Implementation: the main render loop in plugin_config.html splits ordered
properties into basic/advanced tiers; the advanced group reuses the exact
.nested-section/.nested-content/toggleSection() shell that nested object
sections already use, so the settings search's expand-on-match behavior
works on advanced fields with no JS changes. Object-type properties ignore
the flag (they already render as their own collapsible sections). No
backend change needed: jsonschema ignores unknown x-* keywords exactly as
it does for x-widget/x-propertyOrder.

Documented in docs/widget-guide.md alongside the other x-* extensions.

Validation (rendered with real Jinja, not just parsed):
- synthetic schema with 2 advanced fields: basic fields render before the
  section, advanced inside the collapsed shell, count badge correct,
  x-advanced on an object property correctly ignored
- schema without any x-advanced: output is identical to the pre-change
  template (whitespace-normalized diff against git HEAD's version)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): Display Settings basic/advanced split + live total-resolution readout

The Hardware Configuration card showed ~17 fields at equal priority; a new
user only needs 7 of them to get a correctly-sized, correctly-colored image
(rows, cols, chain_length, parallel, brightness, hardware_mapping,
led_rgb_sequence). The other 10 (multiplexing, panel_type, row_address_type,
gpio_slowdown, rp1_rio, scan_mode, pwm_bits, pwm_dither_bits,
pwm_lsb_nanoseconds, limit_refresh_rate_hz) now live in a collapsed
"Advanced Hardware Settings" section using the same nested-section shell as
plugin config forms, so toggleSection() and settings-search auto-expand work
unchanged. led_rgb_sequence moved up beside brightness/hardware_mapping
(2-col grid became 3-col). No field was removed or renamed; the form still
posts the same names to /api/v3/config/main.

Also adds a live "Your display: W x H pixels" readout under the four sizing
fields (width = cols x chain_length, height = rows x parallel - the exact
math the chain-length tooltip describes in prose), recomputed client-side on
every input event, no round-trip.

Deviation from plan, deliberate: disable_hardware_pulsing / inverse_colors /
show_refresh_rate stay in their separate "Display Options" card rather than
moving across cards - relocating fields between form sections risks
regressions for no decluttering gain in the card users complained about.

Validation (real Jinja render): all 17 hardware fields present exactly once,
basic fields render before the advanced section and the 10 advanced fields
inside it, div count balanced (71/71), readout + recompute script present.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): plugin install auto-enables + persistent restart nudge

Getting a plugin onto the display used to take three disconnected manual
steps: install from the store, flip its enable toggle, then restart the
display service - with no in-UI hint that steps 2 and 3 were needed (only
docs/GETTING_STARTED.md mentions it).

- installPlugin() now enables the plugin immediately on successful install
  (owner-confirmed behavior change: always auto-enable, no opt-out; users
  who don't want it running toggle it off as before), then shows a
  persistent toast ("... restart the display to show it") with an inline
  "Restart Now" button wired to the existing restartDisplay() - the same
  function the three existing Restart Display buttons call.
- notification.js: show() accepts optional { actionLabel, onAction } to
  render one inline action button per toast. Callbacks are stored per
  notification id and cleaned up on dismiss; a new triggerAction() public
  method runs the callback and dismisses. The global showNotification()
  shorthand now forwards a full options object as its second argument
  (legacy type-string calls unchanged).

Scope note: applies to the plugin store's install path (window.installPlugin).
The custom-registry install path keeps its existing behavior.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): dismissible Getting Started checklist on Overview

New users land on a dense multi-tab dashboard with no suggested order of
operations (the only guided flow is the WiFi captive portal). This adds a
non-gating checklist card at the top of Overview with five steps, each a
deep link that switches to the right tab (and closes the mobile nav drawer):

1. Set panel size            -> Display tab   (done: rows/cols/chain_length > 0)
2. Set timezone/location     -> General tab   (done: differs from template
                                               defaults America/New_York / Tampa)
3. Install a plugin          -> Plugins tab   (done: /api/v3/plugins/installed
                                               non-empty)
4. Enable a plugin           -> Plugins tab   (done: any installed plugin enabled)
5. Configure it              -> Plugins tab   (done: first enabled plugin has >=1
                                               saved value differing from its
                                               schema defaults)

Steps 1-2 are computed server-side in Jinja from main_config (already in the
partial's context); 3-5 client-side from existing endpoints. No new backend
state: dismissal persists in localStorage (mirroring the reconciliation
banner's sessionStorage pattern one section up); deep links use the same
_x_dataStack app-data access as settings-search.js. Disclosed heuristic
limit: values left at legitimate defaults (a user actually in Tampa) read
as "not done".

Validation: real Jinja render across 3 config variants confirms the
server-side done-flags flip correctly; div balance intact; /plugins/config
response shape (config dict directly in .data) verified against api_v3.py.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(display): drag-and-drop plugin rotation order for the primary display mode

The primary rotation's order was invisible and unconfigurable: modes are
registered in parallel-load COMPLETION order, so rotation order actually
varied between restarts. Only the niche Vegas Scroll mode had a working
order UI. This adds real, persisted ordering end to end:

Backend:
- config.template.json: new display.plugin_rotation_order (default [],
  fully backward compatible).
- display_controller.py: _apply_plugin_rotation_order() rebuilds
  available_modes grouped by plugin per the configured list (each plugin's
  modes keep their declared order; unlisted plugins follow in existing
  relative order; empty config = exact no-op). Applied at startup after
  parallel load and after live enable/disable reconcile (before the
  existing _resync_mode_index_after_change, which preserves the current
  mode). Mirrors vegas_mode get_ordered_plugins() semantics.
- api_v3.py save_main_config: accepts plugin_rotation_order as a JSON
  array (same parse/guard pattern as vegas_plugin_order).

Frontend:
- New shared widget static/v3/js/widgets/plugin-order-list.js: the Vegas
  section's drag-and-drop list factored out verbatim (native HTML5 drag
  events, saved-order-first rendering, hidden-input JSON sync),
  parameterized by container/order-input/optional exclude-checkbox/badge.
- display.html: Vegas section now calls the shared module; its ~130-line
  inline copy of the same logic is deleted.
- durations.html: new "Rotation Order" card above the durations grid using
  the same module, posting plugin_rotation_order with the existing form.

Deviation from plan, deliberate: durations stay as their own mode-keyed
grid rather than inline in the drag rows - verified display_durations keys
are MODE names (display_controller.py resolves duration per mode_key), not
plugin ids, and one plugin can own several modes, so the planned 1:1
inline pairing was wrong.

Validation: py_compile on both Python files; _apply_plugin_rotation_order
unit-tested standalone (configured order applied, empty-config no-op,
unknown ids skipped - 3/3); both templates render with balanced divs, the
hidden input carries the saved order, and the old inline implementation is
confirmed gone; config.template.json parses.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): serve the interface at / — /v3 kept as a legacy alias

The user-visible URL no longer carries the interface version: the pages
blueprint is now registered un-prefixed (primary) AND at /v3 (second
registration, name='pages_v3_legacy'), so:

- http://<device>/ serves the interface directly (the old @app.route('/')
  redirect is removed — the blueprint's own index takes its place)
- every existing /v3/... bookmark and all the hardcoded /v3/partials/...
  fetches in templates/JS keep working verbatim through the alias mount —
  zero template/JS churn, zero broken links
- url_for('pages_v3.*') resolves against the primary registration, so all
  server-side redirects (captive portal detection endpoints) now emit
  un-prefixed URLs
- the AP-mode captive-portal allowlist learned the un-prefixed page paths
  (/setup, /partials/, /settings/, /plugin-ui/) so setup-mode requests
  don't redirect-loop
- /api/v3 and the templates/v3, static/v3 directories are deliberately
  untouched (internal, invisible to users; owner-confirmed scope)

Validation: dual registration mechanics tested against real Flask (test
client): /, /v3, /v3/ redirect, partials and /setup reachable on both
mounts, url_for yields un-prefixed paths; py_compile passes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* perf(web): stream the preview PNG raw instead of PIL decode + re-encode

display_preview_generator() opened each changed snapshot with PIL and
re-encoded it to PNG just to base64 it — but /tmp/led_matrix_preview.png
already IS a PNG, written atomically by the display service (tmp file +
os.replace in display_manager.py), so a partially-written file can never be
observed. Read the bytes and base64 them directly: identical payload
(front-end consumes data:image/png;base64 — verified in base.html), one
full image decode+encode per frame less on the same Pi that's driving the
matrix. The existing mtime skip and viewer-marker throttling are unchanged
(they already covered the "skip unchanged frames" concern).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* perf(web): gzip response compression via flask-compress

The interface ships a ~5,000-line HTML shell and >20k lines of JS
uncompressed; on phone/WiFi that dominates load time. Flask-Compress
gzips/brotlis compressible responses transparently.

- Optional dependency, same graceful pattern as flask-limiter: missing
  package = uncompressed responses, no crash.
- SSE safety verified empirically against the real package (1.24): an
  actual streamed text/event-stream response comes back with no
  Content-Encoding while a large HTML response gzips — the display
  preview / stats / logs streams are unaffected.
- Added flask-compress>=1.14 to web_interface/requirements.txt.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* perf(web): gate verbose console logging behind the existing pluginDebug switch

plugins_manager.js, base.html's inline scripts, and app.js emitted 198
console.log calls in production - including per-interaction [DEBUG] dumps -
costing main-thread time and drowning real errors in noise.

- New window.debugLog() gate defined in base.html's first inline script
  (before any other script runs): forwards to console.log only when
  localStorage.pluginDebug === 'true' - the SAME switch plugins_manager.js
  already used for its _PLUGIN_DEBUG_EARLY logs, so existing debug workflow
  docs stay valid. Exposed as window.LEDMATRIX_DEBUG for other scripts.
- Mechanically rewrote console.log( -> debugLog( in plugins_manager.js
  (127), base.html (64), app.js (7). Verified no occurrences lived inside
  string literals before rewriting; console.error/console.warn untouched.
- app.js's no-Alpine showNotification fallback restored to console.info -
  it's a user-facing last resort, not debug output.

Both load paths are safe: the gate is the first inline <script> in <head>,
and every rewritten file loads deferred after it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* perf(web): vendor CDN assets locally — LAN-speed loads, fully offline-capable

Font Awesome, CodeMirror, and htmx were fetched from cdnjs/unpkg on every
fresh page load, adding third-party round-trips on a device that often
lives on a local network (and htmx was local-only in AP mode, meaning two
different loading behaviors to reason about).

- Vendored pinned copies under static/v3/vendor/: Font Awesome 6.0.0
  (css/all.min.css + the 8 webfonts it references relatively) and
  CodeMirror 5.65.2 (core, javascript mode, closebrackets/matchbrackets
  addons, base + monokai css) - ~1.1 MB total, exact versions the CDN tags
  pinned.
- htmx + sse + json-enc extensions now load from the existing local copies
  (verified 1.9.10, matching the CDN pin) on EVERY network, not just AP
  mode; the pinned CDN copies remain as a one-shot rescue fallback,
  mirroring the pattern Alpine already used. The convoluted isAPMode
  source-flipping logic collapses away.
- Dropped the CDN preconnect/dns-prefetch hints (no longer on the critical
  path).
- Fixed a latent bug while relinking CodeMirror: the loader requested
  mode/json/json.min.js, which does not exist on cdnjs (HTTP 404 verified)
  - it 404'd on every JSON-editor open. JSON highlighting comes from the
  javascript mode; the phantom entry is removed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* perf(web): extract 3,850 lines of inline JS from base.html to cacheable static files

base.html shipped ~4,200 lines of inline JavaScript inside the HTML
document, re-downloaded and re-parsed on every page load (gzip helps the
transfer, but inline scripts can never be browser-cached). The four
largest blocks - none containing any Jinja syntax, verified by scanning
every inline block for {{ }} / {% %} - now live as static files served
with the app's existing mtime-versioned immutable caching:

- js/htmx-config.js (246 lines): HTMX swap/script-execution config,
  toggleSection helpers
- js/app-early.js (346 lines): early helpers + the app() stub that must
  precede Alpine init
- js/app-shell.js (2,997 lines): SSE wiring + the full Alpine app()
  implementation and tab logic
- js/custom-feeds-helpers.js (262 lines): custom-feeds table helpers

Each replacement <script src> is CLASSIC (no defer/async) at the exact
position of the inline block it replaces - identical execution timing and
DOM visibility to inline scripts, so relative ordering with the deferred
scripts and with each other is unchanged. base.html drops from ~4,940 to
1,079 lines.

Validation: extraction proven lossless by programmatically reassembling
the four files back into the template and comparing against git HEAD -
byte-for-byte identical. Jinja parse passes; script open/close tags
balanced (53/53, after excluding a literal "<script>" inside an HTML
comment).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): size the live preview from the PNG's real dimensions, not config

The initial SSE render sized the preview image (and both overlay canvases)
from the server-reported config dimensions (cols x chain_length,
rows x parallel), while the scale slider's re-render path sized from
img.naturalWidth/naturalHeight. Whenever the snapshot PNG's actual size
disagrees with the config (stale config, display service not restarted
after a hardware change), the initial render stretched the image at a
fractional ratio - blurry despite image-rendering: pixelated - and
touching the scale slider "fixed" it. Reported live on the devpi test rig.

Both paths now size from the loaded image's natural dimensions inside
img.onload (which also removes a transient wrong-size flash between
src assignment and load). The meta label now reports the true snapshot
size. The preview card also gets overflow-x-auto so on narrow screens a
wide preview scrolls at its exact pixel-perfect size instead of being
squeezed into the viewport (fractional downscaling of pixel art also
reads as blur).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): fold Display Options into the advanced dropdown; Vegas above Double-Sided

Owner-requested layout refinement of the Display Settings tab:

- The "Display Options" card (disable_hardware_pulsing, inverse_colors,
  show_refresh_rate, use_short_date_format, Dynamic Duration) moves inside
  the collapsed advanced section, now titled "Advanced Hardware & Display
  Options (15)". Hidden form fields still submit with the form, and
  settings search still auto-expands the section on match, so nothing is
  lost - the tab just leads with the essentials.
- The "Vegas Scroll Mode" section moves above "Double-Sided Display".
  New section order: Hardware (+ advanced dropdown) > Vegas Scroll >
  Double-Sided > Multi-Display Sync.

Validation (real Jinja render): all 23 field names present exactly once,
divs balanced (70/70), the five Display Options fields render inside the
advanced section's bounds, and section markers confirm the new order.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): Getting Started items are manually checkable; card auto-hides when complete

Two gaps reported from live testing on the devpi rig:

1. The timezone/location step never showed done for a user whose real
   timezone IS the shipped default (America/New_York) - the heuristic can
   only detect difference-from-default, not "user saved this". Clicking an
   item's checkbox now toggles it done manually (persisted per browser in
   localStorage), so any heuristic false-negative is one tap to clear.
   Clicking the item text still deep-links to its tab.
2. The card now hides itself automatically once every step is done
   (auto-detected or manually checked) - previously it stayed until the X
   was clicked.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* chore(web): CI cleanup — declare debugLog global, fix entity-unescape order, drop v3 from UI branding

- Add debugLog to the /* global */ headers of the six JS files that call it
  (defined in base.html's first inline script) — resolves the wall of
  "'debugLog' is not defined" ESLint errors failing the Codacy check.
- Fix the two js/double-escaping CodeQL alerts in app-shell.js: the
  entity-unescape chains decoded &amp; before &lt;/&gt;, so a value
  containing a pre-escaped "&amp;lt;" wrongly double-decoded to "<".
  &amp; now decodes last (standard order). Pre-existing bug, made visible
  when the inline scripts moved into scannable .js files.
- Page title / header drop the "- v3" suffix, matching the de-versioned
  user-facing URL.

The remaining 7 CodeQL alerts are pre-existing patterns newly visible to
scanning (CodeQL doesn't see inline template JS): 4 github.com/htmx.org
URL-substring checks (the htmx ones match error-message text, not URLs —
false positives in context) and 1 innerHTML XSS-through-DOM in the GitHub
install flow. Triage/fix deferred to a focused follow-up rather than
expanding this PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): add the missing nav tab for the Rotation & Durations page

The durations partial (/partials/durations) has existed as a route with no
nav tab and no content panel referencing it - an orphaned page. That made
the new rotation-order UI unreachable through the interface (caught by the
owner testing on the rig; my endpoint-level tests fetched the partial by
URL and never noticed the missing entry point).

- New "Rotation" tab (fa-rotate icon, verified present in the vendored
  FA 6.0.0 css) between Display and Backup & Restore, wired exactly like
  the other tabs (#durations-content + hx-get + loadtab; loadTabContent()
  is fully generic, so no JS changes needed).
- Page heading updated from "Display Durations" to "Rotation & Durations"
  to match its content since the rotation-order card landed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): Rotation & Durations page lists every enabled plugin's screens

The durations grid looped over display.display_durations, which nothing has
ever populated (verified {} on a real production install) - so the page
rendered no duration fields at all. Worse, its inputs posted bare mode
names, which save_main_config's endswith('_duration') filter silently
dropped: the page was broken in both directions, unnoticed because it was
also unreachable (previous commit).

- pages_v3._load_durations_partial now builds one entry per display mode of
  every ENABLED plugin via plugin_manager.get_plugin_display_modes()
  (falling back to the plugin id), overlaid with saved values, defaulting
  to the display controller's 30s. Grouped per plugin, sorted by name.
  Saved keys not owned by any enabled plugin stay visible under "Other
  saved entries" instead of vanishing.
- durations.html renders the grouped inputs, named duration__<mode_key>
  (mode keys are arbitrary, so they can't use the *_duration suffix
  convention), with an explanatory empty state when no plugins are enabled.
- api_v3.save_main_config accepts the new duration__<mode> fields and
  writes them into display.display_durations under the bare mode key -
  exactly what the display controller reads
  (display_durations.get(mode_key, 30)).

Validation: py_compile both blueprints; Jinja render with 3 groups asserts
grouped inputs, saved-value overlay, stale-entry group, empty state, and
div balance.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): restart-pending banner, unsaved-changes guard, installable web app

Three usability improvements from live testing feedback:

- Restart-pending banner: every successful POST to /api/v3/config/main
  (display hardware, rotation/durations, general settings) now raises a
  persistent banner - "Configuration saved, restart the display to apply" -
  with a Restart Now button that posts restart_display_service directly.
  Backed by sessionStorage so it survives tab switches and reloads until
  restarted or dismissed. Plugin config saves are deliberately excluded:
  they apply live via the display process's config watcher.
- Unsaved-changes guard: plugin config panels are Alpine x-if templates,
  so navigating away destroys the panel and revisiting re-fetches it -
  edits were silently discarded. Forms now mark themselves dirty on input
  (cleared on successful submit), a capture-phase click handler confirms
  before a lossy tab switch, and beforeunload guards full page unloads.
  System tabs (x-show, persistent DOM) are exempt - no false prompts.
- Installable web app: manifest.json (standalone display, dark theme) +
  generated LED-grid icons (192/512 maskable + 180 apple-touch), linked
  from base.html. "Add to Home Screen" now yields an app-like fullscreen
  experience; no service worker, so zero behavioral risk.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* test(web): smoke tests + static-analysis audits for the web UI

Guardrails so this branch's fix classes can't regress silently:

- test_web_smoke.py (24 tests): boots the pages blueprint with the same
  dual registration app.py uses and asserts every page/partial returns 200
  with its load-bearing markers (nav wiring, getting-started card, advanced
  section, rotation order card, per-mode duration inputs), the /v3 legacy
  alias serves everything, all critical static assets (incl. vendored
  fontawesome/codemirror, PWA manifest/icons) are served, durations group
  per plugin with the leftover bucket, and the advanced-hardware section
  really contains the tuning fields. Would have caught this session's
  unreachable-durations-page and orphaned-tab bugs instantly.
- test_web_static_audit.py (3 tests): (1) every responsive utility class
  referenced in templates is actually defined in app.css - the
  silently-no-op class bug that left the header search box invisible at
  every width; (2) every url_for('static', ...) reference points to a real
  file; (3) any JS file calling the debugLog global declares it in a
  /* global */ header.

All 40 web tests pass (24 + 3 new, 13 existing) under pytest + Flask.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): floating live preview + per-plugin "Preview on display", drawer a11y

Preview-while-configuring:
- Floating mini preview (fixed, bottom-right) available on every tab except
  Overview, fed from the same SSE display stream by updateDisplayPreview -
  no new connections. Collapses to a round toggle button; open/closed state
  persists in localStorage; hides on Overview where the full preview lives.
- "Preview on display" button on every plugin config page header: runs that
  plugin on the real display for 60 seconds via the existing
  /display/on-demand/start API and opens the floating preview, closing the
  configure -> see-the-result loop.

Drawer/nav accessibility:
- aria-current="page" tracks the active tab (system + dynamic plugin tabs,
  matched via their Alpine @click expression), updated from the activeTab
  watcher so search deep-links and checklist navigation are covered too.
- Escape closes the mobile drawer and returns focus to the hamburger;
  opening the drawer moves focus to its first tab.

Validation: all 40 web tests pass; Jinja parse + div balance on both
touched templates.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): resolve Codacy findings — DOM building over innerHTML, Map callbacks, misc lint

Verified each of the 27 reported findings against current code; all fixed
except one rule class skipped with reason below.

- app-early.js: plugin tab buttons are now built with createElement/
  createTextNode instead of innerHTML template strings (icon class and name
  come from plugin manifests - semi-trusted input; the old code escaped the
  name but interpolated the icon class). Both construction sites. Also the
  forEach arrow no longer returns tab.remove()'s value.
- plugin-order-list.js: rows, empty state, and error state all built with
  DOM APIs - the file no longer contains innerHTML at all (the now-unneeded
  escapeHtml/escapeAttr helpers are removed); MODE_LABELS is a Map so the
  vegas-mode lookup can't hit prototype properties.
- notification.js: actionCallbacks is a Map (get/set/delete) instead of a
  plain object - resolves the object-injection-sink and dynamic-delete
  findings; triggerAction also type-checks the callback.
- htmx-config.js: unused catch binding dropped; var -> const in the
  afterSettle handler; the swapped-<script> re-execution reads/writes
  textContent instead of innerHTML; the diagnostic form payload uses a
  null-prototype object so a field named __proto__ can't pollute.
- custom-feeds-helpers.js DELETED (with its script tag): all three of its
  functions (addCustomFeedRow, removeCustomFeedRow,
  handleCustomFeedLogoUpload) are shadowed by the deferred
  widgets/custom-feeds.js window assignments, which always win at call time
  - the copies were dead even when they lived inline in base.html. This
  also resolves the unused-function and unused-variable findings there.

Skipped: 4x "Non-serializable expression must be wrapped with $(...)" in
app-early.js - that rule targets code crossing a browser-automation
serialization boundary (e.g. page.evaluate); these are ordinary arrow
functions in plain browser code with no such boundary.

Validation: all 40 web tests pass (incl. the static-asset reference audit,
which confirms no template still points at the deleted file); Jinja parse OK.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web): update flow installs changed Python dependencies automatically

The in-app updater (Update Now banner -> git_pull action) stashed, pulled,
and purged plugins - but never touched Python dependencies. Any release
adding a package (e.g. this branch's flask-compress, which lives in
web_interface/requirements.txt) silently required an SSH session and a
manual pip install that most users will never do.

- git_pull now records HEAD before pulling; after a successful pull it
  diffs old..new and, if requirements.txt or web_interface/requirements.txt
  changed, installs exactly those via _pip_install_requirements - the same
  vetted root-visible sudo path the Tools-tab buttons use (with its
  existing graceful fallback when the sudo wrapper isn't configured).
  Results are appended to the update toast; a failure points the user at
  the Tools-tab button instead of failing the whole update.
- install_base_requirements (Tools tab) now also installs
  web_interface/requirements.txt - previously it only covered the root
  file, so web-only dependencies were unreachable from the UI entirely.

No install happens when the pull was already-up-to-date or when no
requirements file changed, so routine updates stay fast.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): address CodeRabbit review — validation, a11y, perf, and privacy fixes

Verified each finding against current code. Fixed:

- api_v3: plugin_rotation_order is now strictly validated (JSON list of
  strings, 400 with a descriptive message otherwise) and popped from the
  payload before any further handling.
- display_controller: _apply_plugin_rotation_order defensively ignores a
  non-list value (keeps the existing rotation, logs a warning) and drops
  non-string entries; new logs carry the [DisplayController] prefix.
  Unit-tested both defensive paths.
- app.py: snapshot-read handler narrowed to OSError with debug logging;
  flask-compress ImportError now emits one structured warning with the
  install remedy.
- htmx-config: the response-error logger prints form FIELD NAMES only -
  values (API keys, passwords) never reach the console.
- plugin-order-list: saved order/exclusions normalized with Array.isArray
  (a saved "null" previously crashed .forEach); each row gained
  keyboard/touch-accessible move-up/move-down buttons (HTML5 drag events
  don't fire on most mobile browsers) that reorder and syncInputs()
  immediately alongside native drag.
- app-shell: window.installedPlugins setter always takes the new list
  (same-ID metadata/enabled updates were silently dropped); tab rebuild
  stays gated on ID changes. LED dot renderer reads the frame with ONE
  getImageData call instead of one per pixel (~9,200/frame at 192x48).
- plugins_manager: togglePlugin returns its request promise resolving the
  API outcome; the install flow now shows the "installed and enabled"
  toast (with Restart Now) only after enablement succeeds, and a warning
  without a restart offer when it fails.
- a11y: hamburger aria-label flips Open/Close with drawer state; both
  Advanced-section toggle buttons declare aria-controls/aria-expanded and
  the shared toggleSection() keeps aria-expanded in sync; move buttons
  have per-plugin aria-labels.
- Rotation/Vegas order-list bootstraps cap their retries (~5s) and show a
  reload hint instead of spinning forever; Alpine app-state lookups prefer
  [x-data="app()"] with a generic fallback.

Skipped, with reasons:
- executePluginAction arg order: caller (plugin_config.html) already
  passes (actionId, index, pluginId) matching the signature exactly.
- generateFieldHtml XSS, entity-unescape blocks, dotToNested pollution,
  and "app.loadInstalledPlugins" in app-shell: all inside the legacy
  client-side config cluster whose entry points are shadowed by
  plugins_manager.js / replaced by server-rendered forms (zero live
  callers, verified) - queued for wholesale deletion in the follow-up
  rather than patching dead code.
- custom-feeds-helpers.js findings (3): file was deleted in a prior commit.
- console.error/warn override removal and afterSwap script re-execution
  removal: deliberate pre-existing workarounds every partial's inline
  init currently depends on; reworking them safely needs isolated testing
  (follow-up), and the error suppression is already double-gated
  (insertBefore AND htmx match).
- "move durations bootstrap into a bundle": inline partial-scoped init is
  the established pattern for HTMX partials in this codebase.

Validation: all 40 web tests pass; py_compile on all touched Python; all
touched templates parse; rotation-order defensive paths unit-tested.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* chore(web): fix remaining real Codacy findings (2 of 6)

- htmx-config.js: two more unused catch bindings dropped (optional catch
  binding), matching the earlier fix.
- app-early.js: second forEach arrow (the stub updatePluginTabs copy)
  braced so the callback no longer returns tab.remove()'s value.

The other 4 findings ("Non-serializable expression must be wrapped with
$(...)") are deliberately NOT "fixed": that rule belongs to a
browser-automation (WebdriverIO-style) lint context and is misfiring on
ordinary arrow-function constants. Converting them to function
declarations would look compliant but BREAK the code - all four arrows
intentionally capture the enclosing Alpine component's `this` for the
stub-to-full enhancement logic. The right remedy is disabling that
pattern for this repo in Codacy's Code Patterns settings (or dismissing
the four findings), not a code change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): floating preview shows a frame immediately on open + is resizable

The floating preview opened empty and stayed empty until the display next
CHANGED - the SSE stream only pushes frames on change, and the panel only
consumed frames while already open, so the connection's initial frame
(sent while the panel was closed) was dropped. Reported from mobile
testing as "the button doesn't work".

- updateDisplayPreview now caches the latest frame globally regardless of
  panel state; opening the panel populates the image from that cache
  instantly, then live frames take over.
- Resizable: a size button cycles 192/256/384/512px presets (persisted per
  browser; works on touch), and desktop additionally gets a native drag
  handle (CSS resize: both). The image is fluid within the panel; on
  phones the panel is capped to the viewport width. The size icon
  (fa-up-right-and-down-left-from-center) is verified present in the
  vendored FA 6.0.0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): stop duration fields leaking into config root; isolate per-file dependency installs; fix test fixture leak

Verified each finding against current code.

- api_v3 save_main_config: both duration blocks (*_duration suffix fields
  and the newer duration__<mode> fields) only READ from `data`, never
  removed the keys. The generic "remaining keys" merge later in the same
  function has no skip-list entry for either pattern, so every duration
  field was ALSO written a second time as a bogus top-level config key
  (e.g. "clock_duration": 30 and "duration__mlb_live": 42 sitting at
  config root, alongside the correct nested
  display.display_durations.<key>). Confirmed by tracing the full
  function. Fixed by popping each handled key from `data` (same pattern
  already used for plugin_rotation_order) and validating strictly: a
  non-integer duration now returns 400 with a message naming the
  offending field/mode instead of silently logging and moving on (for the
  *_duration fields, which previously had zero validation at all).
- api_v3 dependency-install loops (git_pull's post-update sync and
  install_base_requirements): _pip_install_requirements can raise
  subprocess.TimeoutExpired or OSError (confirmed: install_requirements_file
  in permission_utils.py never catches either internally, despite its
  docstring's "never raises on non-zero exit" only covering return codes).
  Both loops previously let one file's exception either abort the whole
  try block (skipping the second requirements file entirely) or propagate
  uncaught. Each file's install is now in its own try/except, so a timeout
  or OSError on one file is recorded as a labeled failure and the loop
  continues to the next file.
- test_web_smoke.py: the `client` fixture mutated the module-level
  pages_v3 Blueprint singleton's config_manager/plugin_manager directly
  with no teardown - since pages_v3 is shared across the whole pytest
  process (test_web_settings_ui.py touches the same attributes), this
  fixture's mocks could leak into whichever test ran next. Now saves the
  originals, yields the client, and restores them in a finally block.

Validation: py_compile passes; all 40 web tests pass with the now-generator
fixture.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): repair garbled Advanced Hardware section description

An earlier sed-based text update concatenated the old and new copies of
this description instead of replacing one with the other, leaving a
duplicated sentence with the &mdash; entity broken into ".mdash;" (visible
as literal "mdash;" text on the page). Restored to one clean sentence.

Other findings from this review were already fixed in a prior commit
(installedPlugins setter) or are confirmed dead code with zero live
callers (executePluginAction/dotToNested/entity-unescape/generateFieldHtml,
all reachable only from the two unused savePluginConfig copies in
app-shell.js - grepped every template, no references) - same legacy
cluster flagged in earlier review passes, still queued for a dedicated
deletion follow-up rather than patched in place here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): remove redundant htmx:afterSwap script re-execution (was double-executing every partial's inline script)

Re-verified this CodeRabbit finding, previously deferred as "needs isolated
testing" - traced it to a confirmed, active bug rather than a style
concern:

htmx 1.9.10's own config defaults to allowScriptTags: true (confirmed in
the vendored htmx.min.js, which itself contains the same clone-and-reinsert
<script> mechanism internally). This means htmx ALREADY re-executes every
<script> tag in swapped content by default, exactly like a browser
navigating to a new page. The custom htmx:afterSwap listener in
htmx-config.js did the identical clone-and-reinsert a SECOND time on top of
htmx's own handling - so every inline <script> block in every HTMX-loaded
partial (overview, display, durations, plugin config, etc. - most partials
have one) executed twice per load.

Confirmed safe to delete outright, not just narrow: grepped every hand-written
JS file for a manual `dispatchEvent(... 'htmx:afterSwap' ...)` that might
have relied on this handler for a non-htmx code path (e.g. the direct-fetch
fallbacks like loadOverviewDirect) - none exists, so nothing depended on
this listener specifically; htmx's native handling covers every real
htmx-driven swap on its own.

Left in place, unchanged: the console.error/console.warn global override
a few lines up in the same file, which suppresses known-noisy
HTMX-timing-race messages. That one is a legitimate anti-pattern too
(broad substring matching can mask unrelated errors) but redesigning it
needs care to preserve real diagnostics while still hiding the specific
harmless races it targets - a scoped follow-up, not a same-day deletion
like this confirmed-duplicate handler.

Validation: all 27 fast web tests pass; JS brace/paren balance sanity
checked (no local Node/browser available in this sandbox to execute the
file directly - verify manually in-browser before merge).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* chore(display): add missing [DisplayController] prefix to the reconcile-complete log

Re-verifying the full CodeRabbit findings list against current code
surfaced one still-open item: the nitpick asked for the prefix on BOTH
rotation-related log lines, but only "Applied plugin rotation order" got
it in the earlier pass - "Plugin reconcile complete" was missed. No
message/argument/level change, matching the finding's own scope.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web): repair dead /v3/logs link on the display hardware-error banner

The "Logs tab" link in the display-settings simulation-mode banner was a
real <a href> to /v3/logs, but no such route has ever existed (log content
is loaded client-side via activeTab, not a dedicated page route) - the link
404'd regardless of the /v3 prefix change. Switch it to the same
activeTab-switching pattern the real nav uses.

* fix(web): raise display Rows field max from 64 to 128

Cols already allowed up to 128; Rows was capped at 64, which rejects
valid larger panel configurations (e.g. 128-row tile chains). No
server-side schema enforces a rows max, so this was purely an
overly-strict HTML input attribute.

* fix(web): remove redundant htmx.org substring check flagged by CodeQL

CodeQL flags .includes('htmx.org') as "incomplete URL substring
sanitization" - a false positive here, since this string is only ever
matched against console.error/warn message text to decide whether to
suppress a known-harmless HTMX timing-race log line, not used for any
URL-trust/redirect decision. The check was also redundant: 'htmx' is
already a substring of 'htmx.org', so the plain .includes('htmx') check
right next to it already covers every case the removed check did.

* fix(web): restore htmx script re-execution timing that Alpine x-data depends on

Removing the custom htmx:afterSwap script-reexecution handler (in a prior
commit, as a "duplicate execution" cleanup) broke every partial whose Alpine
x-data component function is defined by an inline <script> in that same
partial (e.g. wifi.html's wifiSetup()) - confirmed live via browser console:
"Alpine Expression Error: wifiSetup is not defined" on every field in the
WiFi tab.

Root cause: htmx's own native script execution runs during its "settle"
phase (~20ms after swap, per htmx's own defaultSettleDelay), but Alpine's
MutationObserver evaluates x-data on newly-inserted elements synchronously,
right as the swap lands - before settle. So the inline <script> defining
wifiSetup() was still un-run when Alpine tried to call it, and Alpine does
not retry a failed x-data evaluation later once the function does become
defined.

Fix: re-execute swapped <script> tags ourselves on htmx:afterSwap (which
fires synchronously, before settle, beating Alpine's observer), and disable
htmx's own native script re-execution (htmx.config.allowScriptTags = false)
so the same script doesn't also run a second time during settle - restoring
correct timing without reintroducing the original double-execution bug.

Also in this commit:
- fix XSS: unescaped repoUrl in a title attribute in renderSavedRepositories
- replace .includes('github.com') substring checks with real URL hostname
  validation (CodeQL: incomplete URL substring sanitization)

* fix(web): wait for async plugin install to finish before auto-enabling it

Confirmed live: installing hockey-scoreboard logged "installation queued"
(success) immediately followed by "enabling it failed" with a 404 "Plugin
not found" from /api/v3/plugins/toggle.

/api/v3/plugins/install runs the actual clone + plugin-manager discovery
asynchronously via an operation queue when one is configured - the response
installPlugin() was checking only means the operation was queued, not that
the plugin is installed yet. Calling togglePlugin() right after that
response 404s because plugin_manager hasn't discovered the new plugin.

Fix: reuse the same operation-polling mechanism uninstallPlugin() already
has (generalized pollOperationStatus() to take onComplete/onFailed/onTimeout
callbacks instead of hardcoding uninstall behavior) so installPlugin() waits
for the operation to actually complete before enabling it. Falls back to
enabling immediately when no operation_id is returned (direct/synchronous
install path, no queue configured).

* fix(web): resolve remaining valid findings from latest review pass

- custom-feeds.js: fix asset-upload contract mismatch (field name "file" ->
  "files", response read from top-level "uploaded_files" not "data.files") -
  same bug already fixed in this file on a separate branch/PR (#420), which
  this branch never received since they're independent PRs off main
- custom-feeds.js: add aria-label to the two icon-only "remove feed" buttons
- custom-feeds.js: move file-input reset into .finally() so a failed upload
  doesn't leave the input stuck holding the file, blocking retry of the
  same file
- app-shell.js: fix executePluginAction(pluginId, actionId) parameter
  order/count mismatch vs. its callers' (actionId, index, pluginId) -
  currently masked by plugins_manager.js's correct version overwriting this
  one at load time (classic vs. deferred script order), but worth fixing
  outright since it's an isolated, self-contained reassignment (not inside
  the Alpine app() object literal) and removes a latent footgun
- overview.html: align Alpine-state resolution with settings-search.js's
  two-tier getAppData() (also check appEl.__x.$data, not just _x_dataStack)

Verified already addressed by earlier passes (no change needed):
plugin_rotation_order validation, DisplayController log prefixes, togglePlugin
returning its promise for install-flow chaining, installedPlugins setter
always updating state, mobile-nav aria-label, toggleSection aria-expanded
sync, PluginOrderList bounded init retries (both display.html and
durations.html), plugin-order-list.js Array.isArray validation, batched
getImageData in the LED-dot preview renderer, app.py exception narrowing/
logging, form-submission log redaction.

Confirmed dead code, skipped (unreachable - zero template/JS callers,
verified via full-repo grep): dotToNested prototype-pollution hardening,
generateFieldHtml HTML-injection hardening, and the HTML-entity-unescape
block in JSON parsing - all three live only inside app-shell.js's two
legacy savePluginConfig implementations (one Alpine-method, one standalone),
neither of which any template or script calls. The real, live plugin-config
path is server-rendered via GET /partials/plugin-config/<id>.

Explicitly NOT reverted: the htmx:afterSwap script-execution listener. An
earlier finding batch asked to remove it as "duplicate" htmx behavior; that
was tried and reverted this session after live testing on hardware proved
it broke every partial whose Alpine x-data depends on an inline <script>
in the same partial (confirmed: WiFi tab hard-failed with "wifiSetup is not
defined"). Removing it again would reintroduce that regression.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-16 20:32:07 -04:00
ChuckandClaude Sonnet 5 66f9950a30 chore(ci): add security-audit tooling scripts (workflow files pushed separately) (#414)
* chore(ci): add security-audit workflow and plugin security-proof scripts

- scripts/prove_security.py, audit_plugins.py, generate_report.py --
  automated checks (dangerous eval()/exec() calls, dependency scanning,
  report generation) for plugins.
- .github/workflows/security-audit.yml + bandit.yaml -- CI wiring for
  the above plus gitleaks secret scanning and bandit static analysis.
- .github/workflows/tests.yml -- pytest matrix across Python 3.10-3.12.

Also fixes two Codacy findings while these files are freshly landing:
- prove_security.py: dropped a pointless f-string prefix with no
  placeholders.
- security-audit.yml: pinned gitleaks/gitleaks-action to a full commit
  SHA (matching this repo's existing pinning convention in test.yml)
  instead of the floating v2 tag.

Split out of the original chore/dead-code-removal commit, which had
accidentally bundled this in alongside unrelated dead-code deletions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* chore: drop workflow files -- pushed separately (needs workflow OAuth scope)

* fix(security-tooling): address PR review findings across bandit.yaml, audit_plugins.py, generate_report.py, prove_security.py

bandit.yaml:
- Removed scripts/prove_security.py's file-level exclusion. Ran bandit
  directly to get ground truth: the real false positive is B105 (dict key
  "PASS" misread as password-like), not the eval/exec pattern the old
  comment claimed. Added a targeted # nosec B105 there, and found+fixed
  the identical pattern already present in generate_report.py.
- Left the repo-wide B607 skip as-is: confirmed via AST scan that properly
  narrowing it touches 100+ bare-name subprocess call sites across
  wifi_manager.py, store_manager.py, permission_utils.py, app.py, and
  start.py -- none of which are part of this PR. Out of proportion to fix
  here; flagged as a dedicated follow-up.

scripts/audit_plugins.py:
- SyntaxError/OSError while scanning a plugin file now report CRITICAL
  (blocking) instead of WARNING/INFO -- a file that couldn't be parsed or
  read was never actually checked for danger, so it must not silently
  pass the audit.
- --plugin <name> now tracks whether the requested plugin was found across
  all PLUGIN_BASE_DIRS and exits 1 with a clear error if not, instead of
  silently scanning zero plugins and reporting success.
- The AST visitor now tracks import aliases (import subprocess as sp;
  from builtins import eval as e) and resolves them before checking
  against dangerous APIs, closing a straightforward evasion of every
  PLUGIN-001 through PLUGIN-005 check. Verified against both aliased and
  unaliased evasion patterns.

scripts/generate_report.py:
- _md_table_row now escapes pipe characters and normalizes newlines in
  every cell, so scanner-controlled content (a matched secret, a bandit
  issue_text) can't corrupt the Markdown table structure.
- _load now distinguishes "artifact missing/malformed" from "valid empty
  result": each summarizer returns an availability flag, and main() now
  reports INCOMPLETE (not PASSED) with exit code 1 when any artifact is
  unavailable, instead of silently folding it in as 0 findings.
- Gitleaks suppression now uses exact-match placeholder values (pulled
  from the actual config_secrets.template.json) plus a template-path
  allowlist, replacing broad substring checks that could hide a real
  secret containing something like "example.com" as part of its value.

scripts/prove_security.py:
- T1b (dangerous plugin calls): a file that fails to parse/read now
  reports CRITICAL with the exception details instead of being silently
  swallowed by `except (SyntaxError, OSError): pass`.
- T6 (Docker hardening): base images must now be pinned to an @sha256
  digest; a specific tag like python:3.12 is mutable and is now correctly
  flagged as unpinned, not just missing tags or :latest.
- T2a (API surface): no config mechanism for enforcing local-only access
  exists in this codebase today (app.py hardcodes host='0.0.0.0'), so the
  "environment-aware" check as described isn't buildable without adding
  new config infrastructure -- out of scope here. Applied the achievable
  part: upgraded from INFO to WARNING, since enforcement can never
  currently be confirmed.
- T1a (zip-slip): replaced the whole-file substring check with an AST
  walk that finds every extract()/extractall() call and confirms an
  is_relative_to() guard + "Zip-slip detected" log precede it in the same
  function. Verified it still passes on the real store_manager.py (both
  the per-member and validate-then-bulk-extract call sites) and correctly
  flags a synthetic unguarded extractall().
- T3a (hardcoded secrets): violation details no longer include the
  matched credential text -- only file, line, pattern type, and a
  redacted SHA-256 fingerprint, so a real finding doesn't get published
  into CI logs/artifacts/PR comments with wider exposure than the
  original leak. Verified with a synthetic secret that no raw content
  reaches the output.

Validated: all four files compile; bandit scans all three scripts clean
(2 legitimate targeted suppressions, 0 unaddressed findings); each new/
changed code path exercised directly (alias evasion, unmatched --plugin,
missing/malformed/valid-empty artifacts, digest-pinning, zip-slip
guard/no-guard, secret redaction); full audit_plugins.py -> prove_security.py
-> generate_report.py pipeline run end-to-end producing a correct report.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(security-tooling): address follow-up review findings on PR #414

scripts/generate_report.py:
- Added an explicit `object` type annotation to _md_sanitize_cell's value
  parameter -- it deliberately accepts any stringifiable value (calls
  str(value) unconditionally), so `object` reflects its actual contract
  more accurately than leaving it untyped.

scripts/prove_security.py:
- Dockerfile FROM-line parsing: renamed the comprehension variable `l` to
  `line` (ambiguous single-letter name). More importantly, fixed a real
  false-positive: `FROM --platform=<platform> <image>` was reading the
  --platform= flag itself as the image token, so a properly digest-pinned
  image behind a platform flag was incorrectly reported as unpinned.
  Verified against platform+digest, platform+tag-only, and digest+AS-alias
  Dockerfiles.

scripts/audit_plugins.py:
- Consolidated visit_Call's dangerous-API detection: previously, alias
  resolution only covered ast.Name calls for eval/exec/compile and
  ast.Attribute calls for subprocess/os.system, missing from-imported
  subprocess/os functions called as bare names (from subprocess import
  run as prun; prun(cmd, shell=True) or from os import system as s;
  s(cmd)). Added _resolve_call_target() to resolve both call shapes to a
  single fully-qualified target, then run all five PLUGIN-00x checks
  against that one resolved value. Verified against 10 evasion
  combinations (from-import aliases, direct/attribute calls, aliased
  module imports) and confirmed zero false positives on benign os/
  subprocess usage without shell=True.

Validated: all three files compile, bandit scans clean (same 2 legitimate
suppressions as before, 0 new findings), audit_plugins.py/prove_security.py
re-run against the real repo with no regressions from the prior fix pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-15 14:33:17 -04:00
ChuckandClaude 4abcd0e4f9 Fix PermissionError reading config_secrets.json in web interface (#416)
ledmatrix.service (main display) runs as root while ledmatrix-web.service
runs as the non-root install user (install_web_service.sh). Both
config_manager.py and config_manager_atomic.py only chmod'd
config_secrets.json to 0o640 without ever fixing its group, so a file
written by the root service ended up group-owned by root and unreadable
by the web user, crashing the settings page with a raw PermissionError.

Add ensure_shared_group_ownership() to chgrp secrets/config files (best
effort, root-only) to the project directory's owning group whenever they
are created or saved, and self-heal existing files on load. Also make
get_raw_file_content() tolerate an unreadable secrets file the same way
load_config() already does, degrading to empty secrets instead of a 500.

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-15 14:25:58 -04:00
2a1c47fa76 chore: remove dead modules and unused dependencies (~1,180 LOC) (#412)
* chore: remove dead modules and unused dependencies (~1,180 LOC)

Deletions, each re-verified with a fresh repo-wide grep (core, web,
scripts, docs, plugin monorepo) immediately before removal:

Modules with zero live importers:
- src/background_cache_mixin.py + src/generic_cache_mixin.py (134+150
  LOC — referenced only by each other)
- src/font_test_manager.py (134 LOC)
- src/image_utils.py (22 LOC, self-documented deprecated)
- src/layout_manager.py (408 LOC — only its own test imported it) +
  test/test_layout_manager.py
- src/common/basketball_plugin_example.py (328 LOC sample)

requirements.txt entries with zero importers in core (pre-plugin-era
manager deps): icalevents, geopy, timezonefinder, unidecode. Plus the
google-auth trio (google-auth-oauthlib, google-auth-httplib2,
google-api-python-client): their only importer is the calendar PLUGIN,
which declares all three in its own requirements.txt (verified in the
monorepo and on an installed copy) — the plugin dependency installer
owns them. Existing venvs are unaffected (removal doesn't uninstall);
fresh installs get them when calendar is installed.

Two stale references cleaned (a comment in test_pillow_compat.py, a
directory listing in HOW_TO_RUN_TESTS.md). Full suite green except the
two documented pre-existing failures (circuit_breaker mock drift, fixed
in #400; clock-simple 64x32 overflow, pre-dates this series); all core
entry modules verified importing cleanly under the emulator.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix: remove content accidentally bundled into the dead-code-removal commit

The previous commit's git add/commit swept in a lot of unrelated,
unreviewed work alongside the intended dead-code deletions: a new Plugin
Composer web UI, new security-audit/CI tooling, a new march-madness
plugin, and 23 local-development-only symlinks under plugin-repos/ (per
scripts/setup_plugin_repos.py's own docstring, these are meant to be
generated locally, never committed -- .gitignore has no entry for them,
which is how they slipped in).

Removed here, split into their own PRs instead (except plugin-repos/*
symlinks and march-madness/ncaa_logos, which are dropped rather than
carried forward -- see PR discussion):
- web_interface/blueprints/composer.py + composer-app.js +
  composer-canvas.js + composer.html + manager.py.j2
- scripts/prove_security.py, audit_plugins.py, generate_report.py
- .github/workflows/security-audit.yml, .github/workflows/tests.yml,
  bandit.yaml
- All plugin-repos/* symlinks (local dev artifacts, not meant to be
  committed at all)
- plugin-repos/march-madness/* and the 4 new assets/sports/ncaa_logos/*
  PNGs it needed (left out of every split PR pending a decision on
  whether march-madness belongs in the core repo or the plugin monorepo)

This PR now contains only what its title describes: the dead-code
removal from the previous commit, untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 09:58:31 -04:00
ChuckandClaude Sonnet 5 9db1d2391a chore(assets): add 4 NCAA team logos needed by march-madness (COLGATE, LEHIGH, MICHIGAN, RUTGERS) (#415)
march-madness (already live in ledmatrix-plugins) loads team logos from
this shared assets/sports/ncaa_logos/<ABBR>.png cache at runtime
(manager.py:233) -- these 4 were missing.

Split out of PR #412 (chore/dead-code-removal), which had accidentally
bundled these in alongside unrelated dead-code deletions.


Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-15 09:57:52 -04:00
14a59c863c perf(wifi): one status fetch per monitor tick instead of three (#411)
* perf(wifi): one status fetch per monitor tick instead of three

The wifi monitor daemon fetched WiFi status + ethernet state before its
AP-mode check, the check internally fetched the same state again (with
retry), and the daemon fetched a third time afterwards — each fetch is
several nmcli subprocess forks, every 30s, forever, even on a perfectly
healthy link.

check_and_manage_ap_mode's decision logic is extracted to
_manage_ap_mode(status, ethernet, ap_active); the new
check_and_manage_ap_mode_with_state() runs the single (retrying) fetch
battery and returns (changed, status, ethernet, ap_active_after) — the
post-state is derivable because state only ever flips via one enable or
one disable. The original bool-returning method delegates, so existing
callers are untouched. The daemon's dead pre-fetch is removed and its
post-check reads use the returned state; the AP-enable retry semantics
(_get_wifi_status_with_retry) are preserved exactly. ~10-15 forks/30s
drops to ~4-6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(wifi): add missing type hints on AP-mode state helpers, fix unused-var lint in tests

- check_and_manage_ap_mode_with_state now declares its
  Tuple[bool, WiFiStatus, bool, bool] return type instead of being untyped.
- _manage_ap_mode's status/ethernet_connected/ap_active parameters are now
  typed (WiFiStatus, bool, bool), matching its existing -> bool return hint.
- test_wifi_check_state.py: rename unused unpacked variables (status/
  ethernet/ap_after) to underscore-prefixed names at the three call sites
  that don't use them, leaving the used ones untouched.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 11:22:37 -04:00
bff13129c4 perf(config): mtime-signature fast path for load_config (#410)
load_config re-read and re-parsed config.json, the secrets file, AND the
template (running the recursive migration diff) on every call — with
~30 call sites in web request handlers, some hit 2-3x per request.

Fast path: stat all three files (mtime_ns + size); when unchanged since
the last successful load, return the already-parsed self.config (same
aliasing semantics as before). The signature is taken AFTER load +
migration so a migration write-back doesn't retrigger, and both save
paths refresh it. Cross-process freshness is preserved by construction:
a save from the other process bumps the file mtime, so the next load
here re-reads — verified by a dedicated test. Same-second edits are
caught by mtime_ns plus a size check.


Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 08:49:43 -04:00
ChuckandClaude Sonnet 5 6499794c12 fix(testing): add MockCacheManager.get_cached_data_with_strategy/save_cache (#409)
* fix(testing): add MockCacheManager.get_cached_data_with_strategy/save_cache

ledmatrix-leaderboard's data_fetcher.py calls these two real-CacheManager
methods (src/cache_manager.py:313,817), but MockCacheManager had neither --
update() always hit an AttributeError, caught by a broad except and logged,
so the harness rendered an empty-but-green leaderboard on every test run
without ever exercising real standings data.

Both mocks delegate to the existing get()/set() -- a mock doesn't need the
real strategy's per-data-type max_age/market-hours timing, plugins under
test just need the methods to exist and round-trip whatever was cached.

* fix(testing): clear get_cached_data_with_strategy_calls in MockCacheManager.reset()

reset() cleared get_calls/set_calls/delete_calls but not the newer
get_cached_data_with_strategy_calls tracker, so a reused mock (e.g. across
test cases sharing a fixture) retained stale strategy-call records after
reset().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-14 08:33:08 -04:00
Chuck 3d347a368a fix(testing): stop plugin enabled:false schema defaults from silently disabling harness tests (#408)
check_plugin.py, render_plugin.py, and the pytest plugin matrix each built
config as {"enabled": True} then merged in config_schema.json's defaults on
top, letting a plugin's own enabled:false default (a reasonable choice for
a seasonal/opt-in plugin -- 15 of 23 real plugins ship one) silently win.
Every harness/CI render of those plugins was testing "disabled, do
nothing" rather than real behavior.

Extract build_full_config() into testing/loading.py (already the shared
home for plugin-discovery/config-default logic) and use it from all three
call sites: schema defaults, then a forced enabled=True, then harness.json's
config, then the caller's explicit config -- so a test can still
deliberately disable a plugin on purpose, it just can't happen by accident
via the plugin's own shipped schema default anymore.
2026-07-14 08:21:35 -04:00
0aca40cf3a perf(plugins): run scheduled updates off the render thread (#407)
* perf(plugins): run scheduled updates off the render thread

plugin update() executed inline in the render loop — execute_update's
internal thread.join(timeout=30) blocked it, so one slow plugin HTTP
fetch froze scrolling for the whole fetch (up to 30s; DNS-retry storms
made this a regular occurrence on flaky networks).

Scheduling stays on the render thread and keeps every existing gate
(enabled, circuit breaker, can_execute, interval); due updates are now
enqueued to a single background worker (serialized — same one-at-a-time
execution as before, no thundering herd). RUNNING is set at enqueue so
can_execute blocks re-entry alongside the pending-set dedup.

Per-plugin locks make the old implicit update/display no-overlap
guarantee explicit: the worker holds the plugin's lock through its
update; the display side try-locks and, when the plugin is mid-update,
holds the last frame for that iteration — reported as success so a
mid-update skip never advances the rotation. Unlike before, the
guarantee now also holds across the post-timeout window (previously the
lingering update thread overlapped display()). Deadlock-free by
construction: the worker takes one lock; display never blocks.

Timeout semantics unchanged (lingering daemon thread documented).
Kill switch: plugin_system.synchronous_updates: true restores the
inline path.

8 new concurrency tests (non-blocking scheduler, overlap assertion
under a hammering display loop, lock release on failure/timeout paths,
dedup, kill switch); 4-min devpi soak clean (updates completing,
rotation advancing, no stuck RUNNING states).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(plugins): fix skipped-frame health/force_change tracking, config validation, and lock-lifetime gaps in async updates

Addresses PR #407 review findings:
- display_controller: only clear force_change / record health success
  when display() actually ran this frame, not when the frame was skipped
  because the plugin's lock was busy (a skip must preserve a pending
  mode-switch force_clear).
- plugin_manager: replace bool() coercion of synchronous_updates with
  explicit isinstance validation of plugin_system/synchronous_updates,
  failing safe to synchronous mode (with a logged reason) on malformed
  config instead of silently defaulting to async.
- plugin_manager: _update_worker_loop now acquires the plugin lock before
  looking up its instance and re-checks under the lock, so an unloaded
  plugin's lifecycle state is never resurrected to ENABLED.
- plugin_manager + display_controller: move lock ownership (and, for
  updates, RUNNING/pending lifecycle bookkeeping) into the actual update()/
  display() call itself rather than the timeout-wrapped caller, so the
  lock stays held for the real operation's duration even after
  PluginExecutor's own join(timeout) elapses and a lingering daemon thread
  keeps running in the background.
- DisplayController.cleanup() now stops the update worker before tearing
  down display/cache resources; stop_update_worker() logs when the join
  times out instead of failing silently.
- test_async_plugin_updates: rewrite test_unloaded_while_queued_is_harmless
  to exercise the public unload_plugin() lifecycle (via a deterministic
  blocker) instead of deleting pm.plugins directly, and add a regression
  test proving the plugin lock stays held through PluginExecutor's own
  timeout while the real update() call is still running.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 08:08:51 -04:00
9837315308 perf(display): dirty tracking in update_display + plugin FPS declaration (#406)
* perf(display): dirty tracking in update_display + plugin FPS declaration

update_display now skips SetImage+SwapOnVSync when the frame is
byte-identical to the last pushed one (adler32 digest) AND brightness
is unchanged — brightness is part of the digest, and set_brightness
additionally resets it, so a dim-schedule change can never be skipped.
clear() resets the digest (it writes to the matrix directly). Skipping
a swap is hardware-safe: the panel refreshes the current frame from the
driver's own thread; swaps only change content.

Kill switch: display.dirty_tracking: false restores always-push.

display_controller's high-FPS decision gains a precedence step: a
plugin exposing needs_high_fps is honored first (so static-image can
declare False for still PNGs and stop burning a 125fps loop on them);
static-image without the attribute keeps its historical forced
high-FPS (GIF back-compat); scrolling logic is otherwise unchanged.

Verified with 7 tests against the real DisplayManager on
RGBMatrixEmulator (identical-frame skip, pixel-change push, clear and
brightness invalidation, snapshot-through-skip, kill switch) plus the
202-test display/controller/vegas suites, and a clean devpi deploy.
Audit: every SetImage/SwapOnVSync/Clear/brightness call site is inside
display_manager — no external writer can bypass the digest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(display): serialize update_display, narrow brightness exception, log fixes

CodeRabbit review on #406, verified against current code:

- update_display() can genuinely be called from background threads (some
  sports base classes call it directly from inside update() for an
  immediate "live" refresh), not just the render loop — confirmed via the
  existing follower-mode gating wrapper in display_controller.py, which
  exists specifically because "background plugin threads" can reach it.
  Without a lock, two callers could both pass the digest check before
  either writes _last_pushed_digest back, causing a redundant push, or
  interleave the offscreen/current canvas swap. Added self._update_lock
  (RLock, in case of re-entrant callers) around the full method body so
  every call site is automatically covered — no caller changes needed.
  (No prior lock existed to reuse on DisplayManager; this adds one.)
- Narrowed the brightness-read exception handler to AttributeError,
  matching the established pattern in get_brightness()/set_brightness()
  — a getattr() with a default already swallows AttributeError, so the
  only case this guards is the property getter itself raising, and the
  established pattern treats that as an expected, specific failure mode
  rather than something to blanket-catch.
- FPS-check debug log now includes the plugin_id already in scope
  (previously only active_mode) and a "[DisplayController]" prefix for
  grep-ability, matching the sibling log two lines below it.
- test_display_dirty_tracking.py: dm fixture and test_config_flag_wires_through
  now reset the DisplayManager singleton on teardown, matching the pattern
  test_display_manager.py already uses elsewhere in the same file family.
- test_snapshot_still_written_on_skip previously only exercised the
  non-skip (push) path despite its name; now performs a second update that
  meets the skip conditions (identical frame) and asserts the snapshot is
  still written even though the panel push itself is skipped.

All 7 dirty-tracking tests pass, plus the full display_manager/
display_controller/vegas suite (140 passed). Full repo suite has only the
5 known pre-existing failures (double-sided config x2, state_reconciliation
x2, and test_circuit_breaker's conftest.py mock signature drift — the
latter fixed in #400, which this branch's base predates).

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 10:31:20 -04:00
c1fa5094be fix(store): plugin updates keep the old install until the new one succeeds (#405)
* fix(store): plugin updates keep the old install until the new one succeeds

Both reinstall paths in update_plugin — the monorepo-migration remote
switch AND the routine archive update every store user hits — deleted
the installed plugin directory BEFORE downloading its replacement. A
mid-update failure (bad network, registry error) permanently destroyed
the plugin. Seen in the field: a Pi with broken DNS lost 12 plugins in
one update pass during the monorepo migration.

New _reinstall_with_rollback: rename the old install aside (using the
'.standalone-backup-' name pattern plugin discovery already excludes),
run install_plugin, remove the aside on success — restore it on ANY
failure, clearing partial-download debris first. A stale aside from a
previous crash is cleared before starting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(store): serialize concurrent updates per plugin, check cleanup results

CodeRabbit review on #405 flagged two things in
_reinstall_with_rollback, both verified against current code:

- Real race: the web UI runs Flask with threaded=True and there's a
  single update route, so two overlapping requests for the same
  plugin_id (double-click, two tabs) can interleave. The loser could
  rename the winner's in-progress install aside mid-download, deleting
  its own rollback safety net — worse than the bug this function
  exists to fix. Added a lazy per-plugin_id lock dict (mirrors the
  plugin_manager per-plugin lock pattern) held for the whole function.
- _safe_remove_directory's return value was ignored at both call
  sites. Stale-aside cleanup failure now aborts cleanly instead of
  falling through to a rename that would fail anyway with a less
  useful error; post-success backup-removal failure now logs instead
  of failing silently (still returns True — the update itself
  succeeded, and the next update self-heals the leftover aside).

Left the third nitpick (test_stale_aside_from_previous_crash_is_cleared)
addressed by asserting the stale dir is actually gone and that
install_plugin was reached, rather than just the end-to-end result.

Added a concurrency regression test asserting install_plugin never
runs for the same plugin_id while another call is in flight.

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 09:32:37 -04:00
4d49b0f892 perf(display): snapshot mirror — viewer gating, digest skip, keepalive (#404)
The display service PNG-encoded its frame to /tmp/led_matrix_preview.png
at 5 fps, 24/7 — identical frames, no viewers, per-call imports and a
chmod every write. On the devpi baseline the display service idles at
~92% CPU; this was one of its biggest fixed costs.

- New pure policy (src/common/snapshot_policy.py, unit-tested off-Pi):
  WRITE changed frames at full rate only while a viewer is watching,
  at a 30s idle cadence otherwise; NEVER re-encode unchanged frames —
  bump mtime (os.utime) every 20s instead, keeping the health check's
  snapshot-age liveness proxy (60s threshold in api_v3) green. Cross-
  referencing comments guard the two constants.
- Viewer detection: the web SSE display broadcaster (which only runs
  while browsers are subscribed) touches /tmp/led_matrix_preview_viewer
  each loop; the display service stats it at most 1/s. On viewer
  arrival the write clock resets so the first frame lands within ~1s.
- Hoisted the per-call pathlib/permission_utils imports; directory
  permissions ensured once instead of every frame.


Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 09:28:30 -04:00
efe76d3add perf: hot-path micro fixes in the render loop (#403)
* perf: hot-path micro fixes in the render loop

- _check_wifi_status_message stat'd the status file on every render
  iteration (60+ fps) for a message whose lifetime is seconds; throttle
  the check to 1 Hz with a cached result.
- Demote the per-iteration "Display active, processing mode" INFO to
  DEBUG and convert the remaining eager f-string logs to lazy % args —
  the devpi baseline showed ~9 journald lines/sec, which is both noise
  and SD-card wear.
- Vegas cycle-end blank frame: hoist the inline PIL import and reuse a
  preallocated buffer instead of allocating per cycle wrap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix: initialise wifi-status throttle state in __init__

Codacy (pylint access-member-before-definition) on #403: the throttled
early-return read _wifi_status_last_result relying on the non-local
invariant that the first call always passes the throttle window and
assigns it. Correct at runtime, but fragile — initialise both throttle
fields in the constructor and drop the getattr fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 09:23:25 -04:00
273d9962d1 fix(cache): stop fsync-hammering the SD card on unchanged data (#402)
DiskCache.set wrote every key as mkstemp -> json.dump(indent=4) ->
flush+fsync -> replace -> chmod, on the persistent cache dir — dozens
of force-flushed SD writes per minute on an API-heavy install, mostly
rewriting identical data every plugin update cycle.

- Serialize once, compact (no indent): cache files are machine-read
  only; indenting multiplied the bytes written.
- Skip the disk when the payload for a key is unchanged (adler32 map,
  per-process); refresh the file mtime instead so records relying on
  mtime for TTL don't expire early. Self-heals if the file was removed
  externally (expiry cleanup).
- Drop the per-write fsync: os.replace already guarantees readers never
  see a torn file, and cache data is re-fetchable — the flush bought
  nothing but card wear.

API unchanged; DateTimeEncoder round-trip covered by tests.


Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 09:00:35 -04:00
9e3b5f366e fix(core): harden text-measurement caches; surface snapshot failures (#400)
* fix(core): harden text-measurement caches; surface snapshot failures

Deep-dive findings, all three latent on every 24/7 install:

- font_manager.metrics_cache and display_manager._text_width_cache were
  unbounded dicts keyed by (text, id(font)). Two problems: keys embed
  the measured TEXT, so ever-changing strings (a clock, a live score, a
  ticker) grow them without limit; and id()-keying without holding a
  reference means a garbage-collected font's id can be recycled by a
  DIFFERENT font, silently returning wrong widths/metrics (classic
  plugins create fonts per render, so this is reachable). Both caches
  are now LRU-bounded (1024) and pin the font in the entry so its id
  stays valid. metrics_cache also keyed on the text itself instead of
  hash(text), removing a collision path.

- _write_snapshot_if_due logged failures at DEBUG — invisible at the
  default level. The snapshot's mtime is the web UI's display mirror
  AND its hardware-liveness proxy, so a quiet failure freezes the
  mirror and makes health checks lie (seen in the field: a stale
  root-owned /tmp file froze it for a day). Failures now WARN, rate-
  limited to once per 5 minutes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* test: sync mock cache manager signature with CacheManager.get

test_circuit_breaker has been failing on main: plugin_health passes
memory_ttl= to cache_manager.get(), and the conftest mock's signature
was never updated — the same component/double drift class as the
monitored_update bug (#392).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 08:59:41 -04:00
ChuckandClaude 6edd80d9f3 fix(schedule): stop stray 'days' data from overriding Global schedule (#399)
save_schedule_config never persisted the schedule's 'mode' field, and
_check_schedule inferred per-day vs global purely from whether a 'days'
dict was present for the current day. Config migration
(_merge_template_defaults) re-adds the template's 'schedule.days' (all
days disabled by default) whenever it's missing from the user's saved
config - which is exactly the case after saving Global mode, since that
save path intentionally pops 'days'. The result: a user on Global mode
would get their schedule silently reinterpreted as per-day, with today's
day disabled, blanking the display.

Persist 'mode' on save and have _check_schedule honor it explicitly
(mirroring how _check_dim_schedule already does), so a resurrected
'days' dict can't override an explicit Global selection. Falls back to
the old inference behavior only when no 'mode' is recorded.

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-12 10:53:33 -04:00
ChuckandClaude 1c7a0cef66 fix(vegas): lock plugin_last_update snapshot/diff against concurrent mutation (#398)
* fix(vegas): restore live plugin-update refresh dropped by sync refactor

Investigating a user report that Vegas scroll mode doesn't update scores
or game status. Root cause: PR #299 (Mar 28) added a mechanism so a live
score change reached the ticker within a few seconds instead of waiting
for a full scroll cycle -- _tick_plugin_updates_for_vegas() diffed
plugin_last_update timestamps to detect which plugins got fresh data and
called coordinator.mark_plugin_updated() for each, and should_recompose()
checked has_pending_updates_for_visible_segments() to trigger an immediate
hot-swap.

PR #330 (May 14, multi-display wireless sync) refactored both call sites
while adding sync support and silently deleted this entire mechanism --
not just gated it behind the new sync-mode deferral it legitimately
needed, but removed it outright. The result: VegasModeCoordinator.
mark_plugin_updated() and StreamManager.has_pending_updates_for_visible_
segments() have been fully implemented but never called from anywhere
since. Vegas mode's only remaining freshness sources are a 5s content
cache TTL (fine) and full recompose at cycle boundaries, which depending
on min/max_cycle_duration can be minutes away -- so live scores/status
can sit stale far longer than a user would expect from a "live" ticker.

Fix:
- Restored _tick_plugin_updates_for_vegas() in display_controller.py,
  wired as the Vegas coordinator's update callback in place of the plain
  _tick_plugin_updates(). Diffs plugin_last_update before/after the tick
  and calls vegas_coordinator.mark_plugin_updated(plugin_id) for each
  plugin that actually got new data (rather than returning the list, since
  the callback interface no longer consumes a return value).
- Restored the has_pending_updates_for_visible_segments() check in
  render_pipeline.should_recompose(), positioned after (not instead of)
  the sync-mode early return PR #330 added, so standalone installations
  regain immediate refresh while synced leader/follower pairs correctly
  keep deferring hot-swaps to cycle boundaries as PR #330 intended.

Test plan:
- Added test_display_controller_vegas_tick.py and
  test_vegas_render_pipeline_recompose.py -- neither area had any prior
  test coverage, which is very likely why this regression went unnoticed
  for ~2.5 months.
- Verified both new test files fail against the pre-fix code (swapped in
  the current main versions of both files) with exactly the expected
  errors -- AttributeError for the deleted method, and the recompose
  assertion returning False instead of True -- then pass against the fix.
- Confirmed the sync-mode deferral this restoration must not break still
  holds: test_sync_active_defers_pending_updates_to_cycle_boundary.
- Full related suite (test_vegas_plugin_adapter, test_vegas_config,
  test_display_controller_plugin_toggle, test_display_controller_
  optimizations, test_plugin_system): 108 passed, 1 pre-existing failure
  unrelated to this change (test_circuit_breaker, stale mock signature).
- Full CI plugin-safety suite (test_harness, test_visual_rendering,
  test_plugin_matrix): 52 passed, 2 pre-existing skips.

* fix(vegas): lock plugin_last_update snapshot/diff against concurrent mutation

_tick_plugin_updates_for_vegas() snapshotted and later re-iterated
plugin_manager.plugin_last_update from the Vegas background update-tick
thread while the main render loop (or other callers) could mutate the
same dict concurrently — a real race (unprotected dict iteration/mutation
across threads), not just a style nit.

Move the snapshot/update/diff into a new locked
PluginManager.run_scheduled_updates_with_changes() so all reads and
mutations of plugin_last_update happen under one lock, and update
DisplayController to use it. The lock is only held around the dict
accesses, not the update pass itself, so slow plugin update() calls don't
serialize against other callers.

Also add a regression test covering that the Vegas coordinator is wired
to the Vegas-aware tick callback rather than the plain one.

Skipped as not worth the change:
- Narrowing the broad `except Exception` around
  vc.mark_plugin_updated(plugin_id) to specific types: it's a deliberate
  per-plugin isolation boundary (matches the same pattern used elsewhere
  in this file for plugin/coordinator calls) and there's no documented,
  stable set of exceptions that call can raise to narrow to.
- Adding an inactive-DisplaySyncManager test to
  test_vegas_render_pipeline_recompose.py: verified
  VegasModeCoordinator.set_sync_manager() already normalizes a
  SyncRole.STANDALONE manager to None before handing it to the render
  pipeline (src/vegas_mode/coordinator.py:152-156), so should_recompose()'s
  `is not None` check is correct in practice; the suggested case is
  already covered by that normalization.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-12 10:52:18 -04:00
Chuck 6052a60d22 fix(vegas): restore live plugin-update refresh dropped by sync refactor (#395)
Investigating a user report that Vegas scroll mode doesn't update scores
or game status. Root cause: PR #299 (Mar 28) added a mechanism so a live
score change reached the ticker within a few seconds instead of waiting
for a full scroll cycle -- _tick_plugin_updates_for_vegas() diffed
plugin_last_update timestamps to detect which plugins got fresh data and
called coordinator.mark_plugin_updated() for each, and should_recompose()
checked has_pending_updates_for_visible_segments() to trigger an immediate
hot-swap.

PR #330 (May 14, multi-display wireless sync) refactored both call sites
while adding sync support and silently deleted this entire mechanism --
not just gated it behind the new sync-mode deferral it legitimately
needed, but removed it outright. The result: VegasModeCoordinator.
mark_plugin_updated() and StreamManager.has_pending_updates_for_visible_
segments() have been fully implemented but never called from anywhere
since. Vegas mode's only remaining freshness sources are a 5s content
cache TTL (fine) and full recompose at cycle boundaries, which depending
on min/max_cycle_duration can be minutes away -- so live scores/status
can sit stale far longer than a user would expect from a "live" ticker.

Fix:
- Restored _tick_plugin_updates_for_vegas() in display_controller.py,
  wired as the Vegas coordinator's update callback in place of the plain
  _tick_plugin_updates(). Diffs plugin_last_update before/after the tick
  and calls vegas_coordinator.mark_plugin_updated(plugin_id) for each
  plugin that actually got new data (rather than returning the list, since
  the callback interface no longer consumes a return value).
- Restored the has_pending_updates_for_visible_segments() check in
  render_pipeline.should_recompose(), positioned after (not instead of)
  the sync-mode early return PR #330 added, so standalone installations
  regain immediate refresh while synced leader/follower pairs correctly
  keep deferring hot-swaps to cycle boundaries as PR #330 intended.

Test plan:
- Added test_display_controller_vegas_tick.py and
  test_vegas_render_pipeline_recompose.py -- neither area had any prior
  test coverage, which is very likely why this regression went unnoticed
  for ~2.5 months.
- Verified both new test files fail against the pre-fix code (swapped in
  the current main versions of both files) with exactly the expected
  errors -- AttributeError for the deleted method, and the recompose
  assertion returning False instead of True -- then pass against the fix.
- Confirmed the sync-mode deferral this restoration must not break still
  holds: test_sync_active_defers_pending_updates_to_cycle_boundary.
- Full related suite (test_vegas_plugin_adapter, test_vegas_config,
  test_display_controller_plugin_toggle, test_display_controller_
  optimizations, test_plugin_system): 108 passed, 1 pre-existing failure
  unrelated to this change (test_circuit_breaker, stale mock signature).
- Full CI plugin-safety suite (test_harness, test_visual_rendering,
  test_plugin_matrix): 52 passed, 2 pre-existing skips.
2026-07-12 10:40:34 -04:00
7f7f0d6464 feat: adaptive layout system — size-aware regions, crisp font ladders, image fitting (#393)
* feat(layout): adaptive layout & font scaling system for plugins

Add src/adaptive_layout.py — opt-in core helpers so plugins render
legibly on any panel size without hand-tuned per-display layouts:

- Region: integer rect algebra (bands/columns/weighted splits/centering)
  that partitions space so text bands can't overlap by construction
- Font ladders: ordered (family, size) steps known to render crisply
  (LADDER_GRID: X11 BDFs at native sizes; LADDER_ARCADE: PressStart2P at
  8px multiples) — fitting walks the ladder instead of scaling pixel
  fonts fractionally
- LayoutContext: breakpoint tiers, geometry scale vs. a declared design
  size, and cached fit_text/fit_lines/font_for_rows queries

Generalizes the three patterns proven in the field: f1-scoreboard's
scale factor, masters-tournament's tiers, baseball-scoreboard's font
fallback ladder.

Wiring: BasePlugin gains a lazy .layout property and draw_fit();
FontManager gains get_native_bdf_size() and a cache_generation counter;
manifest schema gains display.design_size and requires.display_size
max_width/max_height; 96x48 joins DEFAULT_TEST_SIZES; the bounds-check
harness records negative-coordinate draws; TextHelper's broken
measurement helpers are fixed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): adaptive image fitting + composite region helpers

Add src/adaptive_images.py — the image counterpart to fit_text:
- fit_image(img, box, mode=contain|cover|fill_height|stretch,
  crop_to_ink, anchor, resample, upscale) promoting the proven plugin
  patterns (football's crop-to-ink fill-height logos, masters' cover
  crop + NEAREST flags, static-image's letterbox). Upscales by default —
  thumbnail()'s downscale-only behavior is why imagery stays tiny on
  big panels.
- draw_fitted_image() pastes aligned within a Region with alpha mask.
- One central Pillow>=9.1 RESAMPLE shim replacing ~15 plugin copies.

LayoutContext.fit_image() caches results per (identity, box size,
options) with a 64-entry LRU; id()-keyed entries pin the source image.
BasePlugin.draw_image() is the one-liner adoption path beside draw_fit.

Composites in adaptive_layout.py: Region.offset() (user x/y-offset
passthrough), scoreboard_regions() (the two-logos-plus-score card math
duplicated across six sports plugins, logo_slot = min(H, W//2)), and
media_row() (art-left/text-right).

Fix LogoHelper's size-blind cache key (stale sizes on panel change);
deprecation note on dead image_utils.py.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(harness): scale-up fill check, config variants, multi-size dev gallery

Quality gates for adaptive layout:

- fill_metrics()/check_scale_up() in the safety harness: overflow catches
  content too big for a panel, but nothing caught content that stays tiny
  on panels >= 2x the plugin's declared design size. The check measures
  lit-content extents and warns (or fails, when a plugin opts into
  "fill_check": "strict" in test/harness.json) below 50% coverage on the
  doubled axis. Warn-only by default so no existing plugin breaks.

- harness.json "variants": extra runs with config overlays and their own
  golden dirs, so an opt-in mode (e.g. layout_mode: adaptive) is golden-
  tested beside the classic default. check_plugin.py loops base + variants
  and labels variant results mode@name.

- Dev preview server: GET /api/sizes (harness size sample), POST
  /api/render-matrix (render at up to 12 sizes in one call), size-preset
  dropdown, and an "All Sizes" side-by-side gallery in the preview UI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(plugins): adaptive-lib discoverability + advisory version compat warning

Discoverability: re-export the adaptive layout/image API from src.common
(the blessed-helpers package plugin authors already know) — canonical
paths stay src.adaptive_layout / src.adaptive_images so nothing breaks.
Document it in src/common/README.md and cross-link ADAPTIVE_LAYOUT.md
from the developer docs authors actually read (quick reference, API
reference, advanced dev, font manager, dev preview, plugin dev guide);
ADAPTIVE_LAYOUT.md gains adaptive-images, composite-layouts and
preserving-user-customization sections.

Compat: PluginLoader now logs one advisory warning (never raises) when a
plugin's manifest declares a min LEDMatrix version newer than the running
core, checking the min_ledmatrix_version / requires.* / versions[]
spellings found in the wild. Guarded against stale core version numbers.

src/__init__.py __version__ bumped 1.0.0 -> 3.1.0 to match the latest
release tag (v3.1.0) — it had never been updated and the compat check
needs a truthful number. NOTE: verify this matches the intended release
numbering before the next tag.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): add measure_font_crispness — verify a ladder rung isn't blurry

PIL antialiases TTF outlines by default; a 'pixel-style' font only
rasterizes without antialiasing at specific sizes (for PressStart2P:
exact multiples of its 8px design grid). A ladder rung at an unverified
size silently renders blurry on an LED panel — this exact bug shipped in
both text-display's and football-scoreboard's custom TTF ladders
(non-8-multiple PressStart2P sizes, and '5by7.regular'/'4x6-font' at
sizes that were never actually crisp).

measure_font_crispness(font, sample_text) renders the sample and reports
the fraction of ink-bbox pixels that are neither pure black nor pure
white. BDF fonts (real bitmaps) always score 0.0; TTF ladders should be
verified against this before shipping — see the new
TestFontFitting::test_ladder_arcade_is_crisp pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): add fit_text_proportional — proportional sizing vs. always-maximize

fit_text always picks the largest ladder rung that fits its box. That's
right when an element owns dedicated space, but wrong when several
independently-fitted elements need to stay visually harmonious as the
panel grows: a score's box might have generous room while a neighboring
logo scales by a fixed geometry factor via px() — fit_text lets the score
balloon out of proportion (even overlapping the logo) even though its
individual pick is technically correct.

fit_text_proportional(text, box, base_size_px, ladder) instead targets
base_size_px * self.scale (the same scale factor px() already uses),
picking the nearest ladder rung at or below that target, still capped to
what fits the box, floored at the smallest rung when the target is below
every rung. Refactored the shared largest-that-fits/ellipsize walk into
_walk_ladder() so fit_text and fit_text_proportional don't duplicate it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(layout): fit_text_proportional gains an axis-specific scale override

self.scale (min(width_ratio, height_ratio)) is the right conservative
default for anything whose aspect ratio matters, but a caller whose
surrounding composition already scales along a single axis — e.g.
football-scoreboard's logo_slot = min(height, width // 2), which tracks
height alone — needs text sized the same way, or it reads as
under-scaled next to logos that grew on a panel that only got taller
(128x32 -> 128x64: self.scale stays 1.0 since width didn't grow, but
logos still double).

fit_text_proportional(..., scale=None) now accepts an explicit override;
None keeps the existing self.scale default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(layout): scoreboard_regions reserves real center space at 2:1 aspect ratios

logo_slot = min(height, width // 2) has a blind spot: at exactly 2:1
aspect ratio (width == 2 * height -- a very common shape: two, four, or
more square modules stacked into a taller panel) width // 2 and height
are equal, so the two logo slots claim the ENTIRE width and leave zero
pixels for a center column, no matter how large the panel gets. Not a
'small panel' problem -- 96x48, 128x64, and 256x128 (all exactly 2:1) hit
it identically, while the 128x32 design baseline and panels like 192x48
or 256x32 never do, because height is already the tighter constraint
there.

Two new parameters fix it in the one shared helper every scoreboard-style
plugin composes through:

- min_center_fraction / min_center_design_px reserve at least
  max(width * fraction, design_px * ctx.scale) for the center column,
  capping logo_slot further when needed. The scaled design-px term
  matters on small panels where a flat fraction alone reserves too little
  absolute space.
- score_bleed_fraction extends the score's own fit box (not the logo
  slots themselves) a controlled amount into each side -- the same way
  real broadcast scoreboards let a big score number's edges cross into
  the team marks flanking it. Without this the reserve alone can still be
  too narrow for a short score to render without truncating.

score_area is now genuinely narrower than the full card width (previously
identical to status_band/detail_band, which still span the full width and
overlay the logos -- short text there was never the problem).

Verified against the full harness size spread: a real game score like
'17-21' never needs ellipsis at any tested 2:1-or-tighter aspect ratio
(test_score_never_needs_ellipsis_for_a_short_score), and wide panels
(128x32/192x48/256x32-style) are provably unaffected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs: document scoreboard_regions' center-reserve and score-bleed params

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: address CodeRabbit review on PR #393

- docs: scope the self.layout note to BasePlugin subclasses (others build
  a LayoutContext directly) and make explicit that adaptive layout is
  opt-in — classic rendering stays unless a plugin adopts the APIs.
- dev_server: broaden the render-request catch (a bad manifest.json now
  returns a clean 400 instead of an unhandled 500) and stop echoing raw
  exception text in the loader-failure responses — full tracebacks go to
  the dev server's console log instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): allowlist plugin_id before any path lookup

CodeQL (py/path-injection): plugin_id arrives in request input and flows
into filesystem paths via find_plugin_dir. Gate it with the same
^[a-zA-Z0-9_-]{1,64}$ allowlist the web UI's pages_v3 uses, at the
single choke point every route resolves through.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): lexical containment check on resolved plugin dirs

CodeQL doesn't recognize the interprocedural allowlist as a
path-injection barrier; add the canonical one — normalize (without
following symlinks, since dev plugins are commonly symlinked into
plugins/) and require the result to stay inside the search dir.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): inline normpath containment barrier before render

CodeQL doesn't credit the sanitization inside find_plugin_dir along
this flow; apply its documented barrier (normpath + startswith against
the allowed roots) inline in _parse_render_request, on the exact path
that reaches the render/load sinks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): derive plugin dir from trusted directory listings

CodeQL's barrier-guard recognition doesn't see a startswith check
inside an any() comprehension, so the normalize-and-prefix approach
still flagged. Break the taint outright instead: after lookup, re-derive
the directory by enumerating the search dirs (iterdir) and matching by
path equality — the Path used for all downstream file access is built
solely from trusted listings, never from request input.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

* fix(dev-server): use os.scandir for path-injection barrier, redact stack traces from render responses

CodeQL doesn't model Path.iterdir() as a taint-clearing enumeration the
way it does os.scandir() -- _trusted_plugin_dir's iterdir-based rebuild
still traced plugin_id through to the manifest.json open(). Switched to
scandir, matching the pattern already verified clean on PR #396.

Also stops surfacing raw exception text (update()/display() failures)
in the JSON render response -- logs full detail server-side via
exc_info instead, returning only the exception class name to the
client. And drops path values from three plugin_loader debug/error
logs that CodeQL flags as clear-text-logging of externally-influenced
data, keeping plugin_id (not flagged) for context.

* fix(dev-server): remove conditional-reassignment ambiguity in plugin_dir resolution

CodeQL's path-injection flow still traced through _parse_render_request
after the scandir fix -- the tainted find_plugin_dir() result and the
scandir-derived _trusted_plugin_dir() result shared the same variable
name (plugin_dir), reassigned only on the truthy branch. That merge
point apparently isn't treated as a barrier by the flow analysis, so it
kept tracing the pre-reassignment value through to the manifest open().

Split into two distinct names -- candidate_dir (tainted, used only to
call _trusted_plugin_dir) and trusted_dir (the only name used for any
downstream file access) -- so there's no reassigned variable for the
flow to walk through.

* fix: remove unused imports flagged by Codacy

Union in adaptive_images.py and field in adaptive_layout.py are both
imported but never used -- the last two Codacy findings on this PR,
matching the same fix already applied on PR #396.

* fix(layout): bound the fit cache; never alias the source image in fits

Two latent issues found in a self-review pass:

- LayoutContext._fit_cache was an unbounded dict (the image cache got an
  LRU cap, the text-fit cache didn't). Cache keys embed the fitted TEXT,
  so a plugin fitting changing strings — a live game clock, a ticker —
  on a 24/7 service grows it forever. Now LRU-bounded at 512 entries via
  the same pattern as the image cache.

- fit_image returned the caller's ORIGINAL image object when the source
  was already RGBA at target size (contain/fill_height, no ink crop).
  ImageFitResult is documented as an independent copy, and LayoutContext
  caches results — an aliased image lets later mutations of the source
  corrupt cached fits (or vice versa). Copy in that branch.

Both covered by new regression tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 10:38:52 -04:00
ChuckandClaude Sonnet 5 05e7c43b27 fix(plugins): replace dependency marker files with a real satisfaction check (#390)
* fix(plugins): replace dependency marker files with a real satisfaction check

The .dependencies_installed hash-marker system only tracked "was this exact
requirements.txt hashed before" — not whether the packages it names are
actually present. That made it fragile (a wiped venv, a manually removed
package, or a lost/corrupted marker forces a needless full pip reinstall or,
worse, a false skip) and produced dead weight for the ~10 plugins whose
requirements.txt is comment-only (they still paid a pip subprocess on first
boot before a marker existed).

Replace it with requirements_are_satisfied() in plugin_loader.py, which
checks each real requirement line against importlib.metadata directly, so
install_dependencies() only shells out to pip when something is actually
missing or version-mismatched. Drops the marker file entirely: removed all
marker read/write sites in plugin_loader.py and store_manager.py, the
now-pointless marker-cleanup step in the git-update path, the unused legacy
marker implementation in plugin_manager.py, and the already-stale
clear_dependency_markers.sh script.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(security): close path-injection gap in dependency-satisfaction checks

CodeQL flagged 2 new high-severity "uncontrolled data used in path
expression" alerts at the open() calls inside this PR's new
requirements_has_real_deps()/requirements_are_satisfied() -- both are
reachable from paths that were never run through the basename+trusted-base
sanitiser this codebase already uses elsewhere:

- PluginLoader.install_dependencies() only applied that sanitiser when its
  optional plugins_dir argument was actually passed; the "no plugins_dir"
  branch trusted plugin_dir_real directly. Made plugins_dir required (not
  Optional) so that branch can't exist, and added an explicit guard in
  load_plugin() so install_deps=True without a plugins_dir fails loudly
  instead of silently. Production's only real caller (PluginManager) always
  passes plugins_dir already; the harness/dev-server/render-plugin callers
  all use install_deps=False and are unaffected.

- StoreManager._install_dependencies() never sanitised plugin_path at all,
  and its call sites ultimately derive that path from a plugin's own
  manifest.json "id" field (install_plugin_from_url) -- a malicious plugin
  could otherwise point requirements_file outside plugins_dir. Applied the
  same os.path.basename()-based containment pattern PluginLoader already
  uses (and that CodeQL recognises as a real sanitiser).

Added test_install_dependencies_requires_plugins_dir and
test_install_dependencies_rejects_path_outside_plugins_dir to lock in the
actual security property, not just quiet the scanner. Verified: all 20
tests in test_plugin_loader.py pass, plus the PR's existing test plan
(test_plugin_system.py, test_store_manager_caches.py: 53 passed) and the
full CI plugin-safety suite (test_harness.py, test_visual_rendering.py,
test_plugin_matrix.py: 52 passed, 2 pre-existing skips) all still pass.

* fix(security): replace basename-only sanitiser with a trusted-enumeration check

The previous commit's os.path.basename() + os.path.join() pattern (which a
pre-existing code comment claimed CodeQL recognises as a sanitiser) did not
actually clear the alert -- the next CodeQL run still flagged the same 2
sink lines, plus a new one at the os.path.join() call itself. Taking a
substring of tainted data apparently isn't treated as a barrier by this
query, whatever the comment assumed.

Replaced it with find_trusted_subdir(): enumerate the trusted plugins_dir
via os.scandir() and only use a name that scandir itself produced, matched
by equality against the caller's requested name. The path is then built
from that enumerated entry, not from the caller's string -- a value
sourced from iterating a trusted, non-tainted directory carries no taint
regardless of what it happens to equal, which is a stronger and more
conventional allowlist-style barrier than string-stripping. Applied
identically in both PluginLoader.install_dependencies() and
StoreManager._install_dependencies(), sharing one implementation.

Re-verified: all 65 tests across test_plugin_loader.py (20, including the
2 new security regression tests), test_store_manager_caches.py (35),
test_plugin_system.py (10) pass, plus the full CI plugin-safety suite
(test_harness.py/test_visual_rendering.py/test_plugin_matrix.py: 52
passed, 2 pre-existing skips).

* fix(security): redact URL credentials from pip subprocess output before logging

CodeQL flagged 3 clear-text-logging-of-secrets alerts in
install_requirements_file() (src/common/permission_utils.py:353,360,371).
Pre-existing on main, unrelated to this PR's own diff, but now visible
since the path-injection alerts that previously took priority in the
annotation list are fixed.

The underlying risk is real: pip can echo a private index URL's embedded
basic-auth credentials (from a requirements.txt --index-url line or
PIP_INDEX_URL) back verbatim in its own stderr/stdout on failure, and this
function both logs that output directly and returns it to callers --
store_manager.py's _install_dependencies() logs result.stderr from this
same function too.

Added _redact_url_credentials(), applied immediately after each of the two
subprocess.run() calls (mutating result.stderr/stdout in place) rather
than patching each log call site individually. This closes the leak at
the source: every downstream use -- the three flagged log lines, the
"note" string embedded in the returned stdout, and store_manager.py's own
logging of the returned result -- gets the redacted text for free.

Verified the fixed-phrase "denied" check (`"a password is required" in
result.stderr`) is unaffected, since URL syntax and those phrases don't
overlap -- covered explicitly by
test_does_not_touch_denied_check_phrases. Added
test/test_permission_utils.py (6 tests) covering the redaction helper
directly and both subprocess.run() call sites (the sudo-wrapper branch,
which this repo's scripts/fix_perms/safe_pip_install.sh makes live, and
the no-wrapper fallback branch). All pass.

* fix(security): stop interpolating req_file/pip-output into log calls

The previous commit's redaction (mutating result.stderr/stdout right after
each subprocess.run()) didn't clear CodeQL's clear-text-logging alerts --
same lesson as the path-injection fix earlier in this PR: a static
analyzer can't tell "this value was already sanitised two lines up" from
"this is still the raw tainted value" just by looking at a single log
call in isolation, so it conservatively keeps flagging it regardless of
what the redaction function actually does.

Removed all dynamic interpolation (req_file, result.stderr) from the 3
flagged logger.warning() calls entirely, replacing them with fixed
messages plus (for the one that had it) result.returncode, which is a
plain int with no possible taint. The full redacted detail is still
available where it actually matters -- in the returned
CompletedProcess.stderr/stdout and the "note" text -- just not duplicated
into a log line a scanner has to reason about in isolation.

Re-verified: all 6 test_permission_utils.py tests still pass (they assert
on the returned result, not log call arguments), plus the full
test_plugin_loader.py/test_store_manager_caches.py/test_plugin_system.py
suite (71 passed, 1 pre-existing deselect, 4 subtests).

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-11 08:55:50 -04:00
Chuck 2ffc57cf40 fix(plugin-harness): add no-op process_deferred_updates to test double (#391)
The safety harness's VisualTestDisplayManager (base of
BoundsCheckingDisplayManager) doesn't implement process_deferred_updates(),
which 5 first-party ledmatrix-plugins call unconditionally between
set_scrolling_state() and their scroll-position update: news, odds-ticker,
ledmatrix-leaderboard, stock-news, and ledmatrix-stocks. Any of them fails
the harness with AttributeError the moment it's touched (surfaced when
ledmatrix-plugins#177 had to add a local hasattr guard in ledmatrix-stocks
just to pass CI). Add the method as a no-op, mirroring the existing
"no-op for testing" pattern already used for set_scrolling_state, so these
plugins render under the harness without every touching PR needing its own
guard.
2026-07-11 08:55:03 -04:00
Chuck aab0e9ade0 fix(plugin-manager): fix TypeError breaking every plugin's scheduled update (#392)
run_scheduled_updates()'s resource-monitor branch wrapped the update call
in a closure stored as a *class* attribute on a dynamically-built type
(type('obj', (object,), {'update': monitored_update})()). The descriptor
protocol turns a function found via class-attribute lookup into a bound
method on instance access, silently prepending the synthetic instance as
an implicit first argument -- but monitored_update() takes none, so every
call raised "monitored_update() takes 0 positional arguments but 1 was
given", was caught by run_scheduled_updates' try/except, and recorded as
an update failure.

self.resource_monitor is None by default and was dormant until PR #388
("activate dormant plugin health/metrics subsystem") wired it up in both
display_controller.py and web_interface/app.py -- meaning this bug went
live in every real deployment as of that merge (2026-07-09) despite the
buggy line itself dating back to 2025-12-27. In practice this means no
plugin's update() has succeeded since upgrading past #388: circuit
breakers cycle through half-open -> immediate failure -> reopened every
health-check interval forever, and all plugin data (scores, odds, prices,
etc.) goes stale from whatever was last fetched before the upgrade.
Confirmed live on a running instance: odds-ticker (and stock-news,
ledmatrix-stocks, baseball-scoreboard, ledmatrix-leaderboard, of-the-day)
failing this exact way every 5-minute circuit-breaker retry.

Fixed by using types.SimpleNamespace(update=monitored_update) instead of
a dynamic class: SimpleNamespace stores attributes on the instance
itself, so attribute lookup returns the plain function unchanged --
never routed through the class-attribute descriptor protocol that
injects an implicit self.

Added test_run_scheduled_updates_calls_update_with_resource_monitor to
test/test_plugin_system.py using a real PluginResourceMonitor (not a
mock of it), so the test exercises the actual descriptor-binding
behavior that caused this. Verified the test fails with the exact
reported error against the pre-fix code and passes against the fix.
2026-07-11 08:54:30 -04:00
ChuckandClaude Sonnet 5 978a03b42d Add settings tooltips and search to the web UI (#387)
* Add settings tooltips and search to the web UI

Help users quickly find settings and understand how each one works.

Tooltips: a new delegated controller (static/v3/js/tooltips.js) drives an
accessible (i) info tooltip that appears on hover, keyboard focus, and tap.
A shared `help_tip` Jinja macro (partials/_macros.html) emits the trigger;
the plugin config macro and the core settings partials now surface help
text through it. Per the design, the always-visible field help paragraphs
are folded into the tooltip to declutter the forms, and the hardware/display
settings carry authored detail (default, range, recommendation).

Search: a global header search box finds settings across every settings tab
— even ones not yet opened — via a lazy client-side index built by scanning
the same field markup (static/v3/js/settings-search.js). Selecting a result
switches tabs, waits for the field to load, then scrolls to and flashes it.
A per-tab filter box hides non-matching fields on the current tab.

Plugin settings get tooltips for free by reusing each field's schema
`description`; every settings field also gets a stable `setting-<tab>-<key>`
anchor id for search navigation.

Styling uses the existing --color-* theme vars so light/dark mode both work,
and honors prefers-reduced-motion. Adds Flask render smoke tests that assert
each settings partial ships tooltips, anchors, and a filter box.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Address Codacy static-analysis findings in settings JS

Refactor the two new modules to clear the flagged patterns without any
behavior change:

- Build the search dropdown with DOM nodes + textContent instead of
  innerHTML string concatenation, removing the XSS sinks and the manual
  escapeHtml helper it needed.
- Replace numeric index access (index[i], terms[j], opts[idx],
  currentResults[i]) with array iteration methods, NodeList.item(), and
  Array.prototype.at() to clear detect-object-injection.
- Use === via a shared termsMatch() helper, optional-catch binding, and
  drop a useless initial assignment.

Verified with ESLint (eslint:recommended + eslint-plugin-security) at zero
findings and re-ran the headless-Chromium behavior test (tooltip, per-tab
filter, global search navigation) — all green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Reduce complexity of revealAncestors in settings search

Extract isNodeHidden() and revealNode() helpers so revealAncestors drops
below the cyclomatic-complexity threshold. No behavior change; verified with
the headless-Chromium test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Resolve remaining Codacy findings in settings search

- Validate plugin ids against a strict allowlist (mirroring the server's
  _SAFE_PLUGIN_ID_RE) before they can appear in a fetch path, so the request
  URL is never built from unvalidated input (Codacy: user-controlled URL).
- Document that the fetched HTML is parsed into an inert document (scripts
  never run, never inserted into the live DOM) purely to read field text for
  the search index.
- Declare block-scoped locals with const instead of var where they were
  nested inside conditionals (Codacy: var not at function root).

No behavior change; re-verified with ESLint and the headless-Chromium test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Serve the settings search index from the server as JSON

Move index building off the client so there is no client-side HTML fetching
or DOM parsing (resolves Codacy's variable-fetch and DOMParser flags on the
read-only, same-origin index build).

- Add GET /v3/settings/search-index (pages_v3.py): renders the settings
  partials server-side and extracts each field's anchor id, key, label,
  tooltip, and section with a small stdlib HTMLParser, then caches the result
  keyed on the installed-plugin set. Parsing the rendered HTML keeps anchor
  ids identical to the live DOM, so the index cannot drift.
- settings-search.js: buildIndex() now does a single fetch of the literal
  endpoint + .json(); removed the per-partial fetch loop, DOMParser, scanDoc,
  CORE_TABS, and the plugin-id allowlist. Search, keyboard nav, navigation,
  and the per-tab filter are unchanged.

Net: fewer requests and no client-side HTML parsing. Verified with a new
endpoint test in test_web_settings_ui.py (12 pass) and the headless-Chromium
test (tooltip, filter, search navigate + flash all green).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Address CodeRabbit review on settings search/tooltips

- pages_v3: include Durations tab in the search index so
  setting-durations-* fields are actually indexed
- test_web_settings_ui: assert setting-durations-clock is present in
  the search-index endpoint response
- settings-search.js: on index fetch failure, reset buildPromise
  instead of caching an empty (truthy) index so search can retry
- settings-search.js: filterScope returns null (not document) when no
  tab container matches, and the caller guards, so the per-tab filter
  can't hide fields across unrelated tabs
- settings-search.js: refresh the stale header comment to describe the
  server-side JSON index flow
- app.css: cap #settings-search-results height with overflow-y so the
  dropdown scrolls instead of overflowing small screens

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Fix stuck search dropdown; add plugin-tab nested-settings filter

Global search:
- Close the results dropdown on input blur (guarded, with a short delay)
  so it reliably dismisses when focus leaves — previously it could linger
  because the only outside-close was a document click that Alpine/HTMX
  handlers can swallow.
- Clear the query text after navigating to a result so refocusing the box
  doesn't re-open stale results.
- Also dismiss on htmx:afterSwap (tab changes / navigation).

Per-tab filter (now on plugin tabs too):
- Render the shared settings_filter box in the plugin Configuration panel.
  It auto-wires: the delegated input handler and filterScope already target
  .plugin-config-tab.
- Teach applyTabFilter to reveal matches inside collapsed nested sections
  (render_nested_section defaults them shut), hide nested-section wrappers
  with no matches, and restore the original collapsed layout when cleared
  (only re-collapsing sections the filter itself opened).
- Count a visible nested-section as content for its parent heading so the
  heading isn't hidden while a subsection below still has matches.

Adds a plugin-config render test (filter box + nested anchors + tooltips).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Close search dropdown via capture-phase outside-click

The dropdown could stay open after clicking away because the only
outside-click close was a bubble-phase document listener. The v3 UI is
one Alpine app() component full of HTMX/Alpine/widget click handlers;
when a click lands inside an element that calls stopPropagation(), the
event never bubbles to document and the close never runs.

- Replace the bubble-phase document 'click' close with a capture-phase
  'pointerdown' listener scoped to #settings-search-wrap. Capture runs
  before any bubbling stopPropagation can swallow the event, so it always
  fires; pointerdown also covers touch on the Pi screen. Clicking a result
  stays inside the wrap, so selection is unaffected.
- Guard the debounced input handler so a delayed render can't re-open the
  box after focus has left (type-then-click-away race).

Keeps the existing blur / Escape / htmx:afterSwap closes as secondary paths.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014gZxznuxw8L92FUMBN3Nqz

* Fix settings search dropdown not visually closing

.hidden has no effect in this app: app.css is a hand-picked utility
subset (no Tailwind build step) and never defines .hidden { display:
none }. openResults()/closeResults() only toggled the class, so the
dropdown stayed rendered (display: block) even once closeResults()
ran - confirmed via computed style in a headless browser. Set
style.display directly, matching the fallback already used by
revealNode()/collapseNode() elsewhere in this file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-09 09:22:24 -04:00
ChuckandClaude Opus 4.8 bd9f461f70 Add system diagnostics, power controls, and WiFi radio toggle to Tools tab (#389)
* Add system diagnostics, power controls, and WiFi radio toggle to Tools tab

Expands the web UI Tools tab with safe, purpose-built controls so users can
manage the Pi without SSHing in, instead of an arbitrary-command terminal
(the web UI has no auth and CSRF is disabled, so a shell would be unsafe).

- System Diagnostics card: renders the existing but previously-unused
  GET /api/v3/system/status endpoint (CPU, memory, temp, disk, uptime),
  with a manual refresh and a 10s poll.
- System Power section: reboot/shutdown buttons wired to the existing
  reboot_system / shutdown_system actions, behind a confirm step, with a
  dedicated powerAction() helper that treats the dropped connection as the
  expected "going offline" outcome rather than an error.
- Network Radio section: WiFi on/off toggle backed by new
  GET/POST /api/v3/wifi/radio endpoints and WiFiManager.set_wifi_radio() /
  get_wifi_radio_state(). Disabling WiFi is refused unless a wired
  connection is present (reusing the existing lockout guards), with an
  explicit force-off confirmation for advanced users.

No new privileged commands: uses nmcli radio wifi on|off (already
sudo-allowlisted) and the existing reboot/poweroff grants, so the
sudoers-alignment guard test stays green. Bluetooth toggle intentionally
omitted since the installer removes the BlueZ stack for LED timing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019X3he87Ggr1qt8y7mnFeuE

* Harden WiFi radio endpoint and diagnostics poll (review feedback)

Addresses code review feedback on the Tools tab additions:

- api_v3.py: parse the `enabled` POST field with the same string-aware
  coercion as `force`. A plain bool() cast turned {"enabled":"false"} into
  True (enabling instead of disabling) for any non-UI API caller.
- wifi_manager.set_wifi_radio() now returns a reason code alongside
  (success, message); the /wifi/radio error response includes it. The Tools
  UI only shows the force-off confirmation when reason == 'no_ethernet', so a
  genuine nmcli failure surfaces its real error instead of a misleading
  "no wired connection" prompt that would just retry into the same failure.
- tools.html: gate the 10s diagnostics poll on panel visibility
  (document.hidden / offsetParent), so switching to another tab stops the
  recurring /api/v3/system/status calls instead of churning the Pi off-screen.
  The initial load and manual Refresh remain unconditional.

Verified: {"enabled":"false"} now disables (refused w/ reason:no_ethernet),
{"enabled":"true"} enables; inline JS passes node --check; py_compile clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019X3he87Ggr1qt8y7mnFeuE

* Narrow WiFi-disable fallback except and add stack trace (review)

Addresses a CodeRabbit nitpick: the fallback handler in set_wifi_radio()'s
disable path caught bare Exception and logged without a traceback. Narrow it to
(OSError, subprocess.SubprocessError) — the errors subprocess.run realistically
raises — and log with exc_info=True for full context on the Pi. Anything
genuinely unexpected now propagates to the endpoint's outer handler (500),
matching the codebase's specific-exception convention.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019X3he87Ggr1qt8y7mnFeuE

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-09 09:22:11 -04:00
ChuckandClaude Opus 4.8 3b93024993 feat: activate dormant plugin health/metrics subsystem and surface it in the web UI (#388)
* feat(plugin-system): activate dormant plugin health & metrics subsystem

PluginManager shipped a fully-built health tracker, resource monitor and
circuit breaker that were never instantiated (health_tracker/resource_monitor
were left as None), so the circuit breaker never engaged and the existing
health/metrics API routes always returned "not available".

- DisplayController now wires a PluginHealthTracker and PluginResourceMonitor
  onto the plugin manager, enabling the circuit breaker (a repeatedly-failing
  plugin's update() is skipped after consecutive failures, then retried after
  a cooldown) and per-plugin execution-time metrics. Both persist to the
  shared cache.
- load_plugin() now validates each plugin's config against its JSON schema in
  a strictly warn/degrade-only way: a violation logs a warning and flags the
  plugin degraded in the health tracker, but never changes whether the plugin
  loads or its pass/fail behaviour. Adds PluginHealthTracker.set_degraded(),
  which never touches the circuit breaker.
- ResourceMonitor CPU/memory sampling now reuses a cached psutil.Process and
  reads cpu_percent(interval=None), so monitoring no longer blocks ~100ms per
  call on the display loop's update path.
- Fix DiskCache.get() raising TypeError for max_age=None ("never expires"),
  which silently discarded persisted plugin health/metrics on read and thus
  broke cross-process and post-restart surfacing.
- Fix two dead PluginManager helpers that called non-existent tracker methods.

Tests: new test_resource_monitor, test_plugin_health,
test_plugin_manager_schema_soft; extended test_cache_manager and
test_display_controller.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

* feat(web-ui): surface plugin health, metrics and load state

With the health/metrics subsystem now active in the display service, expose it
in the web UI (which runs as a separate process from the display loop):

- Wire a health tracker / resource monitor backed by the shared on-disk cache
  into the web process so /api/v3/plugins/health and /plugins/metrics read the
  data the display service persists.
- Build those route responses per installed plugin id (the tracker's in-memory
  view is empty in a fresh web process) so cross-process data is included.
- Add state + error_info to /plugins/installed entries so the UI can show why a
  plugin isn't running instead of just loaded:false.
- Add a "Plugin Health" panel to the Tools page (circuit status, avg/max update
  time, update count, last error) plus PluginAPI.getPluginMetrics().

Tests: route-level tests for the health/metrics endpoints in test_web_api.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

* fix(plugin-metrics): refresh cross-process health/metrics reads; type hints

Addresses CodeRabbit review on #388:

- Major: the web process's health/resource trackers cached the first persisted
  read in an in-memory dict (and the CacheManager memory tier held max_age=None
  entries indefinitely), so a long-lived web process showed the first snapshot
  and never reflected the display service's later updates. Add an opt-in
  force_reload path (get_health_summary/get_health_state/_load_health_state and
  get_metrics_summary/get_metrics) that bypasses the in-memory copy and, via a
  new memory_ttl passthrough on CacheManager.get, the cache manager's memory
  tier — so each /plugins/health and /plugins/metrics poll reads fresh persisted
  state. Default behaviour (force_reload=False) is unchanged for the display
  process and existing callers.
- Minor: DiskCache.get type hint is now Optional[int] with the None ("never
  expires") semantics documented, matching MemoryCache.get.

Tests: new force_reload staleness cases in test_plugin_health and
test_resource_monitor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UvTav268UXv44ub9K11LYq

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-09 07:54:18 -04:00
ChuckandClaude Sonnet 5 85d321cf33 fix: plugin_loader retries with --ignore-installed on apt/pip RECORD conflicts (#386)
* fix: plugin_loader retries with --ignore-installed before assuming apt package satisfies pin

install_dependencies treated any "uninstall-no-record-file" pip failure as
"dependency satisfied" and wrote the success marker without ever attempting
--ignore-installed, unlike install_dependencies_apt.py and
safe_pip_install.sh (added in #385 for the Plugin Store/first-time-install
paths). A plugin pinning a newer version of a system-managed package (e.g.
requests) would silently keep running against whatever version apt shipped,
while the marker file claimed the pinned requirement was met.

Now retries the same install with --ignore-installed on that specific
failure so pip actually lays the pinned version down (shadowing the
system-managed copy) before falling back to the prior tolerant behavior if
the retry itself fails too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

* chore: suppress Codacy finding on new retry subprocess.run call

Same generic Bandit/semgrep pattern-match on non-literal subprocess.run
argv flagged in #385's install_requirements_file, now on the new
--ignore-installed retry call added here: list-form argv (no shell=True),
sys.executable is this process's own interpreter, and requirements_file is
built internally by find_plugin_directory, never raw external input.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

* fix: tolerate a timed-out --ignore-installed retry, dedupe marker-write logic

CodeRabbit review caught a real inconsistency: if the --ignore-installed
retry itself timed out, subprocess.TimeoutExpired propagated to the outer
handler and returned False, failing plugin load — contradicting the
intended "tolerate this specific apt/pip conflict" behavior, where a mere
non-zero retry return code already returns True. Wraps the retry in its own
try/except so a timeout is logged and tolerated the same way as any other
retry failure.

Also extracts the marker-writing logic (open/write/chmod, ignoring OSError)
into _write_dependency_marker, since it was duplicated identically between
the direct-success path and the apt-conflict-retry-fallback path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

* fix: revert marker-write dedup helper, resolves CodeQL path-injection alert

CodeQL flagged _write_dependency_marker's open(marker_file, ...) as
"uncontrolled data used in path expression" (high severity) once the
marker-write logic was extracted into its own method. marker_file is
actually safe — it's built from safe_plugin_dir, which install_dependencies
sanitizes via os.path.basename() (CodeQL's own recognized py/path-injection
sanitizer, per the existing comment a few lines above) — but CodeQL's
interprocedural analysis doesn't carry that sanitized status across the new
method boundary, since the sanitizer call and the open() sink were no
longer in the same function.

This exact code produced zero CodeQL findings before the extraction (in two
duplicated inline blocks) and is unchanged in what data reaches it — only
its location moved. Reverting the extraction (keeping CodeRabbit's other,
independent timeout-handling fix) restores the previously-clean shape rather
than trying to convince the analyzer's cross-function taint tracking that a
refactor changed nothing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-08 11:19:37 -04:00
ChuckandClaude Sonnet 5 63a233f3ed fix: dependency installation gaps in Plugin Store and first-time install (#385)
* fix: install plugin dependencies through root-visible installer in Plugin Store

install_plugin/update_plugin (store_manager.py) installed requirements.txt
with a bare `pip3` off PATH, bypassing the root-visible installer added in
#380 for the "Reinstall Plugin Deps" button. Two bugs stacked: (1) `pip3`
can resolve to a different Python install than the one that actually runs
ledmatrix.service, and (2) even when it resolves correctly, ledmatrix-web
runs as a non-root user so the package lands in that user's local
site-packages, invisible to root-run ledmatrix.service. Either way the
install reports success and writes the .dependencies_installed hash marker,
so plugin_loader's own (correct) install-on-load path skips reinstalling —
leaving the dependency permanently missing until a user finds and clicks
the separate "Reinstall Plugin Deps" tool. This is why users kept hitting
"No module named 'astral'" for the weather plugin even after installing it
from the Store.

Extracts the sudo-wrapper-then-fallback install logic from api_v3.py's
_pip_install_requirements into src/common/permission_utils.py as
install_requirements_file, and routes store_manager.py's dependency
installation through it so the automatic Store install/update path now
matches the manual "Reinstall Plugin Deps" path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

* chore: suppress Codacy false-positive on subprocess.run in install_requirements_file

Codacy's generic subprocess-security rule (Bandit B603 equivalent) flagged
the pip/sudo subprocess.run calls in install_requirements_file for lacking a
"static string argument" — the standard pattern-based flag for any
subprocess.run() call with a variable in its argv list. Both calls use
list-form argv (no shell=True, so no shell-injection surface), and the only
dynamic value is req_file, a Path built internally by callers rather than
raw external input; safe_pip_install.sh independently re-validates it before
installing anything as root. Suppresses with inline `# nosec B603` comments
matching this codebase's existing convention (see permission_utils.py's own
PROTECTED_SYSTEM_DIRECTORIES, display_manager.py, sync_manager.py, etc.).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

* fix: first-time install script fails on apt-managed requests package

web_interface/requirements.txt and requirements.txt both pin
requests>=2.33.0,<3.0.0, but Raspberry Pi OS ships an apt-managed
python3-requests with no pip RECORD file. Upgrading it via plain
`pip install` aborts with "uninstall-no-record-file" because pip refuses to
uninstall a package it has no record of, in place — which is exactly the
"Some web interface dependencies failed to install" warning first-time
install hits.

scripts/install_dependencies_apt.py and scripts/fix_perms/safe_pip_install.sh
already work around this with --ignore-installed (lets pip lay the new
version down in /usr/local, shadowing the apt copy, instead of trying to
remove it first). first_time_install.sh's own direct pip invocations —
the per-package requirements.txt loop, the web_interface/requirements.txt
install, and the requirements_web_v2.txt fallback — didn't have it. Adds
--ignore-installed to all three so first-time install no longer fails on
this well-known apt/pip conflict.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

* chore: add nosemgrep to subprocess.run calls Codacy still flagged

The prior # nosec B603 comments suppressed Bandit's check but Codacy's
semgrep-based rule ("subprocess function 'run' without a static string")
kept flagging the same two lines as a critical security issue even after
that fix landed. install_dependencies_apt.py's _run() already needed both
tags together (# nosec B603 B607 ... # nosemgrep) for the identical
subprocess.run pattern, so apply the same double suppression here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

* fix: add --ignore-installed to install_requirements_file fallback path

CodeRabbit review caught this (confirming a gap already flagged in
conversation): the non-sudo fallback pip install in install_requirements_file
was missing --ignore-installed, unlike the sudo-wrapper branch and
safe_pip_install.sh. Without it, the same apt/pip RECORD-file conflict this
PR fixes elsewhere (first_time_install.sh, install_dependencies_apt.py) could
still hit installs that fall back to this path (e.g. a plugin's
requirements.txt on a host where safe_pip_install.sh isn't set up yet).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X1NnDduw53kTe67i5zWwYx

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-08 09:59:12 -04:00
ChuckandClaude 7a9d01342a Fix inconsistent Vegas scroll transition gaps caused by plugin-baked padding (#384)
* Strip plugin-baked scroll padding when capturing content for Vegas mode

Plugins that build their own ticker image via ScrollHelper.create_scrolling_image()
(or that manually pad both ends for a clean standalone loop) carry a solid-black
margin up to display_width wide on one or both edges. Vegas mode already adds its
own configurable gap around every item, so leaving that margin in place stacked an
extra, uncontrolled blank stretch on top of separator_width for whichever plugin
took the ScrollHelper-capture path — producing inconsistent transition gaps between
modules compared to plugins that provide content natively via get_vegas_content().

_get_scroll_helper_content() now detects and crops any such margin before handing
the image to the Vegas render pipeline, so every plugin's gap is governed solely by
vegas_scroll.separator_width regardless of which capture path produced its content.

* Address CodeRabbit nitpick: warn on double-edge padding crop, add unit tests

Logging a double-edge match at warning level (vs. info for a single edge)
makes it easy to spot an unexpected crop in the field, since two edges
matching at once is a much stronger signal of genuine baked-in padding than
one edge coinciding with real all-black content.

Also adds test/test_vegas_plugin_adapter.py covering _strip_scroll_padding's
branch logic: leading-only, trailing-only, both-edges, no-match, degenerate
all-black, missing/undersized display_width, and the info-vs-warning log level.

* Add type hints and docstring to test _solid helper (CodeRabbit nitpick)

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-08 09:43:13 -04:00
ChuckandClaude Sonnet 5 9b2f02681d feat(web-ui): detect and surface Raspberry Pi under-voltage/throttling (#383)
* feat(web-ui): detect and surface Raspberry Pi under-voltage/throttling

Adds a vcgencmd get_throttled check to the system-status SSE stream and
surfaces it in the web UI:

- A header badge (next to CPU/Memory/Temp) that stays hidden when healthy,
  turns red when under-voltage/throttling is happening right now, and
  yellow if it happened earlier this session but has since cleared.
- A dismissible top banner (same pattern as the update-available banner)
  that appears while under-voltage/throttling is actively occurring, with
  guidance to check the power supply. Re-appears on a fresh occurrence
  even if a previous one was dismissed.
- A "Power Supply" card on the Overview tab alongside CPU/Memory/Temp/
  Display Status.

Motivated by a real device showing intermittent brightness flicker that
turned out to be ~1 under-voltage event every 30-90s (visible live via
`vcgencmd get_throttled` and dmesg's "Undervoltage detected!" messages) --
there was no way to see this from the web UI, only by SSHing in.

Returns None on non-Pi platforms (no vcgencmd on PATH), matching the
existing guard pattern used for the CPU temperature read.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* feat(web-ui): add power supply diagnostics detail to Tools tab

The header badge/banner/Overview card added in the previous commit only
show a collapsed "is it bad right now" signal. This adds a "Power Supply"
section to the Tools tab with the full 8-flag breakdown (under-voltage,
throttled, freq-capped, soft-temp-limit -- each split into "right now" vs
"occurred since boot") for actually troubleshooting a recurring issue,
plus a pointer to the README's power supply sizing guidance when something
is or was flagged.

Reuses the existing stats SSE stream (window.statsSource) rather than
adding a new endpoint -- the same payload already drives the header/
banner/Overview card, so this just listens for it too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web-ui): address review findings on power supply monitoring

- Overview card left its "--" placeholder forever on non-Pi platforms since
  the update only handled a truthy data.power. Now explicitly renders "Not
  available" with a neutral icon/color for that case.
- vcgencmd failures logged at debug, invisible in default remote log
  output. Bumped to warning to match the nearby systemctl failure logging.
- The banner/badge/card and the Tools summary line only looked at
  under_voltage_now/throttled_now (+ occurred), silently ignoring
  freq_capped_now/occurred and soft_temp_limit_now/occurred from
  _get_power_status() -- a Pi that's actively soft-thermal-limited or
  frequency-capped showed a green "OK" everywhere except the detailed
  flag table buried in Tools. All four surfaces now fold all four "now"/
  "occurred" flags into the same active/occurred state.
- The banner text was hardcoded to "Under-voltage detected..." even when
  the actual active condition was throttling/freq-capping/thermal limiting.
  Added _activePowerConditionLabels() (shared, non-module global scope) to
  build the banner/tooltip text from whichever flags are actually set.

Skipped: TTL-caching _get_power_status() to avoid "multiplying forks
across browser tabs" -- that premise doesn't hold against this codebase.
_StreamBroadcaster (its own docstring says as much) already runs exactly
one shared generator per tick regardless of client count, identical to
the uncached cpu_temp file-read two lines above it; there's nothing to
multiply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(web-ui): address second round of review findings on power monitoring

- updatePowerStatus's falsy-power branch only hid power-stat, leaving
  power-warning-banner visible with stale text if _get_power_status()
  fails transiently. Now hides the banner too and resets the dismissed
  flag, same as the "not active" branch.
- The Tools tab's status badge (and the pre-existing dirty/clean badge
  right next to it) build class names like bg-${color}-100/text-${color}-800
  at runtime. This project hand-rolls its own Tailwind-named utility
  classes in app.css rather than running a real Tailwind build, and the
  light-mode base rules for bg-red-100/bg-yellow-100/bg-green-100/
  text-red-800/text-yellow-800/text-green-800 were simply never defined --
  only some had dark-mode overrides, which are no-ops without a base rule
  in light mode. Added the missing light-mode bases plus the two missing
  dark-mode overrides (bg-yellow-100/bg-green-100), fixing both badges.
- The "occurred earlier" tooltip was a hardcoded string regardless of
  which flag(s) actually fired. Generalized _activePowerConditionLabels()
  to take a suffix ('_now' or '_occurred') and reused it for both tooltips.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* refactor(web-ui): move Power Supply status off the Overview tab

The Overview tab's "Power Supply" stat card duplicated what the Tools
tab's diagnostics section already shows (summary badge + full flag
breakdown), so drop the card and its now-dead JS rather than keep two
copies in sync. The header badge and warning banner (visible on every
page) are unaffected -- only the Overview-tab card is removed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-07 09:35:55 -04:00
7a6bad29fe feat(display-controller): hot-reload plugin enable/disable without a restart (#374)
* feat(display-controller): hot-reload plugin enable/disable without a restart

Enabling or disabling a plugin in config previously required restarting the
display service: the plugin list and available_modes were built once at init
and the run loop never revisited them. (Per-plugin config *values* already
hot-reloaded; only the enabled set was restart-only.)

Now the controller reconciles its running plugins against the config's enabled
set whenever that set changes:

- The ConfigService watcher thread only sets a `_pending_plugin_reconcile`
  flag (via a cheap enabled-set diff). It never mutates loop state.
- The run loop applies the reconcile on its own thread (top of each
  iteration, deferred while on-demand is active), so loading/unloading and
  rebuilding available_modes can't race with rendering.
- `_reconcile_enabled_plugins` diffs desired vs running plugins, unloads the
  removed ones (cleanup + on_disable + config-unsubscribe via the new
  `_unregister_plugin`) and loads the added ones, then clamps the rotation
  index so the current mode stays valid.

The per-plugin registration done at startup is extracted into
`_register_loaded_plugin` and reused by the live-enable path so both build
identical state. Extracting it also fixes a latent late-binding bug: the
per-plugin config-change callbacks were closures over the loop variable, so
every plugin's callback targeted the last-loaded instance; each now binds its
own id/instance.

Adds test/test_display_controller_plugin_toggle.py covering live enable,
live disable, index clamping, no-op when unchanged, and the enabled-set diff.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(display-controller): don't exit on empty available_modes, guard rotation modulo

Hot-reload means available_modes can legitimately be empty at startup (no
plugins enabled yet) and become non-empty later via the web UI, or vice
versa mid-run. Fix four issues found reviewing this PR:

- run() exited the process entirely when available_modes was empty at
  startup instead of idling, permanently defeating the point of live
  enable/disable for anyone who starts with zero plugins enabled.
- The mode-rotation step divided by len(available_modes) unconditionally,
  raising ZeroDivisionError if the last enabled plugin is disabled between
  frames.
- _reconcile_enabled_plugins() called .get('enabled', False) on a config
  section without checking it was a dict first, raising AttributeError on
  a malformed config value.
- Minor: pop the config-change callback only after attempting to
  unsubscribe it, and log the exception in the config-read fallback
  instead of swallowing it silently.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix(display-controller): address review findings on the hot-reload PR

- Idle-wait tick was a fixed 30s sleep, delaying pickup of a plugin
  enabled via the web UI while no modes were active. Shortened to ~1s so
  it's roughly as responsive as the per-frame check once modes exist.
- _unregister_plugin popped the config-change callback from
  _plugin_config_callbacks even when config_service.unsubscribe() raised,
  losing the only reference to it. Now only pops on a successful
  unsubscribe.
- _pending_plugin_reconcile was cleared before _reconcile_enabled_plugins()
  ran, so a retryable failure (e.g. plugin discovery erroring) silently
  dropped the enable/disable request. _reconcile_enabled_plugins() now
  returns True/False and the caller only clears the flag on True.
- Added a warning log for the malformed-config case (a plugin's config
  section present but not a dict) so it's actually visible, and updated
  the existing test to assert it via caplog.

Left the broad `except Exception` around config_service.unsubscribe() as
Exception -- the current implementation is a simple lock+dict/list op that
doesn't document or realistically raise a narrower type, so this is a
defensive catch-all, not user error handling.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: ChuckBuilds <charlesmynard@gmail.com>
2026-07-06 17:37:01 -04:00
Chuck bea00448d3 update mlb logos (#382)
update mlb logos

Signed-off-by: Chuck <33324927+ChuckBuilds@users.noreply.github.com>
2026-07-06 08:55:43 -04:00
ChuckandClaude Sonnet 5 deaa3d7a98 fix: array-table boolean columns always saved as unchecked (#381)
Reported bug: countdown plugin always shows "No Active" even with a
countdown configured and enabled — because every save silently unchecked
it. Root cause: the boolean column's hidden-sentinel input (needed so
unchecked checkboxes still submit "false", since browsers omit unchecked
checkboxes entirely) was hardcoded to value="false" and shared the same
`name` as the checkbox, with no sync between them. Whichever of the two
same-named inputs the save request's form-collection happens to prefer,
the hidden's stale "false" could silently override an actually-checked
box on every save, independent of what the user set.

Fixed in both places this pattern is rendered:
- plugin_config.html: the initial server-rendered row for every array-table
  boolean column (affects countdown's per-countdown "enabled" and the
  custom-feeds widget's per-feed "enabled") — now sets the hidden's initial
  value from the actual data, and syncs it via onchange on every toggle.
- array-table.js: the client-side row renderer used when adding/rebuilding
  rows in the browser — same fix, keeping both inputs in sync at render
  time and via a change listener.

Verified other checkbox/hidden-input patterns in plugin_config.html and
confirmed they're unrelated/already-safe: the default single-checkbox
renderer (no hidden pair), checkbox-group (array-bracket names with its
own explicit sync), the array-table advanced-props modal editor (single
hidden per name, explicitly written on Save), the top-level plugin
enable/disable toggle (separate immediate hx-post), and the no-schema
fallback checkbox. custom-feeds.js's own client-side row renderer never
creates a hidden sentinel at all, so it isn't affected either.

Verified the fix's rendered output directly with an isolated Jinja render
of the exact snippet: hidden and checkbox now agree ("true"/checked or
"false"/unchecked) for both states, where before the hidden was always
"false" regardless of the actual value.


Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-06 08:05:06 -04:00
ChuckandClaude Sonnet 5 cbb8ec41e8 Fix: plugin/base requirements installed as web-user, invisible to root-run display service (#380)
* fix: install plugin/base requirements as root so ledmatrix.service can see them

ledmatrix-web.service runs as a non-root user, so "Reinstall plugin
requirements" installed packages into that user's ~/.local site-packages.
ledmatrix.service (the actual display, which loads and runs plugin code)
runs as root and can't see another user's user-site packages, so plugins
with dependencies not already present system-wide would silently fail at
runtime with ModuleNotFoundError even after a "successful" reinstall.
Reproduced and fixed live against a real device (weather plugin's astral
dependency, used for moon-phase data): confirmed the exact failure
("No module named 'astral'" on every almanac cycle) and confirmed it's
gone after this fix.

Adds scripts/fix_perms/safe_pip_install.sh, a root-owned wrapper (mirroring
the existing safe_plugin_rm.sh pattern) that validates the target is
requirements.txt at the project root or under plugin-repos/ or plugins/
before running pip install as root. configure_web_sudo.sh provisions a
narrowly-scoped sudoers rule for it. api_v3.py's install_base_requirements
and install_plugin_requirements actions now use it via `sudo -n`, falling
back to today's current-user-only install (with an explanatory note) if
the wrapper isn't set up yet, so existing installs don't regress.

Also uses --ignore-installed in the wrapper: root's site-packages often has
apt/dpkg-managed copies of common libraries (requests, etc.) with no pip
RECORD file, which pip refuses to upgrade in place and aborts the *entire*
requirements.txt install over — discovered this while testing the fix live,
since a plugin's other already-satisfied-for-the-web-user dependencies had
never actually been attempted as root before.

Also fixes a pre-existing bug in configure_web_sudo.sh where the
display_controller.py/start_display.sh/stop_display.sh sudoers entries used
PROJECT_DIR (scripts/install/, where this script lives) instead of
PROJECT_ROOT (where those files actually live) — visible as the script's
own "File access test" self-check failing. Verified fixed live.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix: invoke safe_pip_install.sh via explicit bash, matching sudoers rule

CodeRabbit caught this on review: the sudoers rule configure_web_sudo.sh
provisions is scoped to "$BASH_PATH $SAFE_PIP_INSTALL_PATH *" (matching
the existing safe_plugin_rm.sh precedent in
src/common/permission_utils.py), but _pip_install_requirements() called
`sudo -n <wrapper> <req_file>` directly, relying on the script's shebang
instead of an explicit bash prefix. sudo matches the literal command line,
so this never matched the allowlisted rule on an install with only the
specific sudoers entries this script provisions — it silently fell back
to the non-root install path every time, which is the exact bug this PR
set out to fix.

This wasn't caught by live testing on ledpi.local because that device
also has a broader, non-standard "NOPASSWD: ALL" grant which masked the
mismatch. Confirmed the fix is correct by reading sudo's documented
command-matching semantics and mirroring the already-proven-working
bash-prefix pattern from permission_utils.py's safe_plugin_rm.sh call
exactly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

* fix: harden safe_pip_install.sh invocation against bash-path drift + address CodeRabbit nitpicks

CodeRabbit follow-up findings on the bash-prefix fix (7558aaab):

1. (Actionable) shutil.which('bash') at runtime could in principle resolve
   to a different absolute path than configure_web_sudo.sh's `command -v
   bash`, which is resolved once at setup time and baked into the static
   sudoers file as a literal string — sudo requires an exact match. Now
   tries /usr/bin/bash and /bin/bash (the standard Debian/Raspberry Pi OS
   locations, matching what the setup script virtually always produces)
   before falling back to this process's own PATH resolution, so a
   divergence in just one of them doesn't break the install.

2. (Nitpick) Any nonzero returncode was treated as "sudo denied", so a
   real pip failure (bad package, build error) would trigger a pointless
   duplicate non-root install attempt and a misleading error message.
   Now distinguishes "sudo -n rejected this exact command line" from
   "sudo ran it but the command itself failed" via sudo's own diagnostic
   text, and surfaces genuine failures immediately without retrying other
   bash candidates or falling back.

3. (Nitpick) Added structured logging for every fallback/failure path
   (wrapper missing, sudo denied, real install failure), previously only
   visible via the returned stdout note — needed for remote debugging on
   a headless Pi.

Verified: function-level smoke test confirms a real failure (this sandbox's
system python3 lacking pip) is now correctly classified as non-denial and
returned immediately without retrying candidates or double-installing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-06 08:04:49 -04:00
c6ce332d49 fix(web): restore missing brace in Tools tab HTMX-fallback path (#378)
The `else if (++tries > 100)` block added by #373 was missing its
closing `}`, leaving the setInterval arrow function syntactically
unclosed. This caused a JS parse error that silenced the entire
1400-line script block — including the EventSource setup — so the
connection-status indicator never left its default "Disconnected"
state for all users after updating.

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-30 10:05:33 -04:00
Chuck 8e5f66501a Add Claude Code GitHub Workflow (#377)
* "Claude PR Assistant workflow"

* "Claude Code Review workflow"
2026-06-29 14:30:41 -04:00
639e1c3a93 fix(web): repair news ticker custom-feeds save for JSON path (#376)
* feat(install): surface root cause of web dependency install failures

install_dependencies_apt.py previously reported only which packages
failed, not why - the actual apt/pip error was discarded (apt) or
could scroll out of the on_error log tail (pip), leaving "Step 7:
Install web interface dependencies (line 915)" as the only visible
detail.

Capture command output for each install attempt and print a compact
DEPENDENCY INSTALLATION FAILURES summary with the last lines of error
output per package. Also run the installer with `python3 -u` for
real-time, correctly-ordered logging, and widen the on_error tail from
50 to 100 lines so the summary isn't cut off.

* fix(web): repair news ticker custom-feeds save for JSON path

The JS dotToNested() helper converts indexed form fields like
feeds.custom_feeds.0.name into a dict {'0': {name:...}} rather than a
proper array. The form-data path already had fix_array_structures() to
convert those dicts back to arrays before schema validation, but the
JSON path (used by all web-UI saves) never ran that fix, so saving any
custom feed produced a schema validation error: "Expected type array,
got object".

Add _fix_json_arrays() immediately after schema loading on the JSON
path, mirroring the existing fix_array_structures() logic.

Also fix custom-feeds.js getValue() to omit the logo key entirely when
no logo is present instead of returning logo:null, which would fail
schema validation (logo expects type object).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(install,tools): address PR 376 review findings

- first_time_install.sh: add _clone_rpi_rgb() wrapper so retry() cleans
  up any partial rpi-rgb-led-matrix-master dir before each clone attempt
- first_time_install.sh: use apt-get -o DPkg::Lock::Timeout=180 so apt
  handles lock contention natively instead of relying solely on flock TOCTOU check
- install_dependencies_apt.py: pass DPkg::Lock::Timeout=180 to apt-get
  install to avoid failing when unattended-upgrades holds the lock
- install_dependencies_apt.py: add type annotations to all public helpers
- api_v3.py: fix install_plugin_requirements to read plugin_manager from
  api_v3 blueprint attribute instead of the always-None module variable
- tools.html: loadGitInfo() now checks r.ok before parsing JSON and
  surfaces d.status === 'error' with the server's message in the panel

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools,api): address three additional review findings

- api_v3.py install_plugin_requirements: replace hardcoded plugin-repos
  fallback with config-driven resolution (plugin_system.plugins_directory),
  matching the pattern used elsewhere in the module
- api_v3.py _fix_json_arrays: recurse into converted and existing array
  elements when items.type is object, so nested numeric-keyed dicts inside
  array items are also normalized
- tools.html toolsAction: check r.ok before r.json() and recover
  gracefully from non-JSON error bodies (HTML 500 pages), consistent
  with the existing loadGitInfo guard

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 14:21:05 -04:00
6096a22c3d feat(web): add Tools tab and row address type display setting (#373)
* feat(web): add Tools tab and row address type setting

Adds a Tools/Utilities tab to the web interface with one-click
maintenance buttons that previously required SSH:
- Git status panel (branch, dirty state, recent commits)
- Pull latest (rebase) and force reset to origin/main
- Reinstall base requirements (pip, with output)
- Reinstall per-plugin requirements (pass/fail per plugin)
- Clear __pycache__ directories
- Quick-access restart for display and web services

Also exposes the hzeller row_address_type option (0–4) in the
Display settings tab. The backend already read this value from
config; the UI, API field list, and validation were missing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): address code review findings

- Add _GIT = shutil.which('git') alongside _SUDO/_JOURNALCTL; return
  503 in force_git_reset and get_git_info if git is unavailable
- Check git branch/status returncodes in get_git_info(); return a clear
  500 error instead of silently treating a failed run as a clean repo
- Cap pip stdout+stderr at 50 KB via _truncate_output() helper to
  avoid OOM on verbose dependency resolution or build failures
- Scrub embedded HTTPS credentials from remote_url via
  _scrub_git_remote_url() using urllib.parse before returning to UI
- Fix clear_pycache to track and report failed deletions separately
  instead of counting them as successes (removed ignore_errors=True,
  wrapped in try/except OSError)

Skipped: plugin_manager-vs-api_v3.plugin_manager (api_v3 is the
Blueprint object; accessing .plugin_manager on it would fail — module-
level variable is the correct pattern used throughout this blueprint);
pages_v3 broad-except (identical to every other _load_*_partial in the
file); base.html HTMX fallback (loadTabContent handles all tabs
generically; named fallbacks only exist for tabs needing JS re-init);
tools.html auth (pre-existing architectural decision — reboot/shutdown
on the same endpoint are also unauthenticated).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): resolve remaining PR review comments

- api_v3: use getattr(api_v3, 'plugin_manager', None) instead of the
  module-level plugin_manager (always None); app.py sets the blueprint
  attribute, not the module global, so the fallback to plugin-repos was
  always taken
- pages_v3: replace broad except Exception in _load_tools_partial with
  specific TemplateNotFound / OSError handlers and add [Pages V3][Tools]
  context prefix to log messages and error responses for easier Pi
  debugging
- base.html: add Tools tab branch to the HTMX-unavailable fallback block
  in loadTabContent so the tab loads gracefully via direct fetch if HTMX
  never initialises

Skipped: auth on execute_system_action — pre-existing app-wide design;
reboot/shutdown and all other system actions share the same exposure.
An app-level auth layer is the correct fix and is out of scope here.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): resolve second-pass review findings

- Wrap per-plugin subprocess.run in try/except TimeoutExpired/OSError so
  one plugin's failure appends a result entry and continues the loop
  rather than collapsing the whole batch into a 500
- Validate double_sided_copies divisibility against chain_length
  (horizontal axis) or parallel (vertical axis) after the range check;
  reads effective axis from the current request or stored config
- Exclude double_sided_fields from the generic key-merge loop so
  double_sided_enabled/copies/axis are never written as root-level keys
- Fix tools.html copy: "then restores the stash" removed — git_pull
  stashes changes but never pops them
- Check r.ok and d.status in loadGitInfo before building the panel;
  backend error messages now surface instead of silently showing a
  false-clean state

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(tools-tab): don't expose filesystem paths in OSError messages

CodeQL flagged str(exc) flowing into the JSON response for the
install_plugin_requirements action. Use exc.strerror instead, which
gives the OS error description ("No such file or directory",
"Permission denied") without the internal filesystem path.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-29 12:19:54 -04:00
Ron PierceandClaude Opus 4.8 fefc2d44a2 feat(display): double-sided mode — mirror one screen across the panel chain (#375)
* feat(display): add double-sided mode to mirror one screen across the panel chain

Renders a plugin once at a logical (per-screen) size, then tiles the
rendered frame across the full physical chain so two (or more) panels show
identical content. A 128x32 chain configured with 2 copies drives two 64x32
screens; vertical axis splits parallel outputs instead of the chain.

Plugins size themselves from matrix.width/height, so a thin _LogicalMatrix
proxy reports the logical size while delegating every real operation
(CreateFrameCanvas, SwapOnVSync, brightness, Clear) to the physical matrix —
no plugin changes required. Duplication is a single PIL paste per copy in
update_display(), so render cost is unchanged.

Config: display.double_sided { enabled, copies, axis }. Invalid config
(non-divisible dimension, bad axis/copies) logs a warning and falls back to
single-screen rather than failing to light up.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(web): expose double-sided display config in the settings UI

Adds a Double-Sided Display section to the Display settings page (enabled
checkbox, copies, horizontal/vertical axis) and wires the save handler to
persist it under display.double_sided. Validates copies (2-8) and axis,
returning 400 on bad input; an omitted checkbox is saved as disabled.
Like the other hardware fields, changes take effect after a display restart.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(display): add type hints and docstrings to double-sided proxy

Addresses CodeRabbit nits: sort _LogicalMatrix.__slots__ (Ruff RUF023),
annotate the proxy's __init__/properties/dunders and _resolve_double_sided's
return type, and add docstrings to the property/dunder methods.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 15:24:53 -04:00
Ron PierceandClaude Opus 4.8 d297dd6217 feat(display-controller): round-robin between simultaneous live-priority games (#372)
_check_live_priority() was stateless first-match-wins: it returned the
first plugin in registration order with live content, and the post-dwell
hold pinned the carousel to it, so when two games were live at once (e.g.
a baseball game and a soccer match) the second never showed until the
first ended.

Add _collect_live_modes() (all currently-live modes, deduped, in
registration order) and give _check_live_priority an 'advance' flag. The
main rotation calls it with advance=True, which returns the live mode
after the one currently shown -- using current_display_mode as the cursor
-- so each dwell advances to the next live game and they take turns. The
Vegas coordinator and the vegas-active check keep the default
non-advancing peek (advance=False), so they only report whether any game
is live without spinning the cursor. should_rotate and _apply_live_priority
are unchanged; a single live game still holds as before.

Adds regression tests to TestDisplayControllerLivePriority.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 14:07:21 -04:00
Ron PierceandClaude Opus 4.8 974d7ea57a fix(install): avoid apt-package uninstall failure during web dep install (#371)
* fix(install): avoid apt-package uninstall failure on web deps

On a fresh Pi install, requests is installed via apt (python3-requests),
which ships no pip RECORD file. When pip later installs
google-api-python-client, its dependency tree pulls a newer requests and
attempts to uninstall the apt copy, failing with "uninstall-no-record-file"
and aborting the whole install at step 7 (web interface dependencies).

Add --ignore-installed to install_via_pip so pip lays the new version down
in /usr/local (shadowing the apt copy) instead of trying to remove an
apt-managed package. This resolves the failure for any transitive
dependency pip needs to upgrade over an apt-installed package, not just
requests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(install): also pass --ignore-installed for local rgbmatrix install

Keeps the rgbmatrix pip install consistent with install_via_pip so it
won't fail trying to uninstall an apt-managed dependency either.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 14:06:56 -04:00
Ron PierceandClaude Opus 4.8 ab0cfd2362 fix(web): preserve dotted schema keys when saving plugin config (#370)
The plugin config form posts form-data with dot-notation paths
(e.g. "leagues.fifa.world.enabled"). _get_schema_property and
_set_nested_value split those paths on every dot, so a schema key that
itself contains a dot (soccer league keys like "fifa.world", "eng.1")
was mistaken for nested "fifa" -> "world" objects. Per-league edits
(enable, favorite_teams, nested booleans) were written to a fabricated
"leagues.fifa.world" branch while the real league object was never
updated, so saves silently dropped the change and produced a
byte-identical config.

Both helpers now greedily match the longest path segment that exists in
the schema (_get_schema_property) or the config being updated
(_set_nested_value), mirroring the frontend's dotted-key handling.

Adds regression tests covering schema lookup, value typing, and writes
under dotted league keys, plus a guard that plain nested paths still work.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 14:06:32 -04:00
Ron PierceandClaude Opus 4.8 d22d0a3754 fix(plugins): stop core updates from resurrecting uninstalled built-in plugins (#368)
* fix(plugins): stop core updates from resurrecting uninstalled built-in plugins

Built-in plugins (e.g. web-ui-info, starlark-apps) are committed into the
repo under plugin-repos/. When a user uninstalls one, a subsequent core
`git pull` update restores the committed files, so the plugin reappears on
every update. The update endpoint stashes the deletion and never pops it,
and `git pull` faithfully restores any committed file whose deletion was
never committed — so excluding plugin-repos/ from the stash can't fix this
(it would only make `git pull --rebase` fail on a dirty tree).

Add a persistent uninstall registry (config/uninstalled_plugins.json,
gitignored) that survives restarts, unlike the existing in-memory tombstone:

- Uninstall records the plugin id; install clears it.
- purge_uninstalled_plugins() re-removes any recorded plugin whose directory
  reappears on disk; called after a successful git-pull update and at web
  startup (covers manual `git pull` on the Pi too).
- The state reconciler also refuses to auto-repair a persistently
  uninstalled plugin.

Wires up mark_recently_uninstalled in the uninstall flow (previously only
referenced by tests) via the new persistent record.

Adds regression tests covering record/forget/purge lifecycle, persistence
across manager instances, and corrupt-registry tolerance.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(plugins): validate uninstall-registry ids and lock registry writes

Address review feedback on the persistent uninstall registry:

- Critical: validate plugin ids on read/record and add a containment guard
  in purge_uninstalled_plugins. A corrupt or hand-edited registry entry of
  "" resolves to the plugins root, so purge could have deleted every plugin;
  traversal ids ("..", "../x") could target paths outside the root. Invalid
  ids are now dropped on read, refused on record, and never removed unless
  the path is a direct child of the plugins directory.
- Major: guard record/forget read-modify-write with a lock so concurrent
  install/uninstall requests can't lose updates.
- Minor: narrow the startup and post-update purge exception handlers from
  bare Exception to (OSError, RuntimeError).

Adds regression tests for empty-id, traversal-id, and invalid-record cases.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 18:18:28 -04:00
ChuckandChuck 5beef0aa01 Improve first-time install error diagnostics and resilience (#369)
* fix(install): don't let outer ERR trap mask first_time_install.sh failures

set +e alone doesn't suppress bash's ERR trap, so any non-zero exit from
first_time_install.sh inside the one-shot installer immediately triggered
the outer on_error handler with a generic "Main installation, line 370"
message — before the script could report the real exit code or point to
logs/. Suspend the trap for that block so the existing if/else handling
runs instead.

* feat(install): surface root cause of web dependency install failures

install_dependencies_apt.py previously reported only which packages
failed, not why - the actual apt/pip error was discarded (apt) or
could scroll out of the on_error log tail (pip), leaving "Step 7:
Install web interface dependencies (line 915)" as the only visible
detail.

Capture command output for each install attempt and print a compact
DEPENDENCY INSTALLATION FAILURES summary with the last lines of error
output per package. Also run the installer with `python3 -u` for
real-time, correctly-ordered logging, and widen the on_error tail from
50 to 100 lines so the summary isn't cut off.

* feat(install): harden first-time install against common Pi failure modes

- wait_for_apt_lock: apt_update/apt_install now wait (up to 3min) for
  unattended-upgrades to release the dpkg lock instead of failing
  outright with "Command failed after 3 attempts" right after first boot.
- check_disk_space: new pre-flight check (Step 1) so a full SD card fails
  fast with a clear message instead of a cryptic mid-build error.
- Step 6: wrap rpi-rgb-led-matrix git clone/submodule operations in retry
  for resilience to transient network issues.
- Step 6: capture `pip install .` build output and print the last 50
  lines on failure, so the actual cmake/compiler error is visible instead
  of just "Failed to install rpi-rgb-led-matrix Python package".

* fix(install): bound subprocess output and dedupe apt update in dependency installer

Address coderabbitai review on PR #369:
- _run() now streams combined stdout/stderr to a temp file and returns
  only the last ERROR_TAIL_LINES lines, instead of buffering full
  output in memory (Codacy also flagged the previous capture_output
  call as a subprocess-without-static-string security issue; the new
  call is annotated as safe since cmd is built from hardcoded args).
- `apt update` now runs once in main() instead of once per package
  needing an apt fallback.

* fix(install): suppress remaining Codacy subprocess false-positive

Codacy's Semgrep-based check still flagged the cmd-built subprocess.run
call as "without a static string" even with the Bandit nosec applied.
Add a nosemgrep marker alongside it - cmd is always a hardcoded
apt/pip argument list, never user input.

* fix(install): correctly detect already-installed dateutil/websocket-client

Address remaining coderabbitai findings on PR #369:
- check_package_installed() did __import__(package_name) directly, but
  python-dateutil and websocket-client import as dateutil/websocket. Both
  always failed the "already installed" check and were reinstalled on
  every run. Add an IMPORT_NAME_MAP for the mismatched names.
- _run() still read the entire temp file into memory before slicing the
  tail. Stream it line-by-line into a deque(maxlen=ERROR_TAIL_LINES)
  instead so memory use stays bounded for very chatty commands.

---------

Co-authored-by: Chuck <chuck@example.com>
2026-06-11 18:12:35 -04:00
cf28a8c0d5 fix(display): restore early-continue guard for mid-loop mode changes (#367)
When the display loop breaks early because current_display_mode changed
(on-demand activation, live priority, etc.), it would fall through to the
"honour minimum duration" sleep for the *previous* mode — blocking for up
to that mode's full display_duration (default 30s) without polling
on-demand requests or re-checking the mode. New modes could sit unrendered
for up to 30s, or get clobbered by a queued stop request before ever
displaying.

This guard was added in #298 to fix #196 (live priority not interrupting
long display durations) and was accidentally dropped in #330 as collateral
damage of an unrelated time.monotonic() -> time.time() cleanup in the same
diff hunk. Restoring it fixes both the original #196 regression and a new
symptom found via the on-air MQTT plugin, where ON/OFF toggles could be
delayed by up to 30s or missed entirely depending on timing within the
previous mode's display cycle.

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 10:16:30 -04:00
ChuckandChuck a06682981c fix(web): allow up to 24 panels in chain length config (#366)
Raises the Chain Length input's max from 8 to 24 to support longer
LED panel strings.

Co-authored-by: Chuck <chuck@example.com>
2026-06-09 21:31:07 -04:00
Ron PierceandClaude Opus 4.8 bc027c921d fix: check_plugin.py honors per-plugin test/harness.json (#365)
check_one() always compares the render against committed golden images, but
the CLI never loaded the plugin's test/harness.json — so the deterministic
settings the goldens were generated with (config, mock data, frozen time,
sizes) weren't applied. For any time/data-dependent plugin this means the CLI
(and the plugins-repo CI workflow that calls it) renders live data and the
golden drifts on every run, even with no real regression. The pytest matrix
path already reads harness.json via load_harness_spec; the CLI now does too.

- check_one loads load_harness_spec(plugin_dir) and layers it under explicit
  CLI flags: config = schema defaults < harness.json < --config; sizes =
  --sizes > LEDMATRIX_TEST_SIZES env > harness.json > default sample;
  mock_data/freeze_time/skip_update fall back to harness.json when not given
  on the CLI.
- parse_sizes returns None (not DEFAULT_TEST_SIZES) when --sizes is omitted,
  so the env/harness.json/default fallback chain in resolve_test_sizes applies.
- Regression tests: harness.json supplies render settings, and CLI flags
  override it. Use a temp fixture plugin so they run in core CI (no plugins).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 12:52:33 -04:00
Ron PierceandClaude Opus 4.8 e0bd7088fa fix: make requirements-test.txt installable alongside requirements.txt (#364)
requirements.txt already pins pytest>=9.0.3,<10 (from #331), but
requirements-test.txt re-pinned pytest>=7.4,<9. The two ranges are
disjoint, so `pip install -r requirements.txt -r requirements-test.txt`
fails with ResolutionImpossible — breaking the core test workflow and the
plugins-repo safety workflow that installs both files.

pytest, pytest-cov, pytest-mock, and jsonschema are all already pinned
with major-version caps in requirements.txt, so drop them from
requirements-test.txt and keep only freezegun (the one test dep
requirements.txt doesn't provide).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 10:32:05 -04:00
Ron PierceandClaude Opus 4.8 313e35a98f Add cross-size/cross-screen plugin safety harness (#361)
* feat(testing): add cross-size/cross-screen plugin safety harness

Render every plugin across all supported matrix sizes (64x32, 128x32,
128x64, 256x32) and every declared screen, failing on crashes, content
drawn past the panel edge, or visual drift vs committed golden images.

- BoundsCheckingDisplayManager: oversized-canvas overflow detection
- harness.py: multi-size/multi-screen render engine + golden compare
- scripts/check_plugin.py: CLI (functional+bounds, --out-dir, --update-golden,
  --freeze-time); render_plugin.py refactored onto shared loading helpers
- test/plugins/test_harness.py + test_plugin_matrix.py (parametrized,
  honors per-plugin test/harness.json; skips when no plugins present)
- MockCacheManager.cache_dir so cache-dir-using plugins load headlessly
- .github/workflows/test.yml + docs/plugin-safety-harness.md

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(testing): address PR review feedback on plugin safety harness

- check_plugin: friendly error for non-numeric --sizes; reject non-object
  --config / --mock-data JSON; sanitize plugin mode before using as a
  filename; stop --update-golden from masking crash/overflow failures
- bounds_display_manager: pad the canvas out to the largest supported panel
  (not a fixed 16px) so far-overshoot coordinates are caught, not clipped
- harness: merge config_schema defaults inside render_plugin_matrix; surface
  update() failures as a non-fatal warning + result field instead of a debug
  log; sanitize mode in golden_path
- loading: fail fast when harness.json references a missing mock_data fixture
- mocks: clean up the per-instance temp cache dir via weakref.finalize
- test_plugin_matrix: add a discovery guard that fails when
  LEDMATRIX_REQUIRE_PLUGINS=1 but none found (still skips locally); type hints
- bound test deps with upper version pins for deterministic CI

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(testing): render plugins across arbitrary panel sizes, not a fixed list

Addresses maintainer feedback that there is no canonical set of supported
panel sizes — a build can be any size/configuration (square, 2x2, 4x4, 8x2,
long strips, tall stacks).

- sizes.py: SUPPORTED_SIZES -> DEFAULT_TEST_SIZES (back-compat alias kept),
  reframed as a representative SAMPLE of real panel-grid arrangements rather
  than an authoritative list; add parse_size_token / coerce_sizes /
  resolve_test_sizes helpers
- sizes are now fully overridable: LEDMATRIX_TEST_SIZES env (global, e.g. test
  on your exact hardware) > per-plugin harness.json "sizes" > default sample;
  CLI --sizes unchanged
- bounds_display_manager: pad the canvas to the largest panel IN THE CURRENT
  RUN (via overflow_extent) instead of a hardcoded max, so cross-size overflow
  detection scales to whatever sizes a run uses
- harness: compute per-run extent and thread it into the bounds manager
- tests: arbitrary-shape + size-parsing/precedence coverage
- docs: rewrite "Supported sizes" -> "Sizes: a sample, not a fixed list"

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(testing): fail the harness on non-connectivity update() errors

Addresses the remaining review thread: recording every update() exception as a
non-fatal warning still let a real update() regression pass green as long as
display() survived. Now update() failures are classified — a tolerated set of
connectivity errors (ConnectionError/TimeoutError/socket/ssl/urllib/http/
requests) is recorded non-fatally (expected with no network in CI), while any
other exception is treated as a genuine bug and fails that render.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* ci(security): pin actions to SHAs and disable checkout credential persistence

Addresses the CodeRabbit/zizmor workflow-hardening finding: pin
actions/checkout and actions/setup-python to full commit SHAs and set
persist-credentials: false on checkout to reduce supply-chain and
token-exposure risk.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(testing): validate positive sizes; narrow requests import except

Two review findings:
- sizes.py: parse_size_token / coerce_sizes now reject non-positive
  dimensions (0x32, -64x32) with a clear message instead of passing invalid
  sizes downstream (CodeRabbit).
- harness.py: the optional `requests` import now catches ImportError
  specifically and logs instead of `except Exception: pass`, clearing the
  Codacy medium "Try, Except, Pass" (harness.py L52) and Ruff S110/BLE001.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 14:32:52 -04:00
Ron PierceandClaude Opus 4.8 122e6d6863 fix(web): use fully-qualified .service unit names for privileged systemctl (#360)
The web interface runs headless, so every privileged systemctl call must be
covered by a NOPASSWD rule in /etc/sudoers.d/ledmatrix_web. The sudo command
matches the command line exactly, but the code called 'systemctl start
ledmatrix' while configure_web_sudo.sh grants 'systemctl start
ledmatrix.service'. The rule never matched, so start/stop/enable/disable/
restart fell back to a password prompt and failed with 'a terminal is
required to read the password'.

Align all privileged systemctl calls on the fully-qualified unit names the
sudoers grants use. Add a regression test that cross-checks api_v3.py calls
against the grants in configure_web_sudo.sh.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 15:17:00 -04:00
d488e8a2ad fix(api): don't coerce all-digit strings to int when schema type is string (#363)
* docs(core): add module and class docstrings to the 5 undocumented core files

Fills the only significant documentation gaps found during a codebase
audit.  All other core files (plugin_system/, logging_config.py, etc.)
already have complete module, class, and function docstrings.

Files changed (documentation only — zero logic changes):

  display_controller.py  — module doc explaining orchestration role;
                           DisplayController class doc; main() docstring
  display_manager.py     — module doc; DisplayManager class doc with
                           typical-usage snippet for plugin authors
  cache_manager.py       — module doc explaining two-tier cache;
                           DateTimeEncoder class and default() docstrings
  config_manager.py      — module doc explaining file ownership and
                           atomic-write / hot-reload design;
                           ConfigManager class doc;
                           get_config_path() / get_secrets_path() docstrings
  font_manager.py        — module doc (class docstring already existed)

Also noted (but not changed to avoid behaviour risk):
  display_manager.py and font_manager.py use logging.getLogger() directly
  instead of the project's get_logger() wrapper.  display_manager.py also
  calls setLevel(logging.INFO) immediately after, which would be lost if
  switched to get_logger().

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(display_controller): three targeted hot-path optimizations

Opt 1 — cache inspect.signature() per plugin_id
  inspect.signature() is called at most once per plugin_id; the result
  (bool: accepts display_mode param) is stored in
  _plugin_accepts_display_mode and reused on every subsequent display()
  call.  Eliminates all reflection from the display path at runtime.
  Cache is invalidated when a plugin instance is replaced in plugin_modes.

Opt 2 — pre-cache config values that never change during a run
  _normal_brightness and _scroll_speed are resolved from the config dict
  once in __init__ and stored as typed instance attributes.
  - Removes 2+ chained dict.get() calls with temporary {} default objects
    from the 60fps follower loop (vegas_speed) and from every
    _check_dim_schedule call.
  - current_brightness init now uses _normal_brightness directly.

Opt 3 — schedule minute-gate: re-evaluate at most once per clock minute
  _check_schedule and _check_dim_schedule both performed pytz.timezone(),
  datetime.now(), strftime(), and datetime.strptime() on every outer loop
  call.  Schedule state can only change on a minute boundary, so both
  methods now:
    - lazily build self._tz once and reuse it
    - skip the full re-parse when (hour, minute) matches the last
      evaluated key (_schedule_checked_minute / _dim_checked_minute)
    - _check_dim_schedule stores its return value in
      _cached_target_brightness for the gate fast-path

Tests: 23 new tests in test_display_controller_optimizations.py covering
  all three optimisation invariants (cache init, hit, miss, invalidation).
  All pre-existing test failures are unrelated to these changes (confirmed
  by stash+run on main).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: resolve 22 pre-existing test failures across 6 groups

Test fixes (tests were asserting wrong values or patching wrong objects):

  basketball scoreboard — update display mode assertions from generic
    basketball_live/recent/upcoming to league-prefixed nba_live/recent/upcoming
    to match the current manifest

  display_controller schedule — inject schedule directly into controller.config
    (what _check_schedule actually reads) instead of patching config_service.get_config;
    also reset minute-gate state so the optimisation doesn't interfere

  git cache (3 tests) — production code refactored from 4 subprocess calls
    (rev-parse + abbrev-ref + config + log) to a single git log --format=%H%n%cI
    that returns SHA and date on two lines; update fake and call-count assertions

  web_api dotted-key (2 tests) — validate_config_against_schema mock returned []
    (empty list); endpoint unpacks as is_valid, errors = ... causing ValueError;
    fix: return_value = (True, [])

  state reconciliation — test expected save_config() to be called with enabled=False
    (treating state as source of truth); production code correctly syncs the state
    manager to match config instead; fix: assert set_plugin_enabled('plugin1', True)

Production fixes (production code had bugs or missing features):

  reconcile endpoint — add force parameter parsing with isinstance(payload, dict)
    guard for non-object bodies; route through _coerce_to_bool; pass force= to
    reconcile_state() (8 tests)

  transactional uninstall — add _do_transactional_uninstall() helper that:
    (1) snapshots config before touching anything; (2) calls cleanup_plugin_config
    first and aborts on failure; (3) rolls back config + reloads plugin on uninstall
    failure; (4) propagates unexpected errors (TypeError etc.) instead of swallowing
    them (6 tests)

  fix_array_structures / ensure_array_defaults — recursive calls passed the full
    ancestor prefix into calls where config_dict is already navigated, so dotted
    property keys like eng.1 caused parent_parts.split('.') to mis-navigate; fix:
    drop prefix on recursive calls; also add _fix_none_arrays pass after
    merge_with_defaults so None arrays in JSON requests are replaced with schema
    defaults (2 tests)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf: four targeted optimizations across the display pipeline

Opt 1 — cache data-fetch interval per plugin (plugin_manager.py)
  _get_plugin_update_interval fell back to config_manager.get_config()
  (a full dict copy) when the manifest lacked an interval.  Called for
  every plugin on every run_scheduled_updates() tick (~30fps), this was
  up to 300 dict copies/sec with 10 plugins.
  Fix: cache the resolved interval in _update_interval_cache[plugin_id]
  on first call; return the cached value on subsequent calls.  Cache is
  cleared on load_plugin and unload_plugin.

Opt 2 — demote noisy per-cycle INFO logs to DEBUG (display_controller.py)
  Four logger.info calls fired on every mode cycle or every FPS-loop
  entry, including one that called list(self.plugin_modes.keys())
  unconditionally (allocating a list every outer loop iteration).
  - "Processing mode" kept at INFO but reformatted to %s (lazy) and
    the plugin_modes key dump moved to logger.debug
  - "Attempting/Got cycle duration" → logger.debug
  - "Entering high/normal FPS loop" → logger.debug
  Mode name at INFO is preserved for black-screen troubleshooting.

Opt 3 — use Image.frombytes instead of Image.fromarray in scroll hot path
  (scroll_helper.py)
  Image.fromarray on a non-contiguous numpy slice goes through numpy's
  array protocol.  Image.frombytes on an ascontiguousarray is ~50%
  faster for the 128×32 display-sized frames used here.  Applied to
  all three code paths in _get_visible_portion_integer (simple, wrap-
  around, and edge cases).

Opt 5 — cache get_text_width per (text, font) pair (display_manager.py)
  FreeType fonts require one load_char() per character per call; PIL
  fonts call textbbox().  Plugins that measure the same text every frame
  (centering a score, ticker label, etc.) were re-measuring from scratch
  on every display() call.
  Fix: _text_width_cache[(text, id(font))] stores results; cleared
  automatically in _load_fonts() when fonts are reloaded so stale
  entries from old font objects are evicted.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(scroll_helper): fix edge-case bug exposed by frombytes switch

The previous commit replaced Image.fromarray with Image.frombytes in
_get_visible_portion_integer.  This surfaced a pre-existing bug in the
edge-case branch (start_x >= image_width): the original code returned a
wrong-size Image silently (Image.fromarray accepts a too-short array);
Image.frombytes raises ValueError instead.

Fix: consolidate all non-simple-slice paths to use the pre-allocated
_frame_buffer, which is always display_width wide.  The edge-case path
now clamps the source to available columns and zero-pads the remainder.

Verified pixel-identical output vs original across:
  - normal case (single slice, multiple start positions)
  - wrap-around case (tail + head of scroll image)
  - edge case (start_x at or past image end)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: address CodeRabbit review comments on PR #358

1. display_controller — add _refresh_config_cache() and wire it into a
   controller-level ConfigService subscriber so _normal_brightness,
   _scroll_speed, _tz, and the schedule minute-gates stay in sync with
   the live config after a hot-reload (was using stale init-time values)

2. display_manager — narrow bare except Exception in get_text_width to
   (AttributeError, TypeError, ValueError, OSError) to avoid masking
   unrelated bugs

3. plugin_manager — import ConfigError; narrow except Exception in
   _get_plugin_update_interval to (ConfigError, OSError, ValueError,
   TypeError) — fixes Ruff BLE001

4. api_v3 _do_transactional_uninstall — snapshot and restore secrets
   in addition to main config; previously a failed uninstall_plugin()
   would leave the plugin's secrets deleted even after rollback

5. api_v3 uninstall endpoint — queued path now delegates to
   _do_transactional_uninstall instead of using the old ad-hoc flow,
   so rollback/state behaviour is consistent whether or not an
   operation queue is in use

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(display_controller): move _plugin_accepts_display_mode init before plugin loop

Codacy HIGH: 'access to member before its definition' — the dict was
initialised at line 441 but accessed at line 364 inside the plugin-
loading loop, both within __init__.

Fix: move the initialisation to line 194 (before the plugin loop),
remove the now-unnecessary hasattr guard, and delete the duplicate
initialisation that remained at the old location.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(api): don't coerce all-digit strings to int when schema type is string

_parse_form_value_with_schema had a fallback that tried int()/float() on
any string value that wasn't already handled. Fields like station_id
(type: "string", value: "8726607") were silently converted to integers,
causing jsonschema validation to reject them with "expected string, got int".

Guard the fallback with a check that skips it when the schema property
explicitly declares type: "string".

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-04 15:01:42 -04:00
Ron PierceandClaude Opus 4.8 b9dcbb5152 fix(display): resume rotation where it left off after live priority ends (#362)
When a live-priority plugin (e.g. live sports, flights overhead) preempted
the rotation, the controller overwrote current_mode_index with the live
plugin's index. Once live priority ended, rotation continued from after the
live plugin's mode, skipping every mode between the interrupted position and
the live plugin. With a live plugin late in the order, modes just before it
were starved indefinitely.

Save the rotation position on the initial live-priority switch and restore it
when live priority ends, in a new _apply_live_priority() helper. Add
regression tests covering resume, no-double-save during the hold, and the
idle no-op.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 10:56:02 -04:00
Ron PierceandClaude Opus 4.8 f27fd260f7 fix(web-ui): load v3 tab content deterministically (#359)
* fix(web-ui): load v3 tab content deterministically

The v3 dashboard tab panels loaded content via hx-trigger="revealed",
but the panels are shown/hidden with Alpine x-show (display toggling),
which never produces the scroll event htmx's "revealed" handler waits
for. loadTabContent tried to force it with htmx.trigger(el, 'revealed'),
but "revealed" is a synthetic scroll/observer trigger, not a dispatchable
event, so that call is a no-op. The result was an intermittently blank
panel - content appeared only when htmx's native reveal scan happened to
fire on its own.

- Replace the trigger with a custom "loadtab" event that nothing fires
  spontaneously (0% native firing).
- Load panels via htmx.ajax, which issues the request directly and works
  even before htmx has processed the element's triggers - unlike
  htmx.trigger, which is lost if dispatched before processing.
- Poll for htmx when it hasn't finished loading from the CDN instead of
  relying on a one-shot htmx:ready event that can be missed.
- Stamp data-loaded on the request promise so each panel loads once.

Verified in the emulator web UI: overview loads on every reload, tabs
lazy-load on demand, and revisiting a tab does not refetch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(web-ui): guard tab loads against stale pollers and re-entry

Address review feedback. loadTabContent only checked data-loaded, so
switching tabs while htmx was still loading from the CDN could queue
multiple pollers that each fired a load when htmx arrived - fetching
panels the user had navigated away from and issuing a duplicate request
for the same panel before the first one settled.

Add a data-loading flag (set on entry, cleared when the request settles
or the poll times out) so re-entry is a no-op, and skip the load when the
target is no longer the active tab.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 12:07:40 -04:00
eedf680a8c perf: display pipeline optimizations — caching, logging, scroll, text width (#358)
* docs(core): add module and class docstrings to the 5 undocumented core files

Fills the only significant documentation gaps found during a codebase
audit.  All other core files (plugin_system/, logging_config.py, etc.)
already have complete module, class, and function docstrings.

Files changed (documentation only — zero logic changes):

  display_controller.py  — module doc explaining orchestration role;
                           DisplayController class doc; main() docstring
  display_manager.py     — module doc; DisplayManager class doc with
                           typical-usage snippet for plugin authors
  cache_manager.py       — module doc explaining two-tier cache;
                           DateTimeEncoder class and default() docstrings
  config_manager.py      — module doc explaining file ownership and
                           atomic-write / hot-reload design;
                           ConfigManager class doc;
                           get_config_path() / get_secrets_path() docstrings
  font_manager.py        — module doc (class docstring already existed)

Also noted (but not changed to avoid behaviour risk):
  display_manager.py and font_manager.py use logging.getLogger() directly
  instead of the project's get_logger() wrapper.  display_manager.py also
  calls setLevel(logging.INFO) immediately after, which would be lost if
  switched to get_logger().

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(display_controller): three targeted hot-path optimizations

Opt 1 — cache inspect.signature() per plugin_id
  inspect.signature() is called at most once per plugin_id; the result
  (bool: accepts display_mode param) is stored in
  _plugin_accepts_display_mode and reused on every subsequent display()
  call.  Eliminates all reflection from the display path at runtime.
  Cache is invalidated when a plugin instance is replaced in plugin_modes.

Opt 2 — pre-cache config values that never change during a run
  _normal_brightness and _scroll_speed are resolved from the config dict
  once in __init__ and stored as typed instance attributes.
  - Removes 2+ chained dict.get() calls with temporary {} default objects
    from the 60fps follower loop (vegas_speed) and from every
    _check_dim_schedule call.
  - current_brightness init now uses _normal_brightness directly.

Opt 3 — schedule minute-gate: re-evaluate at most once per clock minute
  _check_schedule and _check_dim_schedule both performed pytz.timezone(),
  datetime.now(), strftime(), and datetime.strptime() on every outer loop
  call.  Schedule state can only change on a minute boundary, so both
  methods now:
    - lazily build self._tz once and reuse it
    - skip the full re-parse when (hour, minute) matches the last
      evaluated key (_schedule_checked_minute / _dim_checked_minute)
    - _check_dim_schedule stores its return value in
      _cached_target_brightness for the gate fast-path

Tests: 23 new tests in test_display_controller_optimizations.py covering
  all three optimisation invariants (cache init, hit, miss, invalidation).
  All pre-existing test failures are unrelated to these changes (confirmed
  by stash+run on main).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: resolve 22 pre-existing test failures across 6 groups

Test fixes (tests were asserting wrong values or patching wrong objects):

  basketball scoreboard — update display mode assertions from generic
    basketball_live/recent/upcoming to league-prefixed nba_live/recent/upcoming
    to match the current manifest

  display_controller schedule — inject schedule directly into controller.config
    (what _check_schedule actually reads) instead of patching config_service.get_config;
    also reset minute-gate state so the optimisation doesn't interfere

  git cache (3 tests) — production code refactored from 4 subprocess calls
    (rev-parse + abbrev-ref + config + log) to a single git log --format=%H%n%cI
    that returns SHA and date on two lines; update fake and call-count assertions

  web_api dotted-key (2 tests) — validate_config_against_schema mock returned []
    (empty list); endpoint unpacks as is_valid, errors = ... causing ValueError;
    fix: return_value = (True, [])

  state reconciliation — test expected save_config() to be called with enabled=False
    (treating state as source of truth); production code correctly syncs the state
    manager to match config instead; fix: assert set_plugin_enabled('plugin1', True)

Production fixes (production code had bugs or missing features):

  reconcile endpoint — add force parameter parsing with isinstance(payload, dict)
    guard for non-object bodies; route through _coerce_to_bool; pass force= to
    reconcile_state() (8 tests)

  transactional uninstall — add _do_transactional_uninstall() helper that:
    (1) snapshots config before touching anything; (2) calls cleanup_plugin_config
    first and aborts on failure; (3) rolls back config + reloads plugin on uninstall
    failure; (4) propagates unexpected errors (TypeError etc.) instead of swallowing
    them (6 tests)

  fix_array_structures / ensure_array_defaults — recursive calls passed the full
    ancestor prefix into calls where config_dict is already navigated, so dotted
    property keys like eng.1 caused parent_parts.split('.') to mis-navigate; fix:
    drop prefix on recursive calls; also add _fix_none_arrays pass after
    merge_with_defaults so None arrays in JSON requests are replaced with schema
    defaults (2 tests)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf: four targeted optimizations across the display pipeline

Opt 1 — cache data-fetch interval per plugin (plugin_manager.py)
  _get_plugin_update_interval fell back to config_manager.get_config()
  (a full dict copy) when the manifest lacked an interval.  Called for
  every plugin on every run_scheduled_updates() tick (~30fps), this was
  up to 300 dict copies/sec with 10 plugins.
  Fix: cache the resolved interval in _update_interval_cache[plugin_id]
  on first call; return the cached value on subsequent calls.  Cache is
  cleared on load_plugin and unload_plugin.

Opt 2 — demote noisy per-cycle INFO logs to DEBUG (display_controller.py)
  Four logger.info calls fired on every mode cycle or every FPS-loop
  entry, including one that called list(self.plugin_modes.keys())
  unconditionally (allocating a list every outer loop iteration).
  - "Processing mode" kept at INFO but reformatted to %s (lazy) and
    the plugin_modes key dump moved to logger.debug
  - "Attempting/Got cycle duration" → logger.debug
  - "Entering high/normal FPS loop" → logger.debug
  Mode name at INFO is preserved for black-screen troubleshooting.

Opt 3 — use Image.frombytes instead of Image.fromarray in scroll hot path
  (scroll_helper.py)
  Image.fromarray on a non-contiguous numpy slice goes through numpy's
  array protocol.  Image.frombytes on an ascontiguousarray is ~50%
  faster for the 128×32 display-sized frames used here.  Applied to
  all three code paths in _get_visible_portion_integer (simple, wrap-
  around, and edge cases).

Opt 5 — cache get_text_width per (text, font) pair (display_manager.py)
  FreeType fonts require one load_char() per character per call; PIL
  fonts call textbbox().  Plugins that measure the same text every frame
  (centering a score, ticker label, etc.) were re-measuring from scratch
  on every display() call.
  Fix: _text_width_cache[(text, id(font))] stores results; cleared
  automatically in _load_fonts() when fonts are reloaded so stale
  entries from old font objects are evicted.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(scroll_helper): fix edge-case bug exposed by frombytes switch

The previous commit replaced Image.fromarray with Image.frombytes in
_get_visible_portion_integer.  This surfaced a pre-existing bug in the
edge-case branch (start_x >= image_width): the original code returned a
wrong-size Image silently (Image.fromarray accepts a too-short array);
Image.frombytes raises ValueError instead.

Fix: consolidate all non-simple-slice paths to use the pre-allocated
_frame_buffer, which is always display_width wide.  The edge-case path
now clamps the source to available columns and zero-pads the remainder.

Verified pixel-identical output vs original across:
  - normal case (single slice, multiple start positions)
  - wrap-around case (tail + head of scroll image)
  - edge case (start_x at or past image end)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: address CodeRabbit review comments on PR #358

1. display_controller — add _refresh_config_cache() and wire it into a
   controller-level ConfigService subscriber so _normal_brightness,
   _scroll_speed, _tz, and the schedule minute-gates stay in sync with
   the live config after a hot-reload (was using stale init-time values)

2. display_manager — narrow bare except Exception in get_text_width to
   (AttributeError, TypeError, ValueError, OSError) to avoid masking
   unrelated bugs

3. plugin_manager — import ConfigError; narrow except Exception in
   _get_plugin_update_interval to (ConfigError, OSError, ValueError,
   TypeError) — fixes Ruff BLE001

4. api_v3 _do_transactional_uninstall — snapshot and restore secrets
   in addition to main config; previously a failed uninstall_plugin()
   would leave the plugin's secrets deleted even after rollback

5. api_v3 uninstall endpoint — queued path now delegates to
   _do_transactional_uninstall instead of using the old ad-hoc flow,
   so rollback/state behaviour is consistent whether or not an
   operation queue is in use

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(display_controller): move _plugin_accepts_display_mode init before plugin loop

Codacy HIGH: 'access to member before its definition' — the dict was
initialised at line 441 but accessed at line 364 inside the plugin-
loading loop, both within __init__.

Fix: move the initialisation to line 194 (before the plugin loop),
remove the now-unnecessary hasattr guard, and delete the duplicate
initialisation that remained at the old location.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Chuck <chuck@example.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-01 11:58:21 -04:00
Ron PierceandClaude Opus 4.8 ac3a15bfaa fix(web): repair array-table.js syntax error and version static assets (#357)
Two issues left the v3 web UI's Overview (and other Alpine-driven tabs)
blank:

1. array-table.js had two safeSetHTML(target, `...`) calls that closed the
   template-literal argument with `; instead of `); — a SyntaxError that
   aborts the script and halts widget registration / Alpine initialization.

2. Static assets are served `Cache-Control: public, max-age=31536000,
   immutable` but were referenced without a cache-busting version (the header
   comment assumed "versioning via query params", which was only ever applied
   by hand to app.css). So edited JS/CSS never reached browsers — including
   fix #1.

Add a Flask url_defaults hook that appends each static file's mtime as a ?v=
param to every url_for('static', ...), so changed files get a new URL and are
refetched while unchanged files keep the long immutable cache. Drop the now
redundant manual ?v= on app.css.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 11:00:40 -04:00
970 changed files with 172549 additions and 64941 deletions
-1
View File
@@ -4,4 +4,3 @@ exclude_paths:
- "plugins/**"
- "assets/**"
- "test/**"
- "scripts/debug/**"
-145
View File
@@ -1,145 +0,0 @@
# Cursor Helper Files for LEDMatrix Plugin Development
This directory contains Cursor-specific helper files to assist with plugin development in the LEDMatrix project.
## Files Overview
### `.cursorrules`
Comprehensive rules file that Cursor uses to understand plugin development patterns, best practices, and workflows. This file is automatically loaded by Cursor and helps guide AI-assisted development.
### `plugins_guide.md`
Detailed guide covering:
- Plugin system overview
- Creating new plugins
- Running plugins (emulator and hardware)
- Loading and configuring plugins
- Development workflow
- Testing strategies
- Troubleshooting
### `plugin_templates/`
Template files for quick plugin creation:
- `manifest.json.template` - Plugin metadata template
- `manager.py.template` - Plugin class template
- `config_schema.json.template` - Configuration schema template
- `README.md.template` - Plugin documentation template
- `requirements.txt.template` - Dependencies template
- `QUICK_START.md` - Quick start guide for using templates
## Quick Reference
### Creating a New Plugin
1. **Using templates** (recommended):
```bash
# See QUICK_START.md in plugin_templates/
cd plugins
mkdir my-plugin
cd my-plugin
cp ../../.cursor/plugin_templates/*.template .
# Edit files, replacing PLUGIN_ID and other placeholders
```
2. **Using dev_plugin_setup.sh**:
```bash
# Link from GitHub
./scripts/dev/dev_plugin_setup.sh link-github my-plugin
# Link local repo
./scripts/dev/dev_plugin_setup.sh link my-plugin /path/to/repo
```
### Running the Display
```bash
# Emulator mode (development, no hardware required)
python3 run.py --emulator
# (equivalent: EMULATOR=true python3 run.py)
# Hardware (production, requires the rpi-rgb-led-matrix submodule built)
python3 run.py
# As a systemd service
sudo systemctl start ledmatrix
# Dev preview server (renders plugins to a browser without running run.py)
python3 scripts/dev_server.py # then open http://localhost:5001
```
The `-e`/`--emulator` CLI flag is defined in `run.py:19-20` and
sets `os.environ["EMULATOR"] = "true"` before any display imports,
which `src/display_manager.py:2` then reads to switch between the
hardware and emulator backends.
### Managing Plugins
```bash
# List plugins
./scripts/dev/dev_plugin_setup.sh list
# Check status
./scripts/dev/dev_plugin_setup.sh status
# Update plugin(s)
./scripts/dev/dev_plugin_setup.sh update [plugin-name]
# Unlink plugin
./scripts/dev/dev_plugin_setup.sh unlink <plugin-name>
```
## Using These Files with Cursor
### `.cursorrules`
Cursor automatically reads this file to understand:
- Plugin structure and requirements
- Development workflows
- Best practices
- Common patterns
- API reference
When asking Cursor to help with plugins, it will use this context to provide better assistance.
### Plugin Templates
Use templates when creating new plugins:
1. Copy templates from `.cursor/plugin_templates/`
2. Replace placeholders (PLUGIN_ID, PluginClassName, etc.)
3. Customize for your plugin's needs
4. Follow the guide in `plugins_guide.md`
### Documentation
Refer to `plugins_guide.md` for:
- Detailed explanations
- Troubleshooting steps
- Best practices
- Examples and patterns
## Plugin Development Workflow
1. **Plan**: Determine plugin functionality and requirements
2. **Create**: Use templates or dev_plugin_setup.sh to create plugin structure
3. **Develop**: Implement plugin logic following BasePlugin interface
4. **Test**: Test with emulator first, then on hardware
5. **Configure**: Add plugin config to config/config.json
6. **Iterate**: Refine based on testing and feedback
## Resources
- **Plugin System**: `src/plugin_system/`
- **Base Plugin**: `src/plugin_system/base_plugin.py`
- **Plugin Manager**: `src/plugin_system/plugin_manager.py`
- **Example Plugins**: see the
[`ledmatrix-plugins`](https://github.com/ChuckBuilds/ledmatrix-plugins)
repo for canonical sources (e.g. `plugins/hockey-scoreboard/`,
`plugins/football-scoreboard/`). Installed plugins land in
`plugin-repos/` (default) or `plugins/` (dev fallback).
- **Architecture Docs**: `docs/PLUGIN_ARCHITECTURE_SPEC.md`
- **Development Setup**: `scripts/dev/dev_plugin_setup.sh`
## Getting Help
1. Check `plugins_guide.md` for detailed documentation
2. Review `.cursorrules` for development patterns
3. Look at existing plugins for examples
4. Check logs for error messages
5. Review plugin system code in `src/plugin_system/`
-202
View File
@@ -1,202 +0,0 @@
# Implementation Plan: Fix Config Schema Validation Issues
Based on audit results showing 186 issues across 20 plugins.
## Overview
Three priority fixes identified from audit:
1. **Priority 1 (HIGH)**: Remove core properties from required array - will fix ~150 issues
2. **Priority 2 (MEDIUM)**: Verify default merging logic - will fix remaining required field issues
3. **Priority 3 (LOW)**: Calendar plugin schema cleanup - will fix 3 extra field warnings
## Priority 1: Remove Core Properties from Required Array
### Problem
Core properties (`enabled`, `display_duration`, `live_priority`) are system-managed but listed in schema `required` arrays. SchemaManager injects them into properties but doesn't remove them from `required`, causing validation failures.
### Solution
**File**: `src/plugin_system/schema_manager.py`
**Location**: `validate_config_against_schema()` method, after line 295
### Implementation Steps
1. **Add code to remove core properties from required array**:
```python
# After injecting core properties (around line 295), add:
# Remove core properties from required array (they're system-managed)
if "required" in enhanced_schema:
core_prop_names = list(core_properties.keys())
enhanced_schema["required"] = [
field for field in enhanced_schema["required"]
if field not in core_prop_names
]
```
2. **Add logging for debugging** (optional but helpful):
```python
if "required" in enhanced_schema and core_prop_names:
removed_from_required = [
field for field in enhanced_schema.get("required", [])
if field in core_prop_names
]
if removed_from_required and plugin_id:
self.logger.debug(
f"Removed core properties from required array for {plugin_id}: {removed_from_required}"
)
```
3. **Test the fix**:
- Run audit script: `python scripts/audit_plugin_configs.py`
- Expected: Issue count drops from 186 to ~30-40
- All "enabled" related errors should be eliminated
### Expected Outcome
- All 20 plugins should no longer fail validation due to missing `enabled` field
- ~150 issues resolved (all enabled-related validation errors)
## Priority 2: Verify Default Merging Logic
### Problem
Some plugins have required fields with defaults that should be applied before validation. Need to verify the default merging happens correctly and handles nested objects.
### Solution
**File**: `web_interface/blueprints/api_v3.py`
**Location**: `save_plugin_config()` method, around lines 3218-3221
### Implementation Steps
1. **Review current default merging logic**:
- Check that `merge_with_defaults()` is called before validation (line 3220)
- Verify it's called after preserving enabled state but before validation
2. **Verify merge_with_defaults handles nested objects**:
- Check `src/plugin_system/schema_manager.py` → `merge_with_defaults()` method
- Ensure it recursively merges nested objects (it does use deep_merge)
- Test with plugins that have nested required fields
3. **Check if defaults are applied for nested required fields**:
- Review how `generate_default_config()` extracts defaults from nested schemas
- Verify nested required fields with defaults are included
4. **Test with problematic plugins**:
- `ledmatrix-weather`: required fields `api_key`, `location_city` (check if defaults exist)
- `mqtt-notifications`: required field `mqtt` object (check if default exists)
- `text-display`: required field `text` (check if default exists)
- `ledmatrix-music`: required field `preferred_source` (check if default exists)
5. **If defaults don't exist in schemas**:
- Either add defaults to schemas, OR
- Make fields optional in schemas if they're truly optional
### Expected Outcome
- Plugins with required fields that have schema defaults should pass validation
- Issue count further reduced from ~30-40 to ~5-10
## Priority 3: Calendar Plugin Schema Cleanup
### Problem
Calendar plugin config has fields not in schema:
- `show_all_day` (config) but schema has `show_all_day_events` (field name mismatch)
- `date_format` (not in schema, not used in manager.py)
- `time_format` (not in schema, not used in manager.py)
### Investigation Results
- Schema defines: `show_all_day_events` (boolean, default: true)
- Manager.py uses: `show_all_day_events` (line 82: `config.get('show_all_day_events', True)`)
- Config has: `show_all_day` (wrong field name - should be `show_all_day_events`)
- `date_format` and `time_format` appear to be deprecated (not used in manager.py)
### Solution
**File**: `config/config.json` → `calendar` section
### Implementation Steps
1. **Fix field name mismatch**:
- Rename `show_all_day` → `show_all_day_events` in config.json
- This matches the schema and manager.py code
2. **Remove deprecated fields**:
- Remove `date_format` from config (not used in code)
- Remove `time_format` from config (not used in code)
3. **Alternative (if fields are needed)**: Add `date_format` and `time_format` to schema
- Only if these fields should be supported
- Check if they're used anywhere else in the codebase
4. **Test calendar plugin**:
- Run audit for calendar plugin specifically
- Verify no extra field warnings remain
- Test calendar plugin functionality to ensure it still works
### Expected Outcome
- Calendar plugin shows 0 extra field warnings
- Final issue count: ~3-5 (only edge cases remain)
## Testing Strategy
### After Each Priority Fix
1. **Run local audit**:
```bash
python scripts/audit_plugin_configs.py
```
2. **Check issue count reduction**:
- Priority 1: Should drop from 186 to ~30-40
- Priority 2: Should drop from ~30-40 to ~5-10
- Priority 3: Should drop from ~5-10 to ~3-5
3. **Review specific plugin results**:
```bash
python scripts/audit_plugin_configs.py --plugin <plugin-id>
```
### After All Fixes
1. **Full audit run**:
```bash
python scripts/audit_plugin_configs.py
```
2. **Deploy to Pi**:
```bash
./scripts/deploy_to_pi.sh src/plugin_system/schema_manager.py web_interface/blueprints/api_v3.py
```
3. **Run audit on Pi**:
```bash
./scripts/run_audit_on_pi.sh
```
4. **Manual web interface testing**:
- Access each problematic plugin's config page
- Try saving configuration
- Verify no validation errors appear
- Check that configs save successfully
## Success Criteria
- [ ] Priority 1: All "enabled" related validation errors eliminated
- [ ] Priority 1: Issue count reduced from 186 to ~30-40
- [ ] Priority 2: Plugins with required fields + defaults pass validation
- [ ] Priority 2: Issue count reduced to ~5-10
- [ ] Priority 3: Calendar plugin extra field warnings resolved
- [ ] Priority 3: Final issue count at ~3-5 (only edge cases)
- [ ] All fixes work on Pi (not just local)
- [ ] Web interface saves configs without validation errors
## Files to Modify
1. `src/plugin_system/schema_manager.py` - Remove core properties from required array
2. `plugins/calendar/config_schema.json` OR `config/config.json` - Calendar cleanup (if needed)
3. `web_interface/blueprints/api_v3.py` - May need minor adjustments for default merging (if needed)
## Risk Assessment
**Priority 1**: Low risk - Only affects validation logic, doesn't change behavior
**Priority 2**: Low risk - Only ensures defaults are applied (already intended behavior)
**Priority 3**: Very low risk - Only affects calendar plugin, cosmetic issue
All changes are backward compatible and improve the system rather than changing core functionality.
-247
View File
@@ -1,247 +0,0 @@
# Quick Start: Creating a New Plugin
This guide will help you create a new plugin using the templates in `.cursor/plugin_templates/`.
## Step 1: Create Plugin Directory
```bash
cd /path/to/LEDMatrix
mkdir -p plugins/my-plugin
cd plugins/my-plugin
```
## Step 2: Copy Templates
```bash
# Copy all template files
cp ../../.cursor/plugin_templates/manifest.json.template ./manifest.json
cp ../../.cursor/plugin_templates/manager.py.template ./manager.py
cp ../../.cursor/plugin_templates/config_schema.json.template ./config_schema.json
cp ../../.cursor/plugin_templates/README.md.template ./README.md
cp ../../.cursor/plugin_templates/requirements.txt.template ./requirements.txt
```
## Step 3: Customize Files
### manifest.json
Replace placeholders:
- `PLUGIN_ID` → `my-plugin` (lowercase, use hyphens)
- `Plugin Name` → Your plugin's display name
- `PluginClassName` → `MyPlugin` (PascalCase)
- Update description, author, homepage, etc.
### manager.py
Replace placeholders:
- `PluginClassName` → `MyPlugin` (must match manifest)
- Implement `_fetch_data()` method
- Implement `_render_content()` method
- Add any custom validation in `validate_config()`
### config_schema.json
Customize:
- Update description
- Add/remove configuration properties
- Set default values
- Add validation rules
### README.md
Replace placeholders:
- `PLUGIN_ID` → `my-plugin`
- `Plugin Name` → Your plugin's name
- Fill in features, installation, configuration sections
### requirements.txt
Add your plugin's dependencies:
```txt
requests>=2.28.0
pillow>=9.0.0
```
## Step 4: Enable Plugin
Edit `config/config.json`:
```json
{
"my-plugin": {
"enabled": true,
"display_duration": 15
}
}
```
## Step 5: Test Plugin
### Test with Emulator
```bash
cd /path/to/LEDMatrix
python run.py --emulator
```
### Check Plugin Loading
Look for logs like:
```
[INFO] Discovered 1 plugin(s)
[INFO] Loaded plugin: my-plugin v1.0.0
[INFO] Added plugin mode: my-plugin
```
### Test Plugin Display
The plugin should appear in the display rotation. Check logs for any errors.
## Step 6: Develop and Iterate
1. Edit `manager.py` to implement your plugin logic
2. Test with emulator: `python run.py --emulator`
3. Check logs for errors
4. Iterate until working correctly
## Step 7: Test on Hardware (Optional)
When ready, test on Raspberry Pi:
```bash
# Deploy to Pi
rsync -avz plugins/my-plugin/ pi@raspberrypi:/path/to/LEDMatrix/plugins/my-plugin/
# Or if using git
ssh pi@raspberrypi "cd /path/to/LEDMatrix/plugins/my-plugin && git pull"
# Restart service
ssh pi@raspberrypi "sudo systemctl restart ledmatrix"
```
## Common Customizations
### Adding API Integration
1. Add API key to `config_schema.json`:
```json
{
"api_key": {
"type": "string",
"description": "API key for service"
}
}
```
2. Implement API call in `_fetch_data()`:
```python
import requests
def _fetch_data(self):
response = requests.get(
"https://api.example.com/data",
headers={"Authorization": f"Bearer {self.api_key}"}
)
return response.json()
```
3. Store API key in `config/config_secrets.json`:
```json
{
"my-plugin": {
"api_key": "your-secret-key"
}
}
```
### Adding Image Rendering
There is no `draw_image()` helper on `DisplayManager`. To render an
image, paste it directly onto the underlying PIL `Image`
(`display_manager.image`) and then call `update_display()`:
```python
def _render_content(self):
# Load and paste image onto the display canvas
image = Image.open("assets/logo.png").convert("RGB")
self.display_manager.image.paste(image, (0, 0))
# Draw text overlay
self.display_manager.draw_text(
"Text",
x=10, y=20,
color=(255, 255, 255)
)
self.display_manager.update_display()
```
For transparency, paste with a mask:
```python
icon = Image.open("assets/icon.png").convert("RGBA")
self.display_manager.image.paste(icon, (5, 5), icon)
```
### Adding Live Priority
1. Enable in config:
```json
{
"my-plugin": {
"live_priority": true
}
}
```
2. Implement `has_live_content()`:
```python
def has_live_content(self) -> bool:
return self.data and self.data.get("is_live", False)
```
3. Override `get_live_modes()` if needed:
```python
def get_live_modes(self) -> list:
return ["my_plugin_live_mode"]
```
## Troubleshooting
### Plugin Not Loading
- Check `manifest.json` syntax (must be valid JSON)
- Verify `entry_point` file exists
- Ensure `class_name` matches class name in manager.py
- Check for import errors in logs
### Configuration Errors
- Validate config against `config_schema.json`
- Check required fields are present
- Verify data types match schema
### Display Issues
- Check display dimensions: `display_manager.width`, `display_manager.height`
- Verify coordinates are within bounds
- Ensure `update_display()` is called
- Test with emulator first
## Next Steps
- Review existing plugins for patterns:
- `plugins/hockey-scoreboard/` - Sports scoreboard example
- `plugins/ledmatrix-music/` - Real-time data example
- `plugins/ledmatrix-stocks/` - Data display example
- Read full documentation:
- `.cursor/plugins_guide.md` - Comprehensive guide
- `docs/PLUGIN_ARCHITECTURE_SPEC.md` - Architecture details
- `.cursorrules` - Development rules
- Check plugin system code:
- `src/plugin_system/base_plugin.py` - Base class
- `src/plugin_system/plugin_manager.py` - Plugin manager
-156
View File
@@ -1,156 +0,0 @@
# Plugin Name
Brief description of what this plugin does.
## Features
- Feature 1
- Feature 2
- Feature 3
## Installation
1. Link the plugin to your LEDMatrix installation:
```bash
cd /path/to/LEDMatrix
./scripts/dev/dev_plugin_setup.sh link-github PLUGIN_ID
```
Or for local development:
```bash
./scripts/dev/dev_plugin_setup.sh link PLUGIN_ID /path/to/plugin/repo
```
2. Install dependencies:
```bash
pip install -r plugins/PLUGIN_ID/requirements.txt
```
3. Configure the plugin in `config/config.json`:
```json
{
"PLUGIN_ID": {
"enabled": true,
"display_duration": 15
}
}
```
**Note:** API keys and other sensitive credentials must be stored in `config/config_secrets.json`, not in `config/config.json`.
4. Store API keys in `config/config_secrets.json`:
```json
{
"PLUGIN_ID": {
"api_key": "your-secret-api-key"
}
}
```
## Configuration
### Required Settings
- `enabled` (boolean): Enable or disable the plugin
- `api_key` (string): API key for external service (if required)
### Optional Settings
- `display_duration` (number): How long to display this plugin (default: 15 seconds)
- `refresh_interval` (integer): How often to refresh data in seconds (default: 60)
- `live_priority` (boolean): Enable live priority takeover (default: false)
## Display Modes
This plugin provides the following display modes:
- `PLUGIN_ID`: Main display mode
## API Requirements
This plugin requires:
- **API Name**: Description of API requirements
- URL: https://api.example.com
- Rate Limit: X requests per minute
- Authentication: API key required
## Development
### Running Tests
```bash
cd plugins/PLUGIN_ID
python test_PLUGIN_ID.py
```
### Testing with Emulator
```bash
cd /path/to/LEDMatrix
python run.py --emulator
```
### Debugging
Enable debug logging in `config/config.json`:
```json
{
"logging": {
"level": "DEBUG"
}
}
```
Check logs:
```bash
# On Raspberry Pi (if running as service)
journalctl -u ledmatrix -f
# Direct execution
python run.py
```
## Troubleshooting
### Plugin Not Loading
1. Check that `manifest.json` exists and is valid
2. Verify `entry_point` file exists
3. Check that `class_name` matches the class in manager.py
4. Review logs for import errors
### Configuration Errors
1. Validate config against `config_schema.json`
2. Check required fields are present
3. Verify data types match schema
### API Errors
1. Verify API key is correct
2. Check API rate limits
3. Review network connectivity
4. Check API service status
## License
[License information]
## Author
Your Name
## Links
- GitHub: https://github.com/username/ledmatrix-PLUGIN_ID
- Documentation: [Link to docs]
- Issues: https://github.com/username/ledmatrix-PLUGIN_ID/issues
@@ -1,44 +0,0 @@
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"title": "Plugin Configuration Schema",
"description": "Configuration schema for Plugin Name",
"properties": {
"enabled": {
"type": "boolean",
"default": true,
"description": "Enable or disable this plugin"
},
"display_duration": {
"type": "number",
"default": 15,
"minimum": 1,
"maximum": 300,
"description": "How long to display this plugin in seconds"
},
"live_priority": {
"type": "boolean",
"default": false,
"description": "Enable live priority takeover when plugin has live content"
},
"refresh_interval": {
"type": "integer",
"default": 60,
"minimum": 1,
"description": "How often to refresh data in seconds"
},
"api_key": {
"type": "string",
"description": "API key for external service (store in config_secrets.json)",
"default": ""
},
"custom_setting": {
"type": "string",
"description": "Example custom setting - replace with your plugin's settings",
"default": "default_value"
}
},
"required": ["enabled"],
"additionalProperties": false
}
@@ -1,226 +0,0 @@
"""
Plugin Name
Brief description of what this plugin does.
API Version: 1.0.0
"""
from src.plugin_system.base_plugin import BasePlugin
from PIL import Image
from typing import Dict, Any, Optional
import logging
import time
class PluginClassName(BasePlugin):
"""
Plugin class that inherits from BasePlugin.
This plugin demonstrates the basic structure and common patterns
for LEDMatrix plugins.
"""
def __init__(
self,
plugin_id: str,
config: Dict[str, Any],
display_manager,
cache_manager,
plugin_manager,
):
"""Initialize the plugin."""
super().__init__(plugin_id, config, display_manager, cache_manager, plugin_manager)
# Initialize plugin-specific data
self.data = None
self.last_update_time = None
# Load configuration values
self.api_key = config.get("api_key", "")
self.refresh_interval = config.get("refresh_interval", 60)
self.logger.info(f"Plugin {plugin_id} initialized")
def update(self) -> None:
"""
Fetch/update data for this plugin.
This method is called periodically based on update_interval
specified in the manifest. Use cache_manager to avoid
excessive API calls.
"""
cache_key = f"{self.plugin_id}_data"
# Check cache first
cached = self.cache_manager.get(cache_key, max_age=self.refresh_interval)
if cached:
self.data = cached
self.logger.debug("Using cached data")
return
try:
# Fetch new data
self.data = self._fetch_data()
# Cache the data
self.cache_manager.set(cache_key, self.data, ttl=self.refresh_interval)
self.last_update_time = time.time()
self.logger.info("Data updated successfully")
except Exception as e:
self.logger.error(f"Failed to update data: {e}")
# Use cached data if available, even if expired
# Use a very large max_age (1 year) to effectively bypass expiration for fallback
expired_cached = self.cache_manager.get(cache_key, max_age=31536000)
if expired_cached:
self.data = expired_cached
self.logger.warning("Using expired cache due to update failure")
def display(self, force_clear: bool = False) -> None:
"""
Render this plugin's display.
Args:
force_clear: If True, clear display before rendering
"""
if force_clear:
self.display_manager.clear()
# Check if we have data to display
if not self.data:
self._display_error("No data available")
return
try:
# Render plugin content
self._render_content()
# Update the display
self.display_manager.update_display()
except Exception as e:
self.logger.error(f"Display error: {e}")
self._display_error("Display error")
def _fetch_data(self) -> Dict[str, Any]:
"""
Fetch data from external source.
Returns:
Dictionary containing fetched data
"""
# TODO: Implement data fetching logic
# Example:
# import requests
# response = requests.get("https://api.example.com/data",
# headers={"Authorization": f"Bearer {self.api_key}"})
# return response.json()
# Placeholder
return {
"message": "Hello, World!",
"timestamp": time.time()
}
def _render_content(self) -> None:
"""Render the plugin content on the display."""
# Get display dimensions
width = self.display_manager.width
height = self.display_manager.height
# Example: Draw text
text = self.data.get("message", "No data")
x = 5
y = height // 2
self.display_manager.draw_text(
text,
x=x,
y=y,
color=(255, 255, 255) # White
)
# Example: Draw image
# if hasattr(self, 'logo_image'):
# self.display_manager.draw_image(
# self.logo_image,
# x=0,
# y=0
# )
def _display_error(self, message: str) -> None:
"""Display an error message."""
self.display_manager.clear()
width = self.display_manager.width
height = self.display_manager.height
self.display_manager.draw_text(
message,
x=5,
y=height // 2,
color=(255, 0, 0) # Red
)
self.display_manager.update_display()
def validate_config(self) -> bool:
"""
Validate plugin configuration.
Returns:
True if config is valid, False otherwise
"""
# Call parent validation first
if not super().validate_config():
return False
# Add custom validation
# Example: Check for required API key
# if self.config.get("require_api_key", True):
# if not self.api_key:
# self.logger.error("API key is required but not provided")
# return False
return True
def has_live_content(self) -> bool:
"""
Check if plugin has live content to display.
Override this method to enable live priority features.
Returns:
True if plugin has live content, False otherwise
"""
# Example: Check if there's live data
# return self.data and self.data.get("is_live", False)
return False
def get_info(self) -> Dict[str, Any]:
"""
Return plugin info for display in web UI.
Returns:
Dictionary with plugin information
"""
info = super().get_info()
# Add plugin-specific info
info.update({
"data_available": self.data is not None,
"last_update": self.last_update_time,
# Add more info as needed
})
return info
def cleanup(self) -> None:
"""Cleanup resources when plugin is unloaded."""
# Clean up any resources (threads, connections, etc.)
# Example:
# if hasattr(self, 'api_client'):
# self.api_client.close()
super().cleanup()
@@ -1,55 +0,0 @@
{
"id": "PLUGIN_ID",
"name": "Plugin Name",
"version": "1.0.0",
"author": "Your Name",
"description": "Brief description of what this plugin does",
"homepage": "https://github.com/username/ledmatrix-PLUGIN_ID",
"entry_point": "manager.py",
"class_name": "PluginClassName",
"category": "custom",
"tags": ["custom", "example"],
"icon": "fas fa-icon-name",
"compatible_versions": [">=2.0.0"],
"min_ledmatrix_version": "2.0.0",
"max_ledmatrix_version": "3.0.0",
"requires": {
"python": ">=3.9",
"display_size": {
"min_width": 64,
"min_height": 32
}
},
"config_schema": "config_schema.json",
"assets": {
"logos": "Optional: Description of asset requirements"
},
"update_interval": 60,
"default_duration": 15,
"display_modes": [
"PLUGIN_ID"
],
"api_requirements": [
{
"name": "API Name",
"required": false,
"description": "Description of API requirements",
"url": "https://api.example.com",
"rate_limit": "Rate limit information"
}
],
"download_url_template": "https://github.com/username/ledmatrix-PLUGIN_ID/archive/refs/tags/v{version}.zip",
"versions": [
{
"released": "2025-01-01",
"version": "1.0.0",
"ledmatrix_min_version": "2.0.0"
}
],
"last_updated": "2025-01-01",
"stars": 0,
"downloads": 0,
"verified": false,
"screenshot": ""
}
@@ -1,13 +0,0 @@
# Plugin Dependencies
# Add your plugin's Python dependencies here
# Example dependencies (uncomment and modify as needed):
# requests>=2.28.0
# pillow>=9.0.0
# python-dateutil>=2.8.0
# Note: Core LEDMatrix dependencies are already available:
# - PIL/Pillow (for image handling)
# - Core plugin system classes
# - Display manager, cache manager, config manager
@@ -1,136 +0,0 @@
"""
Test file for Plugin Name plugin.
This file provides example unit tests for your plugin.
Run tests with: python -m pytest test_manager.py
Or: python test_manager.py
"""
import unittest
import sys
from pathlib import Path
# Add project root to path
PROJECT_ROOT = Path(__file__).resolve().parent.parent.parent
if str(PROJECT_ROOT) not in sys.path:
sys.path.insert(0, str(PROJECT_ROOT))
from src.plugin_system.testing import PluginTestCase
from manager import PluginClassName
class TestPluginClassName(PluginTestCase):
"""Test cases for PluginClassName plugin."""
def setUp(self):
"""Set up test fixtures."""
super().setUp()
# Update plugin_id to match the plugin being tested
self.plugin_id = 'PLUGIN_ID'
# Create plugin instance
self.plugin = self.create_plugin_instance(
PluginClassName,
plugin_id='PLUGIN_ID',
config=self.get_mock_config()
)
def test_plugin_initialization(self):
"""Test that plugin initializes correctly."""
self.assert_plugin_initialized(self.plugin)
self.assertTrue(self.plugin.enabled)
def test_config_validation(self):
"""Test configuration validation."""
# Valid config should pass
self.assertTrue(self.plugin.validate_config())
# Test with invalid config if applicable
# invalid_config = self.get_mock_config(enabled='not-a-boolean')
# invalid_plugin = self.create_plugin_instance(
# PluginClassName,
# config=invalid_config
# )
# self.assertFalse(invalid_plugin.validate_config())
def test_update_method(self):
"""Test the update() method."""
# Reset mocks
self.cache_manager.reset()
# Call update
self.plugin.update()
# Assertions
# Example: Check that cache was used
# self.assert_cache_get('PLUGIN_ID_data')
# Example: Check that data was fetched and cached
# self.assert_cache_set('PLUGIN_ID_data')
def test_display_method(self):
"""Test the display() method."""
# Ensure plugin has data (call update first if needed)
# self.plugin.update()
# Call display
self.plugin.display(force_clear=True)
# Assertions
self.assert_display_cleared()
self.assert_display_updated()
# Example: Check that text was drawn
# self.assert_text_drawn("Expected Text")
# Example: Check that image was drawn
# self.assert_image_drawn()
def test_display_without_data(self):
"""Test display() behavior when no data is available."""
# Clear any cached data
self.cache_manager.reset()
# Call display
self.plugin.display()
# Should handle gracefully (no exceptions)
# May show error message or fallback content
self.assert_display_updated()
def test_get_display_duration(self):
"""Test display duration configuration."""
duration = self.plugin.get_display_duration()
self.assertIsInstance(duration, (int, float))
self.assertGreater(duration, 0)
# Test with custom duration
custom_config = self.get_mock_config(display_duration=30.0)
custom_plugin = self.create_plugin_instance(
PluginClassName,
config=custom_config
)
self.assertEqual(custom_plugin.get_display_duration(), 30.0)
def test_enable_disable(self):
"""Test plugin enable/disable functionality."""
self.assertTrue(self.plugin.enabled)
self.plugin.on_disable()
self.assertFalse(self.plugin.enabled)
self.plugin.on_enable()
self.assertTrue(self.plugin.enabled)
def test_config_change(self):
"""Test configuration change handling."""
new_config = self.get_mock_config(display_duration=20.0)
self.plugin.on_config_change(new_config)
self.assertEqual(self.plugin.config.get('display_duration'), 20.0)
if __name__ == '__main__':
unittest.main()
-751
View File
@@ -1,751 +0,0 @@
# LEDMatrix Plugin Development Guide
This guide provides comprehensive instructions for creating, running, and loading plugins in the LEDMatrix project.
## Table of Contents
1. [Plugin System Overview](#plugin-system-overview)
2. [Creating a New Plugin](#creating-a-new-plugin)
3. [Running Plugins](#running-plugins)
4. [Loading Plugins](#loading-plugins)
5. [Plugin Development Workflow](#plugin-development-workflow)
6. [Testing Plugins](#testing-plugins)
7. [Troubleshooting](#troubleshooting)
---
## Plugin System Overview
The LEDMatrix project uses a plugin-based architecture where all display functionality (except core calendar) is implemented as plugins. Plugins are dynamically loaded from the `plugins/` directory and integrated into the display rotation.
### Plugin Architecture
```
LEDMatrix Core
├── Plugin Manager (discovers, loads, manages plugins)
├── Display Manager (handles LED matrix rendering)
├── Cache Manager (data persistence)
├── Config Manager (configuration management)
└── Plugins/ (plugin directory)
├── plugin-1/
├── plugin-2/
└── ...
```
### Plugin Lifecycle
1. **Discovery**: PluginManager scans `plugins/` for directories with `manifest.json`
2. **Loading**: Plugin module is imported and class is instantiated
3. **Configuration**: Plugin config is loaded from `config/config.json`
4. **Validation**: `validate_config()` is called to verify configuration
5. **Registration**: Plugin is added to available display modes
6. **Execution**: `update()` is called periodically, `display()` is called during rotation
---
## Creating a New Plugin
### Method 1: Using dev_plugin_setup.sh (Recommended)
This method is best for plugins stored in separate Git repositories.
#### From GitHub Repository
```bash
# Link a plugin from GitHub (auto-detects URL)
./scripts/dev/dev_plugin_setup.sh link-github <plugin-name>
# Example: Link hockey-scoreboard plugin
./scripts/dev/dev_plugin_setup.sh link-github hockey-scoreboard
# With custom URL
./scripts/dev/dev_plugin_setup.sh link-github <plugin-name> https://github.com/user/repo.git
```
The script will:
- Clone the repository to `~/.ledmatrix-dev-plugins/` (or configured directory)
- Create a symlink in `plugins/<plugin-name>/` pointing to the cloned repo
- Validate the plugin structure
#### From Local Repository
```bash
# Link a local plugin repository
./scripts/dev/dev_plugin_setup.sh link <plugin-name> <path-to-repo>
# Example: Link a local plugin
./scripts/dev/dev_plugin_setup.sh link my-plugin ../ledmatrix-my-plugin
```
### Method 2: Manual Plugin Creation
1. **Create Plugin Directory**
```bash
mkdir -p plugins/my-plugin
cd plugins/my-plugin
```
2. **Create manifest.json**
```json
{
"id": "my-plugin",
"name": "My Plugin",
"version": "1.0.0",
"author": "Your Name",
"description": "Description of what this plugin does",
"entry_point": "manager.py",
"class_name": "MyPlugin",
"category": "custom",
"tags": ["custom", "example"],
"display_modes": ["my_plugin"],
"update_interval": 60,
"default_duration": 15,
"requires": {
"python": ">=3.9"
},
"config_schema": "config_schema.json"
}
```
3. **Create manager.py**
```python
from src.plugin_system.base_plugin import BasePlugin
from PIL import Image
import logging
class MyPlugin(BasePlugin):
"""My custom plugin implementation."""
def update(self):
"""Fetch/update data for this plugin."""
# Fetch data from API, files, etc.
# Use self.cache_manager for caching
cache_key = f"{self.plugin_id}_data"
cached = self.cache_manager.get(cache_key, max_age=3600)
if cached:
self.data = cached
return
# Fetch new data
self.data = self._fetch_data()
self.cache_manager.set(cache_key, self.data)
def display(self, force_clear=False):
"""Render this plugin's display."""
if force_clear:
self.display_manager.clear()
# Render content using display_manager
self.display_manager.draw_text(
"Hello, World!",
x=10, y=15,
color=(255, 255, 255)
)
self.display_manager.update_display()
def _fetch_data(self):
"""Fetch data from external source."""
# Implement your data fetching logic
return {"message": "Hello, World!"}
def validate_config(self):
"""Validate plugin configuration."""
# Check required config fields
if not super().validate_config():
return False
# Add custom validation
required_fields = ['api_key'] # Example
for field in required_fields:
if field not in self.config:
self.logger.error(f"Missing required field: {field}")
return False
return True
```
4. **Create config_schema.json**
```json
{
"type": "object",
"properties": {
"enabled": {
"type": "boolean",
"default": true,
"description": "Enable or disable this plugin"
},
"display_duration": {
"type": "number",
"default": 15,
"minimum": 1,
"description": "How long to display this plugin (seconds)"
},
"api_key": {
"type": "string",
"description": "API key for external service"
}
},
"required": ["enabled"]
}
```
5. **Create requirements.txt** (if needed)
```
requests>=2.28.0
pillow>=9.0.0
```
6. **Create README.md**
Document your plugin's functionality, configuration options, and usage.
---
## Running Plugins
### Development Mode (Emulator)
Run the LEDMatrix system with emulator for plugin testing:
```bash
# Using run.py
python run.py --emulator
# Using emulator script
./run_emulator.sh
```
The emulator will:
- Load all enabled plugins
- Display plugin content in a window (simulating LED matrix)
- Show logs for plugin loading and execution
- Allow testing without Raspberry Pi hardware
### Production Mode (Raspberry Pi)
Run on actual Raspberry Pi hardware:
```bash
# Direct execution
python run.py
# As systemd service
sudo systemctl start ledmatrix
sudo systemctl status ledmatrix
sudo journalctl -u ledmatrix -f # View logs
```
### Plugin-Specific Testing
Test individual plugin loading:
```python
# test_my_plugin.py
from src.plugin_system.plugin_manager import PluginManager
from src.config_manager import ConfigManager
from src.display_manager import DisplayManager
from src.cache_manager import CacheManager
# Initialize managers
config_manager = ConfigManager()
config = config_manager.load_config()
display_manager = DisplayManager(config)
cache_manager = CacheManager()
# Initialize plugin manager
plugin_manager = PluginManager(
plugins_dir="plugins",
config_manager=config_manager,
display_manager=display_manager,
cache_manager=cache_manager
)
# Discover and load plugin
plugins = plugin_manager.discover_plugins()
print(f"Discovered plugins: {plugins}")
if "my-plugin" in plugins:
if plugin_manager.load_plugin("my-plugin"):
plugin = plugin_manager.get_plugin("my-plugin")
plugin.update()
plugin.display()
print("Plugin loaded and displayed successfully!")
else:
print("Failed to load plugin")
```
---
## Loading Plugins
### Enabling Plugins
Plugins are enabled/disabled in `config/config.json`:
```json
{
"my-plugin": {
"enabled": true,
"display_duration": 15,
"api_key": "your-api-key-here"
}
}
```
### Plugin Configuration Structure
Each plugin has its own section in `config/config.json`:
```json
{
"<plugin-id>": {
"enabled": true, // Enable/disable plugin
"display_duration": 15, // Display duration in seconds
"live_priority": false, // Enable live priority takeover
"high_performance_transitions": false, // Use 120 FPS transitions
"transition": { // Transition configuration
"type": "redraw", // Transition type
"speed": 2, // Transition speed
"enabled": true // Enable transitions
},
// ... plugin-specific configuration
}
}
```
### Secrets Management
Store sensitive data (API keys, tokens) in `config/config_secrets.json`
under the same plugin id you use in `config/config.json`:
```json
{
"my-plugin": {
"api_key": "secret-api-key-here"
}
}
```
At load time, the config manager deep-merges `config_secrets.json` into
the main config (verified at `src/config_manager.py:162-172`). So in
your plugin's code:
```python
class MyPlugin(BasePlugin):
def __init__(self, plugin_id, config, display_manager, cache_manager, plugin_manager):
super().__init__(plugin_id, config, display_manager, cache_manager, plugin_manager)
self.api_key = config.get("api_key") # already merged from secrets
```
There is no separate `config_secrets` reference field — just put the
secret value under the same plugin namespace and read it from the
merged config.
### Plugin Discovery
Plugins are automatically discovered when:
- Directory exists in `plugins/`
- Directory contains `manifest.json`
- Manifest has required fields (`id`, `entry_point`, `class_name`)
Check discovered plugins:
```bash
# Using dev_plugin_setup.sh
./scripts/dev/dev_plugin_setup.sh list
# Output shows:
# ✓ plugin-name (symlink)
# → /path/to/repo
# ✓ Git repo is clean (branch: main)
```
### Plugin Status
Check plugin status and git information:
```bash
./scripts/dev/dev_plugin_setup.sh status
# Output shows:
# ✓ plugin-name
# Path: /path/to/repo
# Branch: main
# Remote: https://github.com/user/repo.git
# Status: Clean and up to date
```
---
## Plugin Development Workflow
### 1. Initial Setup
```bash
# Create or clone plugin repository
git clone https://github.com/user/ledmatrix-my-plugin.git
cd ledmatrix-my-plugin
# Link to LEDMatrix project
cd /path/to/LEDMatrix
./scripts/dev/dev_plugin_setup.sh link my-plugin ../ledmatrix-my-plugin
```
### 2. Development Cycle
1. **Edit plugin code** in linked repository
2. **Test with the dev preview server**:
`python3 scripts/dev_server.py` (then open `http://localhost:5001`).
Or run the full display in emulator mode with
`python3 run.py --emulator` (or equivalently
`EMULATOR=true python3 run.py`). The `-e`/`--emulator` CLI flag is
defined in `run.py:19-20` and sets the same `EMULATOR` environment
variable internally.
3. **Check logs** for errors or warnings
4. **Update configuration** in `config/config.json` if needed
5. **Iterate** until plugin works correctly
### 3. Testing on Hardware
```bash
# Deploy to Raspberry Pi
rsync -avz plugins/my-plugin/ ledpi@your-pi-ip:/path/to/LEDMatrix/plugins/my-plugin/
# Or if using git, pull on Pi
ssh ledpi@your-pi-ip "cd /path/to/LEDMatrix/plugins/my-plugin && git pull"
# Restart service
ssh ledpi@your-pi-ip "sudo systemctl restart ledmatrix"
```
### 4. Updating Plugins
```bash
# Update single plugin from git
./scripts/dev/dev_plugin_setup.sh update my-plugin
# Update all linked plugins
./scripts/dev/dev_plugin_setup.sh update
```
### 5. Unlinking Plugins
```bash
# Remove symlink (preserves repository)
./scripts/dev/dev_plugin_setup.sh unlink my-plugin
```
---
## Testing Plugins
### Unit Testing
Create test files in plugin directory:
```python
# plugins/my-plugin/test_my_plugin.py
import unittest
from unittest.mock import Mock, MagicMock
from manager import MyPlugin
class TestMyPlugin(unittest.TestCase):
def setUp(self):
self.config = {"enabled": True}
self.display_manager = Mock()
self.cache_manager = Mock()
self.plugin_manager = Mock()
self.plugin = MyPlugin(
plugin_id="my-plugin",
config=self.config,
display_manager=self.display_manager,
cache_manager=self.cache_manager,
plugin_manager=self.plugin_manager
)
def test_plugin_initialization(self):
self.assertEqual(self.plugin.plugin_id, "my-plugin")
self.assertTrue(self.plugin.enabled)
def test_config_validation(self):
self.assertTrue(self.plugin.validate_config())
def test_update(self):
self.cache_manager.get.return_value = None
self.plugin.update()
# Assert data was fetched and cached
def test_display(self):
self.plugin.display()
self.display_manager.draw_text.assert_called()
self.display_manager.update_display.assert_called()
if __name__ == '__main__':
unittest.main()
```
Run tests:
```bash
cd plugins/my-plugin
python -m pytest test_my_plugin.py
# or
python test_my_plugin.py
```
### Integration Testing
Test plugin with actual managers:
```python
# test_plugin_integration.py
from src.plugin_system.plugin_manager import PluginManager
from src.config_manager import ConfigManager
from src.display_manager import DisplayManager
from src.cache_manager import CacheManager
def test_plugin_loading():
config_manager = ConfigManager()
config = config_manager.load_config()
display_manager = DisplayManager(config)
cache_manager = CacheManager()
plugin_manager = PluginManager(
plugins_dir="plugins",
config_manager=config_manager,
display_manager=display_manager,
cache_manager=cache_manager
)
plugins = plugin_manager.discover_plugins()
assert "my-plugin" in plugins
assert plugin_manager.load_plugin("my-plugin")
plugin = plugin_manager.get_plugin("my-plugin")
assert plugin is not None
assert plugin.enabled
plugin.update()
plugin.display()
```
### Emulator Testing
Test plugin rendering visually:
```bash
# Run with emulator
python run.py --emulator
# Plugin should appear in display rotation
# Check logs for plugin loading and execution
```
### Hardware Testing
1. Deploy plugin to Raspberry Pi
2. Enable in `config/config.json`
3. Restart LEDMatrix service
4. Observe LED matrix display
5. Check logs: `journalctl -u ledmatrix -f`
---
## Troubleshooting
### Plugin Not Loading
**Symptoms**: Plugin doesn't appear in available modes, no logs about plugin
**Solutions**:
1. Check plugin directory exists: `ls plugins/my-plugin/`
2. Verify `manifest.json` exists and is valid JSON
3. Check manifest has required fields: `id`, `entry_point`, `class_name`
4. Verify entry_point file exists: `ls plugins/my-plugin/manager.py`
5. Check class name matches: `grep "class.*Plugin" plugins/my-plugin/manager.py`
6. Review logs for import errors
### Plugin Loading but Not Displaying
**Symptoms**: Plugin loads successfully but doesn't appear in rotation
**Solutions**:
1. Check plugin is enabled: `config/config.json` has `"enabled": true`
2. Verify display_modes in manifest match config
3. Check plugin is in rotation schedule
4. Review `display()` method for errors
5. Check logs for runtime errors
### Configuration Errors
**Symptoms**: Plugin fails to load, validation errors in logs
**Solutions**:
1. Validate config against `config_schema.json`
2. Check required fields are present
3. Verify data types match schema
4. Check for typos in config keys
5. Review `validate_config()` method
### Import Errors
**Symptoms**: ModuleNotFoundError or ImportError in logs
**Solutions**:
1. Install plugin dependencies: `pip install -r plugins/my-plugin/requirements.txt`
2. Check Python path includes plugin directory
3. Verify relative imports are correct
4. Check for circular import issues
5. Ensure all dependencies are in requirements.txt
### Display Issues
**Symptoms**: Plugin renders incorrectly or not at all
**Solutions**:
1. Check display dimensions: `display_manager.width`, `display_manager.height`
2. Verify coordinates are within display bounds
3. Check color values are valid (0-255)
4. Ensure `update_display()` is called after rendering
5. Test with emulator first to debug rendering
### Performance Issues
**Symptoms**: Slow display updates, high CPU usage
**Solutions**:
1. Use `cache_manager` to avoid excessive API calls
2. Implement background data fetching
3. Optimize rendering code
4. Consider using `high_performance_transitions`
5. Profile plugin code to identify bottlenecks
### Git/Symlink Issues
**Symptoms**: Plugin changes not appearing, broken symlinks
**Solutions**:
1. Check symlink: `ls -la plugins/my-plugin`
2. Verify target exists: `readlink -f plugins/my-plugin`
3. Update plugin: `./scripts/dev/dev_plugin_setup.sh update my-plugin`
4. Re-link plugin if needed: `./scripts/dev/dev_plugin_setup.sh unlink my-plugin && ./scripts/dev/dev_plugin_setup.sh link my-plugin <path>`
5. Check git status: `cd plugins/my-plugin && git status`
---
## Best Practices
### Code Organization
- Keep plugin code in `plugins/<plugin-id>/` directory
- Use descriptive class and method names
- Follow existing plugin patterns
- Place shared utilities in `src/common/` if reusable
### Configuration
- Always use `config_schema.json` for validation
- Store secrets in `config_secrets.json`
- Provide sensible defaults
- Document all configuration options in README
### Error Handling
- Use plugin logger for all logging
- Handle API failures gracefully
- Provide fallback displays when data unavailable
- Cache data to avoid excessive requests
### Performance
- Cache API responses appropriately
- Use background data fetching for long operations
- Optimize rendering for Pi's limited resources
- Test performance on actual hardware
### Testing
- Write unit tests for core logic
- Test with emulator before hardware
- Test on Raspberry Pi before deploying
- Test with other plugins enabled
### Documentation
- Document plugin functionality in README
- Include configuration examples
- Document API requirements and rate limits
- Provide usage examples
---
## Resources
- **Plugin System Documentation**: `docs/PLUGIN_ARCHITECTURE_SPEC.md`
- **Base Plugin Class**: `src/plugin_system/base_plugin.py`
- **Plugin Manager**: `src/plugin_system/plugin_manager.py`
- **Example Plugins**:
- `plugins/hockey-scoreboard/` - Sports scoreboard example
- `plugins/football-scoreboard/` - Complex multi-league example
- `plugins/ledmatrix-music/` - Real-time data example
- **Development Setup**: `dev_plugin_setup.sh`
- **Example Config**: `dev_plugins.json.example`
---
## Quick Reference
### Common Commands
```bash
# Link plugin from GitHub
./scripts/dev/dev_plugin_setup.sh link-github <name>
# Link local plugin
./scripts/dev/dev_plugin_setup.sh link <name> <path>
# List all plugins
./scripts/dev/dev_plugin_setup.sh list
# Check plugin status
./scripts/dev/dev_plugin_setup.sh status
# Update plugin(s)
./scripts/dev/dev_plugin_setup.sh update [name]
# Unlink plugin
./scripts/dev/dev_plugin_setup.sh unlink <name>
# Run with emulator
python run.py --emulator
# Run on Pi
python run.py
```
### Plugin File Structure
```
plugins/my-plugin/
├── manifest.json # Required: Plugin metadata
├── manager.py # Required: Plugin class
├── config_schema.json # Required: Config validation
├── requirements.txt # Optional: Dependencies
├── README.md # Optional: Documentation
└── ... # Plugin-specific files
```
### Required Manifest Fields
- `id`: Plugin identifier
- `entry_point`: Python file (usually "manager.py")
- `class_name`: Plugin class name
- `display_modes`: Array of mode names
-38
View File
@@ -1,38 +0,0 @@
---
globs: *.py
---
# Python Coding Standards
## Code Quality Principles
- **Simplicity First**: Prefer clear, readable code over clever optimizations
- **Explicit over Implicit**: Make intentions clear through naming and structure
- **Fail Fast**: Validate inputs and handle errors early
- **Documentation**: Use docstrings for classes and complex functions
## Naming Conventions
- **Classes**: PascalCase (e.g., `NHLRecentManager`)
- **Functions/Variables**: snake_case (e.g., `fetch_game_data`)
- **Constants**: UPPER_SNAKE_CASE (e.g., `ESPN_NHL_SCOREBOARD_URL`)
- **Private methods**: Leading underscore (e.g., `_fetch_data`)
## Error Handling
- **Logging**: Use structured logging with context (e.g., `[NHL Recent]`)
- **Exceptions**: Catch specific exceptions, not bare `except:`
- **User-friendly messages**: Explain what went wrong and potential solutions
- **Graceful degradation**: Continue operation when non-critical features fail
## Manager Pattern
All sports managers should follow this structure:
```python
class BaseManager:
def __init__(self, config, display_manager, cache_manager)
def update(self) # Fetch and process data
def display(self, force_clear=False) # Render to display
```
## Configuration Management
- **Type hints**: Use for function parameters and return values
- **Configuration validation**: Check required fields on initialization
- **Default values**: Provide sensible defaults in code, not config
- **Environment awareness**: Handle different deployment contexts
@@ -1,42 +0,0 @@
---
globs: config/*.json,src/*.py
---
# Configuration Management
## Configuration Structure
- **Main config**: [config/config.json](mdc:config/config.json) - Primary configuration
- **Secrets**: [config/config_secrets.json](mdc:config/config_secrets.json) - API keys and sensitive data
- **Templates**: [config/config.template.json](mdc:config/config.template.json) - Default values
## Configuration Principles
- **Validation**: Check required fields and data types on startup
- **Defaults**: Provide sensible defaults in code, not just config
- **Environment awareness**: Handle development vs production differences
- **Security**: Never commit secrets to version control
## Manager Configuration Pattern
```python
def __init__(self, config, display_manager, cache_manager):
self.mode_config = config.get("sport_scoreboard", {})
self.favorite_teams = self.mode_config.get("favorite_teams", [])
self.show_favorite_only = self.mode_config.get("show_favorite_teams_only", False)
```
## Required Configuration Sections
- **Display settings**: Update intervals, display durations
- **API settings**: Timeouts, retry logic, rate limiting
- **Background service**: Threading, caching, priority settings
- **Team preferences**: Favorite teams, filtering options
## Configuration Validation
- **Type checking**: Ensure numeric values are numbers, lists are lists
- **Range validation**: Check that intervals are reasonable
- **Dependency checking**: Verify required services are available
- **Fallback values**: Provide defaults when config is missing or invalid
## Best Practices
- **Documentation**: Comment complex configuration options
- **Examples**: Provide working examples in templates
- **Migration**: Handle configuration changes between versions
- **Testing**: Validate configuration in test environments
-50
View File
@@ -1,50 +0,0 @@
---
globs: src/*.py
---
# Error Handling and Logging
## Logging Standards
- **Structured prefixes**: Use consistent tags like `[NHL Recent]`, `[NFL Live]`
- **Context information**: Include relevant details (team names, game status, dates)
- **Appropriate levels**:
- `info`: Normal operations and status updates
- `debug`: Detailed information for troubleshooting
- `warning`: Non-critical issues that should be noted
- `error`: Problems that need attention
## Error Handling Patterns
```python
try:
data = self._fetch_data()
if not data or 'events' not in data:
self.logger.warning("[Manager] No events found in API response")
return
except requests.exceptions.RequestException as e:
self.logger.error(f"[Manager] API error: {e}")
return None
```
## User-Friendly Messages
- **Explain the situation**: "No games available during off-season"
- **Provide context**: "NHL season typically runs October-June"
- **Suggest solutions**: "Check back when season starts"
- **Distinguish issues**: API problems vs no data vs filtering results
## Graceful Degradation
- **Fallback content**: Show alternative games when favorites unavailable
- **Cached data**: Use cached data when API fails
- **Service continuity**: Continue operation when non-critical features fail
- **Clear communication**: Explain what's happening to users
## Debugging Support
- **Comprehensive logging**: Log API responses, filtering results, display updates
- **State tracking**: Log current state and transitions
- **Performance monitoring**: Track timing and resource usage
- **Error context**: Include stack traces for debugging
## Off-Season Awareness
- **Seasonal messaging**: Different messages for different times of year
- **Helpful context**: Explain why no games are available
- **Future planning**: Mention when season starts
- **Realistic expectations**: Set appropriate expectations during off-season
-51
View File
@@ -1,51 +0,0 @@
---
alwaysApply: true
---
# Git Workflow and Branching
## Branch Naming Conventions
- **Features**: `feature/description-of-feature` (e.g., `feature/weather-forecast-improvements`)
- **Bug fixes**: `fix/description-of-bug` (e.g., `fix/nhl-manager-improvements`)
- **Hotfixes**: `hotfix/critical-issue-description`
- **Refactoring**: `refactor/description-of-refactor`
## Commit Message Format
```
type(scope): description
[optional body]
[optional footer]
```
**Types**: feat, fix, docs, style, refactor, test, chore
**Examples**:
- `feat(nhl): Add enhanced logging for data visibility`
- `fix(display): Resolve rendering performance issue`
- `docs(api): Update ESPN API integration guide`
## Pull Request Guidelines
- **Self-review**: Review your own PR before requesting review
- **Testing**: Test thoroughly on Raspberry Pi hardware
- **Documentation**: Update relevant documentation if needed
- **Clean history**: Squash commits if necessary for clean history
## Code Review Checklist
- **Code Quality**: Proper error handling, logging, type hints
- **Architecture**: Follows project patterns, doesn't break existing functionality
- **Performance**: No negative impact on display performance
- **Testing**: Works on Raspberry Pi hardware
- **Documentation**: Comments added for complex logic
## Merge Strategies
- **Squash and Merge**: Preferred for feature branches and bug fixes
- **Merge Commit**: For complex features with multiple logical commits
- **Rebase and Merge**: For simple, single-commit changes
## Best Practices
- **Keep branches small and focused**
- **Commit frequently with meaningful messages**
- **Update branch regularly with main**
- **Test changes incrementally**
- **Delete feature branches after merge**
-213
View File
@@ -1,213 +0,0 @@
---
description: GitHub branching and pull request best practices for LEDMatrix project
globs: ["**/*.py", "**/*.md", "**/*.json", "**/*.sh"]
alwaysApply: true
---
# GitHub Branching and Pull Request Guidelines
## Branch Naming Conventions
### Feature Branches
- **Format**: `feature/description-of-feature`
- **Examples**:
- `feature/weather-forecast-improvements`
- `feature/stock-api-integration`
- `feature/nba-live-scores`
### Bug Fix Branches
- **Format**: `fix/description-of-bug`
- **Examples**:
- `fix/leaderboard-scrolling-performance`
- `fix/weather-api-timeout`
- `fix/display-rendering-issue`
### Hotfix Branches
- **Format**: `hotfix/critical-issue-description`
- **Examples**:
- `hotfix/display-crash-fix`
- `hotfix/api-rate-limit-fix`
### Refactoring Branches
- **Format**: `refactor/description-of-refactor`
- **Examples**:
- `refactor/sports-manager-architecture`
- `refactor/cache-management-system`
## Branch Management Rules
### Main Branch Protection
- **`main`** branch is protected and requires PR reviews
- Never commit directly to `main`
- All changes must go through pull requests
### Branch Lifecycle
1. **Create** branch from `main` when starting work
2. **Keep** branch up-to-date with `main` regularly
3. **Test** thoroughly before creating PR
4. **Delete** branch after successful merge
### Branch Updates
```bash
# Before starting new work
git checkout main
git pull origin main
# Create new branch
git checkout -b feature/your-feature-name
# Keep branch updated during development
git checkout main
git pull origin main
git checkout feature/your-feature-name
git merge main
```
## Pull Request Guidelines
### PR Title Format
- **Feature**: `feat: Add weather forecast improvements`
- **Fix**: `fix: Resolve leaderboard scrolling performance issue`
- **Refactor**: `refactor: Improve sports manager architecture`
- **Docs**: `docs: Update API integration guide`
- **Test**: `test: Add unit tests for weather manager`
### PR Description Template
```markdown
## Description
Brief description of changes and motivation.
## Type of Change
- [ ] Bug fix (non-breaking change)
- [ ] New feature (non-breaking change)
- [ ] Breaking change (fix or feature that would cause existing functionality to not work as expected)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Refactoring
## Testing
- [ ] Tested on Raspberry Pi hardware
- [ ] Verified display rendering works correctly
- [ ] Checked API integration functionality
- [ ] Tested error handling scenarios
## Screenshots/Videos
(If applicable, add screenshots or videos of the changes)
## Checklist
- [ ] Code follows project style guidelines
- [ ] Self-review completed
- [ ] Comments added for complex logic
- [ ] No hardcoded values or API keys
- [ ] Error handling implemented
- [ ] Logging added where appropriate
```
### PR Review Requirements
#### For Reviewers
- **Code Quality**: Check for proper error handling, logging, and type hints
- **Architecture**: Ensure changes follow project patterns and don't break existing functionality
- **Performance**: Verify changes don't negatively impact display performance
- **Testing**: Confirm changes work on Raspberry Pi hardware
- **Documentation**: Check if documentation needs updates
#### For Authors
- **Self-Review**: Review your own PR before requesting review
- **Testing**: Test thoroughly on Pi hardware before submitting
- **Documentation**: Update relevant documentation if needed
- **Clean History**: Squash commits if necessary for clean history
## Commit Message Guidelines
### Format
```
type(scope): description
[optional body]
[optional footer]
```
### Types
- **feat**: New feature
- **fix**: Bug fix
- **docs**: Documentation changes
- **style**: Code style changes (formatting, etc.)
- **refactor**: Code refactoring
- **test**: Adding or updating tests
- **chore**: Maintenance tasks
### Examples
```
feat(weather): Add hourly forecast display
fix(nba): Resolve live score update issue
docs(api): Update ESPN API integration guide
refactor(sports): Improve base class architecture
```
## Merge Strategies
### Squash and Merge (Preferred)
- Use for feature branches and bug fixes
- Creates clean, linear history
- Combines all commits into single commit
### Merge Commit
- Use for complex features with multiple logical commits
- Preserves commit history
- Use when commit messages are meaningful
### Rebase and Merge
- Use sparingly for simple, single-commit changes
- Creates linear history without merge commits
## Release Management
### Version Tags
- Use semantic versioning: `v1.2.3`
- Tag releases on `main` branch
- Create release notes with technical details
### Release Branches
- **Format**: `release/v1.2.3`
- Use for release preparation
- Include version bumps and final testing
## Emergency Procedures
### Hotfix Process
1. Create `hotfix/` branch from `main`
2. Make minimal fix
3. Test thoroughly
4. Create PR with expedited review
5. Merge to `main` and tag release
6. Cherry-pick to other branches if needed
### Rollback Process
1. Identify last known good commit
2. Create revert PR if possible
3. Use `git revert` for clean rollback
4. Tag rollback release
5. Document issue and resolution
## Best Practices
### Before Creating PR
- [ ] Run all tests locally
- [ ] Test on Raspberry Pi hardware
- [ ] Check for linting errors
- [ ] Update documentation if needed
- [ ] Ensure commit messages are clear
### During Development
- [ ] Keep branches small and focused
- [ ] Commit frequently with meaningful messages
- [ ] Update branch regularly with main
- [ ] Test changes incrementally
### After PR Approval
- [ ] Delete feature branch after merge
- [ ] Update local main branch
- [ ] Verify changes work in production
- [ ] Update any related documentation
-23
View File
@@ -1,23 +0,0 @@
---
alwaysApply: true
---
# LEDMatrix Project Structure
## Core Architecture
- **Main entry point**: [run.py](mdc:run.py) - Primary application launcher
- **Configuration**: [config/config.json](mdc:config/config.json) - Main configuration file
- **Display management**: [src/display_controller.py](mdc:src/display_controller.py) - Core display logic
- **Web interface**: [web_interface_v2.py](mdc:web_interface_v2.py) - Modern web UI
## Source Code Organization
- **Managers**: [src/](mdc:src/) - All sports/weather/stock managers
- **Assets**: [assets/](mdc:assets/) - Logos, fonts, and static resources
- **Tests**: [test/](mdc:test/) - Unit and integration tests
- **Documentation**: [LEDMatrix.wiki/](mdc:LEDMatrix.wiki/) - Comprehensive guides
## Key Design Principles
- **Single Responsibility**: Each manager handles one sport/domain
- **Consistent Patterns**: All managers follow similar structure
- **Configuration-Driven**: Behavior controlled via [config/config.json](mdc:config/config.json)
- **Raspberry Pi Focus**: Optimized for Pi hardware, not Windows development
@@ -1,41 +0,0 @@
---
alwaysApply: true
---
# Raspberry Pi Development Guidelines
## Hardware Constraints
- **Pi-only execution**: Code must run on Raspberry Pi, not Windows development machine
- **LED matrix library**: Uses [rpi-rgb-led-matrix-master/](mdc:rpi-rgb-led-matrix-master/) for hardware control
- **Memory limitations**: Optimize for Pi's limited RAM
- **Performance**: Consider Pi's CPU capabilities in design
## Development Workflow
- **Local development**: Write and test code on Windows
- **Pi deployment**: Deploy and test on actual Pi hardware
- **SSH access**: Use SSH for Pi-based testing and debugging
- **Service management**: Use systemd services for production deployment
## Testing Strategy
- **Unit tests**: Test logic without hardware dependencies
- **Integration tests**: Test with mock display managers
- **Hardware tests**: Validate on actual Pi with LED matrix
- **Performance tests**: Monitor memory and CPU usage
## Deployment Considerations
- **Service files**: [ledmatrix.service](mdc:ledmatrix.service), [ledmatrix-web.service](mdc:ledmatrix-web.service)
- **Installation scripts**: [first_time_install.sh](mdc:first_time_install.sh), [install_service.sh](mdc:install_service.sh)
- **Dependencies**: [requirements.txt](mdc:requirements.txt) for Pi environment
- **Permissions**: Handle file permissions for Pi user
## Performance Optimization
- **Caching**: Use [src/cache_manager.py](mdc:src/cache_manager.py) for data persistence
- **Background services**: Non-blocking data fetching
- **Memory management**: Clean up resources regularly
- **Display optimization**: Minimize unnecessary redraws
## Debugging on Pi
- **Logging**: Comprehensive logging for remote debugging
- **Error reporting**: Clear error messages for troubleshooting
- **Status monitoring**: Health checks and status reporting
- **Remote access**: Web interface for configuration and monitoring
-42
View File
@@ -1,42 +0,0 @@
---
globs: src/*_managers.py
---
# Sports Manager Development
## Manager Architecture
All sports managers inherit from base classes and follow consistent patterns:
- **Base classes**: [src/nhl_managers.py](mdc:src/nhl_managers.py), [src/nfl_managers.py](mdc:src/nfl_managers.py)
- **Common functionality**: Data fetching, caching, display rendering
- **Configuration-driven**: Behavior controlled via config sections
## Required Methods
```python
def __init__(self, config, display_manager, cache_manager)
def update(self) # Fetch fresh data
def display(self, force_clear=False) # Render current data
```
## Data Flow Pattern
1. **Fetch**: Get data from API (with caching)
2. **Process**: Extract relevant game information
3. **Filter**: Apply favorite team preferences
4. **Display**: Render to LED matrix
## Logging Standards
- **Structured prefixes**: `[NHL Recent]`, `[NFL Live]`, etc.
- **Context information**: Include team names, game status, dates
- **Debug levels**: Use appropriate log levels (info, debug, warning, error)
- **User-friendly messages**: Explain what's happening and why
## Error Handling
- **API failures**: Log and continue with cached data if available
- **No data scenarios**: Distinguish between API issues vs no games available
- **Off-season awareness**: Provide helpful context during non-active periods
- **Fallback behavior**: Show alternative content when preferred content unavailable
## Configuration Integration
- **Required settings**: Validate on initialization
- **Optional settings**: Provide sensible defaults
- **Background service**: Use for non-blocking data fetching
- **Caching strategy**: Implement intelligent cache management
-51
View File
@@ -1,51 +0,0 @@
---
globs: test/*.py,src/*.py
---
# Testing Standards
## Test Organization
- **Test directory**: [test/](mdc:test/) - All test files
- **Unit tests**: Test individual components in isolation
- **Integration tests**: Test component interactions
- **Hardware tests**: Validate on Raspberry Pi with actual LED matrix
## Testing Principles
- **Test behavior, not implementation**: Focus on what the code does, not how
- **Mock external dependencies**: Use mocks for APIs, display managers, cache
- **Test edge cases**: Empty data, API failures, configuration errors
- **Pi-specific testing**: Validate hardware integration
## Test Structure
```python
def test_manager_initialization():
"""Test that manager initializes with valid config"""
config = {"sport_scoreboard": {"enabled": True}}
manager = ManagerClass(config, mock_display, mock_cache)
assert manager.enabled == True
def test_api_failure_handling():
"""Test graceful handling of API failures"""
# Test that system continues when API fails
# Verify fallback to cached data
# Check appropriate error logging
```
## Mock Patterns
- **Display Manager**: Mock for testing without hardware
- **Cache Manager**: Mock for testing data persistence
- **API responses**: Mock for consistent test data
- **Configuration**: Use test-specific configs
## Test Categories
- **Unit tests**: Individual manager methods
- **Integration tests**: Manager interactions with services
- **Configuration tests**: Validate config loading and validation
- **Error handling tests**: API failures, invalid data, edge cases
## Testing Best Practices
- **Descriptive names**: Test names should explain what they test
- **Single responsibility**: Each test should verify one thing
- **Independent tests**: Tests should not depend on each other
- **Clean setup/teardown**: Reset state between tests
- **Pi compatibility**: Ensure tests work in Pi environment
-1
View File
@@ -1 +0,0 @@
# Add directories or file patterns to ignore during indexing (e.g. foo/ or *.csv)
-364
View File
@@ -1,364 +0,0 @@
# LEDMatrix Plugin Development Rules
## Plugin System Overview
The LEDMatrix project uses a plugin-based architecture. All display
functionality (except core calendar) is implemented as plugins that are
dynamically loaded from the directory configured by
`plugin_system.plugins_directory` in `config.json` — the default is
`plugin-repos/` (per `config/config.template.json:130`).
> **Fallback note (scoped):** `PluginManager.discover_plugins()`
> (`src/plugin_system/plugin_manager.py:154`) only scans the
> configured directory — there is no fallback to `plugins/` in the
> main discovery path. A fallback to `plugins/` does exist in two
> narrower places:
> - `store_manager.py:1700-1718` — store operations (install/update/
> uninstall) check `plugins/` if the plugin isn't found in the
> configured directory, so plugin-store flows work even when your
> dev symlinks live in `plugins/`.
> - `schema_manager.py:70-80` — `get_schema_path()` probes both
> `plugins/` and `plugin-repos/` for `config_schema.json` so the
> web UI form generation finds the schema regardless of where the
> plugin lives.
>
> The dev workflow in `scripts/dev/dev_plugin_setup.sh` creates
> symlinks under `plugins/`, which is why the store and schema
> fallbacks exist. For day-to-day development, set
> `plugin_system.plugins_directory` to `plugins` so the main
> discovery path picks up your symlinks.
## Plugin Structure
### Required Files
- **manifest.json**: Plugin metadata, entry point, class name, dependencies
- **manager.py**: Main plugin class (must inherit from `BasePlugin`)
- **config_schema.json**: JSON schema for plugin configuration validation
- **requirements.txt**: Python dependencies (if any)
- **README.md**: Plugin documentation
### Plugin Class Requirements
- Must inherit from `src.plugin_system.base_plugin.BasePlugin`
- Must implement `update()` method for data fetching
- Must implement `display()` method for rendering
- Should implement `validate_config()` for configuration validation
- Optional: Override `has_live_content()` for live priority features
## Plugin Development Workflow
### 1. Creating a New Plugin
**Option A: Use dev_plugin_setup.sh (Recommended)**
```bash
# Link from GitHub
./scripts/dev/dev_plugin_setup.sh link-github <plugin-name>
# Link local repository
./scripts/dev/dev_plugin_setup.sh link <plugin-name> <path-to-repo>
```
**Option B: Manual Setup**
1. Create directory in `plugin-repos/<plugin-id>/` (or `plugins/<plugin-id>/`
if you're using the dev fallback location)
2. Add `manifest.json` with required fields
3. Create `manager.py` with plugin class
4. Add `config_schema.json` for configuration
5. Enable plugin in `config/config.json` under `"<plugin-id>": {"enabled": true}`
### 2. Plugin Configuration
Plugins are configured in `config/config.json`:
```json
{
"<plugin-id>": {
"enabled": true,
"display_duration": 15,
"live_priority": false,
"high_performance_transitions": false,
"transition": {
"type": "redraw",
"speed": 2,
"enabled": true
},
// ... plugin-specific config
}
}
```
### 3. Testing Plugins
**On Development Machine:**
- Run the dev preview server: `python3 scripts/dev_server.py` (then
open `http://localhost:5001`) — renders plugins in the browser
without running the full display loop
- Or run the full display in emulator mode:
`python3 run.py --emulator` (or equivalently
`EMULATOR=true python3 run.py`, or `./scripts/dev/run_emulator.sh`).
The `-e`/`--emulator` CLI flag is defined in `run.py:19-20`.
- Test plugin loading: Check logs for plugin discovery and loading
- Validate configuration: Ensure config matches `config_schema.json`
**On Raspberry Pi:**
- Deploy and test on actual hardware
- Monitor logs: `journalctl -u ledmatrix -f` (if running as service)
- Check plugin status in web interface
### 4. Plugin Development Best Practices
**Code Organization:**
- Keep plugin code in `plugin-repos/<plugin-id>/` (or its dev-time
symlink in `plugins/<plugin-id>/`)
- Use shared assets from `assets/` directory when possible
- Follow existing plugin patterns — canonical sources live in the
[`ledmatrix-plugins`](https://github.com/ChuckBuilds/ledmatrix-plugins)
repo (`plugins/hockey-scoreboard/`, `plugins/football-scoreboard/`,
`plugins/clock-simple/`, etc.)
- Place shared utilities in `src/common/` if reusable across plugins
**Configuration Management:**
- Use `config_schema.json` for validation
- Store secrets in `config/config_secrets.json` under the same plugin
id namespace as the main config — they're deep-merged into the main
config at load time (`src/config_manager.py:162-172`), so plugin
code reads them directly from `config.get(...)` like any other key
- There is no separate `config_secrets` reference field
- Validate all required fields in `validate_config()`
**Error Handling:**
- Use plugin's logger: `self.logger.info/error/warning()`
- Handle API failures gracefully
- Cache data to avoid excessive API calls
- Provide fallback displays when data unavailable
**Performance:**
- Use `cache_manager` for API response caching
- Implement background data fetching if needed
- Use `high_performance_transitions` for smoother animations
- Optimize rendering for Pi's limited resources
**Display Rendering:**
- Use `display_manager` for all drawing operations
- Support different display sizes (check `display_manager.width/height`)
- Use `apply_transition()` for smooth transitions between displays
- Clear display before rendering: `display_manager.clear()`
- Always call `display_manager.update_display()` after rendering
## Plugin API Reference
### BasePlugin Class
Located in: `src/plugin_system/base_plugin.py`
**Required Methods:**
- `update()`: Fetch/update data (called based on `update_interval` in manifest)
- `display(force_clear=False)`: Render plugin content
**Optional Methods:**
- `validate_config()`: Validate plugin configuration
- `has_live_content()`: Return True if plugin has live/urgent content
- `get_live_modes()`: Return list of modes for live priority
- `cleanup()`: Clean up resources on unload
- `on_config_change(new_config)`: Handle config updates
- `on_enable()`: Called when plugin enabled
- `on_disable()`: Called when plugin disabled
**Available Properties:**
- `self.plugin_id`: Plugin identifier
- `self.config`: Plugin configuration dict
- `self.display_manager`: Display manager instance
- `self.cache_manager`: Cache manager instance
- `self.plugin_manager`: Plugin manager reference
- `self.logger`: Plugin-specific logger
- `self.enabled`: Boolean enabled status
- `self.transition_manager`: Transition system (if available)
### Display Manager
Located in: `src/display_manager.py`
**Key Methods:**
- `clear()`: Clear the display
- `draw_text(text, x, y, color, font, small_font, centered)`: Draw text
- `update_display()`: Push the buffer to the physical display
- `draw_weather_icon(condition, x, y, size)`: Draw a weather icon
- `width`, `height`: Display dimensions
**Image rendering**: there is no `draw_image()` helper. Paste directly
onto the underlying PIL Image:
```python
self.display_manager.image.paste(pil_image, (x, y))
self.display_manager.update_display()
```
For transparency, paste with a mask: `image.paste(rgba, (x, y), rgba)`.
### Cache Manager
Located in: `src/cache_manager.py`
**Key Methods:**
- `get(key, max_age=300)`: Get cached value (returns None if missing/stale)
- `set(key, value, ttl=None)`: Cache a value
- `delete(key)` / `clear_cache(key=None)`: Remove a single cache entry,
or (for `clear_cache` with no argument) every cached entry. `delete`
is an alias for `clear_cache(key)`.
- `get_cached_data_with_strategy(key, data_type)`: Cache get with
data-type-aware TTL strategy
- `get_background_cached_data(key, sport_key)`: Cache get for the
background-fetch service path
## Plugin Manifest Schema
Required fields in `manifest.json`:
- `id`: Unique plugin identifier (matches directory name)
- `name`: Human-readable plugin name
- `version`: Semantic version (e.g., "1.0.0")
- `entry_point`: Python file (usually "manager.py")
- `class_name`: Plugin class name (must match class in entry_point)
- `display_modes`: Array of mode names this plugin provides
Common optional fields:
- `description`: Plugin description
- `author`: Plugin author
- `homepage`: Plugin homepage URL
- `category`: Plugin category (e.g., "sports", "weather")
- `tags`: Array of tags
- `update_interval`: Seconds between update() calls (default: 60)
- `default_duration`: Default display duration (default: 15)
- `requires`: Python version, display size requirements
- `config_schema`: Path to config schema file
- `api_requirements`: API dependencies and rate limits
## Plugin Loading Process
1. **Discovery**: PluginManager scans `plugins/` directory for directories containing `manifest.json`
2. **Validation**: Validates manifest structure and required fields
3. **Loading**: Imports plugin module and instantiates plugin class
4. **Configuration**: Loads plugin config from `config/config.json`
5. **Validation**: Calls `validate_config()` on plugin instance
6. **Registration**: Adds plugin to available modes and stores instance
7. **Enablement**: Calls `on_enable()` if plugin is enabled
## Common Plugin Patterns
### Sports Scoreboard Plugin
- Use `background_data_service.py` pattern for API fetching
- Implement live/recent/upcoming game modes
- Use `scoreboard_renderer.py` for consistent rendering
- Support team filtering and game filtering
- Use shared sports logos from `assets/sports/`
### Data Display Plugin
- Fetch data in `update()` method
- Cache API responses using `cache_manager`
- Render in `display()` method
- Handle API errors gracefully
- Provide configuration for refresh intervals
### Real-time Content Plugin
- Implement `has_live_content()` for live priority
- Use `get_live_modes()` to specify which modes are live
- Set `live_priority: true` in config to enable live takeover
- Update data frequently when live content exists
## Debugging Plugins
**Check Plugin Loading:**
- Review logs for plugin discovery messages
- Verify manifest.json syntax is valid JSON
- Check that class_name matches actual class name
- Ensure entry_point file exists and is importable
**Check Plugin Execution:**
- Add logging statements in `update()` and `display()`
- Use `self.logger` for plugin-specific logging
- Check cache_manager for cached data
- Verify display_manager is rendering correctly
**Common Issues:**
- Import errors: Check Python path and dependencies
- Config errors: Validate against config_schema.json
- Display issues: Check display dimensions and coordinate calculations
- Performance: Monitor CPU/memory usage on Pi
## Plugin Testing
**Unit Tests:**
- Test plugin class instantiation
- Test `update()` data fetching logic
- Test `display()` rendering logic
- Test `validate_config()` with various configs
- Mock `display_manager` and `cache_manager` for testing
**Integration Tests:**
- Test plugin loading via PluginManager
- Test plugin with actual config
- Test plugin with emulator display
- Test plugin with cache_manager
**Hardware Tests:**
- Test on Raspberry Pi with LED matrix
- Verify display rendering on actual hardware
- Test performance under load
- Test with other plugins enabled
## File Organization
```
plugins/
<plugin-id>/
manifest.json # Plugin metadata
manager.py # Main plugin class
config_schema.json # Config validation schema
requirements.txt # Python dependencies
README.md # Plugin documentation
# Plugin-specific files
data_manager.py
renderer.py
etc.
```
## Git Workflow for Plugins
**Plugin Development:**
- Plugins are typically separate repositories
- Use `dev_plugin_setup.sh` to link plugins for development
- Symlinks are used to connect plugin repos to `plugins/` directory
- Plugin repos follow naming: `ledmatrix-<plugin-name>`
**Branching:**
- Develop plugins in feature branches
- Follow project branching conventions
- Test plugins before merging to main
**Automatic Version Bumping:**
- **Automatic Version Management**: Version bumping is handled automatically via the pre-push git hook - no manual version bumping is required for normal development workflows
- **GitHub as Source of Truth**: Plugin store always fetches latest versions from GitHub (releases/tags/manifest/commit)
- **Pre-Push Hook**: Automatically bumps patch version and creates git tags when pushing code changes
- The hook is self-contained (no external dependencies) and works on any dev machine
- Installation: Copy the hook from LEDMatrix repo to your plugin repo:
```bash
# From your plugin repository directory
cp /path/to/LEDMatrix/scripts/git-hooks/pre-push-plugin-version .git/hooks/pre-push
chmod +x .git/hooks/pre-push
```
- Or use the installer script from the main LEDMatrix repo (one-time setup)
- The hook automatically:
1. Bumps the patch version (x.y.Z) in manifest.json when code changes are detected
2. Creates a git tag (v{version}) for the new version
3. Stages manifest.json for commit
- Skip auto-tagging: Set `SKIP_TAG=1` environment variable before pushing
- **Manual Version Bumping (Edge Cases Only)**: Manual version bumps are only needed in rare circumstances:
- CI/CD pipelines that bypass git hooks
- Forked repositories without the pre-push hook installed
- Major/minor version bumps (hook only handles patch versions)
- When skipping auto-tagging but still needing a version bump
- For manual bumps, use the standalone script: `scripts/bump_plugin_version.py`
- **Registry**: The plugin registry (plugins.json) stores only metadata (name, description, repo URL) - no versions
- **Version Priority**: Plugin store checks versions in this order: GitHub Releases → GitHub Tags → Manifest from branch → Git commit hash
## Resources
- Plugin System Docs: `docs/PLUGIN_ARCHITECTURE_SPEC.md`
- Plugin Examples: `plugins/hockey-scoreboard/`, `plugins/football-scoreboard/`
- Base Plugin: `src/plugin_system/base_plugin.py`
- Plugin Manager: `src/plugin_system/plugin_manager.py`
- Development Setup: `dev_plugin_setup.sh`
- Example Config: `dev_plugins.json.example`
+10
View File
@@ -1,2 +1,12 @@
# Auto detect text files and perform LF normalization
* text=auto
# Files the Pi executes must stay LF even in a Windows checkout with
# core.autocrlf=true: a CRLF shebang line fails with "bad interpreter",
# and systemd rejects CRLF unit files.
*.sh text eol=lf
*.service text eol=lf
# Generated by scripts/build_css.py; collapsed in diffs, not hand-edited.
web_interface/static/v3/tailwind.css linguist-generated=true
web_interface/static/v3/plugin-frame.css linguist-generated=true
+40
View File
@@ -0,0 +1,40 @@
name: Claude Code Review
on:
pull_request:
types: [opened, synchronize, ready_for_review, reopened]
jobs:
claude-review:
# Pull requests from forks get no repository secrets, so without this
# guard every outside contributor's PR showed this check red for a reason
# they can't fix. Skipped checks don't block merging.
if: github.event.pull_request.head.repo.full_name == github.repository
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
issues: read
id-token: write
steps:
- name: Checkout repository
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
fetch-depth: 1
- name: Run Claude Code Review
id: claude-review
uses: anthropics/claude-code-action@756cc22e19660d20e8cc9496b4f242475a7f7790 # v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
# Review PRs opened by the Claude GitHub App. Without this the action
# aborts before reading the diff ("Workflow initiated by non-human
# actor"), so every such PR shows this check red. Named rather than
# '*': the allow-list is matched against the triggering actor, so
# this admits claude[bot] alone and no other app.
allowed_bots: 'claude'
plugin_marketplaces: 'https://github.com/anthropics/claude-code.git'
plugins: 'code-review@claude-code-plugins'
prompt: '/code-review:code-review ${{ github.repository }}/pull/${{ github.event.pull_request.number }}'
+42
View File
@@ -0,0 +1,42 @@
name: Claude Code
on:
issue_comment:
types: [created]
pull_request_review_comment:
types: [created]
issues:
types: [opened, assigned]
pull_request_review:
types: [submitted]
jobs:
claude:
if: |
(github.event_name == 'issue_comment' && contains(github.event.comment.body, '@claude')) ||
(github.event_name == 'pull_request_review_comment' && contains(github.event.comment.body, '@claude')) ||
(github.event_name == 'pull_request_review' && contains(github.event.review.body, '@claude')) ||
(github.event_name == 'issues' && (contains(github.event.issue.body, '@claude') || contains(github.event.issue.title, '@claude')))
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: read
issues: read
id-token: write
actions: read # Required for Claude to read CI results on PRs
steps:
- name: Checkout repository
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
fetch-depth: 1
- name: Run Claude Code
id: claude
uses: anthropics/claude-code-action@756cc22e19660d20e8cc9496b4f242475a7f7790 # v1
with:
claude_code_oauth_token: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }}
# This is an optional setting that allows Claude to read CI results on PRs
additional_permissions: |
actions: read
@@ -0,0 +1,40 @@
name: Release version check
# A release tag, the CHANGELOG, and src.__version__ must agree. They have not
# always: v3.1.0 was tagged while src/__init__.py still said "1.0.0", which
# silently exempted every device installed from that release from plugin
# compatibility warnings. See docs/SPORTS_UNIFICATION.md (phase B4).
on:
push:
tags: ["v*"]
release:
types: [published]
# Pre-flight: run this against the tag you are about to create.
workflow_dispatch:
inputs:
tag:
description: "Tag to check (e.g. v3.2.0)"
required: true
type: string
permissions:
contents: read
jobs:
version-matches-tag:
name: Tag matches src.__version__
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
# No dependencies: the script reads src/__init__.py and CHANGELOG.md only.
- name: Assert the tag, CHANGELOG and src.__version__ agree
run: python scripts/check_release_version.py "${TAG}"
env:
TAG: ${{ inputs.tag || github.ref_name }}
+197
View File
@@ -0,0 +1,197 @@
name: Tests
on:
pull_request:
push:
branches: [main]
# Manual runs against any branch — useful when a PR's automatic run
# needs a re-run or didn't get created.
workflow_dispatch:
# The jobs only check out the repo and run the tests.
permissions:
contents: read
jobs:
plugin-safety:
name: Plugin safety harness + unit tests
runs-on: ubuntu-latest
env:
# The bundled fixture plugin gives the harness at least one real plugin
# to render, and REQUIRE_PLUGINS turns "discovered zero plugins" into a
# hard failure instead of a silent all-skip green run.
LEDMATRIX_PLUGINS_DIR: test/fixtures/plugins
LEDMATRIX_REQUIRE_PLUGINS: "1"
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
cache: pip
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt -r web_interface/requirements.txt -r requirements-test.txt
pip install RGBMatrixEmulator
- name: Run plugin safety harness
run: |
pytest --no-cov test/plugins/
unit-tests:
name: Core unit tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
cache: pip
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt -r web_interface/requirements.txt -r requirements-test.txt
pip install RGBMatrixEmulator
# Run the ENTIRE test tree (except test/plugins, which the
# plugin-safety job owns). New test files are enrolled automatically;
# excluding anything requires a visible, commented --ignore here.
# Coverage is measured and enforced only in this step — pytest.ini
# deliberately carries no coverage flags so local runs stay fast.
- name: Run core unit suites
run: |
pytest -m "not hardware" test/ \
--ignore=test/plugins \
--cov=src --cov=web_interface \
--cov-report=term \
--cov-fail-under=52
js-tests:
name: Web UI JS tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
cache: pip
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0
with:
node-version: "22"
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt -r web_interface/requirements.txt
npm install --no-audit --no-fund --prefix test/js
# The DOM suites test the real server-rendered pages and API, so they
# need the web interface running. REQUIRE_DOM turns "couldn't reach it"
# into a failure instead of a silent skip.
- name: Start the web interface
run: |
EMULATOR=true python -c "from web_interface.app import app; app.run(host='127.0.0.1', port=5000, threaded=True)" > web.log 2>&1 &
for i in $(seq 60); do curl -sf -o /dev/null http://127.0.0.1:5000/ && exit 0; sleep 1; done
cat web.log
exit 1
- name: Run JS suites
env:
BASE: http://127.0.0.1:5000
REQUIRE_DOM: "1"
run: node test/js/run_all.js
css-build:
name: Tailwind CSS is up to date
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
# Downloads the pinned standalone Tailwind CLI (SHA-256 checked; no
# Node), rebuilds static/v3/tailwind.css and plugin-frame.css from the
# templates and JS, and fails if the committed files differ. Fix a
# failure by running `python3 scripts/build_css.py` and committing.
- name: Check the committed CSS matches a fresh build
run: python scripts/build_css.py --check
type-check:
name: Type check (mypy ratchet)
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
cache: pip
# The runtime requirements are installed so mypy sees the real types of
# PIL, requests, psutil and friends -- missing, they'd be Any and the
# result would differ from a developer's machine. mypy and the stubs are
# pinned so a new release can't turn this red without a code change.
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements.txt -r web_interface/requirements.txt
pip install mypy==1.20.2 types-requests==2.33.0.20260906 types-pytz==2026.4.0.20260926
# mypy on exactly the modules in mypy-clean.txt; fails on any error in
# them, or if a listed file is missing. See CONTRIBUTING.md.
- name: Run mypy on the ratchet list
run: python scripts/check_types.py
sports-drift-report:
name: Sports drift report (report only)
runs-on: ubuntu-latest
# A progress measure for docs/SPORTS_UNIFICATION.md, never a gate: the
# monorepo's own check_sports_drift.py is the gate. The step summary shows
# how many bodies each scoreboard method family still has.
continue-on-error: true
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
persist-credentials: false
- name: Check out ledmatrix-plugins (main)
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
repository: ChuckBuilds/ledmatrix-plugins
path: ledmatrix-plugins
persist-credentials: false
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with:
python-version: "3.12"
# Stdlib only; exits 0 whatever it finds.
- name: Report method-family drift across the nine scoreboards
run: |
python scripts/sports_drift_report.py --plugins ledmatrix-plugins \
--markdown --json sports-drift.json >> "$GITHUB_STEP_SUMMARY"
python scripts/sports_drift_report.py --plugins ledmatrix-plugins
- name: Upload the full report
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: sports-drift-report
path: sports-drift.json
+48 -8
View File
@@ -3,11 +3,13 @@ __pycache__/
*.py[cod]
*$py.class
# Secrets
config/config_secrets.json
config/config.json
config/config.json.backup
config/wifi_config.json
# Secrets and per-device state. Everything the software writes into config/
# is local to one device -- config.json, config_secrets.json, wifi_config.json,
# ytm_auth.json (a login session), saved_repositories.json, font_overrides.json,
# and the temp files atomic writes leave behind when interrupted -- so only
# the templates are tracked. Listing files one by one missed several.
config/*
!config/*.template.json
credentials.json
token.pickle
@@ -35,11 +37,12 @@ htmlcov/
# Cache directory (root level only, not src/cache which is source code)
/cache/
# Development plugins directory
# Plugins are managed as separate repositories via multi-root workspace
# See docs/MULTI_ROOT_WORKSPACE_SETUP.md for details
# Development plugins directory: symlinks into a ledmatrix-plugins checkout
# See docs/PLUGIN_DEVELOPMENT_GUIDE.md and docs/MULTI_ROOT_WORKSPACE_SETUP.md
plugins/*
!plugins/.gitkeep
# Local settings for scripts/dev/dev_plugin_setup.sh (template: dev_plugins.json.example)
/dev_plugins.json
# Binary files and backups
bin/pixlet/
@@ -47,3 +50,40 @@ config/backups/
# Starlark apps runtime storage (installed .star files and cached renders)
/starlark-apps/
# JS test deps (test/js)
node_modules/
package-lock.json
# Team logos fetched at runtime.
#
# src/logo_downloader.py and LogoHelper write into assets/sports/<league>_logos/
# whenever a plugin meets a team whose logo is not on disk. Those directories are
# also tracked -- 209 NCAA logos and 153 soccer ones ship with the repo -- so
# every rig accumulated untracked files it was never meant to commit and
# `git status` was permanently dirty. That noise is not harmless: it trains
# everyone to ignore the one signal that says a checkout is not what you think
# it is, which is how a stale tree sat unnoticed on a rig until a restart
# surfaced four dead sports plugins.
#
# Ignoring a directory does not untrack what is already in it, so the logos that
# ship keep shipping. Only new downloads are hidden.
#
# Adding a logo on purpose is rare and deliberate -- the last time was #415, four
# named NCAA logos a plugin needed, and there has been no other in a year. Do it
# with an explicit override:
# git add -f assets/sports/ncaa_logos/DUKE.png
assets/sports/*_logos/
assets/stocks/ticker_icons/
assets/stocks/crypto_icons/
# Plugin operation state written at runtime.
#
# web_interface/app.py writes data/plugin_state.json and data/operation_history.json
# (older releases also data/plugin_operations.json) as the web interface runs, into
# a directory that ships tracked (data/.gitkeep). Unignored, every rig that ever
# opened the web UI -- and every test run that constructs the app -- would leave
# untracked files behind and a permanently dirty `git status`. Same reasoning as
# the logo rule above: a checkout that is always dirty is a checkout nobody reads.
data/*
!data/.gitkeep
+13 -5
View File
@@ -37,14 +37,22 @@ repos:
types: [python]
pass_filenames: false
- repo: https://github.com/pre-commit/mirrors-mypy
rev: v1.8.0
# The mypy ratchet -- the same check as CI's "Type check (mypy ratchet)"
# job: mypy on exactly the modules listed in mypy-clean.txt. Run it with
# pre-commit run mypy --hook-stage manual
# A local hook rather than mirrors-mypy so mypy sees the packages installed
# from requirements.txt, as CI does; an isolated hook env without them types
# PIL, requests and friends as Any and reports different errors. Needs
# mypy==1.20.2 (the version CI pins) in the environment you commit from.
- repo: local
hooks:
- id: mypy
additional_dependencies: [types-requests, types-pytz]
args: [--ignore-missing-imports, --no-error-summary]
name: mypy (ratchet, mypy-clean.txt)
entry: python scripts/check_types.py
language: system
pass_filenames: false
files: ^src/
always_run: true
stages: [manual]
- repo: https://github.com/PyCQA/bandit
rev: 1.8.3
+2172
View File
File diff suppressed because it is too large Load Diff
+31 -8
View File
@@ -6,32 +6,55 @@
- `config/config.json` — User plugin configuration (persists across plugin reinstalls)
- `plugin-repos/` — **Default** plugin install directory used by the
Plugin Store, set by `plugin_system.plugins_directory` in
`config.json` (default per `config/config.template.json:130`).
`config.json` (default per `config/config.template.json`).
Not gitignored.
- `plugins/` — Legacy/dev plugin location. Gitignored (`plugins/*`).
Used by `scripts/dev/dev_plugin_setup.sh` for symlinks. The plugin
loader falls back to it when something isn't found in `plugin-repos/`
(`src/plugin_system/schema_manager.py:77`).
loader does NOT fall back to it — `PluginManager.discover_plugins()`
(`src/plugin_system/plugin_manager.py`) scans only the configured
directory. Fallbacks exist in two narrower places: store operations
(`PluginStoreManager._find_plugin_path()` in `store_manager.py`, which
searches `store_search_dirs()` from `plugin_dirs.py`) and schema lookup
(`SchemaManager.get_schema_path()` in `schema_manager.py`, which probes
`plugins/` *before* `plugin-repos/`).
- `src/plugin_system/plugin_dirs.py` — the one resolver for "which directory
holds plugin X" (manifest `id` first, then `<id>` / `ledmatrix-<id>`)
## Plugin System
- Plugins inherit from `BasePlugin` in `src/plugin_system/base_plugin.py`
- Required abstract methods: `update()`, `display(force_clear=False)`
- Each plugin needs: `manifest.json`, `config_schema.json`, `manager.py`, `requirements.txt`
- Each plugin needs: `manifest.json`, `config_schema.json`, and the entry point (`manager.py` by default); `requirements.txt` if it has dependencies. Required manifest fields: `docs/PLUGIN_API_REFERENCE.md#manifest-required-fields`
- Plugin instantiation args: `plugin_id, config, display_manager, cache_manager, plugin_manager`
- Config schemas use JSON Schema Draft-7
- Display dimensions: always read dynamically from `self.display_manager.matrix.width/height`
- Display dimensions: always read dynamically from `self.display_manager.width/height` — not `display_manager.matrix.width/height`, because `matrix` is `None` when hardware init fails (the properties fall back to the canvas size)
- Secrets: namespaced by plugin id in `config/config_secrets.json`, declared
via `"x-secret": true` in the plugin's config schema, and deep-merged into
the plugin's config dict at load time — plugins read them with plain
`config.get(...)`, never a separate accessor
## Dev Workflow
- Link a plugin for development: `./scripts/dev/dev_plugin_setup.sh link-github <name>` clones the `ledmatrix-plugins` monorepo into `~/.ledmatrix-dev-plugins/` and links its `plugins/<name>` under the manifest id (add a repo URL for a plugin with its own repo; or `link <name> <path>`); symlinks land in `plugins/` — set `plugin_system.plugins_directory` to `plugins` so discovery picks them up. Fork/location overrides: `dev_plugins.json` (from `dev_plugins.json.example`)
- Browser preview without the display loop: `python3 scripts/dev_server.py` → http://localhost:5001
- Full display in emulator mode: `python3 run.py -e` (or `EMULATOR=true python3 run.py`)
- Validate one plugin headlessly: `python3 scripts/check_plugin.py --plugin <id>`
- Soak a rig for frame timing (on the Pi, service running): `python3 scripts/frame_soak.py --preview` — late-frame rate across every scroller; see `docs/SCROLL_PERFORMANCE.md`
## Plugin Store Architecture
- Official plugins live in the `ledmatrix-plugins` monorepo (not individual repos)
- Plugin repo naming convention: `ledmatrix-<plugin-id>` (e.g., `ledmatrix-football-scoreboard`)
- `plugins.json` registry at `https://raw.githubusercontent.com/ChuckBuilds/ledmatrix-plugins/main/plugins.json`
- Store manager (`src/plugin_system/store_manager.py`) handles install/update/uninstall
- Monorepo plugins are installed via ZIP extraction (no `.git` directory)
- Store manager (`PluginStoreManager` in `src/plugin_system/store_manager.py`) handles install/update/uninstall
- Monorepo plugins are installed without a `.git` directory: GitHub Trees API + raw downloads, falling back to ZIP extraction
- Update detection for monorepo plugins uses version comparison (manifest version vs registry latest_version)
- Optional registry entry fields (`store_registry.py`): `ledmatrix_min_version` refuses an incompatible install/update before the download (the post-download manifest gate stays as the fallback); `aliases` are the entry's other ids (manifest id `ledmatrix-weather` for `weather`), used with the `plugin_path` name by update/uninstall/reinstall to find the install — only this registry proof counts, never a bare `ledmatrix-<id>` folder (owner decision, #686; such a folder is only logged); `commit` is informational. An older plugins.json has none of them
- Plugin configs stored in `config/config.json`, NOT in plugin directories — safe across reinstalls
- Third-party plugins can use their own repo URL with empty `plugin_path`
## Common Pitfalls
- paho-mqtt 2.x needs `callback_api_version=mqtt.CallbackAPIVersion.VERSION1` for v1 compat
- paho-mqtt 2.x requires a `CallbackAPIVersion` argument: `VERSION1` for code written against v1 callback signatures (the MQTT bridge uses `VERSION2`)
- BasePlugin uses `get_logger()` from `src.logging_config`, not standard `logging.getLogger()`
- `DisplayManager` has no `draw_image()` — paste onto the PIL image directly:
`self.display_manager.image.paste(img, (x, y))` then `update_display()`
(use a mask for transparency: `image.paste(rgba, (x, y), rgba)`)
- When modifying a plugin in the monorepo, you MUST bump `version` in its `manifest.json` and run `python update_registry.py` — otherwise users won't receive the update
- `src/pi5_matrix_support.py` hardcodes what the pinned `rpi-rgb-led-matrix-master` can drive on a Raspberry Pi 5 (`Rp1PioConfigSupported()` in `lib/rp1/rp1_pio_backend.cc`). Re-check it whenever the submodule is bumped: a stale rule blocks Pi 5 settings the new library supports, and a missing one lets the display service crash-loop. `src/matrix_support.py` holds the same kind of rules for every board (rows, chain length, mapping names, parallel per mapping) and needs the same re-check
+1 -1
View File
@@ -63,7 +63,7 @@ ChuckBuilds, and any other forums hosted by or affiliated with the project.
Instances of abusive, harassing, or otherwise unacceptable behavior may be
reported to the community leaders responsible for enforcement on the
[LEDMatrix Discord](https://discord.gg/uW36dVAtcT) (DM a moderator or
[LEDMatrix Discord](https://discord.gg/RdrC37rEag) (DM a moderator or
ChuckBuilds directly) or by opening a private GitHub Security Advisory if
the issue involves account safety. All complaints will be reviewed and
investigated promptly and fairly.
+21 -5
View File
@@ -9,7 +9,7 @@ improvements, and code changes.
- **Bugs / feature requests**: open an issue using one of the templates
in [`.github/ISSUE_TEMPLATE/`](.github/ISSUE_TEMPLATE/).
- **Real-time discussion**: the
[LEDMatrix Discord](https://discord.gg/uW36dVAtcT).
[LEDMatrix Discord](https://discord.gg/RdrC37rEag).
- **Plugin development**:
[`docs/PLUGIN_DEVELOPMENT_GUIDE.md`](docs/PLUGIN_DEVELOPMENT_GUIDE.md)
and the [`ledmatrix-plugins`](https://github.com/ChuckBuilds/ledmatrix-plugins)
@@ -40,7 +40,7 @@ improvements, and code changes.
## Running the tests
```bash
pip install -r requirements.txt
pip install -r requirements.txt -r requirements-test.txt
pytest
```
@@ -57,9 +57,25 @@ integration tests.
`docs/<short-description>`.
3. **Keep PRs focused.** One conceptual change per PR. If you find
adjacent bugs while working, fix them in a separate PR.
4. **Follow the existing code style.** Python code uses standard
`black`/`ruff` conventions; HTML/JS in `web_interface/` follows the
patterns already in `templates/v3/` and `static/v3/`.
4. **Follow the existing code style.** The pre-commit hooks run
`flake8` (E9, F63, F7, F82 plus bugbear `B` checks), `bandit`,
and `gitleaks` — install the CLI with
`python -m pip install pre-commit`, then run
`pre-commit install` so they run on every commit. Type checking
is a ratchet while the existing mypy errors in `src/` are paid
down: `mypy-clean.txt` lists the modules that type-check clean, and
CI runs `python scripts/check_types.py` (also the manual hook
`pre-commit run mypy --hook-stage manual`) to keep every listed
module clean. When you make another module clean, add it to the
list (sorted); don't take one off to get CI green. Keep type fixes
annotation-only where you can -- widen a hint rather than delete a
defensive runtime check mypy calls unreachable. HTML/JS in
`web_interface/` follows the patterns already in `templates/v3/`
and `static/v3/`. If you change a template or a static JS file,
run `python3 scripts/build_css.py` and commit the regenerated
`static/v3/tailwind.css` with it -- CI fails when the committed CSS
is out of date. It needs no Node; see
[`web_interface/README.md`](web_interface/README.md#styling-tailwind-css).
5. **Update documentation** alongside code changes. If you add a
config key, document it in the relevant `*.md` file (or, for
plugins, in `config_schema.json` so the form is auto-generated).
+74
View File
@@ -0,0 +1,74 @@
# Product
<!-- impeccable:product-schema 1 -->
## Platform
web
## Users
Designed novice-first, with power tools kept within reach.
- **Primary: hobbyist builders.** People who assembled an LED matrix panel on a Raspberry Pi, often by following the install video, and are frequently new to Linux and the Pi. They set the display up once (panel size, timezone, WiFi), install and enable a few plugins, then come back occasionally to tweak what the panel shows. They usually reach the control panel from a phone or laptop on their home network, sometimes as an installed home-screen app.
- **Secondary: tinkerers and plugin developers.** Comfortable with SSH, `config.json`, and GitHub. They lean on the Config Editor, Logs, Cache, Operation History, Tools, GitHub-repo installs, and per-plugin config while building or debugging. Their tools must stay reachable without sitting in the novice's path.
## Product Purpose
LEDMatrix turns a Raspberry Pi and an RGB LED matrix panel into an information-rich display (clock, weather, calendar, sports scores, stocks, music, and more) through a plugin platform. The web control panel ("LED Matrix Control") is where the display gets configured, extended, and kept healthy.
Success means a builder gets from a freshly flashed Pi to a working, personalized display without needing a terminal, and can keep it running (updates, recovery, troubleshooting) the same way.
## Positioning
Four strengths define LEDMatrix, and future work must protect all of them:
1. **Plugin ecosystem.** The core ships only `starlark-apps` and `web-ui-info`; everything else comes from the built-in Plugin Store (the official `ledmatrix-plugins` monorepo), third-party GitHub repos, or Starlark (Tidbyt-style) apps. Each installed plugin gets its own configuration tab, generated from its schema.
2. **Runs on tiny Pis.** The UI is served by the same device that drives the matrix, on boards as small as the Pi Zero 2 W (512 MB), Pi 3/3B+, and the 1 GB Pi 4.
3. **Recovers without SSH.** WiFi access-point fallback with a captive setup page, backup & restore, in-UI updates, live logs, diagnostics, service control, and plugin health let users fix problems from the browser.
4. **Open and community-led.** GPL-3.0, a Discord community, and contributions welcome. The maintainer (ChuckBuilds) builds in public and openly relies on AI development tools.
## Operating Context
- **Access.** Served on the local network at `http://<pi-ip>:5000` by the `ledmatrix-web` service. It is installable as a PWA (`web_interface/static/v3/manifest.json`, short name "LEDMatrix").
- **First run.** When the Pi has no network it creates its own WiFi access point, so the captive setup page (`templates/v3/captive_setup.html`) may be the very first screen a user sees, on a phone, with no internet connection.
- **Navigation.**
- System tabs: Overview, General, WiFi, Schedule, Display, Rotation, Config Editor, Backup & Restore, Fonts, Logs, Cache, Operation History, Tools.
- A second row holds Plugin Manager (with the Plugin Store), Starlark Apps, and one tab per installed plugin.
- **Live data.** The Overview shows system stats (CPU, memory, temperature, power/throttling) and a live display preview, streamed over SSE.
- **Getting Started checklist.** The Overview's first-run checklist runs: set panel size → set timezone → install a plugin → enable it → configure it.
- **Development.** `python3 scripts/dev_server.py` gives a browser preview without the display loop; `python3 run.py -e` runs the full display in emulator mode.
## Capabilities and Constraints
- **Hard constraint: plugin UI compatibility.** Third-party plugins rely on JSON Schema (Draft-7) generated config forms, the widget registry (`static/v3/js/widgets/`), `x-secret` fields, and plugin web-UI actions. UI changes must keep these working.
- **Config storage.** Plugin configuration lives in `config/config.json` and secrets in `config/config_secrets.json`, never in plugin directories, so configs survive reinstalls.
- **Stack.** An existing Flask + HTMX + Alpine.js app with Jinja templates (`web_interface/templates/v3/`) and static JS/CSS (`web_interface/static/v3/`), with self-hosted vendor assets.
- **Terminology.** Plugin, Plugin Store, Starlark app, rotation, display duration, Vegas Scroll Mode, on-demand, AP mode.
- **Open decisions** (offered during init, not adopted as constraints):
- Whether the UI must work fully offline, with no CDN fallbacks at runtime.
- Whether a Node/CSS build step is acceptable for contributors.
- Whether a formal accessibility standard (e.g. WCAG 2.2 AA) is a requirement.
## Brand Commitments
- **Names.** The product is "LEDMatrix" and the web UI is titled "LED Matrix Control". The maintainer brand is ChuckBuilds.
- **Voice.** Friendly, honest, and learning-in-public, as in the README.
- **App icons.** They live in `web_interface/static/v3/icons/`.
No other visual identity has been made binding.
## Evidence on Hand
- **Photos.** Real photographs of running displays are linked in `README.md` (clock, weather, calendar, NHL/MLB/NFL/NCAA, stocks, music).
- **Video.** YouTube install and walkthrough videos from ChuckBuilds.
- **Docs.** Extensive documentation in `docs/`, e.g. `WEB_INTERFACE_GUIDE.md`, `GETTING_STARTED.md`, `WIFI_NETWORK_SETUP.md`, `LOW_MEMORY_BOARDS.md`, `PLUGIN_STORE_GUIDE.md`.
- **Absences.** There are no testimonials, user counts, or benchmark figures. Do not fabricate them.
## Product Principles
1. **Novice path first, power one click away.** Default views serve the first-time builder, while advanced tools stay discoverable for tinkerers.
2. **Never strand the user at a terminal.** Every setup, recovery, and troubleshooting task has a browser path, including from the AP-mode captive page.
3. **Respect the Pi.** Every feature is paid for in memory and CPU on a Pi Zero 2 W that is also driving the display.
4. **The ecosystem is the product.** Plugins, including third-party ones, must feel first-class and keep working across core UI changes.
5. **Honest and welcoming.** Plain language, truthful status, and no overstated claims, in keeping with an open, community-built project.
+171 -86
View File
@@ -33,7 +33,7 @@ I'm trying to be open to constructive criticism and support, as long as it's a r
- Show support on Youtube: https://www.youtube.com/@ChuckBuilds
- Check out the write-up on my website: https://www.chuck-builds.com/led-matrix/
- Stay in touch on Instagram: https://www.instagram.com/ChuckBuilds/
- Want to chat? Reach out on the LEDMatrix Discord: [https://discord.com/invite/uW36dVAtcT](https://discord.gg/dfFwsasa6W)
- Want to chat? Reach out on the LEDMatrix Discord: [https://discord.gg/RdrC37rEag](https://discord.gg/RdrC37rEag)
- Feeling Generous? Consider sponsoring this project or sending a donation (these AI credits aren't cheap!)
-----------------------------------------------------------------------------------
@@ -50,7 +50,15 @@ I'm trying to be open to constructive criticism and support, as long as it's a r
<details>
<summary>Core Features</summary>
The following plugins are available inside of the LEDMatrix project. These modular, rotating Displays that can be individually enabled or disabled per the user's needs with some configuration around display durations, teams, stocks, weather, timezones, and more. Displays include:
LEDMatrix is a plugin platform: the displays below are plugins installed
from the built-in Plugin Store (web interface → Plugins), where each can be
individually enabled, ordered, and configured — display durations, teams,
stocks, weather, timezones, and more. The core repo ships with just two
bundled plugins (`starlark-apps` and `web-ui-info`); the official plugins
live in the [ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins)
monorepo and install with one click, and third-party plugins can be
installed from their own GitHub repositories. Displays available in the
store include:
### Time and Weather
- Real-time clock display (2x 64x32 Displays 4mm Pixel Pitch)
@@ -132,29 +140,31 @@ The system supports live, recent, and upcoming game information for multiple spo
| This project can be finnicky! RGB LED Matrix displays are not built the same or to a high-quality standard. We have seen many displays arrive dead or partially working in our discord. Please purchase from a reputable vendor. |
### Raspberry Pi
- Raspberry Pi Zero's don't have enough processing power for this project.
- **Raspberry Pi 3B, 4, or 5**
- **Raspberry Pi 3B, 4, or 5** (a Pi Zero 2 W also works, with the limits described under the 1GB/low-memory bullet below; the original Pi Zero / Zero W doesn't have enough processing power for this project)
[Amazon Affiliate Link – Raspberry Pi 4 4GB RAM](https://amzn.to/4dJixuX)
[Amazon Affiliate Link – Raspberry Pi 4 8GB RAM](https://amzn.to/4qbqY7F)
- **Pi 5 users**: the installer automatically detects Pi 5 and builds the `rpi-rgb-led-matrix` library with RP1 support. If you previously installed on a Pi 4 and migrated the SD card, or if you see `mmap` errors in the logs, force a fresh library build:
```bash
sudo RPI_RGB_FORCE_REBUILD=1 ./first_time_install.sh
```
- Pi 5 config: leave `rp1_rio` at `0` (PIO mode, default) and set `gpio_slowdown` to `1` or `2`.
- Pi 5 config: leave `rp1_rio` at `0` (PIO mode, default) and start `gpio_slowdown` at `1`, raising it a step at a time if the image flickers or shows garbage (see `gpio_slowdown` under Display Settings).
- **1GB models (Pi 3B / 3B+), the 512MB Pi Zero 2 W and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`. Once running, keep an eye on memory: see [docs/LOW_MEMORY_BOARDS.md](docs/LOW_MEMORY_BOARDS.md).
### RGB Matrix Bonnet / HAT
- [Adafruit RGB Matrix Bonnet/HAT](https://www.adafruit.com/product/3211) – supports one “chain” of horizontally connected displays
- [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular-pi1` as hardware mapping)*
- [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular` as hardware mapping)*
- [Electrodragon RGB HAT](https://www.electrodragon.com/product/rgb-matrix-panel-drive-board-raspberry-pi/) – supports up to 3 vertical “chains”
- [Seengreat Matrix Adapter Board](https://amzn.to/3KsnT3j) – single-chain LED Matrix *(use `regular` as hardware mapping)*
### LED Matrix Panels
(2x in a horizontal chain is recommended)
- [Adafruit 64×32](https://www.adafruit.com/product/2278) – designed for 128×32 but works with dynamic scaling on many displays (pixel pitch is user preference)
**Warning: Lately the Waveshare Panels have had different variations - only some are compatible with this project. I hope to identify what is different to fix it but so far there is a decent chance you get a mis-matched set of panels if you don't buy them all at once! **
- [Waveshare 64×32](https://amzn.to/3Kw55jK) - Does not require E addressable pad
- [Waveshare 96×48](https://amzn.to/4bydNcv) – higher resolution, requires soldering the **E addressable pad** on the [Adafruit RGB Bonnet](https://www.adafruit.com/product/3211) to “8” **OR** toggling the DIP switch on the Adafruit Triple LED Matrix Bonnet *(no soldering required!)*
> Amazon Affiliate Link – ChuckBuilds receives a small commission on purchases
- [Waveshare 96×48](https://amzn.to/4bydNcv) – higher resolution, requires soldering the **E addressable pad** on the [Adafruit RGB Bonnet](https://www.adafruit.com/product/3211) to “8” **OR** toggling the DIP switch on the Adafruit Triple LED Matrix Bonnet *(no soldering required!)*
- There are some Panels on Aliexpress that have worked fine for me, shop around! I think Adafruit is probably the "safest" but they do have some limitation on resolution and layout.
> Amazon Affiliate Links – ChuckBuilds receives a small commission on purchases
### Power Supply
- [5V 4A DC Power Supply](https://www.adafruit.com/product/658) (good for 2 -3 displays, depending on brightness and pixel density, you'll need higher amperage for more)
@@ -162,7 +172,7 @@ The system supports live, recent, and upcoming game information for multiple spo
## Optional but recommended mod for Adafruit RGB Matrix Bonnet
- By soldering a jumper between pins 4 and 18, you can run a specialized command for polling the matrix display. This provides better brightness, less flicker, and better color.
- If you do the mod, we will use the default config with led-gpio-mapping=adafruit-hat-pwm, otherwise just adjust your mapping in config.json to adafruit-hat
- The default config uses `hardware_mapping` `adafruit-hat`. If you do the mod, change it to `adafruit-hat-pwm` (Display settings in the web interface, or `config.json`)
- More information available: https://github.com/hzeller/rpi-rgb-led-matrix/tree/master?tab=readme-ov-file
![DSC00079](https://github.com/user-attachments/assets/4282d07d-dfa2-4546-8422-ff1f3a9c0703)
@@ -314,12 +324,13 @@ curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/
```
This one-shot installer will automatically:
- Check system prerequisites (network, disk space, sudo access)
- Check system prerequisites (network, disk space, memory, sudo access)
- Install required system packages (git, python3, build tools, etc.)
- Clone or update the LEDMatrix repository
- Run the complete first-time installation script
- Print the web interface address, then **reboot the Pi automatically** (your SSH session will disconnect; give it a few minutes to come back)
The installation process typically takes 10-30 minutes depending on your internet connection and Pi model. All errors are reported explicitly with actionable fixes.
The installation process typically takes 10-30 minutes depending on your internet connection and Pi model. Pi 3B/3B+ and other 1GB boards land at the top of that range, because the C++ library is compiled serially to stay within available memory. All errors are reported explicitly with actionable fixes.
**Note:** The script is safe to run multiple times and will handle existing installations gracefully.
@@ -336,10 +347,10 @@ If you prefer to install manually or the one-shot installer doesn't work for you
ssh ledpi@ledpi
```
2. Update repositories, upgrade Raspberry Pi OS, and install prerequisites:
2. Update repositories, upgrade Raspberry Pi OS, and install git (`first_time_install.sh` installs the build dependencies itself: `python3-pip`, `python-dev-is-python3`, `build-essential`, `cmake`, `ninja-build` and the rest):
```bash
sudo apt update && sudo apt upgrade -y
sudo apt install -y git python3-pip cython3 build-essential python3-dev python3-pillow scons
sudo apt install -y git
```
3. Clone this repository:
@@ -356,6 +367,12 @@ sudo bash ./first_time_install.sh
This single script installs services, dependencies, configures permissions and sudoers, and validates the setup.
It finishes by asking whether to reboot. If you run it non-interactively — piped, over a script, or with `-y` — there is no one to ask, so **it reboots immediately without prompting**. Pass `--no-reboot-prompt` to install without rebooting:
```bash
sudo bash ./first_time_install.sh -y --no-reboot-prompt
```
</details>
</details>
@@ -371,6 +388,10 @@ This single script installs services, dependencies, configures permissions and s
### Initial Setup
For a complete list of every key in `config.json` and
`config_secrets.json`, see
[docs/CONFIG_REFERENCE.md](docs/CONFIG_REFERENCE.md).
For most settings I recommend using the web interface:
Edit the project via the web interface at http://[IP ADDRESS or HOSTNAME]:5000 or http://ledpi:5000 .
@@ -379,7 +400,7 @@ If you need to manually edit your config file, you can follow the steps below:
<summary>Manual Config.json editing </summary>
1. **First-time setup**:
The previous "First_time_install.sh" script should've already copied the template to create your config.json:
The previous `first_time_install.sh` script should've already copied the template to create your config.json:
2. **Edit your configuration**:
```bash
@@ -416,7 +437,7 @@ I recommend using the web-ui "Quick Actions" to control the Display.
## Plugins
<details>
LEDMatrix uses a plugin-based architecture where all display functionality (except the core calendar) is implemented as plugins. All managers that were previously built into the core system are now available as plugins through the Plugin Store.
LEDMatrix uses a plugin-based architecture where all display functionality is implemented as plugins. All managers that were previously built into the core system are now available as plugins through the Plugin Store.
### Plugin Store
See the [Plugin Store documentation](https://github.com/ChuckBuilds/ledmatrix-plugins) for detailed installation instructions.
@@ -438,9 +459,9 @@ You can also install plugins directly from GitHub repositories:
See the [Plugin Store documentation](https://github.com/ChuckBuilds/ledmatrix-plugins) for detailed installation instructions.
For plugin development, check out the [Hello World Plugin](https://github.com/ChuckBuilds/ledmatrix-hello-world) repository as a starter template.
For plugin development, the `plugins/hello-world/` plugin in the [ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins) repository is a starter template.
2. **Built-in Managers Deprecated**: The built-in managers (hockey, football, stocks, etc.) are now deprecated and have been moved to the plugin system. **You must install replacement plugins from the Plugin Store** in the web interface instead. The plugin system provides the same functionality with better maintainability and extensibility.
**Built-in Managers Deprecated**: The built-in managers (hockey, football, stocks, etc.) are now deprecated and have been moved to the plugin system. **You must install replacement plugins from the Plugin Store** in the web interface instead. The plugin system provides the same functionality with better maintainability and extensibility.
</details>
## Detailed Information
@@ -455,6 +476,10 @@ If you are copying my exact setup, you can likely leave the defaults alone. Howe
The display settings are located in `config/config.json` under the `"display"` key and are organized into three main sections: `hardware`, `runtime`, and `display_durations`.
The defaults below are the values in `config/config.template.json`. They are what applies when you haven't set a key: on every load, LEDMatrix adds any key your `config.json` lacks from the template, so `DisplayManager`'s own fallbacks are never reached on a normal install.
The web UI and the config API refuse values the rgbmatrix library can't start with. If one is written into `config.json` by hand anyway, the display logs which setting it is (`Failed to initialize RGB Matrix` in `sudo journalctl -u ledmatrix`), runs in fallback mode, and the Display tab shows the message.
### Hardware Configuration (`display.hardware`)
These settings control the physical hardware configuration and how the matrix is driven.
@@ -464,15 +489,18 @@ These settings control the physical hardware configuration and how the matrix is
- **`rows`** (integer, default: 32)
- Number of LED rows (vertical pixels) in each panel
- Common values: 16, 32, 48, 64
- An even number from 8 to 64, the most the rgbmatrix library drives per panel
- Must match your physical panel configuration
- **`cols`** (integer, default: 64)
- Number of LED columns (horizontal pixels) in each panel
- Common values: 32, 64, 96, 128
- At least 16, with no upper limit
- Must match your physical panel configuration
- **`chain_length`** (integer, default: 2)
- Number of LED panels chained together horizontally
- 1 to 255 (the library's Python binding stores it in one byte); longer chains lower the refresh rate
- If you have 2 panels side-by-side, set to 2
- If you have 4 panels in a row, set to 4
- Total display width = `cols × chain_length`
@@ -481,68 +509,70 @@ These settings control the physical hardware configuration and how the matrix is
- Number of parallel chains (panels stacked vertically)
- Use 1 for a single row of panels
- Use 2 if you have panels stacked in two rows
- 1–3, and no more than your `hardware_mapping` has outputs: `regular` and `classic` have 3 (e.g. the Adafruit Triple LED Matrix Bonnet); `adafruit-hat`, `adafruit-hat-pwm`, `regular-pi1` and `classic-pi1` have 1. The library stops the display service outright on a mismatch, so it is refused
- Total display height = `rows × parallel`
#### Brightness and Visual Settings
- **`brightness`** (integer, 0-100, default: 90)
- **`brightness`** (integer, 1-100, default: 90)
- Display brightness level
- Lower values (0-50) are dimmer, higher values (50-100) are brighter
- Lower values (1-50) are dimmer, higher values (50-100) are brighter
- Recommended: 70-90 for indoor use, 90-100 for bright environments
- Very high brightness may cause distortion or require more power
#### Hardware Mapping
- **`hardware_mapping`** (string, default: "adafruit-hat-pwm")
- **`hardware_mapping`** (string, default: "adafruit-hat")
- Specifies which GPIO pin mapping to use for your hardware
- **`"adafruit-hat-pwm"`**: Use this for Adafruit RGB Matrix Bonnet/HAT WITH the jumper mod (PWM enabled). This is the recommended setting for Adafruit hardware with the PWM jumper soldered.
- **`"adafruit-hat"`**: Use this for Adafruit RGB Matrix Bonnet/HAT WITHOUT the jumper mod (no PWM). Remove `-pwm` from the value if you did not solder the jumper.
- **`"regular"`**: Standard GPIO pin mapping for direct GPIO connections (Generic)
- **`"regular"`**: Standard GPIO pin mapping for direct GPIO connections (Generic). Also the right choice for the Adafruit Triple LED Matrix Bonnet
- **`"regular-pi1"`**: Standard GPIO pin mapping for Raspberry Pi 1 (older hardware or non-standard hat mapping)
- **`"classic"`** / **`"classic-pi1"`**: the library's original pin-outs, for old adapter boards wired to them. Not used by current HATs
- Any other name is refused. `compute-module` is only compiled in when the library is built with `ENABLE_WIDE_GPIO_COMPUTE_MODULE`, which the installer doesn't do. On a Raspberry Pi 5, `classic-pi1` isn't supported
- Choose the option that matches your specific hardware setup, if aren't sure try them all.
- Hardware pulsing (see `disable_hardware_pulsing`) needs the panel's OE line on GPIO 18, which `adafruit-hat-pwm` and `regular` provide and `adafruit-hat` does not
#### PWM (Pulse Width Modulation) Settings
These settings affect color fidelity and smoothness of color transitions:
- **`pwm_bits`** (integer, default: 9)
- Number of bits used for PWM (affects color depth)
- Higher values (9-11) = more color levels, smoother gradients
- Lower values (7-8) = fewer color levels, but may improve stability on some hardware
- Range: 1-11, recommended: 9-10
- **`pwm_bits`** (integer, 1-11, default: 9)
- Color depth per channel: how many brightness levels each LED gets
- Higher values (9-11) = more color levels, smoother gradients, lower refresh rate
- Lower values (7-8) = the subtlest shades are dropped for a higher refresh rate; `1` gives 8 colors
- Recommended: 9-10
- **`pwm_dither_bits`** (integer, default: 1)
- Additional dithering bits for smoother color transitions
- Helps reduce color banding in gradients
- Higher values (1-2) = smoother gradients but may impact performance
- Range: 0-2, recommended: 1
- **`pwm_dither_bits`** (integer, 0-2, default: 1)
- Time-dithers the lowest color bits: their brightness comes from showing them on only some frames
- Raises the refresh rate; the cost is that dark shades can shimmer slightly
- `0` = steadiest dim colors, `2` = fastest
- The rgbmatrix library accepts only 0-2; a higher value stops the display starting
- **`pwm_lsb_nanoseconds`** (integer, default: 130)
- Least significant bit timing in nanoseconds
- Controls the base timing for PWM signals
- Lower values = faster PWM, higher values = slower PWM
- **`pwm_lsb_nanoseconds`** (integer, 50-3000, default: 130)
- On-time of the least significant color bit; each higher bit doubles it
- Lower values = higher refresh rate, but can cost color accuracy or add ghosting on some panels
- Higher values = less ghosting (faint trails behind bright text on black), lower refresh rate
- Typical range: 100-300 nanoseconds
- May need adjustment if you see flickering or color issues
#### Advanced Hardware Settings
- **`scan_mode`** (integer, default: 0)
- Panel scan mode (how rows are addressed)
- Common values: 0 (progressive), 1 (interlaced)
- Most panels use 0, but some require 1
- Check your panel datasheet if colors appear incorrect
- **`scan_mode`** (integer, 0-1, default: 0)
- Order the rows are refreshed in: `0` = progressive, `1` = interlaced
- Interlaced can look a little smoother when the refresh rate is very low, but usually shows a comb effect on anything moving
- Leave at `0` unless you are tuning a slow setup
- **`limit_refresh_rate_hz`** (integer, default: 100)
- Maximum refresh rate in Hz (frames per second)
- Caps the refresh rate for better stability
- Lower values (60-80) = more stable, less CPU usage
- Higher values (100-120) = smoother animations, more CPU usage
- Recommended: 80-100 for most setups
- Caps the panel refresh rate in Hz; `0` = no cap
- A steady cap reduces flicker caused by other activity on the Pi, and in camera recordings
- Scroll speeds are worked out against this value (against 100 Hz when it is `0`), so a cap the panel can actually hold keeps scrolling even
- Recommended: 80-120. `sudo python3 scripts/scroll_speeds.py --measure` reports the rate your panel really achieves
- **`disable_hardware_pulsing`** (boolean, default: false)
- Disables hardware pulsing (usually leave as false)
- Set to `true` only if you experience timing issues
- Most users should leave this as `false`
- `false` = the Pi's hardware PWM times each brightness pulse; `true` = software timing
- Leave `false` where possible. Software timing is less exact, so a row, or the whole panel, can briefly flash brighter
- Hardware pulsing needs the panel's OE line on GPIO 18 (`adafruit-hat-pwm`, `regular`, the Adafruit Triple LED Matrix Bonnet). With `adafruit-hat` the library uses software timing anyway
- It also needs the Pi's onboard sound driver (`snd_bcm2835`) disabled, which `first_time_install.sh` does. Set `true` only if you need the Pi's own audio
- **`inverse_colors`** (boolean, default: false)
- Inverts all colors (red becomes cyan, etc.)
@@ -550,9 +580,9 @@ These settings affect color fidelity and smoothness of color transitions:
- Set to `true` only if colors appear inverted
- **`show_refresh_rate`** (boolean, default: false)
- Displays the current refresh rate on the matrix (for debugging)
- Set to `true` to see FPS on the display
- Useful for troubleshooting performance issues
- Prints the live refresh rate to the console; nothing is drawn on the panel
- Readable when you stop the service and run `sudo python3 run.py` in a terminal; under the service the output is buffered
- `sudo python3 scripts/scroll_speeds.py --measure` is an easier way to see the real refresh rate
#### Advanced Panel Configuration (Advanced Users Only)
@@ -562,6 +592,7 @@ These settings are typically only needed for non-standard panels or custom confi
- Color channel order for your LED panel
- Common values: "RGB", "RBG", "GRB", "GBR", "BRG", "BGR"
- Most panels use "RGB", but some use "GRB" or other orders
- If red shows as blue, try "BGR" (the Waveshare 96x48 V2 needs it)
- Check your panel datasheet if colors appear wrong
- **`pixel_mapper_config`** (string, default: "")
@@ -571,40 +602,76 @@ These settings are typically only needed for non-standard panels or custom confi
- Leave empty unless you need custom mapping
- See rpi-rgb-led-matrix documentation for full options
- **`orientation`** (string, default: "normal")
- Rotates the rendered image to match how the panel is physically mounted
- Set to `"180"` (or use the "Upside Down" option in the web UI's Display
settings) if the panel is mounted upside down — useful for optimizing
where the Raspberry Pi and wiring sit relative to the mounting location
- `"90"` and `"270"` are for a panel mounted on its side; they swap the
display's width and height
- Applied independently of `pixel_mapper_config` (appended as a trailing
`Rotate:<degrees>` mapper), so custom mapper configs keep working alongside it
- **`row_address_type`** (integer, default: 0)
- How rows are addressed on the panel
- Most panels use 0 (direct addressing)
- Some panels require 1 (AB addressing) or 2 (ABC addressing)
- 1 = AB-addressed, 2 = direct row select, 3 = ABC-addressed,
4 = ABC shift + DE direct (SM5266), 5 = SM5368 / B707 row shift register
- ABC panels (no E line, e.g. many 128x64 FM6124 panels) use 3
- Panels with SM5368 row drivers use 5 with `led_rgb_sequence` `"BGR"` —
e.g. the Waveshare 96x48 V2 (back silkscreen `24S-A1`; the V1, `24S-A2.1`,
uses the defaults). This is what Waveshare's `96X48_1_24_SM5368` panel
type sets in their library fork.
- SM5368 row drivers are timing-sensitive: if rows jump up and down or the
bottom row shows a copy of other rows, raise `gpio_slowdown`. On a Pi 4
with an Adafruit Triple LED Matrix Bonnet, 4 left rows jumping; 6–8 gave a
stable image.
- On a Raspberry Pi 5 the rgbmatrix library currently supports only 0 and 2
(and `parallel` 1-3). Anything else would crash the display service, so on
a Pi 5 the web UI offers only 0 and 2, the config API refuses the others,
and if one is set in `config.json` anyway the display logs why and runs in
fallback mode
- Check your panel datasheet if display appears corrupted
- **`multiplexing`** (integer, default: 0)
- Panel multiplexing type
- 0 = no multiplexing (standard panels)
- Higher values for panels with different multiplexing schemes
- Check your panel datasheet for the correct value
- **`multiplexing`** (integer, 0-22, default: 0)
- How pixels are wired on outdoor/specialty panels (P10, P8, P4 and P3 outdoor modules and similar) whose LEDs aren't laid out in straight rows
- `0` = direct (standard indoor panels)
- `1` Stripe, `2` Checkered, `3` Spiral, `4` ZStripe, `5` ZnMirrorZStripe,
`6` Coreman, `7` Kaler2Scan, `8` ZStripeUneven, `9` P10-128x4-Z,
`10` QiangLiQ8, `11` InversedZStripe, `12`–`14` P10Outdoor1R1G1B v1–v3,
`15` P10CoremanMapper, `16` P8Outdoor1R1G1B, `17` FlippedStripe,
`18` P10-32x16-HalfScan, `19` P10-32x16-QuarterScan, `20` P3Outdoor-64x64,
`21` DoubleZMultiplex, `22` P4Outdoor-80x40
- If the image is scrambled in a repeating pattern, try the value named after your panel first
- **`panel_type`** (string, default: `""`)
- Sends a start-up initialization sequence to driver chips that need one
- `""` = Standard (no initialization) — right for most panels, including FM6124 / FM6124D / FM6124DJ
- `"FM6126A"` or `"FM6127"` for panels with those chips; try `"FM6126A"` if the panel stays dark or lights only the first pixel on Standard
### Runtime Configuration (`display.runtime`)
These settings control runtime behavior and GPIO timing:
- **`gpio_slowdown`** (integer, default: 3)
- GPIO timing slowdown factor
- **Critical setting**: Must match your Raspberry Pi model for stability
- **Raspberry Pi 3**: Use 3
- **Raspberry Pi 4**: Use 4
- **Raspberry Pi 5**: Use 1–2 in PIO mode (`rp1_rio: 0`, the default); start with `1` and increase if you see flickering
- **Raspberry Pi Zero/1**: Use 1-2
- Incorrect values can cause display corruption, flickering, or system instability
- GPIO timing slowdown factor (0-10): slows GPIO writes so the panel electronics keep up. Higher is more reliable but lowers the refresh rate
- **Critical setting**: depends on your Raspberry Pi model and your panel
- **Raspberry Pi Zero/1**: 0-1
- **Raspberry Pi 2/3**: 1-3
- **Raspberry Pi 4**: 2-4 (the config template ships 3)
- **Raspberry Pi 5**: 1–3 in PIO mode (`rp1_rio: 0`, the default). Start at `1` (the library treats `0` as `1` there) and raise it a step at a time if the image flickers or shows garbage — chained panels are the likeliest to need it
- Panels on `row_address_type` 5 (SM5368 row drivers) can need 6-8 on a Pi 4
- Too low: garbage, flicker or rows jumping. Too high: a lower refresh rate
- If you experience issues, try adjusting this value up or down by 1
- **`rp1_rio`** (integer, 0 or 1, default: 0) — Raspberry Pi 5 only
- Which driver the Pi 5's RP1 chip uses: `0` = PIO (default, less CPU), `1` = RIO (registered I/O, can reach a higher refresh rate)
- In RIO mode the effect of `gpio_slowdown` is inverted: higher values may be faster
- Ignored on a Pi 0-4, and applied only if the installed rgbmatrix library supports it
### Display Durations (`display.display_durations`)
Controls how long each display module stays visible in seconds before switching to the next one.
- **`calendar`** (integer, default: 30)
- Duration in seconds for the calendar display
- Increase for more time to read dates/events
- Decrease to cycle through other displays faster
Controls how long each installed plugin stays visible in seconds before switching to the next one, keyed by plugin id.
- **Plugin-specific durations**
- Each plugin can have its own duration setting
@@ -623,9 +690,10 @@ Controls how long each display module stays visible in seconds before switching
### Display Format Settings
- **`use_short_date_format`** (boolean, default: true)
- Use short date format (e.g., "Jan 15") instead of long format (e.g., "January 15th")
- Set to `false` for longer, more readable dates
- Set to `true` to save space and show more information
- Currently has no effect. The web UI still saves it, but no core code
reads it. Scoreboard plugins that offer a short date format read the
setting from their own plugin config instead. See
[CONFIG_REFERENCE.md](docs/CONFIG_REFERENCE.md#display--other-keys).
### Dynamic Duration Settings (`display.dynamic_duration`)
@@ -634,7 +702,7 @@ Controls how long each display module stays visible in seconds before switching
- Some plugins can automatically adjust their display time based on content
- This setting limits how long they can extend (prevents one display from dominating)
- Example: If set to 60, a plugin can extend up to 60 seconds even if it requests longer
- Leave unset to use the default cap (typically 90 seconds)
- Leave unset to use the default cap (180 seconds; the web UI accepts 30-1800)
### Example Configuration
@@ -681,6 +749,14 @@ Controls how long each display module stays visible in seconds before switching
- Verify `hardware_mapping` matches your HAT/connection type
- Try adjusting `gpio_slowdown`
- Ensure your display doesn't need the E-Addressable line
- If it went blank right after a settings change, the Display tab shows a "simulation mode" banner, and `sudo journalctl -u ledmatrix` shows `Failed to initialize RGB Matrix` followed by the reason. When LEDMatrix refused the settings (for example more than 64 `rows`, `parallel` 2 on an `adafruit-hat` mapping, a misspelled `hardware_mapping`, or on a Raspberry Pi 5 a `row_address_type` other than 0 or 2), the message names each one: change them, save, and restart the display service. Otherwise the library itself failed, and its own message just before names the problem
- A repeating scramble points at `row_address_type` or `multiplexing`; a panel that stays dark, at `panel_type`
**Rows jump up and down, or the bottom row repeats other rows:**
- Raise `gpio_slowdown` a step at a time (SM5368 panels on `row_address_type` 5 can need 6-8 on a Pi 4)
**A row or the whole panel briefly flashes brighter:**
- Set `disable_hardware_pulsing` to `false` (needs the OE line on GPIO 18; see `hardware_mapping`)
**Colors are wrong or inverted:**
- Check `led_rgb_sequence` (try "GRB" if "RGB" doesn't work)
@@ -705,15 +781,21 @@ Controls how long each display module stays visible in seconds before switching
<details>
<summary>Manual SSH Commands (for reference)</summary>
The quick actions essentially just execute the following commands on the Pi.
The web interface's quick actions (Start/Stop/Restart Display) call
`sudo systemctl start|stop|restart ledmatrix.service` — see
`execute_system_action()` in
[`web_interface/blueprints/api_v3/system.py`](web_interface/blueprints/api_v3/system.py).
The service runs [`run.py`](run.py) as root.
From the project root directory (ex: /home/ledpi/LEDMatrix):
To run the display in the foreground instead (for debugging), stop the service
first, then from the project root (e.g. `/home/ledpi/LEDMatrix`):
```bash
sudo python3 display_controller.py
sudo systemctl stop ledmatrix.service
sudo python3 run.py # add -d for debug logging
```
This will start the display cycle but only stays active as long as your ssh session is active.
This only runs as long as your SSH session stays open.
### Convenience Scripts
@@ -755,9 +837,11 @@ sudo ./scripts/install/install_service.sh
The script will:
- Detect your user account and home directory
- Install the service file with the correct paths
- Enable the service to start on boot
- Start the service immediately
- Install `ledmatrix.service` (display, runs as root), `ledmatrix-web.service`
(web interface, runs as your user) and the `ledmatrix-update-verify` units,
with the correct paths
- Enable them to start on boot
- Start them immediately
### Managing the Service
@@ -864,7 +948,7 @@ sudo systemctl enable ledmatrix-web.service
- **On-Demand Controls**: Start specific displays (weather, stocks, sports) on demand
- **Service Management**: Start/stop the main display service
- **System Controls**: Restart, update code, and manage the system
- **API Metrics**: Monitor API usage and system performance
- **System Stats**: CPU, memory and temperature on the Overview tab
- **Logs**: View system logs in real-time
### Troubleshooting Web Interface
@@ -881,9 +965,10 @@ sudo systemctl enable ledmatrix-web.service
3. Check if another service is using port 5000
**Service Fails to Start:**
1. Check Python dependencies are installed
2. Verify the virtual environment is set up correctly
3. Check file permissions and ownership
1. Check Python dependencies are installed. The installer puts them in the
system Python with `pip install --break-system-packages` (there is no
virtual environment), so `python3 -c "import flask"` should succeed.
2. Check file permissions and ownership
</details>
+26 -3
View File
@@ -16,7 +16,7 @@ Use one of these channels, in order of preference:
maintainer.
- Direct link: <https://github.com/ChuckBuilds/LEDMatrix/security/advisories/new>
2. **Discord DM**. Send a direct message to a moderator on the
[LEDMatrix Discord](https://discord.gg/uW36dVAtcT). Don't post in
[LEDMatrix Discord](https://discord.gg/RdrC37rEag). Don't post in
public channels.
Please include:
@@ -61,8 +61,31 @@ Out of scope (please report upstream):
LEDMatrix is designed for trusted local networks. Several limitations
are intentional rather than vulnerabilities:
- **No web UI authentication.** The web interface assumes the network
it's running on is trusted. Don't expose port 5000 to the internet.
- **Web UI authentication is optional and off by default.** Out of the
box the web interface assumes the network it's running on is trusted.
Setting a password under **General > Security** makes every page and
API route require a login or an API token (`Authorization: Bearer`),
with wrong passwords rate-limited per address
(`web_interface/auth.py`). Deliberately left open even then: requests
from the Pi itself (loopback without proxy headers; a reverse proxy on
the Pi must add `X-Forwarded-For`, or every request it relays counts as
local), the Wi-Fi setup flow while the Pi is in access-point mode,
static files, and a status-only `/api/v3/health`. The password is a
werkzeug hash and tokens are stored as SHA-256, in
`config/config_secrets.json`, which no API returns. There is no TLS:
over plain HTTP the password and tokens cross the LAN in the clear, so
still don't expose port 5000 to the internet; put a TLS reverse proxy
or a VPN in front for remote access. Anyone with shell access to the Pi
can turn login off (`scripts/reset_web_password.py`), which is the
documented recovery path.
"Trusted network" does not mean "trusted websites", though: any page
a LAN user opens could make their browser POST to the Pi. So the
interface refuses a `POST`/`PUT`/`PATCH`/`DELETE` whose `Origin` (or
`Referer`) header names another site (`web_interface/origin_guard.py`),
and `/api/v3/system/action` only accepts JSON or HTMX requests. Tools
that send neither header (curl, Home Assistant, the MQTT bridge) are
unaffected. Not covered: DNS rebinding, and anyone who can reach the
port directly.
- **Plugins run unsandboxed.** Installed plugins execute in the same
Python process as the display loop with full file-system and
network access. Review plugin code (especially third-party plugins
+22
View File
@@ -0,0 +1,22 @@
# assets/
Static assets bundled with LEDMatrix. **Do not delete these directories** —
several look unused from core code alone but are resolved at runtime by
installed store plugins.
| Directory | Used by |
|---|---|
| `fonts/` | Core (`FontManager`, `DisplayManager`) and most plugins |
| `sports/` | Core logo tooling (`src/logo_downloader.py`) and the sports scoreboard plugins; team logos are downloaded here on demand |
| `stocks/` | `ledmatrix-stocks` plugin (`crypto_icons/`, `ticker_icons/`) |
| `weather/` | `ledmatrix-weather` plugin (weather icons) |
| `news_logos/` | `news` plugin |
| `broadcast_logos/` | `news` and `odds-ticker` plugins |
| `static_images/` | Legacy examples referenced in the `static-image` plugin's docs; the plugin itself stores uploads under `assets/plugins/<plugin-id>/uploads/` |
| `plugins/` | Per-plugin uploaded files (`assets/plugins/<plugin-id>/uploads/`), served by the web interface |
Plugins resolve these paths relative to the LEDMatrix install directory, so
the directories are part of the de-facto plugin API even where no file in
this repo references them. New plugins should bundle their own assets or
use the per-plugin upload directory instead of adding top-level
directories here.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 48 KiB

After

Width:  |  Height:  |  Size: 102 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 90 KiB

After

Width:  |  Height:  |  Size: 111 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 76 KiB

After

Width:  |  Height:  |  Size: 96 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 52 KiB

After

Width:  |  Height:  |  Size: 109 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 43 KiB

After

Width:  |  Height:  |  Size: 98 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 46 KiB

After

Width:  |  Height:  |  Size: 93 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 69 KiB

After

Width:  |  Height:  |  Size: 120 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 46 KiB

After

Width:  |  Height:  |  Size: 55 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 77 KiB

After

Width:  |  Height:  |  Size: 66 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 40 KiB

After

Width:  |  Height:  |  Size: 70 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 20 KiB

After

Width:  |  Height:  |  Size: 58 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 90 KiB

After

Width:  |  Height:  |  Size: 105 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 27 KiB

After

Width:  |  Height:  |  Size: 50 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 42 KiB

After

Width:  |  Height:  |  Size: 87 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 68 KiB

After

Width:  |  Height:  |  Size: 25 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 96 KiB

After

Width:  |  Height:  |  Size: 140 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 91 KiB

After

Width:  |  Height:  |  Size: 102 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 153 KiB

After

Width:  |  Height:  |  Size: 64 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 89 KiB

After

Width:  |  Height:  |  Size: 91 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 101 KiB

After

Width:  |  Height:  |  Size: 54 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 55 KiB

After

Width:  |  Height:  |  Size: 60 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 9.8 KiB

After

Width:  |  Height:  |  Size: 29 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 18 KiB

After

Width:  |  Height:  |  Size: 30 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 103 KiB

After

Width:  |  Height:  |  Size: 126 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 94 KiB

After

Width:  |  Height:  |  Size: 54 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 92 KiB

After

Width:  |  Height:  |  Size: 93 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 59 KiB

After

Width:  |  Height:  |  Size: 60 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 80 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 38 KiB

After

Width:  |  Height:  |  Size: 77 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 111 KiB

After

Width:  |  Height:  |  Size: 140 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 16 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 23 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 32 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 467 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 33 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 20 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 16 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 39 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 105 KiB

+29
View File
@@ -0,0 +1,29 @@
# bandit.yaml — LEDMatrix bandit configuration
# https://bandit.readthedocs.io/en/latest/config.html
#
# Skips are justified by the specific codebase context documented below.
# Do not remove skips without updating the justification comment.
skips:
# B104: Binding to all interfaces (0.0.0.0)
# Intentional — the Flask server binds 0.0.0.0 for LAN access on a Raspberry Pi.
# This is not internet-facing and is documented in web_interface/app.py.
- B104
# B603: subprocess call without shell=True
# All subprocess.run() calls in this codebase use list arguments (confirmed by
# grep — zero uses of shell=True in src/ or web_interface/). List args prevent
# shell injection. See src/common/permission_utils.py for the primary usage.
- B603
# B607: Starting a process with a partial executable path
# The subprocess calls invoke system utilities (systemctl, sudo, git) by name.
# These are fixed-list invocations, not user-controlled, and rely on PATH.
- B607
exclude_dirs:
- tests
- test
- venv
- .venv
- rpi-rgb-led-matrix-master
+42 -5
View File
@@ -1,5 +1,9 @@
{
"web_display_autostart": true,
"auto_update": {
"enabled": false,
"channel": "stable"
},
"schedule": {
"enabled": false,
"mode": "per-day",
@@ -88,6 +92,7 @@
}
},
"timezone": "America/New_York",
"target_fps": 100,
"location": {
"city": "Tampa",
"state": "Florida",
@@ -109,22 +114,56 @@
"inverse_colors": false,
"show_refresh_rate": false,
"led_rgb_sequence": "RGB",
"limit_refresh_rate_hz": 100
"limit_refresh_rate_hz": 100,
"pixel_mapper_config": "",
"orientation": "normal",
"row_address_type": 0,
"multiplexing": 0,
"panel_type": ""
},
"runtime": {
"gpio_slowdown": 3,
"rp1_rio": 0
},
"double_sided": {
"enabled": false,
"copies": 2,
"axis": "horizontal"
},
"display_durations": {},
"plugin_rotation_order": [],
"use_short_date_format": true,
"vegas_scroll": {
"live_in_ticker": true,
"live_weight": 3,
"favorite_live_weight": 5,
"enabled": false,
"scroll_speed": 50,
"separator_width": 32,
"plugin_order": [],
"excluded_plugins": [],
"target_fps": 125,
"buffer_ahead": 2
"buffer_ahead": 2,
"intra_plugin_gap": 8,
"render_width_pct": 100,
"min_content_separation": 24,
"min_cut_gap": 6,
"continuous_scroll": true,
"smooth_scroll": true,
"extend_threshold_screens": 2.0,
"auto_trim": true,
"trim_threshold": 10,
"content_padding": 8,
"min_plugin_width": 8,
"lead_in_width": 0,
"plugins_per_cycle": 6,
"max_plugin_width_ratio": 0.0,
"overflow_mode": "rotate",
"dynamic_duration_enabled": true,
"min_cycle_duration": 60,
"max_cycle_duration": 240,
"frame_based_scrolling": true,
"scroll_delay": 0.02
}
},
"sync": {
@@ -133,9 +172,7 @@
"follower_position": "left"
},
"plugin_system": {
"plugins_directory": "plugin-repos",
"auto_discover": true,
"auto_load_enabled": true
"plugins_directory": "plugin-repos"
},
"web-ui-info": {
"enabled": true,
+1 -5
View File
@@ -1,9 +1,5 @@
{
"youtube": {
"api_key": "YOUR_YOUTUBE_API_KEY",
"channel_id": "YOUR_YOUTUBE_CHANNEL_ID"
},
"github": {
"api_token": "YOUR_GITHUB_PERSONAL_ACCESS_TOKEN"
}
}
}
+6
View File
@@ -0,0 +1,6 @@
{
"dev_plugins_dir": "~/.ledmatrix-dev-plugins",
"github_user": "ChuckBuilds",
"plugins_repo": "ledmatrix-plugins",
"plugins_branch": "main"
}
+15 -7
View File
@@ -1,12 +1,20 @@
#!/usr/bin/env python3
"""Legacy entry point: runs ``run.py``, which is the one to use.
``python3 run.py`` (``-e`` for the emulator, ``-d`` for debug logging) is how
the display service and the docs start LEDMatrix. This file used to import
``src.display_controller.main`` directly, which skipped what run.py sets up
first -- ``sys.dont_write_bytecode`` (root-owned ``__pycache__`` in plugin
directories blocks the web service from updating them), the ``-e``/``-d``
flags, and the logging configuration. It now runs run.py exactly as
``python3 run.py`` would, with the same arguments.
"""
import os
import sys
# Add the project root directory to Python path
sys.path.append(os.path.dirname(os.path.abspath(__file__)))
from src.display_controller import main
import runpy
if __name__ == "__main__":
main()
runpy.run_path(
os.path.join(os.path.dirname(os.path.abspath(__file__)), "run.py"),
run_name="__main__",
)
+234
View File
@@ -0,0 +1,234 @@
# Adaptive Layout & Font Scaling
`src/adaptive_layout.py` lets a plugin render legibly on **any** panel size
(64x32, 128x32, 96x48, 128x64, 256x64, ...) without hand-tuned per-display
layouts. It is **opt-in**: nothing changes for plugins that don't use it.
It generalizes three patterns proven in the plugin ecosystem:
| Pattern | Origin | Core API |
|---|---|---|
| Geometry scale factor vs. a design size | f1-scoreboard | `ctx.px(base)` / `ctx.scale` |
| Breakpoint tiers | masters-tournament | `ctx.tier` / `ctx.by_tier({...})` |
| "Largest crisp font that fits" ladder | baseball-scoreboard | `ctx.fit_text(...)` and friends |
## Quick start
Every `BasePlugin` has a lazy `self.layout` (a `LayoutContext` for the
current logical display size, rebuilt automatically if the size changes)
and a one-liner `self.draw_fit(...)`:
```python
def display(self, force_clear=False):
from src.adaptive_layout import LADDER_ARCADE
b = self.layout.bounds.inset(1) # Region(0,0,W,H) minus 1px margin
rows = b.split_v(3, 1, 1, gap=1) # 3/5 for time, 1/5 each for the rest
self.draw_fit(self.time_str, rows[0], ladder=LADDER_ARCADE)
self.draw_fit(self.weekday, rows[1]) # default LADDER_GRID
self.draw_fit(self.date_str, rows[2])
self.display_manager.update_display()
```
On 128x64 the time renders at press_start 24px; on 64x32 it steps down to
8px. The rows partition the height, so bands can never overlap — no more
`y = height - 7` magic numbers.
## Region — rect algebra
`Region(x, y, w, h)` is a frozen dataclass. All carving clamps to
non-negative dimensions, so degenerate panels behave.
- Carving: `inset(dx, dy)`, `top_band(h)`, `bottom_band(h)`,
`middle(top_h, bottom_h)`, `left_col(w)`, `right_col(w)`,
`split_h(*weights, gap=0)`, `split_v(*weights, gap=0)`
- Placement: `align_xy(w, h, align, valign)`, `center_xy(w, h)`,
`contains(w, h)`, `.center`, `.right`, `.bottom`
Scoreboard-style layout:
```python
b = self.layout.bounds
status = b.top_band(self.layout.px(7))
detail = b.bottom_band(self.layout.px(7))
score_area = b.middle(status.h, detail.h)
away_slot, home_slot = b.left_col(b.h), b.right_col(b.h)
```
## Font ladders — discrete, never fractional
Pixel fonts (BDF, PressStart2P) only look right at native/integer sizes, so
fonts are never scaled continuously. A `FontLadder` is an ordered tuple of
`FontStep(family, size_px)` rungs, largest first; fitting walks down until
the measured text fits.
- `LADDER_GRID` (default): X11 BDFs at native sizes — 10x20 → 9x18 → 9x15 →
8x13 → 7x13 → 6x13 → 6x12 → 6x10 → 6x9 → 5x8 → 5x7 → 4x6 → tom-thumb.
Body text, labels, multi-row content.
- `LADDER_ARCADE`: PressStart2P at 32/24/16/8 (integer multiples of its 8px
grid). Headline text: clocks, scores.
Custom ladders are just tuples — e.g. to add your plugin's registered font
on top: `(FontStep("myplugin::digits", 16),) + LADDER_GRID`.
## LayoutContext
Built per (width, height); exposes facts and fit queries:
- `bounds`, `width`, `height`, `aspect`
- `tier` by height (`xs`≤16, `sm`≤32, `md`≤48, `lg`≤64, `xl`) and
`width_tier` (`narrow`≤64, `normal`≤128, `wide`≤256, `ultrawide`)
- `is_wide_short` — aspect ≥ 2.5 and height ≤ 32 (the classic 128x32 shape)
- `scale` — `min(w/design_w, h/design_h)` vs. your manifest's
`display.design_size` (default 128x32). **Geometry only** — gaps, icon
and logo sizes via `px(base, minimum, maximum)`; fonts use ladders.
- `by_tier({"sm": 10, "lg": 18})` — value for the nearest defined tier
at-or-below the panel's tier.
- `fit_text(text, box, ladder, ellipsis=True)` → `FitResult` — largest rung
that fits; ellipsizes as a last resort. Cached per (text, box, ladder).
- `fit_text_proportional(text, box, base_size_px, ladder, ellipsis=True, scale=None)` —
rung closest to (not exceeding) `base_size_px * scale`, still capped to
what fits the box. Use this instead of `fit_text` when several
independently-fitted elements need to stay visually harmonious as the
panel grows — `fit_text` maximizes *each one* within its own region,
which can make one element (e.g. a score with a generous box) balloon
out of proportion to a neighbor that scales by geometry (e.g. logos
sized via `px()`), even though each individual pick is "correct" in
isolation. `base_size_px` is normally the element's existing classic/
fixed font size. `scale` defaults to `self.scale` (the conservative
min-of-both-axes factor `px()` uses); pass an axis-specific value when
the surrounding composition already scales that way — e.g. a scoreboard
whose logo slots track height alone (`min(height, width // 2)`) should
size its text by `height / design_height` too, or the text reads as
under-scaled next to bigger logos on a panel that only grew taller.
- `fit_lines(lines, box, ladder, spacing)` — every line fits the width and
the stack fits the height (measures the actual strings).
- `font_for_rows(rows, box_h, ladder)` — largest rung whose line height
fits `rows` rows.
`FitResult` carries the ready-to-use `font` (drops straight into
`display_manager.draw_text(font=...)`), the possibly-ellipsized `text`,
ink `width`/`height`, `baseline`, `y_offset`, `line_height`, and `fits`.
## Adaptive images
`src/adaptive_images.py` is the image counterpart to `fit_text`, exposed as
`self.layout.fit_image(...)` (cached per panel size) and the one-liner
`self.draw_image(...)`:
```python
# Team logo: trim its transparent padding, fill the slot height (the
# football/hockey pattern), cached across frames by a stable key
self.draw_image(logo, regs.away_slot, mode="fill_height",
crop_to_ink=True, cache_key=f"logo:{abbr}")
# Album art: cover-crop a square, faces kept by the top anchor
self.draw_image(art, row.art, mode="cover", anchor="top")
# Pixel flags / sprite icons: NEAREST keeps hard edges
from src.adaptive_images import RESAMPLE_NEAREST
self.draw_image(flag, box, resample=RESAMPLE_NEAREST)
```
Modes: `contain` (letterbox, default), `cover` (crop-to-fill),
`fill_height` (logo-style), `stretch`. Unlike PIL's `thumbnail()`
(downscale-only — why imagery stays tiny on big panels) fitting **upscales
by default**; pass `upscale=False` for the legacy behavior. Results are
cached per (image, box size, options) with a bounded LRU — always pass a
stable `cache_key` (e.g. `"logo:KC"`) for images you reload. The module
also exports the Pillow-compat `RESAMPLE_LANCZOS`/`RESAMPLE_NEAREST`
constants so plugins can drop their local shims.
## Composite layouts
Pre-carved Region arrangements for the layouts plugins keep rebuilding:
```python
from src.adaptive_layout import scoreboard_regions, media_row
regs = scoreboard_regions(self.layout.bounds, ctx=self.layout)
# regs.away_slot / home_slot — logo slots (logo_slot = min(H, W // 2),
# capped so a center reserve always exists —
# see below)
# regs.status_band — top band (replaces the magic y = 1)
# regs.score_area — center gap, plus a controlled bleed into
# each logo slot (replaces y = H//2 - 3)
# regs.detail_band — bottom band (replaces y = H - 7)
# regs.bottom_left / bottom_right — record/timeout corners
row = media_row(self.layout.bounds, ctx=self.layout) # art left, text right
```
Both work on the full panel or on a scroll-mode card Region. They return
Regions and never draw — compose them with `draw_fit`/`draw_image`.
**`scoreboard_regions`'s center reserve.** The raw `logo_slot = min(H, W//2)`
formula has a blind spot: at exactly 2:1 aspect ratio (width = 2×height —
two, four, or more square modules stacked into a taller panel, e.g.
96x48, 128x64, 256x128) the two logo slots mathematically claim the
*entire* width, leaving zero pixels for a center column no matter how
big the panel gets. Wide panels (the 128x32 design baseline, 192x48,
256x32) never hit this, since height is already the tighter constraint
there. Two parameters fix it without any plugin-side code:
`min_center_fraction`/`min_center_design_px` guarantee a real minimum
center reserve at any aspect ratio, and `score_bleed_fraction` lets the
score's *fit box* extend a controlled amount into each logo slot — the
same way a real broadcast scoreboard's numbers cross slightly into the
team marks flanking them — so a short score string never has to truncate
even on the tightest aspect ratios. All three have sane defaults; override
them per call if a plugin's card proportions genuinely differ.
## Preserving user customization
Adaptive layout supplies *defaults*; explicit user configuration wins:
- **User-set fonts win.** If the plugin's config has an explicit
`font`/`font_size` for an element, load it as before and skip the ladder —
fit only when the user hasn't overridden (see the football-scoreboard
`_resolve_element_fit` pattern).
- **Offsets apply on top.** `customization.layout.<element>.{x_offset,y_offset}`
style knobs translate the *computed* region as a final step:
`region.offset(user_dx, user_dy)`. `draw_image(..., offset=(dx, dy))`
does the same for images.
- **Colors pass through.** `draw_fit`/`draw_fitted_text` take explicit
`color=` params; adaptive mode never repaints semantic or user-chosen
colors.
## Manifest declaration
Declare the size your layout was authored against so `ctx.scale` means
something:
```json
"display": { "design_size": { "width": 128, "height": 32 } }
```
Also available under `requires.display_size`: `min_width`, `min_height`,
`max_width`, `max_height`.
## Performance notes (Pi)
Fit queries are cached, so cost is O(unique strings). For per-second text
(clocks, live scores), fit on a **shape placeholder** and reuse the font:
```python
fit = self.layout.fit_text("00:00", box, ladder=LADDER_ARCADE) # cached once
self.display_manager.draw_text(current_time, font=fit.font, ...)
```
## Testing across sizes
The harness already renders every plugin at a spread of sizes (now
including 96x48):
```bash
python scripts/check_plugin.py --plugin <plugin-id> --sizes 64x32,128x32,96x48,128x64,256x64
python scripts/render_plugin.py --plugin <plugin-id> --width 96 --height 48
```
`BoundsCheckingDisplayManager` flags right/bottom overflow and now records
mediated draw calls with negative coordinates in
`negative_coordinate_calls` (raw-PIL draws remain uncovered).
Reference migration: the **text-display** plugin's `font_mode: "auto"`.
+298 -164
View File
@@ -10,22 +10,27 @@ This guide covers advanced LEDMatrix features for users and developers, includin
Vegas scroll mode displays content from multiple plugins in a continuous horizontal scroll, similar to news tickers seen in Las Vegas casinos. Plugins contribute content segments that flow across the display in a seamless ticker-style presentation.
### Display Modes
### How a Plugin Takes Part
**SCROLL (Continuous Scrolling):**
- Content scrolls continuously left
- Smooth, fluid motion
- Best for news-ticker style displays
Each plugin has a *Vegas participation*:
**FIXED_SEGMENT (Fixed-Width Block):**
- Plugin gets fixed-width block on display
- Content doesn't scroll out of its segment
- Multiple plugins can share the display simultaneously
**`scroll` (the default):**
- The plugin's content scrolls by with everyone else's
- Best for news-ticker style content: scores, headlines, prices, the time
**STATIC (Scroll Pauses):**
- Scrolling pauses when content is fully visible
- Displays for specified duration, then resumes scrolling
- Best for content that needs to be fully read
**`pause`:**
- The scroll stops when the plugin's turn comes round
- The plugin draws the whole panel for its display duration, then the
scroll resumes
- Best for content that needs to be read in full, or alerts
**`exclude`:**
- The plugin is left out of Vegas mode
A plugin declares its default; set `vegas_participation` in a plugin's
config to override it (see [Per-Plugin Configuration](#per-plugin-configuration)).
Older documentation also describes a *fixed segment* mode; Vegas never
implemented one, and it has always behaved exactly like `scroll`.
### Configuration
@@ -47,6 +52,11 @@ Enable Vegas mode in `config/config.json`:
}
```
Vegas mode can also be configured entirely from the web UI — the
**Display** tab has a Vegas Scroll Mode section (enable toggle, scroll
speed, separator width, dynamic duration, and more), so hand-editing
JSON is optional.
**Configuration Options:**
| Setting | Default | Description |
@@ -57,7 +67,113 @@ Enable Vegas mode in `config/config.json`:
| `plugin_order` | `[]` | Plugin display order (empty = auto) |
| `excluded_plugins` | `[]` | Plugins to exclude from Vegas mode |
| `target_fps` | `125` | Target frame rate |
| `buffer_ahead` | `2` | Number of panels to render ahead |
| `buffer_ahead` | `2` | Number of plugins buffered ahead |
This table is a subset — `display.vegas_scroll` supports 30 keys in
total. See the full list in
[CONFIG_REFERENCE.md](CONFIG_REFERENCE.md#displayvegas_scroll--continuous-scroll-mode).
### Live Content in the Ticker
By default (since 3.8.0) live content **stays in the ticker** and takes
**extra turns inside it**, and a scoreboard that supports live cards updates
the score on a card already crossing the screen (`live_refresh`, "Update live
content while it scrolls").
To get the old behaviour back -- live content **preempts** Vegas mode: while
any plugin reports live priority the ticker stops and that plugin's
full-screen display is shown instead -- untick **Keep live games in the
ticker** under Vegas mode, or set `live_in_ticker` to `false`:
```json
"vegas_scroll": {
"live_in_ticker": false
}
```
Until 3.8.0 `false` was the default and every config held it, copied from
the template. The first start on 3.8.0 turns it on once (a backup of the
config is kept as `config.json.backup`, and `live_in_ticker_migrated` records
that it ran); a `false` set after that is left alone.
The weights below apply while live content is in the ticker:
```json
"vegas_scroll": {
"live_weight": 3,
"favorite_live_weight": 5
}
```
#### Why weights exist
The rotation is otherwise a strict round robin — every plugin appears exactly
once per cycle. With a dozen plugins enabled, a live score comes round once a
lap and can be minutes old by the time you see it. A weight of *N* gives a
plugin *N* slots per cycle.
The slots are placed by **Smooth Weighted Round-Robin**, the same scheduler
the sports plugins use internally to rotate their own games. The important
property is that repeats are *spread through the cycle* rather than clumped:
three appearances in a row followed by a long silence would be worse than not
boosting at all.
Twelve plugins, with a favorite's baseball game and an ordinary live hockey
game (`live_weight: 3`, `favorite_live_weight: 5`):
```
baseball > hockey > weather > clock > baseball
stocks > news > flights > baseball > hockey
calendar > f1 > music > baseball > tides
birds > hockey > baseball
```
18 slots for 12 plugins. Baseball appears 5 times, hockey 3, everything else
once, and no plugin ever appears twice in a row — **including across the seam**
where the cycle loops back on itself. Smooth Weighted Round-Robin schedules the
heaviest item first and usually last as well, so the strip would otherwise show
it twice running at exactly the one join a within-cycle check cannot see. The
trailing repeat is moved into the widest remaining gap. Where a double is
unavoidable — a plugin holding most of the slots has to neighbour itself — the
schedule is left as it is.
#### Where the weight comes from
For each plugin in the rotation, in order:
1. **The plugin's own answer.** If it implements
`get_vegas_priority_weight()` and returns a number, that wins. This is the
only route for favorite-team awareness — the core can see *that* a game is
live, but not *whose*, so a scoreboard has to say so itself.
2. **The core's default.** When the plugin returns `None` (the base-class
default), a plugin where both `has_live_priority()` and `has_live_content()`
are true gets `live_weight`.
3. **Everything else** gets 1.
Because of step 2, **existing plugins need no changes** — any scoreboard with
`live_priority` enabled already gets extra turns. Step 1 is opt-in, for
plugins that want to distinguish a favorite's game from any other live game.
Weights are clamped to 1–10. A weight of 1 is no boost; a weight below 1 would
drop the plugin from the rotation entirely, which is never what is meant.
#### Things worth knowing
- **Weights are per plugin, not per game.** A scoreboard showing four live
games still occupies one slot at a time, rotating its own games within that
slot using its own `favorite_live_boost`. This controls how often the
*plugin* comes round.
- **The ticker is zero-sum.** Giving baseball 5 slots does not make the cycle
faster; it makes the cycle *longer* and everything else proportionally
rarer. If you want live scores sooner in wall-clock terms, pair this with a
smaller `plugins_per_cycle`.
- **Frequency is not freshness.** Each appearance redraws from the plugin's
current data (`refresh_updated_plugins()` drops cached content when a
plugin's data changes), but how current that data is depends on the
plugin's own `live_update_interval`. Showing a stale score five times a lap
is no better than showing it once.
- **Everything still appears.** A boost never starves another plugin out of
the cycle; low-weight plugins keep their single slot.
### Per-Plugin Configuration
@@ -67,8 +183,7 @@ Override Vegas behavior for specific plugins:
{
"my_plugin": {
"enabled": true,
"vegas_mode": "scroll",
"vegas_panel_count": 2,
"vegas_participation": "pause",
"display_duration": 10
}
}
@@ -78,77 +193,80 @@ Override Vegas behavior for specific plugins:
| Setting | Values | Description |
|---------|--------|-------------|
| `vegas_mode` | `scroll`, `fixed`, `static` | Display mode for this plugin |
| `vegas_panel_count` | `1-10` | Width in panels (1 panel = display width) |
| `display_duration` | seconds | Pause duration for STATIC mode |
| `vegas_participation` | `scroll`, `pause`, `exclude` | How this plugin takes part: its content scrolls by, the scroll pauses for its turn and shows it full screen, or it is left out. Unset uses the plugin's own default |
| `display_duration` | seconds | How long a `pause` plugin holds the screen |
| `vegas_width_pct` | 10–100 | Width of this plugin's card, as a percentage of the panel |
| `vegas_overflow` | `rotate`, `truncate` | What to do when its content is wider than its allowance |
| `vegas_max_width_screens` | number of screens | The widest its card may be |
These are core-owned settings (see
[PLUGIN_CONFIG_CORE_PROPERTIES.md](PLUGIN_CONFIG_CORE_PROPERTIES.md)): every
plugin accepts them whether or not its own schema lists them. Set them in
the plugin's section of config.json, in the web UI's **Config Editor**
tab.
Some plugins also offer a `vegas_mode` setting of their own (`scroll`,
`fixed` or `static`). It still works — `static` pauses, the other two scroll
— but `vegas_participation` takes precedence, and `fixed` has never done
anything different from `scroll`. The old `vegas_panel_count` setting never
had an effect and is deprecated (removed in 3.9.0).
### Plugin Integration (Developer Guide)
All of these have defaults in
[`BasePlugin`](../src/plugin_system/base_plugin.py); override only what you
need. The reference is
[PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md#vegas-scroll-hooks).
**1. Implement Content Method:**
```python
def get_vegas_content(self):
"""
Return PIL Image or list of Images for Vegas mode.
Returns:
PIL.Image or list[PIL.Image]: Content to display
- Single image: fixed-width content
- List of images: multiple segments
- None: skip this cycle
"""
# Example: Return single wide image
img = Image.new('RGB', (256, 32))
# ... render your content ...
return img
# Example: Return multiple segments
return [image1, image2, image3]
# Return a PIL Image, a list of Images, or None.
# A single image is one block; a list becomes one item per image.
return [self._render_game(game) for game in self.games]
```
**2. Specify Content Type:**
If it returns `None` (the default), Vegas falls back to the plugin's
`scroll_helper` image, then to capturing `display()` output
(`PluginAdapter.get_content()` in
[`src/vegas_mode/plugin_adapter.py`](../src/vegas_mode/plugin_adapter.py)).
**2. Declare how the plugin takes part:**
Most plugins need nothing: the default is `scroll`. A plugin that should
pause the scroll, or stay out of Vegas, says so in `manifest.json`:
```json
{
"vegas_participation": "pause"
}
```
The user's own `vegas_participation` setting overrides the manifest. When
the answer depends on state, override the method instead:
```python
def get_vegas_content_type(self):
"""
Specify how content should be handled.
Returns:
str: 'multi' | 'static' | 'none'
"""
return 'multi' # Default for most plugins
def get_vegas_participation(self):
# 'scroll' | 'pause' | 'exclude'
return 'pause' if self._alert_is_live() else 'scroll'
```
**3. Optionally Specify Display Mode:**
```python
def get_vegas_display_mode(self):
"""
Preferred display mode for this plugin.
Returns:
str: 'scroll' | 'fixed' | 'static'
"""
return 'scroll'
def get_supported_vegas_modes(self):
"""
List of supported modes.
Returns:
list: ['scroll', 'fixed', 'static']
"""
return ['scroll', 'static']
```
A plugin written for an older core that declares nothing keeps its
behaviour: `get_vegas_display_mode()` returning `VegasDisplayMode.STATIC`
pauses, `get_vegas_content_type()` returning `'none'` excludes, and
everything else scrolls. `get_supported_vegas_modes()`,
`get_vegas_segment_width()` and the SCROLL / FIXED_SEGMENT distinction are
deprecated (removed in 3.9.0): Vegas never read them.
### Content Rendering Guidelines
**Image Dimensions:**
- **Height:** Must match display height (typically 32 pixels)
- **Width:** Varies by mode:
- SCROLL: Any width (recommended 64-512 pixels)
- FIXED_SEGMENT: `panel_count * display_width`
- STATIC: Any width, optimized for readability
- **Width:** Any width for `scroll` (recommended 64-512 pixels);
`get_vegas_render_width()` is the width Vegas would like, and it narrows
`display_manager` to match while it asks. A `pause` plugin draws the
whole panel in `display()`.
**Color Mode:**
- Use RGB color mode
@@ -198,17 +316,10 @@ class WeatherPlugin(BasePlugin):
def get_vegas_content(self):
"""Return cached Vegas image"""
return self.vegas_image
def get_vegas_content_type(self):
return 'multi'
def get_vegas_display_mode(self):
return 'scroll'
def get_supported_vegas_modes(self):
return ['scroll', 'static']
```
It scrolls, the default participation, so it declares nothing else.
### System Architecture
Vegas mode consists of four core components working together to provide smooth 125 FPS continuous scrolling:
@@ -276,15 +387,23 @@ Vegas mode consists of four core components working together to provide smooth 1
5. Compose into continuous stream with separators
**Key Methods:**
- `get_stream_content()` - Returns current stream content as PIL Image
- `advance_stream(pixels)` - Advances stream by N pixels
- `refresh_stream()` - Regenerates stream from current plugins
- `get_next_segment()` - Returns the next buffered `ContentSegment` (or `None`)
- `take_next_group(count=None, offscreen_only=False)` - Hands over the next
slice of the rotation as `(plugin_id, images)` groups
- `get_grouped_content_for_composition()` - Buffered images grouped by plugin
- `mark_plugin_updated(plugin_id)` / `process_updates()` - Refresh one
plugin's segment in place when its data changes
- `refresh()` - Re-read the plugin list and config
- `advance_cycle()` - Clear the active buffer when a scroll cycle completes
(`src/vegas_mode/stream_manager.py`)
#### 3. PluginAdapter
**Responsibilities:**
- Convert plugin content to scrollable images
- Handle different Vegas display modes (SCROLL, FIXED, STATIC)
- Fetch the content of `scroll` plugins (a `pause` plugin is drawn by
its own `display()` when the scroll pauses; see StreamManager)
- Manage fallback for plugins without Vegas support
- Cache plugin content for performance
@@ -293,21 +412,16 @@ Vegas mode consists of four core components working together to provide smooth 1
- Calls `get_vegas_content()` if available
- Falls back to `display()` method if not
2. **Handle display mode:**
- SCROLL: Returns image as-is for continuous scrolling
- FIXED_SEGMENT: Creates fixed-width block (panel_count * display_width)
- STATIC: Marks content for pause-when-visible behavior
3. **Content type handling:**
- `multi`: Multiple segments (list of images)
- `static`: Single static image
- `none`: Skip this plugin in current cycle
2. **Participation** is decided by the StreamManager, not here
(`resolve_vegas_participation()` in
[`base_plugin.py`](../src/plugin_system/base_plugin.py)): `exclude`
plugins never reach the adapter, and `pause` plugins are not fetched.
**Fallback Behavior:**
- If plugin doesn't implement Vegas methods:
- Calls plugin's `display()` method
- Captures rendered display as static image
- Treats as fixed segment
- Scrolls it by as one block
- Ensures all plugins work in Vegas mode without explicit support
#### 4. RenderPipeline
@@ -332,10 +446,14 @@ Vegas mode consists of four core components working together to provide smooth 1
- **Frame Rate Control:** Precise timing to maintain 125 FPS
- **Pre-rendered Content:** Plugins pre-render during update()
**Scroll Speed Calculation:**
**Scroll Speed Calculation:** motion is by elapsed time; `target_fps` paces
the render loop, not the speed.
```python
pixels_per_frame = (scroll_speed / target_fps)
scroll_position += pixels_per_frame * elapsed_time
# frame_based_scrolling: false
scroll_position += scroll_speed * elapsed_time # scroll_speed in px/s
# frame_based_scrolling: true (the default) -- not stepping, just a clamp
applied = clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay
scroll_position += applied * elapsed_time
```
#### Component Interactions
@@ -406,7 +524,7 @@ All components use thread-safe patterns:
If a plugin doesn't implement Vegas methods:
- System calls the plugin's `display()` method
- Captures the rendered display as a static image
- Treats it as a fixed segment
- Scrolls it by as one block
This ensures all plugins work in Vegas mode, even without explicit support.
@@ -451,7 +569,8 @@ time when something is active.
### REST API Reference
The API is mounted at `/api/v3` (`web_interface/app.py:144`).
The API is mounted at `/api/v3` (the `api_v3` blueprint, registered in
`web_interface/app.py`). Full details: [REST_API_REFERENCE.md](REST_API_REFERENCE.md#display-control).
#### Start On-Demand Display
@@ -507,24 +626,36 @@ curl http://localhost:5000/api/v3/display/on-demand/status
# Response:
{
"active": true,
"plugin_id": "weather",
"mode": "weather",
"remaining": 25.5,
"pinned": false,
"status": "active"
"status": "success",
"data": {
"state": {
"active": true,
"plugin_id": "weather",
"mode": "weather",
"duration": 30,
"pinned": false,
"status": "running",
"last_updated": 1234567890.1
},
"service": {"active": true, "returncode": 0, "stdout": "active", "stderr": ""}
}
}
```
When nothing is running on demand, `data.state` is
`{"active": false, "status": "idle", "last_updated": null}`.
> There is no public Python on-demand API. The display controller's
> on-demand machinery is internal — drive it through the REST endpoints
> above (or the web UI buttons), which write a request into the cache
> manager under the `display_on_demand_request` key
> (`web_interface/blueprints/api_v3.py:1622,1687`) that the controller
> polls at `src/display_controller.py:921`. A separate
> above (or the web UI buttons). The API handlers
> (`start_on_demand_display()` / `stop_on_demand_display()` in
> `web_interface/blueprints/api_v3/display.py`) write a request into the cache
> manager under the `display_on_demand_request` key, which
> `DisplayController._poll_on_demand_requests()`
> (`src/display_controller.py`) picks up. A separate
> `display_on_demand_config` key is used by the controller itself
> during activation to track what's currently running (written at
> `display_controller.py:1195`, cleared at `:1221`).
> during activation (`_activate_on_demand()`) to track what's
> currently running, and is cleared by `_clear_on_demand()`.
### Duration Modes
@@ -646,13 +777,13 @@ keys helps troubleshoot stuck states.
**When Set:** Every display loop iteration
**Auto-Cleared:** Never (continuously updated)
**4. display_on_demand_processed_id** (TTL: 5 minutes)
```
**4. display_on_demand_processed_id** (TTL: 1 hour)
```text
"uuid-string-of-last-processed-request"
```
**Purpose:** Prevents duplicate request processing
**When Set:** After processing request
**Auto-Cleared:** After 5 minutes TTL
**Auto-Cleared:** After 1 hour TTL
### When Manual Clearing is Needed
@@ -685,9 +816,9 @@ keys helps troubleshoot stuck states.
The cache is stored as JSON files under one of:
- `/var/cache/ledmatrix/` (preferred when the service has permission)
- `~/.cache/ledmatrix/`
- `~/.ledmatrix_cache/`
- `/opt/ledmatrix/cache/`
- `/tmp/ledmatrix-cache/` (fallback)
- `$TMPDIR/ledmatrix_cache/` (fallback)
```bash
# Find the cache dir actually in use
@@ -711,8 +842,9 @@ cache.clear_cache('display_on_demand_request')
cache.clear_cache('display_on_demand_processed_id')
```
> The actual public method is `clear_cache(key=None)` — there is no
> `delete()` method on `CacheManager`.
> `CacheManager` also has a `delete(key)` method — a thin wrapper over
> `clear_cache(key)` — so `cache.delete('display_on_demand_config')`
> works equally well.
### Cache Impact on Running Service
@@ -730,7 +862,7 @@ The display controller automatically handles cleanup:
- **Config key**: Cleared when on-demand stops
- **State key**: Updated every display loop iteration
- **Request key**: Expires after 1 hour TTL (or after processing)
- **Processed ID**: Expires after 5 minutes TTL
- **Processed ID**: Expires after 1 hour TTL
---
@@ -760,7 +892,13 @@ Cache Check → Background Fetch → Partial Data → Completion → Cache
### Configuration
Enable background service per plugin in `config/config.json`:
Core does not read a `background_service` config block: the service itself
(`src/background_data_service.py`) is a process-wide singleton, and its
worker count is whatever the first caller of `get_background_service()`
passes. The sports scoreboard plugins read their own
`background_service` settings and pass them to it, so the exact keys and
where they sit (top level or per league) are defined by each plugin's
`config_schema.json`. A typical block looks like:
```json
{
@@ -781,11 +919,11 @@ Enable background service per plugin in `config/config.json`:
| Setting | Default | Description |
|---------|---------|-------------|
| `enabled` | `false` | Enable background service for this plugin |
| `enabled` | plugin-defined | Use the background service for this plugin's fetches |
| `max_workers` | `3` | Max concurrent background tasks |
| `request_timeout` | `30` | Timeout per API request (seconds) |
| `max_retries` | `3` | Retry attempts on failure |
| `priority` | `1` | Task priority (1=highest, 10=lowest) |
| `priority` | `1` | Stored on each request (higher number = higher priority, per `FetchRequest`), but the service runs requests in submission order; it does not reorder by priority |
### Performance Impact
@@ -802,9 +940,9 @@ Enable background service per plugin in `config/config.json`:
The background data service is used by all of the sports scoreboard
plugins (football, hockey, baseball/MLB, basketball, soccer, lacrosse,
F1, UFC), the odds ticker, and the leaderboard plugin. Each plugin's
`background_service` block (under its own config namespace) follows the
same shape as the example above.
F1, UFC), the odds ticker, and the leaderboard plugin. Each plugin reads
its own `background_service` block (under its own config namespace); check
that plugin's `config_schema.json` for the keys it accepts.
### Error Handling & Fallback
@@ -821,9 +959,6 @@ same shape as the example above.
### Testing
```bash
# Run background service test
python test_background_service.py
# Check logs for background operations
sudo journalctl -u ledmatrix -f | grep "background"
```
@@ -832,15 +967,21 @@ sudo journalctl -u ledmatrix -f | grep "background"
**View Statistics:**
```python
from src.background_data_service import BackgroundDataService
from src.background_data_service import get_background_service
from src.cache_manager import CacheManager
service = BackgroundDataService()
service = get_background_service(CacheManager())
stats = service.get_statistics()
print(f"Active tasks: {stats['active_tasks']}")
print(f"Completed: {stats['completed']}")
print(f"Failed: {stats['failed']}")
print(f"Active: {stats['active_requests']}")
print(f"Completed: {stats['completed_requests']}")
print(f"Failed: {stats['failed_requests']}")
```
Other keys: `total_requests`, `cached_hits`, `cache_misses`,
`average_fetch_time`, `completed_requests_count` (results currently held in
memory) — see `BackgroundDataService.get_statistics()` in
[`src/background_data_service.py`](../src/background_data_service.py).
**Enable Debug Logging:**
```python
import logging
@@ -851,6 +992,10 @@ logging.getLogger('src.background_data_service').setLevel(logging.DEBUG)
## 5. Permission Management
Ownership, modes, sudo rules and the repair scripts are listed in
[PERMISSIONS.md](PERMISSIONS.md). This section covers the helpers code uses
to keep files shareable.
### Overview
LEDMatrix uses a dual-user architecture: the display service runs as root (hardware access), while the web interface runs as a non-privileged user. Centralized permission management ensures both can access necessary files.
@@ -875,6 +1020,7 @@ from src.common.permission_utils import (
ensure_file_permissions,
get_config_file_mode,
get_assets_file_mode,
get_assets_dir_mode,
get_plugin_file_mode,
get_cache_dir_mode
)
@@ -883,7 +1029,10 @@ from src.common.permission_utils import (
ensure_directory_permissions(Path("assets/sports"), get_assets_dir_mode())
# Set file permissions after writing
ensure_file_permissions(Path("config/config.json"), get_config_file_mode())
# (get_config_file_mode requires the file path — secrets files get a
# stricter mode than the main config)
config_path = Path("config/config.json")
ensure_file_permissions(config_path, get_config_file_mode(config_path))
```
### When to Use Utilities
@@ -910,7 +1059,7 @@ ensure_file_permissions(Path("config/config.json"), get_config_file_mode())
| Config (secrets) | `rw-r-----` | `0o640` | Owner write, group read |
| Assets | `rw-rw-r--` | `0o664` | Owner/group write, all read |
| Plugins | `rw-rw-r--` | `0o664` | Owner/group write, all read |
| Cache files | `rw-rw-r--` | `0o664` | Owner/group write, all read |
| Cache files | `rw-rw----` | `0o660` | Owner/group write, no world access (`_CACHE_FILE_MODE` in `src/cache/disk_cache.py`) |
**Directory Permissions:**
@@ -938,7 +1087,7 @@ from src.common.permission_utils import ensure_file_permissions, get_config_file
config_path = Path("config/config.json")
with open(config_path, 'w') as f:
json.dump(data, f)
ensure_file_permissions(config_path, get_config_file_mode())
ensure_file_permissions(config_path, get_config_file_mode(config_path))
```
**Pattern 3: Downloading Logo**
@@ -981,43 +1130,28 @@ These core utilities **already handle permissions** - you don't need to call per
### Manual Fixes
If you encounter permission issues:
[PERMISSIONS.md](PERMISSIONS.md) lists who owns what on an installed system,
the expected modes, and which `scripts/fix_perms/` script to run as which
user. In short:
```bash
# Fix all permissions at once
sudo ./scripts/fix_permissions.sh
- `fix_assets_permissions.sh`, `fix_cache_permissions.sh` and
`fix_plugin_permissions.sh` are run with `sudo`.
- `fix_web_permissions.sh` is run as the web interface user, without
`sudo` (it refuses to run as root and calls `sudo` itself where needed).
It resets project file ownership for that user, then makes the two
helper scripts the web user may run as root (`safe_plugin_rm.sh`,
`safe_pip_install.sh`) root-owned again and restores `config_secrets.json`
to its owner, the `ledmatrix` group and mode `640`. It does not write
sudoers rules; `scripts/install/configure_web_sudo.sh` does that.
# Fix specific directory
sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/config
sudo chmod -R 2775 /home/ledpi/LEDMatrix/config
sudo find /home/ledpi/LEDMatrix/config -type f -exec chmod 664 {} \;
# Verify permissions
ls -la config/
ls -la assets/
```
### Verification
```bash
# Check directory has setgid bit
ls -ld assets/
# Should show: drwxrwsr-x (note the 's')
# Check file has correct group
ls -l assets/logo.png
# Should show group 'ledpi'
# Check file permissions
stat -c "%a %n" config/config.json
# Should show: 644 config/config.json
```
Do not `chmod` the whole `config/` directory: `config_secrets.json` must stay
`640`.
---
## Related Documentation
- [PLUGIN_DEVELOPMENT.md](PLUGIN_DEVELOPMENT.md) - Creating plugins with Vegas/on-demand support
- [PLUGIN_DEVELOPMENT_GUIDE.md](PLUGIN_DEVELOPMENT_GUIDE.md) - Creating plugins with Vegas/on-demand support
- [WEB_INTERFACE_GUIDE.md](WEB_INTERFACE_GUIDE.md) - Using on-demand controls in web UI
- [PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md) - Complete API documentation
- [DEVELOPMENT.md](DEVELOPMENT.md) - Development environment and testing
+68 -102
View File
@@ -2,12 +2,18 @@
Advanced patterns, examples, and best practices for developing LEDMatrix plugins.
> **Adaptive layout:** for plugins that should render legibly on any panel
> size (fonts that grow on big panels, layouts that degrade gracefully on
> small ones), use the adaptive layout system — `self.layout`, `draw_fit`,
> `draw_image`, `scoreboard_regions` — documented in
> [ADAPTIVE_LAYOUT.md](ADAPTIVE_LAYOUT.md).
## Table of Contents
- [Using Weather Icons](#using-weather-icons)
- [Implementing Scrolling with Deferred Updates](#implementing-scrolling-with-deferred-updates)
- [Cache Strategy Patterns](#cache-strategy-patterns)
- [Font Management and Overrides](#font-management-and-overrides)
- [Font Management](#font-management)
- [Error Handling Best Practices](#error-handling-best-practices)
- [Performance Optimization](#performance-optimization)
- [Testing Plugins with Mocks](#testing-plugins-with-mocks)
@@ -19,69 +25,12 @@ Advanced patterns, examples, and best practices for developing LEDMatrix plugins
## Using Weather Icons
The Display Manager provides built-in weather icon drawing methods for easy visual representation of weather conditions.
### Basic Weather Icon Usage
```python
def display(self, force_clear=False):
if force_clear:
self.display_manager.clear()
# Draw weather icon based on condition
condition = self.data.get('condition', 'clear')
self.display_manager.draw_weather_icon(condition, x=5, y=5, size=16)
# Draw temperature next to icon
temp = self.data.get('temp', 72)
self.display_manager.draw_text(
f"{temp}°F",
x=25, y=10,
color=(255, 255, 255)
)
self.display_manager.update_display()
```
### Supported Weather Conditions
The `draw_weather_icon()` method automatically maps condition strings to appropriate icons:
- `"clear"`, `"sunny"` → Sun icon
- `"clouds"`, `"cloudy"`, `"partly cloudy"` → Cloud icon
- `"rain"`, `"drizzle"`, `"shower"` → Rain icon
- `"snow"`, `"sleet"`, `"hail"` → Snow icon
- `"thunderstorm"`, `"storm"` → Storm icon
### Custom Weather Icons
For more control, use individual icon methods:
```python
# Draw specific icons
self.display_manager.draw_sun(x=10, y=10, size=16)
self.display_manager.draw_cloud(x=10, y=10, size=16, color=(150, 150, 150))
self.display_manager.draw_rain(x=10, y=10, size=16)
self.display_manager.draw_snow(x=10, y=10, size=16)
```
### Text with Weather Icons
Use `draw_text_with_icons()` to combine text and icons:
```python
icons = [
("sun", 5, 5), # Sun icon at (5, 5)
("cloud", 100, 5) # Cloud icon at (100, 5)
]
self.display_manager.draw_text_with_icons(
"Weather: Sunny, Cloudy",
icons=icons,
x=10, y=20,
color=(255, 255, 255)
)
```
The Display Manager's icon methods — `draw_weather_icon()`, `draw_sun()`,
`draw_cloud()`, `draw_rain()`, `draw_snow()` and `draw_text_with_icons()` —
were removed in 3.8.0. Draw your own icons instead: render them
onto a PIL image and paste it onto `self.display_manager.image`, or ship
icon images with the plugin. The weather plugin's `WeatherIcons` class is an
example. See [Deprecated APIs](PLUGIN_API_REFERENCE.md#deprecated-apis).
---
@@ -91,31 +40,53 @@ For plugins that scroll content (tickers, news feeds, etc.), use scrolling state
### Basic Scrolling Implementation
Scroll with `ScrollHelper`, configured by `src.common.scroll_config`, and
render one frame per `display()` call. Don't pace the scroll with
`time.sleep()`: `update_display()` blocks on the panel's
vsync, which is what paces a scroll. Pass the `frame_hold` that
`scroll_config.configure()` returned to `set_scrolling_state()`, or the
scroll runs faster than the configured speed (see
`set_scrolling_state()` in [PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md)).
```python
from PIL import Image, ImageDraw
from src.common import scroll_config
from src.common.scroll_helper import ScrollHelper
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.scroll_helper = ScrollHelper(
self.display_manager.width, self.display_manager.height, self.logger)
self.scroll_settings = scroll_config.configure(
self.scroll_helper,
plugin_config=self.config,
global_config=self.global_config,
display_manager=self.display_manager,
plugin_logger=self.logger,
)
def _build_scroll_image(self, text):
font = self.display_manager.regular_font
width = self.display_manager.get_text_width(text, font)
img = Image.new("RGB", (width, self.display_manager.height))
ImageDraw.Draw(img).text((0, 0), text, font=font, fill=(255, 255, 255))
self.scroll_helper.set_scrolling_image(img)
def display(self, force_clear=False):
if force_clear:
self.display_manager.clear()
# Mark as scrolling
self.display_manager.set_scrolling_state(True)
try:
# Scroll content
text = "This is a long scrolling message that needs to scroll across the display..."
text_width = self.display_manager.get_text_width(text, self.display_manager.regular_font)
display_width = self.display_manager.width
# Scroll from right to left
for x in range(display_width, -text_width, -2):
self.display_manager.clear()
self.display_manager.draw_text(text, x=x, y=16, color=(255, 255, 255))
self.display_manager.update_display()
time.sleep(0.05)
# Update scroll activity timestamp
self.display_manager.set_scrolling_state(True)
finally:
# Always mark as not scrolling when done
if force_clear or self.scroll_helper.cached_image is None:
self._build_scroll_image(
"This is a long scrolling message that needs to scroll across the display...")
# Mark as scrolling (calling it every frame is fine)
self.display_manager.set_scrolling_state(
True, frame_hold=self.scroll_settings.frame_hold)
self.scroll_helper.update_scroll_position()
self.display_manager.image = self.scroll_helper.get_visible_portion()
self.display_manager.update_display()
if self.scroll_helper.is_scroll_complete():
# Mark as not scrolling when done
self.display_manager.set_scrolling_state(False)
```
@@ -223,11 +194,8 @@ def update(self):
sport_key = "nhl"
cache_key = f"{self.plugin_id}_{sport_key}_games"
# Uses sport-specific live_update_interval from config
cached = self.cache_manager.get_background_cached_data(
cache_key,
sport_key=sport_key
)
# get_background_cached_data() was removed in 3.8.0 — use get()
cached = self.cache_manager.get(cache_key, max_age=60)
if cached:
self.games = cached
@@ -254,9 +222,9 @@ def on_config_change(self, new_config):
---
## Font Management and Overrides
## Font Management
Use the Font Manager for advanced font handling and user customization.
The display manager's built-in fonts and text measurement. For fonts shipped with a plugin, see [FONT_MANAGER.md](FONT_MANAGER.md).
### Using Different Fonts
@@ -628,14 +596,12 @@ def update(self):
```python
def update(self):
# Check if another plugin is enabled
enabled_plugins = self.plugin_manager.get_enabled_plugins()
if "weather" in enabled_plugins:
# Weather plugin is available
weather_plugin = self.plugin_manager.get_plugin("weather")
if weather_plugin:
# Use weather data
pass
# get_enabled_plugins() was removed in 3.8.0 — check the instance's
# `enabled` flag instead
weather_plugin = self.plugin_manager.get_plugin("weather")
if weather_plugin is not None and weather_plugin.enabled:
# Use weather data
pass
```
### Sharing Data Between Plugins
+398
View File
@@ -0,0 +1,398 @@
# Architecture
A map of the codebase for a new contributor: which process does what, how
they talk to each other, and where to start reading for common changes.
## Processes
| systemd unit | Runs as | Runs | Installed by |
|---|---|---|---|
| `ledmatrix.service` | root | [`run.py`](../run.py) → `DisplayController` | [`install_service.sh`](../scripts/install/install_service.sh) |
| `ledmatrix-web.service` | the installing user | [`start_web_conditionally.py`](../scripts/utils/start_web_conditionally.py) → [`web_interface/start.py`](../web_interface/start.py) (Flask, port 5000) | `install_service.sh`, [`install_web_service.sh`](../scripts/install/install_web_service.sh) |
| `ledmatrix-update-verify.path` / `.service` | the web user | Health check after an automatic update | the same installers, or [`src/auto_update_setup.py`](../src/auto_update_setup.py) at runtime |
| `ledmatrix-wifi-monitor.service` | root | [`wifi_monitor_daemon.py`](../scripts/utils/wifi_monitor_daemon.py) | [`install_wifi_monitor.sh`](../scripts/install/install_wifi_monitor.sh) |
| `ledmatrix-mqtt-bridge.service` | root | [MQTT bridge](../integrations/mqtt_bridge/README.md) (optional) | [`install_mqtt_bridge.sh`](../scripts/install/install_mqtt_bridge.sh) |
| `ledmatrix-dns-fix.service` | root | DNS workaround (optional) | [`install_dns_fix.sh`](../scripts/install/install_dns_fix.sh) |
Unit templates are in [`systemd/`](../systemd/README.md). The display runs as
root because the LED matrix library needs direct GPIO access. The web
interface runs unprivileged and uses a fixed list of `sudo` rules for the
few privileged things it does; see [PERMISSIONS.md](PERMISSIONS.md).
`start_web_conditionally.py` exits without starting Flask when
`web_display_autostart` is explicitly false in `config.json`.
## How the two main processes share state
The display and the web interface are separate processes that never call
each other. They share three things:
1. **`config/config.json` and `config/config_secrets.json`.** The web
interface writes them through `ConfigManager`
([`src/config_manager.py`](../src/config_manager.py)); the display
notices through `ConfigService` (below).
2. **The disk cache**, `/var/cache/ledmatrix` (owned `root:ledmatrix`,
setgid, files `0660`), read and written through `CacheManager`
([`src/cache_manager.py`](../src/cache_manager.py),
[`src/cache/disk_cache.py`](../src/cache/disk_cache.py)). Readers in the
other process pass `memory_ttl=0` so they do not serve a stale in-memory
copy.
3. **A few files in `/tmp`.**
| State | Where | Written by | Read by |
|---|---|---|---|
| On-demand command | control socket `/run/ledmatrix/control.sock` ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)) | web: `start_on_demand_display()` / `stop_on_demand_display()` in [`api_v3/display.py`](../web_interface/blueprints/api_v3/display.py), via [`src/ipc/client.py`](../src/ipc/client.py) | display: [`src/ipc/server.py`](../src/ipc/server.py) acks; the render thread applies it in `_poll_on_demand_requests()` |
| On-demand request (fallback) | cache `display_on_demand_request` | web, when the socket fails; four plugins write it directly | display: `_poll_on_demand_requests()` |
| On-demand state | cache `display_on_demand_state` | display: `_publish_on_demand_state()` | web: `/api/v3/display/on-demand/status` |
| Current screen | cache `display_current_state` | display | web: `/api/v3/display/current-status` |
| Plugin errors | cache `plugin_error_snapshot` | display: `ErrorSnapshotPublisher` ([`src/error_aggregator.py`](../src/error_aggregator.py)) | web: `read_error_report()` for `/api/v3/errors/*` |
| Error clear | cache `plugin_error_clear_request` | web | display |
| Font usage | cache `font_usage_snapshot` | display: `FontUsagePublisher` ([`src/font_usage.py`](../src/font_usage.py)) | web: Fonts tab |
| Plugin health | cache `plugin_health:<id>` | display (web writes on reset) | web: `/api/v3/plugins/health` |
| Plugin runtime (loaded, state, last error, version) | cache `plugin_runtime_snapshot` | display: `PluginRuntimePublisher` ([`src/plugin_system/plugin_runtime.py`](../src/plugin_system/plugin_runtime.py)) | web: `read_plugin_runtime()` for `/api/v3/plugins/installed`, `/plugins/state`, reconciliation |
| Preview frame | `/tmp/led_matrix_preview.png` | display: `DisplayManager`, gated by [`snapshot_policy`](../src/common/snapshot_policy.py) | web: display SSE stream, `/api/v3/health` (file age) |
| Preview viewer marker | `/tmp/led_matrix_preview_viewer` | web, while a preview is open | display: writes full-rate snapshots only while it is fresh |
| Hardware init status | `/tmp/led_matrix_hw_status.json` | display | web: `/api/v3/hardware/status` |
| Render-loop heartbeat | `/run/ledmatrix/display-heartbeat.json` (tmpfs) | display: the render thread, via [`display_watchdog`](../src/display_watchdog.py) | web: `/api/v3/health` (`checks.display_loop`); the update health check |
The on-demand start route starts `ledmatrix.service` when it is not running
(`start_service`, on by default) but never restarts a running one. The routes
send the command over the display's control socket and get an ack; when that
fails (a stopped display, one older than the socket) they write the mailbox
instead, which the display reads every `ON_DEMAND_POLL_INTERVAL` (0.25s), from
its dwell sleep, its render loops and Vegas's interrupt check as well as the
main loop. Both ways end in the same handler, `_handle_on_demand_request()`.
The socket's handlers only queue; see [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)
for the protocol, the permission model and the plan to retire the mailboxes.
### Web and display processes: who runs plugins
Only the display process imports plugin code, instantiates plugins and calls
their lifecycle hooks (`update`, `display`, `on_config_change`, `on_enable`,
`on_disable`). The web process is metadata-only: it reads plugins as files
through `PluginCatalog`
([`src/plugin_system/plugin_catalog.py`](../src/plugin_system/plugin_catalog.py))
-- manifests, config schemas (through `SchemaManager`), each plugin's
section of `config.json`, and installed versions. The catalog keeps the
read-only method names of `PluginManager` and has nothing that can run a
plugin (no `load_plugin`, `get_plugin` or `plugins`).
How a web-side change reaches the running plugins:
| Change | How the display picks it up |
|---|---|
| Plugin settings saved, config reset | `ConfigService` sees the new `config.json` and calls the plugin's `on_config_change` with the prepared section |
| Plugin enabled or disabled | `ConfigService` → `_controller_config_change` flags a reconcile; `_reconcile_enabled_plugins` loads it (fresh from disk) or unloads it on the render thread |
| Plugin uninstalled (config removed) | the removed section flips its `enabled` flag, and the reconcile unloads it |
| Plugin installed, not enabled | nothing to do until it is enabled, which loads it |
| Plugin installed while already enabled, updated while enabled, or uninstalled with its config kept | **not picked up**: the display keeps running what it loaded. The route answers `restart_required: true` and the UI shows its restart banner |
`display_restart_required()` in `plugin_catalog.py` holds that last rule;
routes return it as `restart_required` (with the banner's wording in
`restart_message`), and `window.noteRestartRequired()` in
`static/v3/app.js` raises the banner for any response that carries it,
`POST /api/v3/config/main` included.
Runtime state shown in the UI comes from what the display publishes to the
shared cache: health and metrics (`/api/v3/plugins/health`,
`/plugins/metrics`), errors (`/api/v3/errors/*`), the current mode, and the
plugin runtime snapshot described below. `enabled` is read from
`config.json` by the display's rule (a missing flag is disabled).
Plugin code still runs in the web process in one place,
`_import_plugin_code_in_web_process()` in
[`api_v3/__init__.py`](../web_interface/blueprints/api_v3/__init__.py): the
Starlark routes import the starlark-apps plugin's `tronbyte_repository` and
`pixlet_renderer` helper modules (never the plugin class), and a web-UI
action with `oauth_flow` imports its script for `get_auth_url()`. Every
other web-UI action runs its script as a subprocess. A later, explicit
**plugin web-entry contract** -- a declared entry point for plugin web code
-- replaces that function.
Next stages: a **control socket** from the web process to the display
(reload one plugin, ask for its state) in place of `restart_required` and
the cache-key mailboxes, and the plugin web-entry contract above.
### Plugin state: desired, observed, and who owns it
There is one plugin state machine, and the display owns it:
`PluginStateManager` in
[`plugin_state.py`](../src/plugin_system/plugin_state.py) (unloaded →
loaded → enabled ⇄ running, error, disabled), held by the display's
`PluginManager`. It also records, per loaded plugin, the manifest version it
loaded and when. Nothing else keeps plugin state:
| Question | Answered by |
|---|---|
| Is it installed, at which version? | the plugins directory (`manifest.json`) |
| Should it run? | `config.json` (`<id>.enabled`, missing = disabled) |
| Has the user uninstalled it for good? | the store's uninstalled-plugins record |
| Is the display running it, at which version, and why not? | the display's runtime snapshot |
**The runtime snapshot.** `PluginRuntimePublisher`
([`plugin_runtime.py`](../src/plugin_system/plugin_runtime.py)), started by
`DisplayController` right after it creates the `PluginManager`, writes the
cache key `plugin_runtime_snapshot`: per plugin `loaded`, `state`, `error`
(type, a redacted message of at most 200 characters, when, recoverable),
`version` and `loaded_at`, plus `published_at`, `stale_after` and `running`.
The cache is on disk, usually the SD card, so it writes when something a
reader sees changes -- throttled to once per 10 s -- and otherwise once a
minute as a heartbeat. RUNNING, which every `update()` passes through, is
published as ENABLED, so plugin updates alone never cause a write.
`cleanup()` publishes `running: false`.
**Reading it.** `read_plugin_runtime()` judges the snapshot before anyone
uses it: `live` (fresh, from a running display), `stale` (older than
`stale_after`, 3 minutes: a hung or crashed display), `stopped` or
`unknown` (none, unreadable, or another schema). Only a live view reports
per-plugin facts; every other status answers `null` for them, so stale
truth cannot leak into a response. `/api/v3/plugins/installed` returns
`loaded`, `state`, `error_info`, `loaded_version` and `loaded_at` per
plugin and `data.runtime` (`status`, `published_at`, `age_seconds`);
`/api/v3/plugins/state` returns the same beside the desired state.
**Reconciliation**
([`state_reconciliation.py`](../src/plugin_system/state_reconciliation.py))
compares desired state (config + disk) with observed state (the snapshot).
It fixes desired-state gaps -- a plugin on disk with no config section gets
`{"enabled": false}`, a configured plugin missing from disk is reinstalled
unless the user uninstalled it -- and only reports observed-state gaps
(enabled but not loaded, loaded at an older version): the display loads and
unloads by config on its own, and a version gap needs a restart.
**`data/plugin_state.json` is retired.** The web process used to keep a
second `PluginStateManager` (`state_manager.py`) persisted to that file:
per plugin an enabled flag copied from config, a version copied from the
manifest (when set at all), a status derived from those, and install/update
timestamps. Reconciliation mostly synced it back to config and backups
merged it into their plugin list. Every field is derivable (the timestamps
from the operation history), so nothing is migrated: no code reads or
writes the file, and a copy left on a device is inert and safe to delete.
The two classes shared a name but not a concern -- a persisted install
record versus the live lifecycle -- so they were not merged; the persisted
one had nothing left to hold and was removed.
## Display loop
[`src/display_controller.py`](../src/display_controller.py), class
`DisplayController`. `__init__` loads config, starts the cache and the
error-snapshot publisher, runs the startup validator, creates the
`DisplayManager` ([`src/display_manager.py`](../src/display_manager.py)),
`FontManager` and `PluginManager`, loads the enabled plugins in parallel,
runs an initial `update()` pass within a 20-second budget
(`_INITIAL_UPDATE_BUDGET_SECONDS`; a plugin that misses it is deferred to
the scheduler), and sets up Vegas mode.
`run()` is the main loop. Each pass, in order: apply a pending plugin
enable/disable, poll on-demand requests, run scheduled plugin updates, check
the on/off schedule and brightness, then show one screen. Priority is
on-demand, then WiFi status messages, then live priority, then Vegas mode,
then normal rotation.
- **Rotation.** `available_modes` is the ordered list of display modes;
`current_mode_index` advances after each screen.
`_apply_plugin_rotation_order()` applies `display.plugin_rotation_order`.
- **Durations.** `_get_display_duration()`: `display.display_durations[mode]`,
else the plugin's `get_display_duration()`, else 30 s. Plugins that
support dynamic duration run until `is_cycle_complete()`, capped by
`display.dynamic_duration.max_duration_seconds` (default 180 s).
- **On-demand.** A request from the web interface pins one plugin (or mode)
for a duration. `_activate_on_demand()` / `_clear_on_demand()`; the
session is saved under `display_on_demand_config` so it survives a
restart. It also keeps the display on during scheduled off hours. A
request for a plugin that is disabled in config loads it live
(`_load_plugin_for_on_demand()`, `load_plugin(force_enabled=True)`)
without writing `config.json`; the main loop unloads it once on-demand
moves off it (`_release_on_demand_plugins()`).
- **Live priority.** `_check_live_priority()` looks for a plugin whose
`has_live_priority()` and `has_live_content()` are both true and switches
to it, rotating between several live games.
- **Schedule and dim schedule.** `_check_schedule()` reads `schedule`;
`_check_dim_schedule()` reads `dim_schedule` and
`display.hardware.brightness`. Both are re-evaluated once a minute.
- **Long screens.** While a screen is showing (a dwell, a scroll, a Vegas
iteration), `_service_pending_changes()` repeats the on-demand, schedule
and brightness checks every 0.25 s, so a change does not wait for the
screen to end.
- **Config hot reload.** `ConfigService`
([`src/config_service.py`](../src/config_service.py)) polls the config and
secrets files' mtimes every 2 s and notifies subscribers when the content
changes. The controller refreshes its cached settings; enabling or
disabling a plugin queues `_reconcile_enabled_plugins()`, which loads or
unloads it on the display thread; each plugin gets `on_config_change()`
for its own section, under its plugin lock
(`PluginManager.apply_config_change()`). Set `LEDMATRIX_HOT_RELOAD=false` to turn this off.
Matrix hardware settings are only read at start-up.
- **Vegas mode.** [`src/vegas_mode/`](../src/vegas_mode/): the display loop
calls `VegasModeCoordinator.run_iteration()`
([`coordinator.py`](../src/vegas_mode/coordinator.py)) when
`display.vegas_scroll.enabled` is set. `PluginAdapter` gets each plugin's
content (`get_vegas_content()`, else its `scroll_helper` image, else a
capture of `display()`), `StreamManager` orders it and `RenderPipeline`
scrolls it. See [ADVANCED_FEATURES.md](ADVANCED_FEATURES.md).
- **Multi-display sync.** `DisplaySyncManager`
([`src/common/sync_manager.py`](../src/common/sync_manager.py)), enabled by
`sync.role`: a leader sends a follower its share of each frame over UDP
(port 5765).
### Liveness
A render thread stuck inside a plugin leaves the service "active" and the
panel frozen, so liveness is reported by the render thread itself
([`src/display_watchdog.py`](../src/display_watchdog.py), standard library
only). `beat()` from any other thread is ignored: the update worker, Vegas's
tick thread and the prefetcher keep running while the render thread is stuck,
and must not vouch for it.
- **Check-in points.** The top of `run()`'s loop (`loop_pass()`), every
dwell second (`_sleep_with_plugin_updates`), every frame of the per-screen
loops (`_display_once`), every frame of Vegas's own loop and static pause
(`coordinator.run_iteration`), each plugin fetched for a Vegas cycle
(`StreamManager._fetch_plugin_content`), each update on the
`synchronous_updates` path, and every frame pushed
(`DisplayManager.update_display` -> `note_frame()`). Beats are
rate-limited to one ping and one heartbeat write every 5 s.
- **systemd watchdog.** `ledmatrix.service` is `Type=simple` with
`WatchdogSec=120` and `NotifyAccess=main`. `run.py` sends
`WATCHDOG_USEC` = 15 minutes before importing anything heavy (start-up loads
plugins and runs the 20 s update budget, and the watchdog clock starts with
the process). After the first frame -- or the first full pass, when there is
nothing to draw -- the loop sends `READY=1`, restores the unit's 120 s and
pings. `PluginManager.load_plugin()` on the render thread (a plugin enabled
from the web UI, or loaded for on-demand) gets 15 minutes again, since it
can run pip. A missed deadline is a SIGABRT; faulthandler, enabled on
arming, dumps every thread's stack to the journal.
- **Heartbeat.** `/run/ledmatrix/display-heartbeat.json`
(`{"pid", "mono", "wall"}`; `RuntimeDirectory=ledmatrix`, 0755, file 0644 so
the web user can read it). Readers compare `mono` with their own
`time.monotonic()` -- CLOCK_MONOTONIC is shared by every process and does not
jump when NTP first sets an RTC-less Pi's clock. `/api/v3/health` calls it
`stalled` past 60 s; no file is `not_reported` and changes nothing. A clean
stop removes it. Without `RuntimeDirectory=` (an older unit) the display,
as root, creates the directory itself; off Linux, or without root, there
is no heartbeat.
## Plugin system
[`src/plugin_system/`](../src/plugin_system/):
| Area | Where |
|---|---|
| Base class plugins implement | [`base_plugin.py`](../src/plugin_system/base_plugin.py) (`BasePlugin`, `VegasDisplayMode`) |
| Finding a plugin's directory | [`plugin_dirs.py`](../src/plugin_system/plugin_dirs.py): manifest `id` first, then directory `<id>` or `ledmatrix-<id>` |
| Discovery, load, unload, scheduled updates (display process) | [`plugin_manager.py`](../src/plugin_system/plugin_manager.py) (`PluginManager`) |
| Manifest, schema, config and version reads (web process) | [`plugin_catalog.py`](../src/plugin_system/plugin_catalog.py) (`PluginCatalog`; see [who runs plugins](#web-and-display-processes-who-runs-plugins)) |
| Import and instantiate | [`plugin_loader.py`](../src/plugin_system/plugin_loader.py) (`PluginLoader.load_plugin()`: dependencies, module, class) |
| Timeouts | [`plugin_executor.py`](../src/plugin_system/plugin_executor.py) (`PluginExecutor`, 30 s default; a timed-out thread is abandoned, not killed) |
| Circuit breaker | [`plugin_health.py`](../src/plugin_system/plugin_health.py) (`PluginHealthTracker`: 3 consecutive failures open the circuit for 300 s) |
| Resource metrics | [`resource_monitor.py`](../src/plugin_system/resource_monitor.py) |
| Config schemas and defaults | [`schema_manager.py`](../src/plugin_system/schema_manager.py) |
| Install, update, uninstall | [`store_manager.py`](../src/plugin_system/store_manager.py) (`PluginStoreManager`), with its methods split across [`store_registry.py`](../src/plugin_system/store_registry.py) (registry, GitHub), [`store_install.py`](../src/plugin_system/store_install.py) and [`store_update.py`](../src/plugin_system/store_update.py) |
| Core-version gate | [`compatibility.py`](../src/plugin_system/compatibility.py) |
Discovery scans only `plugin_system.plugins_directory` (default
`plugin-repos/`). Scheduled `update()` calls run on one background worker
thread; a per-plugin lock keeps `display()` from running during an update.
**Store flow.** `install_plugin()` renames any existing copy aside
(`<id>.standalone-backup-preinstall`), installs the new one, and puts the old
copy back if the install fails. Monorepo plugins come from the GitHub Trees
API, falling back to the repository ZIP; other plugins by `git clone` or
download. The manifest is checked (see
[required fields](PLUGIN_API_REFERENCE.md#manifest-required-fields)), the core
version gate runs, then dependencies are installed as root through
`scripts/fix_perms/safe_pip_install.sh`. `update_plugin()` pulls git
installs, undoing a pull whose new version is incompatible, and reinstalls
everything else through `_reinstall_with_rollback()`.
## Web interface
- **App.** [`web_interface/app.py`](../web_interface/app.py) builds the
Flask `app` at import time, creates the managers -- a `PluginCatalog`,
never a `PluginManager` -- and registers two blueprints.
`web_interface/start.py` runs it on port 5000.
- **Pages.** [`blueprints/pages_v3.py`](../web_interface/blueprints/pages_v3.py)
serves the shell `templates/v3/base.html` at `/` and each tab as a
partial at `/partials/<name>` (templates in
`web_interface/templates/v3/partials/`). Plugin configuration tabs are
rendered from the plugin's schema by `plugin_config.html`.
- **API.** [`blueprints/api_v3/`](../web_interface/blueprints/api_v3/) is one
blueprint at `/api/v3`, split by area: `backup.py`, `config.py`,
`display.py`, `fonts.py`, `misc.py` (health, logs, errors, cache, sync),
`starlark.py`, `system.py` (service actions, updates, git), `wifi.py`, and
the plugin routes: `plugins.py` (installed list, enable/disable, plugin
actions), `plugin_store.py` (install, update, uninstall, store),
`plugin_config.py` (config, schema, reset), `plugin_assets.py` (uploads,
plugin static files), `plugin_health.py` (health, metrics, limits),
`plugin_operations.py` (operation history, state reconciliation) and
`plugin_calendar.py`. `__init__.py` defines the blueprint and shared helpers and
imports the modules so their routes register. Endpoints are listed in
[REST_API_REFERENCE.md](REST_API_REFERENCE.md).
- **Front end.** HTMX loads each tab's partial on first open
(`hx-trigger="loadtab"`); Alpine.js holds page state. Scripts are in
`web_interface/static/v3/js/`; form widgets are bundled from
[`js/widgets/`](../web_interface/static/v3/js/widgets/README.md).
- **Server-sent events** (`app.py`): `/api/v3/stream/stats` (CPU, memory,
temperature, service state, every 10 s), `/api/v3/stream/display` (preview
frames when the PNG changes) and `/api/v3/stream/logs` (journal of both
services). One generator thread per stream is shared by all clients.
## Updates
- **Update Code** on the Overview tab and the automatic updater both call
`perform_core_update()` in
[`api_v3/system.py`](../web_interface/blueprints/api_v3/system.py):
fetch branches and tags, move the checkout for the update channel, reinstall
changed requirement files, report whether a restart is needed.
- **Update channels** (`auto_update.channel`):
[`web_interface/update_channel.py`](../web_interface/update_channel.py)
decides the move. `stable` checks out the newest `vX.Y.Z` tag (detached
HEAD) when it contains the current commit; `beta` is
`git pull --rebase --autostash` on the current branch, and leaves a
detached release for `main` first. A stable device newer than the newest
release keeps pulling `main` until a release contains its commit, so no
update ever moves backwards; a config without the key is written as
`stable` once the device reaches a release. Checkouts carry uncommitted
edits across with `git stash create`/`apply`, and keep them in the stash
list if they no longer apply.
- **Automatic updates** (`auto_update.enabled`, off by default):
`AutoUpdater` in [`web_interface/auto_update.py`](../web_interface/auto_update.py)
runs in the web process, checks every 30 minutes, and updates at most
weekly between 02:00 and 05:00. Before pulling it copies
[`scripts/utils/auto_update_verify.py`](../scripts/utils/auto_update_verify.py)
to `data/auto_update_verifier.py`, then writes
`data/auto_update_verify.request`. That file triggers
`ledmatrix-update-verify.path`, which runs the verifier as a separate unit
(so restarting the web service does not kill it). The verifier restarts
both services, waits for the web API to answer and the display service to
stay up -- and, when the display wrote a heartbeat before the update, to
keep one fresh from the restarted process (see Liveness) -- and on failure
returns to where HEAD was (the branch, or detached on the previous
release; `old_ref` in the pending file), resets to the previous commit
and restarts again.
Plugin updates run only after a verified core update. State is in
`data/auto_update_state.json` and `data/auto_update_pending.json`.
- **Startup validator.** `StartupValidator`
([`src/startup_validator.py`](../src/startup_validator.py)) runs twice in
`DisplayController.__init__`: config and cache directory first, then
enabled plugins once the plugin manager exists. It also warns when an
installed systemd unit differs from its template in `systemd/`. Results
are logged; startup continues either way. Nothing rewrites installed units
on update: a unit change such as the watchdog reaches an existing install
only when `install_service.sh` is re-run.
## Where to start reading
| Task | Start with |
|---|---|
| Change rotation, durations or priorities | `DisplayController.run()` and `_get_display_duration()` in [`display_controller.py`](../src/display_controller.py) |
| Add a config key | [CONFIG_REFERENCE.md](CONFIG_REFERENCE.md), [`config/config.template.json`](../config/config.template.json), the tab's partial and `api_v3/config.py` |
| Change drawing or fonts | [`display_manager.py`](../src/display_manager.py), [`font_manager.py`](../src/font_manager.py), [`src/common/bdf_font.py`](../src/common/bdf_font.py) |
| Add a plugin-facing API | [`base_plugin.py`](../src/plugin_system/base_plugin.py) or [`src/common/`](../src/common/README.md); document it in [PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md) |
| Plugin install/update bugs | `PluginStoreManager` in [`store_manager.py`](../src/plugin_system/store_manager.py) |
| A plugin that won't load | `PluginManager.load_plugin()` and `PluginLoader.load_plugin()`; `python3 scripts/check_plugin.py --plugin <id>` |
| Add an API endpoint | the matching module in [`api_v3/`](../web_interface/blueprints/api_v3/) |
| Add a web UI tab or control | `templates/v3/base.html`, the tab's partial, `pages_v3.py` |
| Vegas scroll | [`src/vegas_mode/coordinator.py`](../src/vegas_mode/coordinator.py) |
| Installer or permissions | [`first_time_install.sh`](../first_time_install.sh), [`scripts/install/`](../scripts/install/), [PERMISSIONS.md](PERMISSIONS.md) |
| Work without a Pi | [DEV_PREVIEW.md](DEV_PREVIEW.md), [EMULATOR_SETUP_GUIDE.md](EMULATOR_SETUP_GUIDE.md), [HOW_TO_RUN_TESTS.md](HOW_TO_RUN_TESTS.md) |
+31 -13
View File
@@ -172,10 +172,14 @@ ERROR - Plugin football-scoreboard configuration validation failed: 'api_key' is
### Enable Debug Logging
Set environment variable:
Run the display in the foreground with `-d`, or set `LEDMATRIX_DEBUG=true`
(the value must be `true`; `1` is ignored — see `setup_logging()` in
[`src/logging_config.py`](../src/logging_config.py)):
```bash
export LEDMATRIX_DEBUG=1
python run.py
sudo systemctl stop ledmatrix.service
sudo python3 run.py -d
# or
sudo LEDMATRIX_DEBUG=true python3 run.py
```
### Check Merged Configuration
@@ -250,14 +254,21 @@ WARNING - Plugin ID 'Football-Scoreboard' may conflict with 'football-scoreboard
## Checking Configuration via API
The API blueprint mounts at `/api/v3` (`web_interface/app.py:144`).
The API blueprint (`web_interface/blueprints/api_v3/`) is registered at
`/api/v3` in `web_interface/app.py`.
```bash
# Get full main config (includes all plugin sections)
# Get full main config (includes all plugin sections; credential-named
# fields are blanked in the response)
curl http://localhost:5000/api/v3/config/main
# Save updated main config
# Change some settings: only the keys you send are changed
curl -X POST http://localhost:5000/api/v3/config/main \
-H "Content-Type: application/json" \
-d '{"timezone": "America/Chicago", "brightness": 80}'
# Replace config.json wholesale (advanced)
curl -X POST http://localhost:5000/api/v3/config/raw/main \
-H "Content-Type: application/json" \
-d @new-config.json
@@ -269,8 +280,10 @@ curl "http://localhost:5000/api/v3/plugins/config?plugin_id=football-scoreboard"
```
> There is no dedicated `/config/plugin/<id>` or `/config/validate`
> endpoint — config validation runs server-side automatically when you
> POST to `/config/main` or `/plugins/config`. See
> endpoint. `POST /plugins/config` validates against the plugin's schema
> and rejects an invalid config with `400`; `POST /config/main` checks the
> individual fields it knows (display hardware values, durations, Vegas
> and sync settings). See
> [REST_API_REFERENCE.md](REST_API_REFERENCE.md) for the full list.
## Backup and Recovery
@@ -283,9 +296,12 @@ cp config/config.json config/config.backup.json
### Automatic Backups
LEDMatrix creates backups before saves:
LEDMatrix creates backups before saves (`src/config_manager_atomic.py`):
- Location: `config/backups/`
- Format: `config_YYYYMMDD_HHMMSS.json`
- Format: `config.json.backup.YYYYMMDD_HHMMSS_ffffff` (microseconds last),
plus a matching `config_secrets.json.backup.<timestamp>` when a secrets
file exists
- The five most recent are kept
### Recovery
@@ -294,7 +310,7 @@ LEDMatrix creates backups before saves:
ls -la config/backups/
# Restore from backup
cp config/backups/config_20240115_120000.json config/config.json
cp config/backups/config.json.backup.20240115_120000_000000 config/config.json
```
## Troubleshooting Checklist
@@ -309,8 +325,10 @@ cp config/backups/config_20240115_120000.json config/config.json
## Getting Help
1. Check logs: `tail -f logs/ledmatrix.log`
2. Enable debug: `LEDMATRIX_DEBUG=1`
1. Check logs. Both services log to journald, not to a file:
`sudo journalctl -u ledmatrix.service -f` (display) and
`sudo journalctl -u ledmatrix-web.service -f` (web interface)
2. Enable debug: `LEDMATRIX_DEBUG=true` or `python3 run.py -d`
3. Check error dashboard: `/api/v3/errors/summary`
4. Validate JSON: https://jsonlint.com/
5. File an issue: https://github.com/ChuckBuilds/LEDMatrix/issues
+189
View File
@@ -0,0 +1,189 @@
# Configuration Reference
Every key in `config/config.json`, what it does, its default, and where the
code reads it. The file is created from `config/config.template.json` on
first run, and `ConfigManager._migrate_config()` merges any template keys
added by later releases into your existing config (your values are never
overwritten). Secrets live in `config/config_secrets.json` and are merged
into the config at load time.
Most settings are editable from the web interface; this page documents the
underlying keys for people editing `config.json` directly or writing
tooling against it.
## Top level
| Key | Type / default | Meaning | Read by |
|---|---|---|---|
| `web_display_autostart` | bool, `true` | Whether the web interface service starts with the system | `scripts/utils/start_web_conditionally.py` |
| `auto_update.enabled` | bool, `false` | Weekly automatic updates: LEDMatrix code first (health-checked, rolled back on failure), then installed plugins. Toggle in the General tab or install with `first_time_install.sh --enable-auto-update` | `web_interface/auto_update.py`, `src/auto_update_setup.py` (`is_enabled()`) |
| `auto_update.channel` | `"stable"` or `"beta"`, `"stable"` (template) | What Update Code and the weekly update install. `stable`: the newest `vX.Y.Z` release tag (pre-releases ignored), checked out with a detached HEAD. `beta`: `main`. Never moves a device backwards: one newer than the newest release keeps following `main` until a release contains its commit. Missing (configs from before channels) behaves like `stable` and is saved as `stable` once the device is on a release. General tab, Update Channel | `web_interface/update_channel.py` (`resolve()`) |
| `timezone` | string, `"America/New_York"` | IANA timezone for schedules and displays | `ConfigManager.get_timezone()` |
| `target_fps` | int, `100` | Legacy "Scroll Frame Rate". Core scrolling no longer reads it: scroll frames are presented at `display.hardware.limit_refresh_rate_hz` divided by each scroll's frame hold, and speed comes from each plugin's scroll settings. Still exposed to plugins via `BasePlugin.global_config` | `src/plugin_system/base_plugin.py` |
| `location` | object | `city` / `state` / `country`. Supplies the **default** for a plugin's own `location_city` / `location_state` / `location_country` setting, so weather, radar and friends follow this device without being configured twice. A value saved on the plugin itself still overrides it. Starlark (Tidbyt) apps get the same treatment: a `Location` field left blank on the app renders at this city (geocoded once via Open-Meteo, coordinates cached permanently) instead of the app author's default, which is usually San Francisco. If the city can't be looked up (no match, or the geocoder is unreachable; retried after 30 minutes), the app keeps its own default. | `SchemaManager.apply_device_location()`, then plugins via merged config; `src/device_location.py` for Starlark apps |
## `schedule` — display on/off hours
| Key | Type / default | Meaning |
|---|---|---|
| `enabled` | bool, `false` | Master switch for scheduled display on/off |
| `mode` | `"global"` or `"per-day"`, template uses `"per-day"` | Whether one time range applies to all days or each day has its own |
| `start_time` / `end_time` | `"HH:MM"`, `07:00`–`23:00` | Global-mode on/off times |
| `days.<weekday>.{enabled,start_time,end_time}` | per-day objects | Per-day-mode overrides |
Read by `DisplayController._check_schedule()` (`src/display_controller.py`).
Managed in the web UI under Schedule.
## `dim_schedule` — scheduled brightness dimming
Same shape as `schedule` (the template sets its `mode` to `"global"`), plus:
| Key | Type / default | Meaning |
|---|---|---|
| `dim_brightness` | int, `30` | Brightness percentage applied while the dim window is active |
Read by `DisplayController._check_dim_schedule()` (`src/display_controller.py`;
saved via `POST /api/v3/config/dim-schedule`). The display returns to
`display.hardware.brightness` outside the window.
## `display.hardware` — matrix panel hardware
All keys map to the corresponding `rpi-rgb-led-matrix` options and are read
in `DisplayManager._setup_matrix` (`src/display_manager.py`). Defaults are the
`config/config.template.json` values: `ConfigManager` adds any key missing from
`config.json` from the template on load, so `DisplayManager`'s own fallbacks
don't apply on a normal install.
The ranges are what the pinned rgbmatrix library and its Python binding accept
(`src/matrix_support.py`). The config API refuses anything else; a value
hand-edited into `config.json` makes the display log the setting and run in
fallback mode instead of starting the matrix.
| Key | Type / default |
|---|---|
| `rows` / `cols` | int, `32` / `64` — rows: even, 8–64; cols: at least 16 |
| `chain_length` | int, `2` — 1–255 (the Python binding stores it in one byte) |
| `parallel` | int, `1` — 1–3, and no more than `hardware_mapping` has outputs (`regular`, `classic`: 3; the others: 1) |
| `brightness` | int, `90` — 1–100 |
| `hardware_mapping` | string, `"adafruit-hat"` — `"adafruit-hat-pwm"`, `"adafruit-hat"`, `"regular"`, `"regular-pi1"`, `"classic"` or `"classic-pi1"` (case-insensitive; `compute-module` isn't in the installed build). A Pi 5 doesn't support `"classic-pi1"` |
| `scan_mode` | int, `0` — `0` progressive, `1` interlaced |
| `pwm_bits` | int, `9` — 1–11 |
| `pwm_dither_bits` | int, `1` — 0–2 |
| `pwm_lsb_nanoseconds` | int, `130` — 50–3000 |
| `disable_hardware_pulsing` | bool, `false` — `true` times brightness pulses in software (less exact); hardware pulsing needs the OE line on GPIO 18 and the Pi's onboard sound driver off |
| `inverse_colors` | bool, `false` |
| `show_refresh_rate` | bool, `false` — prints the refresh rate to stdout; draws nothing on the panel |
| `led_rgb_sequence` | string, `"RGB"` — `"RGB"`, `"RBG"`, `"GRB"`, `"GBR"`, `"BRG"` or `"BGR"` |
| `limit_refresh_rate_hz` | int, `100` — `0` = no cap; scroll timing assumes 100 Hz when `0` |
| `pixel_mapper_config` | string, `""` — e.g. `"U-mapper"` / `"Rotate:90"`; mappers that rotate or fold the chain change the display size plugins and the web preview see |
| `orientation` | string, `"normal"` — `"180"` rotates the rendered image 180° for panels physically mounted upside down (e.g. to move the Pi/wiring to a more convenient side); `"90"` / `"270"` for a panel on its side, swapping width and height; composed onto `pixel_mapper_config` as a trailing `Rotate:<degrees>` mapper, so it stays independent of any custom `pixel_mapper_config` value |
| `row_address_type` | int, `0` — non-standard panel row addressing: `1` AB, `2` direct row select, `3` ABC, `4` ABC shift + DE direct, `5` SM5368 / B707 row shift register (e.g. Waveshare 96x48 V2, with `led_rgb_sequence` `"BGR"`). On a Pi 5 the library supports only `0` and `2`, and LEDMatrix enforces that (`src/pi5_matrix_support.py`) |
| `multiplexing` | int, `0` — 0–22, pixel wiring scheme for outdoor/specialty panels (names listed in the README) |
| `panel_type` | string, `""` — set to `"FM6126A"` or `"FM6127"` for panels needing init; FM6124 / FM6124D / FM6124DJ panels need none, so leave it `""` |
## `display.runtime`
| Key | Type / default | Meaning |
|---|---|---|
| `gpio_slowdown` | int, `3` | GPIO timing slowdown for faster Pis (0–10). On a Pi 5 in PIO mode start at `1` (`0` acts as `1`) and raise it if the image flickers or shows garbage. Panels on `row_address_type` `5` (SM5368 row drivers) can need 6–8 on a Pi 4 — lower values make rows jump |
| `rp1_rio` | int, `0` | Pi 5 only: `0` = PIO (less CPU), `1` = RIO (higher refresh; `gpio_slowdown` effect inverted). Applied only if the installed matrix library supports it |
## `display.double_sided`
Drives `_LogicalMatrix` in `src/display_manager.py` — renders the same
logical image to multiple chained physical panels.
| Key | Type / default | Meaning |
|---|---|---|
| `enabled` | bool, `false` | Mirror output across panel copies |
| `copies` | int, `2` | Number of physical copies in the chain |
| `axis` | `"horizontal"`, default | Axis along which panels are chained |
## `display` — other keys
| Key | Type / default | Meaning | Read by |
|---|---|---|---|
| `display_durations` | object, `{}` | Per-plugin display duration in seconds, keyed by plugin id (e.g. `"clock": 15`) | `DisplayController._get_display_duration()` (`src/display_controller.py`) |
| `plugin_rotation_order` | array, `[]` | Explicit rotation order of plugin ids; empty = all enabled plugins in discovery order | `DisplayController._apply_plugin_rotation_order()` (`src/display_controller.py`) |
| `use_short_date_format` | bool, `true` | Compact date rendering in sports scoreboards | Nothing since `src/base_classes` was removed; scoreboards read `display.use_short_date_format` from their own plugin config |
| `scan_order_compensation` | string, `"auto"` | `"auto"` shows one half of each panel a refresh behind while something scrolls at one frame per refresh, which removes the 1px step a 1:N-scan panel shows across its middle; `"off"` disables it. Applies only to layouts whose row order is known: plain or parallel chains, 0 or 180 degree orientation, `multiplexing` 0, `scan_mode` 0, and not in the emulator | `DisplayManager._setup_scan_order_compensation()` (`src/display_manager.py`, `src/scan_order.py`) |
| `dynamic_duration.max_duration_seconds` | int, optional | Cap for plugins that request dynamic display time | `DisplayController._get_global_dynamic_cap()` (`src/display_controller.py`) |
## `display.vegas_scroll` — continuous scroll mode
Read by `src/vegas_mode/config.py` (`VegasScrollConfig.from_config`). See
[ADVANCED_FEATURES.md](ADVANCED_FEATURES.md) for behavior details, including
[live content in the ticker](ADVANCED_FEATURES.md#live-content-in-the-ticker).
| Key | Type / default |
|---|---|
| `enabled` | bool, `false` |
| `scroll_speed` | int, `50` (px/s) |
| `separator_width` | int, `32` |
| `plugin_order` | array, `[]` |
| `excluded_plugins` | array, `[]` |
| `target_fps` | int, `125` |
| `buffer_ahead` | int, `2` |
| `intra_plugin_gap` | int, `8` |
| `render_width_pct` | int, `100` |
| `min_content_separation` | int, `24` |
| `min_cut_gap` | int, `6` |
| `continuous_scroll` | bool, `true` |
| `offscreen_prefetch` | bool, `true` — render every plugin's ticker content on the background thread, each on its own canvas. `false` restores handing canvas-bound plugins to the render thread, one pause at a time. Temporary; see [OFFSCREEN_RENDERING.md](OFFSCREEN_RENDERING.md) |
| `prefetch_gate` | bool, `true` — let that background thread run Python only while the render thread is waiting for the panel, so the render thread never waits for the GIL when a refresh comes round. Only takes effect with the rebuilt rgbmatrix binding (`scripts/build_rgbmatrix_nogil.sh`). See [OFFSCREEN_RENDERING.md](OFFSCREEN_RENDERING.md) |
| `switch_interval_ms` | float, `0` — experimental: shorten Python's GIL switch interval to this many ms while Vegas runs. `0` leaves the default (5 ms) alone |
| `live_refresh` | bool, `true` — live elements: a plugin that supports them (scores, the flight map) has what is already scrolling updated when its data changes, instead of freezing each card as it was drawn. Always off under multi-display sync, in swap mode and with `offscreen_prefetch` off. `false` restores the frozen behaviour exactly. Per plugin: `vegas_live` in the plugin's section |
| `live_max_hz` | float, `5` (0–10) — ceiling on how often an animated live element (a moving aircraft) is redrawn; `0` keeps data updates and turns animation off. Capped at 1 Hz without the rebuilt rgbmatrix binding |
| `live_min_interval` | float, `2` (0.5–60) — shortest time between two data redraws of one plugin; a faster plugin is redrawn at this rate, never skipped |
| `live_lead_screens` | float, `1` (0–5) — how far ahead of the screen, in screen widths, an animated element starts being redrawn |
| `smooth_scroll` | bool, `true` — move a whole number of pixels per panel refresh, locked to vsync. `scroll_speed` is snapped to the nearest speed the panel can show that way (at 95Hz: 95, 47.5, 31.7 px/s…), measured against the panel's real refresh rate once scrolling starts |
| `sub_pixel_blend` | bool, `false` — the older smoothing: advance by elapsed time and blend neighbouring pixel columns. Looks anti-aliased in the web preview but shimmers on the panel and is not locked to the refresh. Overrides `smooth_scroll` when on |
| `extend_threshold_screens` | float, `2.0` |
| `auto_trim` | bool, `true` |
| `trim_threshold` | int, `10` |
| `content_padding` | int, `8` |
| `min_plugin_width` | int, `8` |
| `lead_in_width` | int, `0` |
| `plugins_per_cycle` | int, `6` |
| `max_plugin_width_ratio` | float, `0.0` |
| `overflow_mode` | string, `"rotate"` |
| `dynamic_duration_enabled` | bool, `true` |
| `min_cycle_duration` | int, `60` |
| `max_cycle_duration` | int, `240` |
| `frame_based_scrolling` | bool, `true` — does not step or set a frame rate; motion is by elapsed time either way. When `true`, `scroll_speed` passes through a clamp of 0.1–5 px per `scroll_delay` (see next row) |
| `scroll_delay` | float, `0.02` — not a frame period. Only used with `frame_based_scrolling`: the applied speed is `clamp(scroll_speed × scroll_delay, 0.1, 5) / scroll_delay` px/s, so at `0.02` speeds under 5 px/s run at 5, and at `0.001` nothing runs slower than 100 px/s |
| `live_in_ticker` | bool, `true` — keep scrolling during live games instead of handing the display to a full-screen scoreboard. `false` was the default before 3.8.0; the first start on 3.8.0 turns a stored `false` on once and sets `live_in_ticker_migrated` |
| `live_weight` | int, `3` (1–10) — slots per cycle for a plugin with live content |
| `favorite_live_weight` | int, `5` (1–10) — slots per cycle when a plugin reports a favorite team is live |
## `sync` — multi-display synchronization
Read by `src/common/sync_manager.py` and `src/display_controller.py`.
| Key | Type / default | Meaning |
|---|---|---|
| `role` | `"standalone"` (default), `"leader"`, or `"follower"` | This device's role in a synced pair |
| `port` | int, `5765` | TCP port used for sync traffic |
| `follower_position` | `"left"` (default) or `"right"` | Which half of the combined image this follower renders (`src/display_controller.py`) |
## `plugin_system`
| Key | Type / default | Meaning |
|---|---|---|
| `plugins_directory` | string, `"plugin-repos"` | Where the Plugin Store installs plugins and the only directory the plugin loader scans. Read by `PluginManager` and `PluginStoreManager` (`src/plugin_system/`); editable under General settings |
| `auto_discover`, `auto_load_enabled`, `development_mode` | bool | **Unused.** Legacy keys, read by nothing and no longer in the template; older configs may still carry them. Plugins are always discovered, and every plugin with `enabled: true` is loaded — to keep a plugin installed but dormant, set its own `enabled` to `false`. Not shown in the web UI; may be left in or removed from config.json |
## Plugin config blocks
Every installed plugin stores its settings under a top-level key equal to
its plugin id (the template ships one for the bundled `web-ui-info`
plugin). The shape of each block is defined by that plugin's
`config_schema.json`; common keys are `enabled` and `display_duration`.
See [PLUGIN_CONFIG_CORE_PROPERTIES.md](PLUGIN_CONFIG_CORE_PROPERTIES.md).
## `config/config_secrets.json`
| Key | Meaning |
|---|---|
| `github.api_token` | Optional GitHub token the Plugin Store uses to avoid API rate limits (`src/plugin_system/store_registry.py`) |
| `<plugin-id>.*` | Secrets a plugin declares with `"x-secret": true` in its config schema; merged into that plugin's config at load time |
+175
View File
@@ -0,0 +1,175 @@
# Deprecated plugin APIs: usage scan
Generated by `scripts/plugin_api_usage.py` — do not edit by hand; re-run it (see [How to re-run](#how-to-re-run)).
- Scanned: 2026-10-01, core 3.7.0
- Monorepo: [ChuckBuilds/ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins) (main @ 4de1d134), 46 plugins
- Third-party plugins: 8 with their own repo in `plugins.json` (f1-live, gif-player, pga-tour-leaderboard, plex-marquee, ledmatrix-dresden-departures, tidbyt-baseball-scoreboard, sleeper-fantasy, ledmatrix-nascar)
**37 deprecated methods: 36 unused, 1 still used, 0 need review.**
Counted per plugin: a *call* is `<receiver>.method` on an object named like the owner (`cache_manager`, `display_manager`, `font_manager`, `plugin_manager`), or on `self`/`super()` in a subclass; an *override* is `def method` in a subclass of the owner. *Review* hits are `.method` on a receiver whose type the scan cannot tell. *Internal* hits sit inside another deprecated core method and go with it. *Unrelated* hits are a different class's own method with the same name (a name collision), and never block removal; neither do hits in test files.
| Method | Removal | Core | Plugins (calls / overrides) | Name collisions & tests | Verdict |
|---|---|---|---|---|---|
| `CacheManager.has_data_changed` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.update_cache` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.setup_persistent_cache` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.get_sport_live_interval` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.get_sport_key_from_cache_key` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.get_background_cached_data` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.is_background_data_available` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.record_cache_hit` | 3.8.0 | core (1 internal) | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.record_cache_miss` | 3.8.0 | core (1 internal) | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.record_fetch_time` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.get_cache_metrics` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.log_cache_metrics` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `CacheManager.get_memory_cache_stats` | 3.8.0 | core tests (3 test calls) | — | — | unused — safe to remove in 3.8.0 |
| `DisplayManager.draw_sun` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `DisplayManager.draw_cloud` | 3.8.0 | core (2 internals) | — | ledmatrix-weather (1 unrelated) | unused — safe to remove in 3.8.0 |
| `DisplayManager.draw_rain` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `DisplayManager.draw_snow` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `DisplayManager.draw_weather_icon` | 3.8.0 | core (1 internal) | — | ledmatrix-weather (5 unrelateds) | unused — safe to remove in 3.8.0 |
| `DisplayManager.draw_text_with_icons` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `DisplayManager.get_scrolling_stats` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_manager_fonts` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_detected_fonts` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.unregister_plugin_fonts` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_plugin_fonts` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.set_override` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.remove_override` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_overrides` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_available_fonts` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_size_tokens` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_performance_stats` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.get_font_catalog` | 3.8.0 | core tests (1 test call) | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.add_font` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.remove_font` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `FontManager.validate_font` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
| `BasePlugin.get_supported_vegas_modes` | 3.9.0 | core tests (2 test reviews) | blackjack (1 call, 1 override); calendar (1 override); olympics (1 override) | — | still used by blackjack, calendar, olympics — keep or migrate first |
| `BasePlugin.get_vegas_segment_width` | 3.9.0 | core tests (1 test review) | — | — | unused — safe to remove in 3.9.0 |
| `PluginManager.get_enabled_plugins` | 3.8.0 | — | — | — | unused — safe to remove in 3.8.0 |
## Unused — safe to remove (36)
`CacheManager.has_data_changed`, `CacheManager.update_cache`, `CacheManager.setup_persistent_cache`, `CacheManager.get_sport_live_interval`, `CacheManager.get_sport_key_from_cache_key`, `CacheManager.get_background_cached_data`, `CacheManager.is_background_data_available`, `CacheManager.record_cache_hit`, `CacheManager.record_cache_miss`, `CacheManager.record_fetch_time`, `CacheManager.get_cache_metrics`, `CacheManager.log_cache_metrics`, `CacheManager.get_memory_cache_stats`, `DisplayManager.draw_sun`, `DisplayManager.draw_cloud`, `DisplayManager.draw_rain`, `DisplayManager.draw_snow`, `DisplayManager.draw_weather_icon`, `DisplayManager.draw_text_with_icons`, `DisplayManager.get_scrolling_stats`, `FontManager.get_manager_fonts`, `FontManager.get_detected_fonts`, `FontManager.unregister_plugin_fonts`, `FontManager.get_plugin_fonts`, `FontManager.set_override`, `FontManager.remove_override`, `FontManager.get_overrides`, `FontManager.get_available_fonts`, `FontManager.get_size_tokens`, `FontManager.get_performance_stats`, `FontManager.get_font_catalog`, `FontManager.add_font`, `FontManager.remove_font`, `FontManager.validate_font`, `BasePlugin.get_vegas_segment_width`, `PluginManager.get_enabled_plugins`
## Still used — keep or migrate first (1)
`BasePlugin.get_supported_vegas_modes`
## Every hit
File paths are relative to the plugin's directory (core: the repo root).
| Method | Where | File:line | Kind | Code |
|---|---|---|---|---|
| `CacheManager.get_sport_live_interval` | core | src/cache/cache_strategy.py:28 | unrelated | `def get_sport_live_interval(self, sport_key: str) -> int:` |
| `CacheManager.get_sport_live_interval` | core | src/cache/cache_strategy.py:60 | unrelated | `live_interval = self.get_sport_live_interval(sport_key)` |
| `CacheManager.get_sport_live_interval` | core | src/cache_manager.py:785 | unrelated | `return self._strategy_component.get_sport_live_interval(sport_key)` |
| `CacheManager.get_sport_key_from_cache_key` | core | src/cache/cache_strategy.py:214 | unrelated | `def get_sport_key_from_cache_key(self, key: str) -> Optional[str]:` |
| `CacheManager.get_sport_key_from_cache_key` | core | src/cache_manager.py:806 | unrelated | `return self._strategy_component.get_sport_key_from_cache_key(key)` |
| `CacheManager.get_sport_key_from_cache_key` | core | src/cache_manager.py:816 | unrelated | `sport_key = self._strategy_component.get_sport_key_from_cache_key(key)` |
| `CacheManager.record_cache_hit` | core | src/cache_manager.py:869 | internal (in `CacheManager.get_background_cached_data`) | `self.record_cache_hit('background')` |
| `CacheManager.record_cache_miss` | core | src/cache_manager.py:876 | internal (in `CacheManager.get_background_cached_data`) | `self.record_cache_miss('background')` |
| `CacheManager.record_fetch_time` | core | src/cache/cache_metrics.py:67 | unrelated | `def record_fetch_time(self, duration: float) -> None:` |
| `CacheManager.record_fetch_time` | core | src/cache_manager.py:922 | unrelated | `self._metrics_component.record_fetch_time(duration)` |
| `CacheManager.get_memory_cache_stats` | core tests | test/test_cache_manager_memory_tier.py:43 | test call | `stats = cm.get_memory_cache_stats()` |
| `CacheManager.get_memory_cache_stats` | core tests | test/test_cache_manager_memory_tier.py:63 | test call | `assert cm.get_memory_cache_stats()["last_cleanup"] >= before` |
| `CacheManager.get_memory_cache_stats` | core tests | test/test_cache_manager_memory_tier.py:68 | test call | `stats = cm.get_memory_cache_stats()` |
| `DisplayManager.draw_sun` | core | src/plugin_system/testing/visual_display_manager.py:417 | unrelated | `def draw_sun(self, x: int, y: int, size: int = 16):` |
| `DisplayManager.draw_cloud` | core | src/display_manager.py:1359 | internal (in `DisplayManager.draw_rain`) | `self.draw_cloud(x, y, size)` |
| `DisplayManager.draw_cloud` | core | src/display_manager.py:1374 | internal (in `DisplayManager.draw_snow`) | `self.draw_cloud(x, y, size)` |
| `DisplayManager.draw_cloud` | core | src/plugin_system/testing/visual_display_manager.py:421 | unrelated | `def draw_cloud(self, x: int, y: int, size: int = 16, color: Tuple[int, int, int] = (200, 200, 200)):` |
| `DisplayManager.draw_cloud` | ledmatrix-weather | weather_icons.py:184 | unrelated | `def draw_cloud(draw: ImageDraw, x: int, y: int, size: int = 16, color: tuple = (200, 200, 200)):` |
| `DisplayManager.draw_rain` | core | src/plugin_system/testing/visual_display_manager.py:425 | unrelated | `def draw_rain(self, x: int, y: int, size: int = 16):` |
| `DisplayManager.draw_snow` | core | src/plugin_system/testing/visual_display_manager.py:429 | unrelated | `def draw_snow(self, x: int, y: int, size: int = 16):` |
| `DisplayManager.draw_weather_icon` | core | src/display_manager.py:1518 | internal (in `DisplayManager.draw_text_with_icons`) | `self.draw_weather_icon(icon_type, icon_x, icon_y)` |
| `DisplayManager.draw_weather_icon` | core | src/plugin_system/testing/visual_display_manager.py:510 | unrelated | `def draw_weather_icon(self, condition: str, x: int, y: int, size: int = 16) -> None:` |
| `DisplayManager.draw_weather_icon` | core | src/plugin_system/testing/visual_display_manager.py:533 | unrelated | `self.draw_weather_icon(icon_type, icon_x, icon_y)` |
| `DisplayManager.draw_weather_icon` | ledmatrix-weather | manager.py:84 | unrelated | `def draw_weather_icon(image, icon_code, x, y, size):` |
| `DisplayManager.draw_weather_icon` | ledmatrix-weather | manager.py:1280 | unrelated | `WeatherIcons.draw_weather_icon(img, icon_code, icon_x, icon_y,` |
| `DisplayManager.draw_weather_icon` | ledmatrix-weather | manager.py:1559 | unrelated | `WeatherIcons.draw_weather_icon(img, forecast['icon'], icon_x, icon_y, icon_size)` |
| `DisplayManager.draw_weather_icon` | ledmatrix-weather | manager.py:1650 | unrelated | `WeatherIcons.draw_weather_icon(img, forecast['icon'], icon_x, icon_y, icon_size)` |
| `DisplayManager.draw_weather_icon` | ledmatrix-weather | weather_icons.py:168 | unrelated | `def draw_weather_icon(image: Image.Image, icon_code: str, x: int, y: int, size: int = DEFAULT_SIZE):` |
| `DisplayManager.draw_text_with_icons` | core | src/plugin_system/testing/visual_display_manager.py:526 | unrelated | `def draw_text_with_icons(self, text: str, icons: List[tuple] = None,` |
| `FontManager.get_font_catalog` | core tests | test/test_deprecation.py:229 | test call | `assert fm.get_font_catalog() == fm.font_catalog` |
| `BasePlugin.get_supported_vegas_modes` | core tests | test/test_vegas_participation.py:356 | test review | `assert plugin.get_supported_vegas_modes() == [` |
| `BasePlugin.get_supported_vegas_modes` | core tests | test/test_vegas_participation.py:358 | test review | `assert plugin.get_supported_vegas_modes()` |
| `BasePlugin.get_supported_vegas_modes` | blackjack | manager.py:732 | override | `def get_supported_vegas_modes(self):` |
| `BasePlugin.get_supported_vegas_modes` | blackjack | manager.py:695 | call | `if mode in self.get_supported_vegas_modes():` |
| `BasePlugin.get_supported_vegas_modes` | calendar | manager.py:875 | override | `def get_supported_vegas_modes(self) -> List[VegasDisplayMode]:` |
| `BasePlugin.get_supported_vegas_modes` | olympics | manager.py:624 | override | `def get_supported_vegas_modes(self) -> List[VegasDisplayMode]:` |
| `BasePlugin.get_vegas_segment_width` | core tests | test/test_vegas_participation.py:359 | test review | `assert plugin.get_vegas_segment_width() == 2` |
## Sources scanned
| Source | Group | Python files | Hits |
|---|---|---|---|
| core | core | 172 | 20 |
| core tests | core-tests | 347 | 17 |
| 7-segment-clock | monorepo | 3 | 0 |
| afl-scoreboard | monorepo | 35 | 0 |
| baseball-scoreboard | monorepo | 61 | 0 |
| basketball-scoreboard | monorepo | 49 | 0 |
| birdnet-go | monorepo | 2 | 0 |
| blackjack | monorepo | 7 | 2 |
| calendar | monorepo | 5 | 1 |
| christmas-countdown | monorepo | 3 | 0 |
| clock-simple | monorepo | 2 | 0 |
| countdown | monorepo | 5 | 0 |
| cricket-scoreboard | monorepo | 8 | 0 |
| f1-scoreboard | monorepo | 15 | 0 |
| fantasy-blitz | monorepo | 13 | 0 |
| football-scoreboard | monorepo | 74 | 0 |
| geochron | monorepo | 10 | 0 |
| hello-world | monorepo | 2 | 0 |
| hockey-scoreboard | monorepo | 52 | 0 |
| incoming-packages | monorepo | 8 | 0 |
| jellyfin-now-playing | monorepo | 4 | 0 |
| lacrosse-scoreboard | monorepo | 40 | 0 |
| ledmatrix-elections | monorepo | 12 | 0 |
| ledmatrix-flights | monorepo | 48 | 0 |
| ledmatrix-leaderboard | monorepo | 9 | 0 |
| ledmatrix-music | monorepo | 11 | 0 |
| ledmatrix-stocks | monorepo | 7 | 0 |
| ledmatrix-weather | monorepo | 15 | 6 |
| march-madness | monorepo | 4 | 0 |
| masters-tournament | monorepo | 10 | 0 |
| mqtt-notifications | monorepo | 4 | 0 |
| news | monorepo | 6 | 0 |
| nfl-draft | monorepo | 3 | 0 |
| nfl-stat-leaders | monorepo | 8 | 0 |
| nrl-scoreboard | monorepo | 30 | 0 |
| odds-ticker | monorepo | 9 | 0 |
| of-the-day | monorepo | 14 | 0 |
| olympics | monorepo | 16 | 1 |
| on-air | monorepo | 2 | 0 |
| pomodoro-timer | monorepo | 3 | 0 |
| soccer-scoreboard | monorepo | 47 | 0 |
| static-image | monorepo | 3 | 0 |
| stock-news | monorepo | 3 | 0 |
| text-display | monorepo | 4 | 0 |
| tide-display | monorepo | 3 | 0 |
| ufc-scoreboard | monorepo | 38 | 0 |
| web-ui-info | monorepo | 2 | 0 |
| youtube-stats | monorepo | 5 | 0 |
| f1-live | third-party | 10 | 0 |
| gif-player | third-party | 1 | 0 |
| pga-tour-leaderboard | third-party | 2 | 0 |
| plex-marquee | third-party | 1 | 0 |
| ledmatrix-dresden-departures | third-party | 1 | 0 |
| tidbyt-baseball-scoreboard | third-party | 2 | 0 |
| sleeper-fantasy | third-party | 1 | 0 |
| ledmatrix-nascar | third-party | 1 | 0 |
## How to re-run
```bash
# Clones the monorepo and each third-party plugin (depth 1) into a temp cache:
python3 scripts/plugin_api_usage.py --output docs/DEPRECATIONS_3.8.md
# Or scan a local monorepo checkout (read only) instead of cloning it:
python3 scripts/plugin_api_usage.py --monorepo ../ledmatrix-plugins
```
Before removing a method in its release, re-run the scan against the current monorepo and registry: a plugin added since this file was generated may have started calling it. Remove only methods the fresh scan reports unused; move the rest to a later release (the test in `test/test_deprecation.py` fails while a marker names a release at or below `src.__version__`).
+20 -10
View File
@@ -31,7 +31,7 @@ POST /api/v3/system/action
**Base URL**: `http://your-pi-ip:5000/api/v3`
See [API_REFERENCE.md](API_REFERENCE.md) for complete documentation.
See [REST_API_REFERENCE.md](REST_API_REFERENCE.md) for complete documentation.
## Display Manager Quick Methods
@@ -48,8 +48,14 @@ display_manager.draw_text("Centered", centered=True) # Auto-center
width = display_manager.get_text_width("Text", font)
height = display_manager.get_font_height(font)
# Weather icons
display_manager.draw_weather_icon("rain", x=10, y=10, size=16)
# Adaptive layout (recommended for multi-size support — text and images
# that scale to any panel; see docs/ADAPTIVE_LAYOUT.md)
rows = self.layout.bounds.inset(1).split_v(3, 1, gap=1)
self.draw_fit("12:34", rows[0]) # largest crisp font that fits
self.draw_image(logo, rows[1], mode="fill_height", crop_to_ink=True)
# Weather icons: draw_weather_icon() was removed in 3.8.0 — draw your
# own icons (the weather plugin ships WeatherIcons)
# Scrolling state
display_manager.set_scrolling_state(True)
@@ -66,20 +72,23 @@ cache_manager.delete("key") # alias for clear_cache(key)
# Advanced caching
data = cache_manager.get_cached_data_with_strategy("key", data_type="weather")
data = cache_manager.get_background_cached_data("key", sport_key="nhl")
# Strategy
strategy = cache_manager.get_cache_strategy("weather")
interval = cache_manager.get_sport_live_interval("nhl")
```
`get_background_cached_data()` (use `get()`) and `get_sport_live_interval()`
were removed in 3.8.0. See
[Deprecated APIs](PLUGIN_API_REFERENCE.md#deprecated-apis).
## Plugin Manager Quick Methods
```python
# Get plugins
plugin = plugin_manager.get_plugin("plugin-id")
all_plugins = plugin_manager.get_all_plugins()
enabled = plugin_manager.get_enabled_plugins()
# get_enabled_plugins() was removed in 3.8.0 — check `enabled` on the
# entries in plugin_manager.plugins
# Get info
info = plugin_manager.get_plugin_info("plugin-id")
@@ -162,7 +171,7 @@ def display(self, force_clear=False):
- [ ] Plugin inherits from `BasePlugin`
- [ ] Implements `update()` and `display()` methods
- [ ] `manifest.json` with required fields
- [ ] `manifest.json` with the [required fields](PLUGIN_API_REFERENCE.md#manifest-required-fields)
- [ ] `config_schema.json` for web UI (recommended)
- [ ] `README.md` with documentation
- [ ] Error handling implemented
@@ -184,12 +193,13 @@ def display(self, force_clear=False):
```
LEDMatrix/
├── plugins/ # Installed plugins
├── plugin-repos/ # Installed plugins (default; plugins/ is only
│ # for dev symlinks via scripts/dev/dev_plugin_setup.sh)
├── config/
│ ├── config.json # Main configuration
│ └── config_secrets.json # API keys and secrets
├── docs/ # Documentation
│ ├── API_REFERENCE.md
│ ├── REST_API_REFERENCE.md
│ ├── PLUGIN_API_REFERENCE.md
│ └── ...
└── src/
@@ -201,7 +211,7 @@ LEDMatrix/
## Quick Links
- [Complete API Reference](API_REFERENCE.md)
- [Complete REST API Reference](REST_API_REFERENCE.md)
- [Plugin API Reference](PLUGIN_API_REFERENCE.md)
- [Plugin Development Guide](PLUGIN_DEVELOPMENT_GUIDE.md)
- [Advanced Patterns](ADVANCED_PLUGIN_DEVELOPMENT.md)
+10 -9
View File
@@ -43,16 +43,21 @@ git submodule update --init --recursive rpi-rgb-led-matrix-master
#### Building the Submodule
After initializing the submodule, you need to build the Python bindings:
After initializing the submodule, build and install the `rgbmatrix` Python
package from the submodule root. Upstream's `pyproject.toml` builds it with
scikit-build-core, CMake and Ninja; there is no separate `make` step:
```bash
cd rpi-rgb-led-matrix-master
make build-python
cd bindings/python
python3 -m pip install --break-system-packages .
```
**Note:** The `first_time_install.sh` script automates this process during installation.
On a board with 1 GB of RAM or less, cap the compile so it doesn't run out of
memory: `CMAKE_BUILD_PARALLEL_LEVEL=1 python3 -m pip install --break-system-packages .`
**Note:** The `first_time_install.sh` script automates this process during
installation, including the parallelism cap and a temporary swapfile on
low-memory boards.
#### Troubleshooting
@@ -69,7 +74,7 @@ git submodule update --init --recursive rpi-rgb-led-matrix-master
**Build fails:**
Ensure you have the required build dependencies installed:
```bash
sudo apt install -y build-essential python3-dev cython3 scons
sudo apt install -y build-essential python-dev-is-python3 cmake ninja-build
```
**Import error for `rgbmatrix` module:**
@@ -97,8 +102,6 @@ When setting up CI/CD pipelines, ensure submodules are initialized before buildi
- name: Build rpi-rgb-led-matrix
run: |
cd rpi-rgb-led-matrix-master
make build-python
cd bindings/python
pip install .
```
@@ -110,8 +113,6 @@ variables:
build:
script:
- cd rpi-rgb-led-matrix-master
- make build-python
- cd bindings/python
- pip install .
```
+6
View File
@@ -6,6 +6,12 @@ Tools for rapid plugin development without deploying to the RPi.
Interactive web UI for tweaking plugin configs and seeing the rendered display in real time.
The size inputs have a preset dropdown with the harness's standard panel
sizes, and the **All Sizes** button renders the current config at every
harness size in a side-by-side gallery (`POST /api/render-matrix`) — the
quickest way to eyeball adaptive-layout behavior across panels
(see [ADAPTIVE_LAYOUT.md](ADAPTIVE_LAYOUT.md)).
### Quick Start
```bash
+90 -82
View File
@@ -17,13 +17,13 @@ The LEDMatrix emulator allows you to run and test LEDMatrix displays on your com
## Prerequisites
### System Requirements
- Python 3.7 or higher
- Python 3.10 or higher
- Windows, macOS, or Linux
- At least 2GB RAM (4GB recommended)
- Internet connection for plugin downloads
### Required Software
- Python 3.7+
- Python 3.10+
- pip (Python package manager)
- Git (for plugin management)
@@ -50,8 +50,7 @@ pip install -r requirements-emulator.txt
```
This installs:
- `RGBMatrixEmulator` - The core emulation library
- Additional dependencies for display adapters
- `RGBMatrixEmulator` - the emulation library (and whatever it depends on)
### 3. Install Standard Dependencies
@@ -63,29 +62,31 @@ pip install -r requirements.txt
### 1. Emulator Configuration File
The emulator uses `emulator_config.json` for configuration. Here's the
default configuration as it ships in the repo:
The emulator uses `emulator_config.json` for configuration. It isn't in
the repo (it's gitignored): RGBMatrixEmulator writes it on first run.
A typical file looks like this:
```json
{
"pixel_outline": 0,
"pixel_size": 5,
"pixel_size": 16,
"pixel_style": "square",
"pixel_glow": 6,
"display_adapter": "pygame",
"display_adapter": "browser",
"allow_adapter_fallback": true,
"icon_path": null,
"emulator_title": null,
"suppress_font_warnings": false,
"suppress_adapter_load_errors": false,
"browser": {
"_comment": "For use with the browser adapter only.",
"port": 8888,
"target_fps": 24,
"target_fps": 60,
"fps_display": false,
"quality": 70,
"image_border": true,
"debug_text": false,
"image_format": "JPEG"
"image_format": "JPEG",
"open_immediately": false
},
"log_level": "info"
}
@@ -96,13 +97,13 @@ default configuration as it ships in the repo:
| Option | Description | Default | Values |
|--------|-------------|---------|--------|
| `pixel_outline` | Pixel border thickness | 0 | 0-5 |
| `pixel_size` | Size of each pixel | 5 | 1-64 (8–16 is typical for testing) |
| `pixel_size` | Size of each pixel | 16 | 1-64 (8–16 is typical for testing) |
| `pixel_style` | Pixel shape | "square" | "square", "circle" |
| `pixel_glow` | Glow effect intensity | 6 | 0-20 |
| `display_adapter` | Display backend | "pygame" | "pygame", "browser" |
| `display_adapter` | Display backend | "browser" | "browser", "pygame" |
| `allow_adapter_fallback` | Fall back to another adapter if the configured one fails to load | true | true/false |
| `emulator_title` | Window title | null | Any string |
| `suppress_font_warnings` | Hide font warnings | false | true/false |
| `suppress_adapter_load_errors` | Hide adapter errors | false | true/false |
### 3. Browser Adapter Configuration
@@ -111,18 +112,32 @@ When using the browser adapter, additional options are available:
| Option | Description | Default |
|--------|-------------|---------|
| `port` | Web server port | 8888 |
| `target_fps` | Target frames per second | 24 |
| `target_fps` | Target frames per second | 60 |
| `fps_display` | Show FPS counter | false |
| `quality` | Image compression quality | 70 |
| `image_border` | Show image border | true |
| `debug_text` | Show debug information | false |
| `image_format` | Image format | "JPEG" |
| `open_immediately` | Open the browser page automatically on start | false |
## Running the Emulator
### 1. Set Environment Variable
### 1. Use the `-e` Flag (Recommended)
Enable emulator mode by setting the `EMULATOR` environment variable:
`run.py` accepts exactly two flags: `-e`/`--emulator` and
`-d`/`--debug`.
```bash
python3 run.py -e
# With verbose logging
python3 run.py -e -d
```
### 2. Alternative: Set the Environment Variable
You can also enable emulator mode via the `EMULATOR` environment
variable:
**Windows (Command Prompt):**
```cmd
@@ -137,15 +152,6 @@ python run.py
```
**Linux/macOS:**
```bash
export EMULATOR=true
python3 run.py
```
### 2. Alternative: Direct Python Execution
You can also run the emulator directly:
```bash
EMULATOR=true python3 run.py
```
@@ -153,7 +159,8 @@ EMULATOR=true python3 run.py
### 3. Verify Emulator Mode
When running in emulator mode, you should see:
- A window displaying the LED matrix simulation
- The emulated matrix — a web page at `http://localhost:8888` with the
default browser adapter, or a desktop window with the pygame adapter
- Console output indicating emulator mode
- No hardware initialization errors
@@ -161,7 +168,36 @@ When running in emulator mode, you should see:
LEDMatrix supports two display adapters for the emulator:
### 1. Pygame Adapter (Default)
### 1. Browser Adapter (Default)
The browser adapter runs a web server and displays the matrix as a web
page at `http://localhost:8888`. This is the adapter the shipped
`emulator_config.json` uses.
**Features:**
- Web-based interface
- Remote access capability
- Mobile-friendly
- Screenshot capture
**Configuration:**
```json
{
"display_adapter": "browser",
"browser": {
"port": 8888,
"target_fps": 60,
"quality": 70
}
}
```
**Usage:**
1. Start the emulator (`python3 run.py -e`)
2. Open browser to `http://localhost:8888`
3. View the LED matrix display
### 2. Pygame Adapter (Alternative)
The pygame adapter provides a native desktop window with real-time display.
@@ -186,33 +222,6 @@ The pygame adapter provides a native desktop window with real-time display.
- `+/-` - Zoom in/out
- `R` - Reset zoom
### 2. Browser Adapter
The browser adapter runs a web server and displays the matrix in a web browser.
**Features:**
- Web-based interface
- Remote access capability
- Mobile-friendly
- Screenshot capture
**Configuration:**
```json
{
"display_adapter": "browser",
"browser": {
"port": 8888,
"target_fps": 24,
"quality": 70
}
}
```
**Usage:**
1. Start the emulator with browser adapter
2. Open browser to `http://localhost:8888`
3. View the LED matrix display
## Troubleshooting
### Common Issues
@@ -274,8 +283,7 @@ Enable debug logging:
```json
{
"log_level": "debug",
"suppress_font_warnings": false,
"suppress_adapter_load_errors": false
"suppress_font_warnings": false
}
```
@@ -299,17 +307,18 @@ Modify the display dimensions in your main config:
### 2. Plugin Development
For plugin development with the emulator:
`run.py` always runs the full rotation — it has no single-plugin flag.
To preview or check one plugin in isolation, use the dev tools:
```bash
# Enable emulator mode
export EMULATOR=true
# Run the full display in emulator mode (optionally with debug logging)
python3 run.py -e -d
# Run with specific plugin
python run.py --plugin my-plugin
# Live single-plugin preview in the browser (port 5001)
python3 scripts/dev_server.py
# Debug mode
python run.py --debug
# Headless render/validation of one plugin
python3 scripts/check_plugin.py --plugin my-plugin
```
### 3. Performance Tuning
@@ -344,11 +353,10 @@ The emulator can work alongside the web interface:
```bash
# Terminal 1: Start emulator
export EMULATOR=true
python run.py
python3 run.py -e
# Terminal 2: Start web interface
python web_interface/app.py
# Terminal 2: Start web interface (supported entry point)
python3 web_interface/start.py
```
Access the web interface at `http://localhost:5000` while the emulator runs.
@@ -365,13 +373,14 @@ Access the web interface at `http://localhost:5000` while the emulator runs.
### 2. Plugin Testing
```bash
# Test specific plugin
export EMULATOR=true
python run.py --plugin clock-simple
# Test a specific plugin (headless check)
python3 scripts/check_plugin.py --plugin clock-simple
# Test all plugins
export EMULATOR=true
python run.py --test-plugins
# Preview a single plugin live in the browser (port 5001)
python3 scripts/dev_server.py
# Test the full rotation in the emulator
python3 run.py -e
```
### 3. Configuration Management
@@ -385,9 +394,8 @@ python run.py --test-plugins
### Basic Clock Display
```bash
# Start emulator with clock
export EMULATOR=true
python run.py
# Start emulator with clock enabled in config.json
python3 run.py -e
```
### Sports Scores
@@ -395,16 +403,16 @@ python run.py
```bash
# Configure for sports display
# Edit config/config.json to enable sports plugins
export EMULATOR=true
python run.py
python3 run.py -e
```
### Custom Text Display
```bash
# Use text display plugin
export EMULATOR=true
python run.py --plugin text-display --text "Hello World"
# Preview the text display plugin on its own
python3 scripts/check_plugin.py --plugin text-display
# or use the live dev preview server
python3 scripts/dev_server.py
```
## Support
+157 -322
View File
@@ -1,167 +1,94 @@
# FontManager Usage Guide
> **Picking a size automatically:** if you want the *largest font that fits
> a given area* rather than a fixed size, use the adaptive layout system's
> font ladders, which resolve through this FontManager. `BasePlugin`
> subclasses get this as `self.layout.fit_text(...)`; other code can build
> a `LayoutContext(width, height, font_manager)` directly — see
> [ADAPTIVE_LAYOUT.md](ADAPTIVE_LAYOUT.md).
## Overview
The enhanced FontManager provides comprehensive font management for the LEDMatrix application with support for:
- Manager font registration and detection
- Plugin font management
- Manual font overrides via web interface
- Performance monitoring and caching
- Dynamic font discovery
[`src/font_manager.py`](../src/font_manager.py) loads and caches the TTF and
BDF fonts in `assets/fonts/`, registers fonts that plugins ship, and records
which plugin uses which font so the web UI can show it.
## Architecture
Several methods were removed in LEDMatrix 3.8.0 after a release of
deprecation warnings; [Removed methods](#removed-methods) below lists them
with what to use instead.
### Manager-Centric Design
## Getting the FontManager
Managers define their own fonts, but the FontManager:
1. **Loads and caches fonts** for performance
2. **Detects font usage** for visibility
3. **Allows manual overrides** when needed
4. **Supports plugin fonts** with namespacing
### Font Resolution Flow
```
Manager requests font → Check manual overrides → Apply manager choice → Cache & return
```
## For Manager Developers
### Basic Font Usage
There is one shared FontManager per display process. The display controller
creates it and hands it to the `PluginManager`, so a plugin reaches it
through its `plugin_manager`:
```python
from src.font_manager import FontManager
class MyManager:
def __init__(self, config, display_manager, cache_manager):
self.font_manager = display_manager.font_manager # Access shared FontManager
self.manager_id = "my_manager"
def display(self):
# Define your font choices
element_key = "my_manager.title"
font_family = "press_start"
font_size_px = 10
color = (255, 255, 255) # RGB white
# Register your font choice (for detection and future overrides)
self.font_manager.register_manager_font(
manager_id=self.manager_id,
element_key=element_key,
family=font_family,
size_px=font_size_px,
color=color
)
# Get the font (checks for manual overrides automatically)
font = self.font_manager.resolve_font(
element_key=element_key,
family=font_family,
size_px=font_size_px
)
# Use the font for rendering
self.display_manager.draw_text(
"Hello World",
x=10, y=10,
color=color,
font=font
)
class MyPlugin(BasePlugin):
def __init__(self, plugin_id, config, display_manager, cache_manager, plugin_manager):
super().__init__(plugin_id, config, display_manager, cache_manager, plugin_manager)
self.font_manager = self._get_font_manager()
```
### Advanced Font Usage
`BasePlugin._get_font_manager()` returns `plugin_manager.font_manager`, or a
standalone FontManager when none is available (test harnesses, mocks).
`DisplayManager` has **no** `font_manager` attribute —
`display_manager.font_manager` raises `AttributeError`.
## Resolving a font
```python
class AdvancedManager:
def __init__(self, config, display_manager, cache_manager):
self.font_manager = display_manager.font_manager
self.manager_id = "advanced_manager"
# Define your font specifications
self.font_specs = {
"title": {"family": "press_start", "size_px": 12, "color": (255, 255, 0)},
"body": {"family": "four_by_six", "size_px": 8, "color": (255, 255, 255)},
"footer": {"family": "five_by_seven", "size_px": 7, "color": (128, 128, 128)}
}
# Register all font specs
for element_type, spec in self.font_specs.items():
element_key = f"{self.manager_id}.{element_type}"
self.font_manager.register_manager_font(
manager_id=self.manager_id,
element_key=element_key,
family=spec["family"],
size_px=spec["size_px"],
color=spec["color"]
)
def get_font(self, element_type: str):
"""Helper method to get fonts with override support."""
spec = self.font_specs[element_type]
element_key = f"{self.manager_id}.{element_type}"
return self.font_manager.resolve_font(
element_key=element_key,
family=spec["family"],
size_px=spec["size_px"]
)
def display(self):
# Get fonts (automatically checks for overrides)
title_font = self.get_font("title")
body_font = self.get_font("body")
footer_font = self.get_font("footer")
# Render with fonts
self.display_manager.draw_text("Title", font=title_font, color=self.font_specs["title"]["color"])
self.display_manager.draw_text("Body Text", font=body_font, color=self.font_specs["body"]["color"])
self.display_manager.draw_text("Footer", font=footer_font, color=self.font_specs["footer"]["color"])
```
element_key = f"{self.plugin_id}.title"
### Using Size Tokens
```python
# Get available size tokens
tokens = self.font_manager.get_size_tokens()
# Returns: {'xs': 6, 'sm': 8, 'md': 10, 'lg': 12, 'xl': 14, 'xxl': 16}
# Use token to get size
size_px = tokens.get('md', 10) # 10px
# Then use in font resolution
font = self.font_manager.resolve_font(
element_key="my_manager.text",
# Register the choice so the web UI's Fonts tab can list it.
self.font_manager.register_manager_font(
manager_id=self.plugin_id,
element_key=element_key,
family="press_start",
size_px=size_px
size_px=10,
color=(255, 255, 255),
)
font = self.font_manager.resolve_font(
element_key=element_key,
family="press_start",
size_px=10,
)
self.display_manager.draw_text("Hello", x=10, y=10, font=font)
```
## For Plugin Developers
`resolve_font()` applies any entry for `element_key` in
`config/font_overrides.json`, maps a plugin-local family to its namespaced
name when `plugin_id` is passed, and then calls `get_font(family, size_px)`.
On error it returns a fallback font rather than raising.
> **Note**: plugins that ship their own fonts via a `"fonts"` block
> in `manifest.json` are registered automatically during plugin load
> (`src/plugin_system/plugin_manager.py` calls
> `FontManager.register_plugin_fonts()`). The `plugin://…` source
> URIs documented below are resolved relative to the plugin's
> install directory.
>
> The **Fonts** tab in the web UI that lists detected
> manager-registered fonts is still a **placeholder
> implementation** — fonts that managers register through
> `register_manager_font()` do not yet appear there. The
> programmatic per-element override workflow described in
> [Manual Font Overrides](#manual-font-overrides) below
> (`set_override()` / `remove_override()` / the
> `config/font_overrides.json` store) **does** work today and is
> the supported way to override a font for an element until the
> Fonts tab is wired up. If you can't wait and need a workaround
> right now, you can also just load the font directly with PIL
> (or `freetype-py` for BDF) inside your plugin's `manager.py`
> and skip the override system entirely.
`get_font(family, size_px)` looks the family up in `font_catalog` and loads
it (cached per family and size).
### Plugin Font Registration
## Font families
In your plugin's `manifest.json`:
At start-up the FontManager scans `assets/fonts/` for `.ttf` and `.bdf`
files. Each becomes a family named after the file, lower-cased and without
the extension (`PressStart2P-Regular.ttf` → `pressstart2p-regular`). Four
aliases are added on top:
| Alias | File |
|---|---|
| `press_start` | `assets/fonts/PressStart2P-Regular.ttf` |
| `four_by_six` | `assets/fonts/4x6-font.ttf` |
| `five_by_seven` | `assets/fonts/5x7.bdf` |
| `tom_thumb` | `assets/fonts/tom-thumb.bdf` |
Read the catalog directly: `font_manager.font_catalog` is a dict of family
name to file path. Files added later are picked up on the next start of the
display service.
## Plugin fonts
Plugins that ship their own fonts declare them in a `"fonts"` block in
`manifest.json`. The plugin manager calls
`FontManager.register_plugin_fonts()` during plugin load. `plugin://…`
sources are resolved relative to the plugin's install directory.
```json
{
@@ -172,216 +99,124 @@ In your plugin's `manifest.json`:
{
"family": "custom_font",
"source": "plugin://fonts/custom.ttf",
"metadata": {
"description": "Custom plugin font",
"license": "MIT"
}
"metadata": {"description": "Custom plugin font", "license": "MIT"}
},
{
"family": "web_font",
"source": "https://example.com/fonts/font.ttf",
"metadata": {
"description": "Downloaded font",
"checksum": "sha256:abc123..."
}
"metadata": {"checksum": "sha256:abc123..."}
}
]
}
}
```
### Using Plugin Fonts
Registered families are namespaced as `<plugin_id>::<family>`. Pass
`plugin_id` to `resolve_font()` to use the short name:
```python
class PluginManager:
def __init__(self, config, display_manager, cache_manager, plugin_id):
self.font_manager = display_manager.font_manager
self.plugin_id = plugin_id
def display(self):
# Use plugin font (automatically namespaced)
font = self.font_manager.resolve_font(
element_key=f"{self.plugin_id}.text",
family="custom_font", # Will be resolved as "my-plugin::custom_font"
size_px=10,
plugin_id=self.plugin_id
)
self.display_manager.draw_text("Plugin Text", font=font)
```
## Manual Font Overrides
Users can override any font through the web interface:
1. Navigate to **Fonts** tab
2. View **Detected Manager Fonts** to see what's currently in use
3. In **Element Overrides** section:
- Select the element (e.g., "nfl.live.score")
- Choose a different font family
- Choose a different size
- Click **Add Override**
Overrides are stored in `config/font_overrides.json` and persist across restarts.
### Programmatic Overrides
```python
# Set override
font_manager.set_override(
element_key="nfl.live.score",
family="four_by_six",
size_px=8
font = self.font_manager.resolve_font(
element_key=f"{self.plugin_id}.text",
family="custom_font", # resolved as "my-plugin::custom_font"
size_px=10,
plugin_id=self.plugin_id,
)
# Remove override
font_manager.remove_override("nfl.live.score")
# Get all overrides
overrides = font_manager.get_overrides()
```
## Font Discovery
## Overrides
### Available Fonts
`resolve_font()` still honours `config/font_overrides.json` (a map of
element key to `family` and/or `size_px`), which is read once at start-up.
The methods that edited it — `set_override()`, `remove_override()`,
`get_overrides()` — were removed in 3.8.0, and there is no web UI or REST endpoint
for overrides (the override editor and `/api/v3/fonts/overrides` were
removed). To let users choose a font, add a field to your plugin's config
schema.
The FontManager automatically scans `assets/fonts/` for TTF and BDF fonts:
## Font usage in the web UI
The web UI's **Fonts** tab lists, uploads, previews and deletes the font
files in `assets/fonts/`. The web interface runs in its own process and has
no FontManager, so the display service publishes which plugin uses which
font ([`src/font_usage.py`](../src/font_usage.py)), and the tab's **Used by**
column reads it:
- **Source**: `register_manager_font()` registrations of the loaded
plugins. `get_font()` and `resolve_font()` do not know the calling plugin
and are not counted, and neither is a plugin that opens a font file
directly with PIL — register the fonts your plugin draws with if you want
them listed.
- **Names**: a family, alias or path is resolved through `font_catalog` to
the file it loads and reported under that file's name without extension
(`PressStart2P-Regular`, `4x6-font`, `5x7`, `tom-thumb`), which is how the
Fonts tab keys its rows. Fonts outside `assets/fonts/` (a plugin's own
`plugin_id::family` fonts) and families that resolve to nothing are left
out.
- **When**: a daemon thread started once plugins have loaded checks every
10 seconds and writes the `font_usage_snapshot` cache key only when the
usage changed (and once a day, so the cache's cleanup never expires it).
Unloading a plugin drops its registrations (`forget_manager_fonts`).
- **Unknown**: until the display service has published, the column reads
"unknown" and `GET /api/v3/fonts/catalog` returns `used_by: null`.
- The tab warns before deleting a font that a loaded plugin registered.
## Text measurement
```python
# Get all available fonts
fonts = font_manager.get_available_fonts()
# Returns: {'press_start': 'assets/fonts/PressStart2P-Regular.ttf', ...}
# Check if font exists
if "my_font" in fonts:
font = font_manager.get_font("my_font", 10)
```
### Adding Custom Fonts
Place font files in `assets/fonts/` directory:
- Supported formats: `.ttf`, `.bdf`
- Font family name is derived from filename (without extension)
- Will be automatically discovered on next initialization
## Performance Monitoring
```python
# Get performance stats
stats = font_manager.get_performance_stats()
print(f"Cache hit rate: {stats['cache_hit_rate']*100:.1f}%")
print(f"Total fonts cached: {stats['total_fonts_cached']}")
print(f"Failed loads: {stats['failed_loads']}")
print(f"Manager fonts: {stats['manager_fonts']}")
print(f"Plugin fonts: {stats['plugin_fonts']}")
```
## Text Measurement
```python
# Measure text dimensions
width, height, baseline = font_manager.measure_text("Hello", font)
# Get font height
font_height = font_manager.get_font_height(font)
```
## Best Practices
## Tips
### For Managers
1. **Register all fonts** you use for visibility
2. **Use consistent element keys** (e.g., `{manager_id}.{element_type}`)
3. **Cache font references** if using same font multiple times
4. **Use `resolve_font()`** not `get_font()` directly to support overrides
5. **Define sensible defaults** that work well on LED matrix
### For Plugins
1. **Use plugin-relative paths** (`plugin://fonts/...`)
2. **Include font metadata** (license, description)
3. **Provide fallback** fonts if custom fonts fail to load
4. **Test with different display sizes**
### General
1. **BDF fonts** are often better for small sizes on LED matrices
2. **TTF fonts** work well for larger sizes
3. **Monospace fonts** are easier to align
4. **Test on actual hardware** - what looks good on screen may not work on LED matrix
## Migration from Old System
### Old Way (Direct Font Loading)
```python
self.font = ImageFont.truetype("assets/fonts/PressStart2P-Regular.ttf", 8)
```
### New Way (FontManager)
```python
element_key = f"{self.manager_id}.text"
self.font_manager.register_manager_font(
manager_id=self.manager_id,
element_key=element_key,
family="pressstart2p-regular",
size_px=8
)
self.font = self.font_manager.resolve_font(
element_key=element_key,
family="pressstart2p-regular",
size_px=8
)
```
- BDF fonts usually look better than TTF at small sizes on LED panels.
- Use `{plugin_id}.{element}` element keys.
- Register the fonts you draw with, so the Fonts tab can warn before one is
deleted.
- Replace direct `ImageFont.truetype("assets/fonts/...", 8)` calls with
`resolve_font()`: it caches, resolves paths against the install directory,
and handles BDF files.
## Troubleshooting
### Font Not Found
- Check font file exists in `assets/fonts/`
- Verify font family name matches filename (without extension, lowercase)
- Check logs for font discovery errors
**Font not found**
- Check the file exists in `assets/fonts/`.
- The family name is the filename without extension, lower-cased.
- Check the display service log for font discovery errors.
### Override Not Working
- Verify element key matches exactly what manager registered
- Check `config/font_overrides.json` for correct syntax
- Restart application to ensure overrides are loaded
**Plugin fonts not loading**
- Check the manifest's `"fonts"` block.
- Check the log for download or registration errors, and that font URLs are
reachable.
### Performance Issues
- Check cache hit rate in performance stats
- Reduce number of unique font/size combinations
- Clear cache if it grows too large: `font_manager.clear_cache()`
## API reference
### Plugin Fonts Not Loading
- Verify plugin manifest syntax
- Check plugin directory structure
- Review logs for download/registration errors
- Ensure font URLs are accessible
Current methods:
## API Reference
| Method | Purpose |
|---|---|
| `register_manager_font(manager_id, element_key, family, size_px, color=None)` | Record a font choice (feeds the Fonts tab) |
| `forget_manager_fonts(manager_id)` | Drop a manager's registrations (core calls it when a plugin unloads) |
| `resolve_font(element_key, family, size_px, plugin_id=None)` | Get a font, applying overrides and plugin namespacing |
| `get_font(family, size_px)` | Get a font directly |
| `get_native_bdf_size(family)` | Native pixel size of a BDF family, or `None` |
| `measure_text(text, font)` | `(width, height, baseline)` |
| `get_font_height(font)` | Line height |
| `register_plugin_fonts(plugin_id, font_manifest)` | Register a plugin's fonts (core calls it at load) |
| `clear_cache()` | Drop cached fonts and metrics |
| `font_catalog` (attribute) | Family name → file path |
### FontManager Methods
### Removed methods
- `register_manager_font(manager_id, element_key, family, size_px, color=None)` - Register font usage
- `resolve_font(element_key, family, size_px, plugin_id=None)` - Get font with override support
- `get_font(family, size_px)` - Get font directly (bypasses overrides)
- `measure_text(text, font)` - Measure text dimensions
- `get_font_height(font)` - Get font height
- `set_override(element_key, family=None, size_px=None)` - Set manual override
- `remove_override(element_key)` - Remove override
- `get_overrides()` - Get all overrides
- `get_detected_fonts()` - Get all detected font usage
- `get_manager_fonts(manager_id=None)` - Get fonts by manager
- `get_available_fonts()` - Get font catalog
- `get_size_tokens()` - Get size token definitions
- `get_performance_stats()` - Get performance metrics
- `clear_cache()` - Clear font cache
- `register_plugin_fonts(plugin_id, font_manifest)` - Register plugin fonts
- `unregister_plugin_fonts(plugin_id)` - Unregister plugin fonts
## Example: Complete Manager Implementation
For a working example of the font manager API in use, see
`src/font_manager.py` itself and the bundled scoreboard base classes
in `src/base_classes/` (e.g., `hockey.py`, `football.py`) which
register and resolve fonts via the patterns documented above.
Removed in 3.8.0, after logging a deprecation warning on first call since
3.5.0.
| Method | Use instead |
|---|---|
| `get_available_fonts()`, `get_font_catalog()` | read `font_catalog` |
| `get_size_tokens()` | pass a pixel size |
| `get_performance_stats()` | — |
| `set_override()`, `remove_override()`, `get_overrides()` | a font field in your plugin's config schema |
| `get_manager_fonts()`, `get_detected_fonts()` | — |
| `get_plugin_fonts()`, `unregister_plugin_fonts()` | — |
| `add_font()`, `remove_font()`, `validate_font()` | the web UI's Fonts tab |
+71 -28
View File
@@ -21,18 +21,30 @@ This guide will help you set up your LEDMatrix display for the first time and ge
---
## Quick Start (5 Minutes)
## Quick Start
### 1. First Boot
### 1. Install LEDMatrix
1. Insert the MicroSD card with LEDMatrix installed
2. Connect the LED matrix to your Raspberry Pi
3. Plug in the power supply
4. Wait for the Pi to boot (about 60 seconds)
There is no prebuilt SD card image — you install LEDMatrix onto stock
Raspberry Pi OS Lite yourself:
**Expected Behavior:**
1. Flash Raspberry Pi OS Lite to the MicroSD card (Raspberry Pi Imager)
2. Connect the LED matrix to your Raspberry Pi, insert the card, and
power on
3. SSH into the Pi and run the one-shot installer:
```bash
curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | bash
```
or clone the repo and run `sudo ./first_time_install.sh` — see the
[README Installation Steps / Quick Install](../README.md#installation-steps)
for full details
**Expected Behavior after install:**
- LED matrix will light up
- Display will show default plugins (clock, weather, etc.)
- A fresh install ships only the bundled `starlark-apps` and
`web-ui-info` plugins — clock, weather, sports, etc. must be
installed from the Plugin Store (web UI → Plugin Manager) before
anything else displays
- Pi creates WiFi network "LEDMatrix-Setup" if not connected
### 2. Connect to WiFi
@@ -71,10 +83,10 @@ You should see:
1. Open the **Display** tab
2. Set your matrix configuration:
- **Rows**: 32 or 64 (match your hardware)
- **Columns**: commonly 64 or 96; the web UI accepts any integer
in the 16–128 range, but 64 and 96 are the values the bundled
panel hardware ships with
- **Rows**: match your panel — commonly 32 or 64; any even number
from 8 to 64
- **Columns**: match your panel — commonly 64 or 96; at least 16,
with no upper limit
- **Chain Length**: Number of panels chained horizontally
- **Hardware Mapping**: usually `adafruit-hat-pwm` (with the PWM jumper
mod) or `adafruit-hat` (without). See the root README for the full list.
@@ -104,8 +116,8 @@ weather and other location-aware plugins.
4. Wait for installation to finish — installed plugins appear in the
**Installed Plugins** section above and get their own tab in the second
nav row
5. Toggle the plugin to enabled
6. From **Overview**, click **Restart Display Service**
5. Toggle the plugin to enabled. The running display loads it within a
few seconds; no restart is needed
You can also install community plugins straight from a GitHub URL using the
**Install from GitHub** section further down the same tab — see
@@ -115,10 +127,15 @@ You can also install community plugins straight from a GitHub URL using the
1. Each installed plugin gets its own tab in the second navigation row
2. Open that plugin's tab to edit its settings (favorite teams, API keys,
update intervals, display duration, etc.)
3. Click **Save**
4. Restart the display service from **Overview** so the new settings take
effect
update intervals, etc.)
3. Click **Save**. The display service watches `config.json` and hands the
new settings to the running plugin, so no restart is needed. If a plugin
still shows old settings, restart the display service from **Overview**
**Note:** how long each plugin stays on screen is not set in the
plugin's own tab — use the **Rotation** tab's **Screen Durations**
section instead (saved to `display.display_durations` in
`config.json`).
**Example: Weather Plugin**
- Set your location (city, state, country)
@@ -180,14 +197,15 @@ The fastest way to verify a plugin works without waiting for the rotation:
**Check:**
1. Plugin is enabled (toggle on the **Plugin Manager** tab)
2. Display service was restarted after enabling
3. Plugin's display duration is non-zero
4. No errors in the **Logs** tab for that plugin
2. Plugin's display duration is non-zero
3. No errors in the **Logs** tab for that plugin. A plugin whose
`validate_config()` fails is not loaded until its settings are fixed
**Fix:**
1. Enable the plugin from **Plugin Manager**
2. Click **Restart Display Service** on **Overview**
3. Check the **Logs** tab for plugin-specific errors
2. Check the **Logs** tab for plugin-specific errors
3. If it still does not appear, click **Restart Display Service** on
**Overview**
### Weather Plugin Shows "No Data"
@@ -208,18 +226,36 @@ The fastest way to verify a plugin works without waiting for the rotation:
### Customize Your Display
**Adjust display durations:**
- Each plugin's tab has a **Display Duration (seconds)** field — set how
long that plugin stays on screen each rotation.
- Open the **Rotation** tab and use the **Screen Durations** section to
set how long each plugin stays on screen per rotation (saved to
`display.display_durations`).
**Organize plugin order:**
- Use the **Plugin Manager** tab to enable/disable plugins. The display
cycles through enabled plugins in the order they appear.
- The **Rotation** tab also has a drag-and-drop **Rotation Order** list
(saved to `display.plugin_rotation_order`). Enable/disable plugins
from the **Plugin Manager** tab.
**Add more plugins:**
- Check the **Plugin Store** section of **Plugin Manager** for new plugins.
- Install community plugins straight from a GitHub URL via
**Install from GitHub** on the same tab.
### Keep LEDMatrix Up to Date
- **Update Code** on the **Overview** tab installs the newest version, and a
banner at the top of the page says when one is available.
- **General → Automatic Updates** does it once a week, overnight, with a
health check that undoes an update that breaks the device.
- **General → Update Channel** picks which version that is. **Stable** (the
default) installs releases, which have been tested and have release
notes. **Beta** installs the newest code as soon as it is written, before
it is released: fixes arrive sooner, and so do new problems.
- Switching to Stable never installs an older version than the one you
have. If your device is already newer than the latest release (which is
normal if it was set up or updated from the newest code), it keeps
getting the newest code until the next release includes it, then follows
releases from there. The General tab says when this is the case.
### Enable Advanced Features
**Vegas Scroll Mode:**
@@ -280,10 +316,14 @@ sudo journalctl -u ledmatrix-web -f
│ ├── config_secrets.json # API keys and secrets
│ └── wifi_config.json # WiFi settings
├── plugin-repos/ # Installed plugins (default location)
├── cache/ # Cached data
└── web_interface/ # Web interface files
```
> Cached data does not live in the project directory — the cache manager
> uses the first writable location among `/var/cache/ledmatrix`,
> `~/.ledmatrix_cache`, `/opt/ledmatrix/cache`, and
> `$TMPDIR/ledmatrix_cache`.
>
> The plugin install location is configurable via
> `plugin_system.plugins_directory` in `config.json`. The default is
> `plugin-repos/`. Plugin discovery (`PluginManager.discover_plugins()`)
@@ -303,11 +343,14 @@ System tabs:
- WiFi Network selection and AP-mode setup
- Schedule Power and dim schedules
- Display Matrix hardware configuration
- Rotation Rotation order (drag-and-drop) and screen durations
- Config Editor Raw config.json editor
- Backup & Restore Config backup and restore
- Fonts Upload and manage fonts
- Logs Real-time log viewing
- Cache Cached data inspection and cleanup
- Operation History Recent service operations
- Tools System diagnostics, updates, dependencies, maintenance
Plugin tabs (second row):
- Plugin Manager Browse the Plugin Store, install/enable plugins
+78 -88
View File
@@ -10,10 +10,7 @@ Make sure you have the testing packages installed:
```bash
# Install all dependencies including test packages
pip install -r requirements.txt
# Or install just the test dependencies
pip install pytest pytest-cov pytest-mock
pip install -r requirements.txt -r requirements-test.txt
```
### 2. Set Environment Variables
@@ -55,28 +52,26 @@ pytest test/test_display_controller.py test/test_plugin_system.py
```bash
# Run a specific test class
pytest test/test_display_controller.py::TestDisplayControllerModeRotation
pytest test/test_display_controller.py::TestDisplayControllerLivePriority
# Run a specific test function
pytest test/test_display_controller.py::TestDisplayControllerModeRotation::test_basic_rotation
pytest test/test_display_controller.py::TestDisplayControllerSchedule::test_active_hours
```
### Run Tests by Marker
The tests use markers to categorize them:
`pytest.ini` declares the markers `unit`, `integration`, `hardware`, `slow`
and `plugin` (with `--strict-markers`, so a typo in a marker name is an
error). Few tests are marked: only a handful carry `unit`, and none currently
carry `integration`, `slow` or `hardware`, so `-m integration` and `-m slow`
select nothing. Select tests by file, directory or `-k` instead.
```bash
# Run only unit tests (fast, isolated)
pytest -m unit
# What CI runs for the core suites (excludes anything marked hardware)
pytest -m "not hardware" test/ --ignore=test/plugins
# Run only integration tests
pytest -m integration
# Run tests that don't require hardware
pytest -m "not hardware"
# Run slow tests
pytest -m slow
# Tests whose name matches an expression
pytest -k "config and not secrets"
```
### Run Tests in a Directory
@@ -103,7 +98,7 @@ When you run `pytest`, you'll see:
```
test/test_display_controller.py::TestDisplayControllerInitialization::test_init_success PASSED
test/test_display_controller.py::TestDisplayControllerModeRotation::test_basic_rotation PASSED
test/test_display_controller.py::TestDisplayControllerOnDemand::test_activate_on_demand PASSED
...
```
@@ -141,71 +136,73 @@ pytest -sv
## Coverage Reports
The test suite is configured to generate coverage reports.
### View Coverage in Terminal
Coverage is not collected by a plain `pytest` run: `pytest.ini` deliberately
has no coverage flags, so local runs stay fast. Ask for it explicitly
(needs `pytest-cov`, which is in `requirements-test.txt`):
```bash
# Coverage is automatically shown when running pytest
pytest
# Terminal summary
pytest --cov=src --cov=web_interface --cov-report=term test/ --ignore=test/plugins
# The output will show something like:
# ----------- coverage: platform linux, python 3.11.5 -----------
# Name Stmts Miss Cover Missing
# ---------------------------------------------------------------------
# src/display_controller.py 450 120 73% 45-67, 89-102
# HTML report in htmlcov/
pytest --cov=src --cov=web_interface --cov-report=html test/ --ignore=test/plugins
```
### Generate HTML Coverage Report
```bash
# HTML report is automatically generated in htmlcov/
pytest
# Then open the report in your browser
# On Linux:
xdg-open htmlcov/index.html
# On macOS:
open htmlcov/index.html
# On Windows:
start htmlcov/index.html
```
The HTML report shows:
- Line-by-line coverage
- Files with low coverage highlighted
- Interactive navigation
Then open `htmlcov/index.html` in your browser (`xdg-open` on Linux, `open`
on macOS, `start` on Windows).
### Coverage Threshold
The tests are configured to fail if coverage drops below 30%. To change this, edit `pytest.ini`:
```ini
--cov-fail-under=30 # Change this value
```
The only threshold is in CI: the core unit-test job in
[`.github/workflows/test.yml`](../.github/workflows/test.yml) runs with
`--cov-fail-under=52`. To check it locally, add that flag to the command
above.
## Common Test Scenarios
### Run Tests After Making Changes
```bash
# Quick test run (just unit tests)
pytest -m unit
# Quick run: just the tests for the area you changed
pytest test/test_config_manager.py
# Full test suite
pytest
```
### Web UI JavaScript Tests
The suites in `test/js` need node; the DOM ones also need jsdom and a running
web interface (details in [`test/js/README.md`](../test/js/README.md)):
```bash
npm install --no-audit --no-fund --prefix test/js # jsdom; node_modules is gitignored
EMULATOR=true python3 web_interface/app.py # in another shell
BASE=http://localhost:5000 REQUIRE_DOM=1 node test/js/run_all.js
```
`pytest test/test_js_unit_suites.py` runs just the unit suites.
### Plugin Config Form Parity
`test/test_field_model_parity.py` checks `build_field_model` against the
`render_field` macro for every plugin schema it finds
([WEB_FRONTEND_ARCHITECTURE.md](WEB_FRONTEND_ARCHITECTURE.md)). It always
covers `plugin-repos/` and the test fixtures; point it at a checkout of the
official plugins to cover those too:
```bash
LEDMATRIX_MONOREPO_PLUGINS=../ledmatrix-plugins/plugins pytest test/test_field_model_parity.py
```
### Debug a Failing Test
```bash
# Run with maximum verbosity and show print statements
pytest -vv -s test/test_display_controller.py::TestDisplayControllerModeRotation::test_basic_rotation
pytest -vv -s test/test_display_controller.py::TestDisplayControllerSchedule::test_active_hours
# Run with Python debugger (pdb)
pytest --pdb test/test_display_controller.py::TestDisplayControllerModeRotation::test_basic_rotation
pytest --pdb test/test_display_controller.py::TestDisplayControllerSchedule::test_active_hours
```
### Run Tests in Parallel (Faster)
@@ -248,22 +245,15 @@ test/
├── test_config_service.py # Config service tests
├── test_config_validation_edge_cases.py # Config edge cases
├── test_font_manager.py # Font manager tests
├── test_layout_manager.py # Layout manager tests
├── test_text_helper.py # Text helper tests
├── test_error_handling.py # Error handling tests
├── test_error_aggregator.py # Error aggregation tests
├── test_schema_manager.py # Schema manager tests
├── test_web_api.py # Web API tests
├── test_nba_*.py # NBA-specific test suites
├── plugins/ # Per-plugin test suites
│ ├── test_clock_simple.py
│ ├── test_calendar.py
│ ├── test_basketball_scoreboard.py
│ ├── test_soccer_scoreboard.py
│ ├── test_odds_ticker.py
│ ├── test_text_display.py
│ ├── test_visual_rendering.py
│ └── test_plugin_base.py
├── plugins/ # Plugin rendering suites
│ ├── test_plugin_matrix.py # Every discovered plugin, across panel sizes
│ ├── test_harness.py
│ └── test_visual_rendering.py
└── web_interface/
├── test_config_manager_atomic.py
├── test_state_reconciliation.py
@@ -288,8 +278,8 @@ test/
If you see import errors:
```bash
# Make sure you're in the project root
cd /home/chuck/Github/LEDMatrix
# Make sure you're in the project root (wherever you cloned it)
cd ~/LEDMatrix
# Check Python path
python -c "import sys; print(sys.path)"
@@ -304,7 +294,7 @@ If tests fail due to missing packages:
```bash
# Install all dependencies
pip install -r requirements.txt
pip install -r requirements.txt -r requirements-test.txt
# Or install specific missing package
pip install <package-name>
@@ -330,37 +320,37 @@ If coverage reports aren't generating:
# Make sure pytest-cov is installed
pip install pytest-cov
# Run with explicit coverage
pytest --cov=src --cov-report=html
# Coverage is opt-in; ask for it explicitly
pytest --cov=src --cov=web_interface --cov-report=html
```
## Continuous Integration
The repo runs
[`.github/workflows/security-audit.yml`](../.github/workflows/security-audit.yml)
(bandit + semgrep) on every push. A pytest CI workflow at
`.github/workflows/tests.yml` is queued to land alongside this
PR ([ChuckBuilds/LEDMatrix#307](https://github.com/ChuckBuilds/LEDMatrix/pull/307));
the workflow file itself was held back from that PR because the
push token lacked the GitHub `workflow` scope, so it needs to be
committed separately by a maintainer. Once it's in, this section
will be updated to describe what the job runs.
The repo runs the pytest suite via
[`.github/workflows/test.yml`](../.github/workflows/test.yml) on every
push and pull request: a plugin-safety job that runs `test/plugins/`, and a
core unit-test job that runs the whole `test/` tree except `test/plugins/`
with `-m "not hardware"` and enforces coverage (`--cov-fail-under=52`). New
test files are picked up automatically. Release version consistency is checked by
[`.github/workflows/release-version-check.yml`](../.github/workflows/release-version-check.yml).
Bandit, flake8, mypy and gitleaks run as pre-commit hooks (see
`.pre-commit-config.yaml`), not in CI.
## Best Practices
1. **Run tests before committing**:
```bash
pytest -m unit # Quick check
pytest test/test_<area>.py # Quick check of what you touched
```
2. **Run full suite before pushing**:
```bash
pytest # Full test suite with coverage
pytest # Full test suite (add --cov flags for coverage)
```
3. **Fix failing tests immediately** - Don't let them accumulate
4. **Keep coverage above threshold** - Aim for 70%+ coverage
4. **Keep coverage above threshold** - CI fails below 52%
5. **Write tests for new features** - Add tests when adding new functionality
@@ -368,9 +358,9 @@ will be updated to describe what the job runs.
```bash
# Most common commands
pytest # Run all tests with coverage
pytest # Run all tests (no coverage)
pytest -v # Verbose output
pytest -m unit # Run only unit tests
pytest test/test_x.py # Run one file
pytest -k "test_name" # Run tests matching pattern
pytest --cov=src # Generate coverage report
pytest -x # Stop on first failure
+250
View File
@@ -0,0 +1,250 @@
# Control socket (web → display)
The display process serves a Unix socket that the web interface uses to send
it commands and get an answer back. It replaces the cache-file "mailboxes" on
the SD card one command at a time. Stage 1, described here, carries on-demand
start/stop/status. The file mailbox stays as a fallback for one release.
| | |
|---|---|
| Socket | `/run/ledmatrix/control.sock` (tmpfs) |
| Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` |
| Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop` |
| Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it |
| Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it |
## Why
Before the socket, the web interface sent commands by writing a cache key
(`display_on_demand_request`) that the display read every 0.25 s.
- **No acknowledgement.** The route answered "success" once the file was
written, whether or not a display was running to read it.
- **Lost requests.** The display had to read the request and then delete it.
A request written between those two steps could be thrown away (see
`_consume_on_demand_request`). The cache has no atomic claim to prevent it.
- **Fragile.** Each channel repeated its own permission, atomic-write,
staleness and in-memory-cache rules. Two of them caused bugs: a `memory_ttl`
bug ignored every on-demand request after the first for an hour, and a
stopped display was still reported as "active" for two minutes.
The socket answers every command, carries one request per message (so nothing
can overwrite it), and belongs to the display process. If the display is not
running, the socket does not exist, and the web interface knows right away.
## Protocol (version 1)
**Framing.** One JSON object per line (newline-delimited JSON), UTF-8, at
most 64 KiB per line (`MAX_MESSAGE_BYTES`). Senders encode with
`ensure_ascii`, so a newline never appears inside a message. A connection
can carry several requests. Each request gets exactly one response, in order.
**Request**
```json
{"v": 1, "id": "5f0c…", "cmd": "on_demand.start",
"args": {"plugin_id": "clock", "mode": null, "duration": 30, "pinned": false}}
```
- `v` is the protocol version.
- `id` is a printable string of 1-128 characters. It is echoed back in the
response, and for on-demand commands it is also the on-demand `request_id`.
- `cmd` is a command name.
- `args` is an object. It may be omitted when a command takes no arguments.
**Response**
```json
{"v": 1, "id": "5f0c…", "ok": true, "result": {"accepted": true, "request_id": "5f0c…", "queued": 1}}
{"v": 1, "id": "5f0c…", "ok": false, "error": {"code": "busy", "message": "…"}}
```
`id` is `null` only when the request could not be parsed far enough to have
one. Clients branch on `error.code`, never on the message text.
**Commands**
| `cmd` | `args` | `result` | Kind |
|---|---|---|---|
| `hello` | `{versions: [int], client?: str}` | `{version, versions, commands, max_message_bytes, server}` | answered directly |
| `ping` | — | `{pong: true}` | answered directly |
| `on_demand.start` | `{plugin_id?, mode?, duration?, pinned?}` (at least one of `plugin_id` and `mode`) | ack | queued |
| `on_demand.stop` | — | ack | queued |
| `on_demand.status` | — | `{on_demand: {...}, current_mode, display_active}` | answered directly |
`duration` is a number of seconds, or a numeric string. `0`, `null` or `""`
mean "until stopped". `pinned` must be a real boolean: the REST route has
already converted strings like `"false"` before it sends the command. The
`on_demand` object in `on_demand.status` is the same dict the display
publishes to `display_on_demand_state`.
**Acknowledgements.** A queued command is *accepted*, not *done*.
`{"accepted": true, "request_id": …}` means the command is waiting in the
render thread's queue, and the render thread will apply it at its next
on-demand check. That is within one frame on a scrolling screen, 0.25 s
during a dwell, and up to 1 s on a static screen, whose frame loop sleeps a
second between frames. Except on a scrolling screen, where the mailbox waits
up to 0.25 s, these are the mailbox's delays too: stage 1 adds
acknowledgements, not speed. Any outcome is published as before
(`display_on_demand_state`, and `status`/`error` for a bad plugin or mode),
and it can be read with `on_demand.status`.
**Versions.** Every request carries `v`. For any command except `hello`, a
`v` the display does not speak gets `unsupported_version`. `hello` is checked
by its `versions` list instead, and its result names the highest version both
sides share, so a client can find out what a display supports before it
relies on anything newer. Stage 1's client sends `v: 1` and falls back to the
mailbox when the display refuses it. It does not send `hello` first, which
saves a round trip.
**Error codes:** `bad_json`, `bad_request`, `message_too_large`,
`unsupported_version`, `unknown_command`, `invalid_args`, `busy` (queue full,
or too many connections), `forbidden` (peer credentials refused), `internal`.
Try it on a device:
```bash
python3 - <<'EOF'
from src.ipc import client # run from the project directory
print(client.on_demand_status())
EOF
```
## How the display applies a command
The server's threads never touch rendering. A connection thread parses the
request, validates it against the contract, and then does one of two things:
- For a command that changes the panel, it puts a `QueuedCommand` on a
bounded queue (16 entries) and answers with the ack.
- For a query, it answers from a status snapshot the display provides
(`DisplayController._control_status`). The snapshot only reads attributes.
The render thread drains the queue in `_poll_on_demand_requests()`, the same
place it reads the mailbox, and hands each command to
`_handle_on_demand_request()`, which is the mailbox's own handler. The two
paths share all of their code: activation, the processed-id guard, error
publishing, and resuming the rotation afterwards. The 0.25 s floor on the
mailbox read does not apply to the queue, because draining it costs no disk
read. A queued command also lets `_service_pending_changes()` skip its own
floor, so a long scrolling screen or a Vegas iteration takes the command at
its next frame.
**Exactly once.** A command and a mailbox write for the same request share
one `request_id`. If the client times out after the display queued the
command and then also writes the mailbox, the display processes the request
once. The existing `on_demand_request_id` and processed-id checks drop the
second copy.
## Robustness
All of this runs inside the display process, so nothing a client does may
block the render loop or crash it:
- **Bounded connections.** Each connection gets its own daemon thread, with
at most 8 at once. One more is answered `busy` and closed.
- **Timeouts.** Each read and write times out after 2 s. A message must
arrive whole within 5 s of its first byte. An idle connection is closed
after 10 s. A slow or stuck client costs one thread for a few seconds.
- **Malformed input.** A line that is not JSON gets `bad_json`, and the
connection carries on. A line longer than 64 KiB gets `message_too_large`,
and the connection is closed, because the next message boundary cannot be
found. A client that disconnects mid-message is dropped silently. No
exception from a handler leaves the connection thread.
- **Full queue.** When the queue is full, the client gets `busy` and falls
back to the mailbox. A full queue means the render thread is stuck, and the
systemd watchdog deals with that.
- **Startup.** The server binds under a temporary name, sets the mode and the
group, then renames the socket into place, so it never appears with the
umask's permissions. It removes a stale socket (a file that nothing is
listening on). It never removes a live socket or a file that is not a
socket. `close()` removes the socket only if it is still the one this
process created.
- **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind
failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs
as before. The web interface then uses the mailbox.
## Security model
The display runs as root and the web interface as the installing user (see
[PERMISSIONS.md](PERMISSIONS.md)). The socket admits exactly those two, plus
anything else in the group they share:
1. **The directory.** `/run/ledmatrix` is created by `RuntimeDirectory=ledmatrix`
in `ledmatrix.service` (#687): root-owned, `0755`, on tmpfs, and removed
when the display stops. Under an older unit, the display creates the
directory itself as root, as it does for the heartbeat. No installer
change is needed.
2. **The socket file.** The file is `root:<shared group>` with mode `0660`,
and the kernel refuses `connect()` to anyone without write permission on
it. The shared group is the cache directory's group whenever that
directory is group-writable. That is `ledmatrix` on an installed device
(`/var/cache/ledmatrix` is `root:ledmatrix 2775`), and it is the same rule
DiskCache uses for every file the two services share. Otherwise the group
is the project directory's (`get_shared_group_gid()`, which config files
use). With neither, the mode is `0600` and only root can connect.
3. **Peer credentials.** Where the kernel reports them (`SO_PEERCRED`, on
Linux), the server checks every connection again. It accepts root, the
display's own user, or a member of the shared group: the peer's primary
gid, or a supplementary group read from `/proc/<pid>/status`. If `/proc`
is unreadable, it uses the group database. Any other peer gets `forbidden`
and is disconnected. This covers a socket mode that someone loosened by
hand.
The commands are deliberately narrow. Stage 1 can start or stop on-demand
display and read its state, which anyone who can reach the web UI can already
do. Nothing on the socket runs a shell, writes a file, or names a path.
**Development.** A display that is not root and cannot write to
`/run/ledmatrix`, such as `python3 run.py -e` from a checkout, serves the
socket at `$TMPDIR/ledmatrix-<uid>/control.sock`. That directory is private
(`0700`), and the server refuses it if another user owns it. The web
interface, run by the same user, looks there after `/run/ledmatrix`. The test
suite sets `LEDMATRIX_CONTROL_SOCKET=off` (`test/conftest.py`), so a run on a
device never touches the live display.
## Stage plan
1. **On-demand, with acks (this stage).** Contract, server, client.
`on_demand.start`/`stop`/`status`, `hello`, `ping`. The REST routes try the
socket first and report `transport: "socket" | "mailbox"` (plus
`socket_error` on fallback). The mailbox is unchanged, and the plugins that
write it directly (birdnet-go, mqtt-notifications, on-air, pomodoro-timer)
keep working.
2. **Commands that are restarts or polls today.**
- `brightness.set`, transient and with no `config.json` write.
- `plugin.reload`, which replaces the `restart_required` answer from #688
with a live reload of the updated plugin on the render thread.
- `config.reload`, which applies a saved config without waiting for the 2 s
mtime poll and acks which sections changed.
- The dwell sleep and the static screen's 1 s frame sleep wait on the
queue instead of sleeping, so a command lands within milliseconds on
every kind of screen. Under WSL, with a static plugin on screen, a stop
takes 1.0 s by either path today.
3. **A state stream.** A `subscribe` command that keeps the connection open
and pushes events: mode changes, on-demand state, plugin runtime state and
the heartbeat. It replaces the polled `display_current_state`,
`plugin_runtime_snapshot` (#690) and `display-heartbeat.json` (#687) for
readers that hold a connection. The web interface relays it to its
existing SSE stream. The files remain for one release for older readers.
4. **Retire the mailboxes.** After a release in which every device has had the
socket, the web interface stops writing `display_on_demand_request`, and
the display stops polling it, logging the plugins that still write it so
they can move to an in-process `request_display()`. The other cache keys
used as messages (`plugin_error_clear_request` and the remaining
`display_*` keys) move to the socket or to tmpfs.
## Checking it on a device
```bash
ls -l /run/ledmatrix/control.sock # srw-rw---- root ledmatrix
sudo journalctl -u ledmatrix | grep "Control socket"
curl -s -X POST localhost:5000/api/v3/display/on-demand/start \
-H 'Content-Type: application/json' -d '{"plugin_id":"clock","duration":20}'
# ... "transport": "socket"
```
If the response says `"transport": "mailbox"`, `socket_error` gives the
reason. `no_socket` means the display is stopped or predates the socket.
`refused` usually means the web user is not in the socket's group, which
takes effect when the web service restarts after the user is added.
+115
View File
@@ -0,0 +1,115 @@
# Running on Low-Memory Boards
Applies to the Pi Zero 2 W (512 MB), Pi 3 / 3B+ (1 GB), and the 1 GB Pi 4.
If your board has 2 GB or more you can skip this document.
## The failure this prevents
The display process is the largest thing on the board. On a 1 GB Pi 3B+ with
around 20 plugins enabled it settles near **600 MB of 905 MB usable**, leaving
under 200 MB of headroom for everything else.
When that headroom runs out, the board does not crash cleanly. `fork()` starts
failing, and because a new process is needed to do almost anything, the
symptoms look nothing like "out of memory":
| What you see | Why |
|---|---|
| SSH accepts the connection then closes it instantly, before any banner | `sshd` forks a session per connection; the fork fails |
| The web UI still responds quickly | Already running, serves from existing threads, forks nothing |
| Ping is perfect, 0% loss | Handled entirely in the kernel |
| The panel is dark | The display process was killed and cannot be respawned |
| The clock is wrong after the next boot | `fake-hwclock`'s periodic save is a scheduled job, and it cannot fork either |
The board looks healthy from the outside and cannot be logged into. Only a
power cycle clears it. If you are here because SSH stopped working, also see
[SSH_UNAVAILABLE_AFTER_INSTALL.md](SSH_UNAVAILABLE_AFTER_INSTALL.md), which
covers the more common cause (AP mode).
## Check your headroom
```bash
free -m
ps -eo rss,comm --sort=-rss | head -5
```
If `MemAvailable` is under ~150 MB while the display is running, you are close
to the edge. To watch it over time:
```bash
watch -n 30 'free -m | head -2'
```
Available memory that falls steadily rather than holding flat means you will
reach the wall; it is a question of when.
## What to do
**1. Enable the memory cgroup controller.** Without it, the `MemoryMax=85%` in
`systemd/ledmatrix.service` is accepted by systemd and silently ignored, so the
service has no ceiling and a runaway takes the whole board down instead of just
restarting. Raspberry Pi firmware disables this controller by default.
`first_time_install.sh` does this for you. To check it took effect:
```bash
grep memory /sys/fs/cgroup/cgroup.controllers
```
If that prints nothing, add `cgroup_enable=memory cgroup_memory=1` to the
kernel command line and reboot. Edit whichever file your image uses —
`/boot/firmware/cmdline.txt` on current Raspberry Pi OS, `/boot/cmdline.txt` on
older layouts (the installer checks the first and falls back to the second).
Everything must stay on a single line.
This changes the failure mode from "the board becomes unreachable" to "the
display service restarts". It is a safety net, not a fix.
**2. Run fewer plugins.** This is the actual remedy. Every enabled plugin costs
memory permanently — its module, its parsed config, and its cached API
responses. On a 512 MB or 1 GB board, keep the enabled set small and prefer
plugins that poll infrequently.
**3. Lower the cache ceiling.** The in-memory cache is sized from total RAM
(150 entries at 1 GB and below, up to 1500 at 8 GB). To go lower still:
```ini
# /etc/systemd/system/ledmatrix.service.d/override.conf
[Service]
Environment=LEDMATRIX_CACHE_MAX_ENTRIES=75
```
Writing the file does not change the running service. Reload systemd and
restart it:
```bash
sudo systemctl daemon-reload
sudo systemctl restart ledmatrix
```
Fewer entries means more API calls, so lower this only while you are actually
short of memory.
**4. Consider `MemoryHigh`.** `MemoryMax` kills and restarts. `MemoryHigh`
throttles and reclaims instead, which is gentler — but on a board where the
process genuinely wants more than the limit, sustained reclaim can stall the
render loop and show as visible stutter on the panel. Add it only if you prefer
degraded output to a restart:
```ini
[Service]
MemoryHigh=70%
```
## Keep your logs
These images default to volatile journald storage, so every reboot destroys the
logs — including the ones explaining why the board rebooted. `first_time_install.sh`
enables persistent storage capped at 64 MB. To confirm:
```bash
journalctl --list-boots
```
More than one boot listed means logs are surviving reboots. If only one is
listed, journald is still writing to `/run` (tmpfs).
+6 -4
View File
@@ -19,7 +19,6 @@ All installation scripts have been moved from the project root to `scripts/insta
| `install_wifi_monitor.sh` | `scripts/install/install_wifi_monitor.sh` |
| `setup_cache.sh` | `scripts/install/setup_cache.sh` |
| `configure_web_sudo.sh` | `scripts/install/configure_web_sudo.sh` |
| `migrate_config.sh` | `scripts/install/migrate_config.sh` |
#### Permission Fix Scripts
@@ -59,9 +58,12 @@ sudo ./scripts/install/install_service.sh
After updating your scripts, verify they still work:
```bash
# Test installation scripts (if needed)
# Check the installation scripts are at their new paths
ls scripts/install/*.sh
sudo ./scripts/install/install_service.sh --help
./scripts/install/install_service.sh --help # prints usage only
# Note: running install_service.sh for real (with sudo, no --help)
# reinstalls, enables and restarts ledmatrix.service, ledmatrix-web.service
# and the update-verify units.
# Test permission scripts
ls scripts/fix_perms/*.sh
@@ -86,7 +88,7 @@ The plugin system has been enhanced but remains backward compatible with existin
If you encounter issues during migration:
1. Check the [README.md](README.md) for current installation and usage instructions
1. Check the [project root README](../README.md) for current installation and usage instructions
2. Review script README files:
- [`scripts/install/README.md`](../scripts/install/README.md) - Installation scripts documentation
- [`scripts/fix_perms/README.md`](../scripts/fix_perms/README.md) - Permission scripts documentation
+91 -111
View File
@@ -1,169 +1,149 @@
# Multi-Root Workspace Setup Guide
This document explains how the LEDMatrix project uses a multi-root workspace to manage plugins as separate Git repositories.
This document explains how to work on LEDMatrix and the official plugins side
by side, with one editor workspace and the plugins loaded straight from your
plugin checkout.
## Overview
The LEDMatrix project has been migrated from a git submodule implementation to a **multi-root workspace** implementation for managing plugins. This allows:
Official plugins live in a single repository,
[ledmatrix-plugins](https://github.com/ChuckBuilds/ledmatrix-plugins), with one
directory per plugin under `plugins/`. There are no separate per-plugin
repositories. For development you clone that monorepo **next to** LEDMatrix
and symlink the plugin directories you are working on into LEDMatrix's
`plugins/` directory with `scripts/dev/dev_plugin_setup.sh`.
- ✅ Plugins to exist as independent Git repositories
- ✅ Updates to plugins without modifying the LEDMatrix project
- ✅ Easy development workflow with all repos in one workspace
- ✅ Plugin system discovers plugins via symlinks in `plugin-repos/`
- ✅ Plugin code stays in the monorepo checkout, with its own git history
- ✅ LEDMatrix discovers the plugins through symlinks in `plugins/`
(git-ignored), so the production `plugin-repos/` directory is untouched
- ✅ `LEDMatrix.code-workspace` opens both repositories in VS Code/Cursor
## Directory Structure
```text
/home/chuck/Github/
├── LEDMatrix/ # Main project
│ ├── plugin-repos/ # Symlinks to actual repos (managed automatically)
│ │ ├── ledmatrix-clock-simple -> ../../ledmatrix-clock-simple
│ │ ├── ledmatrix-weather -> ../../ledmatrix-weather
~/Github/
├── LEDMatrix/ # Main project
│ ├── plugins/ # Dev plugin directory (git-ignored)
│ │ ├── clock-simple -> ~/Github/ledmatrix-plugins/plugins/clock-simple
│ │ ├── ledmatrix-weather -> ~/Github/ledmatrix-plugins/plugins/ledmatrix-weather
│ │ └── ...
│ ├── LEDMatrix.code-workspace # Multi-root workspace configuration
│ ├── plugin-repos/ # Default (Plugin Store) plugin directory
│ ├── LEDMatrix.code-workspace # Opens LEDMatrix and ../ledmatrix-plugins
│ └── ...
├── ledmatrix-clock-simple/ # Plugin repository (actual git repo)
├── ledmatrix-weather/ # Plugin repository (actual git repo)
├── ledmatrix-football-scoreboard/ # Plugin repository (actual git repo)
└── ... # Other plugin repos
└── ledmatrix-plugins/ # Plugin monorepo (git repo)
├── plugins/
│ ├── clock-simple/
│ ├── ledmatrix-weather/
│ └── ...
├── plugins.json # Store registry
└── update_registry.py
```
## How It Works
### 1. Plugin Repositories
### 1. The plugin monorepo
All plugin repositories are cloned to `/home/chuck/Github/` (parent directory of LEDMatrix) as regular Git repositories:
- `ledmatrix-clock-simple/`
- `ledmatrix-weather/`
- `ledmatrix-football-scoreboard/`
- etc.
### 2. Symlinks in plugin-repos/
The `LEDMatrix/plugin-repos/` directory contains symlinks pointing to the actual repositories in the parent directory. This allows the plugin system to discover plugins without modifying the project structure.
### 3. Multi-Root Workspace
The `LEDMatrix.code-workspace` file configures VS Code/Cursor to open all plugin repositories as separate workspace roots, allowing easy development across all repos.
## Setup Scripts
### Initial Setup
If you already have plugin repositories cloned, use the setup script:
Clone ledmatrix-plugins into the same parent directory as LEDMatrix (the
workspace file and `scripts/update_plugin_repos.py` look for
`../ledmatrix-plugins` relative to the LEDMatrix root):
```bash
cd /home/chuck/Github/LEDMatrix
python3 scripts/setup_plugin_repos.py
cd ~/Github
git clone https://github.com/ChuckBuilds/ledmatrix-plugins.git
```
This script:
- Reads the workspace configuration
- Creates symlinks in `plugin-repos/` pointing to actual repos
- Verifies all links are created correctly
### 2. Symlinks in plugins/
`scripts/dev/dev_plugin_setup.sh link <name> <path>` creates
`LEDMatrix/plugins/<name>` as a symlink to a plugin directory. Use the
plugin's manifest `id` as the name: that is the name the loader and
`config.json` use, and the script warns when the two differ.
### 3. Multi-root workspace
`LEDMatrix.code-workspace` has two roots: LEDMatrix itself and
`../ledmatrix-plugins`.
## Setup
### Link plugins
```bash
cd ~/Github/LEDMatrix
./scripts/dev/dev_plugin_setup.sh link clock-simple ../ledmatrix-plugins/plugins/clock-simple
./scripts/dev/dev_plugin_setup.sh list # show what is linked
```
If a real (non-symlink) directory of the same name already exists in
`plugins/`, the script offers to back it up and replace it.
Without a sibling checkout, `./scripts/dev/dev_plugin_setup.sh link-github
<name>` clones the monorepo into `~/.ledmatrix-dev-plugins/` instead and links
the plugin from there. See the
[Plugin Development Guide](PLUGIN_DEVELOPMENT_GUIDE.md).
### Updating Plugins
To update all plugin repositories:
```bash
cd /home/chuck/Github/LEDMatrix
python3 scripts/update_plugin_repos.py
cd ~/Github/LEDMatrix
python3 scripts/update_plugin_repos.py # git pull in ../ledmatrix-plugins
# or
./scripts/dev/dev_plugin_setup.sh update # git pull in every linked checkout
```
This script:
- Finds all plugins in the workspace
- Runs `git pull` on each repository
- Reports which plugins were updated
The symlinks pick up the new code; restart the display to load it.
## Configuration
The plugin system is configured in `config/config.json`:
The loader scans only `plugin_system.plugins_directory` in
`config/config.json` (default `plugin-repos`). Point it at `plugins` so it
finds the links:
```json
{
"plugin_system": {
"plugins_directory": "plugin-repos",
"auto_discover": true,
"auto_load_enabled": true
"plugins_directory": "plugins"
}
}
```
The `plugins_directory` points to `plugin-repos/`, which contains symlinks to the actual repositories.
## Workflow
### Daily Development
1. **Open Workspace**: Open `LEDMatrix.code-workspace` in VS Code/Cursor
2. **All Repos Available**: All plugin repos appear as separate folders in the workspace
3. **Edit Plugins**: Edit plugin code directly in their repositories
4. **Update Plugins**: Run `update_plugin_repos.py` to pull latest changes
2. **Edit Plugins**: Edit code under `ledmatrix-plugins/plugins/<plugin>/`
3. **Test**: `python3 run.py -e` (emulator) or
`python3 scripts/check_plugin.py --plugin <id>` from LEDMatrix
4. **Ship**: Bump `version` in the plugin's `manifest.json`, run
`python update_registry.py` in ledmatrix-plugins, commit there
### Adding New Plugins
1. **Clone Repository**: Clone the new plugin repo to `/home/chuck/Github/`
2. **Add to Workspace**: Add the plugin folder to `LEDMatrix.code-workspace`
3. **Create Symlink**: Run `setup_plugin_repos.py` to create the symlink
### Updating Individual Plugins
Since plugins are regular Git repositories, you can update them individually:
```bash
cd /home/chuck/Github/ledmatrix-weather
git pull origin master
```
Or update all at once:
```bash
cd /home/chuck/Github/LEDMatrix
python3 scripts/update_plugin_repos.py
```
## Benefits
1. **No Submodule Hassle**: No need to update `.gitmodules` or run `git submodule update`
2. **Independent Updates**: Update plugins independently without touching LEDMatrix
3. **Clean Separation**: Each plugin is a separate repository with its own history
4. **Easy Development**: Multi-root workspace makes it easy to work across repos
5. **Automatic Discovery**: Plugin system automatically discovers plugins via symlinks
1. Create `plugins/<your-plugin-id>/` in the monorepo checkout
2. Link it: `./scripts/dev/dev_plugin_setup.sh link <your-plugin-id> ../ledmatrix-plugins/plugins/<your-plugin-id>`
## Troubleshooting
### Symlinks Not Working
If plugins aren't being discovered:
### Plugins not discovered
```bash
cd /home/chuck/Github/LEDMatrix
python3 scripts/setup_plugin_repos.py
cd ~/Github/LEDMatrix
ls -la plugins/ # links present and not broken?
./scripts/dev/dev_plugin_setup.sh status # link targets and git state
```
This will recreate all symlinks.
Also check that `plugin_system.plugins_directory` is `plugins`.
### Missing Plugins
### Plugin updates not showing
If a plugin is in the workspace but not found:
1. Check if the repo exists in `/home/chuck/Github/`
2. Check if the symlink exists in `plugin-repos/`
3. Run `setup_plugin_repos.py` to recreate symlinks
### Plugin Updates Not Showing
If changes to plugins aren't appearing:
1. Verify the symlink points to the correct directory: `ls -la plugin-repos/ledmatrix-weather`
2. Check that you're editing in the actual repo, not a copy
3. Restart the LEDMatrix service if running
1. Verify the link target: `ls -la plugins/<id>`
2. Check that you're editing the monorepo checkout, not a store-installed copy
3. Restart the LEDMatrix service (or `run.py`)
## Notes
- The `plugin-repos/` directory is tracked in git, but only contains symlinks
- Actual plugin code lives in `/home/chuck/Github/ledmatrix-*/`
- Each plugin repo can be updated independently via `git pull`
- The LEDMatrix project doesn't need to be updated when plugins change
- `plugins/` is git-ignored (except `plugins/.gitkeep`); the symlinks are
never committed.
- When changing a plugin in the monorepo, bump its manifest `version` and run
`python update_registry.py`, or users won't receive the update.
+354
View File
@@ -0,0 +1,354 @@
# Offscreen Rendering
**Status (2026-09-30):** offscreen rendering is implemented
(`DisplayManager.offscreen()`, the adapter on the prefetch thread, the plugin
lock), and so are live elements, which grew out of steps 2 and 3 below: see
*Live elements*. The segment strip proposed as step 2 was not needed; *Why not
a SegmentStrip* says why.
First soak of step 1 on hdpi (50 px/s, `pwm_bits` 7, preview open, 8-minute
runs, A/B/B/A):
| build | late | by 1 | 2 | 3–5 | 6+ | freezes | render-thread fetches |
|---|---|---|---|---|---|---|---|
| #628 | 0.53% | 82 | 2 | 3 | 2 | 3 | 6 |
| step 1 | 0.63% | 78 | 63 | 17 | 2 | 1 | 0 |
| step 1 | 0.42% | 77 | 23 | 10 | 0 | 0 | 0 |
| #628 | 0.37% | 84 | 5 | 3 | 3 | 2 | 14 |
It does what it was built to: no plugin is fetched on the render thread, and
freezes fell from 5 to 1. But frames 2–5 refreshes late rose. The rendering
moved to the prefetch thread still needs the GIL, and the render thread waits
for it (risk 5 below). The late rate did not improve overall. The 1–2 s
freezes appear in both builds and have a separate, not yet identified cause.
The GIL fix, measured on hdpi (90 px/s, `pwm_bits` 8, preview open, 8-minute
runs after a 2-minute warm-up, order A B C C B A, 2026-09-24). Each arm pools
two runs, about 81,000 frames:
| arm | late | by 1 | 2 | 3–5 | 6+ | 2+ late per 10k frames | freezes |
|---|---|---|---|---|---|---|---|
| A: step 1 as is | 0.90% | 575 | 64 | 91 | 9 | 20.1 | 0 |
| B: `switch_interval_ms` 1 | 0.78% | 510 | 105 | 23 | 2 | 15.8 | 0 |
| C: `prefetch_gate` | **0.60%** | 471 | 11 | 7 | 2 | **2.5** | 0 |
The gate removes the frames the render thread spent waiting for the GIL, and
it costs the prefetch nothing that shows: it parked the thread for 3–6 s per
run, and the next group was ready at every strip extension in every arm.
`prefetch_gate` is therefore on by default; `switch_interval_ms` stays an
off-by-default experiment. What is left is almost all one refresh late, which
is the per-frame budget (a 6.75 ms p50 blit in a refresh the panel holds at
83–85 Hz while rendering), not contention.
The runs restart the service, so the hourly sports refresh never fell inside
one. That refresh is its own case: about twenty ESPN chunk-fetch threads at
once, which the gate does not cover (it gates only the prefetch thread).
## The problem
Vegas mode builds its ticker from every plugin's content. Most of that work
already happens on a background prefetch thread
(`RenderPipeline.start_prefetch`). But any plugin whose content needs the
**shared display canvas** is deferred to the render thread
(`RenderPipeline.drain_deferred`), one plugin every two seconds. The code's
own comments put each of those at 40–600 ms, and the render thread presents no
frames while one runs.
On hdpi (Pi 4, 512×64) most plugins take that path: geochron, tide-display,
news, hockey-scoreboard, ledmatrix-stocks, incoming-packages, clock-simple,
countdown, birdnet-go, ledmatrix-music and odds-ticker. They arrive in bursts
("Whole group deferred; strip will extend as it drains") every minute or so,
12 fetches in five minutes. That is the "occasional pause" a viewer sees.
An 8-minute soak (`scripts/frame_soak.py --preview`) of the #628 build on
hdpi:
| late by | frames |
|---|---|
| 1 refresh | 238 |
| 2 | 32 |
| 3–5 | 30 |
| 6+ | 5 |
| freezes ≥ 250 ms | 2 (0.97 s total) |
The 3+ rows and the freezes are the pauses. The single-refresh row is a
separate problem: the blit is 6 ms of a 10 ms refresh, so there is little
slack. It is covered under *What this does not fix*.
## Why a plugin is canvas-bound
The plugin-facing canvas is a set of shared attributes on `DisplayManager`:
`image`, `draw`, `matrix`, and the `width`/`height` properties that read from
`matrix`. Three adapter paths (`src/vegas_mode/plugin_adapter.py`) need them,
and each returns `None` under `offscreen_only=True` so the plugin is queued for
the render thread:
1. **Display capture** (`_capture_display_content`): clear the canvas, call
`plugin.display()`, copy `display_manager.image`. Used by any plugin
without `get_vegas_content()` or a populated `scroll_helper`.
2. **Scroll-content generation** (`_trigger_scroll_content_generation`): a
ticker plugin whose `scroll_helper.cached_image` is empty is made to build
it by calling `display(force_clear=True)` or `_create_scrolling_display()`.
Both draw on the canvas.
3. **Narrowed rendering** (`DisplayManager.render_size`): swaps the shared
`matrix`, `image` and `draw` for a narrower set so the plugin lays out for
`render_width_pct`. The render thread would see the swap mid-frame.
The render thread keeps the canvas coherent only because nothing else touches
it at the same time. A background thread can't use it.
## The design: a per-thread render target
`capture_mode()` is already per-thread (#423 made its state a
`threading.local`, so a background capture no longer suppresses the render
loop's pushes). The same move applies to the canvas itself:
```python
with display_manager.offscreen(width=None, height=None) as surface:
plugin.display(force_clear=True)
content = surface.image.copy()
```
For the **calling thread only**, inside the block:
| accessor | resolves to |
|---|---|
| `display_manager.image`, `.draw` | the surface's own image and draw: a fresh black canvas, `fontmode = "1"` |
| `display_manager.matrix` | a logical proxy reporting the surface size, so `width`/`height` and plugins that read `matrix.width` follow it. Hardware calls through it (`SetImage`, `SwapOnVSync`, `Clear`, brightness writes) are inert. |
| `update_display()`, `clear()` | canvas-only: the block implies capture mode, which is already per-thread |
| `set_scrolling_state()`, `set_frame_hold()` | no-ops, so a plugin's `display()` cannot re-pace the live scroll. Today it can, when it is captured on the render thread. |
Every other thread sees the real canvas, unchanged. The render loop in
particular keeps presenting while a plugin draws elsewhere.
### Implementation sketch
- `image`, `draw` and `matrix` become properties over `_image`, `_draw` and
`_matrix`, plus a thread-local current surface. The getter returns the
surface's value when the calling thread has one, else the shared one; setters
mirror that. That costs about 0.1 µs per access, and `update_display()` reads
each a handful of times per frame. Every existing `self.image = ...` in
`DisplayManager` (`clear()`, setup, fallback) keeps working and becomes
thread-correct for free.
- `render_size()` is rebuilt on `offscreen()`: it creates or narrows the
calling thread's surface instead of swapping shared state.
- `offscreen()` nests and always restores on exit, including when the plugin
raises.
- `VisualDisplayManager` (the plugin test harness) gets the same method, for
parity.
### Adapter changes
- `get_content(offscreen_only=True)` stops returning `None` for the three
paths above. Each runs inside `display_manager.offscreen(render_width)`.
- `_capture_display_content` and `_trigger_scroll_content_generation` drop
their "copy the shared image, restore it afterwards" bookkeeping, since the
shared image is never touched.
- **Take the plugin's lock.** `PluginManager.get_plugin_lock()` keeps
`update()` and `display()` mutually exclusive in normal rotation, but Vegas
never takes it, so today's render-thread captures already race
`update()`. Off the render thread the adapter can afford to wait: blocking
acquire with a timeout (proposed 2 s). On timeout it keeps the cached segment
and tries again next group.
- `drain_deferred()` and the deferred queue are deleted. The only render-thread
fetch left is the inline fallback when no prepared group is ready, which in
practice is the first extension. Prefetching at start removes that too.
## Live elements: content that changes while it scrolls
Offscreen rendering is also what makes fresh content possible. A plugin's
segment is drawn when its group is prefetched, and the strip carries
7,000-10,000 px of content ahead of the viewport, so at ~100 px/s a score drawn
then reaches the screen 70-100 seconds later -- and once in the strip it never
changed: when a plugin reported new data, Vegas only dropped its caches, so the
change appeared on the plugin's *next* turn, minutes later.
A plugin can now hand Vegas **live elements** instead of pictures
(`BasePlugin.get_vegas_elements()`, see "Live Vegas elements" in
[PLUGIN_API_REFERENCE.md](PLUGIN_API_REFERENCE.md#live-vegas-elements)):
named, fixed-width pieces of content -- one per game card, one for a map. Vegas
records where each lands in the strip and, when the plugin's data changes,
redraws just the changed ones off the render thread and copies their pixels
over the old ones between two frames. A card already crossing the panel
changes; nothing next to it moves.
### Why not redraw every frame
On a Pi the render thread has about 4 ms of slack per refresh at 512x64 after
the ~6 ms blit, and a scoreboard card is ~29 ms of Pillow work that holds the
GIL. Drawing on the render thread is out of the question at any rate, so the
render thread only ever *copies* pixels that are already drawn. Measured on a
Pi 4 (ledpi): writing a 35 KB card into a 20,000 px strip takes 8.5 µs, a
101 KB map 17 µs, four cards (the per-frame cap) 34 µs -- against 124 µs for
the viewport slice every frame already does.
### How an update reaches the screen
1. A plugin's `update()` completes. The update worker calls
`PluginManager._note_update_completed`, which calls the update listeners
(`add_update_listener`) there and then, with the plugin's lock still held.
Vegas's listener moves the plugin's **epoch** on
(`src/vegas_mode/elements.py`, `LiveEpochs`) and wakes the live worker.
2. The **live worker** (`src/vegas_mode/live_worker.py`), the one background
thread that draws for the strip once it holds a live element, finds the
plugin's elements whose recorded epoch is older than its current one,
nearest the screen first, and calls `get_vegas_elements()` under the
plugin's lock (0.25 s wait, then a 1 s backoff). Elements whose `version`
is unchanged cost nothing; the rest are pinned and checksummed, and each
whose pixels changed becomes a patch in a one-per-element slot (the latest
wins).
3. Between two frames the render thread
(`RenderPipeline.apply_live_patches`, from `coordinator.run_frame`) pops at
most four patches or two screens of bytes and copies each into the strip
with `ScrollHelper.patch_columns`. It takes no lock and draws nothing. A
patch made for an older strip, for an element trimmed away or already
behind the screen, or from older data than the strip shows, is dropped.
End to end, a new score reaches a card already on screen within one poll of
the data source (30 s for live games) plus about a second: the listener is
immediate, and while live elements exist the update tick that schedules
plugins runs every second instead of every four.
Elements that change with **time** rather than data (an aircraft moving
between position reports) ask for `refresh_hz`; the worker calls
`redraw_vegas_element()` -- without the plugin's lock, from state the plugin
publishes in one assignment -- that often while the element is on or within
`live_lead_screens` of the screen, capped by `live_max_hz` (5), at 1 Hz
without the render gate, and halved for an element whose redraws average over
50 ms.
### Geometry
A live element is never trimmed to its ink: the adapter pads it with
`content_padding` black columns either side and pins its width, and a redraw
at any other width is refused (it shows the next time the plugin comes round).
Records keep **absolute** strip columns -- the strip column plus everything
trimmed off the front since the strip was composed -- so a trim moves one
origin rather than every record. Nothing on screen is ever moved, inserted or
resized; a game added to a slate appears on the plugin's next turn.
### Why not a SegmentStrip
The proposal here was to replace the single strip with a list of segments.
In-place patching of the single strip meets every goal without that: the
patch is O(element) and the strip layout never changes. What a segment list
would still buy is cheaper extensions, and most of that came from making the
strip's PIL copy lazy instead (`ScrollHelper.cached_image`: an extension used
to rebuild it twice, 1.7-3.8 ms each on a Pi 4). The `extend` row of
`frame_soak.py`'s "after work" table says whether the rest is worth it.
### When it is off
- `display.vegas_scroll.live_refresh: false` (the kill switch; also in the
web UI), or `vegas_live: false` in one plugin's section.
- Always under multi-display sync: the follower mirrors whole strips only, so
a patch would never reach it. (Continuous-mode sync has a separate problem:
the follower is not sent extensions or trims at all.)
- In swap mode (`continuous_scroll: false`) and with `offscreen_prefetch:
false`.
- For plugins without the hook, which are drawn and placed exactly as before,
and on the paths that fetch without the plugin's lock (the first strip of a
run, the render thread's fallback fetch), which use `get_vegas_content()`.
## Risks, and what was checked
1. **Plugins holding their own reference to the shared `draw` or `image`.**
They would keep drawing into the shared canvas, and routing by thread can't
redirect them. A grep of the 49 plugins installed on hdpi found none storing
`display_manager.draw` or `.image` in an attribute (a pattern search, so
indirect aliasing would slip past it). A plugin that did would
draw into an image nobody displays, which trims to a blank segment. That is
not corruption, and it is no worse than today.
2. **Plugins calling the matrix directly.** None in the audit. Inside
`offscreen()` the proxy makes it inert anyway.
3. **Font thread-safety.** `FontManager` shares font objects across plugins.
Measured on Pillow 12.3, two threads rendering text take 1.94× as long as
one, so text rendering holds the GIL and FreeType is never entered
concurrently. Re-check if Pillow changes that.
4. **Plugin thread-safety.** `display()` moves to the prefetch thread. The
plugin lock makes it exclusive with `update()`, which is more protection
than it has today. Threads a plugin starts itself are not covered, as today.
5. **The GIL.** Moving 40–600 ms of plugin rendering off the render thread
removes the pauses, but the work still needs the GIL. Pillow drawing holds
it, and a waiting thread only gets it back after the switch interval
(default 5 ms). Expect some single-refresh late frames while a prefetch
runs. Measure with the soak. A render process separate from plugin work
is the structural answer (the "native presenter" step). Two experiments
get most of the way first (results under Status, above):
- `vegas_scroll.switch_interval_ms` lowers the switch interval for a Vegas
run (1 ms is the obvious try), so the render thread waits at most that
long behind bytecode. It does nothing for a C call that keeps the GIL.
- `vegas_scroll.prefetch_gate` (`src/common/render_gate.py`) lets the
prefetch thread run Python only while the render thread is blocked in
`SwapOnVSync`, up to just before the refresh the swap returns on, and
parks it the rest of the time. That covers C calls too, since the gate is
checked before each one starts. It never parks the thread while it holds
a lock the render thread takes, and never for more than 50 ms. It needs
the rebuilt binding, which releases the GIL during the swap. On by
default.
## What this does not fix
- **The blit.** Copying a 512×64 frame into the matrix (`SetImage`) is ~6 ms at
8 PWM bits on a Pi 4, leaving ~4 ms of slack per refresh. That is the main
source of the single-refresh late frames. Holding frames for two refreshes
(≈50 px/s) doubles the budget. Cutting the blit itself is the native-presenter
step.
- **Multi-display sync in continuous mode.** The follower is sent the whole
strip only at a new cycle and on connect, never the extensions and trims of
continuous mode, so it drifts from the leader after the first extension.
Live elements stay off under sync for that reason.
## Test plan
- **Unit, `DisplayManager`:** one thread inside `offscreen()` draws while
another reads `image`/`draw`/`matrix`/`width`/`height` and sees the real
canvas. Also: `update_display()` and `set_scrolling_state()` are inert inside;
`render_size()` narrows only the calling thread; nesting and exceptions
restore state.
- **Unit, adapter:** a stub display-capture plugin and a stub scroll-helper
plugin both return content with `offscreen_only=True`, and nothing is queued
for the render thread. The plugin lock is taken, and a timeout keeps the cached
segment.
- **Emulator integration:** a stub canvas-bound plugin whose `display()` sleeps
300 ms. The Vegas render loop never goes a frame without presenting (frame
timing recorder: zero freezes).
- **Live elements** (`test/test_vegas_live_*.py`,
`test/test_vegas_elements_*.py`, `test/test_scroll_helper_patch.py`): every
record points at exactly its element's pixels through any sequence of
compose, extend and trim; a patch changes only its element's columns (a
property test against a twin strip that is never patched); the render
thread's apply takes no lock and draws nothing; the worker's priorities,
floors, backoff and hand-over; and, end to end on the emulator with the stub
plugin (`test/fixtures/plugins/vegas-live-stub`), an update changes a card
already in the strip and an animated element moves with no update at all.
- **Hardware:** an hdpi soak, A/B against the #628 build, alternating order.
Targets: no freezes, an empty 6+ bucket, the 3–5 bucket near zero, and the late
rate below 0.66%. Plus, for freshness: log each segment's age when it enters
the viewport, and compare the median and max before and after.
## Rollout
1. **Offscreen rendering** (shipped): `offscreen()`, the adapter on the
prefetch thread, and the plugin lock. Removed the render-thread pauses.
2. **Measurement and the lazy strip image:** late frames attributed to the
render-thread work before them (`FrameTimingRecorder.note_op`, the "after
work" table), and extensions no longer rebuilding the strip's PIL copy.
3. **Live elements:** the plugin API, the records, the worker and in-place
patches, with the sports scoreboards and the flight map adopting it.
4. **Live games in the ticker by default:** `live_in_ticker` true, so a live
game's cards update in the marquee instead of the full-screen scoreboard
replacing it; existing configs are switched once (`ConfigManager`).
`display.vegas_scroll.offscreen_prefetch` (default `true`) restores the
deferred path when `false`, and `display.vegas_scroll.live_refresh` (default
`true`) turns live elements off. Keep both for one release, then delete the
old paths.
## Open questions
1. Multi-display sync: is there a two-Pi rig to test on? Replaying strip
operations to the follower (append, trim, patch, in absolute columns) would
fix continuous-mode sync and let live elements run under it.
2. Is the `extend` cost worth a segment list after the lazy image? The soak's
"after work" table answers it per rig.

Some files were not shown because too many files have changed in this diff Show More