Commit Graph
193 Commits
Author SHA1 Message Date
ChuckandClaude Opus 5.5 ece416c4e5 refactor(plugins): one plugin-directory resolver (#623)
* refactor(plugins): one resolver for plugin id -> directory

Five places mapped a plugin id to its directory, each with its own rules
and each re-reading manifests per lookup: PluginManager discovery and
get_plugin_directory, PluginLoader.find_plugin_directory,
PluginStoreManager._find_plugin_path / list_installed_plugins, and
state_reconciliation.disk_plugin_ids. They disagreed on backup dirs,
on whether the manifest id or the directory name is the id, on duplicate
ids and on path safety.

src/plugin_system/plugin_dirs.py now holds the rules once:
PluginDirectoryIndex scans one directory and reads each manifest once;
resolve_plugin_dir() searches directories in order. What legitimately
differs per caller is an explicit argument: search dirs (discovery and
the loader: configured dir only; the store: configured then sibling
plugins/), ledmatrix- prefix (not for the store), case folding (loader
only), manifest pass (not for get_plugin_directory, whose discovery map
already holds it).

Behaviour changes, all for layouts installs do not produce:
- a directory whose manifest declares the id beats one merely named for
  it (discovery already worked this way; the loader and store now agree)
- the store searches the configured dir completely before plugins/
- backup and hidden dirs are skipped everywhere (the loader's case and
  manifest scans and list_installed_plugins used to return them)
- duplicate ids resolve deterministically (exact name, then
  ledmatrix-<id>, then by name) with a one-time warning; discovery no
  longer lists the id twice
- disk_plugin_ids / list_installed_plugins report manifest ids, falling
  back to the directory name; auto-update looks the directory up
- ids that are not one plain path segment resolve to nothing in every
  caller (the loader used to truncate them, the store to join them)

The .standalone-backup- marker is one constant, BACKUP_MARKER, used by
store_manager's rename-aside names and every lookup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): one plugin-directory resolver

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:52:52 -04:00
ChuckandClaude Opus 5.5 13bbb537f3 refactor(web): one logging setup and one TTL cache for the web process (#621)
* refactor(web): use src.logging_config in the web process; routine requests to DEBUG

The web interface had its own logging setup (web_interface/logging_config.py)
that replaced the root handlers with a plain stdout formatter. The web
service's journal lines therefore never carried a syslog priority, so
`journalctl -p err -u ledmatrix-web` returned nothing while errors were
logged, and the line shape differed from the display's (the log viewer's
prefix stripping only matched the display format). It also ran after the
module-level managers were built, so their INFO lines at import (including
"Re-removed N uninstalled plugin(s)") were dropped.

app.py now calls src.logging_config.setup_logging() first thing, the same as
run.py: journald priorities under systemd, LEDMATRIX_DEBUG honoured,
LEDMATRIX_JSON_LOGGING still selects JSON.

Per-request logging moves to web_interface/request_logging.py. Every request
used to be logged at INFO, so the UI's polling filled the journal
("GET /api/v3/errors/summary - 200" every minute per tab). Now a successful
GET/HEAD/OPTIONS is DEBUG, a successful write is INFO, 4xx WARNING, 5xx
ERROR. Durations use perf_counter and print to 0.1ms.

The duplicate module is deleted; nothing else imported it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): one thread-safe TTL cache for the web process

web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.

Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
  counted and get_cached defaulted to 60s. An entry now expires after the TTL
  it was stored with; a reader's ttl_seconds can only shorten that. Both
  current callers pass the same value on both sides (fonts_catalog 300s,
  system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
  same expired key could raise KeyError (reproduced), which the endpoints
  turned into a 500.

app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.

Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): web logging and TTL cache

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): only ask systemctl about known units

Codacy flagged the systemctl argv built from a variable. The unit now has
to be one of two literals, and anything else raises.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): response_time_ms reads the same clock request_logging stamps

request_logging now stamps request.start_time from perf_counter, but
success_response still subtracted it from time.time(), so metadata
reported ~1.8e12 ms. Found testing on ledpi.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:52:17 -04:00
ChuckandClaude Opus 5.5 c1ce0b7b04 fix(web): two api_v3 paths called names that no longer exist (#625)
The Pixlet editor stop route restarts the display after a SIGKILL with
_run_systemctl_command, which starlark.py never imported (since #554). The
Starlark device-location resolver fell back to _ensure_cache_manager, which
#609 deleted; the resolver already accepts no cache manager. Both raised
NameError on the rare path that reaches them. pyflakes finds no other
undefined names in src/ or web_interface/.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 15:49:29 -04:00
ChuckandClaude Opus 5.5 f3894916a9 feat(web): show which plugins use each font; warn before deleting one (#619)
* feat(web): show which plugins use each font, warn before deleting one

The Fonts tab lists font files from the web process's own scan, and the
plugins that register fonts run in the display process, so the tab had no
way to say whether a font was in use before deleting it.

The display service now publishes {catalog key: [plugin ids]} to the
shared cache (font_usage_snapshot, src/font_usage.py), built from the
loaded plugins' FontManager.register_manager_font() registrations. A
daemon thread checks every 10 s and writes only when the usage changed
(plus a daily refresh so cache cleanup cannot expire it); it never raises.
Families, aliases (press_start, four_by_six, ...) and paths are resolved
through FontManager's catalog to the file stem the Fonts tab keys rows by;
fonts outside assets/fonts are left out. Unloading a plugin drops its
registrations (new FontManager.forget_manager_fonts).

GET /api/v3/fonts/catalog merges used_by into each row per request (the
5-minute scan cache is copied, never edited): a list of plugin ids, or
null when the display service has not reported. The tab shows a Used by
column ("unknown" / "-" / ids, rendered as text) and deleting an in-use
font names the plugins in the confirmation, from a fresh read. The server
still refuses only system fonts. Catalog fetches bypass the browser's
5-second API cache, which otherwise served the pre-delete list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: call forget_manager_fonts through a hasattr check pylint can follow

getattr(..., None) then callable() is fine at runtime, but pylint's E1102
("not callable") can't see through it, and Codacy fails the check on it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:50:42 -04:00
ChuckandClaude Opus 5.5 9a1f94f793 fix(starlark): blank app locations use the device location, not San Francisco (#617)
* fix(starlark): blank app locations use the device location, not San Francisco

A Starlark (Tidbyt) app whose Location field is blank rendered at its
author's hard-coded DEFAULT_LOCATION -- usually San Francisco -- even with
the device city set under General settings. A user in Charlotte, NC got San
Francisco weather and radar with nothing in config.json to explain it.

src/device_location.py fills unset location fields at render time (display
plugin and the web standalone render): the device city is geocoded once via
Open-Meteo, preferring a match in the configured state/country, and cached
permanently. A saved location always wins; if the lookup fails the field is
dropped so the app uses its own default, and the failure is not retried for
30 minutes.

Also fixes the config form: clearing a location omitted the key, and the
save merges, so the old value could never be removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(starlark): say what happens when the device location can't be used

A blank app Location only renders at the device's city when one is set and
the Open-Meteo lookup finds it. With no city, no match, or the geocoder
unreachable (retried after 30 minutes), the app gets no location and keeps
its author's default. The guide, the config page hint, CONFIG_REFERENCE and
the CHANGELOG entry now say so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:48:02 -04:00
ChuckandClaude Opus 5.5 61e462c635 refactor: remove the skin system and the unused src/base_classes package (#615)
* refactor: remove the skin system

Skins never rendered with the current scoreboard plugins: the only hook was
SportsCore._render_game in src/base_classes, which no plugin builds on, so
the UI and store already treated them as unsupported. The owner decided on
2026-09-23 to remove them outright.

Removed src/skin_system/ (runtime, base class, fixtures), skins/,
scripts/validate_skin.py and their tests; the store's "type": "skin"
installer, uninstaller and hide/refuse filters (the official registry lists
no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins.

Stored skin/skin_options config values are handled in the next commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(config): drop retired skin/skin_options keys instead of validating them

A config.json written while the skin system existed can carry skin and
skin_options in any plugin section, and most plugin schemas set
additionalProperties: false. They are no longer core plugin properties;
RETIRED_PLUGIN_KEYS in schema_manager lists them and
drop_retired_plugin_keys removes them (unless the plugin's own schema
declares the name) in prepare_plugin_config, which loading, hot reload,
GET /plugins/config and both web saves already share, and in
validate_config_against_schema for callers that validate a raw section.
POST /plugins/config and /config/main also drop them from the stored
section they merge into, so they leave config.json on the next save.

Tests cover the load path (real PluginManager.load_plugin: no schema
warning, not degraded), raw and prepared validation,
validate_all_plugin_configs, and the JSON, form and /config/main saves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor: remove the unused src/base_classes package

No scoreboard plugin builds on src.base_classes: the nine monorepo
scoreboards ship their own sports.py and share code through src/common
(docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins
imports it. The one import anywhere, baseball-scoreboard's
rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing
instantiates.

Removed the package and the eight test files that only tested it
(test_api_extractors, test_data_sources, test_sports_base_characterization,
test_sports_capabilities, test_sports_core_promotions,
test_sports_logo_cache_bounded, test_sports_modes_promotions,
test_sports_odds_fanout). test_common_is_hardware_free no longer lists
src.base_classes as a forbidden import, and comments in sports_helpers.py
and base_odds_manager.py stop pointing at it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: drop the skin system and src/base_classes from the docs

Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to
them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins
section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the
term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was
removed and shared code lives in src/common, in the Layering section and
the view-model-contract rule. Other docs stop pointing at the removed
package. CHANGELOG records both removals under Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(store): hide and refuse registry entries that aren't plugins

The skin filters went with the skin system, but a custom registry can still
list "type": "skin" entries, and installing one as a plugin would unpack it
into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing
type means plugin) now hides non-plugin entries from the store and
custom-registry listings, and install refuses them, in the route with a
clear 400 and in _install_plugin_impl for any other caller.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 16:33:24 -04:00
ChuckandClaude Opus 5 4fe3cdd906 fix(starlark): stop the root display service locking the web UI out of starlark-apps (#604)
* fix(starlark): stop the root display service locking the web UI out

Reported after a fresh install: installing an app from the Starlark tab
failed with "install failed: Failed to install from repository", and so did
uploading a .star file and installing from a GitHub directory. The reporter
found the cause only by reading service logs, and fixed it with

    sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/starlark-apps

starlark-apps is gitignored, so it is never checked out -- it is created
lazily by whichever process reaches it first. Those processes run as
different users. systemd/ledmatrix.service is User=root and constructs this
plugin at startup, which is where _get_apps_directory() is called from;
systemd/ledmatrix-web.service runs as the login user and is what actually
installs apps.

The documented first step is to install pixlet and reboot, so on a fresh
machine the display service usually wins that race and mkdir() leaves the
directory root-owned. The web process then fails in _install_star_file() on
app_dir.mkdir(), which catches nothing, so PermissionError reaches the
route's outer `except Exception` and becomes the generic message the user
saw. All three install paths write to the same directory, which is why all
three failed.

The web user cannot repair this -- chown needs root. So root does it, on
every startup, which also heals machines already broken by this without the
owner having to find the chown themselves. It is a no-op when not root, when
the platform has no POSIX ownership, and when the checkout genuinely belongs
to root; a chown that fails warns rather than killing startup.

Also made the failure legible if the handover is ever prevented: a
PermissionError now names the directory, the automatic repair, and the
manual chown, instead of a message that names neither path nor cause.

Verified by mutation: dropping the handover call, chowning a genuinely
root-owned checkout, and letting a non-root process chown each fail their
own test. 121 starlark tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): address the review on the ownership repair

Findings from the automated review of #604.

Symlinks (CWE-59, the serious one). A root chown that follows links is a
privilege-escalation primitive: anyone able to write in starlark-apps could
point a link at a root-owned file and have the repair hand it over. Entries
are now read with os.lstat, symlinks are skipped outright, and the chown
passes follow_symlinks=False. Descendants are processed before the directory
itself, so the container does not change hands while its contents are still
being walked.

install_app() caught PermissionError in its broad handler and returned
False, which both routes report as a generic install failure -- the exact
shape of the bug this PR exists to fix, since the caller could not tell
"this app is broken" from "this process cannot write here". PermissionError
is now re-raised; every other failure still returns False.

The test fixtures skipped on bare Exception, which would have turned a
syntax error or NameError in the plugin into a green run. They now skip only
for a named absent dependency and re-raise anything else.

Also fixed the _Stat stub that failed in CI but passed locally: it carried
only st_uid/st_gid, and pathlib reads st_mode while walking. It now wraps
the real stat result and overrides ownership alone.

NOT taken: the CodeQL "information exposure through an exception" finding on
the hint response. Dropping `details` would contradict this package's
documented rule -- "if it returns 5xx, it says why" -- which
test_no_api_v3_handler_discards_its_exception enforces with an allowance
that may shrink and never grow. The Starlark routes are the ones that policy
was written for: they answered 500 with no detail for three releases.
describe_exception already redacts credentials and truncates. Keeping the
detail is the deliberate trade-off, so the finding is declined rather than
silently worked around.

Verified on hdpi with the updated code: a symlink to /etc/shadow planted in
starlark-apps was skipped while the directory was handed back, and
/etc/shadow stayed root:shadow.

Mutation-checked all three behaviours. The symlink test was vacuous on the
first attempt -- the link already had the target owner, so it was skipped
for the wrong reason and the mutation passed. It now forces the link to look
like it needs handing over, and fails when the check is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 16:33:12 -04:00
ChuckandClaude Opus 5.5 4a1fd7464a fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)
* fix(errors): serve /api/v3/errors/* from the display service's aggregator

The error aggregator is a per-process singleton and only the display
service runs plugins, so only its aggregator records anything. The routes
read the web process's own, empty one and always reported no errors.

The display service now publishes a bounded snapshot of its aggregator to
the shared cache (plugin_error_snapshot) from a daemon thread: at most once
every 10 s and only when something changed, never raising into the caller.
The routes read it and keep their response shapes, adding
snapshot_available, generated_at and clear_pending; exception text has
credentials redacted.

POST /errors/clear writes a clear request (plugin_error_clear_request) that
the display applies on its next 5 s tick via the new clear_before(), which
keeps errors recorded after the cutoff and rebuilds the counts. Until the
snapshot acknowledges the request, reads hide everything before the cutoff,
so a snapshot written just before the click cannot bring errors back. Adds
"all": true; cleared_count is null when only the display can know it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(web): show plugin errors in the Logs tab

A compact panel under the log viewer: per-plugin error counts, repeating
errors (type, count, affected plugins, a sample message, last seen) and a
Clear button, with empty states for "no errors" and "display service
hasn't reported yet". Polls every 15 s while the tab is active; all text
goes through escapeHtml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: describe where plugin error reports come from and how clear works

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(errors): redact the published snapshot before clipping it

Keeping only a traceback's tail (or clipping a message) could cut an
`api_key=` marker off while keeping the secret after it, and the web side's
redaction would then have nothing to match. The display now redacts every
free-text field of the snapshot first. The patterns move to a Flask-free
src/redaction.py so the display service can use them; redact_text in the web
error handler uses the same function, unchanged in behaviour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 14:32:02 -04:00
ChuckandClaude Opus 5.5 f813ea2117 fix(web): define project_root when plugins_directory is absolute (#616)
project_root was only assigned in the relative-path branch, so an absolute
plugin_system.plugins_directory made web_interface/app.py raise NameError
at import (first use: the SchemaManager construction). Define it before the
if/else; plugins_dir resolution is unchanged.

Adds a regression test that imports the real module in a fresh interpreter
with an absolute and a relative plugins_directory.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:54:29 -04:00
ChuckandClaude Opus 5.5 a8b3e86775 refactor(web): delete dead routes, JS files and duplicate definitions (#609)
* refactor(web): drop validators nothing calls

escape_html, validate_image_url, validate_font_awesome_class,
validate_mime_type, validate_numeric_range, validate_string_length and
sanitize_plugin_config had no callers outside their own tests. Only
validate_file_upload (fonts upload) is imported by the web interface.

dedup_unique_arrays is kept: its one caller in save_plugin_config was
removed by the unrelated sync PR (#330), which looks accidental.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(api): remove the music-auth and of-the-day JSON routes

POST /plugins/authenticate/spotify and /plugins/authenticate/ytm had no
caller but their tests: the music plugin authenticates through its
web_ui_actions (authenticate_spotify.py / authenticate_ytm.py) via
/plugins/action.

POST /plugins/of-the-day/json/upload and /json/delete looked the plugin
up by the id ledmatrix-of-the-day (its manifest id is of-the-day), were
reachable only from a file_type "json" upload field that no schema
declares, and put the plugin directory on sys.path per request to
import scripts.update_config. of-the-day manages its files through
plugin-file-manager and its own web_ui_actions.

The of-the-day branch of GET /plugins/config stays: it matches the real
manifest id and still merges the on-disk category files into the form.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(api): read managers only from the blueprints

api_v3/__init__.py and pages_v3.py declared module globals
(plugin_store_manager, saved_repositories_manager, schema_manager,
operation_queue, plugin_state_manager, operation_history, sync_manager,
config_manager, plugin_manager) that nothing assigns: app.py sets the
managers as attributes on the Blueprint objects, and every route reads
them there. The one reader, backup restore's fallback to the module
plugin_store_manager, could only ever fall back to None.

_ensure_cache_manager() built a second CacheManager in the web process
instead of using the one app.py puts on api_v3. The display routes now
read api_v3.cache_manager, creating it on the blueprint only when
nothing set it (the same None handling as the /cache routes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(web): drop run.sh and the unused log_config_change

web_interface/run.sh was referenced only by web_interface/README.md;
the service starts the UI through scripts/utils/start_web_conditionally.py
and the README already documents `python3 web_interface/start.py`.
log_config_change() in web_interface/logging_config.py was never called.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): delete unreferenced store_manager.js, diff_viewer.js, htmx-sse.js

- js/plugins/store_manager.js (window.PluginStoreManager) and
  js/config/diff_viewer.js (window.ConfigDiffViewer) were loaded on every
  page but nothing reads either global.
- htmx-sse.js (plus its CDN fallback) was loaded after HTMX, but no
  template or plugin page uses sse-connect / hx-ext="sse": the live
  streams run through LEDStreams in app-shell.js.

js/plugins/state_manager.js stays: install_manager.js's updateAll()
reads and refreshes window.PluginStateManager.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): remove app.js helpers nothing calls

- hexToRgb, rgbToHex, validateForm, uploadFont and switchTab (whose
  'switch-tab' event had no listener) have no caller in the templates,
  static JS or the plugin monorepo.
- installPlugin: plugins_manager.js (loaded last) assigns
  window.installPlugin, and its own store cards are the only callers.
- The showNotification fallback could never install: app-shell.js is
  deferred ahead of app.js and defines the same fallback at top level.
- performanceMonitor only logged with ?debug=perf and read an unset
  this.measures; the marks it took on every load had no reader.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): drop app-shell.js refreshPlugin

A top-level function in app-shell.js, so a window global, but nothing
calls it (no inline handler, no window lookup, no string-built name).

The other plugin actions in that block stay. updatePlugin is the live
window.updatePlugin: plugins_manager.js only installs its own copy when
none exists. uninstallPlugin/pollUninstallOperation, updateAllPlugins,
executePluginAction and toggleNestedSection are replaced by later
deferred scripts, but a click that lands while those scripts are still
downloading reaches the app-shell copies, so removing them is not a
pure no-op.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): remove definitions plugins_manager.js always overrides

All of these are replaced before anything can call them, checked
against the load order in base.html and the live window.* values:

- openOnDemandModal/requestOnDemandStop stubs: the IIFE later in the
  same script assigns the real functions synchronously.
- updatePlugin and uninstallPlugin stubs (`window.X || stub`): app-shell.js
  already defined both, so the fallback never installed. Same for the
  later updatePlugin override, gated on the live function containing
  '[UPDATE]', which app-shell.js's never does.
- The first addArrayObjectItem/removeArrayObjectItem: reassigned by the
  top-level copies after the IIFE.
- The first `function formatDate` in the IIFE: a later declaration of
  the same name in the same scope wins.
- deleteUploadedImage, getCurrentImages, showUploadProgress,
  formatFileSize and getScheduleSummary: character-for-character
  copies of js/widgets/file-upload.js, which stays the owner.
- `typeof X === 'undefined'` fallbacks and `typeof X !== 'undefined'`
  re-exports after the IIFE: always false, or a self-assignment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(web): render the shell directly and delete index.html

index.html extended base.html with {% block content %}, but base.html
defines no blocks, so none of index.html ever rendered: rendering both
with jinja2 gives byte-identical output. index() still loaded the config,
read config.json and config_secrets.json raw and json.dumps'd them on
every page load for variables base.html never reads, and flashed errors
that base.html never shows. It now renders base.html with no context.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): stop htmx-config.js replacing console.error and console.warn

It swapped both globals for filters that dropped any error mentioning
insertBefore / "Cannot read properties of null" when "htmx" appeared in
the message or stack, and a list of Permissions-Policy warnings. That
hid real errors from every script on the page, and made every logged
error and warning report htmx-config.js as its source. The beforeSwap
target validation above it, which prevents the insertBefore errors in
the first place, stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(web): quiet the widget load announcements and debug logs

About 30 lines hit the console on every page load: one "... widget
registered" per widget file, one "[WidgetRegistry] Registered widget: X"
per registration, plus the registry, base widget and plugin loader
announcing themselves. The load-time announcements are removed; the
per-call ones (registry register, plugin widget loads, "Render called")
now go through the page's debugLog switch (localStorage.pluginDebug),
guarded because the widgets also load in node tests without it.
fonts.html and wifi.html debug logging goes through debugLog as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(api): drop the removed music-auth and of-the-day JSON routes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:36:53 -04:00
ChuckandClaude Opus 5.5 84afa9d64f refactor: delete dead Python code in the core (and stop storing Wi-Fi passwords) (#608)
* refactor(plugins): remove the no-op PluginHealthMonitor

Its monitor loop did nothing (`if callbacks: pass`), register_health_check
had no callers and api_v3.health_monitor was never read by any route. The
live health data comes from PluginHealthTracker, which is untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(store): drop the never-set uninstall tombstones

Nothing in production called mark_recently_uninstalled, so the
reconciler's was_recently_uninstalled check was always False. The
persistent uninstall registry is what actually stops resurrection; the
reconciler test now exercises that gate instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(common): delete unused config/display/game helpers, utils and error_handler

Nothing in core, the web UI, scripts or the plugin monorepo imports
config_helper, display_helper, game_helper, utils or error_handler; only
their own tests did. The error_handler re-exports leave src.common's
__all__; APIHelper, TextHelper, ScrollHelper, LogoHelper and the adaptive
layout exports are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(config): drop ConfigService's unused versioning and save API

ConfigVersion, get_version/get_version_history/get_version_config,
rollback, save_config, reload, get_plugin_config and the backward-compat
load_config/get_config_path/get_secrets_path had no callers. The display
controller only uses get_config, subscribe, unsubscribe and shutdown,
plus the file watcher. Change detection now compares against the
current checksum instead of the last history entry.

The subscriber tests asserted `callback.called or True`; they now
reload the way the watcher does and assert the notification.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): drop unread plugin state history and callbacks

plugin_state.PluginStateManager kept a bounded per-plugin transition
history that only get_state_history (tests only) read; get_state_info
reports a separate lifetime count, which stays. set_error_info and
record_display had no callers, and set_state_with_error's `error`
argument only fed the history.

The web-side state_manager.PluginStateManager loses
subscribe_to_state_changes, _notify_callbacks, set_plugin_error and
get_state_version, none of which had callers; with no subscribers the
old-state copy in update_plugin_state went with them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): remove unused PluginManager methods and attribute guards

update_all_plugins was only called by a test (the display loop uses
run_scheduled_updates); get_plugin_health_metrics,
get_plugin_resource_metrics and get_plugin_state had no callers; and
plugin_modules was written but never read. plugin_directories is now
initialised in __init__, so the hasattr() guards around it go.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(plugins): remove unused executor, loader, store and package helpers

- PluginExecutor.execute_safe: no callers.
- PluginLoader._parse_semver: only its own tests; compatibility.parse_semver
  is the live copy and test_compatibility.py already covers it.
- PluginStoreManager.get_installed_plugin_info: no callers.
- PluginResourceMonitor._local: never read.
- src.plugin_system.get_store_manager and __api_version__: no importers in
  core, scripts or the plugin monorepo.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(wifi): stop storing Wi-Fi passwords in wifi_config.json

WiFiManager appended every joined network's SSID and password, in
plaintext, to saved_networks in config/wifi_config.json, and nothing
(web UI, backup restore, scripts) ever read them back: NetworkManager
keeps its own credentials. The writes are gone, and loading the config
now drops any saved_networks key and rewrites the file, so passwords
already on disk are scrubbed.

Also removes _check_dnsmasq_conflict (never called) and _detect_trixie,
whose result only reached one log line, along with the
NM_CONNECTIONS_PATHS constant only it used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(display): remove unreachable and unused DisplayController code

- _follower_rebuild_scroll_image: never called.
- mode_duration (never read) and last_mode_change (write-only).
- The `chosen_cap <= 0` branch: chosen_cap is either the minimum of
  caps already filtered to > 0 or DEFAULT_DYNAMIC_DURATION_CAP (180).
- The `max_duration < min_duration` branch directly after
  `max_duration = max(min_duration, max_duration)`.
- The circuit-breaker branch's `display_result = False` and
  `manager_to_display = None`: the first is overwritten a few lines
  later, the second is already None there.
- The bool-to-bool conversion of execute_display's result, which is
  always a bool.
- The `loaded_plugins` lookup in _update_modules: PluginManager has no
  such attribute, so it always fell through to `plugins`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(vegas): remove unused config update, boundary finder and refresh

VegasModeConfig.update had no callers outside its own tests (the
coordinator rebuilds the config with from_config on a change);
geometry.find_item_boundary and StreamManager._refresh_plugin_content
had no callers at all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* refactor(run): drop the debug block that pretended to import the plugin system

In debug mode run.py put src/plugin_system itself on sys.path and printed
"Plugin system import successful" without importing anything. Nothing
imports plugin_system modules by bare name, so the path entry did
nothing either.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: delete tests that test nothing

- test/plugins/test_{basketball_scoreboard,calendar,clock_simple,
  odds_ticker,soccer_scoreboard,text_display}.py skip everywhere the named
  plugins are not installed, including CI (LEDMATRIX_PLUGINS_DIR holds only
  the fixture plugin); test_plugin_matrix.py already covers every
  discovered plugin. Their PluginTestBase and the fixtures only it used
  (plugins_dir, mock_display_manager, mock_cache_manager,
  mock_plugin_manager, base_plugin_config in test/plugins/conftest.py) go
  with them.
- test_plugin_system.py: test_discover_plugins (body was `pass`) and
  test_dependency_check (a comment), plus the test_plugin_manager fixture
  only the former requested.
- test_display_manager.py: test_draw_image asserted that an image it had
  just assigned was not None.
- test_display_controller.py: the rotation and schedule-override tests
  re-implemented the run-loop arithmetic inline and asserted on their own
  result without calling the controller.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: expect one plugin_last_update success stamp after update_all_plugins

EveryStampRecordsACompletion required at least two success-path stamps;
the second was update_all_plugins, removed as test-only. The worker and
synchronous paths share the remaining stamp in _execute_update_now, and
the check that every stamp calls _note_update_completed is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 12:36:26 -04:00
ChuckandClaude Opus 5.5 342e9164b8 fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound

Each repeat of an error pattern appended every plugin in the time window to
the pattern's list again, so a plugin failing in a loop grew the display
process's memory without limit: 3,000 errors from three plugins reached 2.5
million entries. Keep the list unique.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(fonts): load a BDF font at its native size instead of PIL's default

FreeType rejects any size but a BDF strike's own, and FontManager answered
that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf
requested at 8 or 10px rendered as PIL's default font. Retry at the native
strike, as element_style already does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): plugin toggle failures no longer claim "operation in progress"

Every exception in POST /plugins/toggle was mapped to
PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation
is already in progress". Report the failure as what it is, and record the
plugin id in the operation history for form posts too (it read a `data`
variable that only the JSON path set).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): route plugin card clicks through handlePluginAction

The document-level delegation checked `typeof handlePluginAction`, which is
scoped inside the plugin-manager IIFE and so never visible to it. Every card
click took a copied fallback that stopped propagation (the grid's own
listener never ran), confirmed an uninstall twice, and sent Starlark app
uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>.
Expose the handler on window and delegate to it.

Also run every test/js/unit suite under pytest: they need only node, but CI
ran one of the eight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(display): apply Rotation durations, WiFi messages and Vegas settings

Three settings the web UI saves never reached the display:

- Rotation & Durations: display.display_durations was never read. Every
  plugin inherits get_display_duration() and the plugin was asked first. A
  saved value now wins. The page shows unsaved screens blank with the
  plugin's own duration as a placeholder, and saving a blank removes the
  override, so one save no longer pins every screen.
- WiFi status overlay: the controller looked for wifi_status.json one
  directory above the repo. Both sides now use
  wifi_manager.get_wifi_status_path(). The message is written by rename so
  the display never reads it half-written, and the resumed plugin redraws the
  whole panel afterwards.
- Vegas: nothing called coordinator.update_config(), so saved Vegas settings
  never reached a running scroll. They are now queued when
  display.vegas_scroll changes, and applied while Vegas is stopped too, so a
  disable then re-enable works. The follower's scroll-speed default (75) now
  matches VegasModeConfig's (50).

Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us
per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows
with each plugin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: keep affected_plugins order when serialized; guard non-Element targets

ErrorPattern.to_dict() ran the now-ordered list through set(), so
get_error_summary() listed plugins in an unstable order. The document-level
card-action listener called event.target.closest() without checking the
target is an Element.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 10:29:16 -04:00
ChuckandClaude Opus 5 19686ab761 fix(web): plugin settings form shows schema defaults for unsaved keys (#597)
The server-rendered plugin settings partial rendered straight from the
saved config, so an option added in a plugin update (geochron 1.2.0's
show_date / show_date_line, default true) drew as an unchecked box, and
the save route's missing-checkbox handling then stored it as false.
Enum dropdowns likewise showed their first option instead of the default.

- _load_plugin_config_partial runs the stored section through
  prepare_plugin_config (as GET /plugins/config does) before masking
  secrets, so a secret's schema default is masked too.
- render_field falls back to the field's own default, covering children
  of objects that declare a default of their own (where the defaults
  extraction stops).
- The legacy-boolean parity test now compares against the config the
  plugin actually runs with (defaults included).

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-18 11:09:25 -04:00
ChuckandClaude Opus 5 116abb0daa fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service

BackgroundDataService always sent a season range first and, on a 400,
fell back to chunks without recording the rejection, so every background
season fetch spent a doomed request and live scoreboards learned nothing
from it (or it from them). The worker now consults and sets the same
6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight
to month/day chunks, and if every chunk fails the range is asked once for
a real error without re-spending the chunks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep plugin asset and action routes inside their directories

POST /plugins/assets/upload, GET /plugins/assets/list and POST
/plugins/assets/delete joined the request's plugin_id onto assets/plugins
unchecked, so '../../config' created, wrote, listed and deleted outside
it. #561 guarded only the route that serves the files. All three now go
through path_safety.resolve_under and answer 400 for anything but a
plain name, and delete only unlinks a metadata path that resolves into
that plugin's uploads directory.

PluginManager.get_plugin_directory refuses ids that are not one plain
path segment, so /plugins/action (which runs a manifest script from the
returned directory) and every other caller get the guard; the action
route also rejects such ids up front, covering its no-manager fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): report a no-op plugin update as already up to date

update_plugin() returns True both for a real update and for "nothing to
do" (a ZIP-installed monorepo plugin already at the registry version, a
bundled plugin). With no git commit to compare, POST /plugins/update
called every such success "updated successfully", so Check & Update All
counted most official plugins as updated on every run.

The route now reads what changed off the plugin itself (commit, else
manifest version, else last_updated) and returns data.update_status
(updated / up_to_date / local_only). The update-all toast is summarised
by PluginInstallManager.summarizeUpdateResults from that status, falling
back to the message for older servers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(sports): scoreboard scroll speed no longer follows target_fps

sports_scroll computed the crisp speed ladder against the global
target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since
frame-locked presentation (#545) the helper steps a fixed number of whole
pixels per presented frame and the panel presents at its real refresh, so
the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a
50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s.

The ladder now uses the display manager's refresh_hz, then
display.hardware.limit_refresh_rate_hz, then the default. target_fps is
not consulted. Docstrings now say scroll_delay is ignored for pacing (no
behaviour change there) and describe the fixed-step model.

Tests: replace the tests that pinned target_fps as the ladder refresh and
described time-based stepping; assert speed independence from target_fps
(unit and end-to-end presented px/s against the real helper), that the
fixed per-frame step is applied, and that scroll_delay does not change
speed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): escape registry and upload values in plugin manager inline handlers

The store, saved-repository and custom-registry buttons built
onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone,
so a custom registry entry whose id contained ' closed the attribute and
added its own handler. One helper, jsStringAttr(), now HTML-escapes the
JSON literal for every one of those handlers, and the store View button
opens only http(s) repo links.

The live window.updateImageList (plugins_manager.js loads last, so its
copy wins over the file-upload widget's) wrote the uploaded file's
original name, path and ids into markup raw; they are escaped now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note plugin asset, action and inline handler guards

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): let the root pip wrapper install web_interface/requirements.txt

Update Code, the automatic update's health check and Install Base
Requirements install web_interface/requirements.txt through
safe_pip_install.sh, which only allowed the root requirements.txt. The
first commit changing that file would fail its dependency install, and
the automatic updater rolls back any update whose dependencies did not
install -- on every device, for every newer commit.

The wrapper now lists both core requirement files. Only their folders
are resolved, so a requirements.txt symlinked out of the project is
compared by its target and refused (previously the root file's own
symlink target was what got allowed). The updater's file list is a
named constant, and a test runs the real wrapper (pip stubbed) on
every file Update Code and the rollback install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): do not retry plugin requests that got an HTTP answer

PluginAPI.request wrapped everything that was not a structured error as
NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a
JSON error without error_code included. Check & Update All retries
NETWORK_ERROR, so those updates were re-sent five more times with
backoff, contrary to the #587 contract that an HTTP error response is
the server's answer.

NETWORK_ERROR now means only that fetch() rejected. Any HTTP response
without an error_code, or with a body that is not JSON, is API_ERROR
with the HTTP status attached. Tested against the shipped api_client.js.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scroll): restart the stats window when an idle gap is dropped by size

#582 dropped an idle gap from the frame stats two ways: the reset_scroll()
sentinel, which also restarts the 5s window timer, and a size guard for
scrollers that never call reset_scroll(), which did not. On that path the
first real frame after the gap found the boundary overdue and logged a
stats line for a one-frame window. Both paths now share one seeding helper.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): leave plugins alone when update_core's own rollback fails

update_core returns rollback_failed directly when a partial pull or an
update whose health check never started cannot be rolled back. run()
only held plugins back for 'verifying', so those devices still got new
plugin versions and a display restart on top of a core in an unknown
state -- the opposite of what the health-check path does, and of the
3.4.0 changelog (plugins are left alone if the rollback fails).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(api): make the REST reference match the api_v3 package

Every documented request body, query parameter and response shape was
re-checked against the handlers in web_interface/blueprints/api_v3/.
Fixes calls that failed as documented (repo_url, action_id/params,
files/image_id, font_file+font_family, ?font=, cache key,
auto_enable_ap_mode, plugin limit keys), removes the font-override
endpoints dropped in #566, corrects response shapes (plugins/config,
plugins/schema, health, metrics, operation history, github-status,
fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds
the 26 routes it omitted (backup, system auto-update/git, wifi radio,
starlark editor, MQTT bridge, status endpoints, skins).

Documents the merge semantics of partial JSON saves to /config/main and
/plugins/config and the dim-schedule POST accepting GET's days shape,
which land in the same change set. Replaces app.py line numbers and the
removed api_v3.py path with file and function names.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): remove the General-tab plugin system toggles that did nothing

plugin_system.auto_discover, auto_load_enabled and development_mode had
General-tab toggles whose help tips promised dormant plugins and verbose
logging, but nothing reads them: every enabled plugin is discovered and
loaded regardless. Remove the three toggles.

The keys stay tolerated in stored configs. The save handler now stores
a flag only when a client sends it; treating a missing key as an
unchecked box would otherwise rewrite all three to false on every
General-tab save, which still posts plugins_directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(scroll): remove dead code left by #523/#570

- Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read
  them since the numpy blend replaced the scipy path.
- Drop ScrollHelper._last_integer_position and frame_time_target, which
  were written but never read.
- Keep target_fps and set_target_fps() but document them as
  informational: nothing paces off them, yet ledmatrix-elections'
  test_scroll_pacing.py reads helper.target_fps back and third-party
  plugins may call the setter.
- Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to
  throttle", set_sub_pixel_scrolling's "default: True", and
  set_frame_based_scrolling's claim that it steps.

The plugins monorepo was grepped for every removed name; none is used.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI

FONT_MANAGER.md told plugins to read display_manager.font_manager, which
does not exist, so a plugin following it failed to load with
AttributeError. The shared FontManager lives on the PluginManager and
BasePlugin._get_font_manager() returns it (with a fallback for harnesses).

Also removes the Fonts-tab override workflow and element-override panels
that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls

/plugins/store/search does not exist (404) and the list endpoint reads
query, not q. The registry guide's curl examples omitted the JSON
Content-Type, so the handlers saw an empty body and answered 400.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(config): use the shared core-key list in the last three private copies

StartupValidator warned "Plugin 'auto_update' is enabled but not found" on
every display start with auto-update or a dim schedule on; the reserved
plugin-id check missed auto_update, sync, location and the rest; and
ConfigManager's (uncalled) orphan cleanup would have deleted display,
schedule and auto_update. All three now read src/core_config_keys.py, which
also gains CORE_SECRETS_KEYS for the github/youtube secrets sections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): partial JSON saves to /config/main change only what they send

A JSON body with one field reset every checkbox in the sections it touched:
the MQTT bridge's brightness slider turned off disable_hardware_pulsing,
inverse_colors, show_refresh_rate and use_short_date_format, and a
timezone-only save turned off web-UI autostart and weekly auto-updates.
Missing-means-unchecked now applies only to form posts: form-encoded bodies
and the v3 forms, which mark themselves with a hidden __form_section input.

Also on the config routes:
- vegas_min/max_cycle_duration no longer match the generic *_duration rule,
  so they stop landing in display_durations and a blank one no longer
  rejects the whole Display save;
- saving from the Raw JSON editor calls start_setup_if_needed like the
  General form, so enabling auto-update there finishes its setup;
- the schedule and dim-schedule POSTs accept the per-day days.<day> shape
  their GETs return, as well as the flat form keys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): install plugin dependencies from the configured plugins directory

install_plugin_dependencies.sh scanned only plugins/, but the Plugin
Store installs into plugin_system.plugins_directory (default
plugin-repos), so the documented "Recommended" fix found 0 plugins on
every store install. It now reads plugins_directory from
config/config.json (relative to the project root or absolute, default
plugin-repos) and also scans plugins/ for dev symlinks, installing a
plugin reached through both only once.

With set -e alone, `pip ... | tee` took tee's exit status, so a failed
pip install was reported as success; set -o pipefail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: replace stale API names, line numbers and the api_v3.py path

- ADVANCED_FEATURES: StreamManager methods that exist
  (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the
  real on-demand status envelope ({status, data: {state, service}})
- app.py:199 / :144 / :607-619 line citations and
  web_interface/blueprints/api_v3.py (now a package) replaced with file and
  function names in ADVANCED_FEATURES, CONFIG_DEBUGGING,
  PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE,
  PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README
- CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use
  /config/raw/main to replace the file; describe where validation runs
- TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints
  usage)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): verify the web interface that actually ships, on port 5000

verify_installation.sh failed every healthy install: it required the
long-removed web_interface_v2.py and looked for a listener on port 5001,
while the web interface binds 5000 (web_interface/start.py). It now
checks the files ledmatrix-web.service runs (start_web_conditionally.py,
web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the
same 5001 port in its listen check, HTTP probe and printed URLs.

Port matches are anchored so :50001 no longer counts as :5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): one display-size contract: display_manager.width/height

CLAUDE.md (#580) says to read display_manager.width/height because
matrix is None when hardware init fails; the development guide, the
safety-harness doc and two DisplayManager docstrings still recommended
matrix.width/height. The bundled starlark-apps plugin read matrix.width
unguarded, so its magnify recommendation and frame scaling raised in
fallback mode (e.g. after the Pi 5 hardware refusal).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): make install_service.sh --help print usage instead of installing

install_service.sh parsed no arguments, so `sudo ./scripts/install/
install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md)
rewrote ledmatrix.service, ledmatrix-web.service and both update-verify
units and enabled/started them. It now handles -h/--help (usage, exit 0,
no changes) and rejects any other argument with exit 2 before doing
anything. Running it with no arguments, as first_time_install.sh does,
is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scroll): describe the fixed-step model and document frame_hold

Since #545 a crisp speed from scroll_config.configure() makes the helper
advance a fixed whole-pixel step per presented frame with no clock, and the
display manager's frame hold is part of the speed. The docs still described
the removed wall-clock model:

- scroll_config's module and configure() docstrings said speed is applied
  in time-based mode and that omitting the hold "falls back to fractional
  pixels"; omitting it actually runs the scroll frame_hold times too fast.
- SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both
  modes, and read a 20 ms stats median as missed refreshes although that
  is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step,
  the hold-dependent healthy median, that target_fps plays no part, and
  that a hand-added scroll_pixels_per_second loses to a schema-default pair.
- PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling)
  without frame_hold; it now documents the parameter (core 3.4.0) with a
  configure() + set_scrolling_state example.
- update_scroll_position/set_scroll_speed and set_scrolling_state
  docstrings say the same.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark target_fps legacy; describe what Vegas scroll_delay does

- General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after
  the sports_scroll fix nothing in core scrolling reads it. The field and
  its API validation stay so saved configs and plugins that read
  global_config['target_fps'] keep working. CONFIG_REFERENCE says the same.
- Vegas frame_based_scrolling/scroll_delay were described as frame-count
  stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based
  mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and
  still advances by elapsed time, so the applied speed is
  clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The
  config comments, render_pipeline comment and CONFIG_REFERENCE rows now
  say so. No behaviour change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(deps): describe how plugin dependencies are really installed

The guides said the web service runs as root, that installs pick --user
from os.geteuid(), and quoted a warning and a
PluginManager._install_plugin_dependencies() method that don't exist. The
web unit runs as the installing user; store installs go through
install_requirements_file() and sudo safe_pip_install.sh (root), with a
user-level fallback that says so, and load-time installs run in the
display service's own (root) interpreter.

Manual paths now use the configured plugins directory (plugin-repos/ by
default) instead of plugins/, which store installs no longer use, and
install_plugin_dependencies.sh is described as scanning that directory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): count local changes one way for the preflight and the pull

The automatic update's preflight ignored mode-only changes and anything
whose status line contained plugins/ or plugin-repos/, then promised
"Automatic updates will not stash your changes". perform_core_update
used plain git status (modes count) and ignored only 'plugins/', then
ran 'git stash push -- :!plugins', which nothing ever pops. So an edit
to a bundled plugin under plugin-repos/, or the installer's chmods on
tracked scripts, passed the preflight and was stashed away for good.

- auto_update.local_changes() is the one predicate both use:
  core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/
  excluded by leading folder rather than substring (a core file under
  web_interface/static/v3/js/plugins/ now counts).
- Update Code's explicit stash leaves out both plugin folders; the
  pull's --autostash carries their edits and mode changes across and
  reapplies them.
- The automatic updater calls perform_core_update(stash_local_changes=
  False), which refuses instead of stashing edits that appeared after
  the preflight; update_core reports that as 'blocked'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): diagnostics follow the web autostart default and api_v3 package

#556 made a missing web_display_autostart mean "start" (only an explicit
false/off keeps the web interface down), but the diagnostics still said
otherwise: diagnose_web_ui.sh reported a missing key as "defaults to
false", diagnose_web_interface.sh said the web interface "will not start
unless this is set to true" and recommended enabling it, and
debug_web_manual.py printed False. Troubleshooting a down web UI pointed
users at a non-cause.

Both shell scripts now evaluate the setting with the launcher's own
autostart_enabled() (inline fallback if it cannot be imported) and report
on / off / not set (on) / unparseable config; debug_web_manual.py uses
the same function. They also check web_interface/blueprints/api_v3/
__init__.py: api_v3.py became a package in #553, so every healthy
checkout was reported as missing a file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(install): what install_service.sh installs; verify script port; no sudo for --help

install_service.sh installs and starts ledmatrix, ledmatrix-web and the
update-verify units, not only ledmatrix.service (systemd/README.md,
README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help'
as a harmless check; it now shows --help without sudo and warns what a
real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh
checks the web interface on port 5000.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): note update-all, plugin system settings and script fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): size the preview after orientation and pixel mappers

display_geometry.physical_size claimed to give DisplayManager's answer but
only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured
after the library's pixel mappers, so a Rotate:90 / orientation 90 chain
previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for
128x64.

Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc
does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown
names leave it alone), and move the orientation composition here so
DisplayManager and the preview share it. The module docstring no longer
claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH.

Audit finding F18.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display): refuse settings the rgbmatrix library aborts on, on every board

The library answers several settings with a NULL matrix or abort() rather
than an error, so the display service crash-looped (Restart=on-failure)
instead of reaching fallback mode: rows above 64, chain_length above 255
(uint8_t binding setter, documented as "no upper limit"), a misspelled
hardware_mapping, and parallel 2-3 on a single-output mapping, reachable
from the Display form on the default adafruit-hat(-pwm) mapping. #586 only
guarded the Pi 5 subset.

- src/matrix_support.py holds the rules for every board (Options::Validate
  ranges, binding integer types, mapping names and outputs from
  lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the
  API's numeric ranges.
- DisplayManager checks them before building options and raises
  MatrixSettingsRefused, so a hand-edited config falls back with a logged,
  reported reason. Emulator mode only warns.
- The config API refuses them with a 400 naming the setting; combinations
  are checked against stored values but reported only when the request
  sets a field involved.
- The hardware status file gains "cause" (settings/library/forced). The
  fallback log and Display banner give the Pi 5 rebuild hint only for a
  library failure instead of rebuild + gpio_slowdown advice for every
  failure; one Pi 5 slowdown recommendation (1-3, start at 1).
- The Display form offers classic/classic-pi1 and orientation 90/270 and
  renders any other stored mapping selected with a warning, so an
  unrelated save no longer rewrites them; the API accepts 90/270.

Audit findings F03, F16, F19, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(display): library limits, template defaults and Pi 5 slowdown

- rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs,
  classic/classic-pi1 mappings and orientation 90/270 documented.
- Defaults are the config.template.json values: config migration adds
  missing keys from the template, so the listed "code defaults" never
  applied.
- One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode,
  starting at 1.
- Troubleshooting describes the refused-settings fallback, and CHANGELOG
  corrects the Unreleased "no upper limit" entry.

Audit findings F19, F20, F21.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py opens the panel with the service's options

--measure and --demo built RGBMatrixOptions from a private copy of the
display service's builder that had drifted: gpio_slowdown came from
display.hardware (default 2) instead of display.runtime (default 3), and
rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors,
pixel_mapper_config and orientation were skipped, with different defaults
(hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown
was measured -- or garbled -- in a setup the service never drives.

The option filling in DisplayManager._setup_matrix moves, unchanged, into
DisplayManager.apply_matrix_options(options, config), which _setup_matrix
calls and the script reuses (overriding only limit_refresh_rate_hz for
--measure). The script now loads the whole config rather than the hardware
block. Tests pin the script's options to the service's attribute for
attribute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scripts): scroll_speeds.py recommends keys the resolver honours

The ladder ended by telling users to set
display_options.scroll_pixels_per_second. scroll_config ranks that key
below the scroll_speed + scroll_delay pair, deliberately, and several
plugin schemas default the pair into config, so the advised key was
silently ignored (a schema-default 1/0.02 pair plus an advised 66 still
resolved to 50 px/s).

The advice is now the pair that selects the crisp speed exactly
(pixels_per_frame every frame_hold/refresh seconds), explains that the
pair outranks scroll_pixels_per_second, and gives the scoreboards'
per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the
printed pair over a schema-default pair and check it lands on the
advertised speed and hold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula

- SPORTS_UNIFICATION.md still presented honouring global target_fps as
  sports_scroll's added behaviour and its one user-visible gain; note that
  it was withdrawn because it had become a speed multiplier.
- ADVANCED_FEATURES.md gave Vegas scrolling as
  (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s
  by elapsed time, through a 0.1-5 px per scroll_delay clamp when
  frame_based_scrolling is on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): scroll model fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(dev): link-github links plugins from the ledmatrix-plugins monorepo

link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git,
and those per-plugin repositories no longer exist: official plugins are
directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the
monorepo once into the dev directory, finds plugins/<name>,
plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and
links it under its manifest id. With an explicit repo URL it still links a
single-repository plugin as before.

dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a
fork), plus plugins_repo and plugins_branch; github_pattern, which was
documented but never read, is dropped and warned about. Ships
dev_plugins.json.example and git-ignores dev_plugins.json, both of which
the guide promised. Reading JSON falls back to python3 when jq is missing
(get_plugin_id silently returned nothing without jq).

update/status/list find the git checkout above a monorepo plugin
directory (its .git is not in the plugin dir), and update pulls a shared
checkout once. status no longer exits 1 when nothing is broken.

Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration,
workflow, store integration, hello-world link, submission), and the
nonexistent scripts/git-hooks/pre-push-plugin-version and
scripts/bump_plugin_version.py replaced with the real rule: bump the
manifest version and run update_registry.py. scripts/dev/README.md and
CLAUDE.md updated to match.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(scripts): monorepo workspace layout; fix_perms and install READMEs

MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin;
setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into
plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the
workspace file opens LEDMatrix plus ../ledmatrix-plugins.

scripts/fix_perms/README.md listed cache directories
fix_cache_permissions.sh never touches and a 'ledmatrix' service user
that doesn't exist (also in scripts/install/README.md); adds
safe_pip_install.sh. install/README: install_service.sh installs the web
and update-verify units too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): keep the rollback's pip retries inside the unit time limit

The health check reinstalled the previous requirements by trying the
next bash path after any failure, including a 600 s pip timeout. Two
files, two paths: up to 40 minutes of pip alone, while systemd stops
ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the
rollback half-way and leaving the update 'verifying' until the web UI
calls it lost.

- Like permission_utils.install_requirements_file, only a sudo refusal
  moves on to the next bash; a pip that ran and failed or timed out is
  not repeated. The refusal wording is one list
  (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only
  verifier and pinned equal by a test.
- All reinstalls in one rollback share a 600 s budget.
- WORST_CASE_SECONDS adds up every timeout on the longest path (27.5
  min); a test holds it under the unit's TimeoutStartSec and that under
  the web UI's VERIFY_LOST_SECONDS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools

Plugin config was prepared differently depending on how it arrived:

- JSON POST /plugins/config built a partial body on schema defaults, so
  {"enabled": true} reset every other setting of the plugin. It now merges
  onto the stored section first, as the form path already did.
- Legacy-boolean normalization (#588) ran only at load: GET /plugins/config
  returned the raw boolean, posting it back failed validation, and hot
  reload handed plugins the raw section (a legacy dynamic_duration: true
  came back as a boolean). schema_manager.prepare_plugin_config (normalize,
  then defaults) is now used by PluginManager.load_plugin, both save paths,
  GET, the save notifications and DisplayController's hot-reload callback.
- The JSON save's filter kept only enabled/display_duration/live_priority
  and dropped a submitted skin, skin_options or vegas_* tuning key. There
  is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES,
  used by validation and by the save filter; PluginManager's
  CORE_OWNED_CONFIG_KEYS is its vegas subset.
- Plugin sections posted to /config/main were stored verbatim, including
  values /plugins/config rejects. They now go through the same preparation
  (_prepare_plugin_config_for_save, extracted from save_plugin_config), and
  a failing section rejects the whole save before anything is written.
- dev_server read only top-level defaults and let a schema enabled:false
  win; build_full_config shallow-merged overrides, dropping sibling
  defaults; the harness extracted defaults differently from the device.
  loading.build_config now uses the device's extraction and preparation,
  and dev_server, check_plugin, render_plugin and the harness all use it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt-bridge): brightness changes apply live and touch nothing else

The display service's hot reload applies a saved brightness within a few
seconds, and /config/main no longer resets other display settings on a
brightness-only JSON body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): automatic update hardening

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI

It described web_interface_v2.py and index_v2.html (both gone), client-side
form generation, one POST per field with {key, value}, and 'no nested
objects'. The v3 UI renders plugin forms server-side from the schema
(pages_v3 partial + plugin_config.html macros, nested sections and
x-widgets), posts the whole form once, and save_plugin_config() merges onto
the stored section, validates, splits x-secret fields and notifies the
plugin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(mqtt): brightness saves apply via hot reload and leave other settings alone

The bridge README said brightness is applied on the display's next
restart; the display controller's config hot reload applies it within
seconds. It also now states that the bridge's partial JSON save changes
only brightness (the /config/main merge fix in this change set).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(update): don't log pip's output from the health check's reinstall

pip can echo a private index URL with embedded credentials;
permission_utils redacts it, the stdlib-only verifier cannot, so it
logs the exit code only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(config): mark the plugin_system toggles as unused legacy keys

auto_discover, auto_load_enabled and development_mode are read by
nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE
said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and
the REST reference listed them as live settings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): docs and developer tools group

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): legacy plugin-system toggles no longer count as a General save

auto_discover, auto_load_enabled and development_mode have left the General
form, so a post carrying only one of them is not a general-settings save and
must not treat web_display_autostart and auto_update as unchecked. The
plugin_system block itself is left as on main for the branch that reworks it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): config-save and plugin-config preparation fixes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(claude): re-check matrix_support.py rules when the library submodule is bumped

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address Codacy findings on the core audit PR

- plugin_manager.prepare_plugin_config: when the fallback legacy-boolean
  pass also fails, log a warning instead of a bare except/pass.
- api_client.js: request() refuses any endpoint that is not a plain path
  under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace,
  control characters) with INVALID_ENDPOINT before calling fetch(), and
  plugin ids are URL-encoded wherever they are put into a URL (also in the
  app-shell batch load).
- test_update_all.js: pins both against the shipped client.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): check endpoint control characters without a control-character regex

Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags
the \x00-\x1f range in checkEndpoint's regex. Test the char codes
instead; the endpoints refused are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(auto-update): make the seed script executable on disk, not only in the index

On Linux Repo.publish() commits with -a, which recorded scripts/run.sh
as 100644 upstream because the seed file was never chmod +x. The pull
then brought in the same mode the installer chmod had made locally, so
installer_chmod saw no mode change left to check. The updater was fine:
with the upstream commit at 100755 the --autostash carries the device's
chmod across. Verified under Linux (WSL, git 2.43): the old helper fails
exactly as CI did, the fixed one passes all 63 tests in the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 16:37:29 -04:00
ChuckandClaude Opus 5 7e5967e160 fix(web): find installed plugins before anything has discovered them (#594)
The web process discovers plugins lazily: plugin_manifests is empty until
some endpoint calls discover_plugins(). Three routes consulted it without
discovering, so they misbehaved for as long as nothing else had run --
which, after every ledmatrix-web restart, is until someone opens the
dashboard:

- POST /display/on-demand/start answered 404 "Plugin <id> not found"
  (or "Mode <mode> not found"). Measured on a rig: 404 for over three
  minutes after a web restart, until GET /plugins/installed ran. The
  browser UI loads the plugin list first, so API-only callers (the Home
  Assistant MQTT bridge, scripts) are the ones who hit it.
- POST /plugins/toggle answered 404 "Plugin not found".
- POST /config/main did not recognise a plugin section, so it skipped
  secret separation and merged the section as-is: the plugin's API key
  was written to config.json in plain text instead of config_secrets.json.

Add _discovered_plugin_manifests(), which discovers when nothing has been
yet, and rescans once when a specific plugin id (or, for on-demand by
mode, a mode) is not found, so a plugin installed since the last scan is
found too. _installed_plugin_ids() now uses it.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 12:38:55 -04:00
ChuckandClaude Opus 5 f475038895 fix(cache): web UI can read what the display service caches again (#593)
* fix(cache): web UI can read what the display service caches again

ledmatrix-web.service carried CacheDirectory=ledmatrix. With User= set to
the installing user, systemd re-owns /var/cache/ledmatrix and everything
in it to that user and its primary group whenever the directory's owner
differs -- for a directory root created, on the first start. That erased
the root:ledmatrix setgid layout the installers set up, so every file the
display service (root) wrote afterwards was root:root 0660 and unreadable
by the web interface:

  WARNING - Permission denied loading cache for display_current_state ...

Since #547 install_service.sh renders the web unit from the template, so
every fresh install hit this. Measured on one rig: 392 unreadable files,
and the web UI's display status, on-demand state and plugin health empty.

Existing installs only receive `git pull`, never a reinstalled unit, so
the fix for them is in the code the root display service runs:

- DiskCache.set gives each file the directory's group (when the directory
  is group-writable) and 0660 on the open descriptor before the rename,
  independent of setgid. This also closes a window where a fresh file was
  visible as mkstemp's 0600.
- DiskCache.share_existing_files repairs files an older version left
  behind, once per process from the cleanup thread. It works through
  O_NOFOLLOW descriptors and skips hard links and other users' files: the
  directory is writable by the web user, and root must not be steered
  into changing a file outside it.

For new installs, the web unit drops CacheDirectory=/CacheDirectoryMode=,
and install_web_service.sh stops replacing an existing directory's
ledmatrix group with the user's group.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): on-demand and current-display status read the display's latest state

Found testing the cache-permission fix on a rig: once the web interface
could read display_on_demand_state at all, /display/on-demand/status kept
answering "active" for over 100 seconds while the file on disk said
"idle". Both status routes read the display service's keys through the
web process's memory tier, which serves the first copy it read for the
full max_age (120s). Read them with memory_ttl=0, as every other
cross-process reader (plugin health/metrics, the on-demand mailbox)
already does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(install): re-group the cache dir whenever the web user is outside its group

install_web_service.sh replaced an existing cache directory's group only
when it was root's. A directory in any other group the web user is not a
member of -- root:ledmatrix, for a user who is not in ledmatrix -- was left
alone, and every file root wrote there stayed unreadable to the web
interface. Replace the group whenever the installing user is not in it.

A directory whose group the user is already in (ledmatrix, or the user's
own group where CacheDirectory= left it) is still left as it is: re-grouping
a working directory strands the files already in it on the old group.

When the group does change and root-owned JSON files carrying the old group
are present, try-restart ledmatrix.service so DiskCache.share_existing_files
re-groups them through its symlink- and hard-link-safe path, rather than a
recursive chgrp.

Verified under WSL's systemd for seven directory states (user group,
ledmatrix member, ledmatrix non-member with and without root files,
root:root, missing, unnamed gid); the previous version left the non-member
case unchanged.

Addresses CodeRabbit review on #593.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 12:38:39 -04:00
ChuckandClaude Opus 5 f9b3d6ae52 fix(web): accept every panel size and row address type the rgbmatrix library does (#586)
* fix(web): accept every panel size and row address type the rgbmatrix library does

The Display form capped columns at 128 and chain length at 24, and its
submit handler (fixInvalidNumberInputs) rewrote anything larger to the cap,
so wide panels and long chains silently saved as the wrong size. The config
API checked none of the hardware numbers, so values the library rejects (odd
rows, parallel 4, PWM dither bits 3) saved and the matrix then refused to
start.

- Form limits now match the pinned library: rows even 8-64, cols >= 16 and
  chain_length >= 1 with no upper bound, parallel 1-3, PWM dither bits 0-2,
  PWM LSB nanoseconds 50-3000.
- save_main_config rejects out-of-range rows, cols, chain_length, parallel,
  brightness, scan_mode, pwm_bits, pwm_dither_bits, pwm_lsb_nanoseconds and
  gpio_slowdown with a 400.
- A stored gpio_slowdown or pwm_dither_bits of 0 renders as 0 instead of the
  default, so saving the tab no longer overwrites it.
- Row Address Type offers 5 (SM5368 / B707 row shift register). Verified on a
  Waveshare 96x48 V2 (24S-A1) on a Pi 4 with the Adafruit Triple LED Matrix
  Bonnet: rows 48, cols 96, row address type 5, BGR, GPIO slowdown 8.
- Help text and docs: FM6124-family panels use Panel Type Standard; on a Pi 5
  the library supports only row address types 0 and 2.

No change to the rpi-rgb-led-matrix submodule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): drop the rows cap and document every display setting accurately

Rows: no upper limit in the form or the API. Still even and at least 8. The
current rgbmatrix library rejects more than 64 per panel, so a larger value
saves but the matrix won't start; the help tip, README, config reference and
troubleshooting section all say so, and nothing here needs changing if the
library lifts the limit.

limit_refresh_rate_hz: the form accepts 0 (the library's "no cap"), a stored
0 no longer renders and re-saves as 120, and the API rejects negatives.

pwm_dither_bits stays 0-2: the library rejects 3 and 4, so the old form's
0-4 only ever let users save a config the display couldn't start with.

Docs and help tips, checked against the pinned library and its README:
- panel_type and rp1_rio get README entries
- show_refresh_rate prints to stdout; it never drew on the panel
- dither bits raise the refresh rate; the tip said they lowered it
- scan_mode is about interlacing at low refresh, not wrong colours
- disable_hardware_pulsing: hardware pulsing needs OE on GPIO 18 and the
  onboard sound driver off; software timing makes rows flash brighter
- gpio_slowdown guidance agrees between the README and the UI
- all 22 multiplexing values listed; every numeric setting states its range
- troubleshooting for a blank panel after a settings change, jumping rows
  and brightness flashes

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): reject true and 5.5 for row_address_type and multiplexing

Both still went straight through int(), so a JSON true saved as 1 and 5.5
as 5. They now use the shared hardware range check like the other panel
fields. Review feedback on #586.

Also: the RP1 Backend tooltip said it is ignored on Pi 3/4 (it is ignored
on every model but the Pi 5), and the README gave the dynamic-duration
default cap as 90s; the code default is 180s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: refuse matrix settings a Raspberry Pi 5 can't drive

On a Pi 5 the pinned rgbmatrix library drives the panel through the RP1
chip, and that path supports only row address types 0 and 2, parallel 1-3
and the regular / regular-pi1 / classic / adafruit-hat(-pwm) mappings
(Rp1PioConfigSupported in lib/rp1/rp1_pio_backend.cc). For anything else
CreateFromOptions returns NULL; the Python binding doesn't check, so the
display process crashed on its first call into the matrix and systemd
restarted it into the same crash every 10 seconds.

- src/pi5_matrix_support.py: the rule and Pi 5 detection, matching the
  library's /proc/device-tree/model check
- DisplayManager raises before creating the matrix, so it is a logged init
  failure (reported by /api/v3/hardware/status) and fallback mode
- the config API rejects those settings on a Pi 5 when a request sets
  row_address_type, parallel or hardware_mapping
- the Display form offers only row address types 0 and 2 on a Pi 5, and
  warns when a stored value can't be used
- CLAUDE.md: re-check the rule whenever the submodule is bumped

Review feedback on #586.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 19:17:48 -04:00
ChuckandClaude Opus 5 9616a5a054 fix(web): auto_update and other core settings are not orphaned plugins (#589)
v3.4.0 shows "Plugin Config Warning - In config but not installed:
auto_update. Reinstall via the Plugin Store, or remove these entries from
config.json." auto_update is the core weekly-update setting from #581.
Reconciliation treated every top-level dict not in its private
_SYSTEM_CONFIG_KEYS list as a plugin id, and #581 could not know to extend
that list.

- Move core top-level keys into src/core_config_keys.py (CORE_CONFIG_KEYS)
  and use it in reconciliation. Tests fail if a config.template.json key or
  a key written by the general-settings save is missing from it.
- A secrets-file key only counts as a non-plugin when no installed plugin
  has that id. Plugin secrets are namespaced by id, so installed plugins
  with secrets were reported as missing from config on every run.
- still_unresolved() drops "not on disk" findings whose id is no longer a
  plugin entry in config, so a stored verdict clears without a restart.
- A plugin whose id is a core key is skipped with a warning, and the fix
  never writes a plugin stub over or in place of a core setting.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:19:39 -04:00
ChuckandClaude Opus 5 c200b5837d fix(plugins): normalize legacy boolean settings before schema validation (#588)
The news plugin's schema turned global.dynamic_duration from a boolean
into an {enabled, min_duration_seconds, ...} object. Installs that have
not saved the news settings since still hold `true`, so every start
logged "Plugin news config does not match its schema (loading anyway):
Field 'global.dynamic_duration': Expected type object, got bool" and
flagged news degraded.

The settings form already reads such a boolean as {"enabled": <bool>}
(render_nested_section in plugin_config.html) and the next save writes
the object. The loader did not. It now applies the same rule before
merging schema defaults and validating, so the defaults fill in the rest
of the object and the plugin receives it in the new shape.

The rule lives in schema_manager.legacy_bool_as_object /
normalize_legacy_booleans. It applies at any depth of nested objects
but not inside arrays, matching the form, and only to a real bool under
an object-typed property with an `enabled` child. Every other mismatch
still warns. A parity test renders the template macro against the helper
so the two cannot drift.

Nothing is written to config.json at load: the normalization is in
memory, and the next save of the plugin's settings persists the object.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:09:23 -04:00
ChuckandClaude Opus 5 9f2743471c fix(web): update-all skips Starlark apps and no longer misses plugins (#587)
* fix(web): update-all skips Starlark apps and no longer misses plugins

Check & Update All posted every entry from /plugins/installed to
POST /plugins/update, including the virtual starlark:<app_id> entries
that list installed Starlark apps. The store manager cannot find those,
so each answered 500 "plugin not found". Update-all now sends only
plugin ids (install_manager.js, and the older app-shell.js copy), and the
route answers a starlark: id with a 400 saying it is a Starlark app.

A request that got no HTTP answer was recorded as failed and never sent
again. On a device, a web-service restart mid-run killed the in-flight
request and refused the next one, stock-news, which was left on 2.6.2
with 2.8.0 available. Such requests are now re-sent with backoff
(about 30s) before being reported as failed. HTTP error answers are not
retried.

Tests: test/js/unit/test_update_all.js (run from pytest via
test/web_interface/test_update_all_plugins.py so CI covers it) and the
route contract for starlark: ids.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): walk update-all retry delays without indexed lookup

Codacy's ESLint security/detect-object-injection rule flagged
retryDelays[attempt] as a High issue. The index was a bounded loop
counter over a fixed array, but shifting a per-plugin copy of the
schedule gives the same backoff without the pattern. No behaviour
change: test/js/unit/test_update_all.js (21) and
test/web_interface/test_update_all_plugins.py (7) pass unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 18:09:11 -04:00
ChuckandClaude Opus 5 fddb0e06db feat(web): honour x-display: hidden in plugin settings (#585)
* feat(web): honour x-display: hidden in plugin settings

Plugins keep deprecated and internal keys declared so stored configs keep
validating (weather api_key/radar_zoom, countdown's auto-generated row id),
but the settings form drew them as live controls.

A property marked "x-display": "hidden" -- or an object whose children are
all hidden -- now gets no control at any depth: top level, nested sections,
Advanced Settings (not counted either), array-table columns and the row
editor. A hidden top-level key is not reported in __rendered_section.

Saving never changes a hidden value. Plain and nested fields aren't posted,
so the save's deep merge keeps them; _set_missing_booleans_to_false skips
hidden booleans at every depth. A posted array row replaces the stored item,
so hidden row properties are carried as JSON-encoded hidden inputs and
decoded exactly on save (an id "1" stays a string). New rows get none.
JSON API saves are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): read hidden row keys without dynamic property access

Build the set of x-display: hidden item properties once and look values up
through Object.entries, instead of indexing objects by a variable key on
the lines this branch added (Codacy: object injection sink, 6 warnings).
Behaviour is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 15:47:06 -04:00
ChuckandClaude Opus 5 869e36fb2f feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback

A General-tab toggle (off by default) checks for and installs LEDMatrix and
plugin updates once a week, overnight in the configured timezone.

- Pre-update checks skip (and report) instead of forcing: local edits or
  commits, merge/live rebase, no upstream, low disk, missing health check, or
  a version that was already rolled back. An abandoned rebase (HEAD back on a
  branch) is cleared, since it would otherwise block every pull.
- The pull reuses the Update Code path (now perform_core_update(), which
  reports dependency install failures as data).
- ledmatrix-update-verify.service, started via a .path unit from a request
  file, restarts the services from its own cgroup, requires them to come up
  and stay up, and otherwise resets to the previous commit and reinstalls the
  previous requirements. It runs a copy of the checker taken before the pull.
- No SSH needed: switching the toggle on restarts the display service, which
  (as root) installs the two units from the repo templates for the web user.
  first_time_install.sh installs them too and takes --enable-auto-update /
  LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh).
- Plugins update after the code passes its check; failures, blocks and
  rollbacks raise an Overview banner and show under the toggle.

Tested end to end on a Pi: web-UI setup, a good update, a broken web service
and a broken display (both rolled back), a blocked local edit, and an
abandoned rebase found on the device.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(auto-update): address static-analysis findings

- Replace the subprocess.CompletedProcess the verifier fabricated for a
  command that could not start with a plain namedtuple; nothing is executed
  there, but the scanner flags any CompletedProcess built from variables.
- Mark the subprocess imports with the repo's standard B404 annotation (all
  calls are list-form argv, no shell).
- Mark the rollback-failed message as not SQL (B608 matched its wording).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): CI failures on Linux

- Keep the setup result when chown fails. CI runs as a non-root user, where
  chown to the web user raises; that discarded the result file, so the
  General tab would never learn whether setup worked. Regression test added.
- Register the two new /api/v3/system/auto-update routes in the URL map
  snapshot.
- Use utility classes app.css defines (space-y-1, hover:text-red-600).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): address review feedback

- Health check: a failed restart command no longer lets the check run
  against the still-running old process; it counts as a failure (and after a
  rollback, as a failed rollback). An unreadable restart count is never
  treated as stable, since a crash loop looks healthy between attempts.
- Installer writes the auto_update setting to a temp file and swaps it in,
  keeping mode and owner, so a running config watcher never reads a
  truncated config.json.
- Verify unit quotes its command-line paths (install folders with spaces);
  setup refuses folder names systemd would reinterpret (%, quotes,
  backslashes, control characters) and says so on the General tab.
- The auto-update status route no longer returns exception text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): keep error detail in the status route's 500

test_web_error_detail requires every 5xx handler to log the traceback and
return describe_exception(e), which redacts credentials, so failures are
diagnosable from the web UI. Dropping it for CodeQL broke that policy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): dismiss route rejects non-object JSON with 400

A JSON array or scalar body made `.get('alert_id')` raise, returning 500.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(auto-update): let the app-wide handler answer status-route errors

CodeQL (py/stack-trace-exposure, #709) flagged the route's own except,
which returned describe_exception(e). web_interface/app.py's error handler
already logs the traceback and returns the same redacted detail for any
unhandled exception, so the local copy is removed: same response, no new
exception-to-response flow, and test_web_error_detail's policy still holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:58:57 -04:00
ChuckandClaude Opus 5 814c21de1c chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0 (#580)
* chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0

Skins: no current scoreboard plugin builds on src.base_classes, so the only
skin hook (SportsCore._render_game) never runs. The plugin schema endpoint no
longer injects the Visual Skin dropdown, the store hides and refuses
"type": "skin" registry entries, and GET /api/v3/skins reports
supported: false with a message. Stored skin config still loads and saves.
src/skin_system/ and its tests are unchanged apart from the support flag.

Docs: check_plugin.py/render_plugin.py examples use --plugin; document
BasePlugin.get_update_interval() and its interaction with the manifest
update_interval; CLAUDE.md drops the stale template line number and
recommends display_manager.width/height.

Preview size: new src/display_geometry.py holds the size computation and
defaults DisplayManager uses (double-sided applied, chain_length default 2).
The web preview, /display/current, Starlark magnify default, sync handshake
and two dev scripts use it.

Release: __version__ 3.4.0, CHANGELOG 3.4.0 section plus a 3.3.0 tag note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address CodeRabbit review on #580

- Preview fallbacks (SSE stream and /display/current) use logical_size({})
  (128x32, the shared default) instead of a hard-coded 128x64.
- display_geometry treats a non-mapping display/hardware block as missing,
  so a malformed config.json falls back to defaults instead of raising
  AttributeError (which turned the Starlark render into an HTTP 500).
- Docs: the static update interval falls back manifest -> plugin config
  -> 60s, in both the API reference and the architecture spec.

Not taken: validating double_sided copies against chain_length/parallel.
An orientation Rotate: or U-mapper pixel mapper decides which axis panels
lie on, so the counts would reject working setups (the existing
vertical-split test is one).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(display_geometry): a non-finite hardware size raises ValueError, not OverflowError

CodeRabbit flagged the Starlark magnify default in
_standalone_render_starlark_app for truthy non-mapping display values. That
case was already handled by a9e1bd0b (_display/_hardware treat a non-mapping
block as missing, covered by test_non_mapping_display_config_uses_the_defaults),
and the magnify it produces from the 128x32 defaults is the same as from 64x32.

Checking the same path found one input that still escaped: Python's JSON
parser accepts Infinity, and int(inf) raises OverflowError, which neither the
Starlark path (TypeError, ValueError) nor the preview stream in app.py caught,
so a hand-edited "rows": Infinity returned HTTP 500. physical_size now raises
ValueError for it, matching its documented contract, so every caller's
existing fallback applies. DisplayManager already caught Exception.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 18:36:58 -04:00
ChuckandClaude Opus 5 6d1cbfb70b fix(web): plugin config page survives stored values the schema outgrew (#578)
Two stored shapes broke the config form:

* A scalar under a field that is now an object. News' dynamic_duration
  was a boolean and is becoming an object; render_nested_section did
  `key in true` and the whole page failed to render. Look into dicts
  only, and carry a legacy boolean over as the object's `enabled`, so
  the next save upgrades it without switching the feature off.
* A custom feed logo with a path but no id. The template always emitted
  an empty `logo.id` input, which the save route parsed to null, failing
  the id's string type on every save. Emit it only when there is an id,
  as custom-feeds.js already does.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 14:03:15 -04:00
ChuckandClaude Opus 5 dac71afedc fix(web): full-height plugin config form and full-width plugin card descriptions (#573)
* fix(web): let the plugin config form use the full page height

The form wrapper has carried `max-h-96 overflow-y-auto` since #145, but
the class was a no-op until #568 defined `.max-h-96` in app.css. That
silently capped the whole config form at 24rem with a nested scrollbar.
Drop the cap so the form flows naturally and the page scrolls.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): give installed plugin card descriptions the full card width

The enable/disable toggle was a flex sibling of the whole text column
(name, metadata, description), so it reserved its width for the full
height of the card body. Descriptions wrapped into a narrow strip,
leaving blank space under the toggle and making cards very tall.

Move the toggle into a header row with just the name and badges, and
render the metadata and description below at full width.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:27:00 -04:00
ChuckandClaude Opus 5 a6b3384032 fix(web): show the action script's error in the file-manager widgets (#574)
A failing plugin action returns a 400 whose JSON body carries the
script's own message, but both file-manager widgets threw it away:
plugin-file-manager's toggle always said "Toggle failed", and
json-file-manager's request helper threw "Server error 400" before
reading the body. That hid of-the-day's "Category ... not found in
config", which is why its toggles looked broken for no reason.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:26:46 -04:00
ChuckandClaude Opus 5 11bf39cd66 fix(web): two plugin config saves that always returned 400 (geochron, news) (#575)
* fix(web): render widget-less arrays of objects as a table, not comma text

An array of objects with no x-widget (geochron's `cities`) fell through to
the comma-separated text input. Jinja joined each item as a Python dict
repr, the save route read them back as a list of strings, and the schema
rejected them -- so every save of the plugin returned 400 "Configuration
validation failed", whatever setting was changed.

Default such arrays to the existing array-table widget, which already
edits arrays of objects and posts `field.N.key` inputs the save route
rebuilds into a list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): don't leave an empty object stub in array items on save

The unchecked-checkbox pass walked into every nested object of an array
item looking for booleans, creating it when absent. A news custom feed
with no logo came out with `logo: {}`, which fails the logo's
`required: [id, path]`, so every save of the news plugin returned 400.

Recurse into a scratch dict instead and attach it only if a boolean was
actually set in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 13:04:11 -04:00
ChuckandClaude Opus 5 f9b1f87e8d fix: clamp colour components, and let the style editor actually take over (#569)
* fix(element-style): clamp out-of-range colour components instead of rejecting

A regression this framework shipped. The eight scoreboards used to read their
colours through sports_card.coerce_rgb, which clamps; routing them through the
shared element_color sent them through _normalize_color, which rejected any
component outside 0..255 and fell back to the default. So a configured
[999, -5, 20] -- a typo'd bright red -- rendered white instead of (255, 0, 20).

Their own test_element_text_colors.py caught it: one case of nineteen, in all
eight plugins, failing only once the core change reached main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): the style editor takes over its own blocks -- and gets to at all

Two defects, both found by rendering the real partial in a browser rather than
by reading the code.

It was losing a race to its own fields. The hand-off guard asked "do any
fallback controls differ from their server-rendered defaults?" as a proxy for
"is someone editing this?". But the fallback holds this block's own font
fields, and the font-selector widget populates them on the same 50ms timer --
so a plain page load, with nobody touching anything, raced into "dirty" and the
editor removed itself, leaving the 701-line accordion form it exists to
replace. Measured: seven customization.*.font selects dirty ~60ms after
injection, clean again by 400ms. The question is whether a *person* typed, and
event.isTrusted answers exactly that; the listeners now go on synchronously,
because the edit worth protecting can happen before initWidget runs.

It took over too much. Taking over removed the whole fallback section, but a
customization block can hold more than styling -- football keeps
favorite_result_colors there -- so that removed the only UI those fields had,
and the editor also rendered them as an element, giving every row an "enabled"
and three colour columns. Core now marks the blocks it recognises as styling
(the compact declaration already did; hand-written adoption did not), the
widget renders only those, and the template drops only the children the widget
reports owning.

Verified on football's real schema: 28 rows across four mode tabs, columns
Element/Font/Size/Colour/X/Y, favorite_result_colors still editable with its
ten inputs, no duplicated field names, no console errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): style editor no longer drops layout-only fields it never rendered

CodeRabbit flagged elementKeys() in style-editor.js: render() claims the
whole customization.layout child as the widget's own (removing it from the
generic fallback renderer, since posting the same offset twice is worse),
but elementKeys() only listed keys that also have their own top-level style
block. A hand-written schema can put a key under layout that never got one
-- a logo, a timeout indicator, a possession arrow with a position but no
font or colour -- and that key's only control silently disappeared: no row
in the style editor's table (elementKeys never listed it) and no fallback
section either (layout was removed wholesale).

elementKeys() now appends any layout-declared key not already covered by a
style element, so table() renders a row for it (layout columns only, no
style columns) and the wholesale layout ownership claim stays truthful.

Verified against current code before fixing. New regression test
(test/js/unit/test_style_editor_element_keys.js, following this repo's
existing eval-extraction pattern for testing widget JS without a browser)
fails against the reverted function and passes with the fix; added to
run_all.js and the suite table in test/js/README.md.

Full pytest suite: 4887 passed, 62 skipped, 2 failed -- both the
pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias gap on this sandbox,
identical on origin/main, unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dpg3HLWohdCUdzz2QNHanm

* fix(web): style editor no longer strands leaf-valued layout fields

A prior fix on this PR made elementKeys() append any layout-only key with
no style block of its own (a logo, a timeout indicator, a possession
arrow), so table() draws a row for it instead of losing it when the
wholesale `layout` claim removes the generic fallback. That covers a
layout-only key shaped like an object (x_offset/y_offset, ...), because
columnsFor() only ever produced columns from a key's *sub-fields*.

It missed the case where the layout-only key's own value is itself a
leaf -- a plain "show_logo" boolean directly under layout, no x/y object
underneath. elementKeys() still lists it (any row: no matching column),
so it renders as an uneditable blank row and its only control -- the
generic fallback checkbox -- is still gone. Confirmed by executing the
real widget's render() against a synthetic schema in Node (a DOM-stub
harness, not committed): the field's name never appeared as an <input>.

columnsFor() now gives such a leaf key a column keyed to itself
('layout-leaf'), and elementRow() binds it to the leaf's own path
(customization.layout.<key>, matching the name the fallback would have
used) instead of leaving every cell blank.

New regression test (test/js/unit/test_style_editor_layout_leaf_columns.js,
following this PR's existing eval-extraction pattern) checks the leaf
column is produced, is self-keyed, doesn't duplicate, and that a schema
with no leaf-valued layout key is unaffected; wired into run_all.js and
the suite table in test/js/README.md.

test/js/run_all.js: 84 + 6 + 6 = all suites passed (jsdom unavailable
here, DOM suites skip as before). Python suite untouched by this change;
test_style_editor_extra_fields.py, test_style_editor_save_roundtrip.py
and the one PIL-dependent style_editor_takeover.py case fail identically
before this commit -- missing flask/PIL in this sandbox, not this PR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): keep layout-leaf style-editor columns distinct from name collisions

columnsFor() keyed a layout-only leaf field's column by its bare field
name. If an unrelated element's style block or another element's layout
axis block happened to declare a sub-field with that same name, the
`!seen.has(key)` guard skipped creating the leaf's column, silently
dropping its only control again -- the same failure the leaf-column fix
was meant to close, just reached through a name collision (CodeRabbit
review on 324a7ea).

Key layout-leaf columns under a namespaced id so they can never be
shadowed by an unrelated column sharing their name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): CSS.escape() the owned key before it becomes a selector

container.dataset.ownedKeys round-trips schema property keys through a
DOM dataset attribute, and the takeover handoff spliced each one
straight into '[data-child-key="' + k + '"]' with no escaping --
inconsistent with this codebase's own convention elsewhere
(plugin-file-manager.js, app-shell.js's escapeCssSelector) for building
a selector from a dynamic value. A key containing a quote or backslash
would break the selector or be steerable; Codacy's static analysis
flagged this pattern (1 high ErrorProne finding on PR #569, current
head at the time) as a new issue, though its dashboard is unreachable
from this sandbox (egress to app.codacy.com is blocked) and the
check-run API returned no detail text -- verified and fixed by reading
the diff directly rather than the tool's own description.

Added a source-assertion regression test alongside this file's
existing ones (this behavior lives in an inline script no Python test
executes).

Full suite: 4888 passed, 62 skipped, 2 failed -- both the pre-existing
Europe/Kiev/Asia/Calcutta tzdata-alias gap in this sandbox, identical
on origin/main, unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): every advertised layout offset gets a control in the style editor

The editor took the whole layout section over but matched offsets to style
rows by exact key. A hand-written schema's two blocks were never named alike --
football styles score_text but positions score -- so of football's eleven
positionable things only status_text had a control. Score, odds, both logos,
timeouts, possession, down-and-distance, date, time and records were options
the schema advertised and the renderer reads, reachable nowhere in the UI.

Core now resolves each style element's layout key through alias_keys, the map
the resolver already reads offsets with, and records it as x-layout-key. The
widget reads that rather than carrying a second copy of the rules, and posts
under the key the schema declares: football's own offset reader looks up
layout.score, so a value saved as layout.score_text would be kept and never
drawn. Layout entries no style element claims get an "Other positions" table
with its own columns, in every mode panel as well as the base one, in the order
the plugin declared them.

Verified in a browser against football's real schema: 92 of 92 layout fields
(23 base, 23 per mode) rendered exactly once under their declared names, none
posted under a style key, no duplicated field names, favorite_result_colors
still editable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:48:36 -04:00
ChuckandClaude Opus 5 7e580dc005 fix(wifi): make Connect work from the setup AP (#571)
* fix(wifi): make Connect work from the setup AP

Joining a network from LEDMatrix-Setup has to take the AP down first, which
drops the phone that sent the request. The connect endpoint answered only
after the attempt finished, so the browser never got a reply and the WiFi
tab's Connect button appeared to do nothing.

- /wifi/connect answers 202 immediately while the AP is active and connects
  in a background thread; the result (never the password) is reported via
  /wifi/status as last_connect_attempt. A second connect while one is
  pending gets 409.
- connect_to_network holds a /tmp flag for the attempt; the monitor daemon
  skips AP management while it is fresh. Previously the daemon's
  disconnected counter, accumulated over the whole AP session, re-enabled
  the AP on its next tick in the middle of the connect.
- The WiFi tab and captive setup page explain the handoff up front, and on
  reopening show why the last attempt failed. The wrong-password message
  now works: the route sets the error_type the captive page checks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(wifi): serialize connect attempts on both paths

Addresses CodeRabbit review on #571:

- Check for a pending attempt before branching on AP state. A background
  attempt takes the AP down long before it finishes, so a second click
  used to bypass the 409 and start a competing synchronous connect.
- Record pending for the synchronous (non-AP) path too, so two requests
  can't overlap and have the first clear the daemon's in-progress flag
  while the second is still connecting.
- Clear the pending state if the background thread fails to start, rather
  than refusing every later request until restart.
- Say the setup network returns "within a few minutes": a stale flag plus
  the daemon's grace period can take longer than one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 12:40:41 -04:00
ChuckandClaude Opus 5 d1e821c625 fix(web): harden, polish and optimize the web UI per the Sept 2026 audit (#568)
* fix(web): harden, polish and optimize the web UI per the September 2026 audit

Works through docs/archive/WEB_UI_AUDIT_2026-09.md (health 8/20).

Implementation integrity (P0)
- app.css now defines every utility class the templates and JS use,
  including .hidden, so the ~145 JS show/hide toggles work. Button reset,
  and base component rules (.btn, .form-control) wrapped in :where() so
  utility classes on the same element win. New static-audit test fails
  when a used utility class has no rule.

Accessibility
- Focus rings render (the old ring rule referenced undefined variables);
  one :focus-visible outline everywhere; skip link; labelled nav landmarks.
- Shared dialog helper (js/utils/dialog.js): role/aria-modal, focus trap,
  Escape, focus return, applied to every modal.
- Named icon-only buttons and labelled ~70 form fields.
- Toasts announced once; errors persist >= 10s; one showNotification.
- Captive WiFi page: live region, timeouts, dark mode, 16px inputs.

Performance (Pi Zero 2 W)
- SSE streams and tab timers pause when hidden or off-tab; the display
  stream only runs while a preview is visible. app-shell.js deferred.
- Widget scripts served as one versioned bundle (/assets/widgets.js):
  52 -> 21 script tags, 66 -> 35 requests on first load.
- Stdlib gzip fallback when flask-compress is missing: first-load JS/CSS
  1358 KB -> 291 KB on the wire. SSE untouched.

Theming and responsive
- File managers, form fields and Fonts upload on theme tokens; bare
  inputs themed in dark mode; no more white surfaces.
- No horizontal overflow at 375px on any tab; 44px touch targets on
  coarse pointers; reduced-motion respected; header title truncates.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings on #568

- json-file-manager: focus-trap releases kept in a Map (no dynamic
  property access or delete; no value-returning forEach callback)
- notification / schedule-picker: style and day-label lookups via Map
- app.js: move the pending-queue assignment out of the expression
- diff_viewer / error_handler: named function declarations instead of
  arrow consts

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: check the OAuth widget ships in the widget bundle

base.html no longer tags widget scripts one by one; they load through
/assets/widgets.js. Assert the page requests the bundle and the bundle
contains google-oauth.js, which is what the test was protecting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): address review feedback on #568

- widget bundle version fingerprints every file (name, mtime_ns, size)
- gzip fallback appends Accept-Encoding to an existing Vary header
- dialog helper: releasing a non-top dialog no longer moves focus out of
  the dialog the user is in
- labels: file-upload targets its file input; fallback config fields get
  label for/id pairs; native color input has a fallback name
- utility audit also reads class names inside bound :class expressions

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): give the native color-picker input an accessible name

CodeRabbit flagged this on PR #568 as an outside-diff finding (never
posted inline, so it was missed in the round of fixes that addressed
the other 6 review comments). The <input type="color"> only carried a
title attribute; screen readers don't reliably announce title, and
there's no other label naming the control when showHexInput is false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear Codacy findings in app-shell.js

- drop the unused catch binding on the SSE JSON parse
- move the pending-notification queue assignment out of the expression

No behavior change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): contain plugin widgets/ dir and bound style-editor retries

From CodeRabbit review on #568 (code that arrived with the main merge):
- serve_plugin_widget resolves widgets/ with resolve_under before
  resolving the manifest script under it, so a symlinked widgets
  directory can't become the containment base (CWE-22). New test.
- style-editor init stops polling after ~10s when the widget never
  registers and leaves the plain fallback fields in place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 09:42:24 -04:00
ChuckandClaude Opus 5 69d408b321 feat(core): one per-element display-customization framework, wired into the web UI (#566)
* fix(sports): rebuild un-shared faces through the pinned layout engine

unshare_element_fonts re-instantiates a duplicate font face so two
elements can be told apart by id(). It did so through bare
ImageFont.truetype, which takes PIL's default layout engine rather than
the one src/common/font_layout.py pins. Raqm and Basic disagree on
fractional advances -- that disagreement is the reason the pin exists,
having broken golden images across machines -- so a rebuilt face could
measure differently from the shared face it replaced, on any host where
Raqm is installed.

These were the only two call sites in src/ bypassing the pin.

The guard asserts that the rebuild goes through the pinned loader rather
than comparing engine values: where Raqm is absent, bare truetype returns
BASIC anyway, so an engine comparison passes whether or not the pin is
honoured. The first draft of this test did exactly that and passed with
the bug reintroduced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): drop the two dead client-side config-form renderers

generateConfigForm and generateSimpleConfigForm (580 lines) were defined
on the Alpine component and never called: server-side Jinja replaced them,
as pages_v3.py:641 records. Nothing in any template invokes them -- there
is no x-html in the templates and no bracket access on the component.

They carried their own x-widget dispatch, which made them an active trap:
the next person adding a widget would reasonably think both renderers
needed updating.

plugins/config_manager.js (PluginConfigManager, 133 lines) goes for the
same reason -- loaded on every page from base.html, referenced only by
itself and by an archived doc.

Kept, having checked them: widgets/example-color-picker.js is the worked
example docs/widget-guide.md points plugin authors at, and
widgets/plugin-loader.js is the client half of a documented feature
(manifest-declared plugin widgets) whose server route is missing --
soccer-scoreboard already ships a widgets/custom-leagues.js that this
loader is meant to fetch. That is an unfinished feature to complete, not
dead code to delete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): serve plugin-declared widgets, and actually ask for them

LEDMatrixWidgets.loadPluginWidget has always fetched
/static/plugin-widgets/<plugin>/<widget>.js, and docs/widget-guide.md has
always documented that path, but nothing served it. soccer-scoreboard has
shipped a 17KB widgets/custom-leagues.js since August that could never
load. Both halves were missing, not just the route:

- serve_plugin_widget serves the script from the plugin's widgets/
  directory as text/javascript. The manifest is the allowlist -- only a
  widget the plugin declares is reachable -- so installing a plugin does
  not publish everything it ships. Path handling mirrors the sibling
  serve_plugin_web_ui: allowlist regexes, os.path.basename, resolve() +
  relative_to() containment, and the ledmatrix- prefix fallback. The
  declared script name is guarded too, since it comes from the plugin
  rather than the request.

- The config form never requested one. Its x-widget dispatch is a
  hardcoded list of core widget names, so a plugin's own widget fell
  through to a plain text input. An unrecognised x-widget on a string
  field now asks ensureWidget() for it. The text input stays as the
  fallback and is removed only once the widget has actually rendered, so
  a missing or broken widget costs the user an editor rather than their
  configured value on the next save.

- manifest_schema.json gains "widgets", so the declaration is validated
  rather than merely tolerated by additionalProperties.

Verified in a browser against the real partial: a declared widget loads,
registers and renders, and its field posts exactly one value; a field
whose widget 404s keeps its text input and still posts its value.

Not addressed: loadPluginWidgetsFromManifest still has no caller. The
per-field ensureWidget path is lazier and is what the form now uses, so
that bulk helper is dead weight -- worth removing, but left alone here
rather than inventing a call site for it.

Known limitation, documented: only string-typed fields take this path.
object/array/boolean/number fields and enums are dispatched by the
template's own branches, which still only know core widgets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(element-style): a wrong-size BDF now keeps its font, not its size

BDF fonts are fixed-size bitmap strikes: FreeType accepts only the pixel
size baked into the file and raises for anything else. 32 of the 35
shipped fonts are BDF, so a size picked in the web UI usually is not a
valid strike -- and load_font caught that failure with its generic
"unloadable font" handler, which substitutes PressStart2P. Asking for
5x7.bdf at size 10 therefore rendered a completely different typeface,
silently.

It now falls back to the file's own native size instead, which is what
SportsCore._load_custom_font_from_element_config has always done. The
native size is read via FontManager._read_bdf_native_size rather than a
fourth copy of that parser, matching how core.py already delegates.

Also here, because they are the same code path:

- native_bdf_size() is exposed for the web UI, which needs to know when a
  size field can take effect at all. None means "free choice".
- ElementStyle.font_size now reports the size actually realised rather
  than the one requested. Callers lay out from it, and reserving space
  for a size nothing was drawn at is how this surfaces.
- The module font cache is a bounded LRU (256) instead of an unbounded
  dict. The display process runs for weeks and every config save can add
  a (font, size) pair; every other hot cache in the codebase is bounded
  this way.

Untouched configs are unaffected: the shipped classic fonts are the three
TTFs, so nothing was hitting the substitution path by default.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): per-mode style and offset overrides

Lets one element be styled differently per situation -- a scoreboard's
live / upcoming / recent cards, weather's current / hourly / daily
screens -- under customization.modes.<mode>.

The mode is bound at construction rather than passed per call. That is
what makes this cheap to adopt: SportsUpcoming and SportsRecent are
already separate instances with distinct SKIN_MODE values, so binding
once makes every existing style()/offset_value() call site mode-aware
without editing any of them. A per-call mode argument exists for the rare
host that renders more than one mode.

The two layers answer different questions, deliberately:

- The base layer keeps the existing "differs from the schema default"
  rule, because the save flow writes the full default object into
  config.json whether or not the user touched it.
- A mode layer is pure override -- its fields default to None, so
  presence is intent. Nothing writes into it unasked, so there is nothing
  for the stricter rule to protect against.

None therefore means inherit, and has to stay distinct from 0: a mode
y_offset of 0 means "sit at the base position", not "no preference".
This is the same distinction scroll_card.switch_* draws with "inherit".

A malformed mode value falls back to the resolved base value rather than
to the caller's default -- caught by the degradation tests, which is what
they are for: resolving the mode first let one bad string in a mode block
silently discard a good base offset.

With no modes block, and for every existing caller, resolution is
unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): declare per-mode overrides in config_schema.json

A plugin adds "x-style-modes": ["live", "upcoming", "recent"] alongside
its x-style-elements declaration and gets a customization.modes.<mode>
group per mode, with every field of every declared element repeated as an
override.

Those override fields are typed nullable and default to null, which is
the whole trick. The save flow writes schema defaults into config.json
wholesale, so giving a mode field the base element's default would make
every mode a frozen copy of the base the first time a user pressed Save,
and the base would stop reaching them. Null means inherit. The mutation
test for this is explicit: with concrete defaults, a base font_size of 14
resolves as 10 with user_forced set.

min/max from the declaration carry into the mode blocks, so an
out-of-range override is rejected by validation rather than clamped
silently at render time.

Also: the emitted font field now carries "x-widget": "font-selector". The
widget already shipped and the config form already allowlisted it -- the
hint was simply never emitted, so the field rendered as a bare text box
that the user had to type a font filename into.

Verified through the real SchemaManager path -- load_schema, defaults
extraction, merge_with_defaults, validation, then resolution -- rather
than against a hand-built dict, since the thing at risk is what that
pipeline does to a null.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): render the config form from the schema the save route validates

The form read config_schema.json with a raw json.load while
api_v3.save_plugin_config went through SchemaManager. Those are not the
same schema: SchemaManager applies expand_style_elements, which turns a
compact customization.x-style-elements declaration into the per-element
blocks the form knows how to render.

Without it, that customization object has an x-style-elements key and no
"properties", so the template's object branch matched nothing and the
section rendered as empty space -- while saving still validated against
the expanded shape. of-the-day ships the compact form, so its
customization section has been invisible in the web UI.

pages_v3 gains a schema_manager the way it already has config_manager and
plugin_manager. use_cache=False matches the save route, so an edited
schema is not served stale during plugin development. The raw read stays
as a fallback for callers that register this blueprint without one.

Checked before making the change: load_schema does nothing here except
read, validate and expand -- inject_skin_selector is a separate method it
does not call -- so this is not a behaviour change for schemas without
the declaration.

The test pair renders the same compact schema with and without a
SchemaManager, so it documents exactly what was broken as well as what is
fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(web): style-editor widget -- a row per element instead of 65 accordions

Rendered element by element, a realistic scoreboard's customization block
is 65 nested sections, and reaching one per-mode font size takes five
levels of expanding. The widget collapses that to one compact row per
element -- font, size, colour, X, Y -- with a tab per declared mode.

It emits ordinary inputs under the same dotted names the generic renderer
would produce, so the save/validate/merge pipeline is untouched: no hidden
JSON blob and no new server-side parsing. It is driven entirely by the
schema block it is handed, so fields added to the schema later appear
without editing the widget. If it fails to load or throws, the generic
nested rendering it replaces is left in place.

Fixing two things the save path got wrong for nullable fields, found by
posting what the widget actually emits:

- The indexed-array recombiner (text_color.0/.1/.2 -> one list) compared
  the declared type to the string 'array', so a per-mode colour, typed
  ["array", "null"], was never reassembled and failed validation on save.
  _parse_form_value_with_schema had the same comparison.
- A blank nullable field became [] rather than None, which then failed the
  minItems the colour array declares. Null is the inherit sentinel, so it
  has to survive.

And two things the widget itself got wrong, found by looking at it:

- An unset base control fell back to the select's first option, so an
  untouched scoreboard claimed every element used 10x20.bdf -- and the
  size box then locked itself to that bitmap font's fixed size. Base
  controls now show the schema default; mode controls stay blank, because
  blank there means inherit.
- Elements arrived alphabetised (Detail and Odds above Score). Flask's
  JSON provider sorts keys, so declaration order has to be stated
  explicitly; expand_style_elements now emits x-propertyOrder, which the
  generic renderer already honoured too.

Size is disabled and shown as fixed for a bitmap font, using the
scalable/native_size the font catalog now reports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): visibility, alignment and scale per element

Completes the customization vocabulary: hide an element, align it, and
resize a logo, alongside the font/size/colour/offset that already existed.
All three per mode.

They resolve to "change nothing" until the user asks for something -- True,
None and 1.0 -- rather than to whatever the schema declares. That is the
same invariant the font fields keep: a caller that honours them still
renders an untouched config exactly as it did before they existed. A
schema default therefore does not count as a choice, which matters because
the save flow writes that default into config either way.

scale sits in the layout block with the offsets rather than in the element
block, because it is geometry: a logo has a scale and no font. The widget's
columns come from the schema, so a logo row shows visibility, offsets and
scale and no empty font cell.

Two bugs found by the tests rather than by reading:

- A nullable enum needs null in its enum list, not just in its type. The
  mode copy of `align` defaulted to null and then failed its own schema, so
  a plugin declaring any enum field with modes could not save at all. Six
  tests failed on this before any of them reached what they were testing.
- defaults_from_schema only ever extracted font/font_size/text_color, so
  the schema defaults for the new fields were invisible to the resolver and
  a declared default read as a user choice.

Widget: the table scrolls horizontally and pins the element-name column.
Nine columns do not fit the config panel, and clipping them hid the offsets
entirely while scrolling them made every row anonymous.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): resolve elements under the names plugins actually use

Two naming conventions collided as the scoreboards grew. Counted across
the published schemas: the style block names elements with a _text suffix
(score_text, status_text, detail_text), while the layout block mostly uses
the bare noun (score, date, time, odds) -- except status_text, which kept
the suffix in seven plugins and lost it in two. records vs record splits
seven to two the same way.

A lookup now tries the exact name first and then the spellings that mean
the same thing. Exact-first is what makes this inert for any config that
already matches; the aliases only decide cases that resolved to nothing
before.

This is also what makes migrating to the compact declaration form safe.
That form uses one key for both blocks, so a scoreboard adopting it asks
for layout.score_text while its users have layout.score saved -- without
the aliases, every offset they had dialled in would silently become 0.

Applies to the style block, the layout block, the schema defaults and the
per-mode overrides, since the drift shows up in all four.

Not attempting to canonicalise on write: renaming keys in config.json
would break the plugins still reading the old spelling from their own
bundled code, and the drift costs a dict miss rather than correctness.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(plugins): BasePlugin.styles -- per-element styling every plugin inherits

Adopting the element-style system meant repeating three things in every
plugin: a guarded import, finding its own config_schema.json, and
rebuilding the resolver when on_config_change swapped the config dict.
This is those three things once, on the class all 45 plugins already
inherit from.

    title = self.styles.style('title_text',
                              classic_font='PressStart2P-Regular.ttf',
                              classic_size=8, classic_color=(255, 255, 255))

The classic_* arguments are the adoption contract: with nothing configured
they come back verbatim, so a plugin that switches to this renders exactly
as before until a user changes something.

A plugin with one instance per display mode sets STYLE_MODE on the class
and every existing lookup becomes mode-aware without a call site changing
-- which is the point of binding the mode to the resolver rather than
passing it per call. styles_for() covers a plugin that renders several
modes from one instance.

Schema discovery reads the concrete class's own module rather than this
file, because this file lives in src/plugin_system where no plugin schema
exists -- the same trap SportsCore._config_schema_path documents. The
first mutation test for that passed anyway: an installed plugin's module
directory and its entry under plugins_dir are the same path, so the test
could not tell the two apart. The case where they diverge is a plugin
symlinked in for development, and the test now forces that shape.

Getting discovery wrong is silent rather than loud: with no schema the
resolver has no defaults to compare against, so every configured value
reads as a deliberate override and the plugin quietly stops honouring its
own shipped styling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(element-style): adopt hand-written customization blocks, and widen the font list

Nineteen plugins spell their style elements out longhand instead of
declaring them -- football's block is 701 lines for seven elements -- and
predate this system entirely. Core now recognises that shape, so they pick
up the row-per-element editor and the real font picker on a core update
rather than on a plugin release. Checked against every published schema:
21 plugins adopt, and the defaults of each still validate against the
schema generated for it.

Detection requires *every* field in a block to be one this system
understands. A looser "has at least one style field" rule sweeps in
baseball's `count`, which carries a text_color beside geometry that means
nothing here. That distinction took three attempts to test: the first two
assertions passed under both rules, because an over-eager rule leaves a
fontless block looking untouched and only surfaces as an extra row in the
editor.

The hardcoded font enum is replaced rather than extended. Football lists
five of the thirty-five installed fonts, which is why a font a user
uploads can never appear in one. It is not a curated safe set -- it omits
some twenty other faces that fit the declared size cap just as well -- it
is the fonts that happened to exist when it was written.

Widening it does need a guard, though, and not the one the schema already
has: a bitmap font ignores font_size and renders at its size baked into
the file, so `maximum: 16` cannot stop a 27px face. The picker now filters
out fixed-size fonts taller than the element's own declared ceiling, which
drops exactly the four that would overflow a 32px panel and keeps the
other thirty.

Per-mode overrides stay opt-in: core cannot invent a plugin's display
modes, so `x-style-modes` remains the one line that unlocks them. Their
layout half covers every positionable element rather than only those with
a style block -- the two namespaces do not line up in a hand-written
schema, and football positions six things (logos, timeouts, possession)
that have no style block at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(web): remove the two Fonts-tab panels that reported invented data

"Element Font Overrides" let a user configure an override, showed a
success toast, and changed nothing. All three endpoints behind it were
stubs -- GET returned a hardcoded {}, POST and DELETE returned success
without calling anything -- each marked "This would integrate with the
actual font system".

Wiring them to FontManager would not have fixed it. The machinery there is
real (_load_overrides/_save_overrides persist config/font_overrides.json,
resolve_font applies them, and the countdown plugin genuinely consumes
it), but the panel's element dropdown offered eleven invented keys --
nfl.live.score, clock.time, weather.current -- that no plugin has ever
read. An override saved against one of those would have persisted
correctly and still done nothing.

"Detected Manager Fonts" goes for the same reason. It claimed to show
"fonts currently in use by managers (auto-detected)"; its own comment said
"we'll simulate this", and it listed every font in the catalog with a
hardcoded usage_count of 1 -- the panel beside it, with fabricated
numbers attached.

Per-element font choice now lives in each plugin's own config editor,
against the elements that plugin actually has, and covers size, colour,
offsets, visibility, alignment and scale rather than family and size.

Kept: the font library (upload, preview, delete), which works, and
/fonts/tokens, which is a stub but genuinely feeds the preview's size
dropdown. FontManager's override methods are untouched -- countdown uses
them.

Verified in a browser with the tab's JS running: no console errors, 35
fonts listed, upload and preview intact. Removing the panel meant unwiring
it from populateFontSelects too, which would otherwise have bailed out
early on the missing select and left the preview dropdown empty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(sports): one reader for element colours and layout offsets

There were two copies of the per-element colour read and three of the
layout-offset read. They had already drifted -- the scroll-card renderer
carries a comment about having ignored offsets its own schema advertised
-- and each new capability had to be added to all of them or silently work
in some places and not others.

All of them now go through src.element_style, which is what carries the
alias handling and the per-mode lookup. That lands immediately for the
nine plugins importing these modules: a scoreboard asking for `score_text`
offsets finds the `layout.score` its users configured, and a Live instance
resolves its own colours through SKIN_MODE without any call site passing a
mode.

_normalize_color learned "#RRGGBB" in the process. The scoreboards' own
readers have always accepted it, so the shared one had to, or consolidating
would have quietly dropped a form users' configs may hold. _coerce_offset
picked up the non-finite guard the scroll-card reader had and the other two
did not.

_get_layout_offset is promoted onto SportsCoreSharedMixin. Each plugin
still carries its own copy in its bundled sports.py, which wins by MRO --
so adopting this is a deletion in the plugin, and until that deletion
nothing changes for it.

Note for whoever runs the suite next: test_display_dirty_tracking.py is
order-dependent. Fifteen of its tests failed in one full run and passed in
the next with no change in between, and pass in isolation. Pre-existing,
unrelated to this, but it makes a full-run diff untrustworthy until it is
fixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(changelog): record the element-style work under Unreleased

This file's own preamble asks for it: a plugin may delete its bundled
fallback copy of a core module only when its manifest floors on the first
release that shipped that module, which requires the additions to be
recorded here against a version.

Names a plugin can now import and floor on -- the stateless layout_offset
and element_color readers, alias_keys, native_bdf_size, the resolver's mode
binding, BasePlugin.styles, and the promoted
SportsCoreSharedMixin._get_layout_offset -- plus the schema and web-UI
changes, the four fixes and the three removals.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(fonts): log the BDF native-size read failure instead of swallowing it

The bdf-native-size lookup in get_fonts_catalog() caught any exception
and silently discarded it. Every other guarded read added in this PR
(the manifest parse in _declared_widget_script, the SchemaManager
fallback in _load_plugin_config_partial) logs before falling through
to the same degraded behavior. This one didn't, which is the shape a
silent-exception-swallow lint rule flags. Behavior is unchanged --
native_size still comes back None -- but a corrupt or unreadable BDF
file now leaves a trace.

Verified: font-related tests (140) and the full suite still pass,
with only the 2 pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias
failures already present on origin/main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: address CodeRabbit findings on the style-editor/font-selector PR

- Fix _load_font_sized double-wrapping the (font, size) tuple on the
  missing-font path, which handed callers a tuple instead of a font.
- Fix _set_nested_value skipping an explicit None when the key already
  existed, which silently kept stale overrides when a user cleared a
  nullable per-mode field or blanked all channels of an indexed color.
- Preserve BDF scalable/native_size metadata through fetchFontCatalog's
  catalog-format mapping so maxFixedSize filtering actually applies.
- Stop caching an empty array on a failed font-catalog fetch so a later
  call can retry instead of being stuck with the failed result.
- Keep a saved font selected in the style editor even when it no longer
  fits a newly declared maxFixedSize, instead of silently deselecting it.
- Don't drop in-progress user edits to fallback fields when a plugin
  widget finishes loading asynchronously and takes over the form.
- Tighten the removed font-override endpoint test to assert 405, not
  just != 200.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): a partial save no longer switches off checkboxes it never showed

An HTML checkbox posts nothing when unchecked, so the save route walked the
schema and forced every boolean missing from the form to False. That is right
for the rendered form and wrong for every other caller: a script, the MQTT
bridge or a curl against the documented endpoint never rendered a checkbox, and
reading its silence as "all off" turns a one-field save into a mass disable.

Found on hardware. Posting four customization.* keys to a live device switched
off nfl.enabled, ncaa_fb.enabled and every display-mode toggle in one request.

The form now reports the top-level sections it drew (__rendered_section), and
inside those an absent checkbox still means unchecked -- including a section
whose only fields are checkboxes that are all off, which no heuristic could
recover. A post with no marker only touches objects it actually posted a field
from. Meta fields are dropped before form keys are treated as config paths,
because unknown keys are otherwise written straight into config.json.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(sports): resolve element colour by name, and honour visible/align/scale

Two of the three gaps this framework shipped with.

Colour by name. A draw resolved its colour by comparing the *identity* of the
font object it was handed, which cannot tell two elements apart when they share
a face -- so those draws went out white. Every bitmap font is in that case,
because a freetype.Face cannot be re-instantiated to un-share it, which is how
an element rendered in any of the 32 shipped BDF fonts silently lost a colour
its picker had offered all along. _draw_text_with_outline now takes
element="score_text" and reads the colour by name; the identity path remains
for un-annotated callers, but narrows before giving up -- one configured colour
among the sharers is the only thing the user can have meant.

Visible, align and scale. The resolver has understood these since the
framework landed and nothing consumed them: an element could be marked hidden
in the web UI and still render. Adds the stateless readers, the mixin
accessors, and a scale parameter on the one shared logo-sizing seam (keyed into
the cache, so two elements scaled differently cannot be served each other's
image). Naming an element in a draw also honours its visibility.

Untouched configs are unaffected: every new parameter defaults to today's
behaviour, and all ten affected plugins render pixel-identically to main across
every harness size.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(plugins): how to declare styleable elements; harden the widget's lookups

The plugin-author guide for the compact x-style-elements declaration -- what
each key does, how to read values back without breaking the "user-forced only
when it differs from the default" rule, and why a hand-written block needs no
changes to be adopted.

Also clears the static-analysis findings on style-editor.js. Every lookup in
that file is keyed by something out of a schema or a saved config, so a key of
__proto__ or constructor would walk the prototype chain and hand back a
function instead of a schema; reads now go through an own-property helper. The
panel registry became a list, and the flagged vars moved to their function
roots.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): clear the remaining static-analysis findings

Five, all on lines this branch touched.

The Python one is not a new defect: _set_missing_booleans_to_false's first
parameter was always named `config`, which shadows the `config` submodule
imported for its side effects at the bottom of this module. Editing the
signature simply put the existing warning on a changed line. The parameter is
the plugin's config dict, so `plugin_config` is what it should have been called
anyway; callers pass it positionally and are unaffected.

The JavaScript ones are the object-injection rule firing on reads keyed by
data. own() now goes through a property descriptor, so the one unavoidable
data-keyed read is no longer a computed member access; at() consumes its path
instead of indexing it; and the column set is a Map, which has no prototype to
pollute and needs no guarded reads at all.

Verified the widget still renders identically against football's real schema:
29 element rows, all four mode tabs, values populated, no console errors.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): drop the hasOwnProperty alias the descriptor read made redundant

own() now reads through Object.getOwnPropertyDescriptor, so the alias it used to call has no remaining reference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 11:53:50 -04:00
ChuckandClaude Opus 5 59997594ac test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths

Two shared-state problems, both of which show up as a permanently dirty
checkout or an unreproducible test failure.

test_display_dirty_tracking.py builds a real DisplayManager, whose
_snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web
UI reads. Every pytest process on the machine shares that one file, so two
concurrent runs -- CI shards, a second worktree, an agent running the suite
alongside -- overwrite each other's snapshot and the mtime assertions stop
meaning anything. The module fixture now points it at a session-unique temp
path; the individual tests that care still override it further.

To be clear about what this does and does not fix: this is a real shared-path
hazard, but it is NOT the cause of the intermittent 15-test failure in that
module. That turned out to be the emulator's fixed TCP port, fixed in the
follow-up commit. This change stands on its own merits.

web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json
and data/operation_history.json as the web interface runs, into a directory
that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig
that ever opened the web UI -- and every test run that constructs the app --
left three untracked files behind and a permanently dirty `git status`. Only
data/.gitkeep is tracked under data/, so the negation keeps it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: stop the emulator binding a fixed port, so concurrent runs can't collide

This is the cause of the intermittent full-suite failures we have been chasing:
runs of identical code landing anywhere between 100 and 130 failures, while
every implicated test passed in isolation.

Six test modules set EMULATOR=true and build a real DisplayManager. The repo's
emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to
serve the dev preview. That port is a machine-wide singleton, so a second pytest
process -- a CI shard, another worktree, an agent running the suite alongside --
loses the bind. RGBMatrix construction then raises, DisplayManager catches it and
falls back to `self.matrix = None`, and every test that subsequently touches the
matrix dies with

    AttributeError: 'NoneType' object has no attribute 'SwapOnVSync'

which names neither a port nor a socket, and points at the wrong file entirely.
Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its
matrix-touching tests fail together or not at all -- the 15-test swing that made
the totals look random.

Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process
and running test_display_dirty_tracking.py:

    without this change    15 failed, 6 passed
    with this change       21 passed

The "raw" adapter renders in memory and binds nothing. Only display_adapter is
overridden, in a throwaway config written per pytest process; the repo's
emulator_config.json is untouched and `run.py -e` still opens the browser
preview on 8888. Nothing in the suite referenced the adapter, and the tests
wrap SwapOnVSync on the matrix object itself, so they are indifferent to what
sits underneath. allow_adapter_fallback is forced off -- falling back would
land us on the browser adapter and its fixed port, which is the whole problem.

CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set
to an absolute path: the previous behaviour depended on where pytest was invoked
from, and silently wrote a default config into whatever directory that was.

Verified no regressions: full suite on this branch and with origin/main's
versions of the touched files, same machine, back to back -- 115 failed /
4347 passed on both sides, zero failures unique to either. That 115 is the
pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell
scripts, Linux-only binaries); CI on Linux remains authoritative.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore: mark the shell entry points executable

Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh`
fails with "Permission denied" and only works if you know to prefix `bash`.
That one matters most: the web UI's own error hint, added in #560, tells users
to run exactly that path when a system action fails for want of passwordless
sudo, and following that instruction verbatim did not work.

All eleven carry a shebang and are invoked directly, never sourced. The two
sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately
left non-executable, which is what distinguishes a library from an entry point.

Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied
with `git update-index --chmod=+x` because this checkout is on Windows, where
core.fileMode is off and the working-tree bit is not tracked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:45:13 -04:00
ChuckandClaude Opus 5 f6367d63ae security: triage the CodeQL backlog — 129 alerts, three of them live (#561)
* fix(web): escape quotes in every HTML escaper, not just & < >

The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:

    value="${escapeHtml(v)}"   title="${escapeHtml(v)}"

so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.

It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.

app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.

cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `&#39;`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.

url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).

test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): stop request-supplied names from reaching paths outside their base

Three of the py/path-injection alerts were live, not lint:

* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
  whose resolved path *string-prefixed* the plugin directory. Flask's
  default converter forbids a slash but not dots, and
  get_plugin_directory('..') returned the parent of the plugins directory
  because it exists -- so every file under the project root then prefixed
  that directory, config/config_secrets.json included. The prefix check
  was also wrong on its own terms: with plugin dir "plugin-repos/foo",
  "../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
  start with "plugin-repos/foo".

* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
  body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
  file_id of "../../../../etc/something" deleted that file. This is the
  one finding in the batch that destroyed data rather than exposing it.

* POST /api/v3/cache/delete passed the body's key through
  CacheManager.clear_cache to DiskCache, which joined it as a filename and
  called os.remove. Same shape, same result. The guard goes in
  DiskCache.get_cache_path, the single choke point get/set/clear share, so
  every caller is covered rather than just this route. Real keys are the
  stems of files already flat in the cache directory -- that is how
  list_cache_files derives them -- so nothing legitimate is turned away.

The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.

Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.

test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): refuse a plugin id that is not a plain name, don't truncate it

pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.

Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(web): say what the plugin web_ui iframe actually is

The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.

Also drops the now-unused os/os.path imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): inline url-input's scheme guard at the previewLink.href sink

CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.

Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.

Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): address CodeRabbit findings on the CodeQL triage PR

- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
  credential-saving or connect flow runs. NetworkManager only accepts
  printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
  was previously saved/attempted before nmcli itself rejected it.

- web_interface/blueprints/api_v3/config.py: fail closed when the
  plugin config schema path can't be resolved under the plugins
  directory (e.g. a symlinked plugin dir). Previously this fell
  through with secret_fields left empty, so submitted credentials for
  that plugin were saved as ordinary, unencrypted configuration.

- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
  splicing the JSON day/column key into an inline oninput="..." handler
  string. escHtml() escapes quotes for a normal HTML attribute, but the
  browser HTML-decodes the attribute before running it as script, which
  undoes that escaping and lets a crafted column name (e.g. from an
  uploaded JSON file) break out of the JS string and execute. Cell
  edits now travel through data-day/data-col attributes read by one
  delegated 'input' listener instead.

  While in this file: fixed 6 pre-existing missing-')' typos on
  multi-line safeSetHTML(...) calls (already flagged by Biome in this
  PR's own CodeRabbit run as syntax errors blocking its lint pass).
  These predate this PR (present on main too) but made the whole file
  fail to parse in any JS engine, which is a bigger problem than the
  XSS finding itself and directly touches the same lines.

Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 15:32:03 -04:00
ChuckandClaude Opus 5 5137e86d16 feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab

PR #544's change, ported onto the api_v3 package split (#553). Identical
behaviour; only the placement of the new code differs.

The original added 508 lines to web_interface/blueprints/api_v3.py, which #553
deletes, so every hunk of it would conflict irreconcilably. Ported by AST:
26 new top-level items sorted to where the split puts each kind --

  __init__.py   2 imports, 11 constants, 7 helpers
  starlark.py   4 routes  (/starlark/editor/{apps,status,start,stop})
  misc.py       2 routes  (/integrations/mqtt-bridge{,/config})

Everything outside api_v3.py -- the Tools partial, the installer scripts, the
JS tests -- applied unchanged.

Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is
regenerated to match, which is exactly what test_api_v3_url_map.py is designed
to make you do when routes are added.

Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR
ships could not be run here -- node is not installed on this machine -- so
test/js/dom/test_tools_sections.js is unverified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST

The AST-based port of #544 onto the api_v3 package split dropped `time`
from starlark.py's import list. start_pixlet_editor() and
stop_pixlet_editor() both call time.time()/time.sleep() directly, so
every start (NameError building `state['started_at']`) and every stop
that has to wait out the EXIT trap crashed with a 500. No test caught
it because the route's own tests mock subprocess.Popen but never
actually invoked it before now.

Also carries over #544's later fix that this port branched before:
env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an
operator who had already pinned PIXLET_EDITOR_HOST to loopback,
forcing the unauthenticated `pixlet serve` process onto the LAN
regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as
api_v3.starlark.py's siblings already do for _pkg-owned names.

Both fixes route the shared _pkg.time reference the rest of the
package's route modules already use for anything a test might need to
patch, rather than a bare `import time` local to this file.

Ported the existing regression test from #544
(TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's
module layout (web_interface.blueprints.api_v3.starlark instead of the
old monolithic api_v3 module), which is what caught the NameError.

Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this
branch and on origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to
this change).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): clear the six lint errors this rebase introduced

All six were introduced by rebasing this branch onto the merged blueprint
split, not by the split itself. Confirmed by diffing pyflakes output against
main with line numbers normalised -- everything else it reports is present on
main too and is the package's deliberate re-export pattern.

starlark.py used _STARLARK_APPS_DIR three times without importing it (F821).
The rebase resolved an import-list conflict as a union of both sides, and that
symbol was on neither side of the conflict hunk, so it was silently lost. It is
defined in __init__.py and is now imported like its neighbours. This was the
only one of the six that would fail at runtime rather than merely lint.

__init__.py imported contextlib twice (F811): the cherry-pick added one next to
the existing import. Removed the duplicate; the original at line 19 is used.

__init__.py imported signal purely to re-export it to starlark.py, so pyflakes
saw it as unused (F401). signal is stdlib and does not need routing through the
blueprint package, so starlark.py imports it directly and __init__.py no longer
does. contextlib stays re-exported because this module genuinely uses it.

_read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule
this module imports at the bottom for its route side effects (F811). Renamed to
`settings`, with a comment saying why, since the name is otherwise the obvious
one to reach for.

Verified: pyflakes now reports nothing on this branch that main does not, the
package imports, all nine route modules load, and 117 routes register, matching
the pinned URL-map snapshot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply

Two CodeRabbit findings on the bridge settings endpoint, both of which returned
200 while doing something other than what the caller asked.

`request.get_json(silent=True) or {}` turned a missing or unparseable body --
and the JSON literals null, [] and false -- into an empty dict, which then
satisfied the isinstance(data, dict) guard on the very next line. The guard was
there to reject exactly those bodies. Dropping the `or {}` lets None fail it.

The same `or {}` on /errors/clear is left alone: its docstring documents the
body as optional, so an absent body legitimately means "use the defaults". The
difference is that saving settings has nothing sensible to do with no body.

`if data.get('clear_password'):` accepted any truthy value, and the string
"false" is truthy in Python -- so a client echoing the field back as a string
wiped a password it meant to keep. Now coerced through the package's existing
_coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else.

test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus
a missing body, and clear_password across truthy and falsy spellings. Verified
against the unfixed code -- reverting the body guard fails 5, reverting the
coercion fails 3.

Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials
when TLS is off (CWE-319). That is a policy decision about the feature rather
than a defect -- unencrypted MQTT on a trusted LAN is common and often
deliberate -- so it is raised on the PR for a maintainer call instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: work through the remaining review findings on the editor and bridge

allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses
the network in cleartext. Refused now rather than merely warned about -- but
refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal
deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults
false, and is coerced like the other booleans so the string "false" cannot
switch the guard off.

starlark.py:796 -- the supported service runs Flask threaded, so two start
requests could each see running=False, each launch an editor, and the second
state write replace the first PID, orphaning a process that holds the display
down with nothing recording it. The check-launch-write sequence now takes a
module-level lock.

starlark.py:848 -- if the state write failed the route returned success with an
editor running and no PID recorded: status and stop both reported no session
while the display stayed down until the timeout expired. It now terminates the
process group and returns an error.

starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so
nothing hands the display back, yet the response said "the display is
restarting". After an escalation the display is now restarted explicitly, and a
failure to do so returns an error naming the manual step instead of a success.

pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no
timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the
failure lands before the display is stopped rather than after.

pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface
address are three cases, not two. Collapsing the last two printed a URL saying
"localhost" whenever PIXLET_EDITOR_HOST named a LAN address.

tools.html:1254 -- escHtml does not encode single quotes, and the app id was
interpolated into an inline onclick="startPixletEditor('...')", so a directory
containing an apostrophe could break out of the JS string and run script. The
handler binds with addEventListener and reads the id from dataset, where it is
only ever parsed as an HTML attribute.

Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in
in both directions. The tools DOM suite gains three guards asserting the edit
buttons carry no inline onclick and pass the id via dataset -- those need jsdom
and did not run here, so CI verifies them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): log the traceback on the editor state-write failure

The 848 fix answers 500 when the session state cannot be written, and logged
that at error level -- but without exc_info, so the traceback never reached the
log. test_web_error_detail.py guards exactly this: a handler returning 5xx must
write an error-level record *with* the traceback and return the sanitized
detail, because checking that merely something was logged is too weak.

Caught by Core unit tests on the previous commit, not locally: the guard parses
every module under web_interface/blueprints/api_v3 as one source, so it only
fires once the whole package is read together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 11:32:12 -04:00
ChuckandClaude Opus 5 bdb9a94033 refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package

web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181
functions in one module -- 9% of the core by line count and three times the
next largest file. It becomes a package of nine route modules grouped by path
segment, plus __init__.py for the shared imports, constants, Blueprint and
helpers.

Every route module decorates the SAME api_v3 Blueprint object, so endpoint
names stay api_v3.<function>, the URL map is unchanged and app.py is untouched.
Verified: 111 routes before, 111 after, byte-identical rules, endpoints and
methods, and every endpoint still on the one blueprint.

  plugins   3,867   config    1,178   starlark  692   system  619
  fonts       452   misc        398   wifi      361   display 326   backup 212
  __init__  1,787 (imports, constants, Blueprint, 56 helpers)

Two things the URL-map check could not catch, both found by running the suite:

1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one
   directory deeper made that resolve to web_interface/ instead of the project
   root. Nothing failed at import; it surfaced as ~110 tests failing with 404s
   and "installation script not found", because every path built from it was
   one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts
   PROJECT_ROOT/run.py exists so the next move cannot repeat it.

2. Module-attribute patching. Tests do
   monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route
   module that binds such a name by value never sees the patch. The shared code
   therefore stays in __init__.py rather than moving to a _common submodule --
   it has to live on the module the tests patch -- and the eleven names tests
   patch are read back through the package (_pkg.X) instead of bound by value.
   Those eleven were found by AST-scanning every setattr in the test tree, not
   by guessing; "time" is among them, used to drive a fake clock through the
   second-resolution credential-backup filenames.

Test changes are confined to what genuinely moved: patch targets that now name
the owning route module, imports of helpers, and six tests that scan the api_v3
source as a file and now read the package directory.

Full suite: 4,278 passed, 68 skipped, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(api-v3): address CodeRabbit findings from the blueprint-split review

Fixes to the api_v3 package split (PR #553), one per finding verified
against the actual code:

- __init__.py: _redact_credentials only blanked scalar values under a
  credential-named key; a bare list of secrets under such a key (e.g.
  tokens: ["a", "b"]) passed through untouched, since the list branch
  recursed with no memory that its key looked like a credential. Nested
  dicts still walk normally (a documented, tested behaviour -- a container
  like secrets: {api_key: ..., note: ...} is a section name, not a value to
  blank outright), but any value reached under a credential-shaped key is
  now actually blanked.

- __init__.py: the OAuth helper script's raw stderr/stdout went to
  logger.error unredacted (CWE-532) right next to a comment claiming this
  was deliberate; the HTTP response already used the existing redact_text
  helper. Routed the log line through the same helper.

- __init__.py / starlark.py: the standalone Starlark manifest fallback
  (used when the plugin instance isn't loaded) read-modified-wrote
  manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe
  (plugin-repos/starlark-apps/manager.py), which already holds an flock for
  the same file when the plugin is loaded. Added _starlark_manifest_lock,
  mirroring that pattern, and wrapped every standalone read-modify-write
  call site in it. The app-config update route also wrote config.json and
  the manifest as two separate, non-transactional writes (a second,
  distinct finding at the same call site); config.json is now rolled back
  if the manifest write that follows it fails.

- backup.py: restore options used bare bool() on values from the request,
  so {"restore_secrets": "false"} restored secrets anyway (bool("false") is
  True). Switched to the existing _coerce_to_bool helper already used for
  this exact purpose elsewhere in the package.

- config.py: an automated import-rewrite mangled four user-facing
  validation strings and their neighbouring comments -- "Invalid start
  time" had become "Invalid start _pkg.time" (and likewise for "end time")
  in both the schedule and dim-schedule per-day validation paths.

- display.py: `import _pkg.time as time_module` -- _pkg is a local alias
  for the package, not a real importable module, so this raised
  ModuleNotFoundError whenever a caller restarted an already-running
  display service via /display/on-demand/start, after the on-demand
  request was already written to cache. Fixed to `import time`. Audited
  the rest of the package for the same `_pkg.<module>` import mistake;
  every other `_pkg.` reference is a legitimate attribute read-through
  (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken
  import statement.

- fonts.py: validate_file_upload's max_size_mb parameter is silently
  unused by that helper (it only checks filename/extension) -- the font
  upload route saved arbitrarily large files as a result. Added the same
  seek-and-check pattern already used for the sibling .star upload.

- wifi.py: two ad hoc, inconsistent bool coercions. POST
  /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled
  it. POST /wifi/radio's enabled/force parsing recognized real bool and
  some strings but not int 1/0 (1 is True is False in Python). Factored one
  small _parse_bool_ish helper local to this file and used it at all three
  sites.

Not changed: the "unknown/misspelled restore option keys default to True"
half of the backup.py finding -- the file's own comment documents that a
missing key deliberately means "restore everything," matching the
already-existing JSON-parse-failure guard a few lines above it; only the
bool-coercion defect was a real bug.

Added or extended regression tests for every fix, following each area's
existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed
on both this branch and origin/main (missing tzdata package breaks two
timezone-alias tests in test_onboarding_checklist.py, unrelated to this
change) -- no new failures.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5

* fix(api-v3): reject unknown restore option keys

CodeRabbit's review of the blueprint split (#553) asked that
POST /backup/restore reject option keys outside RestoreOptions'
known set. The follow-up commit fixed the bool("false")-is-True
bug with _coerce_to_bool but never added the key check: a typo'd
or renamed key (e.g. "restoreSecrets") is silently ignored by
opts_dict.get(key, True), so the flag stays at its True default
and secrets get restored despite the caller's request saying
otherwise -- with no indication anything was wrong.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb

* fix(api-v3): address CodeRabbit findings on the blueprint split

- _redact_credentials: blank scalar descendants of objects reached
  through a credential-owned list (e.g. tokens: [{"value": "secret"}])
  regardless of field name -- the existing name-based walk only
  protected direct dict values under a credential key, not list items.
- wifi.py: reject enabled/force/auto_enable_ap_mode values
  _parse_bool_ish can't recognize (400) instead of silently treating
  them as False, which could disable Wi-Fi or the radio itself.
- Starlark manifest locking: lock a stable manifest.json.lock sidecar
  instead of manifest.json itself, in both the standalone route path
  (_starlark_manifest_lock) and the plugin path
  (StarlarkAppsPlugin._save_manifest / _update_manifest_safe).
  manifest.json is replaced by an atomic rename on every write, which
  swaps in a fresh inode; a lock held on the old inode does not
  exclude a second locker that opens the path afresh right after the
  rename and gets the new inode, so two writers could race despite
  each holding "a lock". A sidecar that no write ever touches always
  resolves to the same inode for every locker.

Skipped as stale: the "serialize the complete manifest
read-modify-write" finding at api_v3/__init__.py -- every standalone
handler that calls _write_starlark_manifest is already wrapped in
_starlark_manifest_lock() on this branch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-v3): re-check reconciliation findings by the reconciler's own rules

Both CodeRabbit findings on the merge commit, verified against the code first.

Major, plugins.py: the stale-findings filter derived its own notion of "in
config" and "on disk", and both were looser than the reconciliation module's.
set(load_config()) also contains system keys, the secrets-file keys load_config()
merges in, and non-dict values; and any directory holding a manifest.json
counted as installed even when that manifest does not parse. Either looseness
clears a finding that is still true -- and a secrets key read as a plugin is the
precise bug the filter exists to stop reporting, so reintroducing that asymmetry
while re-checking was the wrong way round.

The two extractions now live in state_reconciliation.py as config_plugin_ids()
and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys()
alongside. _get_config_state() and _get_disk_state() use them too, so there is
one definition rather than two that can drift. _get_disk_state() re-reads each
manifest for version/name after taking membership from the shared extractor;
that costs one extra small read per plugin on a path that runs once per boot.

Minor, the new test: the fixture assigned api_v3.config_manager and
api_v3.plugin_manager directly. Those live on a module-level blueprint
singleton, so the mocks leaked into every later test that imports api_v3 --
pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr,
which restores them. This is the same pollution class that made an earlier test
in this session break seven unrelated ones, so it is worth getting right.

Five cases added for the parity itself: a secrets key, a system key and a
non-dict value must not clear an "installed but missing from config" finding,
and neither an unparseable manifest nor a .standalone-backup- directory may
count as installed. All five fail against the looser version.

Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and
CodeRabbit all pass. Codacy reads action_required on every commit of this
branch including the first, so it is pre-existing and not from this work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 10:07:38 -04:00
ChuckandClaude Opus 5 0ab95586fb fix(web): say when a system action failed for want of passwordless sudo (#560)
* fix(web): say when a system action failed for want of passwordless sudo

POSTing reboot_system to a Pi returns, in full:

  {"message": "Action failed; see logs for details", "status": "error"}

The cause is that the web interface runs unprivileged, and its
systemctl/reboot/journalctl calls only work once
scripts/install/configure_web_sudo.sh has granted NOPASSWD. first_time_install.sh
never invokes that script and no user-facing doc mentions it, so on a fresh
device every privileged action fails -- start_display, stop_display, the
autostart toggles, reboot, and the log viewer.

That last one closes the loop: "see logs for details" is unreachable advice
when journalctl is refused for the same reason. This is exactly the failure
src/web_interface/error_handler.py's describe_exception() was written to break,
and /system/action's exception handler was still discarding the cause instead
of using the helper the module already imports.

Two changes, no behaviour change when things work:

- The exception path now returns 'details': describe_exception(e), matching how
  the other handlers in this blueprint already report.
- A failure whose stderr or exception text is sudo refusing to prompt ("a
  password is required", "no tty present", "a terminal is required") reports
  what to do about it, naming configure_web_sudo.sh. Unrelated failures keep
  the generic message and their stderr, so a missing unit is not blamed on
  sudo.

Granting the sudo rights is left alone deliberately: auto-running a script that
hands out NOPASSWD is a security decision for the maintainer, not something to
slip into an installer. Making the refusal legible is the part that is
unambiguously an improvement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): apply the sudo hint on the on-demand start_display path too

start_display with a mode builds its own response and returns before the shared
nonzero-result path, so a recognized sudo refusal there reported only "Failed to
start display" and said nothing about the passwordless sudo that refused it --
the exact gap the rest of this PR closes everywhere else.

Raised by CodeRabbit on #560 and verified against the code before fixing: the
branch at api_v3.py:2058 does return early past the shared handler.

Three regression cases: the on-demand branch reports the sudo cause, keeps its
"Display started" message on success, and does not blame an unrelated failure on
sudo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:45:40 -04:00
ChuckandClaude Opus 5 39f27d285d fix(plugins): stop reconciliation inventing plugins and telling users to delete real config (#557)
On a device running four installed, configured, working plugins, the overview
banner read:

  Stale plugin config entries found: football-scoreboard, odds-ticker, data,
  ledmatrix-weather, starlark-apps. Remove them from config.json or reinstall
  via the Plugin Store.

Every claim in that sentence was wrong, and following its advice would have
deleted 4.9KB of working league settings. Four separate defects combined.

1. Secrets keys became phantom plugins. load_config() merges
   config_secrets.json into the config it returns, and the ignore list named
   only 'github' and 'youtube'. A 'data' key in that file therefore read as a
   plugin id and was reported as "in config but not on disk" forever. Read the
   secrets file's own top-level keys instead of hardcoding two of them.

2. The auto-fix clobbered real config. The handler for "on disk but not in
   config" assigned `config[plugin_id] = {'enabled': False}` unconditionally,
   so whenever detection was wrong it replaced a plugin's entire configuration
   with a stub. On the reported device it only failed to do so because the
   write hit EACCES. Now it refuses to overwrite an entry that already exists.

3. The banner gave backwards advice. plugin_missing_in_config ("on disk, not in
   config") and plugin_missing_on_disk ("in config, not on disk") are opposite
   problems, and both were rendered as "stale config entries ... remove them
   from config.json" -- which is correct for the second and destructive for the
   first. They are now reported separately, each with the advice that fits.

4. A stale verdict was served indefinitely. The result is a snapshot written
   once per run to a status file, and a run that fails to apply a fix also
   declares it will not retry. A condition that had since resolved kept being
   reported for hours. The status endpoint now re-checks stored findings
   against current state, dropping only what it can prove stale and keeping
   any kind it cannot re-verify.

The secrets-key lookup is deliberately fail-safe: an unreadable, absent,
malformed or non-path secrets location narrows the ignore set rather than
raising. An earlier revision let TypeError escape, which the broad handler in
_get_config_state() swallowed as "Error reading config state" -- emptying the
config state and making every downstream detection wrong. The existing
reconciliation tests caught it; there is now a regression test for it too.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:45:21 -04:00
ChuckandClaude Opus 5 aba96e25b3 chore: delete three functions nothing calls (#550)
src/base_classes/baseball.py       _get_baseball_display_text   45 lines
  src/web_interface/api_helpers.py   validate_request_params      22
  web_interface/blueprints/api_v3.py _validate_time_range         14

Each has exactly one occurrence across both repositories -- its own
definition. No decorator, no __all__, no getattr dispatch, nothing in
templates or JavaScript.

A fourth candidate was dropped after checking: _unshare_element_fonts in
src/common/sports_shared.py looked unreferenced, but eight scoreboard plugins
call SportsCore._unshare_element_fonts directly from their
test_element_text_colors.py, plus their own copies at runtime. It is live API.
The earlier reading came from a plugins checkout 84 commits behind main, which
is a good argument for re-verifying this kind of claim against a fresh tree
rather than trusting an earlier scan.

Full suite: 4,265 passed, 68 skipped.


Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 08:45:07 -04:00
ChuckandClaude Opus 5 dcd6e39c96 fix(web): report real disk usage and MemAvailable on the live status stream (#558)
The SSE status stream sent 'disk_used_percent': 0 as a literal, so every
consumer of the live view showed 0% disk no matter how full the card was.
/api/v3/system/status computed it correctly; the stream that the dashboard
actually watches did not. On a Pi with a modest SD card that is the warning a
user most needs, and it was guaranteed to never appear.

The stream also omitted memory_available_mb. /api/v3/system/status carries it
with a comment spelling out why it matters: MemAvailable accounts for
reclaimable page cache, so it is what separates a board reading 70% "used" that
is fine from one reading 70% that is about to fail fork(). A 1GB Pi 3B+ can sit
at either. The number that predicts the failure was missing from the live view.

An unreadable disk now reports None rather than 0. The UI already renders null
as '--'; a confident 0 reads as "plenty of room", which is worse than a blank.

Metric collection moves to web_interface/system_metrics.py, with no Flask or app
imports. That is not cosmetic: importing web_interface.app constructs the Flask
application and a CacheManager, and the latter claims the cache directory with a
cleanup thread. The first version of these tests imported the generator directly
and broke test_cache_cleanup_thread_ownership ("one thread per directory") plus
four starlark route tests through that side effect. Reading a CPU percentage
should not boot a web application, and testing it should not either.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-11 08:42:30 -04:00
ChuckandClaude Opus 5 4423ec33d5 feat(plugins): search, filter and sort for Installed Plugins, on a shared ListFilter helper (#540)
* feat(plugins): add search, filter and sort to Installed Plugins, on a shared helper

The Installed Plugins grid had no way to narrow it down: no search, no way to
see only what's enabled, disabled, or out of date. On a rig with a couple dozen
plugins that means scrolling the whole grid to find one.

The two sections below it already solved this, twice, independently — the
Plugin Store and Starlark Apps carried a copy-paste fork of the same ~600 lines
(filter state, apply-filters-and-sort, page renderer, pagination strip,
active-filter badge, listener wiring). Rather than add a third copy, this
extracts the shared machinery and builds the new toolbar on it.

New: web_interface/static/v3/js/plugins/list_filter.js — ListFilter.create()
owns debounced search, filter axes, sort, the active-filter count, Clear, and
optional pagination/persistence. Callers keep their own card markup via a
`render` callback. Three control types cover every axis the page uses: pills
(new), select (store category, starlark author) and cycle (the tri-state
All -> Installed -> Not Installed button).

Installed Plugins gets a compact toolbar: search box, one-click All / Enabled /
Disabled / Updates pills, and a sort dropdown (A-Z, Z-A, updates first,
recently updated, category). Filters reset on load, so you never come back to a
mysteriously short list. No new CSS — this is the first consumer of the
.filter-pill rules already sitting unused in app.css.

renderInstalledPlugins() is split so it still publishes canonical state while
renderInstalledCards() draws only the visible subset; the filtered list is
never assigned to window.installedPlugins, which the toggle handler,
isStorePluginInstalled(), runUpdateAllPlugins() and the Alpine config tabs all
read as their source of truth. Toggling a plugin while filtered pins its card
so it doesn't vanish from under the cursor.

The Store and Starlark migrations are behaviour-preserving: same element ids,
same localStorage keys (storeSort/storePerPage, starlarkSort/starlarkPerPage),
same tri-state button markup, same pagination. Verified by differential tests
that run the old and new implementations side by side against identical
fixtures and compare every observable after each interaction. The only visible
change is the pagination attribute (data-store-page/data-starlark-page ->
data-list-page), which nothing outside its own click handler referenced.

Net -156 lines in plugins_manager.js while adding a feature.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): keep raw search text, and stop the store search refetching

Two review findings from CodeRabbit on #540.

Do not write the trimmed search value back into the input. setSearch() trimmed
before storing, and syncControls() then copied that trimmed value back over what
the user had typed. Pausing longer than the debounce after typing a space
deleted the space (and reset the caret), making multi-word terms effectively
untypable. The raw text is now kept alongside the trimmed one: filtering and
activeCount() still use the trimmed value, while the input keeps exactly what
was typed.

Remove the legacy #plugin-search / #plugin-category listeners in
initializePlugins(). They bound searchPluginStore as the event handler, so the
DOM event arrived as its `fetchCommitInfo` argument — always truthy, which
skipped the cached-filter fast path and refetched /api/v3/plugins/store/list
with commit info on every keystroke burst and category change. The store's
ListFilter controller already filters the cached list, which is what those two
controls should do. This double-binding predates this PR (the old code guarded
with _listenerSetup and _storeFilterInit, two different flags, so both sets
stayed live); it is fixed here because the refactor owns that wiring now.

Both fixes are covered by tests that fail without them: the trailing-space
regressions in the installed-plugins DOM suite, and a new whole-file jsdom test
that counts fetches while typing (1 request at init, 0 thereafter; previously
1 -> 2 -> 3 -> 5).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): build pagination via DOM APIs, drop computed member access

Addresses the five Codacy security findings, all in list_filter.js.

Pagination no longer assembles an HTML string (3 findings: 2 critical + 1 high,
"unsafe assignment to innerHTML"). The interpolated values were only page
integers and local class constants, so there was no injection path, but
concatenating markup into innerHTML is the pattern the scanners flag and
createElement is no less clear. Each button now also owns its click listener
directly instead of the container being re-queried afterwards, and the strip is
cleared with textContent = '' rather than by assigning empty markup. No
innerHTML assignment remains in the file.

haystack() now walks Object.entries(item) and keeps the configured fields,
instead of reading item[field] per field ("generic object injection sink").
Field order no longer drives the haystack order, which is irrelevant to the
substring test. matches() iterates controls with for...of instead of an index
("variable assigned to object injection sink").

The rendered pagination is unchanged: same buttons, labels, page numbers,
disabled states and classes. The old-vs-new differential tests now compare
pagination structurally (tag, text, page, disabled, sorted class list) rather
than as an HTML string, since building nodes legitimately serialises
differently — «/» as characters rather than &laquo;/&raquo;, disabled="" rather
than a bare attribute. That comparison is stronger than the string one it
replaces, and the real-DOM suite still drives the actual page buttons.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* fix(plugins): keep configured field order when building the search haystack

The previous commit swapped item[field] for Object.entries(item) to clear a
static-analysis object-injection warning, and in doing so changed the order of
the haystack: entries follow the object's own key insertion order, not the
configured `fields` order. Since the values are concatenated, that order decides
which values end up adjacent, so a multi-word query spanning a field boundary
matched differently. For store fields [name, description, author, id, ...] and
API objects keyed {id, name, description, author, ...}, "bob plugin-01" matched
before and stopped matching after.

That contradicted the behaviour-preservation claim for the store and starlark
migrations, and the differential tests missed it because every fixture query was
a single word.

Values now come out of a Map built from Object.entries, iterated in `fields`
order: the original haystack is restored, and there is still no computed member
access for the analyser to flag.

Regression coverage for the ordering itself, at both levels:
  - unit: phrases spanning name->id and category->tags, plus the reverse
    (object-key) order asserted NOT to match
  - differential: the same class of query compared old-vs-new, with a guard that
    the phrase actually matches something so a mutual zero-result cannot pass
    vacuously

Verified both fail without this fix (3 unit, 2 differential) and pass with it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* test(web): add JS suites for ListFilter and the plugin-manager grids

No JS toolchain exists in this repo, so these are plain node scripts with no
framework: each prints ok/FAIL lines and exits non-zero. `node test/js/run_all.js`
runs everything, skipping the DOM suites (rather than failing) when jsdom is
absent or nothing is listening, so it stays useful in a bare checkout.

  unit/test_list_filter.js    ListFilter search/filter/sort/count/sticky, driven
                              through the installed-plugins config eval'd
                              verbatim out of plugins_manager.js so the test
                              cannot drift from the real configuration
  unit/test_render_cards.js   renderInstalledCards markup, both empty states,
                              and escaping of hostile plugin metadata
  dom/test_installed_dom.js   the toolbar in a real DOM, including the HTMX
                              partial re-swap and a getComputedStyle check that
                              .filter-pill[data-active] matches what we emit
  dom/test_store_dom.js       store pagination, per-page, category, tri-state
                              Installed button, persistence across a re-boot
  dom/test_no_double_fetch.js loads the whole plugins_manager.js and counts
                              requests, so a keystroke cannot refetch the store

The DOM suites deliberately fetch the partial and the plugin data from a running
web interface instead of using fixtures, so a renamed element id or a changed
payload shape fails them loudly. Point them at a rig with a full plugin set when
it matters (BASE=http://host:5000); a dev box with two plugins installed passes
while exercising very little.

Several assertions exist to stop specific bugs recurring: trailing spaces
surviving the search debounce, a query spanning two adjacent search fields
(haystack field order is load-bearing), and window.installedPlugins staying at
full length while the grid is filtered. Others guard against passing vacuously —
counting only non-skeleton cards, and checking a search phrase matches something
before comparing two result sets.

The old-vs-new differential suites that verified the store and starlark
migrations are not included: they compared against the pre-refactor code, which
now exists only in git history.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 20:04:47 -04:00
ChuckandClaude Opus 5 26769ee37f fix(starlark): the store authenticated with a key nothing writes (#541)
#535 restored the thirteen routes, so the store stopped answering 404 --
and still would not load. Confirmed against a running device before
anything was changed: /repository/browse answers 200 with 1000 apps in
27s, so the routes are fine. Two things underneath them are not.

**The store never used the token the user configured.** The three
repository routes read `github_token` off config.json. Nothing writes
that key -- it is not in config.template.json, no setting offers it, and
it appears nowhere else in the codebase. The configured token goes to
config_secrets.json as `github.api_token`, which PluginStoreManager
loads and every other GitHub caller uses. So the store could never be
authenticated: 60 requests/hour, on the same per-IP budget 48 installed
plugins spend on update checks, while the 5000 the user had already
configured sat unused. On the device, /plugins/store/github-status
reported authenticated with a limit of 5000 at the same moment
/starlark/repository/browse reported 60, with 18 left. The store going
blank was that 60 running out.

**Every failure looked identical.** list_all_apps_cached turned any
listing failure -- rate limit, DNS, timeout, non-200 -- into an empty
app list, and the route sent that out as `status: success`, so a rate
limit and an empty repository drew the same blank grid with no error
anywhere. It now returns the reason, the route answers 502 with it, and
a failure is no longer cached as an empty repository for two hours.

The guard for a bad response was itself a crash: _make_request catches
`(json.JSONDecodeError, ValueError)` but `json` was never imported, so
evaluating the tuple raises NameError and the guard written for exactly
this case never ran. Reachable whenever something on the path answers
with HTML -- a captive portal, a proxy page, a DNS-hijacking router.

Seventeen handlers answered 5xx with no detail at all.
test_no_api_v3_handler_discards_its_exception is meant to prevent that
across api_v3, but it matched one exact message string, and all thirteen
Starlark routes wrote their own wording. The guard now keys on the shape
that matters: if it returns 5xx, it says why. The 15 pre-existing
non-Starlark functions are listed as a set that may shrink, never grow.

**The listing was capped at 1000 and did not say so.** The contents API
truncates a directory silently; tronbyt/apps has 1075 app directories,
so the store showed a truncated repository and looked complete doing it.
Now listed via the git trees API, which reports `truncated`, with the
contents API kept as a fallback.

Not addressed: the 27-second cold load -- 1075 manifests fetched five at
a time behind skeleton placeholders -- which is probably the largest part
of what "does not load" feels like, and wants its own change.

25 new tests.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:04:27 -04:00
ChuckandClaude Opus 5 e23f1f45d3 feat(starlark,on-demand): the third-party fixes worth taking, plus a Home Assistant MQTT bridge (#538)
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge

Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of
patches and services built while running this project on Starlark apps
under MQTT control. Its patches are whole-file copies taken against an
older tree, so applying them as written would revert #523's frame
pacing, #534's display() bool returns and the GitHub token masking in
plugins_manager.js. Three of its claimed fixes are already in main, and
its api_v3 Starlark routes are #535's. What follows is the rest --
verified against current code, and reimplemented where the patch's
approach did not hold up.

**On-demand display.** `pinned` reached the controller from the API, was
stored on it and republished in the status payload, but never narrowed
the rotation -- a pinned request still cycled every mode its plugin
owns. Right for a sports plugin, whose modes are views of one subject;
wrong for a plugin whose modes are unrelated, which is every Starlark
app. Now honoured, and it survives a restart.

Restarting while on-demand was active loaded *only* the on-demand
plugin, so normal rotation had nothing to return to for the life of the
process -- and a restart mid-session is routine, since that is how an
update is applied. The panel came back cycling one plugin's modes with
no way out but clearing the cache by hand. Every enabled plugin loads
now; on-demand still resumes on its saved mode.

Stop requests are exempt from the duplicate guards on purpose, so that a
second click stops a mode a race left running -- which means consuming
the mailbox is the only thing that ends one. It was never consumed, so
the same stop was re-read and re-processed on every poll, forever. Both
paths now share one compare-before-delete helper.

**Starlark rendering.** `extract_schema` parsed the source with a regex,
which can only see option lists written out literally: an app whose
dropdown is filled from a live API call inside `get_schema()` came back
empty, and the config form offered nothing to pick. Now runs `pixlet
schema`, which executes the app, and falls back to the parser when
Pixlet is absent, too old for the subcommand, or the app fails to run.
The third-party patch replaced the parser outright and hardcoded
/usr/local/bin/pixlet; this keeps the fallback and the binary search.

A `|` in a config value was dropped by a shell-metacharacter filter,
though the command is a list with no shell involved -- and apps do use
it as a separator inside one value. The key went missing silently and
the app rendered its own "not configured" screen with nothing to say
why. And a 0-byte render was reported as success: Pixlet exits 0 and
writes nothing when an app has no content, which read downstream as a
working app drawing a black panel.

**Starlark display.** `display()` ignored the mode it was called with,
so a specific app could not be addressed. It now accepts `display_mode`
-- which is the whole mechanism, since the controller inspects the
signature before passing it. Found while there: `_select_next_app` ran
only while `current_app` was unset, so with several apps installed the
first was picked once and shown forever while the rest were rendered on
schedule and never displayed. And `enable_scrolling` was missing, so
multi-frame apps were called once per rotation slot and never advanced
past frame one.

**GET /api/v3/display/modes.** Every mode that can be requested
on-demand, with the plugin that owns it. Nothing exposed this, so
anything driving the display from outside the web UI read each plugin's
manifest.json off disk and reimplemented PluginManager's fallbacks. It
also triggers discovery, which is otherwise lazy and normally happens
because a person opened the dashboard.

**integrations/mqtt_bridge.** Home Assistant control over MQTT
Discovery: a mode select, a stop button, power, brightness. Rewritten
against the API rather than the filesystem, so it needs no read access
to config.json and cannot drift from the web UI. paho-mqtt 2.x
VERSION2, TLS, an availability topic that is also the last will, and
secrets from the environment.

**Two opt-in extras.** A DNS single-request unit, for glibc's parallel
A/AAAA lookup stalling ~5s per name on routers that answer only the A
query -- which makes any plugin calling an external API slow and
Starlark apps, which have a render timeout, fail outright. And a Pixlet
config editor: a script you run and Ctrl+C rather than the third-party
version's always-on unauthenticated Flask service, since it stops the
display for the length of a session. Neither is installed by default.

Long Starlark app names now wrap instead of overflowing their card.

115 new tests across 5 files. Also unblocked
test_starlark_display_contract.py, which was silently skipping wherever
fcntl is absent. Whole suite: no new failures against main.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(mqtt_bridge): the five issues Codacy flagged on this branch

All in the new bridge, all real:

  * requests floor was 2.31.0, which carries CVE-2024-35195,
    CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which
    is what the project's own requirements.txt already pins.
  * `import time` was never used.
  * `"mqtt_password": None` in DEFAULTS read as a hardcoded credential.
    It is the "no password configured" default; marked nosec B105, the
    convention used elsewhere in the repo.

Also dropped an unused `build_app` from the display-modes test imports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: the review findings on this PR

Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and
is answered below.

**One bad config section blanked the whole mode list.**
`/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`,
so a non-dict under a plugin id -- a shape DisplayController already
guards, so it happens -- raised AttributeError mid-loop and answered 500
with no modes at all. Every MQTT bridge entity is built from that list.
Now skipped with a warning.

**The DNS scripts reported success they had not earned.** Three separate
paths: `resolvconf -u` failing was swallowed by `|| true`; the
systemd-resolved branch exited 0 without applying anything, so the
oneshot unit recorded success while the workaround was inactive; and the
installer's `|| echo` turned a failed start into "installation
complete." with exit 0. All three now fail loudly. `single-request` is a
glibc resolv.conf option with no resolved.conf equivalent, so on those
hosts the honest answer is that it cannot be applied.

A NetworkManager-generated resolv.conf is regenerated on connection
changes, not only at boot, and the unit is oneshot with RemainAfterExit
-- so the option can vanish mid-boot with nothing to put it back. Now
detected and stated plainly rather than implied to be permanent.

**`Before=` does not order a manual restart.** It only orders units
already in the same transaction, so `systemctl restart ledmatrix` could
bypass the fix. install_dns_fix.sh now writes a ledmatrix.service
drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround
failing should not stop the display.

**The Pixlet editor's `--lan` is gone.** `pixlet serve` has no
authentication, and a printed warning is not access control. Loopback
only, with the SSH port-forward in the header where the flag used to be
documented -- SSH does the authenticating and nothing is left listening.

**The MQTT example config now defaults to TLS** on 8883. The installer
copies it verbatim, and without TLS the broker password and every
command cross the network in cleartext. A plaintext broker is still
supported and documented, and the bridge warns once at startup when a
password is configured without TLS.

**Not taken: "the upstream Pixlet CLI has no `schema` subcommand."**
Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh`
installs `tronbyt/pixlet`, whose `cmd/schema.go` is
`schema [PATH]` -> JSON on stdout, built on
`runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is
exactly what extract_schema_via_pixlet calls. A binary without the
subcommand exits non-zero and falls back to the source parser, which is
already covered by a test.

**CodeQL stack-trace exposure: not taken either.** I removed `details`
first and that broke
test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception,
which enforces `describe_exception` across all ~75 handlers -- written
because a device with failing storage answered "see logs for details"
from the log viewer itself. describe_exception redacts credentials; the
trade-off is the project's and is already made. Restored, with the
reasoning in a comment.

11 new tests. Whole suite: no new failures against main, 4127 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 08:34:52 -04:00
ChuckandClaude Opus 5 793b988d33 fix(starlark): toggle the app id the list published, and store a relocatable star_file (#537)
The two review nitpicks left over from #535. Both are still on main after
that merge; the five findings alongside them landed with it.

**The toggle could not find what the list had just shown.**
`_starlark_virtual_plugins` publishes the raw manifest key as
`starlark:<key>`, and `_toggle_starlark_app` passed it back through
`_validate_and_sanitize_app_id`, which lowercases and rewrites every
character outside `[a-z0-9_]`. An app stored as `My-App` was listed as
`starlark:My-App` and looked up as `my_app`, so toggling an app the page
had drawn a moment earlier answered 404. Keys written by
`_install_star_file` are already sanitised, so this only shows up for
manifests written by the starlark-apps plugin itself or edited by hand.

`_validate_starlark_app_path` rejects traversal without rewriting, so it
is the check to use here -- listing and toggling now agree on one key.
The updater also uses `setdefault` rather than indexing: the app is
loaded but its on-disk entry need not exist, and `_update_manifest_safe`
does not catch `KeyError`, so that escaped as a 500 rather than writing
the entry.

**`star_file` was stored absolute.** Readers join it to the app's own
directory -- `_standalone_render_starlark_app` does `app_dir /
app_data.get('star_file', f'{app_id}.star')` -- so the key's default is a
bare filename and an absolute value gave it a second meaning. Since
`Path.__truediv__` discards the left side when the right is absolute,
the manifest was pinned to whatever PROJECT_ROOT installed it, and a
moved or redeployed install could not find its own file. Storing
`dest.name` matches the default and stays relocatable. Read paths are
unchanged, so manifests already holding an absolute path keep working.

7 new tests. Whole suite: no new failures against main, 4013 passed
against 4007.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 16:35:56 -04:00
Chuck 50258635a8 fix(starlark): restore the API routes #330 dropped (#535)
The Pixlet install button reported "Pixlet install failed: Resource not found" -- Flask's 404 handler, because the route did not exist. #253 added thirteen Starlark routes; #330 rewrote api_v3.py and dropped all of them, along with the `starlark:<app_id>` entries that surface installed apps in the plugins list and the toggle branch that enables them.

Restores all thirteen routes, the plugin-list entries and the toggle path, so Pixlet installs, the app store browses and installs, and an installed app can be managed like any other plugin.

Not a straight revert. Three error paths stopped returning exception text to the caller; the manifest write moved off a shared temp filename that two concurrent writers could interleave; both dynamic importers stopped leaving half-initialised modules in sys.modules; the config update rolls back when the save fails; the toggle checks that persistence succeeded; and the path check returns the validated path instead of a boolean so callers stop re-joining the raw value. New tests no longer reach GitHub.

Verified on a 256x64 Pi: Pixlet installs and runs (v0.53.1), the store lists 1000 apps, install/toggle/uninstall round-trip, and traversal and command-injection probes are rejected at every entry.

25 CodeQL alerts dismissed as verified false positives -- path-injection where traversal is blocked, and one list-form subprocess with no shell. Both classes already present on main.

Full core suite: 3981 passed.
2026-09-07 16:06:09 -04:00
ChuckandClaude Opus 5 a0d3e64099 fix: ten defects found validating the whole plugin fleet on hardware (#534)
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers

Three independent fixes found while validating every plugin on a 256x64 rig.

FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf
ships with the core, so every plugin offering it logged "Font family
'tom_thumb' not found" (16 warnings per countdown render) and had to carry a
private loader to use a bundled font. Closes #524.

VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter
that DisplayManager gained, so any plugin passing it died with TypeError at
render time and failed every size. Nine plugins now make that call;
ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other
seven only passed because their scroll path was unreachable without data.
Closes #525.

api_v3 declared module-level config_manager/plugin_manager = None that nothing
ever assigned -- app.py sets the blueprint attributes, which the other 150+
call sites use. Three sites read the decoys, so /health reported the config
unreadable and the plugin system uninitialised (making "degraded" permanent and
unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on
every rig. The decoys are removed rather than assigned, so a bare name is now a
NameError at test time instead of a silent None. The same function's first-call
uptime was computed from two separate clock reads and came out negative.
Closes #529.

Verified on the rig: both previously-failing plugins render, the tom_thumb
warnings are gone, /health reports "healthy" with all three checks passing, and
/display/current reports the real 256x64.

* fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark

The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is
world-writable and sticky, and the display service runs as a different user
from the tooling, so a leftover temp owned by anyone else became unopenable
even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a
sticky directory. The preview and the health check's liveness proxy then froze
until someone deleted the file by hand; on the test rig that meant 23 hours of
a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with
cleanup on failure, matching the hardware-status write a few hundred lines
above. Closes #528.

_poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults
the in-memory TTL to max_age -- so the first request was pinned in memory for an
hour and every later poll returned that stale copy. No second on-demand request
was honoured until the service restarted, while the API kept returning 200.
get() already documents memory_ttl=0 for exactly this cross-process case.
The consumed request is also now deleted: leaving it on disk meant a restart
replayed the previous request, activated it, and ignored the one the caller had
just made. Closes #530.

starlark-apps returned None from display() when it has no app to show, which is
the state of every install without Pixlet and of a fresh one before any app is
added. The controller only skips on a boolean False, so that held a black panel
for the full display_duration instead of rotating on. Closes #456 (core side).

Verified on the rig: two consecutive on-demand requests with no restart between
them are both activated, where the second was previously dropped in silence.

* perf(harness): share one cache across a plugin's renders

_instantiate built a fresh MockCacheManager for every (size, mode), and that
mock is a per-instance in-memory dict, so each render was a cold start. A plugin
that fetches per game or per player re-fetched everything N times over --
baseball-scoreboard at one size took 840s for nine renders where the arithmetic
said ~72s, and at eight sizes it exceeded a 900s timeout.

The second and later renders also never exercised the cache-hit path, which is
what a running rig executes almost all of the time, so a caching regression
could not be caught here.

The cache is now built once per render_plugin_matrix call and threaded down.
The display manager stays per-render -- the bounds checking depends on that --
so only fetched data is shared.

Measured on the rig, same render counts and same goldens:
  tide-display         2s -> 1s   (32 renders)
  cricket-scoreboard  10s -> 3s   (24 renders)
No pass/fail change across tide-display, cricket-scoreboard, clock-simple,
geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages.

Closes #533.

* fix(scripts): run standalone plugin tests instead of collecting nothing

run_plugin_tests.py discovered every plugin test file and handed the lot to
pytest. Most plugin tests are standalone scripts -- module-level main() plus an
`if __name__ == "__main__"` guard, signalling through an exit code -- and pytest
collects zero items from those. The run printed how many files it had *found*,
then "no tests ran", and exited without executing any of them. On a rig with all
44 first-party plugins that is 151 of 248 files.

Files are now classified and each kind runs under the right runner: pytest for
real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1
fail convention ledmatrix-plugins' own runner established (a script that wants a
tty or an LED matrix is a skip, not a regression).

Before:
    $ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos
    Found 1 test file(s)
    collected 0 items
    no tests ran in 0.31s                      rc=0

After:
    Found 1 test file(s) -- 0 collectable, 1 standalone script(s)
    1 passed, 0 skipped, 0 failed (scripts)    rc=0

Verified across three shapes: countdown (1 script), jellyfin-now-playing and
pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files
split 4 collectable / 7 scripts, all seven of which had never run).

Closes #532.

Running the flights scripts for the first time also surfaced four genuinely
failing tests there, hidden by the mirror-image bug in the plugins repo's own
runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465.

* fix(harness): give an empty-looking mode a few frames before warning about it

check_plugin's "drew nothing but display() returned X" warning fired on a single
frame, rendered with force_clear=True, under a frozen clock. All three defeat a
scrolling plugin, whose first frame is legitimately its blank scroll-in buffer.
Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which
people stop reading a warning, which matters because the true positives are
real: a mode that draws nothing and does not return False holds a blank panel
for its whole display duration.

An apparently-empty frame is now re-driven for up to 48 more frames with
force_clear=False (force_clear means "reset the scroll", so repeating it would
redraw frame 1 for ever) and with the clock advancing -- freezegun's factory
where time is frozen, a real sleep where it is not, since scroll position is
usually a function of elapsed time. The first frame that draws content replaces
the result.

The clock is moved back afterwards. It is shared by every render in the matrix,
so time borrowed by the probe leaked into later modes and drifted their goldens
-- f1_upcoming picked up 5 spurious drifts before this was restored.

Measured on the rig:

                        empty warns          check
                        before  after
  f1-scoreboard            42      0    48 PASS / 0 FAIL, goldens intact
  ledmatrix-elections      16      0    16 PASS / 0 FAIL
  on-air                    8      8    true positive, kept
  nfl-draft                 8      8    true positive, kept
  clock-simple/geochron/    0      0    unchanged
  christmas-countdown

58 false positives gone, both true positives kept, no golden regressions. Cost
is confined to modes that really are blank: plugins that draw immediately are
unchanged (clock-simple and tide-display still 2s), while on-air -- eight
deliberately blank modes -- goes to 21s.

Closes #527.

* fix(harness): load nested schema defaults, and merge caller config at leaf level

load_config_defaults read only top-level properties. An object property carries
its defaults on its children, not on itself, so everything nested was dropped --
2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of
565. render_plugin_matrix's comment says the plugin then "behaves like a real
install", which for most of the fleet it did not.

_defaults_from_properties now recurses. merge_config deep-merges the caller's
config onto the result so an override lands at the leaf: a shallow merge would
let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard
every other nhl default, which is the same class of bug being fixed here.

Measured before/after across all 49 installed plugins on the rig: **no render
changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact.
Plugins already fall back to the same values internally via config.get(key,
default), so supplying them explicitly agrees with what they were doing. The
defaults really are arriving now:

  ufc-scoreboard        9 -> 87 defaults
  ledmatrix-flights    51 -> 95
  masters-tournament   10 -> 51
  cricket-scoreboard   22 -> 50
  tide-display         12 -> 18

and hockey-scoreboard, which used to load nhl.enabled=None, now gets
nhl.enabled=True with its full display_modes block.

Caveat worth carrying: the eight plugins with the most nested config
(soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of
the 2,386 dropped defaults, 68%) could not be measured. They import
src.common.sports_shared, which the test rig's core branch predates, so they
fail to load there identically before and after. Re-run this comparison against
a core that has that module before trusting the "nothing changed" result for
them; those are exactly the plugins whose renders should change most.

Closes #531.

* refactor: narrow the exception handlers this branch introduced

Codacy flagged the new code; it passes on other recent PRs, so the finding is
mine. Four of the five broad `except Exception` clauses I added were catching
far more than they needed to, which is the same shape as several bugs this
branch fixes -- hello-world's TypeError sat invisible for exactly this reason.

  freezer() / move_to() / tick()   -> (AttributeError, TypeError, ValueError)
  cache_manager.delete()           -> (OSError, AttributeError, KeyError)

The fifth stays broad and now says why: it wraps a call into a plugin's own
display(), which can raise anything, and the first frame has already rendered --
so a failure there must not turn a good result into an error.

Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty
warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows
7 golden drifts both before and after this branch, so it is not from these
changes -- its committed goldens predate #521's 1-bit text rendering.

* fix: resolve CodeRabbit review and Codacy findings on #534

CodeRabbit raised six; all six were real.

The test double had drifted ahead of production. VisualTestDisplayManager
accepted set_scrolling_state(frame_hold=...) while DisplayManager did not,
so such a call passed every harness run and would raise TypeError on the
panel -- the one failure a safety harness exists to prevent. frame_hold
belongs to the change that adds it to DisplayManager (#523), so it moves
there and the double matches main again.

The harness swallowed exceptions from re-rendered frames. _settle_loop
re-renders a mode that came back blank, to give a scroll time to draw;
returning silently on a crash meant a mode that renders one good frame
and then explodes was reported as passing. Recorded on result.error now,
keeping the captured frame so the failure stays inspectable.

starlark-apps display() returned True after _display_frame() failed, so
the controller held a dead frame for the whole display_duration instead
of rotating on. _display_frame now returns bool on all three paths.

run_plugin_tests.py used env.setdefault for PYTHONPATH and
LEDMATRIX_CORE, so an inherited value won and the subprocess imported a
different core than the one under test -- ledmatrix-plugins#467 exactly.
Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally.

The on-demand mailbox is polled after every frame, ~125x/second on a
scrolling mode, and the read is deliberately uncached, so it was that
many disk reads per second to find nothing. Floored at 250ms, which is
imperceptible for a web-UI click. Consuming it also deleted whatever was
present rather than what had just been processed, so a request posted
while the previous one was in flight was thrown away and never ran; the
delete is now keyed by request_id. That narrows the window rather than
closing it -- a true atomic claim needs a primitive the cache layer does
not offer, and the code says so rather than implying otherwise.

Codacy's 2 criticals were bandit B404/B603 on the subprocess call added
to run_plugin_tests.py. Fixed interpreter, argument list, no shell;
annotated with the repo's existing nosec convention. Bandit is clean on
the file.

Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py
(4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of
those fail against the pre-fix code. Full suite: 3961 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: satisfy Codacy's subprocess checks on the new test runner

Codacy runs Bandit and Opengrep (its Semgrep fork). The new
subprocess.run in scripts/run_plugin_tests.py trips three patterns, on
two different lines:

  Bandit   B404 on the import, B603 on the call
  Opengrep dangerous-subprocess-use-audit          on the run( line
           dangerous-subprocess-use-tainted-env-args on the argv line

A nosemgrep applies only to its own line, so the call line and the argv
line each need one; a single comment on the call covered neither rule
fully. Suppression is the right answer here rather than a rewrite: the
interpreter is sys.executable, the arguments are a list, and no shell is
involved, so there is nothing to word-split or expand.

Matches the pair the rest of the repo already uses for this shape --
permission_utils.py, plugin_loader.py, install_dependencies_apt.py.

Codacy: 0 new issues, up to standards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: leave visual_display_manager untouched so #523 can merge

The only change this branch made to that file was a docstring, and it
collided with #523's rewrite of the same method -- so #534 and #523 each
merged cleanly against main but conflicted with each other. Reverted to
main's text; #523 owns this method and adds frame_hold to it.

The note the docstring carried ('frame_hold arrives in #523') would have
been stale the moment #523 landed anyway. The parity test in #523 is
what actually keeps the two signatures honest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

* chore: add the Ruff suppression nosec/nosemgrep do not cover

Ruff reports S603 on the same call Bandit and Opengrep do, and none of
the three suppressions covers the others. Confirmed the precondition
first: path comes from discover_plugin_tests(), which globs test files
inside the repo, and the call is a fixed interpreter with a list argv
and no shell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 13:37:38 -04:00
Chuck 32d637a446 fix(store): read the core version from disk, not from a stale import (#518)
* fix(store): read the core version from disk, not from a stale import

Updating the core to 3.3.0 and then updating plugins refused all eight sports
scoreboards:

    Refusing to install nrl-scoreboard: NRL Scoreboard supports LEDMatrix
    >=3.3.0, but this system is running 3.2.0.

while src/__init__.py on that machine read 3.3.0. Observed on hardware, not
theorised.

The gate ran `from src import __version__ as core_version`, which binds
whatever the process loaded at start. The plugin store's gate lives in the web
UI, a long-lived service of its own, and the update route deliberately restarts
nothing -- it replaces files on disk and asks the user to restart. Its prompt
named only the *display* service, so a user who followed it left the web
process holding the previous number.

Stale by exactly one release is the case that bites: every plugin flooring on
the release you just installed is refused, blaming a core version that is
already correct on disk. It reads as a broken plugin store. 3.3.0 is the first
release where this hits a whole family at once, since all eight scoreboards
floor there.

compatibility.current_core_version() reads the version from the file instead,
falling back to the imported value on any failure -- so it can only ever be as
correct as before, never worse. All four gate call sites use it: three in
store_manager (install, the git-pull update path, install_from_url) and one in
plugin_loader's advisory warning.

The restart prompt now names both services.

Twelve tests, including the hardware failure itself: a process holding 3.2.0
while disk says 3.3.0 refuses hockey, and reading fresh allows it. The inverse
is asserted too -- a genuinely old core still refuses, so the gate has not
become permissive. One test greps both modules for the old import-bound read;
reintroducing that line fails it, which is what stops this coming back.

Not changed: web_interface/__init__.py also imports __version__, but for
display rather than gating, and the API endpoint already reports a fresh
git describe.

* fix: drop the unused os import

Left over from a first draft that joined paths by hand before this used
pathlib. Flagged by CodeRabbit on #518; confirmed dead -- no os. reference
remains in the module.
2026-09-03 16:06:40 -04:00
ChuckandClaude Opus 5 f90638a9ec feat(web): show available memory in Tools diagnostics (#500)
* feat(web): show available memory in Tools diagnostics

System Diagnostics reported memory as used-percent plus used/total GB.
Neither distinguishes a healthy board from one about to fail, because
page cache counts as used and is reclaimable on demand -- a Pi can read
70% used and be fine, or read the same and be minutes from trouble.

MemAvailable is the kernel's own estimate of what a new allocation can
actually obtain, and it is the number that tracked the failure on a 1GB
Pi 3B+: healthy running sat above 500MB, and the crash came at 73MB. By
that point fork() was failing, so sshd could not spawn a session and
systemd could not respawn the display, while the kernel carried on
answering pings at 0% loss. Used-percent gave no warning at any point on
the way there; available memory fell steadily for hours.

/api/v3/system/status now returns memory_available_mb from
psutil.virtual_memory().available, and Tools renders it as its own tile,
coloured against the thresholds that failure implies: red under 150MB,
amber under 300MB, green above.

The existing memory tile is left alone -- used/total is still what you
want when sizing a workload; this answers the different question of how
much room is left right now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(web): round available memory once, so the tile agrees with itself

The colour was classified from the raw value while the label was rounded,
and the API sends one decimal place. At the boundaries the two disagreed:
149.6 rendered as "150 MB" in red, and 299.6 as "300 MB" in amber -- each
contradicting the threshold its own colour claims to apply ("red under
150MB"). A reader checking the tile against the documented thresholds would
conclude the readout was broken.

Rounding once and using that number for both restores agreement. It moves
those two boundary cases up a band, which does not matter: the thresholds
come from a measured failure at 73MB, so which side of the line a spare
0.4MB falls on carries no information. The tile agreeing with itself does.

Null handling and the thresholds themselves are unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 08:40:28 -04:00
ChuckandClaude Opus 5 a4a55a23fc fix(web): reject non-finite JSON numbers instead of raising (#497)
* fix(web): reject non-finite JSON numbers instead of raising

POST /api/v3/config/dim-schedule with {"dim_brightness": Infinity} answered
500. So did /api/v3/errors/clear with max_age_hours, and /api/v3/config/main
with multiplexing or row_address_type.

json.loads accepts Infinity/-Infinity/NaN by default -- they are not valid
JSON, but Python's parser emits them -- and Flask's get_json passes them
straight through. int(float('inf')) raises OverflowError, which is neither
ValueError nor TypeError, so validation blocks that carefully caught those let
it past and Flask turned it into a 500.

The status code was not the real damage. dim-schedule answered with
CONFIG_SAVE_FAILED and suggested "Check file permissions on config directory"
and "Check available disk space" for what was an invalid number. Every one of
these sites already had a correct 400 response written; they just never
reached it.

NaN already returned 400, because int(nan) raises ValueError. That is why this
only ever showed up for the infinities, and why it survived: the obvious test
case passes.

OverflowError is now caught alongside ValueError/TypeError at the 27 sites in
this file whose try block performs a numeric coercion. An AST sweep confirms
no int()/float() of request-derived data is left outside a block that catches
it.

Verified end to end through Flask's test client rather than by reasoning about
the parser: all four routes returned 500 before and 400 after.

Tests: five Infinity cases (which fail against the previous except tuples),
two NaN cases pinned so narrowing the tuple cannot quietly break them, and a
check that ordinary input is not rejected by the widened guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

* test(web): assert 400 exactly, and prove valid input is accepted

Both review points were right, and the first is the failure mode this file
exists to catch.

Accepting any 4xx meant a 404 would have passed. Renaming one of these routes
would have left the test green while it tested nothing -- the same "looks like
coverage, points somewhere safe" shape that hid the composer injections. Now
asserts exactly 400.

Both infinity signs are exercised for every route. int() raises OverflowError
either way, but only +Infinity was in the original report, and a guard that
special-cased the sign would have passed a one-sided test.

The valid-input test previously asserted "not a 400", which did not show what
it claimed: the mocked save path fails for any input, so that assertion held
whether or not validation had accepted the value. It now gives load_config a
real dict and stubs _save_config_atomic, so the endpoint reaches its success
response and the test can assert 200 -- which only happens if the value passed
validation.

8 of the 11 checks fail with OverflowError removed from the except tuples.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:44:23 -04:00
ChuckandClaude Opus 5 5a1f121e6b Stop array-item secrets being wiped, and logging them (#493)
Three review findings from #485 that I missed when addressing that PR;
it has since merged, so they land here.

1. Array-item secrets destroyed by any unrelated save (data loss).

remove_empty_secrets recursed into dicts but let a list fall through to
the scalar branch and kept it verbatim. Lists merge by *replacement*, so
the blanks the masked form posts back went straight over the stored
array:

    stored   [{"name":"a","token":"REAL-A"}, {"name":"b","token":"REAL-B"}]
    posted   [{"name":"a","token":""},       {"name":"b","token":""}]
    merged   [{"name":"a","token":""},       {"name":"b","token":""}]
             -> both credentials gone

Same failure as the scalar api_key case fixed earlier, one container
deeper. Lists now prune element-wise, and a list with nothing real in it
is dropped so the stored one is left alone. Where one entry does change,
the new merge_secrets merges by index instead of replacing.

Two details the first attempt got wrong, both caught by existing tests:

- An emptied dict item must stay {}, not None. ConfigManager's
  _strip_secrets_recursive treats a secrets list as *parallel* to the
  regular one ({} = "item i has no secrets"); a None makes it stop
  looking parallel, and it then drops the whole key from the main config
  -- silently deleting the items' non-secret fields too.
- The incoming list's length wins. The regular config's list is
  authoritative about how many items exist, so preserving surplus stored
  entries would let the two fall out of step and make deleting an entry
  impossible.

2. Submitted credentials written to the journal (security).

save_plugin_config logged `Full config: {plugin_config}` at INFO and
`Config that failed: {plugin_config}` at ERROR. Both run before
separate_secrets, so plugin_config still held the values just typed into
the form. Now keys only. Swept the rest of web_interface/ and src/ for
the same shape -- these were the only two.

3. Restart banner kept stale wording.

showRestartPending() cleared the stored custom text but left the DOM
element alone, so a config save could show the previous update's
message. The default is read back from the server-rendered copy rather
than duplicated in JS, so the template stays the one owner of the string.

Verified: 556 passed, 1 skipped across the web suite. Mutation-checked --
reverting api_v3 fails the logging guard and the array-merge test;
reverting either half of the secret_helpers change fails the unit tests.
New end-to-end coverage drives the real endpoint, not just the helpers.


Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 11:43:08 -04:00
ChuckandClaude Opus 5 5f29243e87 fix(config): make the device location the default for plugin location fields (#490)
* fix(config): make the device location the default for plugin location fields

A user in Kansas City reported their radar centred on Dallas, TX with
nothing in config.json to explain it.

The radar is the `ledmatrix-weather` plugin's `radar` mode, and it centres
on the same coordinates as every other weather mode: `forecast_data`
lat/lon, geocoded from the plugin's own `location_city` /
`location_state` / `location_country`. Those ship with schema defaults of
Dallas / Texas / US. A user who never opened the weather plugin's config
form therefore has no `location_city` on disk, and `PluginManager` merges
the schema default in at load time — so the whole plugin (not just the
radar) silently runs on Dallas. Radar is just the only mode that draws a
recognisable map and gives the mismatch away.

Meanwhile the device-wide `location` block that General settings writes
was read by nothing at all, despite its own help text promising it was
"used for weather, sunrise/sunset, and other location-based content".

`SchemaManager.generate_default_config()` now substitutes the device
`location` into the three fully-namespaced `location_*` keys before
handing defaults back, so the promise holds:

- Only `location_city` / `location_state` / `location_country` are
  substituted. A bare `state` key is left alone — `ledmatrix-elections`
  uses it for a two-letter code, and rewriting it would break that plugin.
- A value the user saved on the plugin still wins: this replaces the
  schema default, and `merge_with_defaults` puts user config on top.
- The substitution is applied on the way out of the defaults cache rather
  than into it, so changing the device location takes effect immediately.
- No config manager, no `location` block, or an unreadable config all
  fall back to the plugin's own schema defaults.

Every caller benefits: the plugin loader, the config form (which now
pre-fills the user's real city), config save, and reset-to-defaults.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GNLrSZ32FNKpHRaduKEJsg

* docs(web): name the exact plugin keys the device location seeds

Review follow-up. The General settings help text said the device location
was "the default for every plugin that asks for a city", which overstates
what the code does: only the fully-namespaced `location_city` /
`location_state` / `location_country` keys are substituted. A plugin with
a bare `city` key gets nothing — deliberately, since `ledmatrix-elections`
uses `state` for a two-letter code. The tips now name the exact keys.

Worth noting for anyone editing these: `ui.help_tip(...)` takes a
single-quoted Jinja string, so an apostrophe in the tip text has to be
escaped or written around. The wording here avoids them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GNLrSZ32FNKpHRaduKEJsg

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-21 16:22:48 -04:00