mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
9a1f94f7939c256259ce9c2ca46ff0b51d7d4c52
125
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9a1f94f793 |
fix(starlark): blank app locations use the device location, not San Francisco (#617)
* fix(starlark): blank app locations use the device location, not San Francisco A Starlark (Tidbyt) app whose Location field is blank rendered at its author's hard-coded DEFAULT_LOCATION -- usually San Francisco -- even with the device city set under General settings. A user in Charlotte, NC got San Francisco weather and radar with nothing in config.json to explain it. src/device_location.py fills unset location fields at render time (display plugin and the web standalone render): the device city is geocoded once via Open-Meteo, preferring a match in the configured state/country, and cached permanently. A saved location always wins; if the lookup fails the field is dropped so the app uses its own default, and the failure is not retried for 30 minutes. Also fixes the config form: clearing a location omitted the key, and the save merges, so the old value could never be removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(starlark): say what happens when the device location can't be used A blank app Location only renders at the device's city when one is set and the Open-Meteo lookup finds it. With no city, no match, or the geocoder unreachable (retried after 30 minutes), the app gets no location and keeps its author's default. The guide, the config page hint, CONFIG_REFERENCE and the CHANGELOG entry now say so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
61e462c635 |
refactor: remove the skin system and the unused src/base_classes package (#615)
* refactor: remove the skin system Skins never rendered with the current scoreboard plugins: the only hook was SportsCore._render_game in src/base_classes, which no plugin builds on, so the UI and store already treated them as unsupported. The owner decided on 2026-09-23 to remove them outright. Removed src/skin_system/ (runtime, base class, fixtures), skins/, scripts/validate_skin.py and their tests; the store's "type": "skin" installer, uninstaller and hide/refuse filters (the official registry lists no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins. Stored skin/skin_options config values are handled in the next commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(config): drop retired skin/skin_options keys instead of validating them A config.json written while the skin system existed can carry skin and skin_options in any plugin section, and most plugin schemas set additionalProperties: false. They are no longer core plugin properties; RETIRED_PLUGIN_KEYS in schema_manager lists them and drop_retired_plugin_keys removes them (unless the plugin's own schema declares the name) in prepare_plugin_config, which loading, hot reload, GET /plugins/config and both web saves already share, and in validate_config_against_schema for callers that validate a raw section. POST /plugins/config and /config/main also drop them from the stored section they merge into, so they leave config.json on the next save. Tests cover the load path (real PluginManager.load_plugin: no schema warning, not degraded), raw and prepared validation, validate_all_plugin_configs, and the JSON, form and /config/main saves. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: remove the unused src/base_classes package No scoreboard plugin builds on src.base_classes: the nine monorepo scoreboards ship their own sports.py and share code through src/common (docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins imports it. The one import anywhere, baseball-scoreboard's rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing instantiates. Removed the package and the eight test files that only tested it (test_api_extractors, test_data_sources, test_sports_base_characterization, test_sports_capabilities, test_sports_core_promotions, test_sports_logo_cache_bounded, test_sports_modes_promotions, test_sports_odds_fanout). test_common_is_hardware_free no longer lists src.base_classes as a forbidden import, and comments in sports_helpers.py and base_odds_manager.py stop pointing at it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: drop the skin system and src/base_classes from the docs Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was removed and shared code lives in src/common, in the Layering section and the view-model-contract rule. Other docs stop pointing at the removed package. CHANGELOG records both removals under Unreleased. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(store): hide and refuse registry entries that aren't plugins The skin filters went with the skin system, but a custom registry can still list "type": "skin" entries, and installing one as a plugin would unpack it into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing type means plugin) now hides non-plugin entries from the store and custom-registry listings, and install refuses them, in the route with a clear 400 and in _install_plugin_impl for any other caller. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4fe3cdd906 |
fix(starlark): stop the root display service locking the web UI out of starlark-apps (#604)
* fix(starlark): stop the root display service locking the web UI out
Reported after a fresh install: installing an app from the Starlark tab
failed with "install failed: Failed to install from repository", and so did
uploading a .star file and installing from a GitHub directory. The reporter
found the cause only by reading service logs, and fixed it with
sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/starlark-apps
starlark-apps is gitignored, so it is never checked out -- it is created
lazily by whichever process reaches it first. Those processes run as
different users. systemd/ledmatrix.service is User=root and constructs this
plugin at startup, which is where _get_apps_directory() is called from;
systemd/ledmatrix-web.service runs as the login user and is what actually
installs apps.
The documented first step is to install pixlet and reboot, so on a fresh
machine the display service usually wins that race and mkdir() leaves the
directory root-owned. The web process then fails in _install_star_file() on
app_dir.mkdir(), which catches nothing, so PermissionError reaches the
route's outer `except Exception` and becomes the generic message the user
saw. All three install paths write to the same directory, which is why all
three failed.
The web user cannot repair this -- chown needs root. So root does it, on
every startup, which also heals machines already broken by this without the
owner having to find the chown themselves. It is a no-op when not root, when
the platform has no POSIX ownership, and when the checkout genuinely belongs
to root; a chown that fails warns rather than killing startup.
Also made the failure legible if the handover is ever prevented: a
PermissionError now names the directory, the automatic repair, and the
manual chown, instead of a message that names neither path nor cause.
Verified by mutation: dropping the handover call, chowning a genuinely
root-owned checkout, and letting a non-root process chown each fail their
own test. 121 starlark tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(starlark): address the review on the ownership repair
Findings from the automated review of #604.
Symlinks (CWE-59, the serious one). A root chown that follows links is a
privilege-escalation primitive: anyone able to write in starlark-apps could
point a link at a root-owned file and have the repair hand it over. Entries
are now read with os.lstat, symlinks are skipped outright, and the chown
passes follow_symlinks=False. Descendants are processed before the directory
itself, so the container does not change hands while its contents are still
being walked.
install_app() caught PermissionError in its broad handler and returned
False, which both routes report as a generic install failure -- the exact
shape of the bug this PR exists to fix, since the caller could not tell
"this app is broken" from "this process cannot write here". PermissionError
is now re-raised; every other failure still returns False.
The test fixtures skipped on bare Exception, which would have turned a
syntax error or NameError in the plugin into a green run. They now skip only
for a named absent dependency and re-raise anything else.
Also fixed the _Stat stub that failed in CI but passed locally: it carried
only st_uid/st_gid, and pathlib reads st_mode while walking. It now wraps
the real stat result and overrides ownership alone.
NOT taken: the CodeQL "information exposure through an exception" finding on
the hint response. Dropping `details` would contradict this package's
documented rule -- "if it returns 5xx, it says why" -- which
test_no_api_v3_handler_discards_its_exception enforces with an allowance
that may shrink and never grow. The Starlark routes are the ones that policy
was written for: they answered 500 with no detail for three releases.
describe_exception already redacts credentials and truncates. Keeping the
detail is the deliberate trade-off, so the finding is declined rather than
silently worked around.
Verified on hdpi with the updated code: a symlink to /etc/shadow planted in
starlark-apps was skipped while the directory was handed back, and
/etc/shadow stayed root:shadow.
Mutation-checked all three behaviours. The symlink test was vacuous on the
first attempt -- the link already had the target owner, so it was skipped
for the wrong reason and the mutation passed. It now forces the link to look
like it needs handing over, and fails when the check is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4a1fd7464a |
fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)
* fix(errors): serve /api/v3/errors/* from the display service's aggregator The error aggregator is a per-process singleton and only the display service runs plugins, so only its aggregator records anything. The routes read the web process's own, empty one and always reported no errors. The display service now publishes a bounded snapshot of its aggregator to the shared cache (plugin_error_snapshot) from a daemon thread: at most once every 10 s and only when something changed, never raising into the caller. The routes read it and keep their response shapes, adding snapshot_available, generated_at and clear_pending; exception text has credentials redacted. POST /errors/clear writes a clear request (plugin_error_clear_request) that the display applies on its next 5 s tick via the new clear_before(), which keeps errors recorded after the cutoff and rebuilds the counts. Until the snapshot acknowledges the request, reads hide everything before the cutoff, so a snapshot written just before the click cannot bring errors back. Adds "all": true; cleared_count is null when only the display can know it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): show plugin errors in the Logs tab A compact panel under the log viewer: per-plugin error counts, repeating errors (type, count, affected plugins, a sample message, last seen) and a Clear button, with empty states for "no errors" and "display service hasn't reported yet". Polls every 15 s while the tab is active; all text goes through escapeHtml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: describe where plugin error reports come from and how clear works Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(errors): redact the published snapshot before clipping it Keeping only a traceback's tail (or clipping a message) could cut an `api_key=` marker off while keeping the secret after it, and the web side's redaction would then have nothing to match. The display now redacts every free-text field of the snapshot first. The patterns move to a Flask-free src/redaction.py so the display service can use them; redact_text in the web error handler uses the same function, unchanged in behaviour. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a8b3e86775 |
refactor(web): delete dead routes, JS files and duplicate definitions (#609)
* refactor(web): drop validators nothing calls escape_html, validate_image_url, validate_font_awesome_class, validate_mime_type, validate_numeric_range, validate_string_length and sanitize_plugin_config had no callers outside their own tests. Only validate_file_upload (fonts upload) is imported by the web interface. dedup_unique_arrays is kept: its one caller in save_plugin_config was removed by the unrelated sync PR (#330), which looks accidental. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(api): remove the music-auth and of-the-day JSON routes POST /plugins/authenticate/spotify and /plugins/authenticate/ytm had no caller but their tests: the music plugin authenticates through its web_ui_actions (authenticate_spotify.py / authenticate_ytm.py) via /plugins/action. POST /plugins/of-the-day/json/upload and /json/delete looked the plugin up by the id ledmatrix-of-the-day (its manifest id is of-the-day), were reachable only from a file_type "json" upload field that no schema declares, and put the plugin directory on sys.path per request to import scripts.update_config. of-the-day manages its files through plugin-file-manager and its own web_ui_actions. The of-the-day branch of GET /plugins/config stays: it matches the real manifest id and still merges the on-disk category files into the form. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(api): read managers only from the blueprints api_v3/__init__.py and pages_v3.py declared module globals (plugin_store_manager, saved_repositories_manager, schema_manager, operation_queue, plugin_state_manager, operation_history, sync_manager, config_manager, plugin_manager) that nothing assigns: app.py sets the managers as attributes on the Blueprint objects, and every route reads them there. The one reader, backup restore's fallback to the module plugin_store_manager, could only ever fall back to None. _ensure_cache_manager() built a second CacheManager in the web process instead of using the one app.py puts on api_v3. The display routes now read api_v3.cache_manager, creating it on the blueprint only when nothing set it (the same None handling as the /cache routes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(web): drop run.sh and the unused log_config_change web_interface/run.sh was referenced only by web_interface/README.md; the service starts the UI through scripts/utils/start_web_conditionally.py and the README already documents `python3 web_interface/start.py`. log_config_change() in web_interface/logging_config.py was never called. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): delete unreferenced store_manager.js, diff_viewer.js, htmx-sse.js - js/plugins/store_manager.js (window.PluginStoreManager) and js/config/diff_viewer.js (window.ConfigDiffViewer) were loaded on every page but nothing reads either global. - htmx-sse.js (plus its CDN fallback) was loaded after HTMX, but no template or plugin page uses sse-connect / hx-ext="sse": the live streams run through LEDStreams in app-shell.js. js/plugins/state_manager.js stays: install_manager.js's updateAll() reads and refreshes window.PluginStateManager. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove app.js helpers nothing calls - hexToRgb, rgbToHex, validateForm, uploadFont and switchTab (whose 'switch-tab' event had no listener) have no caller in the templates, static JS or the plugin monorepo. - installPlugin: plugins_manager.js (loaded last) assigns window.installPlugin, and its own store cards are the only callers. - The showNotification fallback could never install: app-shell.js is deferred ahead of app.js and defines the same fallback at top level. - performanceMonitor only logged with ?debug=perf and read an unset this.measures; the marks it took on every load had no reader. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): drop app-shell.js refreshPlugin A top-level function in app-shell.js, so a window global, but nothing calls it (no inline handler, no window lookup, no string-built name). The other plugin actions in that block stay. updatePlugin is the live window.updatePlugin: plugins_manager.js only installs its own copy when none exists. uninstallPlugin/pollUninstallOperation, updateAllPlugins, executePluginAction and toggleNestedSection are replaced by later deferred scripts, but a click that lands while those scripts are still downloading reaches the app-shell copies, so removing them is not a pure no-op. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove definitions plugins_manager.js always overrides All of these are replaced before anything can call them, checked against the load order in base.html and the live window.* values: - openOnDemandModal/requestOnDemandStop stubs: the IIFE later in the same script assigns the real functions synchronously. - updatePlugin and uninstallPlugin stubs (`window.X || stub`): app-shell.js already defined both, so the fallback never installed. Same for the later updatePlugin override, gated on the live function containing '[UPDATE]', which app-shell.js's never does. - The first addArrayObjectItem/removeArrayObjectItem: reassigned by the top-level copies after the IIFE. - The first `function formatDate` in the IIFE: a later declaration of the same name in the same scope wins. - deleteUploadedImage, getCurrentImages, showUploadProgress, formatFileSize and getScheduleSummary: character-for-character copies of js/widgets/file-upload.js, which stays the owner. - `typeof X === 'undefined'` fallbacks and `typeof X !== 'undefined'` re-exports after the IIFE: always false, or a self-assignment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): render the shell directly and delete index.html index.html extended base.html with {% block content %}, but base.html defines no blocks, so none of index.html ever rendered: rendering both with jinja2 gives byte-identical output. index() still loaded the config, read config.json and config_secrets.json raw and json.dumps'd them on every page load for variables base.html never reads, and flashed errors that base.html never shows. It now renders base.html with no context. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): stop htmx-config.js replacing console.error and console.warn It swapped both globals for filters that dropped any error mentioning insertBefore / "Cannot read properties of null" when "htmx" appeared in the message or stack, and a list of Permissions-Policy warnings. That hid real errors from every script on the page, and made every logged error and warning report htmx-config.js as its source. The beforeSwap target validation above it, which prevents the insertBefore errors in the first place, stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(web): quiet the widget load announcements and debug logs About 30 lines hit the console on every page load: one "... widget registered" per widget file, one "[WidgetRegistry] Registered widget: X" per registration, plus the registry, base widget and plugin loader announcing themselves. The load-time announcements are removed; the per-call ones (registry register, plugin widget loads, "Render called") now go through the page's debugLog switch (localStorage.pluginDebug), guarded because the widgets also load in node tests without it. fonts.html and wifi.html debug logging goes through debugLog as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(api): drop the removed music-auth and of-the-day JSON routes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
342e9164b8 |
fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound Each repeat of an error pattern appended every plugin in the time window to the pattern's list again, so a plugin failing in a loop grew the display process's memory without limit: 3,000 errors from three plugins reached 2.5 million entries. Keep the list unique. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): load a BDF font at its native size instead of PIL's default FreeType rejects any size but a BDF strike's own, and FontManager answered that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf requested at 8 or 10px rendered as PIL's default font. Retry at the native strike, as element_style already does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin toggle failures no longer claim "operation in progress" Every exception in POST /plugins/toggle was mapped to PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation is already in progress". Report the failure as what it is, and record the plugin id in the operation history for form posts too (it read a `data` variable that only the JSON path set). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): route plugin card clicks through handlePluginAction The document-level delegation checked `typeof handlePluginAction`, which is scoped inside the plugin-manager IIFE and so never visible to it. Every card click took a copied fallback that stopped propagation (the grid's own listener never ran), confirmed an uninstall twice, and sent Starlark app uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>. Expose the handler on window and delegate to it. Also run every test/js/unit suite under pytest: they need only node, but CI ran one of the eight. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): apply Rotation durations, WiFi messages and Vegas settings Three settings the web UI saves never reached the display: - Rotation & Durations: display.display_durations was never read. Every plugin inherits get_display_duration() and the plugin was asked first. A saved value now wins. The page shows unsaved screens blank with the plugin's own duration as a placeholder, and saving a blank removes the override, so one save no longer pins every screen. - WiFi status overlay: the controller looked for wifi_status.json one directory above the repo. Both sides now use wifi_manager.get_wifi_status_path(). The message is written by rename so the display never reads it half-written, and the resumed plugin redraws the whole panel afterwards. - Vegas: nothing called coordinator.update_config(), so saved Vegas settings never reached a running scroll. They are now queued when display.vegas_scroll changes, and applied while Vegas is stopped too, so a disable then re-enable works. The follower's scroll-speed default (75) now matches VegasModeConfig's (50). Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows with each plugin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep affected_plugins order when serialized; guard non-Element targets ErrorPattern.to_dict() ran the now-ordered list through set(), so get_error_summary() listed plugins in an unstable order. The document-level card-action listener called event.target.closest() without checking the target is an Element. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
19686ab761 |
fix(web): plugin settings form shows schema defaults for unsaved keys (#597)
The server-rendered plugin settings partial rendered straight from the saved config, so an option added in a plugin update (geochron 1.2.0's show_date / show_date_line, default true) drew as an unchecked box, and the save route's missing-checkbox handling then stored it as false. Enum dropdowns likewise showed their first option instead of the default. - _load_plugin_config_partial runs the stored section through prepare_plugin_config (as GET /plugins/config does) before masking secrets, so a secret's schema default is masked too. - render_field falls back to the field's own default, covering children of objects that declare a default of their own (where the defaults extraction stops). - The legacy-boolean parity test now compares against the config the plugin actually runs with (defaults included). Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
116abb0daa |
fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service BackgroundDataService always sent a season range first and, on a 400, fell back to chunks without recording the rejection, so every background season fetch spent a doomed request and live scoreboards learned nothing from it (or it from them). The worker now consults and sets the same 6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight to month/day chunks, and if every chunk fails the range is asked once for a real error without re-spending the chunks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep plugin asset and action routes inside their directories POST /plugins/assets/upload, GET /plugins/assets/list and POST /plugins/assets/delete joined the request's plugin_id onto assets/plugins unchecked, so '../../config' created, wrote, listed and deleted outside it. #561 guarded only the route that serves the files. All three now go through path_safety.resolve_under and answer 400 for anything but a plain name, and delete only unlinks a metadata path that resolves into that plugin's uploads directory. PluginManager.get_plugin_directory refuses ids that are not one plain path segment, so /plugins/action (which runs a manifest script from the returned directory) and every other caller get the guard; the action route also rejects such ids up front, covering its no-manager fallback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): report a no-op plugin update as already up to date update_plugin() returns True both for a real update and for "nothing to do" (a ZIP-installed monorepo plugin already at the registry version, a bundled plugin). With no git commit to compare, POST /plugins/update called every such success "updated successfully", so Check & Update All counted most official plugins as updated on every run. The route now reads what changed off the plugin itself (commit, else manifest version, else last_updated) and returns data.update_status (updated / up_to_date / local_only). The update-all toast is summarised by PluginInstallManager.summarizeUpdateResults from that status, falling back to the message for older servers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): scoreboard scroll speed no longer follows target_fps sports_scroll computed the crisp speed ladder against the global target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since frame-locked presentation (#545) the helper steps a fixed number of whole pixels per presented frame and the panel presents at its real refresh, so the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a 50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s. The ladder now uses the display manager's refresh_hz, then display.hardware.limit_refresh_rate_hz, then the default. target_fps is not consulted. Docstrings now say scroll_delay is ignored for pacing (no behaviour change there) and describe the fixed-step model. Tests: replace the tests that pinned target_fps as the ladder refresh and described time-based stepping; assert speed independence from target_fps (unit and end-to-end presented px/s against the real helper), that the fixed per-frame step is applied, and that scroll_delay does not change speed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): escape registry and upload values in plugin manager inline handlers The store, saved-repository and custom-registry buttons built onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone, so a custom registry entry whose id contained ' closed the attribute and added its own handler. One helper, jsStringAttr(), now HTML-escapes the JSON literal for every one of those handlers, and the store View button opens only http(s) repo links. The live window.updateImageList (plugins_manager.js loads last, so its copy wins over the file-upload widget's) wrote the uploaded file's original name, path and ids into markup raw; they are escaped now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note plugin asset, action and inline handler guards Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): let the root pip wrapper install web_interface/requirements.txt Update Code, the automatic update's health check and Install Base Requirements install web_interface/requirements.txt through safe_pip_install.sh, which only allowed the root requirements.txt. The first commit changing that file would fail its dependency install, and the automatic updater rolls back any update whose dependencies did not install -- on every device, for every newer commit. The wrapper now lists both core requirement files. Only their folders are resolved, so a requirements.txt symlinked out of the project is compared by its target and refused (previously the root file's own symlink target was what got allowed). The updater's file list is a named constant, and a test runs the real wrapper (pip stubbed) on every file Update Code and the rollback install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): do not retry plugin requests that got an HTTP answer PluginAPI.request wrapped everything that was not a structured error as NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a JSON error without error_code included. Check & Update All retries NETWORK_ERROR, so those updates were re-sent five more times with backoff, contrary to the #587 contract that an HTTP error response is the server's answer. NETWORK_ERROR now means only that fetch() rejected. Any HTTP response without an error_code, or with a body that is not JSON, is API_ERROR with the HTTP status attached. Tested against the shipped api_client.js. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): restart the stats window when an idle gap is dropped by size #582 dropped an idle gap from the frame stats two ways: the reset_scroll() sentinel, which also restarts the 5s window timer, and a size guard for scrollers that never call reset_scroll(), which did not. On that path the first real frame after the gap found the boundary overdue and logged a stats line for a one-frame window. Both paths now share one seeding helper. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): leave plugins alone when update_core's own rollback fails update_core returns rollback_failed directly when a partial pull or an update whose health check never started cannot be rolled back. run() only held plugins back for 'verifying', so those devices still got new plugin versions and a display restart on top of a core in an unknown state -- the opposite of what the health-check path does, and of the 3.4.0 changelog (plugins are left alone if the rollback fails). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(api): make the REST reference match the api_v3 package Every documented request body, query parameter and response shape was re-checked against the handlers in web_interface/blueprints/api_v3/. Fixes calls that failed as documented (repo_url, action_id/params, files/image_id, font_file+font_family, ?font=, cache key, auto_enable_ap_mode, plugin limit keys), removes the font-override endpoints dropped in #566, corrects response shapes (plugins/config, plugins/schema, health, metrics, operation history, github-status, fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds the 26 routes it omitted (backup, system auto-update/git, wifi radio, starlark editor, MQTT bridge, status endpoints, skins). Documents the merge semantics of partial JSON saves to /config/main and /plugins/config and the dim-schedule POST accepting GET's days shape, which land in the same change set. Replaces app.py line numbers and the removed api_v3.py path with file and function names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): remove the General-tab plugin system toggles that did nothing plugin_system.auto_discover, auto_load_enabled and development_mode had General-tab toggles whose help tips promised dormant plugins and verbose logging, but nothing reads them: every enabled plugin is discovered and loaded regardless. Remove the three toggles. The keys stay tolerated in stored configs. The save handler now stores a flag only when a client sends it; treating a missing key as an unchecked box would otherwise rewrite all three to false on every General-tab save, which still posts plugins_directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(scroll): remove dead code left by #523/#570 - Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read them since the numpy blend replaced the scipy path. - Drop ScrollHelper._last_integer_position and frame_time_target, which were written but never read. - Keep target_fps and set_target_fps() but document them as informational: nothing paces off them, yet ledmatrix-elections' test_scroll_pacing.py reads helper.target_fps back and third-party plugins may call the setter. - Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to throttle", set_sub_pixel_scrolling's "default: True", and set_frame_based_scrolling's claim that it steps. The plugins monorepo was grepped for every removed name; none is used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI FONT_MANAGER.md told plugins to read display_manager.font_manager, which does not exist, so a plugin following it failed to load with AttributeError. The shared FontManager lives on the PluginManager and BasePlugin._get_font_manager() returns it (with a fallback for harnesses). Also removes the Fonts-tab override workflow and element-override panels that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls /plugins/store/search does not exist (404) and the list endpoint reads query, not q. The registry guide's curl examples omitted the JSON Content-Type, so the handlers saw an empty body and answered 400. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(config): use the shared core-key list in the last three private copies StartupValidator warned "Plugin 'auto_update' is enabled but not found" on every display start with auto-update or a dim schedule on; the reserved plugin-id check missed auto_update, sync, location and the rest; and ConfigManager's (uncalled) orphan cleanup would have deleted display, schedule and auto_update. All three now read src/core_config_keys.py, which also gains CORE_SECRETS_KEYS for the github/youtube secrets sections. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): partial JSON saves to /config/main change only what they send A JSON body with one field reset every checkbox in the sections it touched: the MQTT bridge's brightness slider turned off disable_hardware_pulsing, inverse_colors, show_refresh_rate and use_short_date_format, and a timezone-only save turned off web-UI autostart and weekly auto-updates. Missing-means-unchecked now applies only to form posts: form-encoded bodies and the v3 forms, which mark themselves with a hidden __form_section input. Also on the config routes: - vegas_min/max_cycle_duration no longer match the generic *_duration rule, so they stop landing in display_durations and a blank one no longer rejects the whole Display save; - saving from the Raw JSON editor calls start_setup_if_needed like the General form, so enabling auto-update there finishes its setup; - the schedule and dim-schedule POSTs accept the per-day days.<day> shape their GETs return, as well as the flat form keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): install plugin dependencies from the configured plugins directory install_plugin_dependencies.sh scanned only plugins/, but the Plugin Store installs into plugin_system.plugins_directory (default plugin-repos), so the documented "Recommended" fix found 0 plugins on every store install. It now reads plugins_directory from config/config.json (relative to the project root or absolute, default plugin-repos) and also scans plugins/ for dev symlinks, installing a plugin reached through both only once. With set -e alone, `pip ... | tee` took tee's exit status, so a failed pip install was reported as success; set -o pipefail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: replace stale API names, line numbers and the api_v3.py path - ADVANCED_FEATURES: StreamManager methods that exist (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the real on-demand status envelope ({status, data: {state, service}}) - app.py:199 / :144 / :607-619 line citations and web_interface/blueprints/api_v3.py (now a package) replaced with file and function names in ADVANCED_FEATURES, CONFIG_DEBUGGING, PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE, PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README - CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use /config/raw/main to replace the file; describe where validation runs - TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints usage) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): verify the web interface that actually ships, on port 5000 verify_installation.sh failed every healthy install: it required the long-removed web_interface_v2.py and looked for a listener on port 5001, while the web interface binds 5000 (web_interface/start.py). It now checks the files ledmatrix-web.service runs (start_web_conditionally.py, web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the same 5001 port in its listen check, HTTP probe and printed URLs. Port matches are anchored so :50001 no longer counts as :5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(plugins): one display-size contract: display_manager.width/height CLAUDE.md (#580) says to read display_manager.width/height because matrix is None when hardware init fails; the development guide, the safety-harness doc and two DisplayManager docstrings still recommended matrix.width/height. The bundled starlark-apps plugin read matrix.width unguarded, so its magnify recommendation and frame scaling raised in fallback mode (e.g. after the Pi 5 hardware refusal). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(install): make install_service.sh --help print usage instead of installing install_service.sh parsed no arguments, so `sudo ./scripts/install/ install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md) rewrote ledmatrix.service, ledmatrix-web.service and both update-verify units and enabled/started them. It now handles -h/--help (usage, exit 0, no changes) and rejects any other argument with exit 2 before doing anything. Running it with no arguments, as first_time_install.sh does, is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scroll): describe the fixed-step model and document frame_hold Since #545 a crisp speed from scroll_config.configure() makes the helper advance a fixed whole-pixel step per presented frame with no clock, and the display manager's frame hold is part of the speed. The docs still described the removed wall-clock model: - scroll_config's module and configure() docstrings said speed is applied in time-based mode and that omitting the hold "falls back to fractional pixels"; omitting it actually runs the scroll frame_hold times too fast. - SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both modes, and read a 20 ms stats median as missed refreshes although that is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step, the hold-dependent healthy median, that target_fps plays no part, and that a hand-added scroll_pixels_per_second loses to a schema-default pair. - PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling) without frame_hold; it now documents the parameter (core 3.4.0) with a configure() + set_scrolling_state example. - update_scroll_position/set_scroll_speed and set_scrolling_state docstrings say the same. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark target_fps legacy; describe what Vegas scroll_delay does - General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after the sports_scroll fix nothing in core scrolling reads it. The field and its API validation stay so saved configs and plugins that read global_config['target_fps'] keep working. CONFIG_REFERENCE says the same. - Vegas frame_based_scrolling/scroll_delay were described as frame-count stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and still advances by elapsed time, so the applied speed is clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The config comments, render_pipeline comment and CONFIG_REFERENCE rows now say so. No behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(deps): describe how plugin dependencies are really installed The guides said the web service runs as root, that installs pick --user from os.geteuid(), and quoted a warning and a PluginManager._install_plugin_dependencies() method that don't exist. The web unit runs as the installing user; store installs go through install_requirements_file() and sudo safe_pip_install.sh (root), with a user-level fallback that says so, and load-time installs run in the display service's own (root) interpreter. Manual paths now use the configured plugins directory (plugin-repos/ by default) instead of plugins/, which store installs no longer use, and install_plugin_dependencies.sh is described as scanning that directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): count local changes one way for the preflight and the pull The automatic update's preflight ignored mode-only changes and anything whose status line contained plugins/ or plugin-repos/, then promised "Automatic updates will not stash your changes". perform_core_update used plain git status (modes count) and ignored only 'plugins/', then ran 'git stash push -- :!plugins', which nothing ever pops. So an edit to a bundled plugin under plugin-repos/, or the installer's chmods on tracked scripts, passed the preflight and was stashed away for good. - auto_update.local_changes() is the one predicate both use: core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/ excluded by leading folder rather than substring (a core file under web_interface/static/v3/js/plugins/ now counts). - Update Code's explicit stash leaves out both plugin folders; the pull's --autostash carries their edits and mode changes across and reapplies them. - The automatic updater calls perform_core_update(stash_local_changes= False), which refuses instead of stashing edits that appeared after the preflight; update_core reports that as 'blocked'. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): diagnostics follow the web autostart default and api_v3 package #556 made a missing web_display_autostart mean "start" (only an explicit false/off keeps the web interface down), but the diagnostics still said otherwise: diagnose_web_ui.sh reported a missing key as "defaults to false", diagnose_web_interface.sh said the web interface "will not start unless this is set to true" and recommended enabling it, and debug_web_manual.py printed False. Troubleshooting a down web UI pointed users at a non-cause. Both shell scripts now evaluate the setting with the launcher's own autostart_enabled() (inline fallback if it cannot be imported) and report on / off / not set (on) / unparseable config; debug_web_manual.py uses the same function. They also check web_interface/blueprints/api_v3/ __init__.py: api_v3.py became a package in #553, so every healthy checkout was reported as missing a file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(install): what install_service.sh installs; verify script port; no sudo for --help install_service.sh installs and starts ledmatrix, ledmatrix-web and the update-verify units, not only ledmatrix.service (systemd/README.md, README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help' as a harmless check; it now shows --help without sudo and warns what a real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh checks the web interface on port 5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note update-all, plugin system settings and script fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): size the preview after orientation and pixel mappers display_geometry.physical_size claimed to give DisplayManager's answer but only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured after the library's pixel mappers, so a Rotate:90 / orientation 90 chain previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for 128x64. Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown names leave it alone), and move the orientation composition here so DisplayManager and the preview share it. The module docstring no longer claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH. Audit finding F18. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): refuse settings the rgbmatrix library aborts on, on every board The library answers several settings with a NULL matrix or abort() rather than an error, so the display service crash-looped (Restart=on-failure) instead of reaching fallback mode: rows above 64, chain_length above 255 (uint8_t binding setter, documented as "no upper limit"), a misspelled hardware_mapping, and parallel 2-3 on a single-output mapping, reachable from the Display form on the default adafruit-hat(-pwm) mapping. #586 only guarded the Pi 5 subset. - src/matrix_support.py holds the rules for every board (Options::Validate ranges, binding integer types, mapping names and outputs from lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the API's numeric ranges. - DisplayManager checks them before building options and raises MatrixSettingsRefused, so a hand-edited config falls back with a logged, reported reason. Emulator mode only warns. - The config API refuses them with a 400 naming the setting; combinations are checked against stored values but reported only when the request sets a field involved. - The hardware status file gains "cause" (settings/library/forced). The fallback log and Display banner give the Pi 5 rebuild hint only for a library failure instead of rebuild + gpio_slowdown advice for every failure; one Pi 5 slowdown recommendation (1-3, start at 1). - The Display form offers classic/classic-pi1 and orientation 90/270 and renders any other stored mapping selected with a warning, so an unrelated save no longer rewrites them; the API accepts 90/270. Audit findings F03, F16, F19, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(display): library limits, template defaults and Pi 5 slowdown - rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs, classic/classic-pi1 mappings and orientation 90/270 documented. - Defaults are the config.template.json values: config migration adds missing keys from the template, so the listed "code defaults" never applied. - One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode, starting at 1. - Troubleshooting describes the refused-settings fallback, and CHANGELOG corrects the Unreleased "no upper limit" entry. Audit findings F19, F20, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py opens the panel with the service's options --measure and --demo built RGBMatrixOptions from a private copy of the display service's builder that had drifted: gpio_slowdown came from display.hardware (default 2) instead of display.runtime (default 3), and rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors, pixel_mapper_config and orientation were skipped, with different defaults (hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown was measured -- or garbled -- in a setup the service never drives. The option filling in DisplayManager._setup_matrix moves, unchanged, into DisplayManager.apply_matrix_options(options, config), which _setup_matrix calls and the script reuses (overriding only limit_refresh_rate_hz for --measure). The script now loads the whole config rather than the hardware block. Tests pin the script's options to the service's attribute for attribute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py recommends keys the resolver honours The ladder ended by telling users to set display_options.scroll_pixels_per_second. scroll_config ranks that key below the scroll_speed + scroll_delay pair, deliberately, and several plugin schemas default the pair into config, so the advised key was silently ignored (a schema-default 1/0.02 pair plus an advised 66 still resolved to 50 px/s). The advice is now the pair that selects the crisp speed exactly (pixels_per_frame every frame_hold/refresh seconds), explains that the pair outranks scroll_pixels_per_second, and gives the scoreboards' per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the printed pair over a schema-default pair and check it lands on the advertised speed and hold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula - SPORTS_UNIFICATION.md still presented honouring global target_fps as sports_scroll's added behaviour and its one user-visible gain; note that it was withdrawn because it had become a speed multiplier. - ADVANCED_FEATURES.md gave Vegas scrolling as (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s by elapsed time, through a 0.1-5 px per scroll_delay clamp when frame_based_scrolling is on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): scroll model fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(dev): link-github links plugins from the ledmatrix-plugins monorepo link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git, and those per-plugin repositories no longer exist: official plugins are directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the monorepo once into the dev directory, finds plugins/<name>, plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and links it under its manifest id. With an explicit repo URL it still links a single-repository plugin as before. dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a fork), plus plugins_repo and plugins_branch; github_pattern, which was documented but never read, is dropped and warned about. Ships dev_plugins.json.example and git-ignores dev_plugins.json, both of which the guide promised. Reading JSON falls back to python3 when jq is missing (get_plugin_id silently returned nothing without jq). update/status/list find the git checkout above a monorepo plugin directory (its .git is not in the plugin dir), and update pulls a shared checkout once. status no longer exits 1 when nothing is broken. Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration, workflow, store integration, hello-world link, submission), and the nonexistent scripts/git-hooks/pre-push-plugin-version and scripts/bump_plugin_version.py replaced with the real rule: bump the manifest version and run update_registry.py. scripts/dev/README.md and CLAUDE.md updated to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scripts): monorepo workspace layout; fix_perms and install READMEs MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin; setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the workspace file opens LEDMatrix plus ../ledmatrix-plugins. scripts/fix_perms/README.md listed cache directories fix_cache_permissions.sh never touches and a 'ledmatrix' service user that doesn't exist (also in scripts/install/README.md); adds safe_pip_install.sh. install/README: install_service.sh installs the web and update-verify units too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): keep the rollback's pip retries inside the unit time limit The health check reinstalled the previous requirements by trying the next bash path after any failure, including a 600 s pip timeout. Two files, two paths: up to 40 minutes of pip alone, while systemd stops ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the rollback half-way and leaving the update 'verifying' until the web UI calls it lost. - Like permission_utils.install_requirements_file, only a sudo refusal moves on to the next bash; a pip that ran and failed or timed out is not repeated. The refusal wording is one list (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only verifier and pinned equal by a test. - All reinstalls in one rollback share a 600 s budget. - WORST_CASE_SECONDS adds up every timeout on the longest path (27.5 min); a test holds it under the unit's TimeoutStartSec and that under the web UI's VERIFY_LOST_SECONDS. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools Plugin config was prepared differently depending on how it arrived: - JSON POST /plugins/config built a partial body on schema defaults, so {"enabled": true} reset every other setting of the plugin. It now merges onto the stored section first, as the form path already did. - Legacy-boolean normalization (#588) ran only at load: GET /plugins/config returned the raw boolean, posting it back failed validation, and hot reload handed plugins the raw section (a legacy dynamic_duration: true came back as a boolean). schema_manager.prepare_plugin_config (normalize, then defaults) is now used by PluginManager.load_plugin, both save paths, GET, the save notifications and DisplayController's hot-reload callback. - The JSON save's filter kept only enabled/display_duration/live_priority and dropped a submitted skin, skin_options or vegas_* tuning key. There is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES, used by validation and by the save filter; PluginManager's CORE_OWNED_CONFIG_KEYS is its vegas subset. - Plugin sections posted to /config/main were stored verbatim, including values /plugins/config rejects. They now go through the same preparation (_prepare_plugin_config_for_save, extracted from save_plugin_config), and a failing section rejects the whole save before anything is written. - dev_server read only top-level defaults and let a schema enabled:false win; build_full_config shallow-merged overrides, dropping sibling defaults; the harness extracted defaults differently from the device. loading.build_config now uses the device's extraction and preparation, and dev_server, check_plugin, render_plugin and the harness all use it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt-bridge): brightness changes apply live and touch nothing else The display service's hot reload applies a saved brightness within a few seconds, and /config/main no longer resets other display settings on a brightness-only JSON body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): automatic update hardening Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI It described web_interface_v2.py and index_v2.html (both gone), client-side form generation, one POST per field with {key, value}, and 'no nested objects'. The v3 UI renders plugin forms server-side from the schema (pages_v3 partial + plugin_config.html macros, nested sections and x-widgets), posts the whole form once, and save_plugin_config() merges onto the stored section, validates, splits x-secret fields and notifies the plugin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt): brightness saves apply via hot reload and leave other settings alone The bridge README said brightness is applied on the display's next restart; the display controller's config hot reload applies it within seconds. It also now states that the bridge's partial JSON save changes only brightness (the /config/main merge fix in this change set). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): don't log pip's output from the health check's reinstall pip can echo a private index URL with embedded credentials; permission_utils redacts it, the stdlib-only verifier cannot, so it logs the exit code only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark the plugin_system toggles as unused legacy keys auto_discover, auto_load_enabled and development_mode are read by nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and the REST reference listed them as live settings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): docs and developer tools group Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): legacy plugin-system toggles no longer count as a General save auto_discover, auto_load_enabled and development_mode have left the General form, so a post carrying only one of them is not a general-settings save and must not treat web_display_autostart and auto_update as unchecked. The plugin_system block itself is left as on main for the branch that reworks it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): config-save and plugin-config preparation fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(claude): re-check matrix_support.py rules when the library submodule is bumped Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address Codacy findings on the core audit PR - plugin_manager.prepare_plugin_config: when the fallback legacy-boolean pass also fails, log a warning instead of a bare except/pass. - api_client.js: request() refuses any endpoint that is not a plain path under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace, control characters) with INVALID_ENDPOINT before calling fetch(), and plugin ids are URL-encoded wherever they are put into a URL (also in the app-shell batch load). - test_update_all.js: pins both against the shipped client. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): check endpoint control characters without a control-character regex Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags the \x00-\x1f range in checkEndpoint's regex. Test the char codes instead; the endpoints refused are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(auto-update): make the seed script executable on disk, not only in the index On Linux Repo.publish() commits with -a, which recorded scripts/run.sh as 100644 upstream because the seed file was never chmod +x. The pull then brought in the same mode the installer chmod had made locally, so installer_chmod saw no mode change left to check. The updater was fine: with the upstream commit at 100755 the --autostash carries the device's chmod across. Verified under Linux (WSL, git 2.43): the old helper fails exactly as CI did, the fixed one passes all 63 tests in the file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7e5967e160 |
fix(web): find installed plugins before anything has discovered them (#594)
The web process discovers plugins lazily: plugin_manifests is empty until some endpoint calls discover_plugins(). Three routes consulted it without discovering, so they misbehaved for as long as nothing else had run -- which, after every ledmatrix-web restart, is until someone opens the dashboard: - POST /display/on-demand/start answered 404 "Plugin <id> not found" (or "Mode <mode> not found"). Measured on a rig: 404 for over three minutes after a web restart, until GET /plugins/installed ran. The browser UI loads the plugin list first, so API-only callers (the Home Assistant MQTT bridge, scripts) are the ones who hit it. - POST /plugins/toggle answered 404 "Plugin not found". - POST /config/main did not recognise a plugin section, so it skipped secret separation and merged the section as-is: the plugin's API key was written to config.json in plain text instead of config_secrets.json. Add _discovered_plugin_manifests(), which discovers when nothing has been yet, and rescans once when a specific plugin id (or, for on-demand by mode, a mode) is not found, so a plugin installed since the last scan is found too. _installed_plugin_ids() now uses it. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f475038895 |
fix(cache): web UI can read what the display service caches again (#593)
* fix(cache): web UI can read what the display service caches again ledmatrix-web.service carried CacheDirectory=ledmatrix. With User= set to the installing user, systemd re-owns /var/cache/ledmatrix and everything in it to that user and its primary group whenever the directory's owner differs -- for a directory root created, on the first start. That erased the root:ledmatrix setgid layout the installers set up, so every file the display service (root) wrote afterwards was root:root 0660 and unreadable by the web interface: WARNING - Permission denied loading cache for display_current_state ... Since #547 install_service.sh renders the web unit from the template, so every fresh install hit this. Measured on one rig: 392 unreadable files, and the web UI's display status, on-demand state and plugin health empty. Existing installs only receive `git pull`, never a reinstalled unit, so the fix for them is in the code the root display service runs: - DiskCache.set gives each file the directory's group (when the directory is group-writable) and 0660 on the open descriptor before the rename, independent of setgid. This also closes a window where a fresh file was visible as mkstemp's 0600. - DiskCache.share_existing_files repairs files an older version left behind, once per process from the cleanup thread. It works through O_NOFOLLOW descriptors and skips hard links and other users' files: the directory is writable by the web user, and root must not be steered into changing a file outside it. For new installs, the web unit drops CacheDirectory=/CacheDirectoryMode=, and install_web_service.sh stops replacing an existing directory's ledmatrix group with the user's group. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): on-demand and current-display status read the display's latest state Found testing the cache-permission fix on a rig: once the web interface could read display_on_demand_state at all, /display/on-demand/status kept answering "active" for over 100 seconds while the file on disk said "idle". Both status routes read the display service's keys through the web process's memory tier, which serves the first copy it read for the full max_age (120s). Read them with memory_ttl=0, as every other cross-process reader (plugin health/metrics, the on-demand mailbox) already does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(install): re-group the cache dir whenever the web user is outside its group install_web_service.sh replaced an existing cache directory's group only when it was root's. A directory in any other group the web user is not a member of -- root:ledmatrix, for a user who is not in ledmatrix -- was left alone, and every file root wrote there stayed unreadable to the web interface. Replace the group whenever the installing user is not in it. A directory whose group the user is already in (ledmatrix, or the user's own group where CacheDirectory= left it) is still left as it is: re-grouping a working directory strands the files already in it on the old group. When the group does change and root-owned JSON files carrying the old group are present, try-restart ledmatrix.service so DiskCache.share_existing_files re-groups them through its symlink- and hard-link-safe path, rather than a recursive chgrp. Verified under WSL's systemd for seven directory states (user group, ledmatrix member, ledmatrix non-member with and without root files, root:root, missing, unnamed gid); the previous version left the non-member case unchanged. Addresses CodeRabbit review on #593. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f9b3d6ae52 |
fix(web): accept every panel size and row address type the rgbmatrix library does (#586)
* fix(web): accept every panel size and row address type the rgbmatrix library does The Display form capped columns at 128 and chain length at 24, and its submit handler (fixInvalidNumberInputs) rewrote anything larger to the cap, so wide panels and long chains silently saved as the wrong size. The config API checked none of the hardware numbers, so values the library rejects (odd rows, parallel 4, PWM dither bits 3) saved and the matrix then refused to start. - Form limits now match the pinned library: rows even 8-64, cols >= 16 and chain_length >= 1 with no upper bound, parallel 1-3, PWM dither bits 0-2, PWM LSB nanoseconds 50-3000. - save_main_config rejects out-of-range rows, cols, chain_length, parallel, brightness, scan_mode, pwm_bits, pwm_dither_bits, pwm_lsb_nanoseconds and gpio_slowdown with a 400. - A stored gpio_slowdown or pwm_dither_bits of 0 renders as 0 instead of the default, so saving the tab no longer overwrites it. - Row Address Type offers 5 (SM5368 / B707 row shift register). Verified on a Waveshare 96x48 V2 (24S-A1) on a Pi 4 with the Adafruit Triple LED Matrix Bonnet: rows 48, cols 96, row address type 5, BGR, GPIO slowdown 8. - Help text and docs: FM6124-family panels use Panel Type Standard; on a Pi 5 the library supports only row address types 0 and 2. No change to the rpi-rgb-led-matrix submodule. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): drop the rows cap and document every display setting accurately Rows: no upper limit in the form or the API. Still even and at least 8. The current rgbmatrix library rejects more than 64 per panel, so a larger value saves but the matrix won't start; the help tip, README, config reference and troubleshooting section all say so, and nothing here needs changing if the library lifts the limit. limit_refresh_rate_hz: the form accepts 0 (the library's "no cap"), a stored 0 no longer renders and re-saves as 120, and the API rejects negatives. pwm_dither_bits stays 0-2: the library rejects 3 and 4, so the old form's 0-4 only ever let users save a config the display couldn't start with. Docs and help tips, checked against the pinned library and its README: - panel_type and rp1_rio get README entries - show_refresh_rate prints to stdout; it never drew on the panel - dither bits raise the refresh rate; the tip said they lowered it - scan_mode is about interlacing at low refresh, not wrong colours - disable_hardware_pulsing: hardware pulsing needs OE on GPIO 18 and the onboard sound driver off; software timing makes rows flash brighter - gpio_slowdown guidance agrees between the README and the UI - all 22 multiplexing values listed; every numeric setting states its range - troubleshooting for a blank panel after a settings change, jumping rows and brightness flashes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): reject true and 5.5 for row_address_type and multiplexing Both still went straight through int(), so a JSON true saved as 1 and 5.5 as 5. They now use the shared hardware range check like the other panel fields. Review feedback on #586. Also: the RP1 Backend tooltip said it is ignored on Pi 3/4 (it is ignored on every model but the Pi 5), and the README gave the dynamic-duration default cap as 90s; the code default is 180s. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: refuse matrix settings a Raspberry Pi 5 can't drive On a Pi 5 the pinned rgbmatrix library drives the panel through the RP1 chip, and that path supports only row address types 0 and 2, parallel 1-3 and the regular / regular-pi1 / classic / adafruit-hat(-pwm) mappings (Rp1PioConfigSupported in lib/rp1/rp1_pio_backend.cc). For anything else CreateFromOptions returns NULL; the Python binding doesn't check, so the display process crashed on its first call into the matrix and systemd restarted it into the same crash every 10 seconds. - src/pi5_matrix_support.py: the rule and Pi 5 detection, matching the library's /proc/device-tree/model check - DisplayManager raises before creating the matrix, so it is a logged init failure (reported by /api/v3/hardware/status) and fallback mode - the config API rejects those settings on a Pi 5 when a request sets row_address_type, parallel or hardware_mapping - the Display form offers only row address types 0 and 2 on a Pi 5, and warns when a stored value can't be used - CLAUDE.md: re-check the rule whenever the submodule is bumped Review feedback on #586. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9616a5a054 |
fix(web): auto_update and other core settings are not orphaned plugins (#589)
v3.4.0 shows "Plugin Config Warning - In config but not installed: auto_update. Reinstall via the Plugin Store, or remove these entries from config.json." auto_update is the core weekly-update setting from #581. Reconciliation treated every top-level dict not in its private _SYSTEM_CONFIG_KEYS list as a plugin id, and #581 could not know to extend that list. - Move core top-level keys into src/core_config_keys.py (CORE_CONFIG_KEYS) and use it in reconciliation. Tests fail if a config.template.json key or a key written by the general-settings save is missing from it. - A secrets-file key only counts as a non-plugin when no installed plugin has that id. Plugin secrets are namespaced by id, so installed plugins with secrets were reported as missing from config on every run. - still_unresolved() drops "not on disk" findings whose id is no longer a plugin entry in config, so a stored verdict clears without a restart. - A plugin whose id is a core key is skipped with a warning, and the fix never writes a plugin stub over or in place of a core setting. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9f2743471c |
fix(web): update-all skips Starlark apps and no longer misses plugins (#587)
* fix(web): update-all skips Starlark apps and no longer misses plugins Check & Update All posted every entry from /plugins/installed to POST /plugins/update, including the virtual starlark:<app_id> entries that list installed Starlark apps. The store manager cannot find those, so each answered 500 "plugin not found". Update-all now sends only plugin ids (install_manager.js, and the older app-shell.js copy), and the route answers a starlark: id with a 400 saying it is a Starlark app. A request that got no HTTP answer was recorded as failed and never sent again. On a device, a web-service restart mid-run killed the in-flight request and refused the next one, stock-news, which was left on 2.6.2 with 2.8.0 available. Such requests are now re-sent with backoff (about 30s) before being reported as failed. HTTP error answers are not retried. Tests: test/js/unit/test_update_all.js (run from pytest via test/web_interface/test_update_all_plugins.py so CI covers it) and the route contract for starlark: ids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): walk update-all retry delays without indexed lookup Codacy's ESLint security/detect-object-injection rule flagged retryDelays[attempt] as a High issue. The index was a bounded loop counter over a fixed array, but shifting a per-plugin copy of the schedule gives the same backoff without the pattern. No behaviour change: test/js/unit/test_update_all.js (21) and test/web_interface/test_update_all_plugins.py (7) pass unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fddb0e06db |
feat(web): honour x-display: hidden in plugin settings (#585)
* feat(web): honour x-display: hidden in plugin settings Plugins keep deprecated and internal keys declared so stored configs keep validating (weather api_key/radar_zoom, countdown's auto-generated row id), but the settings form drew them as live controls. A property marked "x-display": "hidden" -- or an object whose children are all hidden -- now gets no control at any depth: top level, nested sections, Advanced Settings (not counted either), array-table columns and the row editor. A hidden top-level key is not reported in __rendered_section. Saving never changes a hidden value. Plain and nested fields aren't posted, so the save's deep merge keeps them; _set_missing_booleans_to_false skips hidden booleans at every depth. A posted array row replaces the stored item, so hidden row properties are carried as JSON-encoded hidden inputs and decoded exactly on save (an id "1" stays a string). New rows get none. JSON API saves are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): read hidden row keys without dynamic property access Build the set of x-display: hidden item properties once and look values up through Object.entries, instead of indexing objects by a variable key on the lines this branch added (Codacy: object injection sink, 6 warnings). Behaviour is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
869e36fb2f |
feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback A General-tab toggle (off by default) checks for and installs LEDMatrix and plugin updates once a week, overnight in the configured timezone. - Pre-update checks skip (and report) instead of forcing: local edits or commits, merge/live rebase, no upstream, low disk, missing health check, or a version that was already rolled back. An abandoned rebase (HEAD back on a branch) is cleared, since it would otherwise block every pull. - The pull reuses the Update Code path (now perform_core_update(), which reports dependency install failures as data). - ledmatrix-update-verify.service, started via a .path unit from a request file, restarts the services from its own cgroup, requires them to come up and stay up, and otherwise resets to the previous commit and reinstalls the previous requirements. It runs a copy of the checker taken before the pull. - No SSH needed: switching the toggle on restarts the display service, which (as root) installs the two units from the repo templates for the web user. first_time_install.sh installs them too and takes --enable-auto-update / LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh). - Plugins update after the code passes its check; failures, blocks and rollbacks raise an Overview banner and show under the toggle. Tested end to end on a Pi: web-UI setup, a good update, a broken web service and a broken display (both rolled back), a blocked local edit, and an abandoned rebase found on the device. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(auto-update): address static-analysis findings - Replace the subprocess.CompletedProcess the verifier fabricated for a command that could not start with a plain namedtuple; nothing is executed there, but the scanner flags any CompletedProcess built from variables. - Mark the subprocess imports with the repo's standard B404 annotation (all calls are list-form argv, no shell). - Mark the rollback-failed message as not SQL (B608 matched its wording). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): CI failures on Linux - Keep the setup result when chown fails. CI runs as a non-root user, where chown to the web user raises; that discarded the result file, so the General tab would never learn whether setup worked. Regression test added. - Register the two new /api/v3/system/auto-update routes in the URL map snapshot. - Use utility classes app.css defines (space-y-1, hover:text-red-600). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): address review feedback - Health check: a failed restart command no longer lets the check run against the still-running old process; it counts as a failure (and after a rollback, as a failed rollback). An unreadable restart count is never treated as stable, since a crash loop looks healthy between attempts. - Installer writes the auto_update setting to a temp file and swaps it in, keeping mode and owner, so a running config watcher never reads a truncated config.json. - Verify unit quotes its command-line paths (install folders with spaces); setup refuses folder names systemd would reinterpret (%, quotes, backslashes, control characters) and says so on the General tab. - The auto-update status route no longer returns exception text. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): keep error detail in the status route's 500 test_web_error_detail requires every 5xx handler to log the traceback and return describe_exception(e), which redacts credentials, so failures are diagnosable from the web UI. Dropping it for CodeQL broke that policy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): dismiss route rejects non-object JSON with 400 A JSON array or scalar body made `.get('alert_id')` raise, returning 500. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): let the app-wide handler answer status-route errors CodeQL (py/stack-trace-exposure, #709) flagged the route's own except, which returned describe_exception(e). web_interface/app.py's error handler already logs the traceback and returns the same redacted detail for any unhandled exception, so the local copy is removed: same response, no new exception-to-response flow, and test_web_error_detail's policy still holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
814c21de1c |
chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0 (#580)
* chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0
Skins: no current scoreboard plugin builds on src.base_classes, so the only
skin hook (SportsCore._render_game) never runs. The plugin schema endpoint no
longer injects the Visual Skin dropdown, the store hides and refuses
"type": "skin" registry entries, and GET /api/v3/skins reports
supported: false with a message. Stored skin config still loads and saves.
src/skin_system/ and its tests are unchanged apart from the support flag.
Docs: check_plugin.py/render_plugin.py examples use --plugin; document
BasePlugin.get_update_interval() and its interaction with the manifest
update_interval; CLAUDE.md drops the stale template line number and
recommends display_manager.width/height.
Preview size: new src/display_geometry.py holds the size computation and
defaults DisplayManager uses (double-sided applied, chain_length default 2).
The web preview, /display/current, Starlark magnify default, sync handshake
and two dev scripts use it.
Release: __version__ 3.4.0, CHANGELOG 3.4.0 section plus a 3.3.0 tag note.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: address CodeRabbit review on #580
- Preview fallbacks (SSE stream and /display/current) use logical_size({})
(128x32, the shared default) instead of a hard-coded 128x64.
- display_geometry treats a non-mapping display/hardware block as missing,
so a malformed config.json falls back to defaults instead of raising
AttributeError (which turned the Starlark render into an HTTP 500).
- Docs: the static update interval falls back manifest -> plugin config
-> 60s, in both the API reference and the architecture spec.
Not taken: validating double_sided copies against chain_length/parallel.
An orientation Rotate: or U-mapper pixel mapper decides which axis panels
lie on, so the counts would reject working setups (the existing
vertical-split test is one).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display_geometry): a non-finite hardware size raises ValueError, not OverflowError
CodeRabbit flagged the Starlark magnify default in
_standalone_render_starlark_app for truthy non-mapping display values. That
case was already handled by
|
||
|
|
11bf39cd66 |
fix(web): two plugin config saves that always returned 400 (geochron, news) (#575)
* fix(web): render widget-less arrays of objects as a table, not comma text An array of objects with no x-widget (geochron's `cities`) fell through to the comma-separated text input. Jinja joined each item as a Python dict repr, the save route read them back as a list of strings, and the schema rejected them -- so every save of the plugin returned 400 "Configuration validation failed", whatever setting was changed. Default such arrays to the existing array-table widget, which already edits arrays of objects and posts `field.N.key` inputs the save route rebuilds into a list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): don't leave an empty object stub in array items on save The unchecked-checkbox pass walked into every nested object of an array item looking for booleans, creating it when absent. A news custom feed with no logo came out with `logo: {}`, which fails the logo's `required: [id, path]`, so every save of the news plugin returned 400. Recurse into a scratch dict instead and attach it only if a boolean was actually set in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7e580dc005 |
fix(wifi): make Connect work from the setup AP (#571)
* fix(wifi): make Connect work from the setup AP Joining a network from LEDMatrix-Setup has to take the AP down first, which drops the phone that sent the request. The connect endpoint answered only after the attempt finished, so the browser never got a reply and the WiFi tab's Connect button appeared to do nothing. - /wifi/connect answers 202 immediately while the AP is active and connects in a background thread; the result (never the password) is reported via /wifi/status as last_connect_attempt. A second connect while one is pending gets 409. - connect_to_network holds a /tmp flag for the attempt; the monitor daemon skips AP management while it is fresh. Previously the daemon's disconnected counter, accumulated over the whole AP session, re-enabled the AP on its next tick in the middle of the connect. - The WiFi tab and captive setup page explain the handoff up front, and on reopening show why the last attempt failed. The wrong-password message now works: the route sets the error_type the captive page checks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(wifi): serialize connect attempts on both paths Addresses CodeRabbit review on #571: - Check for a pending attempt before branching on AP state. A background attempt takes the AP down long before it finishes, so a second click used to bypass the 409 and start a competing synchronous connect. - Record pending for the synchronous (non-AP) path too, so two requests can't overlap and have the first clear the daemon's in-progress flag while the second is still connecting. - Clear the pending state if the background thread fails to start, rather than refusing every later request until restart. - Say the setup network returns "within a few minutes": a stale flag plus the daemon's grace period can take longer than one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d1e821c625 |
fix(web): harden, polish and optimize the web UI per the Sept 2026 audit (#568)
* fix(web): harden, polish and optimize the web UI per the September 2026 audit Works through docs/archive/WEB_UI_AUDIT_2026-09.md (health 8/20). Implementation integrity (P0) - app.css now defines every utility class the templates and JS use, including .hidden, so the ~145 JS show/hide toggles work. Button reset, and base component rules (.btn, .form-control) wrapped in :where() so utility classes on the same element win. New static-audit test fails when a used utility class has no rule. Accessibility - Focus rings render (the old ring rule referenced undefined variables); one :focus-visible outline everywhere; skip link; labelled nav landmarks. - Shared dialog helper (js/utils/dialog.js): role/aria-modal, focus trap, Escape, focus return, applied to every modal. - Named icon-only buttons and labelled ~70 form fields. - Toasts announced once; errors persist >= 10s; one showNotification. - Captive WiFi page: live region, timeouts, dark mode, 16px inputs. Performance (Pi Zero 2 W) - SSE streams and tab timers pause when hidden or off-tab; the display stream only runs while a preview is visible. app-shell.js deferred. - Widget scripts served as one versioned bundle (/assets/widgets.js): 52 -> 21 script tags, 66 -> 35 requests on first load. - Stdlib gzip fallback when flask-compress is missing: first-load JS/CSS 1358 KB -> 291 KB on the wire. SSE untouched. Theming and responsive - File managers, form fields and Fonts upload on theme tokens; bare inputs themed in dark mode; no more white surfaces. - No horizontal overflow at 375px on any tab; 44px touch targets on coarse pointers; reduced-motion respected; header title truncates. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): clear Codacy findings on #568 - json-file-manager: focus-trap releases kept in a Map (no dynamic property access or delete; no value-returning forEach callback) - notification / schedule-picker: style and day-label lookups via Map - app.js: move the pending-queue assignment out of the expression - diff_viewer / error_handler: named function declarations instead of arrow consts No behavior change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: check the OAuth widget ships in the widget bundle base.html no longer tags widget scripts one by one; they load through /assets/widgets.js. Assert the page requests the bundle and the bundle contains google-oauth.js, which is what the test was protecting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): address review feedback on #568 - widget bundle version fingerprints every file (name, mtime_ns, size) - gzip fallback appends Accept-Encoding to an existing Vary header - dialog helper: releasing a non-top dialog no longer moves focus out of the dialog the user is in - labels: file-upload targets its file input; fallback config fields get label for/id pairs; native color input has a fallback name - utility audit also reads class names inside bound :class expressions Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): give the native color-picker input an accessible name CodeRabbit flagged this on PR #568 as an outside-diff finding (never posted inline, so it was missed in the round of fixes that addressed the other 6 review comments). The <input type="color"> only carried a title attribute; screen readers don't reliably announce title, and there's no other label naming the control when showHexInput is false. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): clear Codacy findings in app-shell.js - drop the unused catch binding on the SSE JSON parse - move the pending-notification queue assignment out of the expression No behavior change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): contain plugin widgets/ dir and bound style-editor retries From CodeRabbit review on #568 (code that arrived with the main merge): - serve_plugin_widget resolves widgets/ with resolve_under before resolving the manifest script under it, so a symlinked widgets directory can't become the containment base (CWE-22). New test. - style-editor init stops polling after ~10s when the widget never registers and leaves the plain fallback fields in place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
69d408b321 |
feat(core): one per-element display-customization framework, wired into the web UI (#566)
* fix(sports): rebuild un-shared faces through the pinned layout engine unshare_element_fonts re-instantiates a duplicate font face so two elements can be told apart by id(). It did so through bare ImageFont.truetype, which takes PIL's default layout engine rather than the one src/common/font_layout.py pins. Raqm and Basic disagree on fractional advances -- that disagreement is the reason the pin exists, having broken golden images across machines -- so a rebuilt face could measure differently from the shared face it replaced, on any host where Raqm is installed. These were the only two call sites in src/ bypassing the pin. The guard asserts that the rebuild goes through the pinned loader rather than comparing engine values: where Raqm is absent, bare truetype returns BASIC anyway, so an engine comparison passes whether or not the pin is honoured. The first draft of this test did exactly that and passed with the bug reintroduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): drop the two dead client-side config-form renderers generateConfigForm and generateSimpleConfigForm (580 lines) were defined on the Alpine component and never called: server-side Jinja replaced them, as pages_v3.py:641 records. Nothing in any template invokes them -- there is no x-html in the templates and no bracket access on the component. They carried their own x-widget dispatch, which made them an active trap: the next person adding a widget would reasonably think both renderers needed updating. plugins/config_manager.js (PluginConfigManager, 133 lines) goes for the same reason -- loaded on every page from base.html, referenced only by itself and by an archived doc. Kept, having checked them: widgets/example-color-picker.js is the worked example docs/widget-guide.md points plugin authors at, and widgets/plugin-loader.js is the client half of a documented feature (manifest-declared plugin widgets) whose server route is missing -- soccer-scoreboard already ships a widgets/custom-leagues.js that this loader is meant to fetch. That is an unfinished feature to complete, not dead code to delete. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(web): serve plugin-declared widgets, and actually ask for them LEDMatrixWidgets.loadPluginWidget has always fetched /static/plugin-widgets/<plugin>/<widget>.js, and docs/widget-guide.md has always documented that path, but nothing served it. soccer-scoreboard has shipped a 17KB widgets/custom-leagues.js since August that could never load. Both halves were missing, not just the route: - serve_plugin_widget serves the script from the plugin's widgets/ directory as text/javascript. The manifest is the allowlist -- only a widget the plugin declares is reachable -- so installing a plugin does not publish everything it ships. Path handling mirrors the sibling serve_plugin_web_ui: allowlist regexes, os.path.basename, resolve() + relative_to() containment, and the ledmatrix- prefix fallback. The declared script name is guarded too, since it comes from the plugin rather than the request. - The config form never requested one. Its x-widget dispatch is a hardcoded list of core widget names, so a plugin's own widget fell through to a plain text input. An unrecognised x-widget on a string field now asks ensureWidget() for it. The text input stays as the fallback and is removed only once the widget has actually rendered, so a missing or broken widget costs the user an editor rather than their configured value on the next save. - manifest_schema.json gains "widgets", so the declaration is validated rather than merely tolerated by additionalProperties. Verified in a browser against the real partial: a declared widget loads, registers and renders, and its field posts exactly one value; a field whose widget 404s keeps its text input and still posts its value. Not addressed: loadPluginWidgetsFromManifest still has no caller. The per-field ensureWidget path is lazier and is what the form now uses, so that bulk helper is dead weight -- worth removing, but left alone here rather than inventing a call site for it. Known limitation, documented: only string-typed fields take this path. object/array/boolean/number fields and enums are dispatched by the template's own branches, which still only know core widgets. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(element-style): a wrong-size BDF now keeps its font, not its size BDF fonts are fixed-size bitmap strikes: FreeType accepts only the pixel size baked into the file and raises for anything else. 32 of the 35 shipped fonts are BDF, so a size picked in the web UI usually is not a valid strike -- and load_font caught that failure with its generic "unloadable font" handler, which substitutes PressStart2P. Asking for 5x7.bdf at size 10 therefore rendered a completely different typeface, silently. It now falls back to the file's own native size instead, which is what SportsCore._load_custom_font_from_element_config has always done. The native size is read via FontManager._read_bdf_native_size rather than a fourth copy of that parser, matching how core.py already delegates. Also here, because they are the same code path: - native_bdf_size() is exposed for the web UI, which needs to know when a size field can take effect at all. None means "free choice". - ElementStyle.font_size now reports the size actually realised rather than the one requested. Callers lay out from it, and reserving space for a size nothing was drawn at is how this surfaces. - The module font cache is a bounded LRU (256) instead of an unbounded dict. The display process runs for weeks and every config save can add a (font, size) pair; every other hot cache in the codebase is bounded this way. Untouched configs are unaffected: the shipped classic fonts are the three TTFs, so nothing was hitting the substitution path by default. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): per-mode style and offset overrides Lets one element be styled differently per situation -- a scoreboard's live / upcoming / recent cards, weather's current / hourly / daily screens -- under customization.modes.<mode>. The mode is bound at construction rather than passed per call. That is what makes this cheap to adopt: SportsUpcoming and SportsRecent are already separate instances with distinct SKIN_MODE values, so binding once makes every existing style()/offset_value() call site mode-aware without editing any of them. A per-call mode argument exists for the rare host that renders more than one mode. The two layers answer different questions, deliberately: - The base layer keeps the existing "differs from the schema default" rule, because the save flow writes the full default object into config.json whether or not the user touched it. - A mode layer is pure override -- its fields default to None, so presence is intent. Nothing writes into it unasked, so there is nothing for the stricter rule to protect against. None therefore means inherit, and has to stay distinct from 0: a mode y_offset of 0 means "sit at the base position", not "no preference". This is the same distinction scroll_card.switch_* draws with "inherit". A malformed mode value falls back to the resolved base value rather than to the caller's default -- caught by the degradation tests, which is what they are for: resolving the mode first let one bad string in a mode block silently discard a good base offset. With no modes block, and for every existing caller, resolution is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): declare per-mode overrides in config_schema.json A plugin adds "x-style-modes": ["live", "upcoming", "recent"] alongside its x-style-elements declaration and gets a customization.modes.<mode> group per mode, with every field of every declared element repeated as an override. Those override fields are typed nullable and default to null, which is the whole trick. The save flow writes schema defaults into config.json wholesale, so giving a mode field the base element's default would make every mode a frozen copy of the base the first time a user pressed Save, and the base would stop reaching them. Null means inherit. The mutation test for this is explicit: with concrete defaults, a base font_size of 14 resolves as 10 with user_forced set. min/max from the declaration carry into the mode blocks, so an out-of-range override is rejected by validation rather than clamped silently at render time. Also: the emitted font field now carries "x-widget": "font-selector". The widget already shipped and the config form already allowlisted it -- the hint was simply never emitted, so the field rendered as a bare text box that the user had to type a font filename into. Verified through the real SchemaManager path -- load_schema, defaults extraction, merge_with_defaults, validation, then resolution -- rather than against a hand-built dict, since the thing at risk is what that pipeline does to a null. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): render the config form from the schema the save route validates The form read config_schema.json with a raw json.load while api_v3.save_plugin_config went through SchemaManager. Those are not the same schema: SchemaManager applies expand_style_elements, which turns a compact customization.x-style-elements declaration into the per-element blocks the form knows how to render. Without it, that customization object has an x-style-elements key and no "properties", so the template's object branch matched nothing and the section rendered as empty space -- while saving still validated against the expanded shape. of-the-day ships the compact form, so its customization section has been invisible in the web UI. pages_v3 gains a schema_manager the way it already has config_manager and plugin_manager. use_cache=False matches the save route, so an edited schema is not served stale during plugin development. The raw read stays as a fallback for callers that register this blueprint without one. Checked before making the change: load_schema does nothing here except read, validate and expand -- inject_skin_selector is a separate method it does not call -- so this is not a behaviour change for schemas without the declaration. The test pair renders the same compact schema with and without a SchemaManager, so it documents exactly what was broken as well as what is fixed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(web): style-editor widget -- a row per element instead of 65 accordions Rendered element by element, a realistic scoreboard's customization block is 65 nested sections, and reaching one per-mode font size takes five levels of expanding. The widget collapses that to one compact row per element -- font, size, colour, X, Y -- with a tab per declared mode. It emits ordinary inputs under the same dotted names the generic renderer would produce, so the save/validate/merge pipeline is untouched: no hidden JSON blob and no new server-side parsing. It is driven entirely by the schema block it is handed, so fields added to the schema later appear without editing the widget. If it fails to load or throws, the generic nested rendering it replaces is left in place. Fixing two things the save path got wrong for nullable fields, found by posting what the widget actually emits: - The indexed-array recombiner (text_color.0/.1/.2 -> one list) compared the declared type to the string 'array', so a per-mode colour, typed ["array", "null"], was never reassembled and failed validation on save. _parse_form_value_with_schema had the same comparison. - A blank nullable field became [] rather than None, which then failed the minItems the colour array declares. Null is the inherit sentinel, so it has to survive. And two things the widget itself got wrong, found by looking at it: - An unset base control fell back to the select's first option, so an untouched scoreboard claimed every element used 10x20.bdf -- and the size box then locked itself to that bitmap font's fixed size. Base controls now show the schema default; mode controls stay blank, because blank there means inherit. - Elements arrived alphabetised (Detail and Odds above Score). Flask's JSON provider sorts keys, so declaration order has to be stated explicitly; expand_style_elements now emits x-propertyOrder, which the generic renderer already honoured too. Size is disabled and shown as fixed for a bitmap font, using the scalable/native_size the font catalog now reports. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): visibility, alignment and scale per element Completes the customization vocabulary: hide an element, align it, and resize a logo, alongside the font/size/colour/offset that already existed. All three per mode. They resolve to "change nothing" until the user asks for something -- True, None and 1.0 -- rather than to whatever the schema declares. That is the same invariant the font fields keep: a caller that honours them still renders an untouched config exactly as it did before they existed. A schema default therefore does not count as a choice, which matters because the save flow writes that default into config either way. scale sits in the layout block with the offsets rather than in the element block, because it is geometry: a logo has a scale and no font. The widget's columns come from the schema, so a logo row shows visibility, offsets and scale and no empty font cell. Two bugs found by the tests rather than by reading: - A nullable enum needs null in its enum list, not just in its type. The mode copy of `align` defaulted to null and then failed its own schema, so a plugin declaring any enum field with modes could not save at all. Six tests failed on this before any of them reached what they were testing. - defaults_from_schema only ever extracted font/font_size/text_color, so the schema defaults for the new fields were invisible to the resolver and a declared default read as a user choice. Widget: the table scrolls horizontally and pins the element-name column. Nine columns do not fit the config panel, and clipping them hid the offsets entirely while scrolling them made every row anonymous. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): resolve elements under the names plugins actually use Two naming conventions collided as the scoreboards grew. Counted across the published schemas: the style block names elements with a _text suffix (score_text, status_text, detail_text), while the layout block mostly uses the bare noun (score, date, time, odds) -- except status_text, which kept the suffix in seven plugins and lost it in two. records vs record splits seven to two the same way. A lookup now tries the exact name first and then the spellings that mean the same thing. Exact-first is what makes this inert for any config that already matches; the aliases only decide cases that resolved to nothing before. This is also what makes migrating to the compact declaration form safe. That form uses one key for both blocks, so a scoreboard adopting it asks for layout.score_text while its users have layout.score saved -- without the aliases, every offset they had dialled in would silently become 0. Applies to the style block, the layout block, the schema defaults and the per-mode overrides, since the drift shows up in all four. Not attempting to canonicalise on write: renaming keys in config.json would break the plugins still reading the old spelling from their own bundled code, and the drift costs a dict miss rather than correctness. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(plugins): BasePlugin.styles -- per-element styling every plugin inherits Adopting the element-style system meant repeating three things in every plugin: a guarded import, finding its own config_schema.json, and rebuilding the resolver when on_config_change swapped the config dict. This is those three things once, on the class all 45 plugins already inherit from. title = self.styles.style('title_text', classic_font='PressStart2P-Regular.ttf', classic_size=8, classic_color=(255, 255, 255)) The classic_* arguments are the adoption contract: with nothing configured they come back verbatim, so a plugin that switches to this renders exactly as before until a user changes something. A plugin with one instance per display mode sets STYLE_MODE on the class and every existing lookup becomes mode-aware without a call site changing -- which is the point of binding the mode to the resolver rather than passing it per call. styles_for() covers a plugin that renders several modes from one instance. Schema discovery reads the concrete class's own module rather than this file, because this file lives in src/plugin_system where no plugin schema exists -- the same trap SportsCore._config_schema_path documents. The first mutation test for that passed anyway: an installed plugin's module directory and its entry under plugins_dir are the same path, so the test could not tell the two apart. The case where they diverge is a plugin symlinked in for development, and the test now forces that shape. Getting discovery wrong is silent rather than loud: with no schema the resolver has no defaults to compare against, so every configured value reads as a deliberate override and the plugin quietly stops honouring its own shipped styling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): adopt hand-written customization blocks, and widen the font list Nineteen plugins spell their style elements out longhand instead of declaring them -- football's block is 701 lines for seven elements -- and predate this system entirely. Core now recognises that shape, so they pick up the row-per-element editor and the real font picker on a core update rather than on a plugin release. Checked against every published schema: 21 plugins adopt, and the defaults of each still validate against the schema generated for it. Detection requires *every* field in a block to be one this system understands. A looser "has at least one style field" rule sweeps in baseball's `count`, which carries a text_color beside geometry that means nothing here. That distinction took three attempts to test: the first two assertions passed under both rules, because an over-eager rule leaves a fontless block looking untouched and only surfaces as an extra row in the editor. The hardcoded font enum is replaced rather than extended. Football lists five of the thirty-five installed fonts, which is why a font a user uploads can never appear in one. It is not a curated safe set -- it omits some twenty other faces that fit the declared size cap just as well -- it is the fonts that happened to exist when it was written. Widening it does need a guard, though, and not the one the schema already has: a bitmap font ignores font_size and renders at its size baked into the file, so `maximum: 16` cannot stop a 27px face. The picker now filters out fixed-size fonts taller than the element's own declared ceiling, which drops exactly the four that would overflow a 32px panel and keeps the other thirty. Per-mode overrides stay opt-in: core cannot invent a plugin's display modes, so `x-style-modes` remains the one line that unlocks them. Their layout half covers every positionable element rather than only those with a style block -- the two namespaces do not line up in a hand-written schema, and football positions six things (logos, timeouts, possession) that have no style block at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): remove the two Fonts-tab panels that reported invented data "Element Font Overrides" let a user configure an override, showed a success toast, and changed nothing. All three endpoints behind it were stubs -- GET returned a hardcoded {}, POST and DELETE returned success without calling anything -- each marked "This would integrate with the actual font system". Wiring them to FontManager would not have fixed it. The machinery there is real (_load_overrides/_save_overrides persist config/font_overrides.json, resolve_font applies them, and the countdown plugin genuinely consumes it), but the panel's element dropdown offered eleven invented keys -- nfl.live.score, clock.time, weather.current -- that no plugin has ever read. An override saved against one of those would have persisted correctly and still done nothing. "Detected Manager Fonts" goes for the same reason. It claimed to show "fonts currently in use by managers (auto-detected)"; its own comment said "we'll simulate this", and it listed every font in the catalog with a hardcoded usage_count of 1 -- the panel beside it, with fabricated numbers attached. Per-element font choice now lives in each plugin's own config editor, against the elements that plugin actually has, and covers size, colour, offsets, visibility, alignment and scale rather than family and size. Kept: the font library (upload, preview, delete), which works, and /fonts/tokens, which is a stub but genuinely feeds the preview's size dropdown. FontManager's override methods are untouched -- countdown uses them. Verified in a browser with the tab's JS running: no console errors, 35 fonts listed, upload and preview intact. Removing the panel meant unwiring it from populateFontSelects too, which would otherwise have bailed out early on the missing select and left the preview dropdown empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(sports): one reader for element colours and layout offsets There were two copies of the per-element colour read and three of the layout-offset read. They had already drifted -- the scroll-card renderer carries a comment about having ignored offsets its own schema advertised -- and each new capability had to be added to all of them or silently work in some places and not others. All of them now go through src.element_style, which is what carries the alias handling and the per-mode lookup. That lands immediately for the nine plugins importing these modules: a scoreboard asking for `score_text` offsets finds the `layout.score` its users configured, and a Live instance resolves its own colours through SKIN_MODE without any call site passing a mode. _normalize_color learned "#RRGGBB" in the process. The scoreboards' own readers have always accepted it, so the shared one had to, or consolidating would have quietly dropped a form users' configs may hold. _coerce_offset picked up the non-finite guard the scroll-card reader had and the other two did not. _get_layout_offset is promoted onto SportsCoreSharedMixin. Each plugin still carries its own copy in its bundled sports.py, which wins by MRO -- so adopting this is a deletion in the plugin, and until that deletion nothing changes for it. Note for whoever runs the suite next: test_display_dirty_tracking.py is order-dependent. Fifteen of its tests failed in one full run and passed in the next with no change in between, and pass in isolation. Pre-existing, unrelated to this, but it makes a full-run diff untrustworthy until it is fixed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): record the element-style work under Unreleased This file's own preamble asks for it: a plugin may delete its bundled fallback copy of a core module only when its manifest floors on the first release that shipped that module, which requires the additions to be recorded here against a version. Names a plugin can now import and floor on -- the stateless layout_offset and element_color readers, alias_keys, native_bdf_size, the resolver's mode binding, BasePlugin.styles, and the promoted SportsCoreSharedMixin._get_layout_offset -- plus the schema and web-UI changes, the four fixes and the three removals. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(fonts): log the BDF native-size read failure instead of swallowing it The bdf-native-size lookup in get_fonts_catalog() caught any exception and silently discarded it. Every other guarded read added in this PR (the manifest parse in _declared_widget_script, the SchemaManager fallback in _load_plugin_config_partial) logs before falling through to the same degraded behavior. This one didn't, which is the shape a silent-exception-swallow lint rule flags. Behavior is unchanged -- native_size still comes back None -- but a corrupt or unreadable BDF file now leaves a trace. Verified: font-related tests (140) and the full suite still pass, with only the 2 pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias failures already present on origin/main. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address CodeRabbit findings on the style-editor/font-selector PR - Fix _load_font_sized double-wrapping the (font, size) tuple on the missing-font path, which handed callers a tuple instead of a font. - Fix _set_nested_value skipping an explicit None when the key already existed, which silently kept stale overrides when a user cleared a nullable per-mode field or blanked all channels of an indexed color. - Preserve BDF scalable/native_size metadata through fetchFontCatalog's catalog-format mapping so maxFixedSize filtering actually applies. - Stop caching an empty array on a failed font-catalog fetch so a later call can retry instead of being stuck with the failed result. - Keep a saved font selected in the style editor even when it no longer fits a newly declared maxFixedSize, instead of silently deselecting it. - Don't drop in-progress user edits to fallback fields when a plugin widget finishes loading asynchronously and takes over the form. - Tighten the removed font-override endpoint test to assert 405, not just != 200. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): a partial save no longer switches off checkboxes it never showed An HTML checkbox posts nothing when unchecked, so the save route walked the schema and forced every boolean missing from the form to False. That is right for the rendered form and wrong for every other caller: a script, the MQTT bridge or a curl against the documented endpoint never rendered a checkbox, and reading its silence as "all off" turns a one-field save into a mass disable. Found on hardware. Posting four customization.* keys to a live device switched off nfl.enabled, ncaa_fb.enabled and every display-mode toggle in one request. The form now reports the top-level sections it drew (__rendered_section), and inside those an absent checkbox still means unchecked -- including a section whose only fields are checkboxes that are all off, which no heuristic could recover. A post with no marker only touches objects it actually posted a field from. Meta fields are dropped before form keys are treated as config paths, because unknown keys are otherwise written straight into config.json. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(sports): resolve element colour by name, and honour visible/align/scale Two of the three gaps this framework shipped with. Colour by name. A draw resolved its colour by comparing the *identity* of the font object it was handed, which cannot tell two elements apart when they share a face -- so those draws went out white. Every bitmap font is in that case, because a freetype.Face cannot be re-instantiated to un-share it, which is how an element rendered in any of the 32 shipped BDF fonts silently lost a colour its picker had offered all along. _draw_text_with_outline now takes element="score_text" and reads the colour by name; the identity path remains for un-annotated callers, but narrows before giving up -- one configured colour among the sharers is the only thing the user can have meant. Visible, align and scale. The resolver has understood these since the framework landed and nothing consumed them: an element could be marked hidden in the web UI and still render. Adds the stateless readers, the mixin accessors, and a scale parameter on the one shared logo-sizing seam (keyed into the cache, so two elements scaled differently cannot be served each other's image). Naming an element in a draw also honours its visibility. Untouched configs are unaffected: every new parameter defaults to today's behaviour, and all ten affected plugins render pixel-identically to main across every harness size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(plugins): how to declare styleable elements; harden the widget's lookups The plugin-author guide for the compact x-style-elements declaration -- what each key does, how to read values back without breaking the "user-forced only when it differs from the default" rule, and why a hand-written block needs no changes to be adopted. Also clears the static-analysis findings on style-editor.js. Every lookup in that file is keyed by something out of a schema or a saved config, so a key of __proto__ or constructor would walk the prototype chain and hand back a function instead of a schema; reads now go through an own-property helper. The panel registry became a list, and the flagged vars moved to their function roots. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): clear the remaining static-analysis findings Five, all on lines this branch touched. The Python one is not a new defect: _set_missing_booleans_to_false's first parameter was always named `config`, which shadows the `config` submodule imported for its side effects at the bottom of this module. Editing the signature simply put the existing warning on a changed line. The parameter is the plugin's config dict, so `plugin_config` is what it should have been called anyway; callers pass it positionally and are unaffected. The JavaScript ones are the object-injection rule firing on reads keyed by data. own() now goes through a property descriptor, so the one unavoidable data-keyed read is no longer a computed member access; at() consumes its path instead of indexing it; and the column set is a Map, which has no prototype to pollute and needs no guarded reads at all. Verified the widget still renders identically against football's real schema: 29 element rows, all four mode tabs, values populated, no console errors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): drop the hasOwnProperty alias the descriptor read made redundant own() now reads through Object.getOwnPropertyDescriptor, so the alias it used to call has no remaining reference. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f6367d63ae |
security: triage the CodeQL backlog — 129 alerts, three of them live (#561)
* fix(web): escape quotes in every HTML escaper, not just & < >
The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:
value="${escapeHtml(v)}" title="${escapeHtml(v)}"
so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.
It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.
app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.
cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `'`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.
url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).
test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): stop request-supplied names from reaching paths outside their base
Three of the py/path-injection alerts were live, not lint:
* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
whose resolved path *string-prefixed* the plugin directory. Flask's
default converter forbids a slash but not dots, and
get_plugin_directory('..') returned the parent of the plugins directory
because it exists -- so every file under the project root then prefixed
that directory, config/config_secrets.json included. The prefix check
was also wrong on its own terms: with plugin dir "plugin-repos/foo",
"../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
start with "plugin-repos/foo".
* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
file_id of "../../../../etc/something" deleted that file. This is the
one finding in the batch that destroyed data rather than exposing it.
* POST /api/v3/cache/delete passed the body's key through
CacheManager.clear_cache to DiskCache, which joined it as a filename and
called os.remove. Same shape, same result. The guard goes in
DiskCache.get_cache_path, the single choke point get/set/clear share, so
every caller is covered rather than just this route. Real keys are the
stems of files already flat in the cache directory -- that is how
list_cache_files derives them -- so nothing legitimate is turned away.
The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.
Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.
test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): refuse a plugin id that is not a plain name, don't truncate it
pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.
Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(web): say what the plugin web_ui iframe actually is
The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.
Also drops the now-unused os/os.path imports.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): inline url-input's scheme guard at the previewLink.href sink
CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.
Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.
Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): address CodeRabbit findings on the CodeQL triage PR
- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
credential-saving or connect flow runs. NetworkManager only accepts
printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
was previously saved/attempted before nmcli itself rejected it.
- web_interface/blueprints/api_v3/config.py: fail closed when the
plugin config schema path can't be resolved under the plugins
directory (e.g. a symlinked plugin dir). Previously this fell
through with secret_fields left empty, so submitted credentials for
that plugin were saved as ordinary, unencrypted configuration.
- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
splicing the JSON day/column key into an inline oninput="..." handler
string. escHtml() escapes quotes for a normal HTML attribute, but the
browser HTML-decodes the attribute before running it as script, which
undoes that escaping and lets a crafted column name (e.g. from an
uploaded JSON file) break out of the JS string and execute. Cell
edits now travel through data-day/data-col attributes read by one
delegated 'input' listener instead.
While in this file: fixed 6 pre-existing missing-')' typos on
multi-line safeSetHTML(...) calls (already flagged by Biome in this
PR's own CodeRabbit run as syntax errors blocking its lint pass).
These predate this PR (present on main too) but made the whole file
fail to parse in any JS engine, which is a bigger problem than the
XSS finding itself and directly touches the same lines.
Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5137e86d16 |
feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab PR #544's change, ported onto the api_v3 package split (#553). Identical behaviour; only the placement of the new code differs. The original added 508 lines to web_interface/blueprints/api_v3.py, which #553 deletes, so every hunk of it would conflict irreconcilably. Ported by AST: 26 new top-level items sorted to where the split puts each kind -- __init__.py 2 imports, 11 constants, 7 helpers starlark.py 4 routes (/starlark/editor/{apps,status,start,stop}) misc.py 2 routes (/integrations/mqtt-bridge{,/config}) Everything outside api_v3.py -- the Tools partial, the installer scripts, the JS tests -- applied unchanged. Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is regenerated to match, which is exactly what test_api_v3_url_map.py is designed to make you do when routes are added. Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR ships could not be run here -- node is not installed on this machine -- so test/js/dom/test_tools_sections.js is unverified. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST The AST-based port of #544 onto the api_v3 package split dropped `time` from starlark.py's import list. start_pixlet_editor() and stop_pixlet_editor() both call time.time()/time.sleep() directly, so every start (NameError building `state['started_at']`) and every stop that has to wait out the EXIT trap crashed with a 500. No test caught it because the route's own tests mock subprocess.Popen but never actually invoked it before now. Also carries over #544's later fix that this port branched before: env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an operator who had already pinned PIXLET_EDITOR_HOST to loopback, forcing the unauthenticated `pixlet serve` process onto the LAN regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as api_v3.starlark.py's siblings already do for _pkg-owned names. Both fixes route the shared _pkg.time reference the rest of the package's route modules already use for anything a test might need to patch, rather than a bare `import time` local to this file. Ported the existing regression test from #544 (TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's module layout (web_interface.blueprints.api_v3.starlark instead of the old monolithic api_v3 module), which is what caught the NameError. Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this branch and on origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): clear the six lint errors this rebase introduced All six were introduced by rebasing this branch onto the merged blueprint split, not by the split itself. Confirmed by diffing pyflakes output against main with line numbers normalised -- everything else it reports is present on main too and is the package's deliberate re-export pattern. starlark.py used _STARLARK_APPS_DIR three times without importing it (F821). The rebase resolved an import-list conflict as a union of both sides, and that symbol was on neither side of the conflict hunk, so it was silently lost. It is defined in __init__.py and is now imported like its neighbours. This was the only one of the six that would fail at runtime rather than merely lint. __init__.py imported contextlib twice (F811): the cherry-pick added one next to the existing import. Removed the duplicate; the original at line 19 is used. __init__.py imported signal purely to re-export it to starlark.py, so pyflakes saw it as unused (F401). signal is stdlib and does not need routing through the blueprint package, so starlark.py imports it directly and __init__.py no longer does. contextlib stays re-exported because this module genuinely uses it. _read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule this module imports at the bottom for its route side effects (F811). Renamed to `settings`, with a comment saying why, since the name is otherwise the obvious one to reach for. Verified: pyflakes now reports nothing on this branch that main does not, the package imports, all nine route modules load, and 117 routes register, matching the pinned URL-map snapshot. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply Two CodeRabbit findings on the bridge settings endpoint, both of which returned 200 while doing something other than what the caller asked. `request.get_json(silent=True) or {}` turned a missing or unparseable body -- and the JSON literals null, [] and false -- into an empty dict, which then satisfied the isinstance(data, dict) guard on the very next line. The guard was there to reject exactly those bodies. Dropping the `or {}` lets None fail it. The same `or {}` on /errors/clear is left alone: its docstring documents the body as optional, so an absent body legitimately means "use the defaults". The difference is that saving settings has nothing sensible to do with no body. `if data.get('clear_password'):` accepted any truthy value, and the string "false" is truthy in Python -- so a client echoing the field back as a string wiped a password it meant to keep. Now coerced through the package's existing _coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else. test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus a missing body, and clear_password across truthy and falsy spellings. Verified against the unfixed code -- reverting the body guard fails 5, reverting the coercion fails 3. Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials when TLS is off (CWE-319). That is a policy decision about the feature rather than a defect -- unencrypted MQTT on a trusted LAN is common and often deliberate -- so it is raised on the PR for a maintainer call instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: work through the remaining review findings on the editor and bridge allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses the network in cleartext. Refused now rather than merely warned about -- but refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults false, and is coerced like the other booleans so the string "false" cannot switch the guard off. starlark.py:796 -- the supported service runs Flask threaded, so two start requests could each see running=False, each launch an editor, and the second state write replace the first PID, orphaning a process that holds the display down with nothing recording it. The check-launch-write sequence now takes a module-level lock. starlark.py:848 -- if the state write failed the route returned success with an editor running and no PID recorded: status and stop both reported no session while the display stayed down until the timeout expired. It now terminates the process group and returns an error. starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so nothing hands the display back, yet the response said "the display is restarting". After an escalation the display is now restarted explicitly, and a failure to do so returns an error naming the manual step instead of a success. pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the failure lands before the display is stopped rather than after. pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface address are three cases, not two. Collapsing the last two printed a URL saying "localhost" whenever PIXLET_EDITOR_HOST named a LAN address. tools.html:1254 -- escHtml does not encode single quotes, and the app id was interpolated into an inline onclick="startPixletEditor('...')", so a directory containing an apostrophe could break out of the JS string and run script. The handler binds with addEventListener and reads the id from dataset, where it is only ever parsed as an HTML attribute. Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in in both directions. The tools DOM suite gains three guards asserting the edit buttons carry no inline onclick and pass the id via dataset -- those need jsdom and did not run here, so CI verifies them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): log the traceback on the editor state-write failure The 848 fix answers 500 when the session state cannot be written, and logged that at error level -- but without exc_info, so the traceback never reached the log. test_web_error_detail.py guards exactly this: a handler returning 5xx must write an error-level record *with* the traceback and return the sanitized detail, because checking that merely something was logged is too weak. Caught by Core unit tests on the previous commit, not locally: the guard parses every module under web_interface/blueprints/api_v3 as one source, so it only fires once the whole package is read together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bdb9a94033 |
refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181 functions in one module -- 9% of the core by line count and three times the next largest file. It becomes a package of nine route modules grouped by path segment, plus __init__.py for the shared imports, constants, Blueprint and helpers. Every route module decorates the SAME api_v3 Blueprint object, so endpoint names stay api_v3.<function>, the URL map is unchanged and app.py is untouched. Verified: 111 routes before, 111 after, byte-identical rules, endpoints and methods, and every endpoint still on the one blueprint. plugins 3,867 config 1,178 starlark 692 system 619 fonts 452 misc 398 wifi 361 display 326 backup 212 __init__ 1,787 (imports, constants, Blueprint, 56 helpers) Two things the URL-map check could not catch, both found by running the suite: 1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one directory deeper made that resolve to web_interface/ instead of the project root. Nothing failed at import; it surfaced as ~110 tests failing with 404s and "installation script not found", because every path built from it was one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts PROJECT_ROOT/run.py exists so the next move cannot repeat it. 2. Module-attribute patching. Tests do monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route module that binds such a name by value never sees the patch. The shared code therefore stays in __init__.py rather than moving to a _common submodule -- it has to live on the module the tests patch -- and the eleven names tests patch are read back through the package (_pkg.X) instead of bound by value. Those eleven were found by AST-scanning every setattr in the test tree, not by guessing; "time" is among them, used to drive a fake clock through the second-resolution credential-backup filenames. Test changes are confined to what genuinely moved: patch targets that now name the owning route module, imports of helpers, and six tests that scan the api_v3 source as a file and now read the package directory. Full suite: 4,278 passed, 68 skipped, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(api-v3): address CodeRabbit findings from the blueprint-split review Fixes to the api_v3 package split (PR #553), one per finding verified against the actual code: - __init__.py: _redact_credentials only blanked scalar values under a credential-named key; a bare list of secrets under such a key (e.g. tokens: ["a", "b"]) passed through untouched, since the list branch recursed with no memory that its key looked like a credential. Nested dicts still walk normally (a documented, tested behaviour -- a container like secrets: {api_key: ..., note: ...} is a section name, not a value to blank outright), but any value reached under a credential-shaped key is now actually blanked. - __init__.py: the OAuth helper script's raw stderr/stdout went to logger.error unredacted (CWE-532) right next to a comment claiming this was deliberate; the HTTP response already used the existing redact_text helper. Routed the log line through the same helper. - __init__.py / starlark.py: the standalone Starlark manifest fallback (used when the plugin instance isn't loaded) read-modified-wrote manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe (plugin-repos/starlark-apps/manager.py), which already holds an flock for the same file when the plugin is loaded. Added _starlark_manifest_lock, mirroring that pattern, and wrapped every standalone read-modify-write call site in it. The app-config update route also wrote config.json and the manifest as two separate, non-transactional writes (a second, distinct finding at the same call site); config.json is now rolled back if the manifest write that follows it fails. - backup.py: restore options used bare bool() on values from the request, so {"restore_secrets": "false"} restored secrets anyway (bool("false") is True). Switched to the existing _coerce_to_bool helper already used for this exact purpose elsewhere in the package. - config.py: an automated import-rewrite mangled four user-facing validation strings and their neighbouring comments -- "Invalid start time" had become "Invalid start _pkg.time" (and likewise for "end time") in both the schedule and dim-schedule per-day validation paths. - display.py: `import _pkg.time as time_module` -- _pkg is a local alias for the package, not a real importable module, so this raised ModuleNotFoundError whenever a caller restarted an already-running display service via /display/on-demand/start, after the on-demand request was already written to cache. Fixed to `import time`. Audited the rest of the package for the same `_pkg.<module>` import mistake; every other `_pkg.` reference is a legitimate attribute read-through (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken import statement. - fonts.py: validate_file_upload's max_size_mb parameter is silently unused by that helper (it only checks filename/extension) -- the font upload route saved arbitrarily large files as a result. Added the same seek-and-check pattern already used for the sibling .star upload. - wifi.py: two ad hoc, inconsistent bool coercions. POST /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled it. POST /wifi/radio's enabled/force parsing recognized real bool and some strings but not int 1/0 (1 is True is False in Python). Factored one small _parse_bool_ish helper local to this file and used it at all three sites. Not changed: the "unknown/misspelled restore option keys default to True" half of the backup.py finding -- the file's own comment documents that a missing key deliberately means "restore everything," matching the already-existing JSON-parse-failure guard a few lines above it; only the bool-coercion defect was a real bug. Added or extended regression tests for every fix, following each area's existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed on both this branch and origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change) -- no new failures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5 * fix(api-v3): reject unknown restore option keys CodeRabbit's review of the blueprint split (#553) asked that POST /backup/restore reject option keys outside RestoreOptions' known set. The follow-up commit fixed the bool("false")-is-True bug with _coerce_to_bool but never added the key check: a typo'd or renamed key (e.g. "restoreSecrets") is silently ignored by opts_dict.get(key, True), so the flag stays at its True default and secrets get restored despite the caller's request saying otherwise -- with no indication anything was wrong. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb * fix(api-v3): address CodeRabbit findings on the blueprint split - _redact_credentials: blank scalar descendants of objects reached through a credential-owned list (e.g. tokens: [{"value": "secret"}]) regardless of field name -- the existing name-based walk only protected direct dict values under a credential key, not list items. - wifi.py: reject enabled/force/auto_enable_ap_mode values _parse_bool_ish can't recognize (400) instead of silently treating them as False, which could disable Wi-Fi or the radio itself. - Starlark manifest locking: lock a stable manifest.json.lock sidecar instead of manifest.json itself, in both the standalone route path (_starlark_manifest_lock) and the plugin path (StarlarkAppsPlugin._save_manifest / _update_manifest_safe). manifest.json is replaced by an atomic rename on every write, which swaps in a fresh inode; a lock held on the old inode does not exclude a second locker that opens the path afresh right after the rename and gets the new inode, so two writers could race despite each holding "a lock". A sidecar that no write ever touches always resolves to the same inode for every locker. Skipped as stale: the "serialize the complete manifest read-modify-write" finding at api_v3/__init__.py -- every standalone handler that calls _write_starlark_manifest is already wrapped in _starlark_manifest_lock() on this branch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): re-check reconciliation findings by the reconciler's own rules Both CodeRabbit findings on the merge commit, verified against the code first. Major, plugins.py: the stale-findings filter derived its own notion of "in config" and "on disk", and both were looser than the reconciliation module's. set(load_config()) also contains system keys, the secrets-file keys load_config() merges in, and non-dict values; and any directory holding a manifest.json counted as installed even when that manifest does not parse. Either looseness clears a finding that is still true -- and a secrets key read as a plugin is the precise bug the filter exists to stop reporting, so reintroducing that asymmetry while re-checking was the wrong way round. The two extractions now live in state_reconciliation.py as config_plugin_ids() and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys() alongside. _get_config_state() and _get_disk_state() use them too, so there is one definition rather than two that can drift. _get_disk_state() re-reads each manifest for version/name after taking membership from the shared extractor; that costs one extra small read per plugin on a path that runs once per boot. Minor, the new test: the fixture assigned api_v3.config_manager and api_v3.plugin_manager directly. Those live on a module-level blueprint singleton, so the mocks leaked into every later test that imports api_v3 -- pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr, which restores them. This is the same pollution class that made an earlier test in this session break seven unrelated ones, so it is worth getting right. Five cases added for the parity itself: a secrets key, a system key and a non-dict value must not clear an "installed but missing from config" finding, and neither an unparseable manifest nor a .standalone-backup- directory may count as installed. All five fail against the looser version. Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and CodeRabbit all pass. Codacy reads action_required on every commit of this branch including the first, so it is pre-existing and not from this work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0ab95586fb |
fix(web): say when a system action failed for want of passwordless sudo (#560)
* fix(web): say when a system action failed for want of passwordless sudo
POSTing reboot_system to a Pi returns, in full:
{"message": "Action failed; see logs for details", "status": "error"}
The cause is that the web interface runs unprivileged, and its
systemctl/reboot/journalctl calls only work once
scripts/install/configure_web_sudo.sh has granted NOPASSWD. first_time_install.sh
never invokes that script and no user-facing doc mentions it, so on a fresh
device every privileged action fails -- start_display, stop_display, the
autostart toggles, reboot, and the log viewer.
That last one closes the loop: "see logs for details" is unreachable advice
when journalctl is refused for the same reason. This is exactly the failure
src/web_interface/error_handler.py's describe_exception() was written to break,
and /system/action's exception handler was still discarding the cause instead
of using the helper the module already imports.
Two changes, no behaviour change when things work:
- The exception path now returns 'details': describe_exception(e), matching how
the other handlers in this blueprint already report.
- A failure whose stderr or exception text is sudo refusing to prompt ("a
password is required", "no tty present", "a terminal is required") reports
what to do about it, naming configure_web_sudo.sh. Unrelated failures keep
the generic message and their stderr, so a missing unit is not blamed on
sudo.
Granting the sudo rights is left alone deliberately: auto-running a script that
hands out NOPASSWD is a security decision for the maintainer, not something to
slip into an installer. Making the refusal legible is the part that is
unambiguously an improvement.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): apply the sudo hint on the on-demand start_display path too
start_display with a mode builds its own response and returns before the shared
nonzero-result path, so a recognized sudo refusal there reported only "Failed to
start display" and said nothing about the passwordless sudo that refused it --
the exact gap the rest of this PR closes everywhere else.
Raised by CodeRabbit on #560 and verified against the code before fixing: the
branch at api_v3.py:2058 does return early past the shared handler.
Three regression cases: the on-demand branch reports the sudo cause, keeps its
"Display started" message on success, and does not blame an unrelated failure on
sudo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
39f27d285d |
fix(plugins): stop reconciliation inventing plugins and telling users to delete real config (#557)
On a device running four installed, configured, working plugins, the overview
banner read:
Stale plugin config entries found: football-scoreboard, odds-ticker, data,
ledmatrix-weather, starlark-apps. Remove them from config.json or reinstall
via the Plugin Store.
Every claim in that sentence was wrong, and following its advice would have
deleted 4.9KB of working league settings. Four separate defects combined.
1. Secrets keys became phantom plugins. load_config() merges
config_secrets.json into the config it returns, and the ignore list named
only 'github' and 'youtube'. A 'data' key in that file therefore read as a
plugin id and was reported as "in config but not on disk" forever. Read the
secrets file's own top-level keys instead of hardcoding two of them.
2. The auto-fix clobbered real config. The handler for "on disk but not in
config" assigned `config[plugin_id] = {'enabled': False}` unconditionally,
so whenever detection was wrong it replaced a plugin's entire configuration
with a stub. On the reported device it only failed to do so because the
write hit EACCES. Now it refuses to overwrite an entry that already exists.
3. The banner gave backwards advice. plugin_missing_in_config ("on disk, not in
config") and plugin_missing_on_disk ("in config, not on disk") are opposite
problems, and both were rendered as "stale config entries ... remove them
from config.json" -- which is correct for the second and destructive for the
first. They are now reported separately, each with the advice that fits.
4. A stale verdict was served indefinitely. The result is a snapshot written
once per run to a status file, and a run that fails to apply a fix also
declares it will not retry. A condition that had since resolved kept being
reported for hours. The status endpoint now re-checks stored findings
against current state, dropping only what it can prove stale and keeping
any kind it cannot re-verify.
The secrets-key lookup is deliberately fail-safe: an unreadable, absent,
malformed or non-path secrets location narrows the ignore set rather than
raising. An earlier revision let TypeError escape, which the broad handler in
_get_config_state() swallowed as "Error reading config state" -- emptying the
config state and making every downstream detection wrong. The existing
reconciliation tests caught it; there is now a regression test for it too.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
aba96e25b3 |
chore: delete three functions nothing calls (#550)
src/base_classes/baseball.py _get_baseball_display_text 45 lines src/web_interface/api_helpers.py validate_request_params 22 web_interface/blueprints/api_v3.py _validate_time_range 14 Each has exactly one occurrence across both repositories -- its own definition. No decorator, no __all__, no getattr dispatch, nothing in templates or JavaScript. A fourth candidate was dropped after checking: _unshare_element_fonts in src/common/sports_shared.py looked unreferenced, but eight scoreboard plugins call SportsCore._unshare_element_fonts directly from their test_element_text_colors.py, plus their own copies at runtime. It is live API. The earlier reading came from a plugins checkout 84 commits behind main, which is a good argument for re-verifying this kind of claim against a fresh tree rather than trusting an earlier scan. Full suite: 4,265 passed, 68 skipped. Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
26769ee37f |
fix(starlark): the store authenticated with a key nothing writes (#541)
#535 restored the thirteen routes, so the store stopped answering 404 -- and still would not load. Confirmed against a running device before anything was changed: /repository/browse answers 200 with 1000 apps in 27s, so the routes are fine. Two things underneath them are not. **The store never used the token the user configured.** The three repository routes read `github_token` off config.json. Nothing writes that key -- it is not in config.template.json, no setting offers it, and it appears nowhere else in the codebase. The configured token goes to config_secrets.json as `github.api_token`, which PluginStoreManager loads and every other GitHub caller uses. So the store could never be authenticated: 60 requests/hour, on the same per-IP budget 48 installed plugins spend on update checks, while the 5000 the user had already configured sat unused. On the device, /plugins/store/github-status reported authenticated with a limit of 5000 at the same moment /starlark/repository/browse reported 60, with 18 left. The store going blank was that 60 running out. **Every failure looked identical.** list_all_apps_cached turned any listing failure -- rate limit, DNS, timeout, non-200 -- into an empty app list, and the route sent that out as `status: success`, so a rate limit and an empty repository drew the same blank grid with no error anywhere. It now returns the reason, the route answers 502 with it, and a failure is no longer cached as an empty repository for two hours. The guard for a bad response was itself a crash: _make_request catches `(json.JSONDecodeError, ValueError)` but `json` was never imported, so evaluating the tuple raises NameError and the guard written for exactly this case never ran. Reachable whenever something on the path answers with HTML -- a captive portal, a proxy page, a DNS-hijacking router. Seventeen handlers answered 5xx with no detail at all. test_no_api_v3_handler_discards_its_exception is meant to prevent that across api_v3, but it matched one exact message string, and all thirteen Starlark routes wrote their own wording. The guard now keys on the shape that matters: if it returns 5xx, it says why. The 15 pre-existing non-Starlark functions are listed as a set that may shrink, never grow. **The listing was capped at 1000 and did not say so.** The contents API truncates a directory silently; tronbyt/apps has 1075 app directories, so the store showed a truncated repository and looked complete doing it. Now listed via the git trees API, which reports `truncated`, with the contents API kept as a fallback. Not addressed: the 27-second cold load -- 1075 manifests fetched five at a time behind skeleton placeholders -- which is probably the largest part of what "does not load" feels like, and wants its own change. 25 new tests. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e23f1f45d3 |
feat(starlark,on-demand): the third-party fixes worth taking, plus a Home Assistant MQTT bridge (#538)
* feat(starlark,on-demand): the third-party fixes worth taking, plus an MQTT bridge Analysis of ant456/ledmatrix-fixes-repo, a third-party collection of patches and services built while running this project on Starlark apps under MQTT control. Its patches are whole-file copies taken against an older tree, so applying them as written would revert #523's frame pacing, #534's display() bool returns and the GitHub token masking in plugins_manager.js. Three of its claimed fixes are already in main, and its api_v3 Starlark routes are #535's. What follows is the rest -- verified against current code, and reimplemented where the patch's approach did not hold up. **On-demand display.** `pinned` reached the controller from the API, was stored on it and republished in the status payload, but never narrowed the rotation -- a pinned request still cycled every mode its plugin owns. Right for a sports plugin, whose modes are views of one subject; wrong for a plugin whose modes are unrelated, which is every Starlark app. Now honoured, and it survives a restart. Restarting while on-demand was active loaded *only* the on-demand plugin, so normal rotation had nothing to return to for the life of the process -- and a restart mid-session is routine, since that is how an update is applied. The panel came back cycling one plugin's modes with no way out but clearing the cache by hand. Every enabled plugin loads now; on-demand still resumes on its saved mode. Stop requests are exempt from the duplicate guards on purpose, so that a second click stops a mode a race left running -- which means consuming the mailbox is the only thing that ends one. It was never consumed, so the same stop was re-read and re-processed on every poll, forever. Both paths now share one compare-before-delete helper. **Starlark rendering.** `extract_schema` parsed the source with a regex, which can only see option lists written out literally: an app whose dropdown is filled from a live API call inside `get_schema()` came back empty, and the config form offered nothing to pick. Now runs `pixlet schema`, which executes the app, and falls back to the parser when Pixlet is absent, too old for the subcommand, or the app fails to run. The third-party patch replaced the parser outright and hardcoded /usr/local/bin/pixlet; this keeps the fallback and the binary search. A `|` in a config value was dropped by a shell-metacharacter filter, though the command is a list with no shell involved -- and apps do use it as a separator inside one value. The key went missing silently and the app rendered its own "not configured" screen with nothing to say why. And a 0-byte render was reported as success: Pixlet exits 0 and writes nothing when an app has no content, which read downstream as a working app drawing a black panel. **Starlark display.** `display()` ignored the mode it was called with, so a specific app could not be addressed. It now accepts `display_mode` -- which is the whole mechanism, since the controller inspects the signature before passing it. Found while there: `_select_next_app` ran only while `current_app` was unset, so with several apps installed the first was picked once and shown forever while the rest were rendered on schedule and never displayed. And `enable_scrolling` was missing, so multi-frame apps were called once per rotation slot and never advanced past frame one. **GET /api/v3/display/modes.** Every mode that can be requested on-demand, with the plugin that owns it. Nothing exposed this, so anything driving the display from outside the web UI read each plugin's manifest.json off disk and reimplemented PluginManager's fallbacks. It also triggers discovery, which is otherwise lazy and normally happens because a person opened the dashboard. **integrations/mqtt_bridge.** Home Assistant control over MQTT Discovery: a mode select, a stop button, power, brightness. Rewritten against the API rather than the filesystem, so it needs no read access to config.json and cannot drift from the web UI. paho-mqtt 2.x VERSION2, TLS, an availability topic that is also the last will, and secrets from the environment. **Two opt-in extras.** A DNS single-request unit, for glibc's parallel A/AAAA lookup stalling ~5s per name on routers that answer only the A query -- which makes any plugin calling an external API slow and Starlark apps, which have a render timeout, fail outright. And a Pixlet config editor: a script you run and Ctrl+C rather than the third-party version's always-on unauthenticated Flask service, since it stops the display for the length of a session. Neither is installed by default. Long Starlark app names now wrap instead of overflowing their card. 115 new tests across 5 files. Also unblocked test_starlark_display_contract.py, which was silently skipping wherever fcntl is absent. Whole suite: no new failures against main. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(mqtt_bridge): the five issues Codacy flagged on this branch All in the new bridge, all real: * requests floor was 2.31.0, which carries CVE-2024-35195, CVE-2024-47081 and CVE-2026-25645. Raised to >=2.33.0,<3.0.0, which is what the project's own requirements.txt already pins. * `import time` was never used. * `"mqtt_password": None` in DEFAULTS read as a hardcoded credential. It is the "no password configured" default; marked nosec B105, the convention used elsewhere in the repo. Also dropped an unused `build_app` from the display-modes test imports. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: the review findings on this PR Nine of CodeRabbit's ten, plus the CodeQL alert. The tenth is wrong and is answered below. **One bad config section blanked the whole mode list.** `/display/modes` read `full_config.get(plugin_id, {}).get('enabled')`, so a non-dict under a plugin id -- a shape DisplayController already guards, so it happens -- raised AttributeError mid-loop and answered 500 with no modes at all. Every MQTT bridge entity is built from that list. Now skipped with a warning. **The DNS scripts reported success they had not earned.** Three separate paths: `resolvconf -u` failing was swallowed by `|| true`; the systemd-resolved branch exited 0 without applying anything, so the oneshot unit recorded success while the workaround was inactive; and the installer's `|| echo` turned a failed start into "installation complete." with exit 0. All three now fail loudly. `single-request` is a glibc resolv.conf option with no resolved.conf equivalent, so on those hosts the honest answer is that it cannot be applied. A NetworkManager-generated resolv.conf is regenerated on connection changes, not only at boot, and the unit is oneshot with RemainAfterExit -- so the option can vanish mid-boot with nothing to put it back. Now detected and stated plainly rather than implied to be permanent. **`Before=` does not order a manual restart.** It only orders units already in the same transaction, so `systemctl restart ledmatrix` could bypass the fix. install_dns_fix.sh now writes a ledmatrix.service drop-in with Wants= and After=. Wants=, not Requires=: a DNS workaround failing should not stop the display. **The Pixlet editor's `--lan` is gone.** `pixlet serve` has no authentication, and a printed warning is not access control. Loopback only, with the SSH port-forward in the header where the flag used to be documented -- SSH does the authenticating and nothing is left listening. **The MQTT example config now defaults to TLS** on 8883. The installer copies it verbatim, and without TLS the broker password and every command cross the network in cleartext. A plaintext broker is still supported and documented, and the bridge warns once at startup when a password is configured without TLS. **Not taken: "the upstream Pixlet CLI has no `schema` subcommand."** Upstream tidbyt/pixlet has none, but `scripts/download_pixlet.sh` installs `tronbyt/pixlet`, whose `cmd/schema.go` is `schema [PATH]` -> JSON on stdout, built on `runtime.NewAppletFromPath`, so it does execute `get_schema()`. That is exactly what extract_schema_via_pixlet calls. A binary without the subcommand exits non-zero and falls back to the source parser, which is already covered by a test. **CodeQL stack-trace exposure: not taken either.** I removed `details` first and that broke test_web_error_detail.py::test_no_api_v3_handler_discards_its_exception, which enforces `describe_exception` across all ~75 handlers -- written because a device with failing storage answered "see logs for details" from the log viewer itself. describe_exception redacts credentials; the trade-off is the project's and is already made. Restored, with the reasoning in a comment. 11 new tests. Whole suite: no new failures against main, 4127 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
793b988d33 |
fix(starlark): toggle the app id the list published, and store a relocatable star_file (#537)
The two review nitpicks left over from #535. Both are still on main after that merge; the five findings alongside them landed with it. **The toggle could not find what the list had just shown.** `_starlark_virtual_plugins` publishes the raw manifest key as `starlark:<key>`, and `_toggle_starlark_app` passed it back through `_validate_and_sanitize_app_id`, which lowercases and rewrites every character outside `[a-z0-9_]`. An app stored as `My-App` was listed as `starlark:My-App` and looked up as `my_app`, so toggling an app the page had drawn a moment earlier answered 404. Keys written by `_install_star_file` are already sanitised, so this only shows up for manifests written by the starlark-apps plugin itself or edited by hand. `_validate_starlark_app_path` rejects traversal without rewriting, so it is the check to use here -- listing and toggling now agree on one key. The updater also uses `setdefault` rather than indexing: the app is loaded but its on-disk entry need not exist, and `_update_manifest_safe` does not catch `KeyError`, so that escaped as a 500 rather than writing the entry. **`star_file` was stored absolute.** Readers join it to the app's own directory -- `_standalone_render_starlark_app` does `app_dir / app_data.get('star_file', f'{app_id}.star')` -- so the key's default is a bare filename and an absolute value gave it a second meaning. Since `Path.__truediv__` discards the left side when the right is absolute, the manifest was pinned to whatever PROJECT_ROOT installed it, and a moved or redeployed install could not find its own file. Storing `dest.name` matches the default and stays relocatable. Read paths are unchanged, so manifests already holding an absolute path keep working. 7 new tests. Whole suite: no new failures against main, 4013 passed against 4007. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
50258635a8 |
fix(starlark): restore the API routes #330 dropped (#535)
The Pixlet install button reported "Pixlet install failed: Resource not found" -- Flask's 404 handler, because the route did not exist. #253 added thirteen Starlark routes; #330 rewrote api_v3.py and dropped all of them, along with the `starlark:<app_id>` entries that surface installed apps in the plugins list and the toggle branch that enables them. Restores all thirteen routes, the plugin-list entries and the toggle path, so Pixlet installs, the app store browses and installs, and an installed app can be managed like any other plugin. Not a straight revert. Three error paths stopped returning exception text to the caller; the manifest write moved off a shared temp filename that two concurrent writers could interleave; both dynamic importers stopped leaving half-initialised modules in sys.modules; the config update rolls back when the save fails; the toggle checks that persistence succeeded; and the path check returns the validated path instead of a boolean so callers stop re-joining the raw value. New tests no longer reach GitHub. Verified on a 256x64 Pi: Pixlet installs and runs (v0.53.1), the store lists 1000 apps, install/toggle/uninstall round-trip, and traversal and command-injection probes are rejected at every entry. 25 CodeQL alerts dismissed as verified false positives -- path-injection where traversal is blocked, and one list-form subprocess with no shell. Both classes already present on main. Full core suite: 3981 passed. |
||
|
|
a0d3e64099 |
fix: ten defects found validating the whole plugin fleet on hardware (#534)
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers Three independent fixes found while validating every plugin on a 256x64 rig. FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf ships with the core, so every plugin offering it logged "Font family 'tom_thumb' not found" (16 warnings per countdown render) and had to carry a private loader to use a bundled font. Closes #524. VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter that DisplayManager gained, so any plugin passing it died with TypeError at render time and failed every size. Nine plugins now make that call; ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other seven only passed because their scroll path was unreachable without data. Closes #525. api_v3 declared module-level config_manager/plugin_manager = None that nothing ever assigned -- app.py sets the blueprint attributes, which the other 150+ call sites use. Three sites read the decoys, so /health reported the config unreadable and the plugin system uninitialised (making "degraded" permanent and unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on every rig. The decoys are removed rather than assigned, so a bare name is now a NameError at test time instead of a silent None. The same function's first-call uptime was computed from two separate clock reads and came out negative. Closes #529. Verified on the rig: both previously-failing plugins render, the tom_thumb warnings are gone, /health reports "healthy" with all three checks passing, and /display/current reports the real 256x64. * fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is world-writable and sticky, and the display service runs as a different user from the tooling, so a leftover temp owned by anyone else became unopenable even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a sticky directory. The preview and the health check's liveness proxy then froze until someone deleted the file by hand; on the test rig that meant 23 hours of a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with cleanup on failure, matching the hardware-status write a few hundred lines above. Closes #528. _poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults the in-memory TTL to max_age -- so the first request was pinned in memory for an hour and every later poll returned that stale copy. No second on-demand request was honoured until the service restarted, while the API kept returning 200. get() already documents memory_ttl=0 for exactly this cross-process case. The consumed request is also now deleted: leaving it on disk meant a restart replayed the previous request, activated it, and ignored the one the caller had just made. Closes #530. starlark-apps returned None from display() when it has no app to show, which is the state of every install without Pixlet and of a fresh one before any app is added. The controller only skips on a boolean False, so that held a black panel for the full display_duration instead of rotating on. Closes #456 (core side). Verified on the rig: two consecutive on-demand requests with no restart between them are both activated, where the second was previously dropped in silence. * perf(harness): share one cache across a plugin's renders _instantiate built a fresh MockCacheManager for every (size, mode), and that mock is a per-instance in-memory dict, so each render was a cold start. A plugin that fetches per game or per player re-fetched everything N times over -- baseball-scoreboard at one size took 840s for nine renders where the arithmetic said ~72s, and at eight sizes it exceeded a 900s timeout. The second and later renders also never exercised the cache-hit path, which is what a running rig executes almost all of the time, so a caching regression could not be caught here. The cache is now built once per render_plugin_matrix call and threaded down. The display manager stays per-render -- the bounds checking depends on that -- so only fetched data is shared. Measured on the rig, same render counts and same goldens: tide-display 2s -> 1s (32 renders) cricket-scoreboard 10s -> 3s (24 renders) No pass/fail change across tide-display, cricket-scoreboard, clock-simple, geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages. Closes #533. * fix(scripts): run standalone plugin tests instead of collecting nothing run_plugin_tests.py discovered every plugin test file and handed the lot to pytest. Most plugin tests are standalone scripts -- module-level main() plus an `if __name__ == "__main__"` guard, signalling through an exit code -- and pytest collects zero items from those. The run printed how many files it had *found*, then "no tests ran", and exited without executing any of them. On a rig with all 44 first-party plugins that is 151 of 248 files. Files are now classified and each kind runs under the right runner: pytest for real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1 fail convention ledmatrix-plugins' own runner established (a script that wants a tty or an LED matrix is a skip, not a regression). Before: $ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos Found 1 test file(s) collected 0 items no tests ran in 0.31s rc=0 After: Found 1 test file(s) -- 0 collectable, 1 standalone script(s) 1 passed, 0 skipped, 0 failed (scripts) rc=0 Verified across three shapes: countdown (1 script), jellyfin-now-playing and pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files split 4 collectable / 7 scripts, all seven of which had never run). Closes #532. Running the flights scripts for the first time also surfaced four genuinely failing tests there, hidden by the mirror-image bug in the plugins repo's own runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465. * fix(harness): give an empty-looking mode a few frames before warning about it check_plugin's "drew nothing but display() returned X" warning fired on a single frame, rendered with force_clear=True, under a frozen clock. All three defeat a scrolling plugin, whose first frame is legitimately its blank scroll-in buffer. Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which people stop reading a warning, which matters because the true positives are real: a mode that draws nothing and does not return False holds a blank panel for its whole display duration. An apparently-empty frame is now re-driven for up to 48 more frames with force_clear=False (force_clear means "reset the scroll", so repeating it would redraw frame 1 for ever) and with the clock advancing -- freezegun's factory where time is frozen, a real sleep where it is not, since scroll position is usually a function of elapsed time. The first frame that draws content replaces the result. The clock is moved back afterwards. It is shared by every render in the matrix, so time borrowed by the probe leaked into later modes and drifted their goldens -- f1_upcoming picked up 5 spurious drifts before this was restored. Measured on the rig: empty warns check before after f1-scoreboard 42 0 48 PASS / 0 FAIL, goldens intact ledmatrix-elections 16 0 16 PASS / 0 FAIL on-air 8 8 true positive, kept nfl-draft 8 8 true positive, kept clock-simple/geochron/ 0 0 unchanged christmas-countdown 58 false positives gone, both true positives kept, no golden regressions. Cost is confined to modes that really are blank: plugins that draw immediately are unchanged (clock-simple and tide-display still 2s), while on-air -- eight deliberately blank modes -- goes to 21s. Closes #527. * fix(harness): load nested schema defaults, and merge caller config at leaf level load_config_defaults read only top-level properties. An object property carries its defaults on its children, not on itself, so everything nested was dropped -- 2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of 565. render_plugin_matrix's comment says the plugin then "behaves like a real install", which for most of the fleet it did not. _defaults_from_properties now recurses. merge_config deep-merges the caller's config onto the result so an override lands at the leaf: a shallow merge would let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard every other nhl default, which is the same class of bug being fixed here. Measured before/after across all 49 installed plugins on the rig: **no render changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact. Plugins already fall back to the same values internally via config.get(key, default), so supplying them explicitly agrees with what they were doing. The defaults really are arriving now: ufc-scoreboard 9 -> 87 defaults ledmatrix-flights 51 -> 95 masters-tournament 10 -> 51 cricket-scoreboard 22 -> 50 tide-display 12 -> 18 and hockey-scoreboard, which used to load nhl.enabled=None, now gets nhl.enabled=True with its full display_modes block. Caveat worth carrying: the eight plugins with the most nested config (soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of the 2,386 dropped defaults, 68%) could not be measured. They import src.common.sports_shared, which the test rig's core branch predates, so they fail to load there identically before and after. Re-run this comparison against a core that has that module before trusting the "nothing changed" result for them; those are exactly the plugins whose renders should change most. Closes #531. * refactor: narrow the exception handlers this branch introduced Codacy flagged the new code; it passes on other recent PRs, so the finding is mine. Four of the five broad `except Exception` clauses I added were catching far more than they needed to, which is the same shape as several bugs this branch fixes -- hello-world's TypeError sat invisible for exactly this reason. freezer() / move_to() / tick() -> (AttributeError, TypeError, ValueError) cache_manager.delete() -> (OSError, AttributeError, KeyError) The fifth stays broad and now says why: it wraps a call into a plugin's own display(), which can raise anything, and the first frame has already rendered -- so a failure there must not turn a good result into an error. Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows 7 golden drifts both before and after this branch, so it is not from these changes -- its committed goldens predate #521's 1-bit text rendering. * fix: resolve CodeRabbit review and Codacy findings on #534 CodeRabbit raised six; all six were real. The test double had drifted ahead of production. VisualTestDisplayManager accepted set_scrolling_state(frame_hold=...) while DisplayManager did not, so such a call passed every harness run and would raise TypeError on the panel -- the one failure a safety harness exists to prevent. frame_hold belongs to the change that adds it to DisplayManager (#523), so it moves there and the double matches main again. The harness swallowed exceptions from re-rendered frames. _settle_loop re-renders a mode that came back blank, to give a scroll time to draw; returning silently on a crash meant a mode that renders one good frame and then explodes was reported as passing. Recorded on result.error now, keeping the captured frame so the failure stays inspectable. starlark-apps display() returned True after _display_frame() failed, so the controller held a dead frame for the whole display_duration instead of rotating on. _display_frame now returns bool on all three paths. run_plugin_tests.py used env.setdefault for PYTHONPATH and LEDMATRIX_CORE, so an inherited value won and the subprocess imported a different core than the one under test -- ledmatrix-plugins#467 exactly. Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally. The on-demand mailbox is polled after every frame, ~125x/second on a scrolling mode, and the read is deliberately uncached, so it was that many disk reads per second to find nothing. Floored at 250ms, which is imperceptible for a web-UI click. Consuming it also deleted whatever was present rather than what had just been processed, so a request posted while the previous one was in flight was thrown away and never ran; the delete is now keyed by request_id. That narrows the window rather than closing it -- a true atomic claim needs a primitive the cache layer does not offer, and the code says so rather than implying otherwise. Codacy's 2 criticals were bandit B404/B603 on the subprocess call added to run_plugin_tests.py. Fixed interpreter, argument list, no shell; annotated with the repo's existing nosec convention. Bandit is clean on the file. Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py (4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of those fail against the pre-fix code. Full suite: 3961 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * chore: satisfy Codacy's subprocess checks on the new test runner Codacy runs Bandit and Opengrep (its Semgrep fork). The new subprocess.run in scripts/run_plugin_tests.py trips three patterns, on two different lines: Bandit B404 on the import, B603 on the call Opengrep dangerous-subprocess-use-audit on the run( line dangerous-subprocess-use-tainted-env-args on the argv line A nosemgrep applies only to its own line, so the call line and the argv line each need one; a single comment on the call covered neither rule fully. Suppression is the right answer here rather than a rewrite: the interpreter is sys.executable, the arguments are a list, and no shell is involved, so there is nothing to word-split or expand. Matches the pair the rest of the repo already uses for this shape -- permission_utils.py, plugin_loader.py, install_dependencies_apt.py. Codacy: 0 new issues, up to standards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * chore: leave visual_display_manager untouched so #523 can merge The only change this branch made to that file was a docstring, and it collided with #523's rewrite of the same method -- so #534 and #523 each merged cleanly against main but conflicted with each other. Reverted to main's text; #523 owns this method and adds frame_hold to it. The note the docstring carried ('frame_hold arrives in #523') would have been stale the moment #523 landed anyway. The parity test in #523 is what actually keeps the two signatures honest. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * chore: add the Ruff suppression nosec/nosemgrep do not cover Ruff reports S603 on the same call Bandit and Opengrep do, and none of the three suppressions covers the others. Confirmed the precondition first: path comes from discover_plugin_tests(), which globs test files inside the repo, and the call is a fixed interpreter with a list argv and no shell. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f90638a9ec |
feat(web): show available memory in Tools diagnostics (#500)
* feat(web): show available memory in Tools diagnostics System Diagnostics reported memory as used-percent plus used/total GB. Neither distinguishes a healthy board from one about to fail, because page cache counts as used and is reclaimable on demand -- a Pi can read 70% used and be fine, or read the same and be minutes from trouble. MemAvailable is the kernel's own estimate of what a new allocation can actually obtain, and it is the number that tracked the failure on a 1GB Pi 3B+: healthy running sat above 500MB, and the crash came at 73MB. By that point fork() was failing, so sshd could not spawn a session and systemd could not respawn the display, while the kernel carried on answering pings at 0% loss. Used-percent gave no warning at any point on the way there; available memory fell steadily for hours. /api/v3/system/status now returns memory_available_mb from psutil.virtual_memory().available, and Tools renders it as its own tile, coloured against the thresholds that failure implies: red under 150MB, amber under 300MB, green above. The existing memory tile is left alone -- used/total is still what you want when sizing a workload; this answers the different question of how much room is left right now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): round available memory once, so the tile agrees with itself The colour was classified from the raw value while the label was rounded, and the API sends one decimal place. At the boundaries the two disagreed: 149.6 rendered as "150 MB" in red, and 299.6 as "300 MB" in amber -- each contradicting the threshold its own colour claims to apply ("red under 150MB"). A reader checking the tile against the documented thresholds would conclude the readout was broken. Rounding once and using that number for both restores agreement. It moves those two boundary cases up a band, which does not matter: the thresholds come from a measured failure at 73MB, so which side of the line a spare 0.4MB falls on carries no information. The tile agreeing with itself does. Null handling and the thresholds themselves are unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a4a55a23fc |
fix(web): reject non-finite JSON numbers instead of raising (#497)
* fix(web): reject non-finite JSON numbers instead of raising
POST /api/v3/config/dim-schedule with {"dim_brightness": Infinity} answered
500. So did /api/v3/errors/clear with max_age_hours, and /api/v3/config/main
with multiplexing or row_address_type.
json.loads accepts Infinity/-Infinity/NaN by default -- they are not valid
JSON, but Python's parser emits them -- and Flask's get_json passes them
straight through. int(float('inf')) raises OverflowError, which is neither
ValueError nor TypeError, so validation blocks that carefully caught those let
it past and Flask turned it into a 500.
The status code was not the real damage. dim-schedule answered with
CONFIG_SAVE_FAILED and suggested "Check file permissions on config directory"
and "Check available disk space" for what was an invalid number. Every one of
these sites already had a correct 400 response written; they just never
reached it.
NaN already returned 400, because int(nan) raises ValueError. That is why this
only ever showed up for the infinities, and why it survived: the obvious test
case passes.
OverflowError is now caught alongside ValueError/TypeError at the 27 sites in
this file whose try block performs a numeric coercion. An AST sweep confirms
no int()/float() of request-derived data is left outside a block that catches
it.
Verified end to end through Flask's test client rather than by reasoning about
the parser: all four routes returned 500 before and 400 after.
Tests: five Infinity cases (which fail against the previous except tuples),
two NaN cases pinned so narrowing the tuple cannot quietly break them, and a
check that ordinary input is not rejected by the widened guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* test(web): assert 400 exactly, and prove valid input is accepted
Both review points were right, and the first is the failure mode this file
exists to catch.
Accepting any 4xx meant a 404 would have passed. Renaming one of these routes
would have left the test green while it tested nothing -- the same "looks like
coverage, points somewhere safe" shape that hid the composer injections. Now
asserts exactly 400.
Both infinity signs are exercised for every route. int() raises OverflowError
either way, but only +Infinity was in the original report, and a guard that
special-cased the sign would have passed a one-sided test.
The valid-input test previously asserted "not a 400", which did not show what
it claimed: the mocked save path fails for any input, so that assertion held
whether or not validation had accepted the value. It now gives load_config a
real dict and stubs _save_config_atomic, so the endpoint reaches its success
response and the test can assert 200 -- which only happens if the value passed
validation.
8 of the 11 checks fail with OverflowError removed from the except tuples.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
5a1f121e6b |
Stop array-item secrets being wiped, and logging them (#493)
Three review findings from #485 that I missed when addressing that PR; it has since merged, so they land here. 1. Array-item secrets destroyed by any unrelated save (data loss). remove_empty_secrets recursed into dicts but let a list fall through to the scalar branch and kept it verbatim. Lists merge by *replacement*, so the blanks the masked form posts back went straight over the stored array: stored [{"name":"a","token":"REAL-A"}, {"name":"b","token":"REAL-B"}] posted [{"name":"a","token":""}, {"name":"b","token":""}] merged [{"name":"a","token":""}, {"name":"b","token":""}] -> both credentials gone Same failure as the scalar api_key case fixed earlier, one container deeper. Lists now prune element-wise, and a list with nothing real in it is dropped so the stored one is left alone. Where one entry does change, the new merge_secrets merges by index instead of replacing. Two details the first attempt got wrong, both caught by existing tests: - An emptied dict item must stay {}, not None. ConfigManager's _strip_secrets_recursive treats a secrets list as *parallel* to the regular one ({} = "item i has no secrets"); a None makes it stop looking parallel, and it then drops the whole key from the main config -- silently deleting the items' non-secret fields too. - The incoming list's length wins. The regular config's list is authoritative about how many items exist, so preserving surplus stored entries would let the two fall out of step and make deleting an entry impossible. 2. Submitted credentials written to the journal (security). save_plugin_config logged `Full config: {plugin_config}` at INFO and `Config that failed: {plugin_config}` at ERROR. Both run before separate_secrets, so plugin_config still held the values just typed into the form. Now keys only. Swept the rest of web_interface/ and src/ for the same shape -- these were the only two. 3. Restart banner kept stale wording. showRestartPending() cleared the stored custom text but left the DOM element alone, so a config save could show the previous update's message. The default is read back from the server-rendered copy rather than duplicated in JS, so the template stays the one owner of the string. Verified: 556 passed, 1 skipped across the web suite. Mutation-checked -- reverting api_v3 fails the logging guard and the array-merge test; reverting either half of the secret_helpers change fails the unit tests. New end-to-end coverage drives the real endpoint, not just the helpers. Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9c0c0dc851 |
consolidate(web): credential exposure, secret loss, and the update path (#485)
* fix(web): stop /config/main handing out every credential it holds
The endpoint returned the raw config to anyone who could reach the port, and
this web interface has no authentication of any kind. An unauthenticated
request against a live rig returned:
github.api_token 40 chars
incoming-packages.ha_token 183 chars
jellyfin-now-playing.api_key 32 chars
ledmatrix-weather.api_key 32 chars
on-air.mqtt_password 8 chars
youtube.api_key 20 chars
youtube-stats.api_key 39 chars
A GitHub token and a Home Assistant long-lived token among them. Anything on
that LAN could read them.
The x-secret masking the plugin config endpoints use does not reach here: this
route never consults a schema, and core keys such as github.api_token have no
schema to carry the marker. Several of the fields above *are* tagged x-secret
in their plugin's schema and were still returned in full, which is what rules
out the schema route as the fix for this endpoint.
Credential-named fields are now blanked. Matching on the name is blunt, and
for a whole-config dump that is the right default: anything named like a
credential should not leave the process, and a new plugin adding a
differently-shaped secret is covered without anyone remembering to tag it.
Blanked rather than removed, and safe to blank: POST /config/main merges into
the freshly loaded config and writes only the keys it was given, so a client
that round-trips this response cannot erase a secret it never saw. The web API
suites confirm it -- 81 passing, unchanged.
On the test that matters: the first version of this suite exercised the two
helpers and nothing else, and reverting the single line that wires the
redactor into the route passed all thirty of them. A property asserted on a
helper is not a property asserted on the endpoint, and it is the endpoint that
is exposed to the network. The added test goes through the view function, and
it does fail on that revert.
This also corrects an earlier claim of mine. I reported that GET /api/v3/config
did not expose these values; that path 404s, so the check proved nothing. The
real route is /config/main and it exposed all of them.
* fix(web): stop an unrelated config edit from erasing a plugin's secret
Saving any field on a plugin's config form destroyed that plugin's stored
credential. On a rig with a weather API key, changing the city silently
emptied the key, and the plugin stopped working at the next fetch with no
indication why.
The path had no guard at any step. The config partial masks secrets before
rendering (pages_v3.py:740), so the browser posts them back blank; _parse_value
deliberately preserves "" for optional string fields; separate_secrets routes
that "" into secrets_config, which is a truthy dict; deep_merge writes it over
the stored value; save_raw_file_content persists it.
The blank does not even need the round-trip. merge_with_defaults injects the
schema's api_key default ("") into every save, so a client that never sends
the field at all still erases it. test_secret_count_message_counts_top_level_keys
was counting exactly that injected blank as a saved secret field -- the visible
edge of the bug, pinned as expected behaviour.
remove_empty_secrets() already existed for this, with seven unit tests and a
docstring describing this precise scenario ("clients will send those empty
strings back ... so that existing stored secrets are not overwritten with
blanks"). It was never wired into a call site. This wires it into both save
paths that merge into the secrets file.
A blank now means "unchanged" rather than "delete", which is the same contract
the helper's tests already describe. The cost is that a secret can no longer be
cleared by emptying the field; clearing needs its own affordance, since a
control that erases credentials as a side effect of ordinary edits is not one.
Verified by reverting the guard: the new round-trip test then fails with the
stored key read back as ''. 262 web tests pass with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop dumping the config and request headers to the journal
save_main_config logged its entire POST body and the full request headers at
ERROR on every save. The body is the configuration itself, and the headers
carry the session cookie, so a routine settings change wrote both to the
journal -- at a level that guarantees they survive any sane log filter.
The lines are leftover debug output: they say "DEBUG:" in the message while
calling logging.error, and they went through the root logger rather than the
module logger, bypassing the level configured for this blueprint.
Replaced with a debug-level line recording the shape of the request, which is
the part with diagnostic value. The local `import logging` went with them; it
shadowed a module-level import that was already there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop /config/secrets handing out every credential it holds
GET /api/v3/config/secrets returned config_secrets.json in full to anyone who
could reach the port, and this interface has no authentication. Probed against
a real rig it produced six populated credential fields: a 40-character GitHub
token, a 183-character Home Assistant token, and Jellyfin and weather API keys.
This is the second door onto the same credentials; #477 closes the first.
Masking the response alone would have been worse than the leak. The only
client fetches every secret, edits one field and posts all of them back, and
save_raw_file_content replaces the file wholesale -- so a masked GET followed
by the client's own save would write the mask over every credential the user
had not touched. That is why this was left open when the leak was found; it
needs both halves.
Read side: mask_all_secret_values(), which already existed for exactly this
endpoint -- its docstring names it -- and had never been wired to a call site.
It leaves empty values and YOUR_* placeholders alone, so a client can still
tell "set" from "not set" without being told the secret.
Write side: strip the echoed mask and blanks from the submission, then merge
onto what is stored, so "unchanged" means unchanged. The cost is that a secret
can no longer be cleared by blanking it; that wants its own affordance, since
a control that erases credentials as a side effect of saving an unrelated one
is not one.
Browser side: the token field is now left empty rather than filled from the
response. Filling it with the mask would have stored eight bullet characters
as the token the next time the user pressed Save, and filling it with the real
value is the thing being fixed. It reports whether a token is saved instead.
Verified end to end through the Flask endpoints, not the helpers. Reverting
the masking fails the leak tests; reverting the merge fails the preservation
tests; both halves are independently guarded. 278 web tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop reporting "no update" when the update check could not run
check-update returned update_available=False whenever git failed. The banner
is the only route to the update button, so a checkout git refuses to touch
looked exactly like a current one -- permanently, with nothing on screen to
act on and only a log line recording why.
The common cause is an install performed as root. scripts/install/one-shot-install.sh
clones into ${HOME}/LEDMatrix, never consults SUDO_USER, and contains no chown
at all, while its own error text suggests running the whole thing under sudo.
The result is a root-owned checkout, and on a rig this is what every git
command in it does:
fatal: detected dubious ownership in repository at '...'
including the fetch this endpoint runs. Verified on real hardware rather than
assumed.
A failed check now reports check_failed with a message the user can act on --
for dubious ownership, the chown that fixes it. The banner shows that message
instead of hiding itself, with the update button suppressed since updating
cannot work until the cause is fixed. The success path is untouched.
This does not fix the installer, which is the real cause; it stops the symptom
being invisible. The installer needs SUDO_USER handling and a chown, and its
suggestion to run as root should go.
Reverting the endpoint change fails four of the five new tests; the fifth
guards the success path and correctly does not move.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): stop the installer chmod stripping exec bits on every update
git tracks five scripts as mode 644 that first_time_install.sh then chmods to
755 (start_display.sh, stop_display.sh, the two install_*_service.sh, and
one-shot-install.sh does the same to first_time_install.sh). With
core.fileMode true, the default on Linux, git reports all five as modified
from then on, in files the user never touched.
The update button stashes local changes before pulling, so it is not blocked
by this. But it never pops that stash -- stash pop and stash apply appear
nowhere in the update flow -- so the mode change is stashed away and left
there, and the files revert:
=== file modes after the update button's stash ===
664 first_time_install.sh <- installer had made these 755
664 start_display.sh
664 stop_display.sh
664 scripts/install/install_service.sh
So every web-UI update silently strips the executable bit from the installer's
own scripts, and leaves a stash entry holding the difference. start_display.sh
and stop_display.sh stop working from the shell afterwards.
A manual `git pull --rebase` over SSH fails outright, since nothing stashes for
it: "cannot pull with rebase: You have unstaged changes". That is the likely
source of the reports, since plenty of people update that way.
Tracking the five as 755 -- what they should always have been, as the
installer chmodding them attests -- removes the spurious mode change
entirely: nothing to stash, nothing stripped, no stash entry, and manual
pulls work.
The pull also passes --autostash, for the case the code explicitly tolerates:
when the stash fails it logs a warning and pulls anyway, and that pull is what
then fails. Autostash also pops what it stashes, which the manual stash does
not.
Note that `git add -A` after `git update-index --chmod=+x` silently reverts
the index to the on-disk mode, so the modes here were set by chmodding the
files themselves.
Regression test asserts the five stay tracked executable; reverting any one
of them fails it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* fix(web): ask for the restart that makes an update take effect
The update button pulls new code and restarts nothing. There is no systemctl,
restart, reload or reboot anywhere in the 172-line git_pull handler -- it
stashes, pulls, installs changed requirements, re-removes plugins the user had
uninstalled, and returns "Code updated successfully."
Meanwhile both services go on running the code they loaded at boot. So the
display keeps rendering the old build, the web interface keeps serving the old
build, and the user is told the update worked. Nothing on screen suggests
otherwise, and the next reboot is what actually applies it -- whenever that is.
The affordance for this already exists: the restart-pending banner, raised
after main-config saves, with a Restart Now button wired to the display
service. A code update is a stronger reason to show it than a config save is.
The response now reports restart_required, and applyUpdate raises the banner
with wording for a code update rather than a config save. The banner's message
became a parameter and is persisted next to the flag, since it outlives the
page that raised it.
restart_required is only true when the pull actually moved HEAD. "Already up
to date" is a success too, and prompting after a no-op would train users to
dismiss the prompt unread.
This covers the display service, which is what the Restart Now button drives
and what users notice. The web interface still picks up its own new code on
its next restart; restarting it from inside a request it is serving is a
larger change than this one.
Reverting the flag fails the test that a pull which moved HEAD asks for a
restart. 290 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
* Mask list-shaped secrets element-wise, close two vacuous tests
mask_all_secret_values treated any non-empty list as a scalar, so a
secrets file holding
"accounts": [{"name": "a", "token": "tok-a"}, {...}]
came back as a single "••••••••". The caller could not see how many
entries existed, and the raw editor was handed a string where the file
holds an array. Recurse into lists in both _mask_value and _contains_mask.
Lists merge by replacement, not key-wise, so strip_masked_values now
drops a list outright if any element still carries the mask -- storing a
half-masked list would discard the untouched entries.
Two tests could pass without exercising what they claim to check:
- test_git_pull_resolution asserted modes only for paths git ls-files
returned. A renamed or deleted installer target is simply absent from
that output, so its mode was never checked. Assert every CHMODDED path
is tracked first.
- test_config_secrets_masking never checked the POST status. A 500
leaves the old file in place, which satisfies every assertion that
follows. Assert 200 before reading the file back.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01STMbQE4YctTacQXfbYqKuW
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
10e75b977f |
Cover the next tier of untested modules and endpoints, and fix the 43 bugs that surfaced (#459)
* test(sync): cover the display sync protocol, and fix what that surfaced DisplaySyncManager had no tests at all — it appeared in the suite only as a MagicMock() stand-in, so none of its framing, handshake, or socket handling was ever exercised. Writing that coverage surfaced three bugs. Both receive loops caught the generic Exception and immediately retried. A socket left in a bad state raises on every call, so the thread spun at 100% CPU logging the same line; the reverted-code run of the new regression test takes 24 seconds where the fixed one takes 0.2. Both now back off briefly before retrying. The follower dispatched on `data[:8] == _RAW_MAGIC or len(data) > 512`. That size threshold is not part of either wire format: a control message over 512 bytes — a hello_ack carrying a long incompatibility error, for instance — went to the image decoder and was dropped, and a raw frame under 512 bytes went to the JSON parser. Both formats are already self-describing, so dispatch on the magic prefix and treat a JSON parse failure as the legacy unmarked PNG, with the shared frame bookkeeping factored into _handle_received_frame(). _oversized_frame_warned was created on first use through getattr(self, ..., False) rather than in __init__, alone among the instance attributes. 75 tests: role parsing, the hello compatibility matrix, watchdog timeouts, both receive loops, the TCP image server's length and dimension caps and decompression-bomb guard, status shape per role, and one end-to-end loopback handshake so the wire format is exercised for real and not only against mocks. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(logos): cover LogoHelper, and stop bad downloads poisoning the cache Nothing in test/ referenced logo_helper.py, so its caching, resizing and download-fallback logic was entirely unexercised. Two bugs surfaced. _download_logo wrote response.content to disk with no size limit and no check that the bytes were an image. A logo URL is remote input, so the response chose how much went into the assets directory; worse, an undecodable one stayed there, and because load_logo() only reports the decode failure and returns None, every later call re-read the same corrupt file. The download path never retried, so a single bad response made a logo permanently blank rather than falling back to the placeholder. Cap the response, verify it decodes, and delete it if not, which lets the existing fallback in load_logo_with_download do its job. get_cache_stats() divided by self.cache_size with no guard, so a helper built with cache_size=0 raised ZeroDivisionError from what is only a stats call. 37 tests: size-qualified cache keys, LRU eviction and refresh, the four load_logo_with_download paths, download permissions and timeout, placeholder generation, and the abbreviation normalizer — including a test pinning its deliberate divergence from LogoDownloader.normalize_abbreviation, since logo filenames on existing installs depend on both behaviors staying put. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(web): cover the error and response builders, and stop dropping empty values errors.py and error_handler.py's response builders had no direct tests, though every API response passes through them. Two bugs surfaced. WebInterfaceError set suggested_fixes with `or`, so a caller passing [] to mean "I have no suggestions for this one" got the default list instead. Only None should fall back. create_success_response gated `data` on `is not None` but `message` and `metadata` on truthiness, so an explicitly-passed "" or {} vanished from the response while 0 and False survived — the response shape depended on the value. api_helpers.success_response() then re-gated metadata the same way, which is the path every api_v3 endpoint actually calls, so fixing only the inner function would have changed nothing observable. Both now use `is not None`. That wrapper also merged request timing into the caller's own metadata dict in place. A caller reusing a dict across requests would accumulate previous responses' timings; it now copies before adding. 79 tests: category inference for every error code, mapped vs fallback suggestions, the JSON shape including which keys are omitted when empty, exception-to-code inference, and the success/error builders end to end. Two behaviours are pinned as deliberate rather than fixed: an empty context stays out of the response body, and from_exception's `message` is the fixed per-code string, never the raw exception text. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(web): cover the input validators, and close three holes in them validators.py had tests for dedup_unique_arrays only; the other eight functions were untested. Three bugs surfaced. validate_image_url checked for '..' only inside its relative-path branch, so http://host/../secret passed validation while /../secret was rejected — the traversal check now runs before the branch split, which is where a safety check on the whole URL belongs. validate_file_upload lowercased the uploaded filename's extension but compared it against the caller's list verbatim, so allowed_extensions of ['.TTF'] rejected every valid .ttf file. Both sides are lowercased now. The one in-tree caller passes lowercase already, so this only widens what future callers can hand it. validate_numeric_range accepted True and False, because bool subclasses int; a boolean then compared as 1 or 0 against the range and validated cleanly. Excluded explicitly, matching how base_plugin.py already handles the same trap for display_duration. 84 tests. Two behaviours are pinned rather than changed: sanitize_plugin_config deliberately does not HTML-escape strings, since escaping at this layer would store the escaped form in config.json — the docstring said "prevent injection", which read as a promise it does not keep, and now says what it actually does. validate_font_awesome_class's second 'fa-' check is unreachable behind its own regex; harmless, so characterized rather than removed. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(api): cover wifi and registry endpoints, and fix bodyless POSTs The /wifi/* routes drive the host's real networking and the registry routes reach GitHub, and neither had endpoint-level tests. Covering them surfaced a bug affecting six endpoints. Six handlers read their body as `request.get_json() or {}`. The `or {}` says every field is optional and a missing body should fall back to defaults — but get_json() without silent=True raises UnsupportedMediaType when there is no JSON Content-Type, and it raises before `or {}` is ever evaluated. Each handler's catch-all then reported that as a 500. So POSTing with no body — what curl sends by default, and what a fetch() without options sends — failed on /plugins/store/refresh, /display/on-demand/start, /plugins/config/reset, /plugins/of-the-day/json/delete, /plugins/{id}/limits and /plugins/authenticate/spotify. The shipped UI always sends a JSON object, which is why this stayed hidden. All six now use silent=True. test_api_v3_optional_body.py covers the affected endpoints and adds a source check, since the combination of `or <default>` with a non-silent read is self-contradictory wherever it appears and is easier to catch by inspection than by exercising each endpoint by hand. Also adds test/_api_v3_test_helpers.py: the blueprint holds its managers on a module-level singleton rather than in Flask app state, so a test that mocks them leaks into every later test unless the originals are restored. The existing _make_client() does this for unittest classes; this is the pytest-fixture equivalent, for the five suites still to come. 69 endpoint tests: connect/disconnect/AP/radio including the string-aware boolean coercion these endpoints deliberately use, the radio's lockout-refusal path, registry refresh and fetch-from-URL, and a guard that WiFiManager is never constructed for real. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(api): cover the music auth endpoints, and always clean up the wrapper The Spotify step-2 handler writes a Python wrapper script to a temp file with the user's redirect URL embedded in its source, then executes it. That is the most dangerous shape in the blueprint and had no tests. The wrapper was deleted in the success/failure branch and again in the TimeoutExpired handler. Any other failure from subprocess.run — no interpreter, a fork failure, an interrupted call — reached neither, and left a world-readable temp file containing the user's redirect URL on disk. Cleanup moves to a finally block, which is what "delete this whatever happens" should have been from the start. The injection tests are the point of this file. Eight adversarial redirect URLs (embedded quotes, backslashes, newlines, triple quotes, a full `"; import os; os.system("id"); "`) are each pushed through the endpoint and the generated wrapper is parsed with ast: it must still be valid Python, the URL must still be a single string literal bound to redirect_url, and no os.system call may appear anywhere in the tree. json.dumps holds up, but nothing was checking that it does. 40 tests. Also pins that the two endpoints are not symmetrical despite the matching names — only Spotify has a two-step flow and a wrapper; YTM runs its script directly — so a later change does not "restore" a parity that was never there. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(api): cover the credentials upload, and stop it hoarding secrets The endpoint that receives the user's Google OAuth credentials file had no tests. Two bugs surfaced. The OAuth-shape check ran inside `except Exception: pass`. A JSON document that parses but is not an object — a bare 42, true, null, a list — makes `'installed' not in creds_data` raise TypeError, which the bare except swallowed, and the file was then written out as credentials.json regardless. The check now decides the outcome instead of being advisory, so anything not credentials-shaped is refused up front rather than failing later inside the calendar plugin. Every overwrite copies the old file to credentials.json.backup.<ts> and nothing removed them, so a user who re-uploaded ten times had ten complete sets of OAuth client credentials sitting in the plugin directory, indefinitely. Keep the newest five. Pruning is housekeeping, so a backup that cannot be removed logs and leaves the upload alone. 27 tests: size and extension limits, malformed JSON, the shape check, 0600 permissions on the written file, backup-on-overwrite, and pruning including the repeated-upload case that stays bounded. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(api): cover the install endpoints, and make 14 dead guards reachable /plugins/install and /plugins/install-from-url were tested only at the PluginStoreManager layer, so the route logic — the queue-versus-direct branch, schema invalidation, discovery, state and history recording — was unexercised. Covering them surfaced the wider form of the body-parsing bug fixed for the `or {}` handlers in the previous commit. Fourteen handlers read `data = request.get_json()` and immediately guard with `if not data: return 400, 'No data provided'`. That guard cannot run: get_json() without silent=True raises UnsupportedMediaType for a request with no JSON body, so the catch-all answered 500 "an error occurred; see logs for details" where the handler plainly meant to answer 400 and say which field was missing. Every one of these endpoints told a caller who simply forgot the body to go read the server logs. All fourteen now use silent=True, so the guard each author already wrote is the one that runs. This covers /config/raw/main and /config/raw/secrets among them, whose own bodyless case had the same shape. The two remaining bare reads are left alone: neither declares what a missing body should do, so there is no stated intent to honour. 31 install tests plus 17 body tests. The install pair is checked against each other rather than only individually — the same install logic is written twice, once in the queue callback and once in the fallback, so the tests assert both produce identical schema, discovery, state and history effects. They agree today; the one difference is the success message wording, which is characterized rather than changed. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(api): cover the raw config write endpoints /config/raw/main and /config/raw/secrets write whatever JSON they are given straight to config.json and config_secrets.json, bypassing the secret-separation path the rest of the config surface goes through. Given how carefully that surface keeps secrets out of config.json, the pair that skips it was worth pinning precisely. Backed by a real ConfigManager over tmp_path, so the assertions are against files on disk. 20 tests covering both routes: what lands in which file, that a raw secrets write never touches config.json and vice versa, the GitHub token reload, the uninitialized-manager and empty-body branches, and the ConfigError path that carries config_path through to the response. The bypass itself is pinned as intentional rather than changed — these back the raw JSON editor, so writing the body verbatim is the feature. The test says so explicitly, because the failure mode is someone later routing plugin config through here as a convenience and silently losing secret separation. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(api): cover backup restore and path containment, and fix restore scope Restore is the most destructive thing the web interface can do — it overwrites config, secrets, WiFi settings and fonts, then reinstalls plugins — and neither it nor the file routes beside it had tests. A malformed `options` field fell back to {}. Every RestoreOptions flag defaults to True, so a caller who asked for a narrow restore and mis-serialized the request got a full one instead, secrets included, and was told it succeeded. Valid JSON that is not an object was worse: `"null"` or `"[1,2]"` reached .get() on a non-dict and raised, so the request died as a generic 500. Both are now refused with a 400 that says what was wrong, and restore_backup is never reached. The other file routes take a filename straight out of the URL and turn it into a path — one to read, one to unlink. _safe_backup_path is the only thing keeping those inside the export directory, and it was untested. No bypass was found; the thirteen traversal shapes are pinned so a later loosening of that pattern has to argue with something. The delete route's by-name enumeration is covered too, including that a directory sharing a backup's name is not removed. 84 tests. Two behaviours are pinned as intentional: a failed plugin reinstall turns the whole restore into an error even though file restoration succeeded, and omitting `options` entirely still means restore everything — that is the documented default, and it is only the mis-serialized case that was wrong. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * ci: raise coverage floor to 52% Measured 54.45% after the Tier 1 and Tier 2 suites, up from 50%. Keeping the same two points of headroom the 45 -> 48 ratchet used. The modules this branch set out to cover: sync_manager 0 -> 97%, logo_helper 0 -> 98%, errors and error_handler 0 -> 100%, validators 0 -> 97%. api_v3 moved less in percentage terms because it is 4,341 statements, but the endpoints covered are the destructive and credential-handling ones. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(sync): probe for a free port on loopback, not every interface CodeQL flagged the ephemeral-port probe in the handshake test for binding to all interfaces. The probe only needs a free port number, so loopback is both sufficient and correct — a test should not open a port to the network to discover one. The manager under test still binds to all interfaces, which is deliberate and already marked nosec: a follower has to receive the leader's UDP broadcast. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * fix: bound the logo download, and stop malformed input reading as a fault Review findings on the coverage branch. The download size cap I added checked len(response.content), which has already buffered the whole body -- it stopped the bytes reaching disk but not memory, which was the point. A server that omits Content-Length and never stops sending would still exhaust the process. Stream it instead, counting as it arrives, into a sibling .part file that is replaced over the target only once it decodes. A transfer that dies midway now leaves nothing behind rather than a truncated logo for load_logo() to cache. The follower's control-message handler caught three exception types, but two reachable UDP payloads raise others: a bare JSON scalar makes msg.get() raise AttributeError, and an "sx" carrying a non-numeric x raises ValueError or TypeError from float(). Those escaped to the outer handler, skipping the legacy-PNG fallback and -- since this branch added a backoff there -- charging one malformed packet a 0.1s stall on the receive path. The legacy-PNG path also decoded without the dimension cap its TCP counterpart applies, so a crafted 65KB frame could force a large allocation on the render thread; both paths now share one constant. Three repo_url handlers called .strip() on client input without checking it was a string, so {"repo_url": 12345} answered 500. The credentials upload parsed the same file twice, the second time inside a bare except that a preceding parse had already made unreachable. And both raw-config handlers kept a json.JSONDecodeError arm that get_json(silent=True) had turned into dead code, collapsing "sent something unparseable" into "sent nothing" -- they now say which. Two of the new tests were not testing what they claimed. The pruning round-trip wrote ten backups inside one second, so all ten landed on the same int(time.time()) filename and overwrote each other; it never reached the limit it asserted. And the sync clock helper patched attributes on the stdlib time module, freezing time process-wide for every daemon thread earlier tests had left running. Full suite: 3352 passed, coverage 54%. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(sync): probe broadcast by sending, not by listening The broadcast check added in the previous commit bound INADDR_ANY to receive its own probe datagram, and the free-port probe did the same to pick a port. CodeQL flagged both, correctly: a test suite has no reason to open a socket the whole network can reach. Sending is enough for what the probe is actually for. An environment that refuses broadcast raises on sendto, which is the case that occurs in sandboxes and is the one worth skipping over; confirming delivery would have required the listening socket. A network that accepts the send and silently drops it still reaches the assertion, exactly as it did before either commit. The port probe binds loopback -- it only needs a number, and the manager's own bind is the one that has to succeed, with the retry loop already covering a port taken elsewhere. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * fix: keep callback faults out of the frame-decode fallback Review follow-up on the previous two commits. Widening the control-message except tuple put the callback dispatch inside it, so an _on_new_cycle() that raised ValueError, TypeError or AttributeError sent a perfectly good control packet to the legacy PNG decoder -- which reported it as an image decode error and buried the real fault. Split the two: whether the payload parses as JSON decides frame vs control message, a second guard covers reading the fields of an attacker-shaped body, and the callback fires outside both. It still cannot kill the receive thread; the loop's own handler catches it, and now says what actually went wrong. The logo download's temp file was a fixed "<name>.part". Two plugins asking for the same logo at once would interleave writes into it, publish the mixture, or delete each other's partial. mkstemp gives each download its own name in the same directory, so os.replace stays atomic. Its descriptor is adopted by fdopen before the request runs, since a request that raises before the write would otherwise leak the fd -- quietly, because load_logo_with_download swallows that. Two test fixes: the oversized-frame test replaced PIL.Image.open process-wide, the same hazard the clock helper documents, and Ruff B007 on an unused loop variable. Full suite: 3355 passed, coverage 54%. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test(sync): cover the announce loop, and reject non-finite scroll positions Three review findings from the follower receive path. Non-finite scroll x reached follower rendering. json.loads accepts the bare NaN/Infinity literals and float() accepts them as strings, so "x": NaN arrived as a real float and was stored verbatim. NaN loses every comparison the scroll code makes, so a follower given one sits on a position it can never advance past. It now raises through the existing malformed-control-message guard, which logs and drops the packet and leaves the last good position in place. _broadcast_available() only proves the host accepts sendto() for a broadcast; a network that accepts the send and drops the packet would let TestRealSocketHandshake run to its five-second deadline and fail on assertions the code did not break. The deadline now distinguishes the two: if not one packet crossed in either direction, that is the environment, and the test skips rather than reporting a protocol failure. That skip could hide a real regression in the announcing side, so TestFollowerAnnounceLoop covers it on mock sockets, where no network is involved and nothing can skip: hello carries this display's hardware config and goes to the broadcast address, heartbeats follow, an empty hardware config falls back to 32x64x1, hello is not resent before its interval, and a send failure is swallowed rather than killing the loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
5713fd20a7 |
fix(calendar): implement the OAuth and calendar-listing endpoints (#458)
* fix(calendar): implement the OAuth and calendar-listing endpoints The plugin's config advertised a three-step setup and only step 1 existed. Step 3's picker fetched /api/v3/plugins/calendar/list-calendars, which was never registered, so Flask fell through to the global 404 handler and the user saw "Resource not found" -- a message that names nothing and points nowhere. Step 2 had no endpoint at all, so even a working picker would have found no token to list with. Two routes, following the pattern the spotify and ytm plugins already use for their own auth scripts: POST /plugins/calendar/authenticate two-step Google OAuth GET /plugins/calendar/list-calendars calendars for the picker The authenticate route drives calendar_registration.py, which the plugin already ships and which was written expressly for this -- it reads a redirect URL on stdin and prints one JSON object. It takes two calls because a human has to visit Google in between; the script persists the PKCE verifier from the first call for the second, without which the exchange fails with "Missing code verifier". The listing route reads the token directly rather than shelling out again: the picker is interactive and a subprocess per click is slower than the API call it would wrap. It refreshes an expired token in place, sorts the primary calendar first, and drops entries with no id, which could not be selected anyway. Both name the plugin when it is not installed, rather than reproducing the anonymous 404 that started this. Verified against the live Google API on the dev rig: HTTP 200 with the account's real calendars. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(calendar): one input for the auth code, and a louder warning Two things reported after testing the flow. There were two boxes and no way to tell which to use. The config template's string branch dispatches widgets from an allow-list of names, and anything missing from it falls through to a plain input type=text -- so the field rendered both the widget's own box and a stray one for the same key. google-oauth is now on that list, which is all the widget ever needed to render in place of the fallback rather than beside it. And the warning that the redirect page fails to load was small grey text under a link, which is where it is least likely to be read. It is now an amber callout that leads with "The next page will fail to load. That is expected." The failure lands at exactly the moment the user has to act on it, and it looks precisely like the flow breaking rather than working. The paste box is labelled too, rather than relying on a placeholder that vanishes on focus. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(calendar): redact script diagnostics, and page the calendar list Three findings from CodeRabbit, all valid. Raw subprocess output was being returned to the client -- the script's stderr on one path, and its own error payload on another. CodeQL flagged the same line. That script handles OAuth client secrets and interpolates exceptions into its messages, so either could carry a secret or a path. Both now go to the log unredacted, where they are worth having in full, and reach the client through a redactor. That redactor already existed inside describe_exception, which only takes exceptions. Split out as redact_text: an exception is not the only thing worth returning, and a subprocess's stderr is just as capable of quoting a token. calendarList.list returns 100 entries per page by default, caps at 250, and hands back a nextPageToken when there are more. Reading one page would have hidden calendars from the picker with nothing to say the list was cut short. It now pages, asking for 250 at a time, bounded at ten pages so a malformed token cannot spin. And a test helper was a lambda where ruff wants a def. The five new tests fail against the previous commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(calendar): redact the last two raw exception interpolations CodeQL flagged four exposure paths. Two were mine and genuinely raw: the OSError from failing to spawn calendar_registration.py, which carries the interpreter path and whatever the OS chose to say, and the ImportError for the Google libraries, whose message named the missing module by interpolating the exception directly. Both now go through describe_exception, and the unredacted text goes to the log. The other two are the repo-wide pattern from PR #448 -- 67 handlers on main already return details=describe_exception(e), and these two new handlers follow it. That function is the sanitizer: it strips URL userinfo, auth headers and credential-shaped key=value pairs, collapses to one line and caps the length. CodeQL's taint tracking cannot see a sanitizer it has no model for, so it reports the flow regardless. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(calendar): announce status changes, name the paste box, drop a no-op Three more from the review, all valid. The status line is written after every async call -- the consent link is ready, the exchange failed -- and was a plain paragraph, so a screen reader was told none of it. It is a live region now. The paste box had a visible label that was never associated with it, so its only accessible name was the placeholder, which disappears on focus: precisely when the value is being pasted. The label now points at the input by id. And a conditional in the test helper returned the same value from both branches, which Ruff flags as RUF034. It was left over from making the fake page; one page is all those cases need, and TestPagination builds its own sequences. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * test(calendar): assert the accessibility relationships, not their parts The previous assertions searched for role="status", aria-live, a label `for` and an input `id` independently, so they passed whether or not those belonged together. Two attributes on different elements announce nothing, and a `for` that names something other than the input leaves it just as anonymous. Both attributes are now asserted on the status element itself, and the label and input are checked to go through the same identifier rather than merely both existing. Verified by mutation: a mismatched pair and a displaced aria-live are both caught. Reported by CodeRabbit, against tests I had written two commits earlier. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fce1fdac57 |
Add panel orientation setting for upside-down mounting (#455)
Adds a display.hardware.orientation config field ("normal" / "180")
so panels mounted upside down (e.g. to put the Pi/wiring on a more
convenient side) render correctly without custom pixel_mapper_config
edits. Composes onto the existing pixel_mapper_config as a trailing
"Rotate:180" mapper, so it stays independent of any custom mapper
string (e.g. U-mapper chain layouts) already in use.
Exposed as a "Panel Orientation" dropdown in the web UI's Display
settings, validated server-side, and documented in README and
CONFIG_REFERENCE.
Claude-Session: https://claude.ai/code/session_01FakipqMDHQLpsFjTuBdSFQ
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
44f59ede07 |
fix(web): say what actually went wrong instead of "unknown" (#448)
* fix(web): say what actually went wrong instead of "unknown"
Every failing endpoint returned "An error occurred; see logs for
details" and nothing else. That is survivable until the logs are the
thing you cannot reach: a device whose SD card was failing answered the
restart action, /system/status and /logs with that same sentence -- the
log viewer included, because journalctl could not be executed -- while
the exception underneath said
[Errno 5] Input/output error: 'systemctl'
which names the fault outright. The only endpoint that helped was
/health, and only because it happens to pass a subprocess's stderr
through. Diagnosis came down to guessing which endpoint leaked something.
Add describe_exception(), returning "TypeName: message" on one line, and
populate the `details` field that the response schema has always had and
nothing ever filled. The type alone carries information -- a bare
PermissionError says more than any generic sentence.
Exception text is not automatically safe to echo: a requests error
quotes the URL it failed on, and plugins that authenticate by query
string put their key there. Credential values are redacted while the
parameter name is kept, since knowing which credential was involved is
part of the diagnosis. Length is capped and newlines collapsed so a
parser's context cannot flood a JSON field.
Nine handlers in api_v3 bound the exception and never used it, so the
promised log entry was never written either -- "see logs for details"
was false, not merely unhelpful. Those now log with a traceback and
carry the detail. The other 60 already logged and are unchanged; they
can adopt the helper as they are touched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): redact auth headers and URL userinfo, and cover every handler
Three review findings.
The sanitizer missed two credential shapes that requests puts in its
exception text verbatim: `Authorization: Bearer <token>` and
`https://user:password@host`. Both would have gone straight into a
response. The auth-scheme name and the username are kept -- they say
which credential and whose without being the secret.
The AST test only asked whether *something* had been logged, so a
`logger.info("failed")` satisfied it while discarding the exception just
as completely. It now requires an error-level record carrying exc_info
and `describe_exception()` called on the handler's own bound exception.
Enforcing that revealed the first cut had scoped itself wrongly. I had
converted the nine handlers that logged nothing and left the sixty that
logged, reasoning their detail was at least in the journal. But
/system/status is one of the sixty, and on the failing device it told me
nothing -- the journal was exactly what could not be read. Splitting
them left most of the diagnostic surface unhelpful for the case this
change exists for, so all sixty-nine now carry the detail.
Two handlers had no bound exception name, and three passed the message
through a variable rather than a literal; both shapes needed doing by
hand. Full suite: 2383 passed, one pre-existing unrelated failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): stop reporting client errors as server faults
Werkzeug's HTTPExceptions subclass Exception, so the catch-all handler
saw them too and turned every 405, 400, 413 and 415 into a 500
UNKNOWN_ERROR. A GET on a POST-only route answered "an error occurred;
see logs for details", which tells the caller nothing and blames the
wrong side -- found while probing a device whose POST-only config
endpoints did exactly that.
Hand HTTPExceptions back as themselves, with their own status and
description. A genuine server fault still reports as one, with the
detail this branch adds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
* fix(web): redact any auth scheme, and require the detail in the response
Two review findings.
The auth-header pattern listed Bearer, Basic, Digest and Token, so
`Authorization: ApiKey SECRET` or `Negotiate SECRET` went to the client
intact. A fixed list silently leaks whatever it does not name, and
plugin APIs invent their own schemes, so match any scheme name and keep
it while redacting the credential.
The AST test accepted a describe_exception(e) call anywhere in the
handler, which a handler could satisfy by computing the detail and
dropping it before returning the generic message. It now requires the
call inside every return expression, which is where it has to be to
reach the caller.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
ee59caa577 |
Follow-ups from #441: secret-helper migration, ten more bug fixes, and coverage for every remaining untested module (#444)
* refactor(web): use canonical secret helpers in api_v3; make ConfigManager secret strip/merge array-aware
api_v3.py carried three inline nested copies of find_secret_fields/
separate_secrets (main-config save, plugin-config save, plugin-config
reset). They drifted from each other (one lacked isinstance guards) and
none supported the canonical module's array-item secrets
(accounts[].token). All three endpoints now import from
src/web_interface/secret_helpers.
Adopting the canonical behavior makes array-item secrets reachable, and
their parallel-placeholder shape ([{'token': ...}, {}] alongside the
regular list) was not survivable by ConfigManager's round-trip:
_strip_secrets_recursive dropped the whole key (losing the regular
fields from config.json) and _deep_merge replaced the regular list
wholesale on load. Both are now array-aware:
- strip removes the secret fields from each item and ALWAYS keeps the
list so indices survive for merge-on-load; whole-key secrets (scalar
lists, shape mismatches) still drop the key entirely — never leak.
- merge folds each secrets item into the config item at the same index,
skipping {} placeholders. The regular list's length is authoritative
in both directions: a user deleting an array item never has it
resurrected from a stale secrets entry (extras warn and are ignored).
api_v3's own deep_merge intentionally still replaces lists wholesale —
form posts carry complete arrays and index-merging would resurrect
deleted items; a comment now documents that.
Tests: the parity guard flips from 'exactly 3 inline copies' to 'zero,
and the canonical import must exist'; TestArraySecretStripAndMerge
covers the new strip/merge semantics incl. length-mismatch contracts;
new test_api_v3_secret_roundtrip.py drives all three endpoints through
a Flask client with a REAL ConfigManager+SchemaManager over tmp_path,
proving secrets land in config_secrets.json, config.json stays clean,
and a fresh load merges them back into the right array items.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: repair broken helper paths across display, cache, odds, logging, resolver, repos, config, validator
Nine fixes for bugs surfaced while writing coverage for previously
untested modules (plus the bool-duration quirk pinned in PR #441):
- base_plugin.get_display_duration: exclude bools from both numeric
branches — display_duration=True no longer reads as a 1-second slot;
it falls through to config, then the 15.0 default.
- display_helper: draw_error_message/draw_no_data_message called
_draw_centered_text with the wrong arguments and crashed with
AttributeError — both now delegate to draw_centered_text.
draw_scorebug_layout drew status and clock at the same y, overprinting
each other — they now share one combined top line.
draw_ticker_layout drew its text starting at x=display_width (fully
off-canvas), returning a blank frame every time — now draws at x=0;
scroll_speed stays accepted-but-unused and is documented as such.
- api_helper.clear_cache guarded on a nonexistent CacheManager.clear()
method, silently never clearing anything; it now uses the real surface
(clear_cache/delete/list_cache_files) and no-ops safely otherwise.
- base_odds_manager._extract_espn_data raised AttributeError when ESPN
sent explicit JSON nulls ("homeTeamOdds": null) — every level now
null-safes with 'or {}'. format_odds_summary gated on
is_odds_available, which deliberately ignores money lines, so
ML-only odds formatted as "No odds available" — it now gates only on
empty/no_odds data and formats money lines.
- logging_config.ContextualFormatter mutated record.msg in place, so a
second handler prepended the context prefix twice; it now formats a
copy. log_error hardcoded exc_info=True and raised TypeError when the
caller passed exc_info — now kwargs.setdefault.
- dynamic_team_resolver wrote its "shared" class cache through self,
creating instance shadows — the cache was per-instance and every
scoreboard refetched rankings. Writes now go through the class.
- saved_repositories cleaned URLs with an unanchored .replace('.git','')
that mangled URLs merely containing '.git' (my.github.io -> myhub.io);
now strips only a trailing suffix. add/remove also roll back the
in-memory list when the save fails, so memory always matches disk.
- config_helper.merge_configs shallow-copied the base, aliasing every
un-overridden nested dict into the result — now deep-copies.
- startup_validator.validate_all accumulated errors/warnings across
calls — now resets both lists per run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: cover the previously untested modules
Nine new suites plus an extension, asserting the Phase-1b fixed behavior
and pinning the quirks deliberately left alone:
- test_logging_config.py: formatters (JSON shape, no record mutation,
single prefix through two handlers), PluginLoggerAdapter precedence,
setup_logging handler hygiene and LEDMATRIX_DEBUG, log_error exc_info.
- test_startup_validator.py: exact messages, error-vs-warning split,
accessor split (load_config vs get_config), cache-dir branches with
os.access monkeypatched (root can write anything in CI), idempotence,
raise_on_errors classification precedence.
- test_config_helper.py (full): load/save round trips, dot-notation
get/set incl. silent-failure contract, post-fix no-aliasing merge,
schema validation branches, the '{id}_config' key pin, default-enabled
pin.
- test_saved_repositories.py: three load shapes, bare-list rewrite pin,
trailing-only .git strip (my.github.io regression), save-failure
rollback, type-classification case-sensitivity pin.
- test_api_helper.py: rate-limit math, cache-hit short circuit, ESPN
URL/key formats, exact User-Agent guard, retry adapter, post-fix
clear_cache against the real CacheManager surface, ttl-dropped pin.
- test_base_odds_manager.py: cache-key/URL construction, no_odds
sentinel round trip, stale-cache fallback, null-safe extraction,
ML-only formatting, is_odds_available truth table (ML-blind by
contract), config key/attr mismatch pin.
- test_dynamic_team_resolver.py: expansion/dedup/slicing, dropped
unknown-dynamic names (TOP_ substring hazard pinned), genuinely
shared class cache (second instance: zero HTTP), TTL expiry,
failure degradation without raising.
- test_display_helper.py (full): the fixed error/no-data renders,
combined scorebug top line, non-blank ticker with scroll_speed
no-op pin, composite upconversion, logo bleed positions, square
orientation pin.
- test_skin_runtime_cache.py: discovery-cache hit/invalidation
semantics (manifest mtime, .py edits pinned as non-invalidating),
sys.modules namespacing contract incl. bare-name restore and stdlib
shadowing, entry-module execute-once, API minor-version tolerance,
skin_matches_target table.
- test_sports_capabilities.py (extended): _draw_celebration_layout
executed for real (flash window, matrix-dims fallback, highlight
alternation, logo-failure isolation), _should_celebrate_for direct,
strict duration boundary, score_to_int edges, both-teams-score
precedence, expired-coalesce refire, disabled-win baseline
preservation, id-less prune.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* test: real schedule/dim coverage for DisplayController; fix two vacuous schedule tests
New test_display_controller_schedule.py drives _check_schedule and
_check_dim_schedule on a bare controller stub: same-day and
midnight-crossing windows with inclusive boundaries, global vs per-day vs
legacy-inferred modes (and dim's global-only default — no legacy
inference), per-day disabled days, invalid %H:%M fallbacks, unknown
timezone -> UTC, dim_brightness default 30, inactive-display short
circuit, and the _was_display_active/_was_dimmed transition flags.
test_display_controller.py's test_schedule_disabled and
test_active_hours patched config_service.get_config — which
_check_schedule never reads — so both asserted the init-default value
and could not fail. Rewritten on the test_inactive_hours pattern
(inject controller.config['schedule'], reset the minute gate, flip the
flag to the opposite state first so the assertion has teeth).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* ci: raise coverage floor to 48%
Measured 50% with the new suites in place (was 47% baseline when the
gate was introduced at 45); floor stays two points under measured.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
* fix: address CodeQL alert and review findings
- config_manager: the "secrets list longer than config list" warning now
interpolates only config-side data (no key name or secrets-derived
values), resolving the CodeQL clear-text-logging alert.
- base_plugin: validate_config rejects bool display_duration, matching
get_display_duration (bool is an int subclass and would otherwise pass
as a positive number).
- config_helper: merge_configs deep-copies override values in the
non-recursive branch so mutating the merged result cannot reach back
into override_config.
- saved_repositories: saves are atomic (temp file + fsync + os.replace),
so a failed write can no longer truncate saved_repositories.json.
- tests: regression cases for each fix, plus a pin that whole-item
array secrets (key[] + key[].field both marked) strip to empty {}
skeletons — no secret values can reach config.json.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
fc25a70d75 |
fix(web): make the update button work on branches without tracking (#443)
Reported from a pi whose checkout sat on a local branch:
git pull failed (returncode=1): There is no tracking information for
the current branch. Please specify which branch you want to rebase
against.
The Tools tab reported that as "Update failed; check logs for details",
which tells the user nothing they can act on, and the underlying git
message never reached the UI at all.
A branch with no upstream is easy to end up on — checking one out by
name, restoring a backup, or following a guide that names a branch — and
until now it left the update button permanently broken with no way out
except SSH.
resolve_pull_command() now decides how to pull:
- upstream set -> git pull --rebase, as before
- no upstream, origin/<branch> -> git pull --rebase origin <branch>,
then attach tracking so the next update is a plain pull
- no upstream, no remote branch -> an error naming the branch and
pointing at Switch branch
- detached HEAD -> says so, rather than failing obscurely
That resolution happens BEFORE the stash. Previously the handler stashed
local changes and then discovered it could not pull, putting the user's
work away for an update that was never going to run.
Failures now surface git's own message instead of "check logs".
Adds a branch picker to the Tools tab, backed by GET
/system/git-branches (local + remote-only) and a checkout_branch action.
Switching attaches tracking, so Pull Latest works afterwards. Branch
names are validated against a strict pattern before reaching a subprocess
argument list.
Local edits block a checkout, as they should. Rather than a truncated
one-line error, the response carries git's full list of blocking files
and a can_retry_with_stash flag; the UI then offers "Stash and switch" as
an explicit choice. Stashing is never done unasked — putting someone's
edits away without consent is worse than refusing the switch.
Verified on the pi that produced the report: on its untracked 'audit'
branch the update now returns the actionable message, git-info reports
upstream='' and can_pull=false, and an injected branch name is rejected.
27 tests build real git repositories and cover each path, including the
stash route that could not be exercised safely on the device.
|
||
|
|
31d607f6b3 |
fix(backup): make a restored device match the one that was backed up (#439)
* fix(backup): make a restored device match the one that was backed up Found by wiping a working device and reinstalling from scratch. Every problem here is invisible until you actually do that, which is why a green test suite and eleven hours of uptime had not surfaced any of them. **Four enabled plugins vanished on restore.** Weather, stocks, music and leaderboard have a registry `id` that differs from the `id` in their own manifest: the registry calls them `weather`, everything else calls them `ledmatrix-weather`. Installation already prefers the manifest id for the directory name and warns when the two disagree, so on disk, in config.json and in a backup they are `ledmatrix-weather` -- but nothing resolved that in reverse. Restore asked the store for `ledmatrix-weather` and got "Plugin not found in registry", four times, and the device came back missing four plugins the user had enabled. Registry lookup now falls back to matching `plugin_path`, which already records `plugins/ledmatrix-weather`. Renaming the published ids would have orphaned `plugin_state.json` entries keyed on the old ones. Exact id still wins, so a path that collides with another entry's id cannot shadow it. Against the live registry and a real 28-plugin install this takes unresolvable directories from five to one -- the one being starlark-apps, which is genuinely not in the registry. **Secrets could not be restored at all.** A fresh install left config_secrets.json group-readable but not group-writable, and the web interface -- which is what performs a restore -- does not necessarily run as the owner. Every other file in the backup restored; secrets failed with EACCES. Now group-writable, so the account running the web UI can put them back. **A partial restore reported "Restore had errors" and nothing else.** That is the same message whether the whole thing failed or it quietly dropped your API keys. It now names what was restored, what failed, and which plugins were not reinstalled. **ytm_auth.json was never in the backup.** It sits in config/ beside the three files that are, and is pure device-local auth: losing it silently signs the user out of YouTube Music. Backed up and restored with the wifi config, which it resembles. **Backups were written inside the directory a reinstall deletes.** config/backups/exports is destroyed by the reinstall the user was told to make it before. Exports now go beside the install, falling back to the old path when that is not writable. **The installer reboots without asking in non-interactive mode**, which the README did not mention -- easy to hit when piping the install, and alarming when a device you are installing onto disappears. Documented, with --no-reboot-prompt. Its log also claimed root:ledmatrix while printing a hardcoded group name rather than the one it used. Tests: registry resolution gets its own suite, including the collision case and third-party entries with an empty plugin_path. The existing round-trip test passed throughout this because its fixture plugin has a directory name equal to its id -- the one shape that cannot fail -- so it now carries ytm_auth too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * ci: run the backup suites test_backup_manager.py existed but was never enrolled, so the tests that should have guarded backup and restore have not run on a pull request. That is part of why the restore bugs in the previous commit reached a device: the suite was there, it just was not watching. Adds it alongside the new registry-resolution tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(backup): restore must not depend on owning the file it replaces Follow-up from testing the previous commit on real hardware, where the secrets fix turned out to be both too narrow and slightly wrong. Too narrow: config.json, wifi_config.json and ytm_auth.json are installed root-owned and group-readable exactly like the secrets file, so all four were unrestorable by the web service, not just one. `shutil.copy2` opens the destination for writing, which needs permission on the *existing file*; the web user could create files in that directory all day and still not replace them. Slightly wrong: the previous commit loosened the secrets file to group-writable. That was treating the symptom. The real error was deciding ownership from `ledmatrix.service` -- the display service, which runs as root and only ever *reads* secrets -- when the account that *writes* them is the web interface, which deliberately does not run as root. Ownership now follows the web service's user and the mode stays 640. `_copy_file` writes a temporary file alongside the target and renames over it. That needs only directory permission, so a restore no longer cares who owns the destination, and it is atomic: a crash mid-restore can no longer leave a half-written config. The destination's mode is carried across so restoring secrets does not widen them to the umask, and its owner is carried across too when the OS allows it -- only root can hand a file to another user, so a restore run by the web service keeps its own ownership rather than pretending to preserve root's. Verified on a device with all four config files set root-owned 640 and unwritable by the web user: before, every one failed with EACCES; after, the restore reports success with no errors and all four sections restored, mode still 640, root still able to read them and the web service still able to write them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(backup): address CodeRabbit and CodeQL findings on PR #439 - first_time_install.sh: verify chown/chmod succeed and the final owner/group/mode on config_secrets.json before reporting success; exit with a clear error otherwise instead of swallowing failures. - api_v3.py: replace the predictable .writetest probe with an exclusive NamedTemporaryFile to avoid a race with concurrent resolvers; log the preferred/fallback export path and OSError when falling back to the reinstall-deleted directory. - api_v3.py: mark a restore as failed when plugin reinstalls fail, even if file restoration itself succeeded, so the endpoint no longer reports HTTP 200 success on a partial restore. - api_v3.py: stringify plugin IDs before joining them into the error message so a malformed backup's non-string plugin_id can't raise a TypeError and mask the detailed response. - backup_manager.py / api_v3.py: stop putting raw exception text (originating from a user-controlled backup file) into restore results returned to the client; log full details server-side instead. Addresses the CodeQL "stack trace information exposure" alert. - test coverage: add a test for get_plugin_info() resolving a manifest id, and assert the disabled restore_wifi path also skips and omits ytm_auth.json. Co-authored-by: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d6c5f97c13 |
Test suite overhaul + fixes for the three bugs it uncovered (#441)
* ci: run the whole test tree and make the plugin-safety job assert something real The unit-tests CI job ran an explicit 24-file allowlist that had rotted: 63 of 90 test files (display, vegas, store manager, web API, web_interface) never ran on a PR. The job now runs all of test/ (minus test/plugins, which the plugin-safety job owns) so new test files are enrolled by default and any exclusion needs a visible, commented --ignore. The plugin-safety job was a green no-op: plugins/ is empty in CI, so every test skipped with 'Manifest not found'. It now renders a bundled deterministic fixture plugin (test/fixtures/plugins/ci-fixture-plugin, golden images included for all 8 default sizes) via LEDMATRIX_PLUGINS_DIR, and sets LEDMATRIX_REQUIRE_PLUGINS=1 so discovering zero plugins fails loudly instead of skipping green. The per-plugin suites document that they target dev machines with real plugins installed. Coverage is now measured and enforced in exactly one place — the CI unit-tests step (--cov=src --cov=web_interface --cov-fail-under=45, from a measured 47% baseline). pytest.ini previously declared --cov-fail-under=30 but CI always passed --no-cov, so the gate had never run anywhere; local pytest is now coverage-free and fast. Enabling the 63 unenrolled files surfaced three cases of test rot, fixed here: test_display_controller_vegas_tick.py could not collect without the hardware rgbmatrix module (now uses the emulator convention), the state-reconciliation unrecoverable-cache tests broke when production added the is_plugin_uninstalled tombstone check (bare Mock returned truthy), and test_get_system_status assumed the optional psutil dependency (now installed via requirements-test.txt and guarded by importorskip). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test: replace can't-fail tests with real assertions test_font_manager.py was 5 of 6 tests shaped as 'try: call(); assert True / except: assert True' — running in CI while unable to fail on any regression. Rewritten against the real FontManager API and the bundled assets/fonts: returned font types, cache-hit identity, distinct entries per size, default-font fallback for unknown families and corrupt files (recorded in failed_loads), BDF native-size reading, text measurement, and cache lifecycle. test_display_manager.py's test_draw_text ended in 'assert True'; it now renders onto a known-black canvas and asserts pixels were actually lit — which required un-breaking the fixture's freetype MagicMock so draw_text's isinstance check doesn't silently swallow the draw. test_display_controller.py carried a permanently-skipped test whose skip reason already declared it redundant; deleted. Both display test files now set EMULATOR=true before importing display_manager (the same convention as test_display_dirty_tracking.py) so they collect standalone instead of depending on which test module imports display_manager first. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test: cover the untested fragile logic (compatibility gate, secrets, config merges, durations, skin cards) New unit tests for pure or filesystem-only logic that previously had zero direct coverage: - test_compatibility.py: the semver install gate (parse_semver suffix handling, every range operator, TRUSTWORTHY_FLOOR behavior for cores reporting untrustworthy versions, 'more restrictive wins', and the malformed-manifest shapes that used to raise). - test/web_interface/test_secret_helpers.py: the canonical x-secret helpers — find/separate/mask/remove, array-item secrets, no input mutation, and a separate->recombine round-trip. - test/web_interface/test_api_v3_helpers.py: the module-level helpers behind the plugin config save endpoint (_is_plugin_update_available, _coerce_to_bool including the int==1 quirk, deep_merge including its shared-subtree shallowness, _parse_form_value, dotted-key-aware _get_schema_property/_set_nested_value). - test_base_plugin_duration.py: get_display_duration's full coercion ladder (instance attr -> config -> 15.0), including the bool-is-int quirk where display_duration=True means one second. - test_config_manager_secrets.py: the secrets round-trip — deep-merge on load, strip on save, group pruning, the load fast path — and two characterized sharp edges marked SUSPECTED BUG: an unreadable secrets file at save time writes secrets into config.json in plaintext, and a same-mtime-same-size content swap is served stale. - test_schema_manager_merge.py: merge_with_defaults branch behavior (None replacement vs falsey preservation, dict-vs-scalar mismatches, arrays replaced wholesale, defaults never mutated). - test_skin_system.py (extended): render_skin_card shares _render_game's 3-strike counter but never resets it on success — the asymmetry is pinned in both directions, along with card fallthrough and the disable interaction between the two paths. Suspected bugs are characterized, not fixed — each carries a comment so a future behavior change is deliberate rather than accidental. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test: add drift guards for cross-file contracts Three guard suites that pin contracts spanning multiple files, where one side changing unilaterally breaks the other silently: - test_version_comparison_consistency.py: the repo's four version comparators (compatibility.parse_semver, api_v3's packaging-based _is_plugin_update_available, store_manager update_plugin's raw string equality, skin_runtime._major) answer differently on the same inputs. A table pins each one's verdict; update_plugin is driven through its real code path to show the SUSPECTED BUGs: 'v1.2.0' vs '1.2.0' triggers a full reinstall the UI calls unnecessary, and a locally-ahead plugin gets downgraded. A pairwise-ordering check keeps parse_semver agreeing with packaging on plain X.Y.Z. - test/web_interface/test_secret_separation_parity.py: api_v3.py carries three inline copies of find_secret_fields/separate_secrets that lack the canonical module's array-item support. The copy count is asserted exact (it may only go down; new copies must import src/web_interface/secret_helpers), the missing-array-support gap is asserted so it can't grow silently, and the canonical behavior that migration will adopt is documented executably. - test_discovery_path_contract.py: the three 'where is plugin X' resolvers (PluginManager discovery, StoreManager._find_plugin_path, SchemaManager.get_schema_path) agree on the configured directory, and their divergent fallback chains are characterized. Also pins the .standalone-backup- naming contract shared by store rollback and discovery, and _resolve_skin_target's path-traversal rejection. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * test: address review feedback — fixture lifecycle, test names, ClassVar - ci-fixture-plugin: call display_manager.clear() before rendering (per plugin guidelines — the fixture should model a well-behaved plugin), add a class docstring, and document why Pillow is deliberately not pinned in its requirements.txt (core dependency; harness installs nothing). - Rename two tests whose names contradicted their assertions: test_unparseable_core_version_is_compatible -> test_unparseable_core_with_high_floor_is_blocked, and test_unreadable_secrets_file... -> test_corrupt_secrets_file... - Annotate TestGetSchemaProperty.SCHEMA as ClassVar (RUF012). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * ci: allow manual test.yml runs via workflow_dispatch Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * fix: unify version comparison, refuse secret-leaking saves, reset skin strikes on card success Fixes the three suspected bugs this PR's characterization tests pinned, flipping those tests to assert the corrected behavior: - plugins/store: ONE shared update comparator. New compatibility.is_update_available() (PEP 440 via packaging) is now used by both the web UI's update badge (api_v3._is_plugin_update_available is a thin alias) and store_manager.update_plugin's reinstall decision. Previously update_plugin used raw string equality: 'v1.2.0' vs '1.2.0' triggered a full reinstall the UI called unnecessary, and a locally- ahead plugin (2.0.0 installed, registry 1.9.0) was silently DOWNGRADED. Now equivalent spellings skip the reinstall and locally-ahead versions are never downgraded; unparseable versions still reconcile by reinstalling from the registry. - config: save_config and save_config_atomic now refuse (ConfigError) when config_secrets.json exists but cannot be loaded. Both previously proceeded without stripping, writing the merged secrets into config.json in plaintext. The shared _load_secrets_for_save() helper raises with an actionable message instead; a missing secrets file is still fine (nothing to strip), and _migrate_config's catch-all keeps boot resilient. - skins: render_skin_card resets _skin_failures on both success paths (vegas card returned, or mode renderer handled), mirroring _render_game. Transient card failures no longer accumulate across a session until they permanently disable a working skin. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh * fix: harden shared comparator edges from review - is_update_available: reject truthy non-string versions (a malformed manifest can carry a number; packaging raises TypeError on those) by surfacing the mismatch instead of raising. - store_manager.update_plugin: drop the truthiness gate around the comparator so a missing version on either side follows the shared 'no update' verdict, keeping the store consistent with the UI badge; a missing manifest still uses the reinstall recovery path. - config_manager._load_secrets_for_save: catch only expected read/parse failures (OSError/ValueError/RecursionError) so implementation bugs propagate as themselves, and log with traceback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NohXi78cwsAKtN1sCfxjUh --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
d9683e28be |
Codebase audit: fix shipping bugs, remove verified-dead code, repair doc drift, add regression guards (#438)
* fix(web): implement delete_cached so the font catalog cache actually invalidates
api_v3.py's font upload/delete handlers import delete_cached from
web_interface.cache, but the function was never defined. The surrounding
except ImportError silently swallowed the failure, so the fonts_catalog
cache entry survived uploads/deletes and newly uploaded fonts did not
appear until the TTL expired or the service restarted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(web): remove dead weather/stocks partial routes that returned 500
The partial dispatcher still routed 'weather' and 'stocks' to loaders
rendering v3/partials/weather.html and stocks.html — templates that no
longer exist since weather and stocks became store plugins. Requesting
either partial raised TemplateNotFound, which the catch-all turned into
a 500. No template or JS references these partials (the only 'weather'
hit in the front end is a plugin-store category filter option), so the
branches and both loader functions are removed; unknown partials now
fall through to the existing 404 handler.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(deps): align contradictory psutil/Flask-Limiter/freetype-py pins
requirements.txt's optional-install comment recommended psutil>=5.9,<6.0
while web_interface/requirements.txt hard-requires >=6.0,<7.0 — anyone
following the comment ends up with an unsatisfiable pair. The comment now
recommends the same range the web interface requires (all psutil APIs
used — Process, boot_time, cpu_percent, disk_usage, virtual_memory — are
stable in 6.x). Flask-Limiter gains the same <4.0 cap in both files and
freetype-py the same >=2.5.1 floor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(config): add template keys the code already reads
display.hardware gains pixel_mapper_config, row_address_type,
multiplexing and panel_type (read at display_manager.py with these exact
fallbacks — users on non-standard panels previously had no way to
discover them from the template). vegas_scroll gains
frame_based_scrolling and scroll_delay, the only two of its 27 keys the
template omitted (read in src/vegas_mode/config.py). plugin_system gains
development_mode, which the web UI reads and writes but the template
never declared.
Every added value is byte-identical to the code-side .get() fallback, so
ConfigManager._migrate_config() merging these keys into existing user
configs cannot change behavior on any installed device.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(scripts): repair broken sys.path setup in utility scripts
clear_cache.py and download_nba_logos.py pointed sys.path at a 'src'
directory relative to the script's own folder (scripts/utils/src and
scripts/src — neither exists), so both crashed on import; they now insert
the project root and import via the src package like the other scripts.
debug_web_manual.py resolved 'project root' to scripts/debug/ instead of
two levels up. fix_nhl_cache.sh is removed: it used Python docstring
syntax in a bash script and invoked clear_nhl_cache.py, which does not
exist anywhere in the repo — it cannot ever have worked in its current
location.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: correct stale file:line references and the loader-fallback contradiction
CLAUDE.md and .cursorrules disagreed about plugin-directory fallback
behavior; the code (SchemaManager.get_schema_path) probes plugins/
BEFORE plugin-repos/, and the main discovery path has no fallback at
all — both files now describe the real behavior, preferring symbol names
over line numbers so the references rot slower. REST_API_REFERENCE.md
pointed at app.py:144/:607 for mounts that live at :199/:799 and counted
92 routes where there are 94. PLUGIN_ARCHITECTURE_SPEC.md's historical
banner gains a note that its example imports
(src/plugin_system/base_classes/*_plugin.py) never shipped — the real
base classes are src.base_classes.sports.SportsCore and
src.base_classes.hockey.Hockey.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: fix broken links, phantom script references, and stale CI description
Repairs every broken relative link in active docs (targets renamed or
archived long ago: PLUGIN_DEVELOPMENT.md -> PLUGIN_DEVELOPMENT_GUIDE.md,
API_REFERENCE.md -> REST_API_REFERENCE.md, PLUGIN_STORE_USER_GUIDE.md ->
PLUGIN_STORE_GUIDE.md, plugin_docs/ dir, TROUBLESHOOTING_QUICK_START.md,
and MIGRATION_GUIDE's README link that silently resolved to the docs
index instead of the project README). Replaces commands invoking scripts
that do not exist (scripts/update_stats.py, validate_registry.py,
check_updates.py, fix_permissions.sh) with the real tooling, and
rewrites HOW_TO_RUN_TESTS.md's CI section, which described a
security-audit workflow that was never committed and a pytest workflow
'queued to land' that landed long ago as test.yml.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: complete the docs index and refresh the web interface file tree
docs/README.md's own policy says every page must be linked from the
index, yet five weren't — including the entire skin system
(SKIN_SYSTEM.md, CREATING_SKINS.md), ADAPTIVE_LAYOUT.md,
plugin-safety-harness.md and SPORTS_UNIFICATION.md. Each is now listed
in the section it belongs to, and PLUGIN_ARCHITECTURE_SPEC.md is marked
historical in the index (the doc itself already carries the banner).
web_interface/README.md's static/v3 tree showed only app.css/app.js;
it now reflects the actual contents.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: remove dead modules confirmed unused in-repo and across all store plugins
- src/common/cli.py: imports a 'ledmatrix_common' package that exists
nowhere (not in this repo, any requirements file, or the plugin
monorepo), so it cannot ever have run; its README section claimed
scripts/dev/* used it, which was also untrue.
- src/web_interface/logging_config.py: zero callers — the web app uses
web_interface/logging_config.py (a different module), and nothing
imports the src copy.
- handle_errors decorator in src/web_interface/error_handler.py: zero
call sites (the module's response helpers stay — they are used).
- ConfigManager.get_clock_config(): reads a 'clock' config key that no
longer exists anywhere; only caller was its own unit test.
Deliberately kept despite zero in-repo callers: DisplayError,
src/common/config_helper.py and display_helper.py — all documented as
plugin-facing API (docs/PLUGIN_ERROR_HANDLING.md, src/common/README.md),
and third-party plugins outside the official monorepo cannot be
enumerated. Verified against a fresh clone of ledmatrix-plugins (43
plugins): zero references to any removed symbol.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: remove manager-era NBA test files and one-off debug scripts
The four test_nba_*.py files imported nba_managers, leaderboard_manager
and odds_manager — top-level modules deleted when sports displays became
plugins — inside try/except blocks that swallowed the ImportError, so
they passed while exercising nothing. test_nba_data_structure.py and
debug_nba_api.py (a diagnostic script living in test/) made live ESPN
API calls rather than testing repo code. None were enrolled in CI.
scripts/debug/direct_fix_imports.py and check_imports.py were one-shot
artifacts that edited/inspected a hardcoded ~/LEDMatrix/web_interface/
app.py to fix an import problem solved long ago; nothing references
them.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: remove generate_report.py, which aggregates artifacts of CI jobs that do not exist
The script's only function is to merge JSON artifacts
(bandit/semgrep/pip-audit/safety/gitleaks results) produced by a
security-audit workflow that was never committed —
.github/workflows/ has no such jobs, so there is nothing for it to
aggregate and no way to run it usefully. Its siblings stay:
prove_security.py and audit_plugins.py both run standalone (verified),
and .codacy.yml stays because the Codacy service (README badge) reads it
server-side without a workflow file.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(web): load the three widget scripts store plugins already declare
time-picker.js, file-upload-single.js and plugin-file-manager.js
register widgets that installed store plugins reference in their config
schemas (countdown uses x-widget: time-picker and file-upload-single;
of-the-day uses plugin-file-manager), but base.html never included the
scripts. plugin_config.html renders such fields as an empty container
that polls LEDMatrixWidgets.get(...) on a 50ms loop forever, so those
plugin config fields appeared permanently blank. The audit initially
flagged these files as dead code; the monorepo cross-check proved the
opposite — they were unreachable, not unused.
example-color-picker.js (the documented custom-widget example) gains an
explicit warning that including it in base.html would shadow the
built-in color-picker widget.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore: drop the legacy youtube block from the secrets template
No code reads a top-level youtube secrets key: the youtube-stats plugin
receives its API key namespaced under its own plugin id (declared via
x-secret in its config schema), like every other store plugin. The key
survives only in state_reconciliation.py's non-plugin-key exclusion set,
which stays — existing installs still carry the key in their generated
config_secrets.json, and the exclusion prevents it from being
misclassified as a plugin config. New installs simply stop being asked
for a YouTube API key they have nowhere to use.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* chore(deps): remove packages nothing imports, declare direct imports, move mypy to test deps
Removed from requirements.txt: python-socketio, python-engineio,
websockets, websocket-client — zero imports anywhere in this repo, and
the one store plugin that needs Socket.IO (ledmatrix-music) declares it
in its own requirements.txt, which the plugin store installs. Removed
the same quartet plus timezonefinder, geopy, google-auth-oauthlib,
google-auth-httplib2, google-api-python-client, unidecode, icalevents,
python-dateutil, flask-wtf and the werkzeug pin from
web_interface/requirements.txt — all leftovers from the deleted built-in
weather/calendar/music displays (flask-wtf was doubly dead: app.py
explicitly disables CSRF and sets csrf=None). scripts/
install_dependencies_apt.py, which mirrors these lists for the
first-time installer, drops the same packages.
Added: urllib3 (imported directly in four core modules), jinja2 and
markupsafe (imported directly in pages_v3.py) — previously reachable
only as transitives. mypy moves from runtime requirements to
requirements-test.txt.
Verified in a fresh venv: all four requirements files co-install, pip
check is clean, the full CI-enrolled suite (907 tests) and a Flask boot
smoke pass with the trimmed dependency set.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* refactor: single canonical DateTimeEncoder
src/cache_manager.py and src/cache/disk_cache.py each defined an
identical DateTimeEncoder (datetime -> ISO-8601). The disk_cache copy is
the only one actually used for serialization; cache_manager now
re-exports it instead of defining a twin, so the two can never silently
diverge. Import compatibility is preserved — from src.cache_manager
import DateTimeEncoder still works and is the same class object.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs(code): document deliberate duplicates instead of merging them
The audit surfaced several near-duplicate implementations that turned
out to be either deliberate forks or behaviorally different — merging
any of them would risk changing behavior on installed devices, so each
now carries an explicit comment stating the relationship:
- VisualDisplayManager: headless fork of DisplayManager; header now
lists the ~15 mirrored methods and warns that DisplayManager changes
must be mirrored.
- normalize_abbreviation: LogoDownloader's version (called directly by
nine scoreboard plugins) replaces filesystem-unsafe characters;
LogoHelper's strips spaces. Logo filenames on existing installs
depend on both behaviors staying put.
- The two PluginTestBase classes: the shipped one is plugin-author
API, the repo's own richer harness lives in test/plugins/ — now
cross-referenced.
Also verified (no change needed): ConfigManager's backup/rollback
methods genuinely delegate to AtomicConfigManager, and SportsCore
already delegates _read_bdf_native_size to FontManager.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: add a unified configuration reference
There was no single place documenting what lives in config.json —
display.* keys were scattered across README sections, vegas_scroll lived
in ADVANCED_FEATURES.md, and dim_schedule, display.double_sided,
sync.follower_position, plugin_system.development_mode and the four
newly-templated hardware keys were documented nowhere. CONFIG_REFERENCE.md
now lists every template key plus the code-read-only keys, each with
type, default, and the code location that reads it, and explains the
secrets file's plugin-id namespacing. Linked from the docs index.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: bring the README's feature tour into the plugin era
The Core Features section still presented clock/weather/sports/stocks/
music displays as built into the project, when all of them are store
plugins installed from the ledmatrix-plugins monorepo — only
starlark-apps and web-ui-info ship in this repo. The intro now says so
(the showcase itself is unchanged; those are real displays available in
the store). The display_durations reference drops its built-in-calendar
example in favor of plugin-id keys, and the Configuration section links
the new CONFIG_REFERENCE.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* docs: archive the custom-icons status report, cross-link config docs, document assets/
PLUGIN_CUSTOM_ICONS_FEATURE.md was a 'What Was Implemented' status
report duplicating the actual guide (PLUGIN_CUSTOM_ICONS.md) — moved to
docs/archive/ per the docs index's own policy. The overlapping
plugin-config docs keep their content but PLUGIN_CONFIG_ARCHITECTURE.md
now states up front which doc is canonical for which purpose.
assets/README.md is new and load-bearing: assets/stocks, weather,
news_logos and broadcast_logos have zero references in this repo's code,
which makes them look deletable — but store plugins (ledmatrix-stocks,
ledmatrix-weather, news, odds-ticker) resolve those exact paths at
runtime against the install directory. The README records that evidence
so a future cleanup doesn't break installed plugins.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* test: add regression guards for the bug classes fixed in this PR
Three lightweight static checks, all enrolled in CI's unit-test
allowlist along with the new web-cache test:
- test_template_targets.py: every literal render_template() target must
exist (would have caught the weather/stocks partial 500s at commit
time).
- test_widget_scripts.py: every widget JS file must be script-included
in base.html or explicitly allowlisted with a reason (would have
caught the unloaded time-picker/file-upload-single/plugin-file-manager
widgets), and allowlisted files must NOT be included (prevents the
example widget from shadowing the real color-picker).
- test_doc_links.py: relative markdown links in active docs must
resolve (docs/archive/ exempt).
Each guard was verified to fail against the pre-PR tree and pass now.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SXb4mKcAkVaxkeTb3YnAdr
* fix(deps): restore the werkzeug version floor
Commit
|
||
|
|
83f20b64fe |
Give plugins access to device-wide config, and a global scroll frame rate (#424)
* Give plugins access to device-wide config, and a global scroll frame rate
The sports scoreboards read `getattr(self, 'global_config', {})` to find a
shared scroll frame rate, but nothing ever set that attribute: the loader
constructs plugins with only plugin_id/config/display_manager/cache_manager/
plugin_manager (plugin_loader.py:671), `global_config` appears nowhere in
src/, no plugin manager assigns it, and BasePlugin has no __getattr__ to
synthesize it. The lookup always returned {}, so the ten scroll_display.py
copies that thread target_fps through to ScrollHelper could never fire on any
core. There was also no global target_fps to find -- the only one in the
template is display.vegas_scroll.target_fps, which is Vegas-scoped.
Adds the missing half:
- `BasePlugin.global_config` resolves the full config via
plugin_manager.config_manager, then cache_manager.config_manager, then {}.
Same order the sports timezone helpers already use. Exceptions are swallowed
to debug so an unreadable config can never stop a plugin loading, and a
non-dict result is rejected rather than handed to callers that will .get()
it and feed the result to numeric code.
- A top-level `target_fps` (default 100), exposed on the General tab and
validated 30-200 on save to match ScrollHelper.set_target_fps -- which
clamps silently, so a rejected save reports a value that would otherwise
appear to save and then behave differently.
The property has a setter deliberately. news, stock-news, ledmatrix-stocks,
ledmatrix-elections, ledmatrix-leaderboard and nfl-draft all assign
`self.global_config = config.get('global', {})`; without a setter that raises
"property has no setter" and those six plugins stop loading. Reproduced, then
pinned with a test.
target_fps is also kept out of the `is_general_update` key list: that branch
treats a missing web_display_autostart as an unchecked box, so counting a
target_fps-only POST as a General save would silently switch autostart off.
Verified end to end: config.json -> BasePlugin.global_config ->
scroll_display's existing block -> ScrollHelper.target_fps 120 -> 100, with no
plugin-side change needed. Suite 1441 passed; the 4 failures
(test_display_dirty_tracking, test_web_api::test_get_system_status, two in
test_state_reconciliation) are pre-existing and reproduce identically on a
clean tree.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ
* Address review findings on target_fps validation and config resolution
- Reject floats and bools before int() in the target_fps save path. A JSON
body can carry them, where int(90.5) silently stored 90 and true stored 1.
Form posts send strings, so '90.5' already failed in int().
- Assert the template's target_fps is 100, not merely an int, so the
documented default is actually pinned.
- Empty-config precedence: keeping the `and config` check deliberately, now
spelled out in the comment and covered by a test. Both managers default to
the same config/config.json, so falling through cannot pick up a different
file's settings; treating {} as an answer would instead return {} when the
first manager simply hasn't loaded yet, silently disabling every setting
read through the property -- the failure this property exists to fix.
Suite 1446 passed. The float-rejection test was checked to fail without the
guard.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
5b45f35888 |
Vegas mode: reclaim dead space and pace the rotation (#423)
* Vegas mode: reclaim dead space and pace the rotation On a wide panel Vegas mode spent much of its time showing black. At 50px/s on a 512px display, one display width of blank is 10.2 seconds, which makes several long-standing behaviours expensive: - ScrollHelper prepended a full display width of black as an "initial gap", charged once per cycle — 10.2s of black at the start of every rotation. - Plugins without get_vegas_content() are captured off a full-display canvas, so their blank margins entered the ticker too. Measured: of-the-day drew 35px of "No Data" on a 512px canvas (92% blank), youtube-stats 142px of content with 185px of black either side. Only the scroll_helper path had any trimming. - Cycle transitions deliberately pushed a blank frame and then recomposed synchronously: 84ms at best, 4.8s at worst, every millisecond of it black. - buffer_ahead doubled as the cycle size, so a 21-plugin install showed 3 plugins per cycle and took ~7 cycles to come around. - separator_width was applied between every image rather than at plugin boundaries, so a per-row ticker like the F1 scoreboard (116 images, which it renders 4px apart internally) got a 32px chasm between each row — and the width budget didn't count those gaps, so the plugin quietly occupied far more of the panel than intended. Changes: - src/vegas_mode/geometry.py: numpy column-ink primitives shared by the trimmer and the audit tool, so the number reported is the number acted on. A Python per-column loop over a 17,000px strip is far too slow for the render path. - PluginAdapter trims every content path, not just scroll_helper. Only outer edges are cropped: interior blank columns are the plugin's own layout (logo left, score right) and closing them would corrupt the design. A plugin on a non-black background is inherently unaffected. - ScrollHelper.create_scrolling_image takes an explicit lead_gap, still defaulting to display_width so the many standalone-ticker callers are unchanged. Vegas passes lead_in_width (default 0). - Cycle end holds the last rendered frame instead of blanking, turning the recompose into a brief freeze rather than the panel switching off. - plugins_per_cycle (default 6) is split from buffer_ahead, which goes back to being only a prefetch low-water mark. - max_plugin_width_ratio (default 3x display width) caps one plugin's share of a cycle. Overflow is deferred, not discarded: a rotation offset advances each fetch so later rows appear on subsequent cycles. Single oversized images are cropped at a blank column so the cut misses glyphs. - Composition groups images by plugin: rows are joined by intra_plugin_gap (default 8) and separator_width applies only between plugins. The width budget now counts those gaps. - Plugin data updates no longer run on the Vegas render path. All new settings are user-configurable in Display -> Vegas Scroll, including min/max cycle duration and dynamic duration, which previously existed in code but were reachable only by hand-editing config.json. Measured with scripts/dev/vegas_audit.py on a 512x64 panel: mean ink coverage 42.7% -> 69.4% fully blank 5.9% -> 0% reads as empty 13.6% -> 0% worst blank stretch 4.8s -> 0s full rotation 414s -> 123s plugins per cycle 3 -> 6 Note the metric choice: a "fully blank" scan (>=95% black viewport) reported only 0.4% and badly understated the problem, because two full-width segments with mid-canvas content never fully blank the viewport — they hold it at ~28%. window_coverage_stats grades every viewport position by how much ink it carries, which is what tracks perceived dead time. Known remaining: cycle transitions still freeze ~3.5s while the next cycle is fetched. Fixing that needs background prefetch, which is deferred because the fallback-capture path mutates the shared display_manager.image and racing it against the render loop risks torn frames. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Drop unused Optional import from the vegas audit script Flagged by Codacy (F401). Any, Dict and List are all still used. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Align Vegas API bounds with validate(), fix audit config plumbing Both from review feedback on #423. The web API's accepted ranges disagreed with VegasModeConfig.validate(), which is what actually gates Vegas starting: scroll_speed 1-100 -> 1-200 (a slider value of 150 returned 400) separator_width 0-500 -> 0-128 target_fps 1-200 -> 30-200 buffer_ahead 1-20 -> 1-5 The three loose ones were the dangerous direction: the value saved with a 200, then VegasModeCoordinator.start() failed validation with only a log line, so the ticker silently never ran. The UI already matched validate() in all four cases, so the API was the odd one out. test_vegas_api_bounds_match_validate parses the numeric_fields map out of api_v3 and asserts every bound against validate(), plus that validate() accepts both endpoints and rejects just outside them, so these cannot drift apart again. That test immediately caught a missing upper bound on min_plugin_width, now added — unbounded it would drop every segment and leave a blank ticker. Separately, vegas_audit.py constructed PluginAdapter without the config, so it fell back to VegasModeConfig() defaults and would report trimming and width-budget behaviour that differed from the user's config.json. It now passes the loaded config exactly as the coordinator does. This is the same class of drift the explicit lead_gap and grouping arguments already guard against. Output is unchanged on a rig whose config matches the defaults. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Vegas mode: render plugins narrower, space rows by measured separation Trimming reclaims blank margins but cannot compact a layout that genuinely spans the display — a five-column forecast, a progress bar drawn at 100% width, a stat block with the panel's whole width between its elements. Those need the plugin to make different layout decisions, which means telling it the screen is narrower while it renders. DisplayManager.render_size() presents a smaller logical canvas for the duration of a Vegas content fetch, reusing the same _LogicalMatrix indirection double-sided mode already relies on so plugins see a consistent size from every accessor. Plugins that size themselves from matrix.width need no changes at all; one that wants to be explicit can read the new BasePlugin.get_vegas_render_width(). Width is a percentage so a single setting travels across panel sizes: vegas_scroll.render_width_pct globally, or vegas_width_pct in an individual plugin's config. Measured on a 512x64 panel with real data: ledmatrix-weather 1536px -> 576px (forecast becomes narrow cards) youtube-stats 353px -> 199px (2% blank left, so genuinely compact) geochron 453px -> 153px (ink density rises to 100%) ledmatrix-flights 950px -> 740px The youtube-stats figure is the clearest evidence the layout itself changed rather than being cropped: at full width the content had to be trimmed from 512px to 353px, whereas at 40% it arrives with almost no blank to reclaim. Row spacing is now measured rather than added. A flat gap gets it wrong in both directions at once — content drawn flush to its own edges ends up nearly touching (reported for recent sports scores, which sat 8px apart), while content already carrying wide margins gets pushed even further out. separation_gap() measures the blank each pair already has and adds only the shortfall, up to min_content_separation (default 24). intra_plugin_gap stays as a floor applied regardless. Two tests shipped in the previous commit encoded the old flat-gap arithmetic and are updated to the measured semantics, including one renamed to reflect that zero intra_plugin_gap alone no longer butts rows together. Also fixes a real bug found while testing: the harness display manager had no render_size(), and because the adapter catches broadly that surfaced as "no content" rather than an error, silently dropping five plugins. Added the context to VisualTestDisplayManager for parity, and _render_at() now degrades to a no-op on any display manager lacking it, so a third-party or older harness loses the narrowing rather than the content. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Vegas mode: end cycles before the wrap, keep the width budget honest Three fixes, the first a regression from lead_in_width defaulting to 0. get_visible_portion wraps: once scroll_position + display_width passes the end of the strip it fills the right of the frame from the *head* of the same strip. So the final display_width of travel showed the cycle's first plugin re-entering on the right while its last plugin exited on the left, and the recompose that followed replaced both at once. On a 512px panel at 50px/s that was 10.2s of two plugins on screen at once, ending in a hard cut — reported as the ticker "switching mid-scroll" from F1 to news. That used to be invisible because the strip began with a full display_width of blank, so the wrapped-in region was black. Removing that blank (it was 10s of dead panel per cycle) exposed the wrap. Cycles now end one display width earlier, before any wrapped content appears, clamped for strips no wider than the display so they don't complete instantly and spin the recompose loop. Verified on hardware: a 3936px strip now completes at 68.5s, exactly (3936 - 512) / 50. Second, auto_trim=False also skipped the width budget, which is an unrelated concern — turning off margin cropping should not let one plugin hold the panel for minutes. Seen in the field: the F1 scoreboard contributed 116 images and 14,848px untouched, giving a 33,821px cycle (11 minutes of content). The budget now applies regardless of trimming; with it restored that cycle is 6,362px. Third, the budget accounted for row gaps using the flat intra_plugin_gap while the compositor had moved to measured separation, so it under-counted by up to (min_content_separation - intra_plugin_gap) per row and a many-row plugin overran its cap. Both now use the same separation_gap() rule, and a test asserts the composed block fits the budget end to end rather than trusting the two paths to agree. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Fix IndexError in find_blank_cut when the cut lands on the image edge A cut position after the last column is legitimate — _crop_to_budget asks for min(start + budget, img.width), which equals the width whenever the remaining strip is shorter than the budget. find_blank_cut clamped target to width but then walked leftwards starting at target itself, so ink[width] raised IndexError. Caught on hardware: it killed the ledmatrix-stocks fetch, and because _fetch_plugin_content catches broadly that surfaced as the plugin silently contributing nothing for the cycle. Only reachable on the second or later pass of the rotating window over a single oversized image, which is why the existing tests missed it — they all exercised the first pass, where start is 0 and start + budget is comfortably inside the image. Added TestRotationAcrossMultipleCycles, which walks the window round several times and asserts content is never lost, plus direct coverage of find_blank_cut at and beyond the image edge. Both bounds now stop at width - 1 so neither direction can index past the end. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Only cut oversized segments at real gaps between items The width-budget crop snapped to the nearest blank column, and in rendered text the gap between two characters is a single column. So a cut routinely landed inside a word: the cycle showed "Wednesda" and the orphaned "y" turned up as a lone floating letter in the next cycle, positioned after whatever plugin happened to precede it. Measured on the clock-simple segment to confirm: its blank runs are [1, 1, 1, 1, 1, 8, 8] — five single-column letter gaps, every one of which find_blank_cut would happily have chosen. Cuts now only land in a run of at least min_cut_gap blank columns (default 6), which excludes letter spacing while still finding the gaps plugins put between items (the stocks ticker uses 32px, baseball 48px). Where no boundary falls inside the budget the cut waits for the next one and overruns, because splitting an item is worse than a slightly long segment. Continuous content is treated differently on purpose: an image with no internal gaps is a map or a chart, where any column is as good as another, so it is still cut to the budget exactly. The gap rule protects discrete items; letting a solid image escape the cap in its name would be wrong. blank_runs() is vectorised — 48ms for a 17,000px strip, against seconds for a per-column Python loop. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Hold capture_mode for every plugin render, not just narrowed ones The native content path only entered capture_mode when it was also narrowing the canvas, so at full width — which is every plugin without a vegas_width_pct override, i.e. most of them — a plugin calling update_display() while building its Vegas content wrote straight to the hardware. That is a visible flash mid-scroll, and it lines up with the flash reported at cycle transitions, when several plugins are fetched back to back. Suppression is now unconditional; the narrowing context stays separate because it is already a no-op at full width. Both contexts are reached through helpers that degrade to nullcontext when the display manager lacks them. That matters more than it looks: the adapter's handlers are deliberately broad, so an AttributeError from a missing context does not surface as an error — it surfaces as the plugin contributing nothing. Making the call unconditional without this turned 44 tests red for exactly that reason, all of them reporting lost content rather than the real cause. The test double now provides capture_mode and render_size too, so tests exercise the real contexts instead of silently taking the degraded path. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Vegas mode: one continuous strip instead of swapping cycles A cycle used to be a discrete strip that got replaced: motion stopped, every pixel was substituted at once, and the next group started with the viewport already full. That is the freeze, the flash and the jump. The strip is now extended rather than replaced. ScrollHelper gains append_content(), which adds items on the right without touching scroll_position or total_distance_scrolled, so motion continues and the next group simply arrives from the right. Because completion is measured against total_scroll_width, extending also defers completion — there is no longer a cycle boundary to see. drop_scrolled_prefix() reclaims what has gone past, keeping the strip bounded however long Vegas runs (observed 5,000-11,000px against an unbounded strip otherwise). It shifts total_distance_scrolled and total_scroll_width together so the completion arithmetic is unchanged, and refuses to run while the viewport is wrapping: wrapping reads the head of the strip into the right of the frame, so trimming the head there would visibly change the picture. A test caught that. Groups are prepared off the render thread. The constraint is that the canvas and the matrix proxy are process-wide mutable state, so narrowing or capturing through them from another thread would corrupt the frame the render loop is pushing. get_content() therefore takes offscreen_only: the background thread uses only paths that avoid the canvas, and anything needing it is marked and picked up on the render thread. That puts the expensive work (native renders of leaderboard and baseball cards, seconds each) in the background and leaves the cheap work (display capture, 40-600ms) in the foreground. DisplayManager's capture flag is now thread-local. As a shared flag, a background capture would have suppressed the render loop's own frame pushes for its duration, freezing the panel precisely when the point was to avoid a freeze. Canvas-bound plugins are drained one at a time rather than as a batch: six at once held the render thread for 1.75s. Drains are also spaced by two seconds while the lookahead is healthy, since taking them back to back turns one long stall into a run of short ones. When the strip is genuinely running short the throttle is ignored, because content matters more than smoothness there. Measured on hardware: zero cycle-complete swaps, drains landing 2-4s apart, lookahead holding at 1,200-3,500px, no errors. Set continuous_scroll false to restore the swap behaviour; the old path is intact. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Pace the Vegas frame loop adaptively: 31.5 -> 78.7 fps The loop slept a fixed frame_interval on top of however long the frame took, so at a measured 31.6ms per frame a flat 8ms of that was pure idle — a quarter of the budget spent not rendering. It now sleeps only the remainder of the budget. Measured on hardware: 31.5 fps to 78.7 fps sustained, with CPU going *down* from 150% to 127%. Scroll speed is unchanged at 49.9px/s against a configured 50, because motion is derived from elapsed time rather than frame count — this buys smoothness, not speed. Worth recording what the bottleneck was not: the per-frame render path measures 0.34ms in total (0.18ms for the numpy slice, 0.17ms for the dirty-tracking digest), which is a theoretical 2900 fps. Optimising any of that would have been wasted effort. The frame was idle, not busy. Also nices the prefetch thread. Its work is PIL and numpy that releases the GIL, so the scheduler can act on the priority, and without it the prefetch competes for the same cores as the render loop. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Sub-pixel scrolling: motion at the frame rate, not the pixel rate With integer positioning the number of distinct frames per second equals the scroll speed in px/s, however fast the loop renders. Measured at 50px/s and 78.7fps, 36% of frames were byte-identical: the extra frames cost work and bought no motion, and what was left was 50 discrete 1px steps a second. Two things were wrong with the pre-existing sub-pixel support. get_visible_portion never consulted sub_pixel_scrolling — it always took the integer path, so the flag and _get_visible_portion_subpixel were dead code. And that implementation needed scipy.ndimage.shift, which is not installed on the target devices (HAS_SCIPY is False there), so it would not have interpolated even if reached. Verified both: positions 1000.0 and 1000.5 produced identical frames either way. Blending is now wired up and implemented with numpy. Two details make it affordable: slice cached_array directly instead of building two PIL images only to convert them straight back (the naive version measured 15x the integer path), and use fixed-point uint16 multiply-add rather than float32, which suits the Pi's cores and gives finer weighting than the panel can resolve. Result 0.939ms against 0.237ms — 0.70ms added per frame, a 1065fps ceiling. Measured on hardware: 81.2 fps with blending on, against 78.7 with it off, so no cost within noise — and every frame is now a distinct position rather than one in three being a repeat. The trade is a slight horizontal softening of text, since each frame blends two positions. Set smooth_scroll false for maximum crispness. Also benchmarked and cleared as non-issues: extending the strip costs 9.4ms on an 11,000px strip and trimming 2.5ms, both under one frame at this rate. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Add overflow handling: keep ordered content whole instead of rotating a window The width budget split any oversized plugin by advancing a window each cycle. That is right for interchangeable items — news headlines, odds, stock prices — but wrong for ordered content: a league table showed ranks 1-6, then resumed at 7 two rotations later, which reads as out of order and out of context. Nobody needs rank 23 in a ticker; they need the top of the table, every time. overflow_mode chooses between them: rotate — advance a window each cycle so everything is seen eventually (unchanged default) truncate — always show the start and drop the rest, keeping ordered content coherent. Records no window state, so every pass starts at the top. Per-plugin vegas_overflow overrides the global setting, since one install has both kinds of plugin. Also adds per-plugin vegas_max_width_screens, so content that must stay whole can be given more room — or uncapped with 0 — without lifting the cap on every ticker. Applied on the test rig: f1-scoreboard and ledmatrix-leaderboard set to truncate, and baseball given 4.5 screens because it was showing 8 of 9 games when the whole slate needed only a little more room. Verified: F1 now reports "the first 10 of 116 ... the rest are not shown", baseball has dropped out of the budget log entirely, and stocks, odds-ticker and stock-news still rotate. Also corrects the crop log, which claimed "window advances next cycle" unconditionally and so misreported truncated crops. A test now pins the behaviour behind the message: truncate must leave no offset recorded. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Stop Vegas mode showing last night's games as if they were live A game that was live in the evening was still being drawn as live the next morning. Two faults combined to freeze plugin visuals indefinitely. PR #291 added a call to plugin_adapter.invalidate_plugin_scroll_cache() so a plugin's own cached scroll image would be rebuilt from fresh data. That method was never implemented. hot_swap_content() wraps the call in a broad except, so every hot swap has raised AttributeError and been swallowed silently ever since — which is why the visuals it was meant to keep fresh never were. Continuous scrolling then removed the only path that reached it at all: should_recompose() and hot_swap_content() are called from the non-continuous branch of run_frame(), and continuous_scroll defaults to True. So on a default install the pending-update flags were set by the update tick, never consumed, and grew without bound. Together these froze content completely, because refetching is not enough on its own: the sports plugins' get_vegas_content() regenerates only "if the cache is empty", so take_next_group() kept receiving the same picture however often it asked. Fixed by: - Implementing invalidate_plugin_scroll_cache(). It covers both layouts — a helper directly on the plugin (stocks, news, odds-ticker) and one owned by a scroll-display manager (the sports scoreboards, which is the shape that produced this bug) — and clears cached_image and cached_array together, since the array is the image's numpy mirror. - Adding StreamManager.invalidate_pending_updates() and calling it from the continuous branch. It only drops the caches; the plugin recomposes when it next comes round in the rotation. process_updates() is wrong here: it refetches synchronously and merges into the active buffer that continuous mode bypasses, and hot_swap_content() rebuilds and repositions the whole strip, which is the freeze-and-jump this mode exists to avoid. Tests assert the fix rather than the implementation: 14 of the 17 new tests fail without it. Includes the wiring itself, since the regression was a call that was simply absent, and a check that the scroll position is untouched so this cannot regress into the swap's visible jump. All Vegas suites pass (355 tests). test_display_controller_vegas_tick.py still cannot be collected off-device for want of rgbmatrix, identically with and without this change. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ * Fix two CodeRabbit-flagged test assertions in vegas density tests test_prepared_group_is_used_without_refetching had a tautological final assertion; now checks stream.calls directly. test_no_partial_letter_at_either_edge required both crop edges to be blank, but the left edge here is always the crop's start position with no lead-in gap in word_strip, so it legitimately carries ink — only the right edge is an actual cut and needs the check. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KEZK1P1Q1fu5pcuVrkrCFZ --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
e2acbfb566 |
fix(web): don't validate double-sided settings when the feature is disabled (#422)
* fix(web): don't validate double-sided settings when the feature is disabled
Saving anything on the Display tab failed with a 400 when double-sided
mode was off:
Double-sided copies (2) must divide chain length (3) evenly
The Display form posts every field in one request, including
double_sided_copies (default 2) and double_sided_axis, whether or not
the Enabled checkbox is ticked. The server block was gated only on "is
any double-sided field present in the payload" — it wrote
ds_config['enabled'] but never read it. A user with chain_length: 3 and
the untouched default copies: 2 was locked out of saving any display
setting at all: brightness, GPIO slowdown, Vegas, sync.
Gate the checks on the enabled flag:
- Divisibility against chain_length/parallel is hardware-relational and
only runs when the feature is on.
- Structural checks (copies parses as an int in 2..8, axis in the
whitelist) still 400 when enabled; when disabled they drop the value
and leave the stored one untouched rather than rejecting the save.
The runtime already gated correctly (_resolve_double_sided returns None
when disabled), so nothing there changes.
Also in the Display tab:
- Hide Copies / Split Axis until Enabled is ticked, mirroring the Vegas
Scroll pattern. Hidden rather than disabled, so the fields keep
submitting and the server still sees an 'off' state to persist.
- Fix the save toast: the form's handler read xhr.responseJSON, a jQuery
property that doesn't exist on a native XMLHttpRequest, so it was
always undefined and every save reported a green "Display settings
saved" — even the 400s. Parse responseText and use the real status.
Tests: two existing double-sided tests asserted 200 on payloads that the
divisibility check (added later, in #373) turns into 400s; the first now
supplies matching hardware values and the second passes as written now
that a disabled save skips the check. Added coverage for the reported
regression, for bad values while disabled, and for the check still
firing when enabled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016aUvzTXpEYhNEWYQvFqQU9
* fix(web): treat only 2xx as a successful display save
Review feedback on #422.
- showDisplaySaveResult tested `xhr.status >= 400` for failure, so a
network error — which reports status 0 — was waved through as
"Display settings saved". That's the same class of false-success bug
this branch set out to fix. Test the 2xx range instead, and let a
response body refine a successful verdict without overturning a
failed one.
- Annotate the _copies_fits_hardware helper, matching the annotated
helpers already in api_v3.py.
- Cover the vertical divisibility branch: chain_length 2 would divide
evenly, so only parallel 3 can produce the rejection, which pins the
branch to the right hardware dimension.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016aUvzTXpEYhNEWYQvFqQU9
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
3872a68ff7 |
Surface plugin update availability on the plugins page (#421)
* Surface plugin update availability on the plugins page
The plugin manager page already showed each installed plugin's version and
had an Update button, but nothing told users an update actually existed —
they had to guess. The store manager already compares the installed
manifest version against the registry's latest_version for its reinstall
decision; this surfaces that same signal in the UI.
- api_v3 /plugins/installed now returns `latest_version` (from the registry
cache, no extra network call) and an `update_available` flag computed by a
new semver-aware helper `_is_plugin_update_available`. A locally modified
plugin whose version is ahead of the registry is not flagged.
- The installed-plugin card shows "vX.Y.Z available" next to the installed
version, and the Update button becomes emphasized ("Update to vX.Y.Z" with
a gentle pulse) when a newer version is published — mirroring the app's own
update banner styling.
- Added tests for the helper and the endpoint fields.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012EW8yDbsk8EceDpq6Hjgqj
* Catch InvalidVersion specifically in plugin update comparison
Address review feedback: the version comparison caught a blind `except
Exception` (ruff BLE001). Split the two failure modes and catch each
specifically — ImportError for a missing `packaging` (a core dependency)
and InvalidVersion for an unparseable version string — while preserving
the existing "surface the mismatch" fallback behavior. Added a test for
the unparseable-version path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012EW8yDbsk8EceDpq6Hjgqj
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
cdf03fb107 |
Add skin system: user-installable visual overlays for sports scoreboards (#419)
* Add skin system: user-installable visual overlays for sports scoreboards Skins restyle a scoreboard's live/recent/upcoming rendering while the host plugin keeps doing data fetching, scheduling, caching, live priority, and vegas mode — the anti-fork alternative for users who only want a different layout. - src/skin_system/: ScoreboardSkin API, SkinContext (canvas + adaptive layout + logo/font helpers), discovery/loading runtime with API major version gating and per-skin module namespacing - src/base_classes/sports.py: _render_game() seam at the three _draw_scorebug_layout call sites; skin-first with built-in fallback, 3-strikes session disable, slow-render warning; per-mode skin config - scripts/validate_skin.py: headless multi-mode/multi-size validator with bundled per-sport fixtures (no hardware or network needed) - skins/example-classic-baseball/: working reference skin - Web UI: served-schema Visual Skin dropdown (validation never enum-restricted, so uninstalled skins can't invalidate configs) and GET /api/v3/skins - Store: registry entries with type "skin" install to skins/ - docs/SKIN_SYSTEM.md (architecture), docs/CREATING_SKINS.md (author guide incl. Claude Code prompt), view-model contract locked by tests Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LrCusPasy1qeUN5anK3aA1 * Address review feedback on skin system - skin_runtime: cache the entry module so the 2nd/3rd load of the same skin (live/recent/upcoming hosts) doesn't re-execute it with unbound sibling aliases; rebind cached sibling modules to their bare names around entry import and restore prior bindings after; include per- manifest mtimes in the discovery cache fingerprint so in-place skin updates are picked up - sports.py: count render_skin_card exceptions toward the 3-strike session disable - store_manager: validate skin ids (pattern + resolved-path containment in skins/), reject registry/manifest id mismatches, and stage+validate downloads in a temp sibling before replacing an existing skin - schema_manager: leave the schema untouched when the configured skin value is a per-mode mapping (a string dropdown could overwrite it) - validate_skin.py: reject non-positive sizes and non-object --options at parse time; support --output-dir outside the repo; type annotations - example skin: validate accent_color once at load with logged fallback - fixtures: pregame 0-0 scores in football/hockey upcoming fixtures - api /skins: rely on the self-invalidating discovery cache instead of force_refresh - docs: valid JSON manifest example, load_logo caching semantics spelled out, language ids on fenced blocks - tests: view-model contract test now exercises the real extractor; regression test for repeated same-skin loads with sibling modules Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LrCusPasy1qeUN5anK3aA1 --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
6a9d8014e5 |
Tag plugin logs structurally and surface the active plugin in System Logs (#418)
- get_logger() now returns a PluginLoggerAdapter when given a plugin_id,
so every plugin log call is stamped with plugin_id automatically instead
of only calls that explicitly passed extra={'plugin_id': ...}. This makes
the "[Plugin: x]" prefix reliable in the journalctl-backed log stream.
- display_controller publishes the currently active mode/plugin to the
shared cache whenever it changes, exposed via a new
GET /api/v3/display/current-status endpoint.
- System Logs page: adds a "Now showing" banner backed by that endpoint, a
plugin filter dropdown (populated from parsed log lines), a plugin badge
per log entry, and fixes log parsing to handle the short-iso timestamp
format journalctl actually returns (the old regex only matched syslog
timestamps, so level/plugin extraction silently never ran).
Co-authored-by: Claude <noreply@anthropic.com>
|