mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
c6701ac00d4c16a726d2da80f9ca023c3832823b
19
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7f06cc9c3b |
feat(vegas): keep live games in the ticker by default (#699)
* perf(timing): say which render-thread work a late frame followed The soak already says how often a moving frame reached the panel late, but not what the render thread was doing just before it. Vegas does two kinds of work there between frames -- building its strip (compose, extend) and, with live elements, patching changed pixels into it -- and deciding whether either is affordable needs their own numbers. - FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame. Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind; aggregate() still takes frames without ops. The file schema is unchanged. - Vegas tags compose and every strip extension (with the bytes it copied). - frame_soak prints an "after work" table: frames, late %, freezes and MB moved per kind, only when something tagged its work. - render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes / --patch-every / --patch-where (in-place column writes, as a live element update does) and --extend-every-screens / --extend-width (append + trim on a fixed cadence that holds the strip's width). No runtime behaviour changes: this is the measurement gate for live Vegas elements. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): note the frame-op attribution and bench modes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * perf(scroll): build the strip's PIL image only when something reads it Every Vegas strip extension rebuilt ScrollHelper.cached_image from cached_array in full, twice (append, then trim), on the render thread: Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a Pi 4 (measured on ledpi), about two thirds of an extension's render-thread cost. Nothing on the frame path reads the image's pixels; every frame is cut from the array. cached_image is now a property. append_content and drop_scrolled_prefix defer it; the first read builds it from the array it started with and keeps it only if the strip has not changed meanwhile, so a sync push racing an extension cannot leave a stale image cached. Assigning cached_image stores exactly what was assigned, as before. has_strip() says whether there is a strip without building its image; the helper's frame path, Vegas and the adapter's scroll-cache invalidation use it. The strip is also no longer held in memory twice. In Vegas the image is now built only by a multi-display sync push. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): live elements -- a plugin API for content that changes while it scrolls Vegas bakes each plugin's pictures into one strip, so a card already on its way across the panel keeps what it showed when it was drawn. This adds the API and bookkeeping for content that can be updated in place; the worker that redraws and swaps it follows separately. No shipped plugin implements the hook yet, so nothing changes for users. Plugin API (core 3.8.0), all no-ops by default: - BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version, live, refresh_hz)]: named, fixed-width pieces of Vegas content. - BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free redraw for content that changes with time. - BasePlugin.notify_vegas_data_changed(): data that lands outside update(). - src/plugin_system/vegas_elements.py (VegasElement, re-exported from base_plugin). Core: - PluginAdapter asks a plugin that implements the hook for elements on the background fetch only (under its lock, on its own canvas); every other path keeps get_vegas_content(). Live elements are pinned (padded with content_padding, never trimmed), tagged with their key, digest and data epoch in Image.info so the existing cache and group plumbing carry them unchanged, and untagged if a width budget crops them. - RenderPipeline records where each live element lands (ElementRecord), in absolute strip columns a trim does not move; the block-start arithmetic is shared with the STATIC markers. - PluginManager update listeners (add/remove_update_listener, notify_data_changed): told the moment update() completes, not at the next ~4s Vegas poll. The coordinator uses one to move each plugin's data epoch on. - vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval, live_lead_screens; per-plugin core-owned vegas_live. Live elements are off under multi-display sync, in swap mode and with offscreen_prefetch off. - scripts/check_plugin.py checks the element contract (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub is a working example. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): live elements update in place while they scroll One background worker (src/vegas_mode/live_worker.py) redraws a plugin's live elements when its data epoch moves on (update listener) or on their refresh_hz, nearest the screen first, and hands changed pixels lock-free to the render thread, which copies them into the strip between frames (RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most four patches or two screens of bytes a frame, no drawing or locks there. The worker takes over group prefetch once a live element is placed, runs inside the render gate, and is supervised. Update tick 1s while live elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md describes what was built and why SegmentStrip was not needed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(sports): live Vegas cards for the scoreboards (shared layer) One live element per game, drawn only when what the card shows changes, so a score changes on a card already crossing the panel. The shared part, so each scoreboard adopts it in a few lines: - src/common/sports_vegas.py: game_key, game_fingerprint (the whole game dict, frozen: no drawn field can be missed), dedupe_games, VegasCardCache, StickyOdds (odds a live poll left out stay drawn), finished_games / with_finished_games (a game that just went final keeps its card, after its league's live games; one a heuristic only judged over keeps its live state, so a tied end of regulation never shows FINAL early). - SportsScrollDisplay.make_vegas_renderer() is the override point; build_vegas_elements() and SportsScrollDisplayManager .get_vegas_elements_for() do the rest. A card's version includes its teams' ranks, which the renderer draws from the rankings cache. - SportsLiveSharedMixin._record_finished_game() / finished_games_snapshot(): held for FINISHED_GAME_TTL after it leaves the live list. A sport that does not implement make_vegas_renderer keeps its ordinary Vegas content, so no scoreboard changes until it opts in. scripts/render_plugin.py --vegas renders a plugin's Vegas block as the ticker lays it out, and --timeline stacks it at successive moments as the ticker would update it in place; the join is now render_pipeline.join_plugin_rows(). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): keep live games in the ticker by default display.vegas_scroll.live_in_ticker now defaults to true: through a live game the marquee keeps running and the live scoreboard takes extra turns in it -- its cards updating in place while they scroll -- instead of the ticker giving way to the full-screen scoreboard. The new default would reach nobody on its own: every existing config holds an explicit false copied from the template (there was no control for it), and the template merge only adds missing keys. ConfigManager therefore turns a stored false on once, with a backup, and records live_in_ticker_migrated so a false chosen afterwards stays. The marker is never in the template. A "Keep live games in the ticker" checkbox under Vegas mode sets it. Tests that pin the full-screen takeover now say live_in_ticker=false. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sports): a default _determine_game_type on SportsScrollDisplay render_vegas_card looked the method up with getattr and a None default, which static analysis (Codacy) reports as calling something that may not be callable. The base class now has the default -- the card type from the game's state -- and the plugins that define their own override it as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: review follow-ups on the shared live-card layer - The reused Vegas renderer always gets the current rankings, empty included, so ranks cleared since are not kept drawn. - render_plugin.py: --timeline refuses --no-live (a timeline shows live elements changing), --timeline/--no-live need --vegas, and the Vegas paths create the output's directory like the display path does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f4bda50710 |
feat(vegas): live elements update in place while they scroll (#697)
* perf(timing): say which render-thread work a late frame followed The soak already says how often a moving frame reached the panel late, but not what the render thread was doing just before it. Vegas does two kinds of work there between frames -- building its strip (compose, extend) and, with live elements, patching changed pixels into it -- and deciding whether either is affordable needs their own numbers. - FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame. Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind; aggregate() still takes frames without ops. The file schema is unchanged. - Vegas tags compose and every strip extension (with the bytes it copied). - frame_soak prints an "after work" table: frames, late %, freezes and MB moved per kind, only when something tagged its work. - render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes / --patch-every / --patch-where (in-place column writes, as a live element update does) and --extend-every-screens / --extend-width (append + trim on a fixed cadence that holds the strip's width). No runtime behaviour changes: this is the measurement gate for live Vegas elements. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): note the frame-op attribution and bench modes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * perf(scroll): build the strip's PIL image only when something reads it Every Vegas strip extension rebuilt ScrollHelper.cached_image from cached_array in full, twice (append, then trim), on the render thread: Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a Pi 4 (measured on ledpi), about two thirds of an extension's render-thread cost. Nothing on the frame path reads the image's pixels; every frame is cut from the array. cached_image is now a property. append_content and drop_scrolled_prefix defer it; the first read builds it from the array it started with and keeps it only if the strip has not changed meanwhile, so a sync push racing an extension cannot leave a stale image cached. Assigning cached_image stores exactly what was assigned, as before. has_strip() says whether there is a strip without building its image; the helper's frame path, Vegas and the adapter's scroll-cache invalidation use it. The strip is also no longer held in memory twice. In Vegas the image is now built only by a multi-display sync push. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): live elements -- a plugin API for content that changes while it scrolls Vegas bakes each plugin's pictures into one strip, so a card already on its way across the panel keeps what it showed when it was drawn. This adds the API and bookkeeping for content that can be updated in place; the worker that redraws and swaps it follows separately. No shipped plugin implements the hook yet, so nothing changes for users. Plugin API (core 3.8.0), all no-ops by default: - BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version, live, refresh_hz)]: named, fixed-width pieces of Vegas content. - BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free redraw for content that changes with time. - BasePlugin.notify_vegas_data_changed(): data that lands outside update(). - src/plugin_system/vegas_elements.py (VegasElement, re-exported from base_plugin). Core: - PluginAdapter asks a plugin that implements the hook for elements on the background fetch only (under its lock, on its own canvas); every other path keeps get_vegas_content(). Live elements are pinned (padded with content_padding, never trimmed), tagged with their key, digest and data epoch in Image.info so the existing cache and group plumbing carry them unchanged, and untagged if a width budget crops them. - RenderPipeline records where each live element lands (ElementRecord), in absolute strip columns a trim does not move; the block-start arithmetic is shared with the STATIC markers. - PluginManager update listeners (add/remove_update_listener, notify_data_changed): told the moment update() completes, not at the next ~4s Vegas poll. The coordinator uses one to move each plugin's data epoch on. - vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval, live_lead_screens; per-plugin core-owned vegas_live. Live elements are off under multi-display sync, in swap mode and with offscreen_prefetch off. - scripts/check_plugin.py checks the element contract (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub is a working example. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): live elements update in place while they scroll One background worker (src/vegas_mode/live_worker.py) redraws a plugin's live elements when its data epoch moves on (update listener) or on their refresh_hz, nearest the screen first, and hands changed pixels lock-free to the render thread, which copies them into the strip between frames (RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most four patches or two screens of bytes a frame, no drawing or locks there. The worker takes over group prefetch once a live element is placed, runs inside the render gate, and is supervised. Update tick 1s while live elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md describes what was built and why SegmentStrip was not needed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7804ea8f69 |
feat(update): stable/beta update channel; stable follows release tags (#684)
Adds auto_update.channel: stable follows the newest vX.Y.Z release tag (detached HEAD; pre-releases and other tags ignored), beta follows main as before. Nothing ever moves a device backwards: a checkout newer than the newest release keeps following main (or stays put when detached) until a release contains its commit. Legacy configs migrate to stable when they reach a release. Update Code, the weekly updater's preflight, and the verifier's rollback (back to old_ref: branch or detached release) all honour the channel. General tab Update Channel select, GET/POST /api/v3/system/update-channel, release-aware Overview banner and Tools git panel. New installs default to stable. Rig fix (ledpi): /system/check-update reports update_available: false when the channel's action is none (a detached HEAD newer than the newest release), matching Update Code; the Tools panel no longer calls every detached HEAD "a release". Merged with main through #687 (heartbeat verifier, #683 login, #688 plugin_catalog, #685 Tailwind build). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7ab6fb1aff |
refactor(web): read plugins through a PluginCatalog; only the display runs them (#688)
The web process built its own PluginManager and loaded plugins into itself: store installs and updates loaded or reloaded a web-side copy, and config saves and enable/disable called on_config_change, on_enable and on_disable on it. None of that reached the panel, and /plugins/installed reported runtime state from those copies. - Add PluginCatalog (src/plugin_system/plugin_catalog.py): manifests, directories, display modes, installed version, schema and config reads, with no way to run a plugin. app.py and both blueprints use it; the plugin_manager blueprint attribute is gone. - Remove every lifecycle call from the web routes. Config changes already reach the display through ConfigService (on_config_change) and the enabled-set reconcile. - Health and metrics readers move to api_v3.health_tracker / resource_monitor. /plugins/installed reports loaded/state/error_info as null (the display does not publish them) and enabled by the display's rule. - Store install, update and uninstall answer restart_required when the running display will not pick the change up by itself (display_restart_required). The restart banner follows the flag via window.noteRestartRequired instead of the /config/main URL heuristic; /config/main now sends restart_required: true. - The one remaining in-process import of plugin code (Starlark helper modules, oauth_flow action scripts) goes through _import_plugin_code_in_web_process() until a web-entry contract. - /plugins/installed reports vegas_participation (from #682) from the user's setting or the manifest, with vegas_participation_source; when only the plugin's code decides it, null with source 'runtime', since the web process no longer has plugin instances to ask. - Check & Update All keeps its restart flags when the final list refresh fails, and asks for a restart when an enabled plugin's first request got no answer and the re-sent one found it up to date. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
e3c85cece6 |
feat(web): optional web login and API tokens, off by default (stacked on #674) (#683)
Optional web login, off by default: a device that sets no password behaves exactly as before. Set under General > Security; then every page and API route needs a session login or an API token (Authorization: Bearer). Loopback, the Wi-Fi setup flow in AP mode, static files, captive-portal probes and a reduced /api/v3/health stay open. Secrets live in the web_auth section of config_secrets.json and no API returns them. scripts/reset_web_password.py turns login off. Stacked on #674. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
09103a8a7d |
refactor(web): split api_v3/plugins.py by area (#658)
* refactor(web): split api_v3/plugins.py by area
web_interface/blueprints/api_v3/plugins.py (3,285 lines) becomes:
- plugins.py: installed list, enable/disable, plugin actions
- plugin_store.py: install, update, uninstall, store, saved repositories
- plugin_config.py: config get/save, schema, reset
- plugin_assets.py: asset uploads and plugin static files
- plugin_health.py: health, metrics, limits
- plugin_operations.py: operation history, state reconciliation
- plugin_calendar.py: calendar credentials and auth
Pure move: all 44 functions and 38 route decorators are byte-identical
(checked with ast), URLs and endpoint names are unchanged (url-map test).
Each module imports only what it uses. Tests and config.py that reached
into plugins.py for moved names now import from the new module.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): keep exception text out of calendar responses; annotate moved code
The split made scanners report existing findings in the moved code as new:
- CodeQL: the calendar auth and calendar-list routes returned exception
text (redacted, but still derived from the exception). Both now log the
exception and return a fixed message pointing at the log.
- MD5 in the asset upload only makes a filename unique: usedforsecurity=False.
- pickle reads/writes the calendar plugin's own OAuth token (as before):
annotated. Token-status labels and a log line naming the secrets path are
false positives: annotated with the repo's nosec/nosemgrep convention.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): name uploaded assets with SHA-256 instead of MD5
The hash only makes an uploaded image's filename unique. Codacy flags MD5
even with usedforsecurity=False, and SHA-256 does the job as well; existing
files keep their names.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): keep the redacted exception detail in calendar errors
Reverts the calendar part of
|
||
|
|
6cfcf2e384 |
fix(web): plugin dir resolver in routes, nmcli AP detection, daemon config reload, upload safety, BDF preview (#655)
* fix(web): plugin dir resolver in routes, nmcli AP detection, daemon config reload, upload safety - Route plugin lookups (installed list, update, recorded version, config form, web UI pages) through the plugin manager's resolver so plugins in ledmatrix-<id> directories work. - Captive-portal detection also sees the nmcli fallback AP (cached). - WiFi monitor daemon re-reads wifi_config.json when its mtime changes. - Drop the AP check in disconnect_from_network that could never fire. - LED status file per WiFiManager; config path falls back to this checkout. - BDF font preview via src.common.bdf_font. - Asset uploads validate every file before saving; metadata and calendar credentials written atomically; no absolute path in the response; asset delete answers 400 for a missing body. - Coerce string booleans in plugin toggle, on-demand start and AP force. - SSE broadcaster clears its thread handle before exiting. - start.py log filter handles every exc_info form. - Cleanups: unused plugins/fonts partial work, duplicate backup catch-alls, raw-config error helper, update-route tidy, redundant imports. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): request BDF font previews now that the server renders them The Fonts tab skipped the preview request for .bdf files because the server used to refuse them; /fonts/preview now draws BDF with the shared loader. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): take the update route's plugin directory from a directory listing CodeQL flagged the path built from the request's plugin_id (the id was already validated with safe_path_component, which CodeQL doesn't model; the same flow on main is alerts 738/739). The directory is now the entry of plugins_dir matched by name, so nothing built from user input reaches the filesystem; an id with nothing installed goes to the store manager, which reports it not found as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): read the blueprint's plugin_manager defensively in _plugin_directory _get_plugin_version now goes through _plugin_directory, which read api_v3.plugin_manager directly; the attribute exists only once the app sets it, so test_path_traversal_guards::test_a_real_manifest_is_read failed when run on its own (order-dependent in the full suite). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
bcef1957a9 |
fix(security): refuse unsafe plugin ids, keep secrets private, validate request bodies (#643)
* fix(security): refuse unsafe plugin ids, keep secrets private, validate bodies - install_from_url and the registry install's manifest rename refuse a plugin id that is not a single safe name (no ../ out of plugins_dir). - Uninstall and config reset refuse core config sections and ids with path parts; uninstall of a plugin whose directory is gone still works. - separate_secrets checks a field's own x-secret marker before recursing, so object/array secrets no longer land in config.json. - Backup restore creates missing secrets/wifi/ytm files with mode 640; export skips non-object manifests and no longer collides on same-second exports. - SYSTEM_FONTS includes every bundled font from BUNDLED_FONTS. - Raw config/secrets saves and validate_request_json require a JSON object. - A blank max_dynamic_duration_seconds keeps the stored value; other values are validated to 30-1800 instead of raising a 500. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(security): validate the id before install_plugin moves anything; claim backup names atomically - install_plugin set aside plugins_dir / plugin_id before any id check, so "../x" moved a directory outside the plugins dir (the rollback moved it back, but only if the install path got that far) - two exports finishing in the same second could both see a free name and the later os.replace destroyed the first archive; the name is now claimed with O_EXCL before the archive is swapped in Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3a81f38f09 |
fix(web): uniqueItems saves, /health count, Vegas order wipe; one list-repair helper (#638)
* fix(web): drop repeats from uniqueItems lists before validating a plugin save dedup_unique_arrays lost its only caller in #330, so submitting a value a uniqueItems list already holds (a stock symbol saved once and posted again) failed the whole save with a validation error. _prepare_plugin_config_for_save runs it again just before validation, which covers both POST /plugins/config and plugin sections posted to /config/main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): /health counts the discovered plugins and logs the checks it fails The plugin check counted plugin_manager.get_available_plugins(), which PluginManager does not have, behind a hasattr guard that made plugin_count 0 on every device. It now counts the discovered manifests, discovering first when nothing has been scanned yet. The config, plugin and hardware checks answered "see logs for details" without logging anything. Each now logs a warning with the traceback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): store refresh no longer claims a commit-metadata refresh POST /plugins/store/refresh read fetch_commit_info (or fetch_latest_versions) only to append "(with refreshed commit metadata from GitHub)" to its message. It never fetched any: the route re-downloads the registry and nothing else. search_plugins takes the flag, but it reads commit info through its cache, so passing it on would not refresh anything either. The flag is ignored now and the message says what happened. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): refuse a malformed Vegas plugin order instead of clearing it A vegas_plugin_order or vegas_excluded_plugins value that was not JSON, or not a list, was stored as [] and the save answered 200, so a bad value wiped the saved order or exclusions. Both now answer 400 and save nothing, the way plugin_rotation_order already did; the three share one parser. A list that holds anything but plugin-id strings is refused as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): per-plugin health and metrics read the display service's latest GET /plugins/health/<id> and /plugins/metrics/<id> called get_health_summary and get_metrics_summary without force_reload, so they answered with whatever the web process read first and kept in memory, while the display service kept writing newer state. They now pass force_reload=True, as the list routes do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin config reset saves through the shared atomic save POST /plugins/config/reset called config_manager.save_config directly, so it took no backup, and a failed write escaped as an unhandled exception. It then handed on_config_change the raw stored section, not the prepared config a loaded plugin runs with. It now saves through _save_config_atomic with a backup, answers CONFIG_SAVE_FAILED when that fails, and notifies with _prepared_plugin_config, as POST /plugins/config does. POST /plugins/toggle carried its own copy of _save_config_atomic's save_config_atomic-or-save_config fallback; it calls the shared helper now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): one reading and one "unavailable" for each system metric system_metrics.collect_system_metrics() promised None for a metric it could not read, but returned cpu_temp as 0 off a Pi, and the whole no-psutil fallback as zeros. GET /system/status measured the same numbers a second time with its own code, and answered None there. Now both come from collect_system_metrics(), and "unavailable" is None everywhere. /system/status keeps its 0.1s CPU sample and its 10s cache, and gains nothing it did not already send. Two differences: without psutil it answers 200 with null metrics instead of 503, and a disk it cannot stat is null instead of a 500. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): /display/current sends the snapshot as-is and logs a failed read GET /display/current PIL-decoded the preview snapshot and re-encoded it before base64-ing it, spending CPU on the Pi to send the same picture, and dropped any failure with `except Exception: pass`. The /stream/display SSE stream already passed the PNG's bytes straight through. Both now read through web_interface/display_preview.py and answer with the same payload. A missing snapshot is still a null image; any other read failure is logged as a warning. /health reads the snapshot path from the same module. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): one helper puts a submitted plugin config's lists back The plugin-config save turned position-keyed dicts ({"0": ..., "1": ...}) back into lists in five copies: four in the form path's fix_array_structures (whose prefix branches never ran, since no caller passed one), and _fix_json_arrays on the JSON path. It then force-fixed the news plugin's feeds.custom_feeds by name, in case the generic pass had missed it. src/web_interface/config_arrays.coerce_array_shapes now does it for both paths, custom_feeds included. ensure_array_defaults duplicated _fix_none_arrays and is gone. In the same function: the union-type re-checks that the null handling above them made unreachable, the "(temporary)" random_seed debug log, and a commented-out log line are removed. A failed validation is logged once as a warning, not four ERROR lines and a WARNING. Element types are left to normalize_config_values, which already converted them for both paths. One difference: the form path no longer adds an empty {} for a nested object the post left out that has no defaults. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): import at module top and log through the module logger The web_interface.cache imports in config.py and fonts.py were wrapped in `except ImportError` fallbacks. It is an in-repo module that imports nothing from the project, so it cannot fail to import; it is imported once at module top, as system.py now does. cache.py's docstring said blueprints import it lazily "to avoid circular imports"; it now says why that is unnecessary. Five logging.error calls in the dim-schedule GET and three logging.warning calls in plugins.py went to the root logger; they use the module logger. Function-local re-imports of json, os, shutil, logging and Path, all already imported by the module, are gone. The `import os` inside two except blocks of save_plugin_config also made os a local name for the whole function. execute_plugin_action's step-1 handler gets a comment saying why it stays: it looks like a copy of the blueprint handler, but without it a TimeoutExpired from the plugin's script would reach the route's own `except subprocess.TimeoutExpired` and be answered as a 408. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): app.py loses dead CSRF and reconciliation state, comments fixed - csrf was always None, so `if csrf: csrf.exempt(...)` never ran, and its note that the api_v3 blueprint "is exempted above" named an exemption that does not exist. Both are gone; the reason there is no CSRF protection stays, shortened. - The SSE rate-limit comment called the default "tight" at 20 per minute. The default is 1000 per minute and the streams' 200 is the tighter one; the comment now says so. The limits are unchanged. - _reconciliation_done was written and never read. The docstring that explains why reconciliation runs once keeps its reason, in the present tense. - Removed: a dangling "import cache functions" comment with no import under it, a "security check ... within project_root" label on an existence check, the "(simplified version)" narration, and the note that no redirect route is needed. The preview loop's sleep comment no longer mentions a PIL encode that the loop does not do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(web): api_v3 comments name the package __init__, not a _common module Every route module's docstring said the shared blueprint comes "from ._common", a module the package split never created; they name the package __init__. The PROJECT_ROOT comment described the path from _common.py; it now describes this package and keeps the incident it guards against. The "(corrected) in this commit" note in resolve_pull_command and the /health comment the split's mechanical time -> _pkg.time rewrite garbled ("Stamp the start _pkg.time") read correctly again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): drop hasattr checks for attributes PluginManager always has PluginManager.__init__ sets health_tracker and resource_monitor (to None until they are configured), so the seven hasattr(api_v3.plugin_manager, ...) guards in the health, metrics and limits routes were always true. The falsy checks that do the work stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): pages_v3 dispatches partials from a dict with one error handler load_partial chose a loader through a fourteen-branch if/elif, and thirteen of the loaders then wrapped themselves in the same try/except, logging "Error loading partial" without saying which. The route now looks the name up in _PARTIAL_LOADERS and has the one handler, which logs the partial's name. The loaders just render. _load_tools_partial keeps its own messages. The search index's _partial_html already catches a loader that raises. serve_plugin_web_ui repeated _plugin_dir_for inline (containment plus the ledmatrix- prefix fallback); it calls it now. Also removed: the unused markupsafe.escape import, function-local json/Path re-imports, and unused exception bindings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove unused imports, locals and a try that cannot fail - get_error_aggregator was imported by the api_v3 package and used by no one; seven names config.py imported, and Path in misc.py and logging in plugins.py, likewise. - branch_info in install_plugin was built and never logged; test_config in /health was bound and never read (the load_config call is the check). - An f-string with no placeholders in the asset upload route. - _installed_plugin_ids wrapped list(manifests.keys()) in try/except; _discovered_plugin_manifests always returns a dict. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): start.py logs its startup lines and drops unreachable branches The startup banner went to stdout with print(); it goes through a logger now, which the app import has already configured, so it reaches the journal with a level and timestamp like every other line. The "no addresses" branch is gone: get_local_ips() always returns at least "localhost". The except around app.run re-raised "only if it's not a client disconnection error" from inside the branch that had just established it was one, so that raise could not run. It is one check now, on a named tuple of the errnos, which the werkzeug log filter uses too. The comment on threaded=True counts three SSE endpoints, which is how many there are. Trailing whitespace is stripped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): save_main_config names its General fields once The General tab's field names were listed twice, once to detect a General form post and again, with four more, to keep the remaining-keys merge from storing them as top-level keys. GENERAL_FIELDS and _MAPPED_TOP_LEVEL_FIELDS hold them now, and the four per-section skip checks are one set. The comment on that merge said plugin configs are handled "here too", and "(including plugin keys)". Plugin sections are handled and removed from the body before it runs; the comment says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): plugin directories come from the plugin manager only Six lookups fell back to PROJECT_ROOT/plugins/<id> when there was no plugin manager: GET /plugins/config's of-the-day data, POST /plugins/action, the plugin static-file route, the calendar credentials upload and the calendar OAuth routes. The loader never scans plugins/ (PluginManager.discover_plugins reads only the configured directory, plugin-repos by default), so what they found there was a plugin that never runs. _plugin_directory() asks the manager and answers None without one, which each route already reports as "not found". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): web-backend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4e61d7248a |
refactor(web): one error-response path for api_v3 (#624)
* refactor(web): answer unhandled api_v3 errors from one blueprint handler
Fifty-three api_v3 routes ended in a copy of the same catch-all: log the
traceback, return {status, "An error occurred; see logs for details",
details: describe_exception(e)} with a 500. They are replaced by one
errorhandler on the api_v3 blueprint that returns exactly that body.
It lives on the blueprint rather than falling through to app.py's global
handler because the two answers differ: the global one adds
error_code: UNKNOWN_ERROR, and api_client.js sends a body with an
error_code to the error modal and one without to a plain toast. A
blueprint handler also gives tests that mount api_v3 on a bare Flask app
the same answer the real app gives.
Only handlers that were byte-for-byte that shape were removed (matched on
the AST, and each rewritten function re-parsed and compared). Handlers
with their own message, extra keys, operation-history records or cleanup
stay, as does execute_plugin_action's step-1 handler, which sits inside
an `except subprocess.TimeoutExpired` arm that would otherwise turn a
plugin's timeout into a 408.
HTTPExceptions raised inside a route go back as themselves in the global
handler's 4xx shape. Where a removed catch-all used to swallow one (only
delete_plugin_asset's non-silent get_json() is reachable), a malformed
request now gets its 415/400 instead of a 500.
Most of the diff is re-indentation from unwrapping the try blocks;
`git diff -w` shows the real change.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin action errors name the real failure, not UnboundLocalError
execute_plugin_action bound a local `logger` in its JSON-parsing arm,
which made `logger` local to the whole function. Every other
`logger.error` in it then raised UnboundLocalError, so a failing OAuth
step-1 script was reported as "UnboundLocalError: cannot access local
variable 'logger'" -- from the step-1 handler, and before the previous
commit from the route's outer catch-all too. Use the module logger.
Found by comparing every api_v3 route's forced-failure response before
and after the catch-all consolidation.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop the error category and exception-name code guessing
WebInterfaceError derived an ErrorCategory from every error code and put
it in each structured error body as `error_category`. Nothing reads it:
not the web UI (static/ and templates/), not the tests beyond the ones
pinning the mapping itself, and not any plugin in ledmatrix-plugins. The
enum, the inference table and the JSON key go.
from_exception() could also guess an error code from the exception's
class name ("Config" -> CONFIG_LOAD_FAILED, and so on). Every caller
passes a code, so the guess never ran; error_code is now required.
suggested_fixes stays: the error dialog in static/v3/js/utils/
error_handler.js lists them.
The REST reference loses error_category and says what an unanticipated
exception in an /api/v3 route answers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one call for the from_exception error responses
Nine plugin routes built a structured error by hand:
from src.web_interface.errors import WebInterfaceError
error = WebInterfaceError.from_exception(e, ErrorCode.X)
return error_response(error.error_code, error.message,
details=error.details, context=error.context,
status_code=500)
That is now exception_error_response(e, ErrorCode.X) in api_helpers, so
error_response() is the only structured-error entry point the routes
use. The three operation-history routes never passed the context, and
with_context=False keeps their bodies exactly as they were; a test
compares the helper against the hand-written pair for both forms.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): one api_v3 error-response path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
61e462c635 |
refactor: remove the skin system and the unused src/base_classes package (#615)
* refactor: remove the skin system Skins never rendered with the current scoreboard plugins: the only hook was SportsCore._render_game in src/base_classes, which no plugin builds on, so the UI and store already treated them as unsupported. The owner decided on 2026-09-23 to remove them outright. Removed src/skin_system/ (runtime, base class, fixtures), skins/, scripts/validate_skin.py and their tests; the store's "type": "skin" installer, uninstaller and hide/refuse filters (the official registry lists no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins. Stored skin/skin_options config values are handled in the next commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(config): drop retired skin/skin_options keys instead of validating them A config.json written while the skin system existed can carry skin and skin_options in any plugin section, and most plugin schemas set additionalProperties: false. They are no longer core plugin properties; RETIRED_PLUGIN_KEYS in schema_manager lists them and drop_retired_plugin_keys removes them (unless the plugin's own schema declares the name) in prepare_plugin_config, which loading, hot reload, GET /plugins/config and both web saves already share, and in validate_config_against_schema for callers that validate a raw section. POST /plugins/config and /config/main also drop them from the stored section they merge into, so they leave config.json on the next save. Tests cover the load path (real PluginManager.load_plugin: no schema warning, not degraded), raw and prepared validation, validate_all_plugin_configs, and the JSON, form and /config/main saves. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: remove the unused src/base_classes package No scoreboard plugin builds on src.base_classes: the nine monorepo scoreboards ship their own sports.py and share code through src/common (docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins imports it. The one import anywhere, baseball-scoreboard's rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing instantiates. Removed the package and the eight test files that only tested it (test_api_extractors, test_data_sources, test_sports_base_characterization, test_sports_capabilities, test_sports_core_promotions, test_sports_logo_cache_bounded, test_sports_modes_promotions, test_sports_odds_fanout). test_common_is_hardware_free no longer lists src.base_classes as a forbidden import, and comments in sports_helpers.py and base_odds_manager.py stop pointing at it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: drop the skin system and src/base_classes from the docs Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was removed and shared code lives in src/common, in the Layering section and the view-model-contract rule. Other docs stop pointing at the removed package. CHANGELOG records both removals under Unreleased. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(store): hide and refuse registry entries that aren't plugins The skin filters went with the skin system, but a custom registry can still list "type": "skin" entries, and installing one as a plugin would unpack it into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing type means plugin) now hides non-plugin entries from the store and custom-registry listings, and install refuses them, in the route with a clear 400 and in _install_plugin_impl for any other caller. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
342e9164b8 |
fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound Each repeat of an error pattern appended every plugin in the time window to the pattern's list again, so a plugin failing in a loop grew the display process's memory without limit: 3,000 errors from three plugins reached 2.5 million entries. Keep the list unique. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): load a BDF font at its native size instead of PIL's default FreeType rejects any size but a BDF strike's own, and FontManager answered that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf requested at 8 or 10px rendered as PIL's default font. Retry at the native strike, as element_style already does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin toggle failures no longer claim "operation in progress" Every exception in POST /plugins/toggle was mapped to PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation is already in progress". Report the failure as what it is, and record the plugin id in the operation history for form posts too (it read a `data` variable that only the JSON path set). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): route plugin card clicks through handlePluginAction The document-level delegation checked `typeof handlePluginAction`, which is scoped inside the plugin-manager IIFE and so never visible to it. Every card click took a copied fallback that stopped propagation (the grid's own listener never ran), confirmed an uninstall twice, and sent Starlark app uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>. Expose the handler on window and delegate to it. Also run every test/js/unit suite under pytest: they need only node, but CI ran one of the eight. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): apply Rotation durations, WiFi messages and Vegas settings Three settings the web UI saves never reached the display: - Rotation & Durations: display.display_durations was never read. Every plugin inherits get_display_duration() and the plugin was asked first. A saved value now wins. The page shows unsaved screens blank with the plugin's own duration as a placeholder, and saving a blank removes the override, so one save no longer pins every screen. - WiFi status overlay: the controller looked for wifi_status.json one directory above the repo. Both sides now use wifi_manager.get_wifi_status_path(). The message is written by rename so the display never reads it half-written, and the resumed plugin redraws the whole panel afterwards. - Vegas: nothing called coordinator.update_config(), so saved Vegas settings never reached a running scroll. They are now queued when display.vegas_scroll changes, and applied while Vegas is stopped too, so a disable then re-enable works. The follower's scroll-speed default (75) now matches VegasModeConfig's (50). Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows with each plugin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep affected_plugins order when serialized; guard non-Element targets ErrorPattern.to_dict() ran the now-ordered list through set(), so get_error_summary() listed plugins in an unstable order. The document-level card-action listener called event.target.closest() without checking the target is an Element. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
116abb0daa |
fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service BackgroundDataService always sent a season range first and, on a 400, fell back to chunks without recording the rejection, so every background season fetch spent a doomed request and live scoreboards learned nothing from it (or it from them). The worker now consults and sets the same 6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight to month/day chunks, and if every chunk fails the range is asked once for a real error without re-spending the chunks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep plugin asset and action routes inside their directories POST /plugins/assets/upload, GET /plugins/assets/list and POST /plugins/assets/delete joined the request's plugin_id onto assets/plugins unchecked, so '../../config' created, wrote, listed and deleted outside it. #561 guarded only the route that serves the files. All three now go through path_safety.resolve_under and answer 400 for anything but a plain name, and delete only unlinks a metadata path that resolves into that plugin's uploads directory. PluginManager.get_plugin_directory refuses ids that are not one plain path segment, so /plugins/action (which runs a manifest script from the returned directory) and every other caller get the guard; the action route also rejects such ids up front, covering its no-manager fallback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): report a no-op plugin update as already up to date update_plugin() returns True both for a real update and for "nothing to do" (a ZIP-installed monorepo plugin already at the registry version, a bundled plugin). With no git commit to compare, POST /plugins/update called every such success "updated successfully", so Check & Update All counted most official plugins as updated on every run. The route now reads what changed off the plugin itself (commit, else manifest version, else last_updated) and returns data.update_status (updated / up_to_date / local_only). The update-all toast is summarised by PluginInstallManager.summarizeUpdateResults from that status, falling back to the message for older servers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): scoreboard scroll speed no longer follows target_fps sports_scroll computed the crisp speed ladder against the global target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since frame-locked presentation (#545) the helper steps a fixed number of whole pixels per presented frame and the panel presents at its real refresh, so the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a 50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s. The ladder now uses the display manager's refresh_hz, then display.hardware.limit_refresh_rate_hz, then the default. target_fps is not consulted. Docstrings now say scroll_delay is ignored for pacing (no behaviour change there) and describe the fixed-step model. Tests: replace the tests that pinned target_fps as the ladder refresh and described time-based stepping; assert speed independence from target_fps (unit and end-to-end presented px/s against the real helper), that the fixed per-frame step is applied, and that scroll_delay does not change speed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): escape registry and upload values in plugin manager inline handlers The store, saved-repository and custom-registry buttons built onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone, so a custom registry entry whose id contained ' closed the attribute and added its own handler. One helper, jsStringAttr(), now HTML-escapes the JSON literal for every one of those handlers, and the store View button opens only http(s) repo links. The live window.updateImageList (plugins_manager.js loads last, so its copy wins over the file-upload widget's) wrote the uploaded file's original name, path and ids into markup raw; they are escaped now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note plugin asset, action and inline handler guards Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): let the root pip wrapper install web_interface/requirements.txt Update Code, the automatic update's health check and Install Base Requirements install web_interface/requirements.txt through safe_pip_install.sh, which only allowed the root requirements.txt. The first commit changing that file would fail its dependency install, and the automatic updater rolls back any update whose dependencies did not install -- on every device, for every newer commit. The wrapper now lists both core requirement files. Only their folders are resolved, so a requirements.txt symlinked out of the project is compared by its target and refused (previously the root file's own symlink target was what got allowed). The updater's file list is a named constant, and a test runs the real wrapper (pip stubbed) on every file Update Code and the rollback install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): do not retry plugin requests that got an HTTP answer PluginAPI.request wrapped everything that was not a structured error as NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a JSON error without error_code included. Check & Update All retries NETWORK_ERROR, so those updates were re-sent five more times with backoff, contrary to the #587 contract that an HTTP error response is the server's answer. NETWORK_ERROR now means only that fetch() rejected. Any HTTP response without an error_code, or with a body that is not JSON, is API_ERROR with the HTTP status attached. Tested against the shipped api_client.js. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): restart the stats window when an idle gap is dropped by size #582 dropped an idle gap from the frame stats two ways: the reset_scroll() sentinel, which also restarts the 5s window timer, and a size guard for scrollers that never call reset_scroll(), which did not. On that path the first real frame after the gap found the boundary overdue and logged a stats line for a one-frame window. Both paths now share one seeding helper. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): leave plugins alone when update_core's own rollback fails update_core returns rollback_failed directly when a partial pull or an update whose health check never started cannot be rolled back. run() only held plugins back for 'verifying', so those devices still got new plugin versions and a display restart on top of a core in an unknown state -- the opposite of what the health-check path does, and of the 3.4.0 changelog (plugins are left alone if the rollback fails). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(api): make the REST reference match the api_v3 package Every documented request body, query parameter and response shape was re-checked against the handlers in web_interface/blueprints/api_v3/. Fixes calls that failed as documented (repo_url, action_id/params, files/image_id, font_file+font_family, ?font=, cache key, auto_enable_ap_mode, plugin limit keys), removes the font-override endpoints dropped in #566, corrects response shapes (plugins/config, plugins/schema, health, metrics, operation history, github-status, fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds the 26 routes it omitted (backup, system auto-update/git, wifi radio, starlark editor, MQTT bridge, status endpoints, skins). Documents the merge semantics of partial JSON saves to /config/main and /plugins/config and the dim-schedule POST accepting GET's days shape, which land in the same change set. Replaces app.py line numbers and the removed api_v3.py path with file and function names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): remove the General-tab plugin system toggles that did nothing plugin_system.auto_discover, auto_load_enabled and development_mode had General-tab toggles whose help tips promised dormant plugins and verbose logging, but nothing reads them: every enabled plugin is discovered and loaded regardless. Remove the three toggles. The keys stay tolerated in stored configs. The save handler now stores a flag only when a client sends it; treating a missing key as an unchecked box would otherwise rewrite all three to false on every General-tab save, which still posts plugins_directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(scroll): remove dead code left by #523/#570 - Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read them since the numpy blend replaced the scipy path. - Drop ScrollHelper._last_integer_position and frame_time_target, which were written but never read. - Keep target_fps and set_target_fps() but document them as informational: nothing paces off them, yet ledmatrix-elections' test_scroll_pacing.py reads helper.target_fps back and third-party plugins may call the setter. - Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to throttle", set_sub_pixel_scrolling's "default: True", and set_frame_based_scrolling's claim that it steps. The plugins monorepo was grepped for every removed name; none is used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI FONT_MANAGER.md told plugins to read display_manager.font_manager, which does not exist, so a plugin following it failed to load with AttributeError. The shared FontManager lives on the PluginManager and BasePlugin._get_font_manager() returns it (with a fallback for harnesses). Also removes the Fonts-tab override workflow and element-override panels that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls /plugins/store/search does not exist (404) and the list endpoint reads query, not q. The registry guide's curl examples omitted the JSON Content-Type, so the handlers saw an empty body and answered 400. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(config): use the shared core-key list in the last three private copies StartupValidator warned "Plugin 'auto_update' is enabled but not found" on every display start with auto-update or a dim schedule on; the reserved plugin-id check missed auto_update, sync, location and the rest; and ConfigManager's (uncalled) orphan cleanup would have deleted display, schedule and auto_update. All three now read src/core_config_keys.py, which also gains CORE_SECRETS_KEYS for the github/youtube secrets sections. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): partial JSON saves to /config/main change only what they send A JSON body with one field reset every checkbox in the sections it touched: the MQTT bridge's brightness slider turned off disable_hardware_pulsing, inverse_colors, show_refresh_rate and use_short_date_format, and a timezone-only save turned off web-UI autostart and weekly auto-updates. Missing-means-unchecked now applies only to form posts: form-encoded bodies and the v3 forms, which mark themselves with a hidden __form_section input. Also on the config routes: - vegas_min/max_cycle_duration no longer match the generic *_duration rule, so they stop landing in display_durations and a blank one no longer rejects the whole Display save; - saving from the Raw JSON editor calls start_setup_if_needed like the General form, so enabling auto-update there finishes its setup; - the schedule and dim-schedule POSTs accept the per-day days.<day> shape their GETs return, as well as the flat form keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): install plugin dependencies from the configured plugins directory install_plugin_dependencies.sh scanned only plugins/, but the Plugin Store installs into plugin_system.plugins_directory (default plugin-repos), so the documented "Recommended" fix found 0 plugins on every store install. It now reads plugins_directory from config/config.json (relative to the project root or absolute, default plugin-repos) and also scans plugins/ for dev symlinks, installing a plugin reached through both only once. With set -e alone, `pip ... | tee` took tee's exit status, so a failed pip install was reported as success; set -o pipefail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: replace stale API names, line numbers and the api_v3.py path - ADVANCED_FEATURES: StreamManager methods that exist (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the real on-demand status envelope ({status, data: {state, service}}) - app.py:199 / :144 / :607-619 line citations and web_interface/blueprints/api_v3.py (now a package) replaced with file and function names in ADVANCED_FEATURES, CONFIG_DEBUGGING, PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE, PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README - CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use /config/raw/main to replace the file; describe where validation runs - TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints usage) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): verify the web interface that actually ships, on port 5000 verify_installation.sh failed every healthy install: it required the long-removed web_interface_v2.py and looked for a listener on port 5001, while the web interface binds 5000 (web_interface/start.py). It now checks the files ledmatrix-web.service runs (start_web_conditionally.py, web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the same 5001 port in its listen check, HTTP probe and printed URLs. Port matches are anchored so :50001 no longer counts as :5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(plugins): one display-size contract: display_manager.width/height CLAUDE.md (#580) says to read display_manager.width/height because matrix is None when hardware init fails; the development guide, the safety-harness doc and two DisplayManager docstrings still recommended matrix.width/height. The bundled starlark-apps plugin read matrix.width unguarded, so its magnify recommendation and frame scaling raised in fallback mode (e.g. after the Pi 5 hardware refusal). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(install): make install_service.sh --help print usage instead of installing install_service.sh parsed no arguments, so `sudo ./scripts/install/ install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md) rewrote ledmatrix.service, ledmatrix-web.service and both update-verify units and enabled/started them. It now handles -h/--help (usage, exit 0, no changes) and rejects any other argument with exit 2 before doing anything. Running it with no arguments, as first_time_install.sh does, is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scroll): describe the fixed-step model and document frame_hold Since #545 a crisp speed from scroll_config.configure() makes the helper advance a fixed whole-pixel step per presented frame with no clock, and the display manager's frame hold is part of the speed. The docs still described the removed wall-clock model: - scroll_config's module and configure() docstrings said speed is applied in time-based mode and that omitting the hold "falls back to fractional pixels"; omitting it actually runs the scroll frame_hold times too fast. - SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both modes, and read a 20 ms stats median as missed refreshes although that is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step, the hold-dependent healthy median, that target_fps plays no part, and that a hand-added scroll_pixels_per_second loses to a schema-default pair. - PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling) without frame_hold; it now documents the parameter (core 3.4.0) with a configure() + set_scrolling_state example. - update_scroll_position/set_scroll_speed and set_scrolling_state docstrings say the same. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark target_fps legacy; describe what Vegas scroll_delay does - General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after the sports_scroll fix nothing in core scrolling reads it. The field and its API validation stay so saved configs and plugins that read global_config['target_fps'] keep working. CONFIG_REFERENCE says the same. - Vegas frame_based_scrolling/scroll_delay were described as frame-count stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and still advances by elapsed time, so the applied speed is clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The config comments, render_pipeline comment and CONFIG_REFERENCE rows now say so. No behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(deps): describe how plugin dependencies are really installed The guides said the web service runs as root, that installs pick --user from os.geteuid(), and quoted a warning and a PluginManager._install_plugin_dependencies() method that don't exist. The web unit runs as the installing user; store installs go through install_requirements_file() and sudo safe_pip_install.sh (root), with a user-level fallback that says so, and load-time installs run in the display service's own (root) interpreter. Manual paths now use the configured plugins directory (plugin-repos/ by default) instead of plugins/, which store installs no longer use, and install_plugin_dependencies.sh is described as scanning that directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): count local changes one way for the preflight and the pull The automatic update's preflight ignored mode-only changes and anything whose status line contained plugins/ or plugin-repos/, then promised "Automatic updates will not stash your changes". perform_core_update used plain git status (modes count) and ignored only 'plugins/', then ran 'git stash push -- :!plugins', which nothing ever pops. So an edit to a bundled plugin under plugin-repos/, or the installer's chmods on tracked scripts, passed the preflight and was stashed away for good. - auto_update.local_changes() is the one predicate both use: core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/ excluded by leading folder rather than substring (a core file under web_interface/static/v3/js/plugins/ now counts). - Update Code's explicit stash leaves out both plugin folders; the pull's --autostash carries their edits and mode changes across and reapplies them. - The automatic updater calls perform_core_update(stash_local_changes= False), which refuses instead of stashing edits that appeared after the preflight; update_core reports that as 'blocked'. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): diagnostics follow the web autostart default and api_v3 package #556 made a missing web_display_autostart mean "start" (only an explicit false/off keeps the web interface down), but the diagnostics still said otherwise: diagnose_web_ui.sh reported a missing key as "defaults to false", diagnose_web_interface.sh said the web interface "will not start unless this is set to true" and recommended enabling it, and debug_web_manual.py printed False. Troubleshooting a down web UI pointed users at a non-cause. Both shell scripts now evaluate the setting with the launcher's own autostart_enabled() (inline fallback if it cannot be imported) and report on / off / not set (on) / unparseable config; debug_web_manual.py uses the same function. They also check web_interface/blueprints/api_v3/ __init__.py: api_v3.py became a package in #553, so every healthy checkout was reported as missing a file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(install): what install_service.sh installs; verify script port; no sudo for --help install_service.sh installs and starts ledmatrix, ledmatrix-web and the update-verify units, not only ledmatrix.service (systemd/README.md, README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help' as a harmless check; it now shows --help without sudo and warns what a real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh checks the web interface on port 5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note update-all, plugin system settings and script fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): size the preview after orientation and pixel mappers display_geometry.physical_size claimed to give DisplayManager's answer but only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured after the library's pixel mappers, so a Rotate:90 / orientation 90 chain previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for 128x64. Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown names leave it alone), and move the orientation composition here so DisplayManager and the preview share it. The module docstring no longer claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH. Audit finding F18. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): refuse settings the rgbmatrix library aborts on, on every board The library answers several settings with a NULL matrix or abort() rather than an error, so the display service crash-looped (Restart=on-failure) instead of reaching fallback mode: rows above 64, chain_length above 255 (uint8_t binding setter, documented as "no upper limit"), a misspelled hardware_mapping, and parallel 2-3 on a single-output mapping, reachable from the Display form on the default adafruit-hat(-pwm) mapping. #586 only guarded the Pi 5 subset. - src/matrix_support.py holds the rules for every board (Options::Validate ranges, binding integer types, mapping names and outputs from lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the API's numeric ranges. - DisplayManager checks them before building options and raises MatrixSettingsRefused, so a hand-edited config falls back with a logged, reported reason. Emulator mode only warns. - The config API refuses them with a 400 naming the setting; combinations are checked against stored values but reported only when the request sets a field involved. - The hardware status file gains "cause" (settings/library/forced). The fallback log and Display banner give the Pi 5 rebuild hint only for a library failure instead of rebuild + gpio_slowdown advice for every failure; one Pi 5 slowdown recommendation (1-3, start at 1). - The Display form offers classic/classic-pi1 and orientation 90/270 and renders any other stored mapping selected with a warning, so an unrelated save no longer rewrites them; the API accepts 90/270. Audit findings F03, F16, F19, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(display): library limits, template defaults and Pi 5 slowdown - rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs, classic/classic-pi1 mappings and orientation 90/270 documented. - Defaults are the config.template.json values: config migration adds missing keys from the template, so the listed "code defaults" never applied. - One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode, starting at 1. - Troubleshooting describes the refused-settings fallback, and CHANGELOG corrects the Unreleased "no upper limit" entry. Audit findings F19, F20, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py opens the panel with the service's options --measure and --demo built RGBMatrixOptions from a private copy of the display service's builder that had drifted: gpio_slowdown came from display.hardware (default 2) instead of display.runtime (default 3), and rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors, pixel_mapper_config and orientation were skipped, with different defaults (hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown was measured -- or garbled -- in a setup the service never drives. The option filling in DisplayManager._setup_matrix moves, unchanged, into DisplayManager.apply_matrix_options(options, config), which _setup_matrix calls and the script reuses (overriding only limit_refresh_rate_hz for --measure). The script now loads the whole config rather than the hardware block. Tests pin the script's options to the service's attribute for attribute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py recommends keys the resolver honours The ladder ended by telling users to set display_options.scroll_pixels_per_second. scroll_config ranks that key below the scroll_speed + scroll_delay pair, deliberately, and several plugin schemas default the pair into config, so the advised key was silently ignored (a schema-default 1/0.02 pair plus an advised 66 still resolved to 50 px/s). The advice is now the pair that selects the crisp speed exactly (pixels_per_frame every frame_hold/refresh seconds), explains that the pair outranks scroll_pixels_per_second, and gives the scoreboards' per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the printed pair over a schema-default pair and check it lands on the advertised speed and hold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula - SPORTS_UNIFICATION.md still presented honouring global target_fps as sports_scroll's added behaviour and its one user-visible gain; note that it was withdrawn because it had become a speed multiplier. - ADVANCED_FEATURES.md gave Vegas scrolling as (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s by elapsed time, through a 0.1-5 px per scroll_delay clamp when frame_based_scrolling is on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): scroll model fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(dev): link-github links plugins from the ledmatrix-plugins monorepo link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git, and those per-plugin repositories no longer exist: official plugins are directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the monorepo once into the dev directory, finds plugins/<name>, plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and links it under its manifest id. With an explicit repo URL it still links a single-repository plugin as before. dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a fork), plus plugins_repo and plugins_branch; github_pattern, which was documented but never read, is dropped and warned about. Ships dev_plugins.json.example and git-ignores dev_plugins.json, both of which the guide promised. Reading JSON falls back to python3 when jq is missing (get_plugin_id silently returned nothing without jq). update/status/list find the git checkout above a monorepo plugin directory (its .git is not in the plugin dir), and update pulls a shared checkout once. status no longer exits 1 when nothing is broken. Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration, workflow, store integration, hello-world link, submission), and the nonexistent scripts/git-hooks/pre-push-plugin-version and scripts/bump_plugin_version.py replaced with the real rule: bump the manifest version and run update_registry.py. scripts/dev/README.md and CLAUDE.md updated to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scripts): monorepo workspace layout; fix_perms and install READMEs MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin; setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the workspace file opens LEDMatrix plus ../ledmatrix-plugins. scripts/fix_perms/README.md listed cache directories fix_cache_permissions.sh never touches and a 'ledmatrix' service user that doesn't exist (also in scripts/install/README.md); adds safe_pip_install.sh. install/README: install_service.sh installs the web and update-verify units too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): keep the rollback's pip retries inside the unit time limit The health check reinstalled the previous requirements by trying the next bash path after any failure, including a 600 s pip timeout. Two files, two paths: up to 40 minutes of pip alone, while systemd stops ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the rollback half-way and leaving the update 'verifying' until the web UI calls it lost. - Like permission_utils.install_requirements_file, only a sudo refusal moves on to the next bash; a pip that ran and failed or timed out is not repeated. The refusal wording is one list (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only verifier and pinned equal by a test. - All reinstalls in one rollback share a 600 s budget. - WORST_CASE_SECONDS adds up every timeout on the longest path (27.5 min); a test holds it under the unit's TimeoutStartSec and that under the web UI's VERIFY_LOST_SECONDS. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools Plugin config was prepared differently depending on how it arrived: - JSON POST /plugins/config built a partial body on schema defaults, so {"enabled": true} reset every other setting of the plugin. It now merges onto the stored section first, as the form path already did. - Legacy-boolean normalization (#588) ran only at load: GET /plugins/config returned the raw boolean, posting it back failed validation, and hot reload handed plugins the raw section (a legacy dynamic_duration: true came back as a boolean). schema_manager.prepare_plugin_config (normalize, then defaults) is now used by PluginManager.load_plugin, both save paths, GET, the save notifications and DisplayController's hot-reload callback. - The JSON save's filter kept only enabled/display_duration/live_priority and dropped a submitted skin, skin_options or vegas_* tuning key. There is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES, used by validation and by the save filter; PluginManager's CORE_OWNED_CONFIG_KEYS is its vegas subset. - Plugin sections posted to /config/main were stored verbatim, including values /plugins/config rejects. They now go through the same preparation (_prepare_plugin_config_for_save, extracted from save_plugin_config), and a failing section rejects the whole save before anything is written. - dev_server read only top-level defaults and let a schema enabled:false win; build_full_config shallow-merged overrides, dropping sibling defaults; the harness extracted defaults differently from the device. loading.build_config now uses the device's extraction and preparation, and dev_server, check_plugin, render_plugin and the harness all use it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt-bridge): brightness changes apply live and touch nothing else The display service's hot reload applies a saved brightness within a few seconds, and /config/main no longer resets other display settings on a brightness-only JSON body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): automatic update hardening Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI It described web_interface_v2.py and index_v2.html (both gone), client-side form generation, one POST per field with {key, value}, and 'no nested objects'. The v3 UI renders plugin forms server-side from the schema (pages_v3 partial + plugin_config.html macros, nested sections and x-widgets), posts the whole form once, and save_plugin_config() merges onto the stored section, validates, splits x-secret fields and notifies the plugin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt): brightness saves apply via hot reload and leave other settings alone The bridge README said brightness is applied on the display's next restart; the display controller's config hot reload applies it within seconds. It also now states that the bridge's partial JSON save changes only brightness (the /config/main merge fix in this change set). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): don't log pip's output from the health check's reinstall pip can echo a private index URL with embedded credentials; permission_utils redacts it, the stdlib-only verifier cannot, so it logs the exit code only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark the plugin_system toggles as unused legacy keys auto_discover, auto_load_enabled and development_mode are read by nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and the REST reference listed them as live settings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): docs and developer tools group Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): legacy plugin-system toggles no longer count as a General save auto_discover, auto_load_enabled and development_mode have left the General form, so a post carrying only one of them is not a general-settings save and must not treat web_display_autostart and auto_update as unchecked. The plugin_system block itself is left as on main for the branch that reworks it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): config-save and plugin-config preparation fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(claude): re-check matrix_support.py rules when the library submodule is bumped Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address Codacy findings on the core audit PR - plugin_manager.prepare_plugin_config: when the fallback legacy-boolean pass also fails, log a warning instead of a bare except/pass. - api_client.js: request() refuses any endpoint that is not a plain path under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace, control characters) with INVALID_ENDPOINT before calling fetch(), and plugin ids are URL-encoded wherever they are put into a URL (also in the app-shell batch load). - test_update_all.js: pins both against the shipped client. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): check endpoint control characters without a control-character regex Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags the \x00-\x1f range in checkEndpoint's regex. Test the char codes instead; the endpoints refused are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(auto-update): make the seed script executable on disk, not only in the index On Linux Repo.publish() commits with -a, which recorded scripts/run.sh as 100644 upstream because the seed file was never chmod +x. The pull then brought in the same mode the installer chmod had made locally, so installer_chmod saw no mode change left to check. The updater was fine: with the upstream commit at 100755 the --autostash carries the device's chmod across. Verified under Linux (WSL, git 2.43): the old helper fails exactly as CI did, the fixed one passes all 63 tests in the file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7e5967e160 |
fix(web): find installed plugins before anything has discovered them (#594)
The web process discovers plugins lazily: plugin_manifests is empty until some endpoint calls discover_plugins(). Three routes consulted it without discovering, so they misbehaved for as long as nothing else had run -- which, after every ledmatrix-web restart, is until someone opens the dashboard: - POST /display/on-demand/start answered 404 "Plugin <id> not found" (or "Mode <mode> not found"). Measured on a rig: 404 for over three minutes after a web restart, until GET /plugins/installed ran. The browser UI loads the plugin list first, so API-only callers (the Home Assistant MQTT bridge, scripts) are the ones who hit it. - POST /plugins/toggle answered 404 "Plugin not found". - POST /config/main did not recognise a plugin section, so it skipped secret separation and merged the section as-is: the plugin's API key was written to config.json in plain text instead of config_secrets.json. Add _discovered_plugin_manifests(), which discovers when nothing has been yet, and rescans once when a specific plugin id (or, for on-demand by mode, a mode) is not found, so a plugin installed since the last scan is found too. _installed_plugin_ids() now uses it. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f9b3d6ae52 |
fix(web): accept every panel size and row address type the rgbmatrix library does (#586)
* fix(web): accept every panel size and row address type the rgbmatrix library does The Display form capped columns at 128 and chain length at 24, and its submit handler (fixInvalidNumberInputs) rewrote anything larger to the cap, so wide panels and long chains silently saved as the wrong size. The config API checked none of the hardware numbers, so values the library rejects (odd rows, parallel 4, PWM dither bits 3) saved and the matrix then refused to start. - Form limits now match the pinned library: rows even 8-64, cols >= 16 and chain_length >= 1 with no upper bound, parallel 1-3, PWM dither bits 0-2, PWM LSB nanoseconds 50-3000. - save_main_config rejects out-of-range rows, cols, chain_length, parallel, brightness, scan_mode, pwm_bits, pwm_dither_bits, pwm_lsb_nanoseconds and gpio_slowdown with a 400. - A stored gpio_slowdown or pwm_dither_bits of 0 renders as 0 instead of the default, so saving the tab no longer overwrites it. - Row Address Type offers 5 (SM5368 / B707 row shift register). Verified on a Waveshare 96x48 V2 (24S-A1) on a Pi 4 with the Adafruit Triple LED Matrix Bonnet: rows 48, cols 96, row address type 5, BGR, GPIO slowdown 8. - Help text and docs: FM6124-family panels use Panel Type Standard; on a Pi 5 the library supports only row address types 0 and 2. No change to the rpi-rgb-led-matrix submodule. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): drop the rows cap and document every display setting accurately Rows: no upper limit in the form or the API. Still even and at least 8. The current rgbmatrix library rejects more than 64 per panel, so a larger value saves but the matrix won't start; the help tip, README, config reference and troubleshooting section all say so, and nothing here needs changing if the library lifts the limit. limit_refresh_rate_hz: the form accepts 0 (the library's "no cap"), a stored 0 no longer renders and re-saves as 120, and the API rejects negatives. pwm_dither_bits stays 0-2: the library rejects 3 and 4, so the old form's 0-4 only ever let users save a config the display couldn't start with. Docs and help tips, checked against the pinned library and its README: - panel_type and rp1_rio get README entries - show_refresh_rate prints to stdout; it never drew on the panel - dither bits raise the refresh rate; the tip said they lowered it - scan_mode is about interlacing at low refresh, not wrong colours - disable_hardware_pulsing: hardware pulsing needs OE on GPIO 18 and the onboard sound driver off; software timing makes rows flash brighter - gpio_slowdown guidance agrees between the README and the UI - all 22 multiplexing values listed; every numeric setting states its range - troubleshooting for a blank panel after a settings change, jumping rows and brightness flashes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): reject true and 5.5 for row_address_type and multiplexing Both still went straight through int(), so a JSON true saved as 1 and 5.5 as 5. They now use the shared hardware range check like the other panel fields. Review feedback on #586. Also: the RP1 Backend tooltip said it is ignored on Pi 3/4 (it is ignored on every model but the Pi 5), and the README gave the dynamic-duration default cap as 90s; the code default is 180s. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: refuse matrix settings a Raspberry Pi 5 can't drive On a Pi 5 the pinned rgbmatrix library drives the panel through the RP1 chip, and that path supports only row address types 0 and 2, parallel 1-3 and the regular / regular-pi1 / classic / adafruit-hat(-pwm) mappings (Rp1PioConfigSupported in lib/rp1/rp1_pio_backend.cc). For anything else CreateFromOptions returns NULL; the Python binding doesn't check, so the display process crashed on its first call into the matrix and systemd restarted it into the same crash every 10 seconds. - src/pi5_matrix_support.py: the rule and Pi 5 detection, matching the library's /proc/device-tree/model check - DisplayManager raises before creating the matrix, so it is a logged init failure (reported by /api/v3/hardware/status) and fallback mode - the config API rejects those settings on a Pi 5 when a request sets row_address_type, parallel or hardware_mapping - the Display form offers only row address types 0 and 2 on a Pi 5, and warns when a stored value can't be used - CLAUDE.md: re-check the rule whenever the submodule is bumped Review feedback on #586. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
869e36fb2f |
feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback A General-tab toggle (off by default) checks for and installs LEDMatrix and plugin updates once a week, overnight in the configured timezone. - Pre-update checks skip (and report) instead of forcing: local edits or commits, merge/live rebase, no upstream, low disk, missing health check, or a version that was already rolled back. An abandoned rebase (HEAD back on a branch) is cleared, since it would otherwise block every pull. - The pull reuses the Update Code path (now perform_core_update(), which reports dependency install failures as data). - ledmatrix-update-verify.service, started via a .path unit from a request file, restarts the services from its own cgroup, requires them to come up and stay up, and otherwise resets to the previous commit and reinstalls the previous requirements. It runs a copy of the checker taken before the pull. - No SSH needed: switching the toggle on restarts the display service, which (as root) installs the two units from the repo templates for the web user. first_time_install.sh installs them too and takes --enable-auto-update / LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh). - Plugins update after the code passes its check; failures, blocks and rollbacks raise an Overview banner and show under the toggle. Tested end to end on a Pi: web-UI setup, a good update, a broken web service and a broken display (both rolled back), a blocked local edit, and an abandoned rebase found on the device. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(auto-update): address static-analysis findings - Replace the subprocess.CompletedProcess the verifier fabricated for a command that could not start with a plain namedtuple; nothing is executed there, but the scanner flags any CompletedProcess built from variables. - Mark the subprocess imports with the repo's standard B404 annotation (all calls are list-form argv, no shell). - Mark the rollback-failed message as not SQL (B608 matched its wording). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): CI failures on Linux - Keep the setup result when chown fails. CI runs as a non-root user, where chown to the web user raises; that discarded the result file, so the General tab would never learn whether setup worked. Regression test added. - Register the two new /api/v3/system/auto-update routes in the URL map snapshot. - Use utility classes app.css defines (space-y-1, hover:text-red-600). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): address review feedback - Health check: a failed restart command no longer lets the check run against the still-running old process; it counts as a failure (and after a rollback, as a failed rollback). An unreadable restart count is never treated as stable, since a crash loop looks healthy between attempts. - Installer writes the auto_update setting to a temp file and swaps it in, keeping mode and owner, so a running config watcher never reads a truncated config.json. - Verify unit quotes its command-line paths (install folders with spaces); setup refuses folder names systemd would reinterpret (%, quotes, backslashes, control characters) and says so on the General tab. - The auto-update status route no longer returns exception text. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): keep error detail in the status route's 500 test_web_error_detail requires every 5xx handler to log the traceback and return describe_exception(e), which redacts credentials, so failures are diagnosable from the web UI. Dropping it for CodeQL broke that policy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): dismiss route rejects non-object JSON with 400 A JSON array or scalar body made `.get('alert_id')` raise, returning 500. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): let the app-wide handler answer status-route errors CodeQL (py/stack-trace-exposure, #709) flagged the route's own except, which returned describe_exception(e). web_interface/app.py's error handler already logs the traceback and returns the same redacted detail for any unhandled exception, so the local copy is removed: same response, no new exception-to-response flow, and test_web_error_detail's policy still holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f6367d63ae |
security: triage the CodeQL backlog — 129 alerts, three of them live (#561)
* fix(web): escape quotes in every HTML escaper, not just & < >
The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:
value="${escapeHtml(v)}" title="${escapeHtml(v)}"
so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.
It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.
app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.
cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `'`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.
url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).
test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): stop request-supplied names from reaching paths outside their base
Three of the py/path-injection alerts were live, not lint:
* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
whose resolved path *string-prefixed* the plugin directory. Flask's
default converter forbids a slash but not dots, and
get_plugin_directory('..') returned the parent of the plugins directory
because it exists -- so every file under the project root then prefixed
that directory, config/config_secrets.json included. The prefix check
was also wrong on its own terms: with plugin dir "plugin-repos/foo",
"../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
start with "plugin-repos/foo".
* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
file_id of "../../../../etc/something" deleted that file. This is the
one finding in the batch that destroyed data rather than exposing it.
* POST /api/v3/cache/delete passed the body's key through
CacheManager.clear_cache to DiskCache, which joined it as a filename and
called os.remove. Same shape, same result. The guard goes in
DiskCache.get_cache_path, the single choke point get/set/clear share, so
every caller is covered rather than just this route. Real keys are the
stems of files already flat in the cache directory -- that is how
list_cache_files derives them -- so nothing legitimate is turned away.
The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.
Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.
test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): refuse a plugin id that is not a plain name, don't truncate it
pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.
Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(web): say what the plugin web_ui iframe actually is
The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.
Also drops the now-unused os/os.path imports.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): inline url-input's scheme guard at the previewLink.href sink
CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.
Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.
Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): address CodeRabbit findings on the CodeQL triage PR
- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
credential-saving or connect flow runs. NetworkManager only accepts
printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
was previously saved/attempted before nmcli itself rejected it.
- web_interface/blueprints/api_v3/config.py: fail closed when the
plugin config schema path can't be resolved under the plugins
directory (e.g. a symlinked plugin dir). Previously this fell
through with secret_fields left empty, so submitted credentials for
that plugin were saved as ordinary, unencrypted configuration.
- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
splicing the JSON day/column key into an inline oninput="..." handler
string. escHtml() escapes quotes for a normal HTML attribute, but the
browser HTML-decodes the attribute before running it as script, which
undoes that escaping and lets a crafted column name (e.g. from an
uploaded JSON file) break out of the JS string and execute. Cell
edits now travel through data-day/data-col attributes read by one
delegated 'input' listener instead.
While in this file: fixed 6 pre-existing missing-')' typos on
multi-line safeSetHTML(...) calls (already flagged by Biome in this
PR's own CodeRabbit run as syntax errors blocking its lint pass).
These predate this PR (present on main too) but made the whole file
fail to parse in any JS engine, which is a bigger problem than the
XSS finding itself and directly touches the same lines.
Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5137e86d16 |
feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab PR #544's change, ported onto the api_v3 package split (#553). Identical behaviour; only the placement of the new code differs. The original added 508 lines to web_interface/blueprints/api_v3.py, which #553 deletes, so every hunk of it would conflict irreconcilably. Ported by AST: 26 new top-level items sorted to where the split puts each kind -- __init__.py 2 imports, 11 constants, 7 helpers starlark.py 4 routes (/starlark/editor/{apps,status,start,stop}) misc.py 2 routes (/integrations/mqtt-bridge{,/config}) Everything outside api_v3.py -- the Tools partial, the installer scripts, the JS tests -- applied unchanged. Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is regenerated to match, which is exactly what test_api_v3_url_map.py is designed to make you do when routes are added. Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR ships could not be run here -- node is not installed on this machine -- so test/js/dom/test_tools_sections.js is unverified. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST The AST-based port of #544 onto the api_v3 package split dropped `time` from starlark.py's import list. start_pixlet_editor() and stop_pixlet_editor() both call time.time()/time.sleep() directly, so every start (NameError building `state['started_at']`) and every stop that has to wait out the EXIT trap crashed with a 500. No test caught it because the route's own tests mock subprocess.Popen but never actually invoked it before now. Also carries over #544's later fix that this port branched before: env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an operator who had already pinned PIXLET_EDITOR_HOST to loopback, forcing the unauthenticated `pixlet serve` process onto the LAN regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as api_v3.starlark.py's siblings already do for _pkg-owned names. Both fixes route the shared _pkg.time reference the rest of the package's route modules already use for anything a test might need to patch, rather than a bare `import time` local to this file. Ported the existing regression test from #544 (TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's module layout (web_interface.blueprints.api_v3.starlark instead of the old monolithic api_v3 module), which is what caught the NameError. Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this branch and on origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): clear the six lint errors this rebase introduced All six were introduced by rebasing this branch onto the merged blueprint split, not by the split itself. Confirmed by diffing pyflakes output against main with line numbers normalised -- everything else it reports is present on main too and is the package's deliberate re-export pattern. starlark.py used _STARLARK_APPS_DIR three times without importing it (F821). The rebase resolved an import-list conflict as a union of both sides, and that symbol was on neither side of the conflict hunk, so it was silently lost. It is defined in __init__.py and is now imported like its neighbours. This was the only one of the six that would fail at runtime rather than merely lint. __init__.py imported contextlib twice (F811): the cherry-pick added one next to the existing import. Removed the duplicate; the original at line 19 is used. __init__.py imported signal purely to re-export it to starlark.py, so pyflakes saw it as unused (F401). signal is stdlib and does not need routing through the blueprint package, so starlark.py imports it directly and __init__.py no longer does. contextlib stays re-exported because this module genuinely uses it. _read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule this module imports at the bottom for its route side effects (F811). Renamed to `settings`, with a comment saying why, since the name is otherwise the obvious one to reach for. Verified: pyflakes now reports nothing on this branch that main does not, the package imports, all nine route modules load, and 117 routes register, matching the pinned URL-map snapshot. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply Two CodeRabbit findings on the bridge settings endpoint, both of which returned 200 while doing something other than what the caller asked. `request.get_json(silent=True) or {}` turned a missing or unparseable body -- and the JSON literals null, [] and false -- into an empty dict, which then satisfied the isinstance(data, dict) guard on the very next line. The guard was there to reject exactly those bodies. Dropping the `or {}` lets None fail it. The same `or {}` on /errors/clear is left alone: its docstring documents the body as optional, so an absent body legitimately means "use the defaults". The difference is that saving settings has nothing sensible to do with no body. `if data.get('clear_password'):` accepted any truthy value, and the string "false" is truthy in Python -- so a client echoing the field back as a string wiped a password it meant to keep. Now coerced through the package's existing _coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else. test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus a missing body, and clear_password across truthy and falsy spellings. Verified against the unfixed code -- reverting the body guard fails 5, reverting the coercion fails 3. Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials when TLS is off (CWE-319). That is a policy decision about the feature rather than a defect -- unencrypted MQTT on a trusted LAN is common and often deliberate -- so it is raised on the PR for a maintainer call instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: work through the remaining review findings on the editor and bridge allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses the network in cleartext. Refused now rather than merely warned about -- but refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults false, and is coerced like the other booleans so the string "false" cannot switch the guard off. starlark.py:796 -- the supported service runs Flask threaded, so two start requests could each see running=False, each launch an editor, and the second state write replace the first PID, orphaning a process that holds the display down with nothing recording it. The check-launch-write sequence now takes a module-level lock. starlark.py:848 -- if the state write failed the route returned success with an editor running and no PID recorded: status and stop both reported no session while the display stayed down until the timeout expired. It now terminates the process group and returns an error. starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so nothing hands the display back, yet the response said "the display is restarting". After an escalation the display is now restarted explicitly, and a failure to do so returns an error naming the manual step instead of a success. pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the failure lands before the display is stopped rather than after. pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface address are three cases, not two. Collapsing the last two printed a URL saying "localhost" whenever PIXLET_EDITOR_HOST named a LAN address. tools.html:1254 -- escHtml does not encode single quotes, and the app id was interpolated into an inline onclick="startPixletEditor('...')", so a directory containing an apostrophe could break out of the JS string and run script. The handler binds with addEventListener and reads the id from dataset, where it is only ever parsed as an HTML attribute. Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in in both directions. The tools DOM suite gains three guards asserting the edit buttons carry no inline onclick and pass the id via dataset -- those need jsdom and did not run here, so CI verifies them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): log the traceback on the editor state-write failure The 848 fix answers 500 when the session state cannot be written, and logged that at error level -- but without exc_info, so the traceback never reached the log. test_web_error_detail.py guards exactly this: a handler returning 5xx must write an error-level record *with* the traceback and return the sanitized detail, because checking that merely something was logged is too weak. Caught by Core unit tests on the previous commit, not locally: the guard parses every module under web_interface/blueprints/api_v3 as one source, so it only fires once the whole package is read together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bdb9a94033 |
refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181 functions in one module -- 9% of the core by line count and three times the next largest file. It becomes a package of nine route modules grouped by path segment, plus __init__.py for the shared imports, constants, Blueprint and helpers. Every route module decorates the SAME api_v3 Blueprint object, so endpoint names stay api_v3.<function>, the URL map is unchanged and app.py is untouched. Verified: 111 routes before, 111 after, byte-identical rules, endpoints and methods, and every endpoint still on the one blueprint. plugins 3,867 config 1,178 starlark 692 system 619 fonts 452 misc 398 wifi 361 display 326 backup 212 __init__ 1,787 (imports, constants, Blueprint, 56 helpers) Two things the URL-map check could not catch, both found by running the suite: 1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one directory deeper made that resolve to web_interface/ instead of the project root. Nothing failed at import; it surfaced as ~110 tests failing with 404s and "installation script not found", because every path built from it was one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts PROJECT_ROOT/run.py exists so the next move cannot repeat it. 2. Module-attribute patching. Tests do monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route module that binds such a name by value never sees the patch. The shared code therefore stays in __init__.py rather than moving to a _common submodule -- it has to live on the module the tests patch -- and the eleven names tests patch are read back through the package (_pkg.X) instead of bound by value. Those eleven were found by AST-scanning every setattr in the test tree, not by guessing; "time" is among them, used to drive a fake clock through the second-resolution credential-backup filenames. Test changes are confined to what genuinely moved: patch targets that now name the owning route module, imports of helpers, and six tests that scan the api_v3 source as a file and now read the package directory. Full suite: 4,278 passed, 68 skipped, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(api-v3): address CodeRabbit findings from the blueprint-split review Fixes to the api_v3 package split (PR #553), one per finding verified against the actual code: - __init__.py: _redact_credentials only blanked scalar values under a credential-named key; a bare list of secrets under such a key (e.g. tokens: ["a", "b"]) passed through untouched, since the list branch recursed with no memory that its key looked like a credential. Nested dicts still walk normally (a documented, tested behaviour -- a container like secrets: {api_key: ..., note: ...} is a section name, not a value to blank outright), but any value reached under a credential-shaped key is now actually blanked. - __init__.py: the OAuth helper script's raw stderr/stdout went to logger.error unredacted (CWE-532) right next to a comment claiming this was deliberate; the HTTP response already used the existing redact_text helper. Routed the log line through the same helper. - __init__.py / starlark.py: the standalone Starlark manifest fallback (used when the plugin instance isn't loaded) read-modified-wrote manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe (plugin-repos/starlark-apps/manager.py), which already holds an flock for the same file when the plugin is loaded. Added _starlark_manifest_lock, mirroring that pattern, and wrapped every standalone read-modify-write call site in it. The app-config update route also wrote config.json and the manifest as two separate, non-transactional writes (a second, distinct finding at the same call site); config.json is now rolled back if the manifest write that follows it fails. - backup.py: restore options used bare bool() on values from the request, so {"restore_secrets": "false"} restored secrets anyway (bool("false") is True). Switched to the existing _coerce_to_bool helper already used for this exact purpose elsewhere in the package. - config.py: an automated import-rewrite mangled four user-facing validation strings and their neighbouring comments -- "Invalid start time" had become "Invalid start _pkg.time" (and likewise for "end time") in both the schedule and dim-schedule per-day validation paths. - display.py: `import _pkg.time as time_module` -- _pkg is a local alias for the package, not a real importable module, so this raised ModuleNotFoundError whenever a caller restarted an already-running display service via /display/on-demand/start, after the on-demand request was already written to cache. Fixed to `import time`. Audited the rest of the package for the same `_pkg.<module>` import mistake; every other `_pkg.` reference is a legitimate attribute read-through (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken import statement. - fonts.py: validate_file_upload's max_size_mb parameter is silently unused by that helper (it only checks filename/extension) -- the font upload route saved arbitrarily large files as a result. Added the same seek-and-check pattern already used for the sibling .star upload. - wifi.py: two ad hoc, inconsistent bool coercions. POST /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled it. POST /wifi/radio's enabled/force parsing recognized real bool and some strings but not int 1/0 (1 is True is False in Python). Factored one small _parse_bool_ish helper local to this file and used it at all three sites. Not changed: the "unknown/misspelled restore option keys default to True" half of the backup.py finding -- the file's own comment documents that a missing key deliberately means "restore everything," matching the already-existing JSON-parse-failure guard a few lines above it; only the bool-coercion defect was a real bug. Added or extended regression tests for every fix, following each area's existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed on both this branch and origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change) -- no new failures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5 * fix(api-v3): reject unknown restore option keys CodeRabbit's review of the blueprint split (#553) asked that POST /backup/restore reject option keys outside RestoreOptions' known set. The follow-up commit fixed the bool("false")-is-True bug with _coerce_to_bool but never added the key check: a typo'd or renamed key (e.g. "restoreSecrets") is silently ignored by opts_dict.get(key, True), so the flag stays at its True default and secrets get restored despite the caller's request saying otherwise -- with no indication anything was wrong. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb * fix(api-v3): address CodeRabbit findings on the blueprint split - _redact_credentials: blank scalar descendants of objects reached through a credential-owned list (e.g. tokens: [{"value": "secret"}]) regardless of field name -- the existing name-based walk only protected direct dict values under a credential key, not list items. - wifi.py: reject enabled/force/auto_enable_ap_mode values _parse_bool_ish can't recognize (400) instead of silently treating them as False, which could disable Wi-Fi or the radio itself. - Starlark manifest locking: lock a stable manifest.json.lock sidecar instead of manifest.json itself, in both the standalone route path (_starlark_manifest_lock) and the plugin path (StarlarkAppsPlugin._save_manifest / _update_manifest_safe). manifest.json is replaced by an atomic rename on every write, which swaps in a fresh inode; a lock held on the old inode does not exclude a second locker that opens the path afresh right after the rename and gets the new inode, so two writers could race despite each holding "a lock". A sidecar that no write ever touches always resolves to the same inode for every locker. Skipped as stale: the "serialize the complete manifest read-modify-write" finding at api_v3/__init__.py -- every standalone handler that calls _write_starlark_manifest is already wrapped in _starlark_manifest_lock() on this branch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): re-check reconciliation findings by the reconciler's own rules Both CodeRabbit findings on the merge commit, verified against the code first. Major, plugins.py: the stale-findings filter derived its own notion of "in config" and "on disk", and both were looser than the reconciliation module's. set(load_config()) also contains system keys, the secrets-file keys load_config() merges in, and non-dict values; and any directory holding a manifest.json counted as installed even when that manifest does not parse. Either looseness clears a finding that is still true -- and a secrets key read as a plugin is the precise bug the filter exists to stop reporting, so reintroducing that asymmetry while re-checking was the wrong way round. The two extractions now live in state_reconciliation.py as config_plugin_ids() and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys() alongside. _get_config_state() and _get_disk_state() use them too, so there is one definition rather than two that can drift. _get_disk_state() re-reads each manifest for version/name after taking membership from the shared extractor; that costs one extra small read per plugin on a path that runs once per boot. Minor, the new test: the fixture assigned api_v3.config_manager and api_v3.plugin_manager directly. Those live on a module-level blueprint singleton, so the mocks leaked into every later test that imports api_v3 -- pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr, which restores them. This is the same pollution class that made an earlier test in this session break seven unrelated ones, so it is worth getting right. Five cases added for the parity itself: a secrets key, a system key and a non-dict value must not clear an "installed but missing from config" finding, and neither an unparseable manifest nor a .standalone-backup- directory may count as installed. All five fail against the looser version. Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and CodeRabbit all pass. Codacy reads action_required on every commit of this branch including the first, so it is pre-existing and not from this work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |