mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
f6c0fe55d9ebc12543f2b3497bc00f3590b1c472
248
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f6c0fe55d9 |
fix(core): font zip cache, monotonic timers, resolver back-off, and other core/common fixes (#654)
* fix(core): font zip cache, monotonic timers, resolver back-off, and other core/common fixes - font_manager: a .zip font URL is served as its extracted font after a restart (the cached-file check returned the archive first); downloads use requests with a 30s timeout into a temp file + os.replace. - api_helper / sync_manager: rate-limit and heartbeat/leader timeouts use time.monotonic(); last_request_time and the status file's ts stay wall-clock. set_on_new_cycle docstring no longer claims core uses it. - logo_helper: the placeholder uses the same scaled box as a real logo. - permission_utils: one _sudo_bash_candidates() helper (with the sudoers exact-argv rationale) shared by sudo_remove_directory, which now retries the next bash path on a sudo refusal, and install_requirements_file. - dynamic_team_resolver: failed/empty fetch backs off 5 min; duplicate INFO log and contradictory docstring example fixed. - element_style: scale default looked up through element aliases. - background_data_service: cache-hit callback runs outside the lock. - config_arrays: union-aware type check (["array","null"]); stale dotToNested() reference removed. - auto_update_setup: non-dict auto_update reads as off; temp result file unlinked when the write fails. - exceptions: constructors copy the caller's context dict. - logging_config: StructuredFormatter json.dumps(default=str). - error_aggregator: removed unused export_path/export_to_file/_auto_export. - Docstrings: validate_file_upload max_size_mb, raise_on_errors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(sync): retry the status-file rename like the other atomic writers On Windows os.replace can fail with "Access is denied" while a scanner briefly holds the target open; config_manager_atomic._replace already retries that (and re-raises at once on other platforms). The sync status writer called os.replace directly, which made test_concurrent_writers_each_use_their_own_temp_file flaky on Windows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6f45ff5e63 |
fix(display): thread-safety for deferred updates, BDF faces and follower image; one refresh default (#652)
- DisplayManager.defer_update()/process_deferred_updates(): one lock around every queue mutation (appends from the update thread were lost to the render thread's filter/slice reassignments); callables run outside it. - FontManager and element_style no longer cache BDF freetype.Face objects process-wide (load_bdf_face caches them per thread); element_style's LRU is locked against get/move_to_end vs eviction races. - limit_refresh_rate_hz default is one constant, DEFAULT_REFRESH_LIMIT_HZ = 100 (the template's), for the library options, refresh_hz, the matrix guard, Vegas and scroll_config. Previously a missing key capped the panel at 90 while pacing assumed 100. - Sync follower: the TCP thread queues the leader's scroll image; the render thread swaps image/array/width in between frames. - update_display() error log rate-limited (traceback first, then once a minute with a count); swallowed DisplayController exceptions log at DEBUG. - Root display_controller.py runs run.py via runpy. - stream_manager: correct the RLock release comments; merge duplicate if. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
1e62677257 |
fix(vegas): pause for STATIC plugins where their turn falls in the strip (#651)
The static trigger peeked at the front of StreamManager's segment buffer, which continuous scrolling (the default) never advances -- it extends the strip with take_next_group() -- so the same first segment was examined on every frame. A STATIC plugin paused the scroll only if it was first, once, at startup; otherwise it scrolled past as ordinary content. Swap mode had the same problem for any STATIC plugin not first in its cycle. The render pipeline now records a marker (strip column, plugin id) for each STATIC plugin where the strip is built -- composition and every extension -- shifts the markers when the scrolled prefix is trimmed, and clears them on reset. The coordinator pauses when the scroll reaches the next marker: a tuple comparison per frame instead of a lock, a plugin lookup and a get_vegas_display_mode() call. take_next_group() no longer renders STATIC plugins' content. The pause calls display() under the plugin lock and is timed with the monotonic clock. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
224847cebc |
fix(display): stop the run loop spinning when no mode has anything to show (#649)
* fix(display): stop the run loop spinning when no mode has anything to show A mode whose display() reports nothing rotates to the next at once, with no dwell. With every enabled mode empty (only a sports plugin in its off-season, say) the loop went round with no sleep: on ledpi, 169% CPU and ~1,800 "No content" log lines every 10 seconds. After one full rotation of empty passes it now pauses EMPTY_ROTATION_PAUSE (1s) per pass, servicing plugin updates and returning early on on-demand or schedule changes; live priority is still checked at the top of every pass, and the streak resets as soon as any mode shows something. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): restart the loop if the empty-rotation pause starts on-demand; per-rotation streak - an on-demand request serviced during the pause returned early into the on-demand branch, which advanced past the mode just requested; the loop now restarts when the pause changed the mode, on-demand state or schedule - the streak is reset when the rotation changes (on-demand start/stop, a plugin enabled or disabled), so a streak from one rotation can't make another pause before its own modes are tried - docstring: live content is picked up within the pause, not "at once" Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
aeaeaa4e94 |
chore: make contributor tooling work; fix drifted docs (#648)
- mypy.ini parses again (multi-line exclude and inline value comments made mypy reject the file); the mypy pre-commit hook is manual-only until the ~500 existing errors in src/ are paid down, and CONTRIBUTING says so - .gitignore: ignore all of config/ except the templates (ytm_auth.json and others weren't ignored) - .gitattributes: LF for .sh and .service - claude-code-review: skip fork PRs, which have no secrets - check_system_compatibility.sh: 3.13 supported, <3.10 an error - docs/scripts drift: emulator guide, README API Metrics, route count, docs index, scripts README; pyflakes nits in dev scripts Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7eb7a58d0c |
fix: web UI and src.common bugs (wifi wrong-password, plugin icon, starlark toggle, API caching, scroll/logo/font helpers) (#646)
- wifi: keep the "wrong_password:" prefix through the restore/AP fallback so the UI's incorrect-password prompt fires again. - /plugins/installed returns the manifest's icon (string only). - /starlark/apps/<id>/toggle coerces `enabled` and delegates to _toggle_starlark_app (disk before memory, no KeyError, "false" is false). - /api/v3/ JSON GETs are sent Cache-Control: no-store; non-JSON keeps 5s. - ScrollHelper.set_scrolling_image converts non-RGB input (alpha onto black); create/set_scrolling_image reset last_update_time like reset_scroll. - LogoHelper backs off a failed download per path for MISSING_LOGO_RECHECK_SECONDS; cleared on invalidate/clear_cache. - refresh_placeholder_timestamp saves atomically. - FontManager.clear_cache / _clear_plugin_font_cache bump cache_generation. - Odds manager: per-game logs to DEBUG; JSON decode error caught before RequestException (same cooldown). - element_style mangled continuations; startup validator skips null plugin blocks and reuses the controller's discovery. - src/common/README lists frame_timing, json_body, render_gate. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b518c51679 |
fix(plugin-system): load/enable failures, atomic state files, pip lock, test-double parity (#645)
- load_plugin: an on_enable() that raises unregisters the instance, so the next load retries instead of returning True "already loaded". - get_plugin_info: guard plugin.get_info(); one plugin raising no longer breaks /api/v3/plugins/installed. - plugin_state.json and the operation history are written with atomic_write_text under their lock. - plugin_loader: module-level lock serialises pip installs across the parallel startup loaders. - store_manager._install_via_download: extract dir cleanup moved to finally. - Test doubles: draw_image() warns (DeprecationWarning; the real DisplayManager has none), MockDisplayManager.draw_text accepts the real signature's optional params, VisualTestDisplayManager logs draw errors at WARNING. - Docs/comments: compatibility.py method name, PluginState.LOADED meaning, brittle schema count, why _report_skip_once uses setdefault. - Remove unused PluginOperationQueue.get_active_operations(). Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6bc13a8934 |
fix(display): Vegas resumes after live priority, and six smaller runtime fixes (#644)
- Vegas: a live-priority pause was only lifted from inside run_frame(), which returns before that check while paused, so the ticker never came back until a restart. run_iteration() now resumes it (the controller only calls it when nothing preempts Vegas); start()/stop() clear the pause state. Iteration length is timed with the monotonic clock. - Dim schedule: a per-day disabled day now updates the minute-gate cache, so brightness no longer flips back to dim within each minute. - On-demand: a second request no longer overwrites the rotation resume index with the first request's mode. - Render pipeline: reset() drops the prepared group and deferred queue, and a prefetch in flight across a reset discards its result. - Sync: stop() removes the status file (and the controller's cleanup now calls it), standalone removes a stale one at startup, and writes use a unique mkstemp temp file. - render_gate.swap_releases_gil() delegates to frame_timing. - Stale docstrings/comments corrected. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
bcef1957a9 |
fix(security): refuse unsafe plugin ids, keep secrets private, validate request bodies (#643)
* fix(security): refuse unsafe plugin ids, keep secrets private, validate bodies - install_from_url and the registry install's manifest rename refuse a plugin id that is not a single safe name (no ../ out of plugins_dir). - Uninstall and config reset refuse core config sections and ids with path parts; uninstall of a plugin whose directory is gone still works. - separate_secrets checks a field's own x-secret marker before recursing, so object/array secrets no longer land in config.json. - Backup restore creates missing secrets/wifi/ytm files with mode 640; export skips non-object manifests and no longer collides on same-second exports. - SYSTEM_FONTS includes every bundled font from BUNDLED_FONTS. - Raw config/secrets saves and validate_request_json require a JSON object. - A blank max_dynamic_duration_seconds keeps the stored value; other values are validated to 30-1800 instead of raising a 500. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(security): validate the id before install_plugin moves anything; claim backup names atomically - install_plugin set aside plugins_dir / plugin_id before any id check, so "../x" moved a directory outside the plugins dir (the rollback moved it back, but only if the install path got that far) - two exports finishing in the same second could both see a free name and the later os.replace destroyed the first archive; the name is now claimed with O_EXCL before the archive is swapped in Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
da5937da3d |
fix: six bugs found testing main on a real Pi (#641)
* fix: six bugs found testing main on a real Pi (ledpi) - Stopping the service now runs cleanup. systemd stops ledmatrix.service with SIGTERM, whose default action ended Python before run()'s finally block, so the update worker, Vegas and the panel were never torn down. main() now turns SIGTERM into KeyboardInterrupt, the Ctrl-C path. - "Now showing" no longer turns into "unknown". display_current_state was only written on a mode change and the web UI reads it with max_age=120, so a live game or a single plugin on screen for longer read as unknown. It is republished every 30 s while unchanged. - Switching Vegas on in the web UI works when it was off at startup. The coordinator was only created at startup; the config watcher now flags it and the render thread creates it. - configure_web_sudo.sh finds reboot and poweroff in /usr/sbin. Run as the web user it could not, silently dropped their rules and still said it granted them, so the web UI's Reboot/Shutdown stopped working. - check_system_compatibility.sh reports installed packages as installed. `dpkg -l | grep -q` under pipefail failed when grep exited early. - A network failure fetching GitHub repo info logs a WARNING, not ERROR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): fixes found testing on a Pi Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: run the Linux-only script tests correctly The sbin-lookup test set PATH=/nonexistent and then could not find bash itself; call it by absolute path. The dpkg-query stub read $4, but the package name is the third (last) argument. Both now pass on a Pi. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): cover the follower, long-render and startup cases Review follow-ups on the ledpi fixes: - The pending Vegas start is applied before the sync-follower branch too (_apply_pending_vegas_init), which skips _is_vegas_mode_active() while a follower is connected but needs the coordinator for the leader's image. - _service_pending_changes(), which runs inside Vegas iterations and long screens, republishes a stale display_current_state as well; the main loop alone could be away for a 240 s Vegas iteration. - The SIGTERM handler is installed after DisplayController() is built, so a stop during parallel plugin loading keeps the default immediate exit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c4927e82a3 |
fix(config): normalize nullable arrays and objects instead of refusing them (#642)
* fix(config): normalize nullable arrays and objects instead of refusing them
`element_style._nullable` widens every `customization.modes.<mode>` override
with 'null' so a blank means "inherit the base", which turns a colour declared
"array" into ["array", "null"]. `normalize_config_values` only knew how to
convert null/integer/number/boolean out of a union, so a valid [0, 249, 0]
matched nothing and logged
Could not normalize field customization.modes.upcoming.odds_text.text_color:
value=[0, 249, 0], type=<class 'list'>, schema_type=['array', 'null']
The warning was the harmless half. It `continue`d past the single-type handling
below, where `prop_type == 'array'` coerces items, so a nullable array never had
its items normalized while a plain one did. A form posts numbers as strings, so
["0", "249", "0"] survived to the validator and was rejected with "Expected type
integer, got str" -- setting a per-mode colour in the web UI failed outright.
Every per-mode override of a structural or string type was exposed, across all
eight scoreboard plugins, not only colours.
Re-enter the single-type handling with the matched member rather than bailing,
accept a string that matches, and warn only on a genuine mismatch so the
diagnostic still reaches the validator.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ET8e5weDrb5Ju5QTKLU7zh
* fix(config): convert only integral numbers for integer array items
Review catch on the previous commit. Routing nullable arrays into the shared
item handling made its integer coercion reachable for them, and that coercion
called int(v) on any number: a client sending [2.5, 249, 0] for an RGB array
got 2 stored and a 200 back, so a wrong value was silently corrected into a
valid-looking one rather than refused.
Convert only genuinely integral values, at both the union-item and the plain
'array' item branch so the two cannot drift. A whole float -- 2.0, which is all
JSON can express for an integer -- still converts. This also settles an
inconsistency that predates the change: int('2.5') raises, so the string form
was always preserved and rejected while the numeric form was truncated.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ET8e5weDrb5Ju5QTKLU7zh
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9964dd2183 |
feat(vegas): render plugin content off the render thread, and keep it off the GIL when the panel needs it (#630)
DisplayManager.offscreen() gives a thread its own canvas, so Vegas renders every plugin's ticker content on its prefetch thread instead of pausing the scroll for canvas-bound plugins on the render thread. A render gate (src/common/render_gate.py, vegas_scroll.prefetch_gate, on by default with the GIL-releasing binding) lets the prefetch thread run Python only while the render thread waits in SwapOnVSync: on hdpi, frames 2+ refreshes late fell eightfold and late frames overall from 0.90% to 0.60%. See docs/OFFSCREEN_RENDERING.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
865d62f67b |
feat(display): compensate for the panel's scan order while scrolling (#634)
A 1:N-scan HUB75 panel lights the two rows either side of its middle at opposite ends of each refresh, so a scroll at one pixel per refresh shows a 1px step across the middle of every panel. While something scrolls at one frame per refresh, DisplayManager now shows the half whose seam row lights first one refresh behind the other (src/scan_order.py), which lines the two up again. Only for layouts whose row order is known; display.scan_order_compensation "off" disables it. Confirmed on hdpi before and after. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8ad9d191a7 |
feat(perf): frame timing for every presented frame, a soak tool, a render bench and a stall watchdog (#629)
src/common/frame_timing.py times every frame the display presents, whoever drew it, and writes cumulative counters to /dev/shm. scripts/frame_soak.py grades a running service (late frames, freezes, where the time goes) and scripts/render_bench.py the hardware and render path alone. A stall watchdog logs the stacks behind any scroll held up for 250 ms or more (LEDMATRIX_STALL_WATCHDOG_MS lowers that). See docs/SCROLL_PERFORMANCE.md, "Soaking a rig". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7f9c73e9aa |
fix(vegas): smooth Vegas scroll pacing -- whole pixels per refresh, measured refresh, off-thread preview writes (#628)
Vegas scrolls a whole number of pixels per panel refresh, locked to SwapOnVSync, against the refresh the panel really holds (measured from swap gaps), instead of blending sub-pixel positions against the refresh cap. The web preview PNG is encoded off the render thread while scrolling, with writes ordered and retried. On hdpi, late frames fell from 6.3% to 0.7%. See docs/SCROLL_PERFORMANCE.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b9416ef803 |
fix(display): Vegas teardown and 240s default; remove dead Vegas buffer code (#637)
* fix(display): tear down Vegas mode on controller cleanup DisplayController.cleanup() never called VegasModeCoordinator.cleanup(), so the Vegas teardown (stop, pipeline/stream reset, adapter cache drop) was unreachable. Call it before the display manager is cleaned up, and skip it when Vegas was never created. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): default max_cycle_duration to the documented 240s The template, the web UI help, CONFIG_REFERENCE and the controller all say 240, but the code defaulted to 600 in two places, so a config without the key ran Vegas iterations 2.5x longer than documented. from_config now falls back to the dataclass field defaults instead of repeating each one, so the two copies can no longer drift, and the controller's follower scroll-speed default reads VegasModeConfig's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): let run.py -d show display_manager's DEBUG output display_manager pinned its logger to INFO at import, overriding the root level, so debug mode never showed its DEBUG lines. Use get_logger() from src.logging_config like the rest of the core and leave the level to the logging setup. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): run each startup validation check once StartupValidator.validate_all() ran twice at boot, before and after the plugin manager was created, so every config, cache, display and systemd-unit warning was logged twice. The second pass now runs only the plugin checks. Drop the commented-out raise_on_errors line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): one INFO line per plugin-list refresh StreamManager logged "=" * 60 banners and a line per plugin (INCLUDED, SKIPPED, FETCHING CONTENT, SEGMENT CREATED) at INFO on every refresh and fetch, i.e. at each cycle start and every 30s. Log one INFO summary of the rotation per refresh and move the per-plugin detail, the weighting breakdown and "no content this cycle" to DEBUG (the adapter still warns when every content path fails). Also drop the check/cross marks from the controller's log messages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): drop the per-iteration static-mode plugin scan run_iteration() rebuilt _static_mode_plugins on every iteration, asking every plugin for its display mode and logging the set at INFO, but nothing ever read it: static pauses are triggered by _check_static_plugin_trigger() from the next segment. Delete it, the coordinator's get_ordered_plugins() that only it used, and the write-only _static_pause_plugin / _static_pause_start. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): remove the staging buffer that was never filled StreamManager and RenderPipeline carried a double-buffer design that nothing used: _staging_buffer was only ever cleared or swapped, so swap_buffers() never did anything and should_recompose()'s staging_count > 0 branch was dead, and _active_scroll_image, _staging_scroll_image, _is_rendering, _last_frame_time and _frame_interval were written but never read. Delete the machinery and rewrite the docstrings around what actually carries updates: _pending_updates, consumed by process_updates() in swap mode and invalidate_pending_updates() in continuous mode. should_recompose() no longer builds a buffer-status dict every frame. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): tidy the display controller without changing behaviour - Import VegasModeCoordinator locally instead of through module globals (there is no circular import to avoid). - Drop hasattr() checks on attributes PluginManager.__init__ always sets (plugin_executor, plugin_last_update, get_plugin_lock, run_scheduled_updates*, stop_update_worker) and the dead "older manager" fallbacks; keep the health_tracker None checks, now via _health_tracker(). - Extract _display_once() for the per-frame display call both render loops copied, _advance_on_demand() for the two on-demand rotations, _reset_on_demand_fields() for the error and clear paths, and _timezone() / _in_window() for the two schedule checks. - Remove always-true conditions and the unreachable non-plugin else branch in run(), and read _was_display_active / _last_published_mode / vegas_coordinator directly now that __init__ declares them. - Declare the follower render state in __init__, name its tuning constants, add _follower_sign(), and share the 90/s sync send interval with the render pipeline (SYNC_SEND_INTERVAL). - Delete history narration and the "Opt #N" labels, fix the comment that called _scroll_speed constant (hot reload updates it), and drop a startup timing log that measured nothing. - render_pipeline / plugin_adapter: read display_manager.width/height as the properties they are, drop an empty TYPE_CHECKING block, an aliased threading import and a duplicated `if result and self.sync_manager:`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): trim dead code from display_manager - Add _new_canvas() for the image/draw/fontmode="1" setup that was copied six times. - Call resolve_double_sided() and compose_pixel_mapper_config() directly instead of through a module alias and a passthrough method, and replace the comment that said the passthrough read class attributes. - Delete the unused _initialized flag and _ORIENTATION_ROTATE_DEGREES alias (no core or monorepo reader; tests stop resetting the flag), the test pattern's unreachable no-matrix branch (it only runs once the matrix exists), `del old_image # help GC` (a no-op on a local), a duplicated early return in process_deferred_updates, and stale comments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(vegas): remove unread fields and test-only helpers, fix docstrings - ContentSegment: drop total_width, fetched_at, is_stale, image_count and is_static, none of which is read. - StreamManager: drop _current_index (never advanced) and the test-only get_all_content_for_composition() and has_pending_updates(); VegasModeConfig: drop the test-only is_plugin_included(). - geometry.find_blank_cut() has had no production caller since the crop moved to item boundaries; delete it and its tests. - PluginAdapter: the _finalize docstring described separator_width between every image, and _crop_to_budget's said cuts snap to the nearest blank column; both now describe what the code does. - Coordinator: the static-pause interrupt log no longer blames follower mode for every interrupt, and set_update_callback names the callback the controller actually wires. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(scroll): correct ScrollHelper comments and drop dead branches - Four comments said the strip always starts with display_width of blank; it does only when lead_gap is None (Vegas passes its own). - Delete the "Width calculation mismatch" warning: the image is created at the calculated width, so the two can never differ. - Remove the two scroll_delay <= 0 fallbacks (which disagreed with each other): set_scroll_delay clamps it to at least 0.001 and nothing in core or the plugin monorepo assigns it directly. - Trim the scipy history from the blend docstring. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(run): drop a redundant comment Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): log set_scrolling_state only when it changes Vegas and scrolling plugins set the scrolling state every frame, so once display_manager's DEBUG output became visible in debug mode it printed "Scrolling state set to: True" about 120 times a second. Log only when the value differs from the previous one; the state, activity timestamp and frame hold still update on every call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): display-vegas Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f6afbdbb15 |
fix(web): Cache/Logs error mix-up, store errors, tab fallbacks; remove ~2.5k lines of dead JS (#639)
* fix(web): keep Cache and Logs helpers out of each other's way Both partials declared top-level showError and escapeHtml. Their scripts run at global scope after every HTMX swap, so whichever tab was opened last owned window.showError, and a Cache failure after visiting Logs rendered into the Logs panel (and the other way round). Each script is now an IIFE; Cache still exports deleteCacheFile for its row buttons. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): make the HTMX-failure fallbacks for tab panels actually run - The "HTMX never loaded" fallback read appElement.__x.$data, which is Alpine 2. The page ships Alpine 3, so the check was always false and the Overview never loaded without HTMX. It now reads Alpine.$data(). - The Overview and WiFi panels used hx-on::htmx:response-error, which htmx expands to "htmx:htmx:response-error", an event that never fires. - loadTabContent sent requests with <body> as the source, so htmx fired its events on <body> and no panel's hx-on handler ran at all. The panel is now the source. htmx also resolves its promise on a 4xx/5xx, and the panel was stamped data-loaded anyway, leaving a skeleton that never retried; it is now stamped only when no responseError fired. loadPluginsDirect, loadOverviewDirect and loadWifiDirect are merged into one window.loadPartialDirect(id, url), which also runs the partial's inline scripts before Alpine sees the markup (as htmx-config.js does on htmx:afterSwap). The ~10 s "htmx never arrived" path in loadTabContent uses it for every tab instead of four hard-coded ones. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): store and registry failures no longer wipe the Plugin Manager showError replaced the whole #plugins-content with an error message, so one failed store search, custom-registry install or saved-repository call took the installed list, the store and every control with it, with no way back short of reloading the tab. Those failures are now error notifications. The full-panel message is kept only for a first load of the installed list that failed (nothing to show yet); a failed refresh of an already-rendered list is a notification too. showSuccess's fallback branch, which wrote the message into innerHTML unescaped, is gone: showNotification always exists. The "Please try refreshing your browser" hint tested for the text "Failed to Fetch", which no browser produces (Chrome says "Failed to fetch", Firefox "NetworkError..."), so it never appeared. It now keys on the failure itself: a TypeError from fetch(), or PluginAPI's NETWORK_ERROR wrapper around one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): escape plugin action and install output on every path executePluginAction escaped data.message and data.output when an action failed but put data.message straight into innerHTML when it succeeded, and set the OAuth step-2 button's innerHTML from the manifest's step2_button_text. Plugin actions run plugin code, so that is plugin- or server-controlled markup in the page. Both paths now escape, and the button label is set with textContent. The same pattern sat in the install-from-GitHub-URL status lines (plugin_id, the server's message, and error.message, which can echo a repository URL) and the custom-registry load error; those are escaped too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): file-upload widget owns the image list and schedule editor plugins_manager.js loads after the widget bundle, so its older copies of deleteUploadedFile, updateImageList, hideUploadProgress, formatDate, openImageSchedule, toggleImageScheduleEnabled, updateImageSchedule{Mode, Time,Day} and updateCheckboxGroupData replaced the widget's. They are deleted; the widget files are the only definitions. Before switching over, the two sets were diffed and fixed so nothing regresses: - The old copy labelled the schedule/delete buttons for screen readers and lazy-loaded thumbnails; the widget now does both. - The schedule button did nothing on a card rendered by plugin_config.html whenever the image id is a UUID (every upload): the template turns "-" into "_" in the editor's id, and neither JS copy did. Both now use the template's rule. - The widget's "keep the open editor open" copied the editor's innerHTML into the new list. That dropped its event listeners and showed the old values, so after the first change the editor looked live but ignored input. A schedule edit now saves to the hidden input and updates the card's summary in place without re-rendering the list; a list re-render (upload, delete) rebuilds an open editor from the data. Editor controls are routed by one delegated change listener, so there are no per-element listeners to lose. - The old deleteUploadedFile had a JSON branch that removed a #file_<id> element and skipped the re-render. No template or script renders such an element, and JSON uploads are listed through updateImageList like images, so re-rendering (the widget's behaviour) is the consistent one; the branch was not carried over. - The template always renders the summary line (".image-schedule-summary", "Always shown" when unscheduled) so an edit has a line to update. The inline-handler test evaluated plugins_manager.js's updateImageList; test_file_upload_widget.js now covers the widget's list and editor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): delete the unused handleCredentialsUpload Its last caller went when plugin_config.html switched credential uploads to the file-upload widget's handleSingleFileSelect. Nothing in the web UI, the tests or the plugin monorepo references it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): delete dead and shadowed front-end code Nothing calls any of these (checked across web_interface/, test/ and the ledmatrix-plugins monorepo, including hx-*/x-*/onclick attributes): - app-shell.js: the Alpine methods refreshPlugins (it called a nonexistent this.searchPluginStore), loadPluginConfig, savePluginConfig, getSchemaPropertyType, escapeCssSelector, formatCommitInfo and formatDateInfo, and the top-level copies of savePluginConfig, getSchemaPropertyType, escapeCssSelector, formatCommitInfo, formatDateInfo and togglePluginFromTab. Plugin config forms save through hx-post in plugin_config.html. - window.reconnectSSE (app-shell.js); window.updateArrayTableAddButtonState (array-table.js). - toggleNestedSection, defined twice (app-shell.js and plugins_manager.js) and called from nowhere. - plugins_manager.js: the window.initializePlugins wrapper around an IIFE-local origInit that was always undefined, and __pluginDomReady, which was written but never read. - display.html's fixInvalidNumberInputs fallback: app-shell.js defines it before any partial loads. - base.html's window.loadCodeMirror and the two CodeMirror stylesheet preloads, and the .CodeMirror rules in plugins.html. The raw JSON editor is a plain textarea. Also deleted: app-shell.js definitions that a later script always replaced, so they never ran: executePluginAction (plugins_manager.js assigns its own), uninstallPlugin and its pollUninstallOperation (plugins_manager.js), and updateAllPlugins (install_manager.js). vendor/codemirror stays: test/test_web_smoke.py still requests codemirror.min.js as a sample static asset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): call showNotification without checking it exists app-shell.js defines window.showNotification (a stand-in that queues until the notification widget loads) and base.html runs it, deferred, before every other script that notifies: app.js, the utilities, the widget bundle, plugins_manager.js, and all partials, which HTMX loads after the page. The 81 `typeof showNotification === 'function'` / `!== 'undefined'` checks, the `window.showNotification || console.log` and `|| alert` fallbacks, and their else branches (alert(), console output, and schedule.html's own hand-built toast) could never take the fallback path. They are removed, as is fonts.html's second copy of the queueing stand-in. The stand-in in app-shell.js keeps its guard (it must not replace the widget's implementation if load order ever changes), and BaseWidget's public notify()/getNotificationFunction() keep their shape for widgets that plugins ship. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): one HTML escaper, window.LEDEscape About 30 files each carried their own escapeHtml / escapeAttr / escHtml / _esc / escapeJs. They disagreed: several (notification.js, display.html's escapeHtml, operation_history.html, the app() stub) did not escape quotes, google-calendar-picker.js and tools.html's escHtml left ' alone, and some turned 0 into ''. Most were fine only because the quote-safe widget copies were preferred at runtime. window.LEDEscape now lives at the top of app-early.js, a blocking script in <head>, so it exists before any other script runs: html(v) & < > " ' as entities, null/undefined as '' attr(v) the same, for call sites that want to say "attribute" jsStringAttr(v) a JS string literal safe inside an inline handler Every former copy is now a one-line name for it (kept so call sites do not change), widgets included, with no fallback. plugins_manager.js loses its four escapeJs wrappers (callers use jsStringAttr), the duplicate escapeAttr and escapeHtml inside renderInstalledCards and renderCustomRegistryPlugins, and the window.escapeHtml / window.escapeAttribute exports, which nothing read. addArrayObjectItem's fallback markup (with a sixth hand-written escape chain) is gone too: window.renderArrayObjectItem is defined earlier in the same file, so the fallback could not run. The unused escapeHtml methods on the Alpine app (app-early.js stub and app-shell.js) are deleted. test_html_escaping.js now runs LEDEscape and every remaining name for it, and fails if a hand-rolled escaper reappears anywhere in web_interface/. Suites that evaluate slices of plugins_manager.js or widget files load LEDEscape from app-early.js through test/js/led_escape.js. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): stop htmx re-running partial scripts after every tab load htmx-config.js runs each swapped-in <script> itself on htmx:afterSwap and meant to turn htmx's own script handling off with htmx.config.allowScriptTags = false. It did that once, while setting up, but base.html loads htmx with a dynamic <script>, so htmx was not defined yet and the setting never applied. On every tab load htmx then tried to run each script again in its settle phase, found it already replaced (no parent node) and threw "Cannot read properties of null (reading 'insertBefore')" into the console, which also skipped the rest of that swap's settle tasks. The setting is now applied in the afterSwap handler, which always runs after htmx exists and before htmx settles the same swap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): show "--" for a system stat the server could not read The stats stream and /system/status now send null for a metric they cannot read (cpu_temp off a Pi, for one) instead of 0. updateSystemStats built the header and Overview text as value + unit, so a null showed as "null°C". CPU, memory and temperature, in the header and on the Overview, now render "--" plus the unit for null or a missing field -- the same placeholder the page starts with, and what tools.html already shows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): one Alpine accessor and one plugin-list signal window.getApp() (app-early.js) returns the root <body x-data="app()"> component through Alpine's public Alpine.$data, or null before Alpine has initialised it. It replaces the private el._x_dataStack[0] reads in app.js, app-early.js, app-shell.js, settings-search.js, overview.html and plugins_manager.js, the three local getAppComponent/appData/getAppData copies, and the Alpine 2 el.__x.$data fallbacks, which Alpine 3 never provides. Publishing the installed-plugin list: one load set window.installedPlugins and dispatched pluginsUpdated twice (loadInstalledPlugins, then renderInstalledPlugins), then wrote into the Alpine component through _x_dataStack[0] and called its updatePluginTabs() directly, and app-early.js's global listener set window.installedPlugins a third time and called updatePluginTabs() again. Now renderInstalledPlugins is the one publisher: it sets window.installedPlugins and dispatches pluginsUpdated once, and the full app()'s listener (app-shell.js) is the receiver. The app-early.js listener only builds the tab row while the app is not the full implementation yet. The "grid not loaded yet" case is a normal state (Plugin Manager tab not opened), so it logs through pluginLog instead of console.warn. updatePluginTabs had a "Debounce" comment and clearTimeout over a timer nothing ever set, and two identical branches; it now just calls _doUpdatePluginTabs (app-early.js detects the full implementation by that name in its source, which the new comment says). app()'s unused baseComponent lookup is removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): reload the plugin list after installs and failed toggles Several callers refreshed the installed list with if (typeof loadInstalledPlugins === 'function') loadInstalledPlugins(); else if (typeof window.loadInstalledPlugins === 'function') ... but loadInstalledPlugins is local to the plugin-manager IIFE and window.loadInstalledPlugins is never defined, so from outside that IIFE both tests were false and nothing reloaded: - A failed plugin toggle left the switch drawn in the new state while the data said the old one. It now re-renders from the reverted data. The optimistic in-place edit also has to forget the grid's last-rendered markup, or setGridHtmlIfChanged sees identical HTML and skips the revert. A successful toggle still keeps the switch (and focus) as drawn. - Installing from a GitHub URL (the early handleGitHubPluginInstall), installing or uploading a Starlark app, and toggling a Starlark app on its config tab never refreshed the list, so the new app had no tab or Installed badge until the page was reloaded. They now force a reload through window.pluginManager.loadInstalledPlugins(true), and the Starlark grid redraws when that finishes instead of after a fixed 500 ms. - The Starlark uninstall inside the IIFE reloaded from the 3 s cache, which could still hold the app; it now forces a reload. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): route debug output through debugLog base.html defines window.debugLog, gated on localStorage.pluginDebug. plugins_manager.js read the same key twice more into its own flags (_PLUGIN_DEBUG_EARLY, and PLUGIN_DEBUG behind a pluginLog() wrapper), and api_client.js's RequestThrottler had a separate `debug` property with a setDebug() that nothing called. All of it now goes through debugLog. The "functions defined" dumps with their ✓ lines, and two per-plugin "enabled=" loops that ran on every render, are dropped; "[PLUGINS STUB]" labels on code that has not been a stub for a long time read "[PLUGINS]". Ungated console.log calls that announced normal events on every page load or action (settings search and tooltips registering, every toast repeated to the console, the schedule pickers initialising, widget registry unregister/clear) go through debugLog too. What remains on console.log is the widget registry's on-demand LEDMatrixWidgets.debug() dump and BaseWidget.notify's no-notifier fallback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): drop waits and guards that could never fire - handlePluginAction polled up to 10 x 50 ms for window.togglePlugin, configurePlugin, updatePlugin and uninstallPlugin before calling them. All four are defined when the scripts load, before any card can be clicked, so the poll always succeeded at once; it now calls them. The long thinking-aloud comment over the toggle state is replaced by two lines on why the stored state, not the checkbox, decides. - initializePlugins checked typeof on setupGitHubInstallHandlers and applyStoreFiltersAndSort, function declarations in the same IIFE, and wrapped window.checkGitHubAuthStatus(), which returns a promise with its own .catch, in try/catch. - searchPluginStore wrapped each "#store-count" update (a getElementById and an innerHTML assignment) in try/catch four times; one setStoreCount() helper does it. The store's post-render re-attach of the GitHub token handler dropped its try/catch and existence checks for the same reason. - The load-time fallback outside the IIFE tested typeof initializePluginPageWhenReady, which is IIFE-local and so always undefined there; it calls window.initPluginsPage directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): delete two unused plugin-manager helpers stopOnDemand (IIFE-local; the page's stop button calls window.stopOnDemand from app-shell.js) and debounce had no callers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): document the plugin-config handlers templates call validatePluginConfigForm, handleConfigSave, handleToggleResponse, handlePluginUpdate and refreshPluginConfig each get a JSDoc naming the attribute in partials/plugin_config.html that calls it and what the return value means (only validatePluginConfigForm's matters: false cancels the submit). - The `if (!window.__pluginConfigHandlersInitialized)` wrapper is gone: app-shell.js runs once per page, so it was never false. The block is dedented one level; `git diff -w` shows the real change. - The three handlers read xhr.responseJSON first. XMLHttpRequest has no such property (it is jQuery's), so that branch never ran; one xhrJson(xhr) helper parses responseText for all of them, with the same fallbacks as before. - runPluginOnDemand and stopOnDemand checked that plugins_manager.js's openOnDemandModal/requestOnDemandStop exist; plugins_manager.js is on every page, so they call them. - fixInvalidNumberInputs had a stray "Notification helper function" comment on top of its own; a leftover "section toggle ... duplicate definition removed" note is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): one toast per save, and a failed durations save says so app.js's global htmx:afterRequest listener showed the server's message for every htmx request, and every form and button that posts through htmx (plugin config save/toggle/update, Display, Durations, General, Schedule, Dim schedule, the Overview actions) also reports its own result from hx-on after-request. Each save showed two toasts. The global listener now stays quiet for a request whose element, or its form, has its own after-request handler. That exposed the Rotation & Durations form's handler, which read xhr.responseJSON: XMLHttpRequest has no such property, so it always said "Durations saved" in green, even when the save failed (the global toast had been the only place the error showed). display.html already had a correct version (2xx only counts as saved; the server's message wins; its status may refine success but never overturn failure). That is now window.showSaveResult(xhr, savedText, failedText) in app.js, used by the Display, Durations and General forms; General's inline copy of the same logic is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(web): file headers and comments that say what the code does now - plugins_manager.js, app-shell.js, app.js and app-early.js open with a header: what the file owns, how base.html loads it and in what order relative to the others, and the globals it defines. app-early.js's app() stub also says why it exists and that, with app-shell.js now loaded before Alpine, it does not run in practice. - base.html's note on plugins_manager.js said it must load last to win over same-named functions in app.js/app-shell.js; there are none left, so it now gives the real reason (it uses everything loaded before it). - Change-narration and "already defined at the top, no need to redefine" notes are gone or rewritten as present-tense reasons; comments that were wrong are fixed ("Toggle password visibility" over the function that opens the token panel, "Insert before the closing </nav>" over an appendChild, "(from v2)", the export note that still listed escapeHtml). About forty comments that restated the line below them are removed, and a second window.currentPluginConfig = null outside the IIFE is dropped (the IIFE sets it). - The file-upload, checkbox-group and custom-feeds widgets' render() stubs say plainly that the widget is rendered server-side, instead of "for now" / "placeholder for future client-side rendering". test_plugin_action_delegation.js sliced the source up to one of the removed notes; it now ends the slice at the next section header. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): keep the escapeHtml/escapeAttribute globals for plugin pages 6da77363 removed window.escapeHtml and window.escapeAttribute because nothing in core or the plugin monorepo read them. Plugin web UIs served through serve_plugin_web_ui and third-party plugin pages may still call them, so they come back as aliases of window.LEDEscape.html and .attr, defined in app-early.js before any other script runs. test_html_escaping.js checks the aliases exist. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): web-frontend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): encode the image thumbnail path; match script tags case-insensitively CodeQL flagged the upload widget building an <img> src from a stored path, and the escaper test extracting inline scripts with a case-sensitive regex. Each path segment is now URL-encoded (still a same-origin path, and correct for names with spaces or * fix(web): clear Codacy findings in the escaper, app shell and upload widget - LEDEscape looks entities up in a Map instead of indexing an object. - showNotification is declared as a global for app-shell.js. - openImageSchedule checks the index is a non-negative integer and reads the image with Array.prototype.at. - The schedule editor calls escapeHtml directly and documents why its innerHTML template is safe: every value is escaped or constrained. The remaining rule hits are suppressed on that line with the reason. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): build the image schedule editor with DOM calls Codacy does not honour inline suppressions, and the editor's innerHTML template kept tripping its XSS rules even though every value was escaped. The editor is now built with a small element helper (createElement and setAttribute), so no value is ever parsed as HTML, and the file's own escapeHtml goes away. Also for Codacy: - LEDEscape.attr is its own function rather than a second name for html. - The tab loader records a failed load on the panel (data-load-failed) from a named handler, instead of a closure over a local flag. The fake DOM in test_file_upload_widget.js gains append/replaceChildren, its hostile-id check now asserts the id arrives as attribute data with no innerHTML anywhere in the editor, and test_html_escaping.js drops the file-upload.js escaper it no longer has. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): schedule editor helpers as plain functions Codacy's lint flags arrow functions held in local constants and a forEach callback that returns a value. The editor's pieces are now named function declarations (displayStyle, scheduleModeOption, scheduleRangeTime, scheduleDayTime, scheduleDayRow) taking what they need as arguments, and the element helper loops with for...of. htmx is declared as a global in app-shell.js. Output is unchanged; test_file_upload_widget.js passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3a81f38f09 |
fix(web): uniqueItems saves, /health count, Vegas order wipe; one list-repair helper (#638)
* fix(web): drop repeats from uniqueItems lists before validating a plugin save dedup_unique_arrays lost its only caller in #330, so submitting a value a uniqueItems list already holds (a stock symbol saved once and posted again) failed the whole save with a validation error. _prepare_plugin_config_for_save runs it again just before validation, which covers both POST /plugins/config and plugin sections posted to /config/main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): /health counts the discovered plugins and logs the checks it fails The plugin check counted plugin_manager.get_available_plugins(), which PluginManager does not have, behind a hasattr guard that made plugin_count 0 on every device. It now counts the discovered manifests, discovering first when nothing has been scanned yet. The config, plugin and hardware checks answered "see logs for details" without logging anything. Each now logs a warning with the traceback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): store refresh no longer claims a commit-metadata refresh POST /plugins/store/refresh read fetch_commit_info (or fetch_latest_versions) only to append "(with refreshed commit metadata from GitHub)" to its message. It never fetched any: the route re-downloads the registry and nothing else. search_plugins takes the flag, but it reads commit info through its cache, so passing it on would not refresh anything either. The flag is ignored now and the message says what happened. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): refuse a malformed Vegas plugin order instead of clearing it A vegas_plugin_order or vegas_excluded_plugins value that was not JSON, or not a list, was stored as [] and the save answered 200, so a bad value wiped the saved order or exclusions. Both now answer 400 and save nothing, the way plugin_rotation_order already did; the three share one parser. A list that holds anything but plugin-id strings is refused as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): per-plugin health and metrics read the display service's latest GET /plugins/health/<id> and /plugins/metrics/<id> called get_health_summary and get_metrics_summary without force_reload, so they answered with whatever the web process read first and kept in memory, while the display service kept writing newer state. They now pass force_reload=True, as the list routes do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin config reset saves through the shared atomic save POST /plugins/config/reset called config_manager.save_config directly, so it took no backup, and a failed write escaped as an unhandled exception. It then handed on_config_change the raw stored section, not the prepared config a loaded plugin runs with. It now saves through _save_config_atomic with a backup, answers CONFIG_SAVE_FAILED when that fails, and notifies with _prepared_plugin_config, as POST /plugins/config does. POST /plugins/toggle carried its own copy of _save_config_atomic's save_config_atomic-or-save_config fallback; it calls the shared helper now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): one reading and one "unavailable" for each system metric system_metrics.collect_system_metrics() promised None for a metric it could not read, but returned cpu_temp as 0 off a Pi, and the whole no-psutil fallback as zeros. GET /system/status measured the same numbers a second time with its own code, and answered None there. Now both come from collect_system_metrics(), and "unavailable" is None everywhere. /system/status keeps its 0.1s CPU sample and its 10s cache, and gains nothing it did not already send. Two differences: without psutil it answers 200 with null metrics instead of 503, and a disk it cannot stat is null instead of a 500. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): /display/current sends the snapshot as-is and logs a failed read GET /display/current PIL-decoded the preview snapshot and re-encoded it before base64-ing it, spending CPU on the Pi to send the same picture, and dropped any failure with `except Exception: pass`. The /stream/display SSE stream already passed the PNG's bytes straight through. Both now read through web_interface/display_preview.py and answer with the same payload. A missing snapshot is still a null image; any other read failure is logged as a warning. /health reads the snapshot path from the same module. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): one helper puts a submitted plugin config's lists back The plugin-config save turned position-keyed dicts ({"0": ..., "1": ...}) back into lists in five copies: four in the form path's fix_array_structures (whose prefix branches never ran, since no caller passed one), and _fix_json_arrays on the JSON path. It then force-fixed the news plugin's feeds.custom_feeds by name, in case the generic pass had missed it. src/web_interface/config_arrays.coerce_array_shapes now does it for both paths, custom_feeds included. ensure_array_defaults duplicated _fix_none_arrays and is gone. In the same function: the union-type re-checks that the null handling above them made unreachable, the "(temporary)" random_seed debug log, and a commented-out log line are removed. A failed validation is logged once as a warning, not four ERROR lines and a WARNING. Element types are left to normalize_config_values, which already converted them for both paths. One difference: the form path no longer adds an empty {} for a nested object the post left out that has no defaults. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): import at module top and log through the module logger The web_interface.cache imports in config.py and fonts.py were wrapped in `except ImportError` fallbacks. It is an in-repo module that imports nothing from the project, so it cannot fail to import; it is imported once at module top, as system.py now does. cache.py's docstring said blueprints import it lazily "to avoid circular imports"; it now says why that is unnecessary. Five logging.error calls in the dim-schedule GET and three logging.warning calls in plugins.py went to the root logger; they use the module logger. Function-local re-imports of json, os, shutil, logging and Path, all already imported by the module, are gone. The `import os` inside two except blocks of save_plugin_config also made os a local name for the whole function. execute_plugin_action's step-1 handler gets a comment saying why it stays: it looks like a copy of the blueprint handler, but without it a TimeoutExpired from the plugin's script would reach the route's own `except subprocess.TimeoutExpired` and be answered as a 408. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): app.py loses dead CSRF and reconciliation state, comments fixed - csrf was always None, so `if csrf: csrf.exempt(...)` never ran, and its note that the api_v3 blueprint "is exempted above" named an exemption that does not exist. Both are gone; the reason there is no CSRF protection stays, shortened. - The SSE rate-limit comment called the default "tight" at 20 per minute. The default is 1000 per minute and the streams' 200 is the tighter one; the comment now says so. The limits are unchanged. - _reconciliation_done was written and never read. The docstring that explains why reconciliation runs once keeps its reason, in the present tense. - Removed: a dangling "import cache functions" comment with no import under it, a "security check ... within project_root" label on an existence check, the "(simplified version)" narration, and the note that no redirect route is needed. The preview loop's sleep comment no longer mentions a PIL encode that the loop does not do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(web): api_v3 comments name the package __init__, not a _common module Every route module's docstring said the shared blueprint comes "from ._common", a module the package split never created; they name the package __init__. The PROJECT_ROOT comment described the path from _common.py; it now describes this package and keeps the incident it guards against. The "(corrected) in this commit" note in resolve_pull_command and the /health comment the split's mechanical time -> _pkg.time rewrite garbled ("Stamp the start _pkg.time") read correctly again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): drop hasattr checks for attributes PluginManager always has PluginManager.__init__ sets health_tracker and resource_monitor (to None until they are configured), so the seven hasattr(api_v3.plugin_manager, ...) guards in the health, metrics and limits routes were always true. The falsy checks that do the work stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): pages_v3 dispatches partials from a dict with one error handler load_partial chose a loader through a fourteen-branch if/elif, and thirteen of the loaders then wrapped themselves in the same try/except, logging "Error loading partial" without saying which. The route now looks the name up in _PARTIAL_LOADERS and has the one handler, which logs the partial's name. The loaders just render. _load_tools_partial keeps its own messages. The search index's _partial_html already catches a loader that raises. serve_plugin_web_ui repeated _plugin_dir_for inline (containment plus the ledmatrix- prefix fallback); it calls it now. Also removed: the unused markupsafe.escape import, function-local json/Path re-imports, and unused exception bindings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove unused imports, locals and a try that cannot fail - get_error_aggregator was imported by the api_v3 package and used by no one; seven names config.py imported, and Path in misc.py and logging in plugins.py, likewise. - branch_info in install_plugin was built and never logged; test_config in /health was bound and never read (the load_config call is the check). - An f-string with no placeholders in the asset upload route. - _installed_plugin_ids wrapped list(manifests.keys()) in try/except; _discovered_plugin_manifests always returns a dict. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): start.py logs its startup lines and drops unreachable branches The startup banner went to stdout with print(); it goes through a logger now, which the app import has already configured, so it reaches the journal with a level and timestamp like every other line. The "no addresses" branch is gone: get_local_ips() always returns at least "localhost". The except around app.run re-raised "only if it's not a client disconnection error" from inside the branch that had just established it was one, so that raise could not run. It is one check now, on a named tuple of the errnos, which the werkzeug log filter uses too. The comment on threaded=True counts three SSE endpoints, which is how many there are. Trailing whitespace is stripped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): save_main_config names its General fields once The General tab's field names were listed twice, once to detect a General form post and again, with four more, to keep the remaining-keys merge from storing them as top-level keys. GENERAL_FIELDS and _MAPPED_TOP_LEVEL_FIELDS hold them now, and the four per-section skip checks are one set. The comment on that merge said plugin configs are handled "here too", and "(including plugin keys)". Plugin sections are handled and removed from the body before it runs; the comment says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): plugin directories come from the plugin manager only Six lookups fell back to PROJECT_ROOT/plugins/<id> when there was no plugin manager: GET /plugins/config's of-the-day data, POST /plugins/action, the plugin static-file route, the calendar credentials upload and the calendar OAuth routes. The loader never scans plugins/ (PluginManager.discover_plugins reads only the configured directory, plugin-repos by default), so what they found there was a plugin that never runs. _plugin_directory() asks the manager and answers None without one, which each route already reports as "not found". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): web-backend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7b90759252 |
fix: /errors stack traces, Wi-Fi disconnect and save, plugin fonts, API cache TTL (#636)
* fix(errors): record the exception's own stack trace record_error() called traceback.format_exc(), which only sees an exception while its except block is running. plugin_executor records exceptions caught on a worker thread after that block has ended, so every trace on /errors read "NoneType: None". The trace is now built from the exception's __traceback__. The executor's log call had the same problem with exc_info=True and now passes the exception. record_error() also merged LEDMatrixError context into the caller's dict in place; it now works on a copy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(wifi): point at configure_wifi_permissions.sh instead of a sudoers list The module docstring told users to grant NOPASSWD sudo on iptables and ip. configure_wifi_permissions.sh refuses those grants on purpose: a wildcard rule for either runs an arbitrary program as root. Point at the script and say why it leaves them out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): disconnect finds the saved profile by SSID disconnect_from_network() asked `nmcli -f NAME,802-11-wireless.ssid connection show` for the profile to take down, but nmcli rejects that column for `connection show`, so the lookup always failed and only the device was disconnected. The per-profile lookup _connect_nmcli() already used is now _find_profile_for_ssid(), and both callers share it. It also splits terse output on the last colon and unescapes "\:", so a profile name containing a colon is found. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): write wifi_config.json atomically and report a failed save _save_config() opened the file for writing in place and swallowed any error, so a wifi_config.json left owned by root made the web toggle for auto-enabling AP mode report success while nothing was saved, and a crash mid-write could truncate the file. It now uses atomic_write_json, which also keeps the file's owner and shared group when root saves it, and returns False on failure. POST /wifi/ap/auto-enable answers 500 in that case. The file is now written with indent=4, like the other config files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): resolve plugin:// fonts in the plugin's own directory FontManager looked for a plugin's bundled fonts under Path("plugins") / plugin_id: relative to the process cwd, and not the default install directory (plugin-repos/), so a manifest's plugin:// fonts never loaded. register_plugin_fonts() takes an optional plugin_dir, and PluginManager passes the directory it loaded the plugin from. Callers that omit it get a lookup in the configured plugin_system.plugins_directory, then plugins/, resolved against the install root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(api-helper): cache responses for the requested cache_ttl APIHelper.get(cache_ttl=...) and set_cache(ttl=...) dropped the ttl on the claim that CacheManager does not support one, but CacheManager.set() takes a ttl, stores it with the entry, and both cache tiers honour it over a reader's max_age. Without it every response expired after the 300-second default read age, whatever the plugin asked for. The ttl is now passed through, and the cache read passes cache_ttl as max_age for entries written without one. The class docstring describes what the helper actually does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(style): one scale range for the schema, element_scale and LogoHelper The generated Scale field allowed 0.1 to 10, element_style's reader capped at 10 with no floor, and LogoHelper accepted 0.05 to 8 and reset anything else to 1.0. A logo scale of 9, which the form accepts, drew at the shipped size. MIN_ELEMENT_SCALE / MAX_ELEMENT_SCALE (0.1, 10.0) in src.element_style are now the schema bounds and the clamp every reader applies through coerce_scale(): a positive number outside the range is clamped, and anything that is not a finite positive number means the default. That also stops element_scale() passing NaN through, since min(nan, 10.0) is nan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(logos): placeholder lands at the requested path; empty logos list download_missing_logo() wrote its fallback placeholder to <normalize_abbreviation(abbr)>.png in the logo directory rather than to the logo_path the caller passed, so it could return True while nothing existed where the plugin looks (e.g. "TA&M.png" vs "TAANDM.png"). create_placeholder_logo() takes an optional filepath, and download_missing_logo passes the requested one. download_missing_logo_for_team() only caught KeyError, so a team whose "logos" list is empty raised IndexError; it now treats KeyError, IndexError and TypeError as "no logo URL". The placeholder is drawn with PLACEHOLDER_SIZE / PLACEHOLDER_BG, the constants is_placeholder_logo() recognises it by, instead of repeated literals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): resolve bundled font paths against the install root TextHelper's default font_dir, the logo placeholder's font and FontManager's font_overrides.json were all relative to the process cwd, so a process started anywhere but the install root (the plugin safety harness, a manual run, a unit without WorkingDirectory) drew with PIL's default face and read no overrides. They now go through font_layout.resolve_asset_path; the overrides file sits in the install root's config/. The resolver docstrings described an order the code does not follow: resolve_asset_path never consults the cwd, and sports_shared's _resolve_font_path tries the cwd first. Both docstrings now say what the code does, and _resolve_font_path calls resolve_asset_path instead of probing FontManager for it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(sync): the web UI reads the sync status file the display writes sync_manager writes its status to tempfile.gettempdir(), but GET /api/v3/sync/status read a hardcoded /tmp/led_matrix_sync_status.json and defaulted the port to a literal 5765. Wherever TMPDIR is set (or on any non-/tmp host) the page only ever showed "starting". The endpoint now uses sync_manager.STATUS_FILE and SYNC_PORT. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(http): the rankings resolver sends the project's User-Agent DynamicTeamResolver fetched ESPN rankings with a bare requests.get, so it sent python-requests' default User-Agent, which ESPN rejects; the AP_TOP_N favourites then resolved to nothing. It now sends DEFAULT_HTTP_HEADERS. BaseOddsManager carried its own copy of the User-Agent string and now uses the same shared headers (which also adds Accept-Language). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(backup): record the core release and read the configured plugin dir The manifest's ledmatrix_version came from a VERSION file that does not exist, then from .git/HEAD: a 12-character sha, or "ref: refs/he" when the branch's ref was packed. It is now src.__version__. list_installed_plugins() scanned a hardcoded plugin-repos/, so on an install whose plugin_system.plugins_directory points elsewhere, plugins missing from plugin_state.json were left out of the backup. It now reads the configured directory from config/config.json, defaulting to plugin-repos. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(startup): report a missing display section once A config without a display section produced three errors for the one problem ("Missing required configuration key: display", "Display configuration is missing or empty" and "Display configuration is missing"), and an empty one produced two. _validate_config now reports it once, as a missing key or an empty section, and _validate_display_config leaves it to that. The module docstring said the validator fails fast; nothing in the display service calls raise_on_errors(), so it now says the errors are reported and startup continues. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(wifi): share the copied blocks and name the AP constants - _parse_nmcli_wifi_list() is the one parser behind _scan_nmcli and _scan_nmcli_cached. - _verify_connected(), _wait_for_device_idle(), _failsafe_ap() and _mark_forced() replace blocks that were pasted two or three times in the connect and enable-AP paths. The device-idle wait now checks before its first one-second sleep instead of after it. - _check_command() calls _find_command_path() instead of repeating it. - AP_IP, PORTAL_PORT, AP_PROFILE_NAME and AP_PROFILE_NAMES name values that were spelled out 14, 12, 8 and 2 times; the two deletion loops now walk the same tuple. The iwconfig status path compares the AP address exactly: startswith() also skipped 192.168.4.10-19. - Dropped a second WIFI.SIGNAL query that repeated the first, a no-op "if ssid: continue", the try/except around _connect_wpa_supplicant's constant return, and a second save of a scan scan_networks already saves. - _ensure_wifi_radio_enabled's docstring says it returns True when the radio state cannot be read at all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(config): drop dead branches and history comments in ConfigManager - The module docstring pointed plugin authors at update_plugin_config(), which does not exist; it now names save_config_atomic() and save_raw_file_content(). - load_config's FileNotFoundError handler tested the message for "config_secrets.json", but a missing secrets file is handled where it is read, so only config.json reaches it; the check is gone. - save_raw_file_content's `file_type == "main" or "secrets"` guard was always true (anything else raised earlier). - get_raw_file_content('secrets') already returns {} for a missing file, so the os.path.exists() in front of two calls to it is gone. - Comments that narrated earlier behaviour are rewritten as what the code does now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(background-data): present-tense comments, drop unused API - Comments that told the history of each fix (what "used to" happen, "the old per-delivery release") now state the invariant the code keeps. - get_statistics() no longer reports a constant 'queue_size': 0, and the uncalled clear_completed_requests() is gone (_cleanup_completed_requests does that job on every completion). Neither is referenced in core, the web UI or the plugin monorepo. shutdown_background_service() has no production caller either, but it is the only way to tear down the get_background_service() singleton, which the tests rely on, so it stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(odds): drop the unread cache_ttl and merge the odds_data branches BaseOddsManager loaded base_odds_manager.cache_ttl from config and never used it: cached odds live for the update interval (get_odds' ttl=interval). No core or monorepo code reads the attribute, so it is gone along with its log line. The two consecutive `if odds_data:` blocks are one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(backup): one table for the single-file sections config, secrets, wifi and ytm_auth were each spelled out in create, preview, validate and restore. _SINGLE_FILE_SECTIONS lists them once, with the RestoreOptions flag that restores each, and all four walk it. Restore error messages keep their wording ("Failed to restore <file name>"). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(fonts): drop FontManager's write-only state and duplicate logs - fonts_config, font_metadata and font_dependencies were written and never read; the performance_stats keys font_load_times, render_times, total_renders and the per-call "resolve" timings (_record_performance_metric) likewise. get_performance_stats() reads only the counters that remain. Nothing in core or the plugin monorepo references any of them. - A failed BDF load was logged twice, by _load_bdf_font and again by get_font; get_font's line is the one kept. - Removed "NEW:" and commented-out cozette entries, the "Copy font to assets/fonts" comment on code that copies nothing, and local imports of names the module already imports. The deprecated add_font() now resolves assets/fonts against the install root. The @deprecated methods stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(text-helper): cache loaded fonts; drop the pre-textlength fallback TextHelper declared _font_cache, cleared it and reported its size, but never stored anything in it. load_fonts() now keeps each (file, size) it loads there, so clear_font_cache() and get_font_cache_stats() mean what they say and repeated load_fonts() calls reuse the fonts. get_text_width() no longer catches AttributeError for Pillow releases without ImageDraw.textlength; requirements.txt pins Pillow>=12.2. The class docstring describes what the helper does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(common): fix wrong docstrings in api_helper, permission_utils, snapshot_policy - permission_utils called 0o2775 "sticky bit"; the 2 is setgid, which is what makes new files take the directory's group. - snapshot_policy pointed at web_interface/blueprints/api_v3.py, which is a package now; the health check is in api_v3/misc.py. - APIHelper.clear_cache() lost a history note and a fallback to a clear() method that neither CacheManager nor the testing MockCacheManager has. The session headers are built from DEFAULT_HTTP_HEADERS instead of a copy of them, and the module docstring says what the module offers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(sports): present-tense comments in the shared scoreboard renderers - sports_scroll and sports_game_renderer comments that referred to "this PR", "the old flat 128px card" or what the renderer "previously" did now describe the current behaviour and its reason. - The block explaining why non-finite settings are rejected sat above _score_reserve_width; it describes _center_gap_width and now lives in it. - unshare_element_fonts wrapped its import of font_layout.load_truetype in an `except ImportError` that cannot fire inside core; the import stays at call time so tests can spy on the pinned loader. - sports_card docstrings that told the history of a fix say what the code does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sports-shared): drop dead code, name the ESPN limit - _get_weeks_data asked for limit=1000, which fetch_espn_scoreboard clamps to ESPN_MAX_LIMIT anyway; it now names that constant. Its unused `immediate_events = []` is gone. - _get_season_schedule_dates() returned ("", "") and has no caller in core or the plugin monorepo. - _should_log keeps its warning_type parameter (part of the inherited signature, though nothing in core or the monorepo calls it) and its docstring says the cooldown is shared across types. - An unused ImageFont import is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sync): one follower-mode switch, shared panel defaults - The class docstring said the leader sends PNG frames. Frames go over UDP as raw RGB; PNG is only the Vegas scroll image sent over TCP. It now describes both paths. - _enter_follower_mode() replaces the two copies of "note the leader, switch from standalone to follower, log, write status" in the frame and scroll-position handlers. - The rows/cols fallbacks use DEFAULT_ROWS / DEFAULT_COLS from src.display_geometry, as chain_length already did. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(style): drop _layout_axis, name the layout group title - ElementStyleResolver._layout_axis() had no caller in core or the plugin monorepo. - _element_block_from_spec checked spec['size'] was a dict again after size_spec already had; it reads size_spec. - The "Layout Offsets" title written into three generated schema blocks is _LAYOUT_TITLE. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(logo-helper): say what the placeholder draws; name the 1.5 box factor - _create_placeholder_logo's docstring said it draws the team abbreviation; it draws an outlined grey box and nothing else. The docstring says so, and the "in a real implementation you'd want text" comments are gone. - The 1.5 x panel default logo box, written out six times, is DEFAULT_LOGO_BOX_FACTOR. - ImageDraw is imported with Image at the top of the module. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(logos): drop dead code and a duplicate regex in logo_downloader - _SAFE_LEAGUE_CODE_RE was the same pattern as _SAFE_LEAGUE_RE; both checks use the one. - get_logo_filename_variations reassigned the TA&M case to the list it already had; the function returns the two names directly. - _get_team_name_variations() had no caller in core or the plugin monorepo. - fetch_single_team's docstring was copied from fetch_teams_data; a log message read "for{team_id}". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: drop the Pillow<9.1 resample shim and a catch-and-reraise - adaptive_images fell back to Image.LANCZOS/NEAREST for Pillow < 9.1; requirements.txt pins Pillow>=12.2. RESAMPLE_LANCZOS and RESAMPLE_NEAREST keep their names (src.common re-exports them). - CacheManager.save_cache caught CacheError only to re-raise it; the disk write is now called directly, with the same result. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(api-helper): stop the real CacheManager's cleanup thread The cache-lifetime tests built a CacheManager and left its cleanup thread's class-wide claim on the directory in place, which broke test_cache_cleanup_thread_ownership when it ran later in the session. The fixture now stops the thread on teardown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): core-common Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b11bcfa204 |
fix(plugins): store and plugin-manager bugs; tidy src/plugin_system (#635)
* fix(store): don't read a ZIP-installed plugin's remote from the LEDMatrix repo update_plugin looked up remote.origin.url with `git -C <plugin> config --local` for plugins that are not git checkouts. Under plugin-repos/ git walks up to the enclosing LEDMatrix repository, so the lookup returned LEDMatrix's own URL and a plugin missing from the registry was "reinstalled" from the LEDMatrix repo. Only ask git when the plugin directory has its own .git, the test _get_local_git_info already uses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(schema): report each missing required field once, by name validate_config_against_schema ran its own required-fields loop after Draft7Validator.iter_errors, which already yields one `required` error per missing field, so every missing top-level field was listed twice. The validator's copy also printed the schema's whole `required` list ("Missing required property '['api_key', 'city']'") instead of the field. Drop the loop and take the field name from the error itself. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(store): stop mangling repository URLs that contain ".git" install_from_url and fetch_registry_from_url cleaned URLs with `rstrip('/').replace('.git', '')`, which removes ".git" anywhere: https://github.com/user/my.github.io became .../myhub.io, so installing or browsing that repository asked GitHub for one that does not exist. Add src/plugin_system/repo_urls.py with one anchored normalize_repo_url(), same_repo() for comparisons, github_owner_repo() and github_api_headers(), and use them for the five copies of the owner/repo parsing and GitHub headers in the store and for saved repositories. GitHub URLs are now recognised by urlparse().hostname everywhere: _get_latest_commit_info used a substring test, and _install_from_monorepo_api parsed any host. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(store): install a repository whose only branch is not main/master _install_via_git returned None both when every clone failed and when the last-resort clone of the repository's default branch succeeded. _install_plugin_impl papered over it with `and not plugin_path.exists()`; install_from_url did not, so a repository whose only branch is e.g. `develop` was cloned, then treated as a failure, then "downloaded" from main/master archives that do not exist. After a default-branch clone, return the branch the clone checked out (read from .git/HEAD), so None means failure and nothing else, and give both callers the same `branch_used is None` fallback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): judge the memory limit on each call's own growth monitor_call stores `metrics.memory_mb = max(previous, growth)`, and _check_limits compared that high-water mark with max_memory_mb. It never decreases, so once one update() grew the process past the limit every later call raised ResourceLimitExceeded and the circuit breaker kept reopening. Pass the call's own RSS growth to _check_limits; keep the high-water mark for reporting and document what it measures. Remove ResourceMetrics.update_average_execution_time: nothing called it, and it overwrote the running total with the average. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): reload_plugin re-reads the manifest from the discovered directory reload_plugin read `plugins_dir / plugin_id / "manifest.json"`, ignoring the discovery map and the plugin_dirs rules. For a plugin whose directory name differs from its manifest id the path did not exist, the re-read was skipped without a word, and the reload kept the stale manifest. Resolve the directory with find_plugin_directory, as load_plugin does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): drop the always-null last_display from plugin state info PluginStateManager reported `last_display` from `_last_display`, which nothing ever wrote, so it was null for every plugin. Recording it in PluginExecutor.execute_display would not help: get_state_info's only reader is the web process, whose PluginManager never calls display(). Remove the field, its dict and get_last_display() (no caller in core, the web UI or the plugin monorepo). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(store): share the rollback and requirements helpers, drop dead code - install_plugin and _reinstall_with_rollback set aside, discard and restore the old copy through _set_aside/_discard_backup/_restore_backup instead of two copies of the same blocks. - The loader and the store run the same pre-pip checks through contained_plugin_dir() and requirements_to_install() in plugin_loader. They still invoke pip differently (sys.executable -m pip vs. the sudo wrapper). `except (BrokenPipeError, OSError)` + `isinstance(e, OSError)` becomes `except OSError` checking errno.EPIPE. - load_module never returns None, so load_plugin's check is gone and the docstring says what it raises. - Remove the always-true JSONSCHEMA_AVAILABLE, the inline re-imports of re and permission_utils, the fake status_result object nobody reads, hasattr(git_error, 'cmd'), a redundant "merge conflict" test and `import traceback` (exc_info=True does it). - Correct comments: install_from_url names the directory for the caller's id when given (not always the manifest id), _get_local_git_info saves one git subprocess (not four), _enrich calls two helpers, search_plugins documents all its arguments, _find_plugin_path states its behaviour instead of a TODO, and history narration is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): tidy base_plugin, correct plugin_manager/state comments - base_plugin: drop the unused `import logging`; get_display_duration runs the instance value and the config value through one _positive_seconds() helper instead of two copies of the coercion; the 'static'/'none'/fallback branches of get_vegas_display_mode, which all returned FIXED_SEGMENT, are one; fix the mis-indented validate_config example; say that get_supported_vegas_modes/get_vegas_segment_width are not consulted by core (kept, plugins override them). - schema_manager: import expand_style_elements normally rather than swallowing an ImportError of a core module. - plugin_manager: the plugins directory is the configured one (plugin-repos/ by default), not plugins/; get_config() returns the live dict, not a copy, so the interval cache comments say what it saves. - state_manager: config_version and the file version are not used to detect corruption; say what they are. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): stop writing data/plugin_operations.json PluginOperationQueue wrote its finished-operation history to data/plugin_operations.json after every operation, and read it back only into its own in-memory list, which only get_operation_history() exposes -- and nothing calls that. The operation-history endpoint reads OperationHistory (data/operation_history.json). No code in src/, web_interface/, scripts/ or test/ reads the file. Drop the history_file/lazy_load parameters and the load/save code; the bounded in-memory history stays. web_interface/app.py and the integration test stop passing the removed arguments. An existing data/plugin_operations.json is left in place (data/* is gitignored). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): plugin-system Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3967a6cffc |
fix(security): re-harden root sudo helpers; installer fixes; ARCHITECTURE and PERMISSIONS docs (#640)
* docs: add ARCHITECTURE and PERMISSIONS guides ARCHITECTURE.md maps the processes, the state the display and web services share through the cache, the display loop, the plugin system, the web UI and the update path, with links into the code and a where-to-start table. PERMISSIONS.md lists who owns what after install, both sudoers files (and why iptables is not granted), the polkit rule, and which scripts/fix_perms script to run as which user. Both are linked from the docs index, along with the MQTT bridge README and src/common/README.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: correct stale setup, service and troubleshooting claims - README: quick actions run systemctl on ledmatrix.service (run.py), not display_controller.py; use_short_date_format has no effect; the installer uses system pip with --break-system-packages, not a venv. - CONFIG_DEBUGGING: LEDMATRIX_DEBUG must be "true"; logs are in journald. - GETTING_STARTED, WEB_INTERFACE_GUIDE, TROUBLESHOOTING: enabling a plugin, plugin settings, brightness and Vegas settings apply without a restart; matrix hardware settings still need one. - TROUBLESHOOTING: install dependencies with sudo so the root service sees them; point permission problems at PERMISSIONS.md instead of a project-wide chown. - ADVANCED_FEATURES: real BackgroundDataService stats keys; Vegas hooks return VegasDisplayMode and None falls back to capture; cache files are 0660; fix_web_permissions.sh runs as the web user and does not touch sudoers. - STARLARK_APPS_GUIDE: only the linux-arm64 pixlet binary is downloaded. - HOW_TO_RUN_TESTS: test class examples that exist. - CLAUDE.md: PluginStoreManager, plugin_dirs.py, monorepo installs via the Trees API with ZIP fallback, requirements.txt is optional. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: mark deprecated plugin APIs and state manifest fields once Methods @deprecated("3.7.0") (the set pinned in test_deprecation.py) were shown as current API in the quick reference, API reference, advanced guide, development guide and FONT_MANAGER. Each is now marked deprecated with its replacement. FONT_MANAGER is rewritten around the current API; the override editor is gone and override methods are deprecated. Required manifest fields were stated three different ways. The API reference now has one section: the 7 schema-required fields, the 4 the store refuses without, class_name for the loader, and the 8 to set. The other guides link to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: document every src/common module and every widget - src/common/README.md covered 7 of 17 modules. It now has a table of all of them (purpose, whether plugins import it, release to floor on), a short entry each, and logging advice that matches the code. - SPORTS_UNIFICATION listed two shared modules and called sports_helpers the first; it now lists all six. - The widgets README lists all 28 registered widgets plus the support files, and absorbs the parts that only docs/widget-guide.md had (x-options.labels, x-advanced, x-display hidden, plugin-file-manager). docs/widget-guide.md is now a pointer to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(security): fix_web_permissions.sh re-hardens the root sudo helpers The script chowns the whole project to the web user. That included scripts/fix_perms/safe_plugin_rm.sh and safe_pip_install.sh -- the two helpers /etc/sudoers.d/ledmatrix_web lets the web user run as root -- so running it turned both into a root shell for whoever can edit them. It also re-grouped config_secrets.json away from ledmatrix. After the chown it now does what first_time_install.sh's Steps 11 and 11.1 do: helpers back to root:root 755, and config_secrets.json back to the web unit's User=:ledmatrix 640. Each step is non-fatal and prints the manual command if it fails. Also fixes what the script and its docs claimed: it never configured sudoers, its closing hint pointed at ./configure_web_sudo.sh (wrong path), and the README and ADVANCED_FEATURES.md said to run it with sudo, which it refuses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(security): validate and harden every sudoers drop-in the scripts write configure_wifi_permissions.sh copied its rules into /etc/sudoers.d/ledmatrix_wifi without `visudo -c`. A malformed drop-in makes sudo refuse every command for every user, which on a headless Pi leaves no way back in. It now checks first and leaves the installed file alone when the rules do not parse, as the other two writers do. (It already used mktemp, so that part of the review did not apply.) It also grants the two literal commands wifi_manager.py runs for NetworkManager's shared-mode dnsmasq drop-in -- `cp /tmp/ledmatrix-nm-dnsmasq.conf .../dnsmasq-shared.d/ledmatrix-captive.conf` and `rm -f` of that file. The directory's mkdir was granted, the file was not. Both are pinned in test_sudo_allowlist_covers_calls.py. configure_web_sudo.sh wrote its rules to /tmp/ledmatrix_web_sudoers_$$, a predictable name in a world-writable directory; it now uses mktemp with an EXIT trap, as first_time_install.sh does. It sets mode 440 on the installed file instead of leaving the temp file's mode, and finds visudo in /usr/sbin when that is not on the user's PATH, which skipped the check silently. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(install): escape the project path in the DNS-fix and MQTT unit renderers install_dns_fix.sh and install_mqtt_bridge.sh substituted __PROJECT_ROOT_DIR__ with the raw path, while the other three renderers go through sed_escape_replacement from lib_systemd_render.sh. A checkout under a path containing `&`, `\` or `|` rendered a corrupted unit from these two only. Both now source the helper and use it, and a test checks that every placeholder substitution in scripts/install uses an escaped value. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(install): stop the installer scripts reporting things that are not true - first_time_install.sh printed "Password: ledmatrix123" for the setup access point. wifi_manager creates it as an open network ("No password" on the panel), so it now says so. - Step 10.1 printed "✓ WiFi management permissions configured" straight after its own failure message; install_wifi_monitor.sh printed "✓ Package installation completed" after a failed apt install. The tick now only follows success. - Step 7 printed "Web dependencies already installed ... in Step 5" in the one branch that runs because Step 5 did not install them, then created .web_deps_installed on that basis. It now warns and leaves the marker off so the next run retries, as the comment below it intends. - check_system_compatibility.sh called Debian 12 Bookworm "full compatibility confirmed" while first_time_install.sh refuses anything but Debian 13. Bookworm, older Debian and non-Debian systems are now errors. Its counters used ((X++)), which under `set -e` exits the script at the first warning or error (the expression is 0), so the check never reached its summary on any system with one. - configure_web_sudo.sh and configure_wifi_permissions.sh finished by testing `sudo -n test -f ...` and `sudo -n nmcli device status`, neither of which is granted, so they always reported a failure. They now ask `sudo -n -l` about commands the new rules do grant, which checks the rule without running anything. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(install): print the completion summary before rebooting With -y -- and so for every one-shot `curl | bash` install, which always passes -y -- first_time_install.sh ran `reboot` about 180 lines before its "Installation Complete / Web UI Access" summary. reboot returns at once, so the summary printed while the Pi was going down and the SSH session usually dropped before the web UI address could be read. The reboot block moves, unchanged, to the very end of the script. The interactive prompt now also follows the summary. Because the summary now runs before the -y reboot, its one command that could fail under `set -Eeuo pipefail` (the SSID lookup, when nmcli reports a connected device but no active network line) gets `|| true`; a missing SSID was already handled as "SSID unknown". one-shot-install.sh prints its "Next steps" after the installer returns, by which time the reboot is under way, so it now says so, and README's Quick Install mentions the automatic reboot. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(scripts): correct wrong comments and messages, drop dead code No behaviour change except the output text noted below. - 2775 is setgid, not the sticky bit (first_time_install.sh Step 3.1, fix_plugin_permissions.sh), and root needs no "PWM hardware access" to plugin files. - The 777 comments in first_time_install.sh Step 3's fallback and fix_assets_permissions.sh said root needs it to write. Root ignores mode bits; the comments now say what 777 actually opens. The 777 itself is unchanged. - apt_remove ends in `|| true`, so Step 12's "Some packages could not be removed" branch could never run; it is gone and the helper stays non-fatal. - detect_web_service_user's comment named Step 8 for the web unit (install_service.sh installs it in Step 7.5) and now says which branch actually runs. - Step 5 described an "already installed" check that does not exist; the ACTUAL_USER comment described the re-exec backwards. - on_error printed a literal "\n" before "Common fixes:". - Dead code: one-shot-install.sh's uncalled fix_tmp_permissions, LEDMATRIX_ELEVATED=1 (never read) on the sudo re-exec, and configure_web_sudo.sh's unused PYTHON_PATH, which also made a missing python3 fatal for rules that never mention it. - start_display.sh / stop_display.sh said "for user: <you>"; the service runs as root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(fix_perms): fix_cache_permissions.sh uses setup_cache.sh's model There were two models for /var/cache/ledmatrix. setup_cache.sh (the installer's Step 2) and install_web_service.sh share it through the ledmatrix group: root:ledmatrix, 2775, files 660, which is also what DiskCache relies on to give files the directory's group. fix_cache_permissions.sh instead made it 777 and re-grouped it to the invoking user's group, undoing that. It now runs setup_cache.sh for /var/cache/ledmatrix and keeps its own handling of ~/.ledmatrix_cache. Dropped: /var/cache/ledmatrix/ placeholder_logos (nothing reads it) and the checks against the `daemon` user (no service runs as daemon). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: pin actions/checkout in the Claude workflows, drop template comments claude.yml and claude-code-review.yml used actions/checkout@v4 while test.yml and release-version-check.yml pin the v4.2.2 commit SHA; they now pin the same SHA. The commented-out starter-template settings (prompt, claude_args, paths, author filter) are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(scripts): index every script and list removal candidates New scripts/README.md gives one line per top-level script and scripts directory, marked keep, dev-only or diagnostic, and lists the eight scripts nothing in the repo refers to as candidates for removal (kept for now). The install, utils and dev READMEs now list the files they were missing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: tighten two checks that mutation testing showed were too loose - The wifi sudoers check matched `visudo -c -f "$TEMP_SUDOERS"` in the error report too, so replacing the check with `if false` still passed. It now requires the command as the condition. - The summary test never had the setup access point up, so reinstating the bogus "Password: ledmatrix123" line went unnoticed. A case with hostapd active now checks the AP is described as open. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(permissions): describe the repaired fix_perms scripts and new WiFi grants Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): docs-scripts Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4e61d7248a |
refactor(web): one error-response path for api_v3 (#624)
* refactor(web): answer unhandled api_v3 errors from one blueprint handler
Fifty-three api_v3 routes ended in a copy of the same catch-all: log the
traceback, return {status, "An error occurred; see logs for details",
details: describe_exception(e)} with a 500. They are replaced by one
errorhandler on the api_v3 blueprint that returns exactly that body.
It lives on the blueprint rather than falling through to app.py's global
handler because the two answers differ: the global one adds
error_code: UNKNOWN_ERROR, and api_client.js sends a body with an
error_code to the error modal and one without to a plain toast. A
blueprint handler also gives tests that mount api_v3 on a bare Flask app
the same answer the real app gives.
Only handlers that were byte-for-byte that shape were removed (matched on
the AST, and each rewritten function re-parsed and compared). Handlers
with their own message, extra keys, operation-history records or cleanup
stay, as does execute_plugin_action's step-1 handler, which sits inside
an `except subprocess.TimeoutExpired` arm that would otherwise turn a
plugin's timeout into a 408.
HTTPExceptions raised inside a route go back as themselves in the global
handler's 4xx shape. Where a removed catch-all used to swallow one (only
delete_plugin_asset's non-silent get_json() is reachable), a malformed
request now gets its 415/400 instead of a 500.
Most of the diff is re-indentation from unwrapping the try blocks;
`git diff -w` shows the real change.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin action errors name the real failure, not UnboundLocalError
execute_plugin_action bound a local `logger` in its JSON-parsing arm,
which made `logger` local to the whole function. Every other
`logger.error` in it then raised UnboundLocalError, so a failing OAuth
step-1 script was reported as "UnboundLocalError: cannot access local
variable 'logger'" -- from the step-1 handler, and before the previous
commit from the route's outer catch-all too. Use the module logger.
Found by comparing every api_v3 route's forced-failure response before
and after the catch-all consolidation.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop the error category and exception-name code guessing
WebInterfaceError derived an ErrorCategory from every error code and put
it in each structured error body as `error_category`. Nothing reads it:
not the web UI (static/ and templates/), not the tests beyond the ones
pinning the mapping itself, and not any plugin in ledmatrix-plugins. The
enum, the inference table and the JSON key go.
from_exception() could also guess an error code from the exception's
class name ("Config" -> CONFIG_LOAD_FAILED, and so on). Every caller
passes a code, so the guess never ran; error_code is now required.
suggested_fixes stays: the error dialog in static/v3/js/utils/
error_handler.js lists them.
The REST reference loses error_category and says what an unanticipated
exception in an /api/v3 route answers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one call for the from_exception error responses
Nine plugin routes built a structured error by hand:
from src.web_interface.errors import WebInterfaceError
error = WebInterfaceError.from_exception(e, ErrorCode.X)
return error_response(error.error_code, error.message,
details=error.details, context=error.context,
status_code=500)
That is now exception_error_response(e, ErrorCode.X) in api_helpers, so
error_response() is the only structured-error entry point the routes
use. The three operation-history routes never passed the context, and
with_context=False keeps their bodies exactly as they were; a test
compares the helper against the hand-written pair for both forms.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): one api_v3 error-response path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
ece416c4e5 |
refactor(plugins): one plugin-directory resolver (#623)
* refactor(plugins): one resolver for plugin id -> directory Five places mapped a plugin id to its directory, each with its own rules and each re-reading manifests per lookup: PluginManager discovery and get_plugin_directory, PluginLoader.find_plugin_directory, PluginStoreManager._find_plugin_path / list_installed_plugins, and state_reconciliation.disk_plugin_ids. They disagreed on backup dirs, on whether the manifest id or the directory name is the id, on duplicate ids and on path safety. src/plugin_system/plugin_dirs.py now holds the rules once: PluginDirectoryIndex scans one directory and reads each manifest once; resolve_plugin_dir() searches directories in order. What legitimately differs per caller is an explicit argument: search dirs (discovery and the loader: configured dir only; the store: configured then sibling plugins/), ledmatrix- prefix (not for the store), case folding (loader only), manifest pass (not for get_plugin_directory, whose discovery map already holds it). Behaviour changes, all for layouts installs do not produce: - a directory whose manifest declares the id beats one merely named for it (discovery already worked this way; the loader and store now agree) - the store searches the configured dir completely before plugins/ - backup and hidden dirs are skipped everywhere (the loader's case and manifest scans and list_installed_plugins used to return them) - duplicate ids resolve deterministically (exact name, then ledmatrix-<id>, then by name) with a one-time warning; discovery no longer lists the id twice - disk_plugin_ids / list_installed_plugins report manifest ids, falling back to the directory name; auto-update looks the directory up - ids that are not one plain path segment resolve to nothing in every caller (the loader used to truncate them, the store to join them) The .standalone-backup- marker is one constant, BACKUP_MARKER, used by store_manager's rename-aside names and every lookup. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): one plugin-directory resolver Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
13bbb537f3 |
refactor(web): one logging setup and one TTL cache for the web process (#621)
* refactor(web): use src.logging_config in the web process; routine requests to DEBUG
The web interface had its own logging setup (web_interface/logging_config.py)
that replaced the root handlers with a plain stdout formatter. The web
service's journal lines therefore never carried a syslog priority, so
`journalctl -p err -u ledmatrix-web` returned nothing while errors were
logged, and the line shape differed from the display's (the log viewer's
prefix stripping only matched the display format). It also ran after the
module-level managers were built, so their INFO lines at import (including
"Re-removed N uninstalled plugin(s)") were dropped.
app.py now calls src.logging_config.setup_logging() first thing, the same as
run.py: journald priorities under systemd, LEDMATRIX_DEBUG honoured,
LEDMATRIX_JSON_LOGGING still selects JSON.
Per-request logging moves to web_interface/request_logging.py. Every request
used to be logged at INFO, so the UI's polling filled the journal
("GET /api/v3/errors/summary - 200" every minute per tab). Now a successful
GET/HEAD/OPTIONS is DEBUG, a successful write is INFO, 4xx WARNING, 5xx
ERROR. Durations use perf_counter and print to 0.1ms.
The duplicate module is deleted; nothing else imported it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one thread-safe TTL cache for the web process
web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.
Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
counted and get_cached defaulted to 60s. An entry now expires after the TTL
it was stored with; a reader's ttl_seconds can only shorten that. Both
current callers pass the same value on both sides (fonts_catalog 300s,
system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
same expired key could raise KeyError (reproduced), which the endpoints
turned into a 500.
app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.
Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): web logging and TTL cache
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): only ask systemctl about known units
Codacy flagged the systemctl argv built from a variable. The unit now has
to be one of two literals, and anything else raises.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): response_time_ms reads the same clock request_logging stamps
request_logging now stamps request.start_time from perf_counter, but
success_response still subtracted it from time.time(), so metadata
reported ~1.8e12 ms. Found testing on ledpi.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
afe9001aed |
refactor(fonts): one BDF loader and one BDF rasterizer (#627)
* refactor(fonts): one BDF loader and one BDF rasterizer BDF faces were loaded three ways (FontManager._load_bdf_font, element_style._load_bdf, DisplayManager._load_fonts) and drawn by two copies of the same per-pixel loop (DisplayManager._draw_bdf_text and the plugin test harness's "replicated" copy), which golden images and check_plugin/dev_server previews rely on matching the panel. src/common/bdf_font.py now owns both: - load_bdf_face(path, size) -> (face, realised_px): native-strike fallback for sizes the file lacks, one bounded LRU cache keyed on path, size and mtime. FontManager, element_style and DisplayManager delegate to it; read_bdf_native_size moves here (the old names delegate). - draw_bdf_text(draw, text, x, y, face, color, clip): builds each glyph as a 1-bit mask and fills it with ImageDraw.bitmap instead of a draw.point per pixel. A blending Draw (RGB image, "RGBA" mode) keeps the point path so translucent colours still blend. Pixel-identical: 220,032 renders (every bundled BDF at native and off-strike sizes, 14 strings, 4 colours, clipped on every edge, through each old loader x rasterizer) match origin/main byte for byte. test/test_bdf_font.py keeps a lightweight version against a frozen copy of the old loop. DisplayManager._draw_bdf_text goes from 1.4-23 ms to about 0.1 ms per string (the old loop re-read FreeType's buffer as a Python list for every pixel). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(testing): harness calendar_font is sized like the panel's VisualTestDisplayManager built its 5x7 calendar_font / bdf_5x7_font as a bare freetype.Face. With no size set its ascender reads 0, so BDF text drawn with it landed 6px above where DisplayManager draws it -- entirely off the canvas at y=0 -- and get_font_height() returned 0. Golden images and check_plugin / dev_server previews showed text the panel does not. Load it through load_bdf_face at the panel's 7px, so it is the very face DisplayManager uses. Across the differential run this changes only the cases drawn with the harness's own calendar_font (968 of 220,032), which now match the panel's output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): one BDF face per thread The shared face cache now hands every loader (FontManager, element_style, DisplayManager, the harness) the same freetype.Face. FreeType does not allow two threads to use one face at once, since load_char rewrites its glyph slot, so key the cache by thread as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
abedc46104 |
refactor(sports): merge the sports_shared/sports_card twins that behave identically (#626)
* refactor(sports): wrap the sports_card twins that behave identically SportsCoreSharedMixin (switch mode, via each scoreboard's sports.py) and sports_card (scroll/Vegas mode, via game_renderer.py) carried the same helpers twice. test/test_sports_twins.py now calls every pair with the same inputs -- the eight scoreboards' harness fixture games in flat, flat+nested and nested-only shapes, plus edge cases (favourites by id and abbreviation, NRL's colliding abbreviations, missing and non-numeric scores, bad zones, out-of-range dates, shared font faces). Identical pairs become thin wrappers over the sports_card function: _card_option, _vs_text, _format_game_time, _coerce_rgb, _crisp_size (with the class's own tables), _unshare_element_fonts (with the class's own element map, via a new optional argument), and the colour/month/weekday/ font-grid tables (dicts copied, not aliased). _format_game_date shares the card's formatting body but keeps its own setting, weekday zone and month table; _schema_font_size shares the parser but keeps its per-class cache, because a reloaded plugin gets new classes and a shared path cache would stop it seeing an edited schema. _resolve_font_size agrees but keeps its body so it still dispatches through the overridable hooks. No behaviour change: old and new mixin/card agree on all 22,994 comparisons over the test corpus, and the pairs that do differ (favourite-result colours on nested payloads and by favourites source, the weekday's timezone, the element-name map, per-mode colours) are left alone and pinned in TestPinnedDivergence for an owner decision. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(sports): pin that an ambiguous NRL abbreviation tints in both modes NRL's resolver passes a shared abbreviation ("NEW") through with an error and its _is_favorite_game matches ids only, but both favourite-colour helpers match on abbreviation as well, so both display modes tint a Knights or Warriors result for a user who typed "NEW". The twins agree; neither consults the _favorite_key seam. Pinned so a fix is deliberate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
1fe7237799 |
refactor(install): generate the web sudoers rules in one place (#622)
* refactor(install): generate the web sudoers rules in one place /etc/sudoers.d/ledmatrix_web was written by two copies of the same allow-list: a heredoc in first_time_install.sh Step 10 and a block of echo lines in scripts/install/configure_web_sudo.sh. They drifted before (safe_pip_install.sh was granted by one only), and a test existed just to catch that. Both now call web_sudoers_rules() from the new scripts/install/lib_sudoers.sh and keep their own validate (visudo -c), install and confirm flows. - first_time_install.sh output is byte-for-byte unchanged, so a device re-running the installer gets "already up to date". If the library is missing, Step 10 keeps the installed file and carries on, the same way it handles rules that fail visudo (an empty file would pass visudo). - configure_web_sudo.sh now writes the installer's layout: same 18 rules, different comments and order. It still leaves out reboot, poweroff and journalctl when they are missing; the library does that for both. The drift test now pins the generator's grants, checks that neither installer writes rules of its own, and runs each installer's call line to check the argument order. Tests that read the rule text now read the library. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(install): detect the web service user in one function first_time_install.sh pasted the same WEB_SERVICE_USER detection block three times (Step 3.1's fallback, the plugin-repos setup and Step 11). The copies were identical apart from comments; they now call detect_web_service_user(), whose body is that block unchanged. Behaviour is the same: the function sets the same global and always returns 0, as the inline if-chain did. Checked on Linux against all three original copies across 13 layouts (installed unit with and without User=, the repo as shipped, each grep branch, template placeholders). The comment notes that the install_web_service.sh / install_service.sh greps no longer match anything, so until Step 8 installs the unit the result is "root". That behaviour is left as it was. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
ddf5f085a5 |
perf(cache): tell a stale record from its header instead of parsing it (#633)
The sports plugins cache whole season schedules: 53MB for MLB, 18MB for NHL, 17MB for NCAA baseball. On a Pi 4, orjson.loads of the MLB file takes ~1.8s with the GIL held, and every thread in the display service waits -- the stall watchdog caught the render thread frozen 0.5-1.3s with the interpreter itself blocked, right on these reads. When a season record expired, DiskCache.get paid that whole parse only to find the timestamp too old and throw the result away. CacheManager.set now writes timestamp and ttl ahead of the data, and DiskCache.get reads them from the first 256 bytes of the file, applying the same rule as before (a per-entry ttl wins over max_age; no limit means never stale). A record that is stale is refused without being parsed. Files in the old layout, and records from other writers, don't match the header and are parsed in full as before. Also: ESPN responses in the background data service and espn_dates are parsed with orjson when it is installed (src/common/json_body.py). The stdlib parser behind response.json() takes 3.1s on the MLB season against orjson's 1.8s, both with the GIL held. espn_dates imports it with a fallback, since plugins bundle copies of that module for older cores. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
82f3a3a3e4 |
fix(redaction): make credential redaction linear, not quadratic (#631)
* fix(redaction): make URL-userinfo redaction linear, not quadratic _REDACT_URL_USERINFO could start a match at every letter of a run of scheme characters, and each attempt read to the end of the run looking for `://`. On a long unbroken run of letters or digits (a hex digest, an ID, part of a response body) that is quadratic: 1.6s for 20k characters. The display service redacts every message, stack trace and context value it publishes in the error snapshot, holding the aggregator lock, and re.sub holds the GIL for the whole call, so one such exception stalled every thread, render loop included (~0.5s measured for 20k chars of hex). It also made test_snapshot_stays_small the slowest test in the suite by far: 142s of a 383s run, 139s of it in this one regex. A match may now only start where a run of scheme characters starts (negative lookbehind). Leading digits and `+.-` are captured in group 1 so the substitution restores them, and the scheme still has to start with a letter, so what gets redacted is unchanged: old and new output were identical on 300k fuzzed inputs. 20k chars now take ~0.5ms, 200k ~6ms, and test_snapshot_stays_small takes 0.8s. test/test_redaction.py pins the exact output for schemes that begin after digits or `+.-`, and bounds 50k-character runs at 1s; against the old pattern those timing tests fail at 3-11s each. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK * fix(redaction): make Authorization-header redaction linear too _REDACT_AUTH_HEADER matched the value's opening as `\s*["\']?\s*`: two `\s*` separated only by an optional quote. With no quote, a whitespace run could be split between them in every possible way, and when no credential followed (end of text, or `,` `"` `<` ...) the engine tried them all before giving up: quadratic, 8s for `authorization:` and 20k spaces, 17s with `Proxy-Authorization:` (tried again at the inner `authorization`). Same stall as the URL pattern: re.sub holds the GIL, and the display service redacts everything it publishes. The quote and the whitespace after it are now one optional unit, `\s*(?:["\']\s*)?`, which matches the same strings with only one way to split them. Output is identical to the old pattern on 300k fuzzed inputs; 20k spaces now take ~1.6ms. A scan of all three redaction patterns over prefix/run/suffix shapes finds none left that scales superlinearly. test/test_redaction.py pins exact output for quoted, tabbed, multi-line and credential-less headers, and bounds header + 20k whitespace at 1s; against the previous pattern those fail at 8-17s each. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
c1ce0b7b04 |
fix(web): two api_v3 paths called names that no longer exist (#625)
The Pixlet editor stop route restarts the display after a SIGKILL with _run_systemctl_command, which starlark.py never imported (since #554). The Starlark device-location resolver fell back to _ensure_cache_manager, which #609 deleted; the resolver already accepts no cache manager. Both raised NameError on the rare path that reaches them. pyflakes finds no other undefined names in src/ or web_interface/. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f3894916a9 |
feat(web): show which plugins use each font; warn before deleting one (#619)
* feat(web): show which plugins use each font, warn before deleting one
The Fonts tab lists font files from the web process's own scan, and the
plugins that register fonts run in the display process, so the tab had no
way to say whether a font was in use before deleting it.
The display service now publishes {catalog key: [plugin ids]} to the
shared cache (font_usage_snapshot, src/font_usage.py), built from the
loaded plugins' FontManager.register_manager_font() registrations. A
daemon thread checks every 10 s and writes only when the usage changed
(plus a daily refresh so cache cleanup cannot expire it); it never raises.
Families, aliases (press_start, four_by_six, ...) and paths are resolved
through FontManager's catalog to the file stem the Fonts tab keys rows by;
fonts outside assets/fonts are left out. Unloading a plugin drops its
registrations (new FontManager.forget_manager_fonts).
GET /api/v3/fonts/catalog merges used_by into each row per request (the
5-minute scan cache is copied, never edited): a list of plugin ids, or
null when the display service has not reported. The tab shows a Used by
column ("unknown" / "-" / ids, rendered as text) and deleting an in-use
font names the plugins in the confirmation, from a fresh read. The server
still refuses only system fonts. Catalog fetches bypass the browser's
5-second API cache, which otherwise served the pre-delete list.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix: call forget_manager_fonts through a hasattr check pylint can follow
getattr(..., None) then callable() is fine at runtime, but pylint's E1102
("not callable") can't see through it, and Codacy fails the check on it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
9a1f94f793 |
fix(starlark): blank app locations use the device location, not San Francisco (#617)
* fix(starlark): blank app locations use the device location, not San Francisco A Starlark (Tidbyt) app whose Location field is blank rendered at its author's hard-coded DEFAULT_LOCATION -- usually San Francisco -- even with the device city set under General settings. A user in Charlotte, NC got San Francisco weather and radar with nothing in config.json to explain it. src/device_location.py fills unset location fields at render time (display plugin and the web standalone render): the device city is geocoded once via Open-Meteo, preferring a match in the configured state/country, and cached permanently. A saved location always wins; if the lookup fails the field is dropped so the app uses its own default, and the failure is not retried for 30 minutes. Also fixes the config form: clearing a location omitted the key, and the save merges, so the old value could never be removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(starlark): say what happens when the device location can't be used A blank app Location only renders at the device's city when one is set and the Open-Meteo lookup finds it. With no city, no match, or the geocoder unreachable (retried after 30 minutes), the app gets no location and keeps its author's default. The guide, the config page hint, CONFIG_REFERENCE and the CHANGELOG entry now say so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
61e462c635 |
refactor: remove the skin system and the unused src/base_classes package (#615)
* refactor: remove the skin system Skins never rendered with the current scoreboard plugins: the only hook was SportsCore._render_game in src/base_classes, which no plugin builds on, so the UI and store already treated them as unsupported. The owner decided on 2026-09-23 to remove them outright. Removed src/skin_system/ (runtime, base class, fixtures), skins/, scripts/validate_skin.py and their tests; the store's "type": "skin" installer, uninstaller and hide/refuse filters (the official registry lists no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins. Stored skin/skin_options config values are handled in the next commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(config): drop retired skin/skin_options keys instead of validating them A config.json written while the skin system existed can carry skin and skin_options in any plugin section, and most plugin schemas set additionalProperties: false. They are no longer core plugin properties; RETIRED_PLUGIN_KEYS in schema_manager lists them and drop_retired_plugin_keys removes them (unless the plugin's own schema declares the name) in prepare_plugin_config, which loading, hot reload, GET /plugins/config and both web saves already share, and in validate_config_against_schema for callers that validate a raw section. POST /plugins/config and /config/main also drop them from the stored section they merge into, so they leave config.json on the next save. Tests cover the load path (real PluginManager.load_plugin: no schema warning, not degraded), raw and prepared validation, validate_all_plugin_configs, and the JSON, form and /config/main saves. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: remove the unused src/base_classes package No scoreboard plugin builds on src.base_classes: the nine monorepo scoreboards ship their own sports.py and share code through src/common (docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins imports it. The one import anywhere, baseball-scoreboard's rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing instantiates. Removed the package and the eight test files that only tested it (test_api_extractors, test_data_sources, test_sports_base_characterization, test_sports_capabilities, test_sports_core_promotions, test_sports_logo_cache_bounded, test_sports_modes_promotions, test_sports_odds_fanout). test_common_is_hardware_free no longer lists src.base_classes as a forbidden import, and comments in sports_helpers.py and base_odds_manager.py stop pointing at it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: drop the skin system and src/base_classes from the docs Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was removed and shared code lives in src/common, in the Layering section and the view-model-contract rule. Other docs stop pointing at the removed package. CHANGELOG records both removals under Unreleased. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(store): hide and refuse registry entries that aren't plugins The skin filters went with the skin system, but a custom registry can still list "type": "skin" entries, and installing one as a plugin would unpack it into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing type means plugin) now hides non-plugin entries from the store and custom-registry listings, and install refuses them, in the route with a clear 400 and in _install_plugin_impl for any other caller. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4fe3cdd906 |
fix(starlark): stop the root display service locking the web UI out of starlark-apps (#604)
* fix(starlark): stop the root display service locking the web UI out
Reported after a fresh install: installing an app from the Starlark tab
failed with "install failed: Failed to install from repository", and so did
uploading a .star file and installing from a GitHub directory. The reporter
found the cause only by reading service logs, and fixed it with
sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/starlark-apps
starlark-apps is gitignored, so it is never checked out -- it is created
lazily by whichever process reaches it first. Those processes run as
different users. systemd/ledmatrix.service is User=root and constructs this
plugin at startup, which is where _get_apps_directory() is called from;
systemd/ledmatrix-web.service runs as the login user and is what actually
installs apps.
The documented first step is to install pixlet and reboot, so on a fresh
machine the display service usually wins that race and mkdir() leaves the
directory root-owned. The web process then fails in _install_star_file() on
app_dir.mkdir(), which catches nothing, so PermissionError reaches the
route's outer `except Exception` and becomes the generic message the user
saw. All three install paths write to the same directory, which is why all
three failed.
The web user cannot repair this -- chown needs root. So root does it, on
every startup, which also heals machines already broken by this without the
owner having to find the chown themselves. It is a no-op when not root, when
the platform has no POSIX ownership, and when the checkout genuinely belongs
to root; a chown that fails warns rather than killing startup.
Also made the failure legible if the handover is ever prevented: a
PermissionError now names the directory, the automatic repair, and the
manual chown, instead of a message that names neither path nor cause.
Verified by mutation: dropping the handover call, chowning a genuinely
root-owned checkout, and letting a non-root process chown each fail their
own test. 121 starlark tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(starlark): address the review on the ownership repair
Findings from the automated review of #604.
Symlinks (CWE-59, the serious one). A root chown that follows links is a
privilege-escalation primitive: anyone able to write in starlark-apps could
point a link at a root-owned file and have the repair hand it over. Entries
are now read with os.lstat, symlinks are skipped outright, and the chown
passes follow_symlinks=False. Descendants are processed before the directory
itself, so the container does not change hands while its contents are still
being walked.
install_app() caught PermissionError in its broad handler and returned
False, which both routes report as a generic install failure -- the exact
shape of the bug this PR exists to fix, since the caller could not tell
"this app is broken" from "this process cannot write here". PermissionError
is now re-raised; every other failure still returns False.
The test fixtures skipped on bare Exception, which would have turned a
syntax error or NameError in the plugin into a green run. They now skip only
for a named absent dependency and re-raise anything else.
Also fixed the _Stat stub that failed in CI but passed locally: it carried
only st_uid/st_gid, and pathlib reads st_mode while walking. It now wraps
the real stat result and overrides ownership alone.
NOT taken: the CodeQL "information exposure through an exception" finding on
the hint response. Dropping `details` would contradict this package's
documented rule -- "if it returns 5xx, it says why" -- which
test_no_api_v3_handler_discards_its_exception enforces with an allowance
that may shrink and never grow. The Starlark routes are the ones that policy
was written for: they answered 500 with no detail for three releases.
describe_exception already redacts credentials and truncates. Keeping the
detail is the deliberate trade-off, so the finding is declined rather than
silently worked around.
Verified on hdpi with the updated code: a symlink to /etc/shadow planted in
starlark-apps was skipped while the directory was handed back, and
/etc/shadow stayed root:shadow.
Mutation-checked all three behaviours. The symlink test was vacuous on the
first attempt -- the link already had the target owner, so it was skipped
for the wrong reason and the mutation passed. It now forces the link to look
like it needs handing over, and fails when the check is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4a1fd7464a |
fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)
* fix(errors): serve /api/v3/errors/* from the display service's aggregator The error aggregator is a per-process singleton and only the display service runs plugins, so only its aggregator records anything. The routes read the web process's own, empty one and always reported no errors. The display service now publishes a bounded snapshot of its aggregator to the shared cache (plugin_error_snapshot) from a daemon thread: at most once every 10 s and only when something changed, never raising into the caller. The routes read it and keep their response shapes, adding snapshot_available, generated_at and clear_pending; exception text has credentials redacted. POST /errors/clear writes a clear request (plugin_error_clear_request) that the display applies on its next 5 s tick via the new clear_before(), which keeps errors recorded after the cutoff and rebuilds the counts. Until the snapshot acknowledges the request, reads hide everything before the cutoff, so a snapshot written just before the click cannot bring errors back. Adds "all": true; cleared_count is null when only the display can know it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): show plugin errors in the Logs tab A compact panel under the log viewer: per-plugin error counts, repeating errors (type, count, affected plugins, a sample message, last seen) and a Clear button, with empty states for "no errors" and "display service hasn't reported yet". Polls every 15 s while the tab is active; all text goes through escapeHtml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: describe where plugin error reports come from and how clear works Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(errors): redact the published snapshot before clipping it Keeping only a traceback's tail (or clipping a message) could cut an `api_key=` marker off while keeping the secret after it, and the web side's redaction would then have nothing to match. The display now redacts every free-text field of the snapshot first. The patterns move to a Flask-free src/redaction.py so the display service can use them; redact_text in the web error handler uses the same function, unchanged in behaviour. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
cd5a4e2251 |
fix(display): apply on-demand, brightness and schedule changes mid-screen (#618)
* fix(display): apply on-demand, brightness and schedule changes mid-screen The main loop read the on-demand mailbox, the on/off schedule and the brightness target once per pass -- once per screen. A dwell can be a minute and a Vegas iteration runs for max_cycle_duration (240s), so on a Pi an on-demand request posted at 10:54:27 was activated at 10:57:24, and two brightness saves 12s apart inside one 30s screen never reached the panel. During Vegas nothing read the mailbox at all: _check_vegas_interrupt only checked on_demand_active, which only the main-loop read sets. _service_pending_changes does the main loop's on-demand poll, expiry, schedule and brightness steps, throttled to PENDING_CHANGES_INTERVAL (the existing 0.25s mailbox floor), on the display thread. It runs from the Vegas interrupt checker, the high-FPS and once-a-second render loops (replacing their direct on-demand poll) and _sleep_with_plugin_updates; between passes it costs one monotonic compare. A brightness change re-pushes the current frame, since the panel only shows it from the next push. Callers act on what it leaves behind: Vegas yields on an on-demand start or the display being scheduled off (and the main loop then blanks instead of rendering a screen), the render loops break on a schedule-off as they already did on a mode change, and the dwell sleep returns early on an on-demand start/stop or a schedule flip -- so the 60s scheduled-off sleep now wakes for an on-demand request. The main loop no longer rotates after a dwell that ended that way, which advanced a new on-demand session past the mode that was asked for. A brightness set_brightness() refuses is not retried until the target changes, so the 4Hz pass doesn't log the same failure (fallback mode) four times a second. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(display): a screen scheduled off midway stops rendering Covers the schedule-off break added to the high-FPS and once-a-second render loops: with the display scheduled off halfway through a 120s screen, neither loop renders for more than one redraw plus one service interval past the boundary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3f8edf5113 |
fix(logos): one hardened logo download path; shared, real HTTP headers (#612)
* fix(logos): harden the plugin logo download and share core HTTP headers download_missing_logo / LogoDownloader.download_logo, the path the scoreboard plugins use, read response.content with no size cap and wrote straight to the final path, so a failed or corrupt download could be left in place and cached as the logo. It now goes through fetch_logo: streamed with a 10 MB cap, image/* only, decoded by Pillow, converted to RGBA once, and moved into place atomically. A failure leaves no partial or temp file and keeps any logo already on disk. LogoHelper._download_logo delegates to the same code. Public signatures and return values are unchanged; saved files are pixel-identical to before (RGBA, palette+tRNS, L+tRNS, LA, JPEG). download_missing_logo reuses one downloader per thread instead of a new Session per logo. Per thread rather than behind a lock: Session is not documented thread-safe, and a lock would serialise every plugin's downloads behind the slowest one. Placeholders are written atomically, without the test_write.tmp probe. The logo downloader and background data service now send the real ChuckBuilds User-Agent from src.common.api_helper (USER_AGENT, DEFAULT_HTTP_HEADERS) instead of a yourusername/contact@example.com placeholder, and no longer hand-set Accept-Encoding: br (brotli is not installed). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(http): drop APIHelper's hand-set brotli encoding; LogoHelper sends the real UA APIHelper advertised `br` though brotli isn't installed, so a server that honoured it would send a body requests can't decode. LogoHelper sent a bare `LEDMatrix-Common/1.0`, the kind of User-Agent ESPN has been rejecting. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f813ea2117 |
fix(web): define project_root when plugins_directory is absolute (#616)
project_root was only assigned in the relative-path branch, so an absolute plugin_system.plugins_directory made web_interface/app.py raise NameError at import (first use: the SchemaManager construction). Define it before the if/else; plugins_dir resolution is unchanged. Adds a regression test that imports the real module in a fresh interpreter with an absolute and a relative plugins_directory. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9d024f24ef |
refactor(cache): remove the cache layer's duplicate cleanup and dead lookups (#613)
* refactor(cache): collapse CacheStrategy's all-60 defaults table and twin soccer branch get_sport_live_interval() without a config manager looked the sport up in a table where every value was 60, with 60 as the fallback; it now returns 60. get_data_type_from_key() had an `if 'soccer'` branch returning the same 'sports_live' as its else. test_cache_strategy_intervals pins the returned strategy for every data type x sport key x config-manager shape; it passes unchanged on the old code. A 2,544-entry dump of every CacheStrategy method over a wider grid is identical before and after. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(cache): drop CacheStrategy's `<sport>_scoreboard` config lookup get_sport_live_interval() and get_cache_strategy() read live/recent/ upcoming intervals from config[f"{sport}_scoreboard"]. Those sections belonged to the built-in scoreboards the plugin system replaced; plugin config is keyed by plugin id ("football-scoreboard"), so on a current config the lookup always fell through to the defaults (60 live, 1800 recent, 10800 upcoming), which are now returned directly. The one input where this differs: a config.json upgraded from the pre-plugin era that still carries e.g. an "nfl_scoreboard" section (no code removes them), queried with an explicit sport key. No caller in core or the plugin monorepo passes a sport key here -- get_with_auto_strategy only derives one for keys classed sports_live/live_scores, and its callers (odds managers, odds-ticker) use odds keys -- so the stale section was unreachable in practice. A dump of every CacheStrategy method over 2,544 inputs differs from the previous commit only in those 45 legacy-config entries; the test grid now includes that shape. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * perf(cache): list cache files without holding the memory-tier lock CacheManager.list_cache_files() held the in-memory cache's lock while it listed and stat'd the whole cache directory -- 8,864 files on a real rig -- so every get()/set() from the display loop and plugins waited out the scan. The lock never protected the disk: DiskCache writes and deletes under their own lock, and a file vanishing between listdir and stat was already handled (logged and skipped). The body is unchanged apart from the dedent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(cache): delegate memory-tier cleanup and stats to MemoryCache CacheManager._cleanup_memory_cache() was a line-for-line copy of MemoryCache.cleanup(), and get_memory_cache_stats() a copy of MemoryCache.get_stats(), both reaching into the component's private _cache/_timestamps/_lock through "backward compatibility" aliases bound in __init__. So the component's own cleanup and stats only ever ran in tests, and the aliases went stale whenever the component was swapped (test_cache_ttl_honoured does). Both now delegate, and the aliases are gone: nothing in core, the tests, or the ledmatrix-plugins monorepo reads them. Behaviour is the same. Compared line by line, the two cleanups differ only in the sort key's fallback (0 vs 0.0, which orders identically), range+bounds check vs slice for the eviction, and the logger name on the DEBUG summary line (src.cache_manager -> src.cache.memory_cache). A differential run over 20,000 random memory states (str/None/garbage/ future timestamps, orphan keys, sizes 0-12, forced and throttled runs) gives identical removed counts, resulting dicts and last-cleanup times; the same harness catches each of three seeded mutations of MemoryCache.cleanup. The throttle clock also moves with it: CacheManager kept its own copy of last-cleanup, the component's is used now, and they started equal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(background): inline the sport cache key and drop the unused request queue get_sport_cache_key() constructed a whole CacheManager -- ConfigManager, config parse, cache-dir probing with test-file writes -- to return f"{sport}_{date}". It now builds the key itself in the same format as CacheManager.generate_sport_cache_key() (UTC date, %Y%m%d); tests check the two agree for explicit dates and, with a frozen clock at 03:30 UTC, for the default date. Median per call on Windows: ~0.6 ms -> ~2 us (alternating runs); on a Pi the old path also wrote a probe file per call. request_queue was a PriorityQueue nothing ever put into: requests go straight to the executor, so `priority` never did anything. The queue is gone; the `priority` parameter and FetchRequest field stay (every monorepo scoreboard passes priority=) and are documented as ignored, and get_statistics() keeps reporting queue_size, now a literal 0 as it always was in practice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
269385c97c |
fix(config): write config.json through one durable atomic writer (#611)
save_config() opened config.json with 'w' and streamed json.dump into it, so a power cut or an unencodable value left the file truncated. save_config_atomic() renamed a temp file into place but never fsynced it, rewrote the unchanged secrets file on every save, and re-parsed every backup to rotate them. save_raw_file_content() had its own third copy. All of them, plus rollback and config creation from the template, now go through atomic_write_text(): temp file in the same directory, fsync, final mode set before the rename, rename (retried on Windows while a reader holds the file), directory fsync. A root save copies the previous owner onto the new file so a rename by the display service no longer hands config.json to root; the shared-group fix-up is unchanged. The mode is chosen from the file name, so a "secrets" directory in the install path no longer makes config.json 0640. The secrets file is rewritten only when its content changes, and backup rotation works from filenames alone. Backups keep their names (config/backups/config.json.backup.<version>, paired secrets backup) and the five newest are kept, as before. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
604f58ff07 |
feat: deprecate unused plugin-facing methods for removal in 3.7.0 (#610)
35 methods on CacheManager, DisplayManager, FontManager and PluginManager have no caller in core, the ledmatrix-plugins monorepo or the registry's third-party plugins, but plugins live elsewhere, so they stay for one release. src.deprecation.deprecated logs a warning (and emits a DeprecationWarning) the first time each is called in a process, naming the release that removes it. The list and replacements are in CHANGELOG and PLUGIN_API_REFERENCE's new Deprecated APIs section; a test pins the set. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a8b3e86775 |
refactor(web): delete dead routes, JS files and duplicate definitions (#609)
* refactor(web): drop validators nothing calls escape_html, validate_image_url, validate_font_awesome_class, validate_mime_type, validate_numeric_range, validate_string_length and sanitize_plugin_config had no callers outside their own tests. Only validate_file_upload (fonts upload) is imported by the web interface. dedup_unique_arrays is kept: its one caller in save_plugin_config was removed by the unrelated sync PR (#330), which looks accidental. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(api): remove the music-auth and of-the-day JSON routes POST /plugins/authenticate/spotify and /plugins/authenticate/ytm had no caller but their tests: the music plugin authenticates through its web_ui_actions (authenticate_spotify.py / authenticate_ytm.py) via /plugins/action. POST /plugins/of-the-day/json/upload and /json/delete looked the plugin up by the id ledmatrix-of-the-day (its manifest id is of-the-day), were reachable only from a file_type "json" upload field that no schema declares, and put the plugin directory on sys.path per request to import scripts.update_config. of-the-day manages its files through plugin-file-manager and its own web_ui_actions. The of-the-day branch of GET /plugins/config stays: it matches the real manifest id and still merges the on-disk category files into the form. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(api): read managers only from the blueprints api_v3/__init__.py and pages_v3.py declared module globals (plugin_store_manager, saved_repositories_manager, schema_manager, operation_queue, plugin_state_manager, operation_history, sync_manager, config_manager, plugin_manager) that nothing assigns: app.py sets the managers as attributes on the Blueprint objects, and every route reads them there. The one reader, backup restore's fallback to the module plugin_store_manager, could only ever fall back to None. _ensure_cache_manager() built a second CacheManager in the web process instead of using the one app.py puts on api_v3. The display routes now read api_v3.cache_manager, creating it on the blueprint only when nothing set it (the same None handling as the /cache routes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(web): drop run.sh and the unused log_config_change web_interface/run.sh was referenced only by web_interface/README.md; the service starts the UI through scripts/utils/start_web_conditionally.py and the README already documents `python3 web_interface/start.py`. log_config_change() in web_interface/logging_config.py was never called. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): delete unreferenced store_manager.js, diff_viewer.js, htmx-sse.js - js/plugins/store_manager.js (window.PluginStoreManager) and js/config/diff_viewer.js (window.ConfigDiffViewer) were loaded on every page but nothing reads either global. - htmx-sse.js (plus its CDN fallback) was loaded after HTMX, but no template or plugin page uses sse-connect / hx-ext="sse": the live streams run through LEDStreams in app-shell.js. js/plugins/state_manager.js stays: install_manager.js's updateAll() reads and refreshes window.PluginStateManager. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove app.js helpers nothing calls - hexToRgb, rgbToHex, validateForm, uploadFont and switchTab (whose 'switch-tab' event had no listener) have no caller in the templates, static JS or the plugin monorepo. - installPlugin: plugins_manager.js (loaded last) assigns window.installPlugin, and its own store cards are the only callers. - The showNotification fallback could never install: app-shell.js is deferred ahead of app.js and defines the same fallback at top level. - performanceMonitor only logged with ?debug=perf and read an unset this.measures; the marks it took on every load had no reader. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): drop app-shell.js refreshPlugin A top-level function in app-shell.js, so a window global, but nothing calls it (no inline handler, no window lookup, no string-built name). The other plugin actions in that block stay. updatePlugin is the live window.updatePlugin: plugins_manager.js only installs its own copy when none exists. uninstallPlugin/pollUninstallOperation, updateAllPlugins, executePluginAction and toggleNestedSection are replaced by later deferred scripts, but a click that lands while those scripts are still downloading reaches the app-shell copies, so removing them is not a pure no-op. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove definitions plugins_manager.js always overrides All of these are replaced before anything can call them, checked against the load order in base.html and the live window.* values: - openOnDemandModal/requestOnDemandStop stubs: the IIFE later in the same script assigns the real functions synchronously. - updatePlugin and uninstallPlugin stubs (`window.X || stub`): app-shell.js already defined both, so the fallback never installed. Same for the later updatePlugin override, gated on the live function containing '[UPDATE]', which app-shell.js's never does. - The first addArrayObjectItem/removeArrayObjectItem: reassigned by the top-level copies after the IIFE. - The first `function formatDate` in the IIFE: a later declaration of the same name in the same scope wins. - deleteUploadedImage, getCurrentImages, showUploadProgress, formatFileSize and getScheduleSummary: character-for-character copies of js/widgets/file-upload.js, which stays the owner. - `typeof X === 'undefined'` fallbacks and `typeof X !== 'undefined'` re-exports after the IIFE: always false, or a self-assignment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): render the shell directly and delete index.html index.html extended base.html with {% block content %}, but base.html defines no blocks, so none of index.html ever rendered: rendering both with jinja2 gives byte-identical output. index() still loaded the config, read config.json and config_secrets.json raw and json.dumps'd them on every page load for variables base.html never reads, and flashed errors that base.html never shows. It now renders base.html with no context. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): stop htmx-config.js replacing console.error and console.warn It swapped both globals for filters that dropped any error mentioning insertBefore / "Cannot read properties of null" when "htmx" appeared in the message or stack, and a list of Permissions-Policy warnings. That hid real errors from every script on the page, and made every logged error and warning report htmx-config.js as its source. The beforeSwap target validation above it, which prevents the insertBefore errors in the first place, stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(web): quiet the widget load announcements and debug logs About 30 lines hit the console on every page load: one "... widget registered" per widget file, one "[WidgetRegistry] Registered widget: X" per registration, plus the registry, base widget and plugin loader announcing themselves. The load-time announcements are removed; the per-call ones (registry register, plugin widget loads, "Render called") now go through the page's debugLog switch (localStorage.pluginDebug), guarded because the widgets also load in node tests without it. fonts.html and wifi.html debug logging goes through debugLog as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(api): drop the removed music-auth and of-the-day JSON routes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
84afa9d64f |
refactor: delete dead Python code in the core (and stop storing Wi-Fi passwords) (#608)
* refactor(plugins): remove the no-op PluginHealthMonitor Its monitor loop did nothing (`if callbacks: pass`), register_health_check had no callers and api_v3.health_monitor was never read by any route. The live health data comes from PluginHealthTracker, which is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(store): drop the never-set uninstall tombstones Nothing in production called mark_recently_uninstalled, so the reconciler's was_recently_uninstalled check was always False. The persistent uninstall registry is what actually stops resurrection; the reconciler test now exercises that gate instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(common): delete unused config/display/game helpers, utils and error_handler Nothing in core, the web UI, scripts or the plugin monorepo imports config_helper, display_helper, game_helper, utils or error_handler; only their own tests did. The error_handler re-exports leave src.common's __all__; APIHelper, TextHelper, ScrollHelper, LogoHelper and the adaptive layout exports are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(config): drop ConfigService's unused versioning and save API ConfigVersion, get_version/get_version_history/get_version_config, rollback, save_config, reload, get_plugin_config and the backward-compat load_config/get_config_path/get_secrets_path had no callers. The display controller only uses get_config, subscribe, unsubscribe and shutdown, plus the file watcher. Change detection now compares against the current checksum instead of the last history entry. The subscriber tests asserted `callback.called or True`; they now reload the way the watcher does and assert the notification. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): drop unread plugin state history and callbacks plugin_state.PluginStateManager kept a bounded per-plugin transition history that only get_state_history (tests only) read; get_state_info reports a separate lifetime count, which stays. set_error_info and record_display had no callers, and set_state_with_error's `error` argument only fed the history. The web-side state_manager.PluginStateManager loses subscribe_to_state_changes, _notify_callbacks, set_plugin_error and get_state_version, none of which had callers; with no subscribers the old-state copy in update_plugin_state went with them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): remove unused PluginManager methods and attribute guards update_all_plugins was only called by a test (the display loop uses run_scheduled_updates); get_plugin_health_metrics, get_plugin_resource_metrics and get_plugin_state had no callers; and plugin_modules was written but never read. plugin_directories is now initialised in __init__, so the hasattr() guards around it go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): remove unused executor, loader, store and package helpers - PluginExecutor.execute_safe: no callers. - PluginLoader._parse_semver: only its own tests; compatibility.parse_semver is the live copy and test_compatibility.py already covers it. - PluginStoreManager.get_installed_plugin_info: no callers. - PluginResourceMonitor._local: never read. - src.plugin_system.get_store_manager and __api_version__: no importers in core, scripts or the plugin monorepo. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): stop storing Wi-Fi passwords in wifi_config.json WiFiManager appended every joined network's SSID and password, in plaintext, to saved_networks in config/wifi_config.json, and nothing (web UI, backup restore, scripts) ever read them back: NetworkManager keeps its own credentials. The writes are gone, and loading the config now drops any saved_networks key and rewrites the file, so passwords already on disk are scrubbed. Also removes _check_dnsmasq_conflict (never called) and _detect_trixie, whose result only reached one log line, along with the NM_CONNECTIONS_PATHS constant only it used. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): remove unreachable and unused DisplayController code - _follower_rebuild_scroll_image: never called. - mode_duration (never read) and last_mode_change (write-only). - The `chosen_cap <= 0` branch: chosen_cap is either the minimum of caps already filtered to > 0 or DEFAULT_DYNAMIC_DURATION_CAP (180). - The `max_duration < min_duration` branch directly after `max_duration = max(min_duration, max_duration)`. - The circuit-breaker branch's `display_result = False` and `manager_to_display = None`: the first is overwritten a few lines later, the second is already None there. - The bool-to-bool conversion of execute_display's result, which is always a bool. - The `loaded_plugins` lookup in _update_modules: PluginManager has no such attribute, so it always fell through to `plugins`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(vegas): remove unused config update, boundary finder and refresh VegasModeConfig.update had no callers outside its own tests (the coordinator rebuilds the config with from_config on a change); geometry.find_item_boundary and StreamManager._refresh_plugin_content had no callers at all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(run): drop the debug block that pretended to import the plugin system In debug mode run.py put src/plugin_system itself on sys.path and printed "Plugin system import successful" without importing anything. Nothing imports plugin_system modules by bare name, so the path entry did nothing either. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: delete tests that test nothing - test/plugins/test_{basketball_scoreboard,calendar,clock_simple, odds_ticker,soccer_scoreboard,text_display}.py skip everywhere the named plugins are not installed, including CI (LEDMATRIX_PLUGINS_DIR holds only the fixture plugin); test_plugin_matrix.py already covers every discovered plugin. Their PluginTestBase and the fixtures only it used (plugins_dir, mock_display_manager, mock_cache_manager, mock_plugin_manager, base_plugin_config in test/plugins/conftest.py) go with them. - test_plugin_system.py: test_discover_plugins (body was `pass`) and test_dependency_check (a comment), plus the test_plugin_manager fixture only the former requested. - test_display_manager.py: test_draw_image asserted that an image it had just assigned was not None. - test_display_controller.py: the rotation and schedule-override tests re-implemented the run-loop arithmetic inline and asserted on their own result without calling the controller. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: expect one plugin_last_update success stamp after update_all_plugins EveryStampRecordsACompletion required at least two success-path stamps; the second was update_all_plugins, removed as test-only. The worker and synchronous paths share the remaining stamp in _execute_update_now, and the check that every stamp calls _note_update_completed is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
e1ce7189f1 |
fix(install): make one-shot retry() retry, and drop root grants on user files (#606)
retry() in one-shot-install.sh used `if ! "$@"; then status=$?`, where $? is the status of the negation -- always 0. A failed command was never retried and retry() reported success, so a failed `git clone` carried on until a later check noticed the missing checkout. It now retries (3 attempts) and returns the command's status. The two apt steps stay non-fatal: warning and continuing is what they effectively did before, and making them fatal would stop installs that work today. A clone that keeps failing stops the install, as it already did, just sooner and with the one-shot's own error message. Both installers granted the web user NOPASSWD root on display_controller.py, start_display.sh and stop_display.sh. Those files are owned by the user after Step 11's chown, so the grant let the web user rewrite them and run them as root, and nothing ever ran them through sudo. Removed from both installers, with a test that every project file granted as root is a root-owned fix_perms helper. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
342e9164b8 |
fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound Each repeat of an error pattern appended every plugin in the time window to the pattern's list again, so a plugin failing in a loop grew the display process's memory without limit: 3,000 errors from three plugins reached 2.5 million entries. Keep the list unique. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): load a BDF font at its native size instead of PIL's default FreeType rejects any size but a BDF strike's own, and FontManager answered that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf requested at 8 or 10px rendered as PIL's default font. Retry at the native strike, as element_style already does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin toggle failures no longer claim "operation in progress" Every exception in POST /plugins/toggle was mapped to PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation is already in progress". Report the failure as what it is, and record the plugin id in the operation history for form posts too (it read a `data` variable that only the JSON path set). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): route plugin card clicks through handlePluginAction The document-level delegation checked `typeof handlePluginAction`, which is scoped inside the plugin-manager IIFE and so never visible to it. Every card click took a copied fallback that stopped propagation (the grid's own listener never ran), confirmed an uninstall twice, and sent Starlark app uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>. Expose the handler on window and delegate to it. Also run every test/js/unit suite under pytest: they need only node, but CI ran one of the eight. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): apply Rotation durations, WiFi messages and Vegas settings Three settings the web UI saves never reached the display: - Rotation & Durations: display.display_durations was never read. Every plugin inherits get_display_duration() and the plugin was asked first. A saved value now wins. The page shows unsaved screens blank with the plugin's own duration as a placeholder, and saving a blank removes the override, so one save no longer pins every screen. - WiFi status overlay: the controller looked for wifi_status.json one directory above the repo. Both sides now use wifi_manager.get_wifi_status_path(). The message is written by rename so the display never reads it half-written, and the resumed plugin redraws the whole panel afterwards. - Vegas: nothing called coordinator.update_config(), so saved Vegas settings never reached a running scroll. They are now queued when display.vegas_scroll changes, and applied while Vegas is stopped too, so a disable then re-enable works. The follower's scroll-speed default (75) now matches VegasModeConfig's (50). Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows with each plugin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep affected_plugins order when serialized; guard non-Element targets ErrorPattern.to_dict() ran the now-ordered list through set(), so get_error_summary() listed plugins in an unstable order. The document-level card-action listener called event.target.closest() without checking the target is an Element. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
967f3a0567 |
fix(install): parse the sudoers rules before installing them (#602)
Both installers generated the ledmatrix_web rules and copied them straight into /etc/sudoers.d without ever parsing them. Every rule is built from `which` lookups, so an empty or surprising path produces a malformed drop-in -- and a malformed file in /etc/sudoers.d makes sudo refuse every command for every user. On a headless Pi that is unrecoverable over SSH. first_time_install.sh now runs `visudo -c` on the generated file and, if it does not parse, prints what visudo said and leaves the installed file untouched rather than replacing it with a broken one. configure_web_sudo.sh does the same before it offers the rules for confirmation. first_time_install.sh also built the file at a fixed /tmp path as root; mktemp now picks the name. test/test_sudoers_is_validated.py renders the installer's own sudoers heredoc and checks the result with visudo -- the check neither installer had -- and asserts the install stays gated on it. Claude-Session: https://claude.ai/code/session_01Dby94z9PV3zVM25fqGNXTt Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
cf02538d2e |
fix(sports): ask the endpoint the league actually publishes for standings (#600)
* fix(sports): ask the endpoint the league actually publishes for standings ESPNDataSource.fetch_standings tried /standings first regardless of league and fell back to /rankings only on a 404. College leagues answer /standings with a 200 that carries no poll, so the fallback never fired and the poll came back empty every time. Nothing failed; the rank badge simply never appeared, and anything keyed off rankings quietly did nothing. Endpoints are now ordered by whether the league publishes a poll, a 200 that lacks the key counts as a miss so a league answering both still ends up with whichever one carries the poll, and only a 404 is treated as routine -- it is how a league says it has none. A connection error, a timeout or an unparseable body is logged as an error again. This is the implementation the football, baseball and hockey boards already ship; core was the last copy still on the old one. Verified against live ESPN: mens-college-basketball returns a populated rankings key where it previously returned nothing, and nba still resolves from /standings alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(standings): stop the endpoint handler from swallowing its own bugs Addresses both CodeRabbit findings on #600. The handler caught `Exception`, so an AttributeError or TypeError raised while *inspecting* the payload was indistinguishable from an endpoint that failed. The loop would move on and, if the other endpoint had nothing either, return {} -- silently dropping rankings for a league that has them. That is the precise failure this function was written to fix, so the handler was able to reintroduce it. Only the request is guarded now. `requests.RequestException` covers the transport failures and `ValueError` covers a body that will not parse; payload inspection happens after the handler, where a bug surfaces instead of being logged as a missing poll. A non-dict payload is treated as a miss explicitly rather than by tripping over `.get`. Tests: the fallback paths had no coverage -- the old single-endpoint code would have passed the suite unchanged. Added order assertions for both league kinds, a 200-without-a-poll fall-through, 404 and non-404 recovery, a non-object payload, and a guard proving a bug is no longer swallowed. `test_fetch_standings_returns_empty_on_error` faked a transport failure with a bare `Exception`, which only passed because the handler caught everything. It now raises ConnectionError, which is what actually happens. Verified by mutation: restoring standings-first fails 5 tests, restoring the catch-all fails the bug-not-swallowed guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
81e1bc596f |
fix(sports): stop the idle back-off sleeping through a kickoff (#599)
* fix(sports): stop the idle back-off sleeping through a kickoff A league with no live games backs its poll off as empty checks mount, capped by live_idle_max_interval. The escalation counts empty looks and nothing else, so a league three hours before kickoff is indistinguishable from one three months out of season. Both reach the ceiling -- and the ceiling then *is* the blind spot. Measured on two rigs on 2026-09-19: gaps of up to 928s between looks, ten of them at or above 900s. Reproduced in the wild on 2026-09-20, where an unpatched rig sat for fifteen minutes with eight NFL games in progress and had not noticed any of them. That is the "it doesn't pick up new live games until I restart it" report -- restarting being the one thing that forces an immediate look. The clamp costs no extra request: the live fetch already downloads the whole day's scoreboard, upcoming games included, so the earliest start still ahead of us falls out of the payload the manager already has. Before a kickoff the wait is shortened so it cannot run past it; just after one, the live cadence is held for _KICKOFF_GRACE_SECONDS, because a provider that has not yet flipped the status would otherwise look like another empty check and escalate the back-off again, right when the game is starting. The grace window needed a second pass. A soak caught it as dead code: the just-passed kickoff was replaced by the next fixture on the card the instant it passed, `now < start` went true again, and the back-off returned to its ceiling. Observed live -- the rig polled at 13:00:45, found nothing because ESPN had not flipped the status, then went quiet for a quarter of an hour. A kickoff inside the grace window is now kept. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * test(sports): pin absolute tolerances and correct a wrong grace expectation pytest.approx defaults to a relative tolerance. On a unix timestamp that is roughly 1790 seconds, so every kickoff assertion here was effectively vacuous -- it called a kickoff half an hour away "equal". All seven now pin abs=1. That hid a wrong expectation. test_an_earlier_kickoff_still_wins_during_the_grace asserted a game ten minutes out should displace one that kicked off moments ago. It should not, and the code does not: while the grace holds, the wait is the live cadence (30s), which is strictly tighter than clamping to the nearer kickoff would give (~600s). Letting the candidate win would set a ten-minute wait at the exact moment games are starting -- the dead grace window this branch exists to fix. The test now pins the real behaviour plus the safety property that makes it correct, and is renamed to say what it checks. Reported by CodeRabbit on the PR. The finding was right that code and test disagreed; the suggested fix was the wrong way to resolve it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
19686ab761 |
fix(web): plugin settings form shows schema defaults for unsaved keys (#597)
The server-rendered plugin settings partial rendered straight from the saved config, so an option added in a plugin update (geochron 1.2.0's show_date / show_date_line, default true) drew as an unchecked box, and the save route's missing-checkbox handling then stored it as false. Enum dropdowns likewise showed their first option instead of the default. - _load_plugin_config_partial runs the stored section through prepare_plugin_config (as GET /plugins/config does) before masking secrets, so a secret's schema default is masked too. - render_field falls back to the field's own default, covering children of objects that declare a default of their own (where the defaults extraction stops). - The legacy-boolean parity test now compares against the config the plugin actually runs with (defaults included). Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
92f1960d00 |
perf(sports): fetch ESPN date chunks concurrently (#596)
* perf(sports): fetch ESPN date chunks concurrently Since ESPN started rejecting `dates=YYYYMMDD-YYYYMMDD` on 2026-09-15, one season request became a chunk per month -- and a month over the 500-event cap becomes a request per day. A cold college-baseball season is about 130 requests, and they went out one at a time. That is slower than the 20s budget `_update_plugins()` shares across every plugin at startup, so scoreboards were logging `update() timed out` on first run and being deferred to the scheduled tick with nothing on the panel. Measured on a Pi 4 against live ESPN, March+April college baseball (63 requests, 3101 events): 11.2s sequential, 1.6s concurrent. Over a whole boot that moved football-scoreboard, ledmatrix-flights and birdnet-go inside the budget -- 13 plugins deferred before, 10 after. Chunks now go out six at a time, in two passes: months and edge days first, then the days of any month that came back capped. Six keeps the shared Session under requests' default pool_maxsize of 10, so no connection is discarded. Merged events still follow `espn_date_chunks` order -- a capped month's days are spliced back into its own slot -- so the payload does not depend on which request won the race. Request order is no longer significant, so the three tests that pinned it compare the chunks as a set and keep asserting the merged event order, which is the part callers actually see. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): drop capped month payloads before fetching their days Review of the concurrent chunk fetch found it raised the worst-case peak memory more than the concurrency explains. The old loop discarded a month that came back at the 500-event cap the moment it saw it; the rewrite kept every capped month alive in `results`/`slots` until all of their day requests had finished. Measured on a Pi 4 fetching 20260201-20260531 college baseball (four capped months, 5462 events), peak RSS growth over the call: sequential (main) 83 MB concurrent, months retained 121 MB (+43) concurrent, one worker 108 MB -- the retention alone was +25 concurrent, months dropped 98-100 MB (+16) docs/LOW_MEMORY_BOARDS.md puts a 1 GB Pi 3B+ at under 200 MB of headroom, where running out makes the board unreachable until a power cycle, so the difference matters. The remaining +16 MB is six responses parsing at once; three workers saved about 6 MB more, within run-to-run noise, so the worker count stays at six. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sports): state what ESPN_CHUNK_WORKERS was measured to do, not more The comment claimed the sequential fetch made scoreboards blow the 20s startup update() timeout. A boot on this branch still deferred 12 plugins and timed out baseball-scoreboard while its season fetches took 0.74s and 1.12s: the startup budget is spent on other per-plugin work. Say what was measured -- 17.7s sequential, 2.6-3.3s concurrent -- and nothing else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |