mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-10 17:16:36 +00:00
478516dbd19d0291d88b7cf7ea1c96178404850a
2072
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
478516dbd1 | Merge branch 'claude/frame-timing-harness' into claude/offscreen-rendering | ||
|
|
690407860e |
Merge branch 'claude/hdpi-scroll-performance-antialiasing-4ae609' into claude/offscreen-rendering
# Conflicts: # CHANGELOG.md |
||
|
|
99a3608516 | Merge remote-tracking branch 'origin/main' into claude/frame-timing-harness | ||
|
|
42ff3825b4 | Merge remote-tracking branch 'origin/main' into claude/hdpi-scroll-performance-antialiasing-4ae609 | ||
|
|
edfcd9e2a1 |
Merge remote-tracking branch 'origin/main' into claude/frame-timing-harness
# Conflicts: # CHANGELOG.md |
||
|
|
3a81f38f09 |
fix(web): uniqueItems saves, /health count, Vegas order wipe; one list-repair helper (#638)
* fix(web): drop repeats from uniqueItems lists before validating a plugin save dedup_unique_arrays lost its only caller in #330, so submitting a value a uniqueItems list already holds (a stock symbol saved once and posted again) failed the whole save with a validation error. _prepare_plugin_config_for_save runs it again just before validation, which covers both POST /plugins/config and plugin sections posted to /config/main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): /health counts the discovered plugins and logs the checks it fails The plugin check counted plugin_manager.get_available_plugins(), which PluginManager does not have, behind a hasattr guard that made plugin_count 0 on every device. It now counts the discovered manifests, discovering first when nothing has been scanned yet. The config, plugin and hardware checks answered "see logs for details" without logging anything. Each now logs a warning with the traceback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): store refresh no longer claims a commit-metadata refresh POST /plugins/store/refresh read fetch_commit_info (or fetch_latest_versions) only to append "(with refreshed commit metadata from GitHub)" to its message. It never fetched any: the route re-downloads the registry and nothing else. search_plugins takes the flag, but it reads commit info through its cache, so passing it on would not refresh anything either. The flag is ignored now and the message says what happened. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): refuse a malformed Vegas plugin order instead of clearing it A vegas_plugin_order or vegas_excluded_plugins value that was not JSON, or not a list, was stored as [] and the save answered 200, so a bad value wiped the saved order or exclusions. Both now answer 400 and save nothing, the way plugin_rotation_order already did; the three share one parser. A list that holds anything but plugin-id strings is refused as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): per-plugin health and metrics read the display service's latest GET /plugins/health/<id> and /plugins/metrics/<id> called get_health_summary and get_metrics_summary without force_reload, so they answered with whatever the web process read first and kept in memory, while the display service kept writing newer state. They now pass force_reload=True, as the list routes do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin config reset saves through the shared atomic save POST /plugins/config/reset called config_manager.save_config directly, so it took no backup, and a failed write escaped as an unhandled exception. It then handed on_config_change the raw stored section, not the prepared config a loaded plugin runs with. It now saves through _save_config_atomic with a backup, answers CONFIG_SAVE_FAILED when that fails, and notifies with _prepared_plugin_config, as POST /plugins/config does. POST /plugins/toggle carried its own copy of _save_config_atomic's save_config_atomic-or-save_config fallback; it calls the shared helper now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): one reading and one "unavailable" for each system metric system_metrics.collect_system_metrics() promised None for a metric it could not read, but returned cpu_temp as 0 off a Pi, and the whole no-psutil fallback as zeros. GET /system/status measured the same numbers a second time with its own code, and answered None there. Now both come from collect_system_metrics(), and "unavailable" is None everywhere. /system/status keeps its 0.1s CPU sample and its 10s cache, and gains nothing it did not already send. Two differences: without psutil it answers 200 with null metrics instead of 503, and a disk it cannot stat is null instead of a 500. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): /display/current sends the snapshot as-is and logs a failed read GET /display/current PIL-decoded the preview snapshot and re-encoded it before base64-ing it, spending CPU on the Pi to send the same picture, and dropped any failure with `except Exception: pass`. The /stream/display SSE stream already passed the PNG's bytes straight through. Both now read through web_interface/display_preview.py and answer with the same payload. A missing snapshot is still a null image; any other read failure is logged as a warning. /health reads the snapshot path from the same module. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): one helper puts a submitted plugin config's lists back The plugin-config save turned position-keyed dicts ({"0": ..., "1": ...}) back into lists in five copies: four in the form path's fix_array_structures (whose prefix branches never ran, since no caller passed one), and _fix_json_arrays on the JSON path. It then force-fixed the news plugin's feeds.custom_feeds by name, in case the generic pass had missed it. src/web_interface/config_arrays.coerce_array_shapes now does it for both paths, custom_feeds included. ensure_array_defaults duplicated _fix_none_arrays and is gone. In the same function: the union-type re-checks that the null handling above them made unreachable, the "(temporary)" random_seed debug log, and a commented-out log line are removed. A failed validation is logged once as a warning, not four ERROR lines and a WARNING. Element types are left to normalize_config_values, which already converted them for both paths. One difference: the form path no longer adds an empty {} for a nested object the post left out that has no defaults. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): import at module top and log through the module logger The web_interface.cache imports in config.py and fonts.py were wrapped in `except ImportError` fallbacks. It is an in-repo module that imports nothing from the project, so it cannot fail to import; it is imported once at module top, as system.py now does. cache.py's docstring said blueprints import it lazily "to avoid circular imports"; it now says why that is unnecessary. Five logging.error calls in the dim-schedule GET and three logging.warning calls in plugins.py went to the root logger; they use the module logger. Function-local re-imports of json, os, shutil, logging and Path, all already imported by the module, are gone. The `import os` inside two except blocks of save_plugin_config also made os a local name for the whole function. execute_plugin_action's step-1 handler gets a comment saying why it stays: it looks like a copy of the blueprint handler, but without it a TimeoutExpired from the plugin's script would reach the route's own `except subprocess.TimeoutExpired` and be answered as a 408. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): app.py loses dead CSRF and reconciliation state, comments fixed - csrf was always None, so `if csrf: csrf.exempt(...)` never ran, and its note that the api_v3 blueprint "is exempted above" named an exemption that does not exist. Both are gone; the reason there is no CSRF protection stays, shortened. - The SSE rate-limit comment called the default "tight" at 20 per minute. The default is 1000 per minute and the streams' 200 is the tighter one; the comment now says so. The limits are unchanged. - _reconciliation_done was written and never read. The docstring that explains why reconciliation runs once keeps its reason, in the present tense. - Removed: a dangling "import cache functions" comment with no import under it, a "security check ... within project_root" label on an existence check, the "(simplified version)" narration, and the note that no redirect route is needed. The preview loop's sleep comment no longer mentions a PIL encode that the loop does not do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(web): api_v3 comments name the package __init__, not a _common module Every route module's docstring said the shared blueprint comes "from ._common", a module the package split never created; they name the package __init__. The PROJECT_ROOT comment described the path from _common.py; it now describes this package and keeps the incident it guards against. The "(corrected) in this commit" note in resolve_pull_command and the /health comment the split's mechanical time -> _pkg.time rewrite garbled ("Stamp the start _pkg.time") read correctly again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): drop hasattr checks for attributes PluginManager always has PluginManager.__init__ sets health_tracker and resource_monitor (to None until they are configured), so the seven hasattr(api_v3.plugin_manager, ...) guards in the health, metrics and limits routes were always true. The falsy checks that do the work stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): pages_v3 dispatches partials from a dict with one error handler load_partial chose a loader through a fourteen-branch if/elif, and thirteen of the loaders then wrapped themselves in the same try/except, logging "Error loading partial" without saying which. The route now looks the name up in _PARTIAL_LOADERS and has the one handler, which logs the partial's name. The loaders just render. _load_tools_partial keeps its own messages. The search index's _partial_html already catches a loader that raises. serve_plugin_web_ui repeated _plugin_dir_for inline (containment plus the ledmatrix- prefix fallback); it calls it now. Also removed: the unused markupsafe.escape import, function-local json/Path re-imports, and unused exception bindings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove unused imports, locals and a try that cannot fail - get_error_aggregator was imported by the api_v3 package and used by no one; seven names config.py imported, and Path in misc.py and logging in plugins.py, likewise. - branch_info in install_plugin was built and never logged; test_config in /health was bound and never read (the load_config call is the check). - An f-string with no placeholders in the asset upload route. - _installed_plugin_ids wrapped list(manifests.keys()) in try/except; _discovered_plugin_manifests always returns a dict. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): start.py logs its startup lines and drops unreachable branches The startup banner went to stdout with print(); it goes through a logger now, which the app import has already configured, so it reaches the journal with a level and timestamp like every other line. The "no addresses" branch is gone: get_local_ips() always returns at least "localhost". The except around app.run re-raised "only if it's not a client disconnection error" from inside the branch that had just established it was one, so that raise could not run. It is one check now, on a named tuple of the errnos, which the werkzeug log filter uses too. The comment on threaded=True counts three SSE endpoints, which is how many there are. Trailing whitespace is stripped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): save_main_config names its General fields once The General tab's field names were listed twice, once to detect a General form post and again, with four more, to keep the remaining-keys merge from storing them as top-level keys. GENERAL_FIELDS and _MAPPED_TOP_LEVEL_FIELDS hold them now, and the four per-section skip checks are one set. The comment on that merge said plugin configs are handled "here too", and "(including plugin keys)". Plugin sections are handled and removed from the body before it runs; the comment says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): plugin directories come from the plugin manager only Six lookups fell back to PROJECT_ROOT/plugins/<id> when there was no plugin manager: GET /plugins/config's of-the-day data, POST /plugins/action, the plugin static-file route, the calendar credentials upload and the calendar OAuth routes. The loader never scans plugins/ (PluginManager.discover_plugins reads only the configured directory, plugin-repos by default), so what they found there was a plugin that never runs. _plugin_directory() asks the manager and answers None without one, which each route already reports as "not found". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): web-backend Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
19ee84a760 | Merge remote-tracking branch 'origin/main' into claude/hdpi-scroll-performance-antialiasing-4ae609 | ||
|
|
9a705c2ca4 |
revert(vegas): gate only the prefetch thread; gating ESPN fetches measured worse
|
||
|
|
7b90759252 |
fix: /errors stack traces, Wi-Fi disconnect and save, plugin fonts, API cache TTL (#636)
* fix(errors): record the exception's own stack trace record_error() called traceback.format_exc(), which only sees an exception while its except block is running. plugin_executor records exceptions caught on a worker thread after that block has ended, so every trace on /errors read "NoneType: None". The trace is now built from the exception's __traceback__. The executor's log call had the same problem with exc_info=True and now passes the exception. record_error() also merged LEDMatrixError context into the caller's dict in place; it now works on a copy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(wifi): point at configure_wifi_permissions.sh instead of a sudoers list The module docstring told users to grant NOPASSWD sudo on iptables and ip. configure_wifi_permissions.sh refuses those grants on purpose: a wildcard rule for either runs an arbitrary program as root. Point at the script and say why it leaves them out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): disconnect finds the saved profile by SSID disconnect_from_network() asked `nmcli -f NAME,802-11-wireless.ssid connection show` for the profile to take down, but nmcli rejects that column for `connection show`, so the lookup always failed and only the device was disconnected. The per-profile lookup _connect_nmcli() already used is now _find_profile_for_ssid(), and both callers share it. It also splits terse output on the last colon and unescapes "\:", so a profile name containing a colon is found. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): write wifi_config.json atomically and report a failed save _save_config() opened the file for writing in place and swallowed any error, so a wifi_config.json left owned by root made the web toggle for auto-enabling AP mode report success while nothing was saved, and a crash mid-write could truncate the file. It now uses atomic_write_json, which also keeps the file's owner and shared group when root saves it, and returns False on failure. POST /wifi/ap/auto-enable answers 500 in that case. The file is now written with indent=4, like the other config files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): resolve plugin:// fonts in the plugin's own directory FontManager looked for a plugin's bundled fonts under Path("plugins") / plugin_id: relative to the process cwd, and not the default install directory (plugin-repos/), so a manifest's plugin:// fonts never loaded. register_plugin_fonts() takes an optional plugin_dir, and PluginManager passes the directory it loaded the plugin from. Callers that omit it get a lookup in the configured plugin_system.plugins_directory, then plugins/, resolved against the install root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(api-helper): cache responses for the requested cache_ttl APIHelper.get(cache_ttl=...) and set_cache(ttl=...) dropped the ttl on the claim that CacheManager does not support one, but CacheManager.set() takes a ttl, stores it with the entry, and both cache tiers honour it over a reader's max_age. Without it every response expired after the 300-second default read age, whatever the plugin asked for. The ttl is now passed through, and the cache read passes cache_ttl as max_age for entries written without one. The class docstring describes what the helper actually does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(style): one scale range for the schema, element_scale and LogoHelper The generated Scale field allowed 0.1 to 10, element_style's reader capped at 10 with no floor, and LogoHelper accepted 0.05 to 8 and reset anything else to 1.0. A logo scale of 9, which the form accepts, drew at the shipped size. MIN_ELEMENT_SCALE / MAX_ELEMENT_SCALE (0.1, 10.0) in src.element_style are now the schema bounds and the clamp every reader applies through coerce_scale(): a positive number outside the range is clamped, and anything that is not a finite positive number means the default. That also stops element_scale() passing NaN through, since min(nan, 10.0) is nan. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(logos): placeholder lands at the requested path; empty logos list download_missing_logo() wrote its fallback placeholder to <normalize_abbreviation(abbr)>.png in the logo directory rather than to the logo_path the caller passed, so it could return True while nothing existed where the plugin looks (e.g. "TA&M.png" vs "TAANDM.png"). create_placeholder_logo() takes an optional filepath, and download_missing_logo passes the requested one. download_missing_logo_for_team() only caught KeyError, so a team whose "logos" list is empty raised IndexError; it now treats KeyError, IndexError and TypeError as "no logo URL". The placeholder is drawn with PLACEHOLDER_SIZE / PLACEHOLDER_BG, the constants is_placeholder_logo() recognises it by, instead of repeated literals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): resolve bundled font paths against the install root TextHelper's default font_dir, the logo placeholder's font and FontManager's font_overrides.json were all relative to the process cwd, so a process started anywhere but the install root (the plugin safety harness, a manual run, a unit without WorkingDirectory) drew with PIL's default face and read no overrides. They now go through font_layout.resolve_asset_path; the overrides file sits in the install root's config/. The resolver docstrings described an order the code does not follow: resolve_asset_path never consults the cwd, and sports_shared's _resolve_font_path tries the cwd first. Both docstrings now say what the code does, and _resolve_font_path calls resolve_asset_path instead of probing FontManager for it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(sync): the web UI reads the sync status file the display writes sync_manager writes its status to tempfile.gettempdir(), but GET /api/v3/sync/status read a hardcoded /tmp/led_matrix_sync_status.json and defaulted the port to a literal 5765. Wherever TMPDIR is set (or on any non-/tmp host) the page only ever showed "starting". The endpoint now uses sync_manager.STATUS_FILE and SYNC_PORT. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(http): the rankings resolver sends the project's User-Agent DynamicTeamResolver fetched ESPN rankings with a bare requests.get, so it sent python-requests' default User-Agent, which ESPN rejects; the AP_TOP_N favourites then resolved to nothing. It now sends DEFAULT_HTTP_HEADERS. BaseOddsManager carried its own copy of the User-Agent string and now uses the same shared headers (which also adds Accept-Language). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(backup): record the core release and read the configured plugin dir The manifest's ledmatrix_version came from a VERSION file that does not exist, then from .git/HEAD: a 12-character sha, or "ref: refs/he" when the branch's ref was packed. It is now src.__version__. list_installed_plugins() scanned a hardcoded plugin-repos/, so on an install whose plugin_system.plugins_directory points elsewhere, plugins missing from plugin_state.json were left out of the backup. It now reads the configured directory from config/config.json, defaulting to plugin-repos. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(startup): report a missing display section once A config without a display section produced three errors for the one problem ("Missing required configuration key: display", "Display configuration is missing or empty" and "Display configuration is missing"), and an empty one produced two. _validate_config now reports it once, as a missing key or an empty section, and _validate_display_config leaves it to that. The module docstring said the validator fails fast; nothing in the display service calls raise_on_errors(), so it now says the errors are reported and startup continues. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(wifi): share the copied blocks and name the AP constants - _parse_nmcli_wifi_list() is the one parser behind _scan_nmcli and _scan_nmcli_cached. - _verify_connected(), _wait_for_device_idle(), _failsafe_ap() and _mark_forced() replace blocks that were pasted two or three times in the connect and enable-AP paths. The device-idle wait now checks before its first one-second sleep instead of after it. - _check_command() calls _find_command_path() instead of repeating it. - AP_IP, PORTAL_PORT, AP_PROFILE_NAME and AP_PROFILE_NAMES name values that were spelled out 14, 12, 8 and 2 times; the two deletion loops now walk the same tuple. The iwconfig status path compares the AP address exactly: startswith() also skipped 192.168.4.10-19. - Dropped a second WIFI.SIGNAL query that repeated the first, a no-op "if ssid: continue", the try/except around _connect_wpa_supplicant's constant return, and a second save of a scan scan_networks already saves. - _ensure_wifi_radio_enabled's docstring says it returns True when the radio state cannot be read at all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(config): drop dead branches and history comments in ConfigManager - The module docstring pointed plugin authors at update_plugin_config(), which does not exist; it now names save_config_atomic() and save_raw_file_content(). - load_config's FileNotFoundError handler tested the message for "config_secrets.json", but a missing secrets file is handled where it is read, so only config.json reaches it; the check is gone. - save_raw_file_content's `file_type == "main" or "secrets"` guard was always true (anything else raised earlier). - get_raw_file_content('secrets') already returns {} for a missing file, so the os.path.exists() in front of two calls to it is gone. - Comments that narrated earlier behaviour are rewritten as what the code does now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(background-data): present-tense comments, drop unused API - Comments that told the history of each fix (what "used to" happen, "the old per-delivery release") now state the invariant the code keeps. - get_statistics() no longer reports a constant 'queue_size': 0, and the uncalled clear_completed_requests() is gone (_cleanup_completed_requests does that job on every completion). Neither is referenced in core, the web UI or the plugin monorepo. shutdown_background_service() has no production caller either, but it is the only way to tear down the get_background_service() singleton, which the tests rely on, so it stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(odds): drop the unread cache_ttl and merge the odds_data branches BaseOddsManager loaded base_odds_manager.cache_ttl from config and never used it: cached odds live for the update interval (get_odds' ttl=interval). No core or monorepo code reads the attribute, so it is gone along with its log line. The two consecutive `if odds_data:` blocks are one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(backup): one table for the single-file sections config, secrets, wifi and ytm_auth were each spelled out in create, preview, validate and restore. _SINGLE_FILE_SECTIONS lists them once, with the RestoreOptions flag that restores each, and all four walk it. Restore error messages keep their wording ("Failed to restore <file name>"). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(fonts): drop FontManager's write-only state and duplicate logs - fonts_config, font_metadata and font_dependencies were written and never read; the performance_stats keys font_load_times, render_times, total_renders and the per-call "resolve" timings (_record_performance_metric) likewise. get_performance_stats() reads only the counters that remain. Nothing in core or the plugin monorepo references any of them. - A failed BDF load was logged twice, by _load_bdf_font and again by get_font; get_font's line is the one kept. - Removed "NEW:" and commented-out cozette entries, the "Copy font to assets/fonts" comment on code that copies nothing, and local imports of names the module already imports. The deprecated add_font() now resolves assets/fonts against the install root. The @deprecated methods stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(text-helper): cache loaded fonts; drop the pre-textlength fallback TextHelper declared _font_cache, cleared it and reported its size, but never stored anything in it. load_fonts() now keeps each (file, size) it loads there, so clear_font_cache() and get_font_cache_stats() mean what they say and repeated load_fonts() calls reuse the fonts. get_text_width() no longer catches AttributeError for Pillow releases without ImageDraw.textlength; requirements.txt pins Pillow>=12.2. The class docstring describes what the helper does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(common): fix wrong docstrings in api_helper, permission_utils, snapshot_policy - permission_utils called 0o2775 "sticky bit"; the 2 is setgid, which is what makes new files take the directory's group. - snapshot_policy pointed at web_interface/blueprints/api_v3.py, which is a package now; the health check is in api_v3/misc.py. - APIHelper.clear_cache() lost a history note and a fallback to a clear() method that neither CacheManager nor the testing MockCacheManager has. The session headers are built from DEFAULT_HTTP_HEADERS instead of a copy of them, and the module docstring says what the module offers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(sports): present-tense comments in the shared scoreboard renderers - sports_scroll and sports_game_renderer comments that referred to "this PR", "the old flat 128px card" or what the renderer "previously" did now describe the current behaviour and its reason. - The block explaining why non-finite settings are rejected sat above _score_reserve_width; it describes _center_gap_width and now lives in it. - unshare_element_fonts wrapped its import of font_layout.load_truetype in an `except ImportError` that cannot fire inside core; the import stays at call time so tests can spy on the pinned loader. - sports_card docstrings that told the history of a fix say what the code does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sports-shared): drop dead code, name the ESPN limit - _get_weeks_data asked for limit=1000, which fetch_espn_scoreboard clamps to ESPN_MAX_LIMIT anyway; it now names that constant. Its unused `immediate_events = []` is gone. - _get_season_schedule_dates() returned ("", "") and has no caller in core or the plugin monorepo. - _should_log keeps its warning_type parameter (part of the inherited signature, though nothing in core or the monorepo calls it) and its docstring says the cooldown is shared across types. - An unused ImageFont import is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sync): one follower-mode switch, shared panel defaults - The class docstring said the leader sends PNG frames. Frames go over UDP as raw RGB; PNG is only the Vegas scroll image sent over TCP. It now describes both paths. - _enter_follower_mode() replaces the two copies of "note the leader, switch from standalone to follower, log, write status" in the frame and scroll-position handlers. - The rows/cols fallbacks use DEFAULT_ROWS / DEFAULT_COLS from src.display_geometry, as chain_length already did. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(style): drop _layout_axis, name the layout group title - ElementStyleResolver._layout_axis() had no caller in core or the plugin monorepo. - _element_block_from_spec checked spec['size'] was a dict again after size_spec already had; it reads size_spec. - The "Layout Offsets" title written into three generated schema blocks is _LAYOUT_TITLE. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(logo-helper): say what the placeholder draws; name the 1.5 box factor - _create_placeholder_logo's docstring said it draws the team abbreviation; it draws an outlined grey box and nothing else. The docstring says so, and the "in a real implementation you'd want text" comments are gone. - The 1.5 x panel default logo box, written out six times, is DEFAULT_LOGO_BOX_FACTOR. - ImageDraw is imported with Image at the top of the module. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(logos): drop dead code and a duplicate regex in logo_downloader - _SAFE_LEAGUE_CODE_RE was the same pattern as _SAFE_LEAGUE_RE; both checks use the one. - get_logo_filename_variations reassigned the TA&M case to the list it already had; the function returns the two names directly. - _get_team_name_variations() had no caller in core or the plugin monorepo. - fetch_single_team's docstring was copied from fetch_teams_data; a log message read "for{team_id}". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor: drop the Pillow<9.1 resample shim and a catch-and-reraise - adaptive_images fell back to Image.LANCZOS/NEAREST for Pillow < 9.1; requirements.txt pins Pillow>=12.2. RESAMPLE_LANCZOS and RESAMPLE_NEAREST keep their names (src.common re-exports them). - CacheManager.save_cache caught CacheError only to re-raise it; the disk write is now called directly, with the same result. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(api-helper): stop the real CacheManager's cleanup thread The cache-lifetime tests built a CacheManager and left its cleanup thread's class-wide claim on the directory in place, which broke test_cache_cleanup_thread_ownership when it ran later in the session. The fixture now stops the thread on teardown. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): core-common Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b11bcfa204 |
fix(plugins): store and plugin-manager bugs; tidy src/plugin_system (#635)
* fix(store): don't read a ZIP-installed plugin's remote from the LEDMatrix repo update_plugin looked up remote.origin.url with `git -C <plugin> config --local` for plugins that are not git checkouts. Under plugin-repos/ git walks up to the enclosing LEDMatrix repository, so the lookup returned LEDMatrix's own URL and a plugin missing from the registry was "reinstalled" from the LEDMatrix repo. Only ask git when the plugin directory has its own .git, the test _get_local_git_info already uses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(schema): report each missing required field once, by name validate_config_against_schema ran its own required-fields loop after Draft7Validator.iter_errors, which already yields one `required` error per missing field, so every missing top-level field was listed twice. The validator's copy also printed the schema's whole `required` list ("Missing required property '['api_key', 'city']'") instead of the field. Drop the loop and take the field name from the error itself. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(store): stop mangling repository URLs that contain ".git" install_from_url and fetch_registry_from_url cleaned URLs with `rstrip('/').replace('.git', '')`, which removes ".git" anywhere: https://github.com/user/my.github.io became .../myhub.io, so installing or browsing that repository asked GitHub for one that does not exist. Add src/plugin_system/repo_urls.py with one anchored normalize_repo_url(), same_repo() for comparisons, github_owner_repo() and github_api_headers(), and use them for the five copies of the owner/repo parsing and GitHub headers in the store and for saved repositories. GitHub URLs are now recognised by urlparse().hostname everywhere: _get_latest_commit_info used a substring test, and _install_from_monorepo_api parsed any host. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(store): install a repository whose only branch is not main/master _install_via_git returned None both when every clone failed and when the last-resort clone of the repository's default branch succeeded. _install_plugin_impl papered over it with `and not plugin_path.exists()`; install_from_url did not, so a repository whose only branch is e.g. `develop` was cloned, then treated as a failure, then "downloaded" from main/master archives that do not exist. After a default-branch clone, return the branch the clone checked out (read from .git/HEAD), so None means failure and nothing else, and give both callers the same `branch_used is None` fallback. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): judge the memory limit on each call's own growth monitor_call stores `metrics.memory_mb = max(previous, growth)`, and _check_limits compared that high-water mark with max_memory_mb. It never decreases, so once one update() grew the process past the limit every later call raised ResourceLimitExceeded and the circuit breaker kept reopening. Pass the call's own RSS growth to _check_limits; keep the high-water mark for reporting and document what it measures. Remove ResourceMetrics.update_average_execution_time: nothing called it, and it overwrote the running total with the average. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): reload_plugin re-reads the manifest from the discovered directory reload_plugin read `plugins_dir / plugin_id / "manifest.json"`, ignoring the discovery map and the plugin_dirs rules. For a plugin whose directory name differs from its manifest id the path did not exist, the re-read was skipped without a word, and the reload kept the stale manifest. Resolve the directory with find_plugin_directory, as load_plugin does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): drop the always-null last_display from plugin state info PluginStateManager reported `last_display` from `_last_display`, which nothing ever wrote, so it was null for every plugin. Recording it in PluginExecutor.execute_display would not help: get_state_info's only reader is the web process, whose PluginManager never calls display(). Remove the field, its dict and get_last_display() (no caller in core, the web UI or the plugin monorepo). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(store): share the rollback and requirements helpers, drop dead code - install_plugin and _reinstall_with_rollback set aside, discard and restore the old copy through _set_aside/_discard_backup/_restore_backup instead of two copies of the same blocks. - The loader and the store run the same pre-pip checks through contained_plugin_dir() and requirements_to_install() in plugin_loader. They still invoke pip differently (sys.executable -m pip vs. the sudo wrapper). `except (BrokenPipeError, OSError)` + `isinstance(e, OSError)` becomes `except OSError` checking errno.EPIPE. - load_module never returns None, so load_plugin's check is gone and the docstring says what it raises. - Remove the always-true JSONSCHEMA_AVAILABLE, the inline re-imports of re and permission_utils, the fake status_result object nobody reads, hasattr(git_error, 'cmd'), a redundant "merge conflict" test and `import traceback` (exc_info=True does it). - Correct comments: install_from_url names the directory for the caller's id when given (not always the manifest id), _get_local_git_info saves one git subprocess (not four), _enrich calls two helpers, search_plugins documents all its arguments, _find_plugin_path states its behaviour instead of a TODO, and history narration is gone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): tidy base_plugin, correct plugin_manager/state comments - base_plugin: drop the unused `import logging`; get_display_duration runs the instance value and the config value through one _positive_seconds() helper instead of two copies of the coercion; the 'static'/'none'/fallback branches of get_vegas_display_mode, which all returned FIXED_SEGMENT, are one; fix the mis-indented validate_config example; say that get_supported_vegas_modes/get_vegas_segment_width are not consulted by core (kept, plugins override them). - schema_manager: import expand_style_elements normally rather than swallowing an ImportError of a core module. - plugin_manager: the plugins directory is the configured one (plugin-repos/ by default), not plugins/; get_config() returns the live dict, not a copy, so the interval cache comments say what it saves. - state_manager: config_version and the file version are not used to detect corruption; say what they are. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): stop writing data/plugin_operations.json PluginOperationQueue wrote its finished-operation history to data/plugin_operations.json after every operation, and read it back only into its own in-memory list, which only get_operation_history() exposes -- and nothing calls that. The operation-history endpoint reads OperationHistory (data/operation_history.json). No code in src/, web_interface/, scripts/ or test/ reads the file. Drop the history_file/lazy_load parameters and the load/save code; the bounded in-memory history stays. web_interface/app.py and the integration test stop passing the removed arguments. An existing data/plugin_operations.json is left in place (data/* is gitignored). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): plugin-system Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3967a6cffc |
fix(security): re-harden root sudo helpers; installer fixes; ARCHITECTURE and PERMISSIONS docs (#640)
* docs: add ARCHITECTURE and PERMISSIONS guides ARCHITECTURE.md maps the processes, the state the display and web services share through the cache, the display loop, the plugin system, the web UI and the update path, with links into the code and a where-to-start table. PERMISSIONS.md lists who owns what after install, both sudoers files (and why iptables is not granted), the polkit rule, and which scripts/fix_perms script to run as which user. Both are linked from the docs index, along with the MQTT bridge README and src/common/README.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: correct stale setup, service and troubleshooting claims - README: quick actions run systemctl on ledmatrix.service (run.py), not display_controller.py; use_short_date_format has no effect; the installer uses system pip with --break-system-packages, not a venv. - CONFIG_DEBUGGING: LEDMATRIX_DEBUG must be "true"; logs are in journald. - GETTING_STARTED, WEB_INTERFACE_GUIDE, TROUBLESHOOTING: enabling a plugin, plugin settings, brightness and Vegas settings apply without a restart; matrix hardware settings still need one. - TROUBLESHOOTING: install dependencies with sudo so the root service sees them; point permission problems at PERMISSIONS.md instead of a project-wide chown. - ADVANCED_FEATURES: real BackgroundDataService stats keys; Vegas hooks return VegasDisplayMode and None falls back to capture; cache files are 0660; fix_web_permissions.sh runs as the web user and does not touch sudoers. - STARLARK_APPS_GUIDE: only the linux-arm64 pixlet binary is downloaded. - HOW_TO_RUN_TESTS: test class examples that exist. - CLAUDE.md: PluginStoreManager, plugin_dirs.py, monorepo installs via the Trees API with ZIP fallback, requirements.txt is optional. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: mark deprecated plugin APIs and state manifest fields once Methods @deprecated("3.7.0") (the set pinned in test_deprecation.py) were shown as current API in the quick reference, API reference, advanced guide, development guide and FONT_MANAGER. Each is now marked deprecated with its replacement. FONT_MANAGER is rewritten around the current API; the override editor is gone and override methods are deprecated. Required manifest fields were stated three different ways. The API reference now has one section: the 7 schema-required fields, the 4 the store refuses without, class_name for the loader, and the 8 to set. The other guides link to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: document every src/common module and every widget - src/common/README.md covered 7 of 17 modules. It now has a table of all of them (purpose, whether plugins import it, release to floor on), a short entry each, and logging advice that matches the code. - SPORTS_UNIFICATION listed two shared modules and called sports_helpers the first; it now lists all six. - The widgets README lists all 28 registered widgets plus the support files, and absorbs the parts that only docs/widget-guide.md had (x-options.labels, x-advanced, x-display hidden, plugin-file-manager). docs/widget-guide.md is now a pointer to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(security): fix_web_permissions.sh re-hardens the root sudo helpers The script chowns the whole project to the web user. That included scripts/fix_perms/safe_plugin_rm.sh and safe_pip_install.sh -- the two helpers /etc/sudoers.d/ledmatrix_web lets the web user run as root -- so running it turned both into a root shell for whoever can edit them. It also re-grouped config_secrets.json away from ledmatrix. After the chown it now does what first_time_install.sh's Steps 11 and 11.1 do: helpers back to root:root 755, and config_secrets.json back to the web unit's User=:ledmatrix 640. Each step is non-fatal and prints the manual command if it fails. Also fixes what the script and its docs claimed: it never configured sudoers, its closing hint pointed at ./configure_web_sudo.sh (wrong path), and the README and ADVANCED_FEATURES.md said to run it with sudo, which it refuses. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(security): validate and harden every sudoers drop-in the scripts write configure_wifi_permissions.sh copied its rules into /etc/sudoers.d/ledmatrix_wifi without `visudo -c`. A malformed drop-in makes sudo refuse every command for every user, which on a headless Pi leaves no way back in. It now checks first and leaves the installed file alone when the rules do not parse, as the other two writers do. (It already used mktemp, so that part of the review did not apply.) It also grants the two literal commands wifi_manager.py runs for NetworkManager's shared-mode dnsmasq drop-in -- `cp /tmp/ledmatrix-nm-dnsmasq.conf .../dnsmasq-shared.d/ledmatrix-captive.conf` and `rm -f` of that file. The directory's mkdir was granted, the file was not. Both are pinned in test_sudo_allowlist_covers_calls.py. configure_web_sudo.sh wrote its rules to /tmp/ledmatrix_web_sudoers_$$, a predictable name in a world-writable directory; it now uses mktemp with an EXIT trap, as first_time_install.sh does. It sets mode 440 on the installed file instead of leaving the temp file's mode, and finds visudo in /usr/sbin when that is not on the user's PATH, which skipped the check silently. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(install): escape the project path in the DNS-fix and MQTT unit renderers install_dns_fix.sh and install_mqtt_bridge.sh substituted __PROJECT_ROOT_DIR__ with the raw path, while the other three renderers go through sed_escape_replacement from lib_systemd_render.sh. A checkout under a path containing `&`, `\` or `|` rendered a corrupted unit from these two only. Both now source the helper and use it, and a test checks that every placeholder substitution in scripts/install uses an escaped value. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(install): stop the installer scripts reporting things that are not true - first_time_install.sh printed "Password: ledmatrix123" for the setup access point. wifi_manager creates it as an open network ("No password" on the panel), so it now says so. - Step 10.1 printed "✓ WiFi management permissions configured" straight after its own failure message; install_wifi_monitor.sh printed "✓ Package installation completed" after a failed apt install. The tick now only follows success. - Step 7 printed "Web dependencies already installed ... in Step 5" in the one branch that runs because Step 5 did not install them, then created .web_deps_installed on that basis. It now warns and leaves the marker off so the next run retries, as the comment below it intends. - check_system_compatibility.sh called Debian 12 Bookworm "full compatibility confirmed" while first_time_install.sh refuses anything but Debian 13. Bookworm, older Debian and non-Debian systems are now errors. Its counters used ((X++)), which under `set -e` exits the script at the first warning or error (the expression is 0), so the check never reached its summary on any system with one. - configure_web_sudo.sh and configure_wifi_permissions.sh finished by testing `sudo -n test -f ...` and `sudo -n nmcli device status`, neither of which is granted, so they always reported a failure. They now ask `sudo -n -l` about commands the new rules do grant, which checks the rule without running anything. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(install): print the completion summary before rebooting With -y -- and so for every one-shot `curl | bash` install, which always passes -y -- first_time_install.sh ran `reboot` about 180 lines before its "Installation Complete / Web UI Access" summary. reboot returns at once, so the summary printed while the Pi was going down and the SSH session usually dropped before the web UI address could be read. The reboot block moves, unchanged, to the very end of the script. The interactive prompt now also follows the summary. Because the summary now runs before the -y reboot, its one command that could fail under `set -Eeuo pipefail` (the SSID lookup, when nmcli reports a connected device but no active network line) gets `|| true`; a missing SSID was already handled as "SSID unknown". one-shot-install.sh prints its "Next steps" after the installer returns, by which time the reboot is under way, so it now says so, and README's Quick Install mentions the automatic reboot. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(scripts): correct wrong comments and messages, drop dead code No behaviour change except the output text noted below. - 2775 is setgid, not the sticky bit (first_time_install.sh Step 3.1, fix_plugin_permissions.sh), and root needs no "PWM hardware access" to plugin files. - The 777 comments in first_time_install.sh Step 3's fallback and fix_assets_permissions.sh said root needs it to write. Root ignores mode bits; the comments now say what 777 actually opens. The 777 itself is unchanged. - apt_remove ends in `|| true`, so Step 12's "Some packages could not be removed" branch could never run; it is gone and the helper stays non-fatal. - detect_web_service_user's comment named Step 8 for the web unit (install_service.sh installs it in Step 7.5) and now says which branch actually runs. - Step 5 described an "already installed" check that does not exist; the ACTUAL_USER comment described the re-exec backwards. - on_error printed a literal "\n" before "Common fixes:". - Dead code: one-shot-install.sh's uncalled fix_tmp_permissions, LEDMATRIX_ELEVATED=1 (never read) on the sudo re-exec, and configure_web_sudo.sh's unused PYTHON_PATH, which also made a missing python3 fatal for rules that never mention it. - start_display.sh / stop_display.sh said "for user: <you>"; the service runs as root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(fix_perms): fix_cache_permissions.sh uses setup_cache.sh's model There were two models for /var/cache/ledmatrix. setup_cache.sh (the installer's Step 2) and install_web_service.sh share it through the ledmatrix group: root:ledmatrix, 2775, files 660, which is also what DiskCache relies on to give files the directory's group. fix_cache_permissions.sh instead made it 777 and re-grouped it to the invoking user's group, undoing that. It now runs setup_cache.sh for /var/cache/ledmatrix and keeps its own handling of ~/.ledmatrix_cache. Dropped: /var/cache/ledmatrix/ placeholder_logos (nothing reads it) and the checks against the `daemon` user (no service runs as daemon). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: pin actions/checkout in the Claude workflows, drop template comments claude.yml and claude-code-review.yml used actions/checkout@v4 while test.yml and release-version-check.yml pin the v4.2.2 commit SHA; they now pin the same SHA. The commented-out starter-template settings (prompt, claude_args, paths, author filter) are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(scripts): index every script and list removal candidates New scripts/README.md gives one line per top-level script and scripts directory, marked keep, dev-only or diagnostic, and lists the eight scripts nothing in the repo refers to as candidates for removal (kept for now). The install, utils and dev READMEs now list the files they were missing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: tighten two checks that mutation testing showed were too loose - The wifi sudoers check matched `visudo -c -f "$TEMP_SUDOERS"` in the error report too, so replacing the check with `if false` still passed. It now requires the command as the condition. - The summary test never had the setup access point up, so reinstating the bogus "Password: ledmatrix123" line went unnoticed. A case with hostapd active now checks the AP is described as open. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(permissions): describe the repaired fix_perms scripts and new WiFi grants Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): docs-scripts Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
e39f65fce0 | Merge remote-tracking branch 'origin/claude/frame-timing-harness' into claude/offscreen-rendering | ||
|
|
d37a3a712a |
style(perf): say why render_bench fell back; mark the stats path as a safe fixed name (Codacy)
Codacy (Bandit B110, B108). The bench's silent except now prints why it read config.json directly. The stats file's fixed name in /dev/shm is safe: write() goes through mkstemp and os.replace, which replaces a planted symlink instead of following it; the comment says so and marks it nosec. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
409b63bc17 |
fix(vegas): tell code the gate must not park in by module name, not path
The gate never parks a thread inside logging, threading, importlib or the cache, and it matched those as substrings of each frame's file path. On GitHub's runners Python lives under /opt/hostedtoolcache, so every stdlib frame said "cache" and the gate never parked anything -- three tests failed there and passed here. A virtualenv under ~/.cache would have done the same on a Pi. Match the frame's module name (f_globals['__name__']) instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
84043468ca |
feat(vegas): ESPN fetches give way to the render thread too
The hourly sports refresh froze hdpi's Vegas scroll for 1.9s: about twenty espn-chunk threads fetching and parsing at once, and the render thread (and the stall watchdog) queued behind all of them for the GIL. The prefetch gate only covered the prefetch thread. Vegas now makes its gate the active one (render_gate.set_active), and render_gate.yielding() gives way through it when there is one and does nothing otherwise. espn_dates wraps each chunk fetch in it (behind the same import fallback as json_body, for the copies plugins bundle), and the background data service wraps each worker. The render thread is never gated -- the first thread to swap is exempt, and a plugin pushing a live refresh from its update thread cannot take its place -- and nested blocks keep the outermost frame as the boundary for the lock checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a6648ec991 |
Merge branch 'claude/frame-timing-harness' into claude/offscreen-rendering
# Conflicts: # src/display_manager.py |
||
|
|
1df6def8f7 | Merge branch 'claude/hdpi-scroll-performance-antialiasing-4ae609' into claude/offscreen-rendering | ||
|
|
6279530c16 | Merge remote-tracking branch 'origin/main' into claude/hdpi-scroll-performance-antialiasing-4ae609 | ||
|
|
52bc520335 |
Merge remote-tracking branch 'origin/main' into claude/frame-timing-harness
# Conflicts: # CHANGELOG.md # docs/SCROLL_PERFORMANCE.md |
||
|
|
1afb2383cd |
feat(vegas): prefetch_gate on by default, after an A/B/C soak on hdpi
Two runs per arm, about 81,000 frames each, order A B C C B A: A step 1 as is 0.90% late, 20.1 per 10k two+ refreshes late B switch_interval_ms 1 0.78% late, 15.8 per 10k C prefetch_gate 0.60% late, 2.5 per 10k No freezes in any arm, and the next group was ready at every strip extension, so parking the prefetch thread (3-6s per 8-minute run) cost nothing visible. The gate is now on unless vegas_scroll.prefetch_gate is false; on a stock binding it cannot work and says so at INFO once a run rather than warning on every install. switch_interval_ms stays off. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4e61d7248a |
refactor(web): one error-response path for api_v3 (#624)
* refactor(web): answer unhandled api_v3 errors from one blueprint handler
Fifty-three api_v3 routes ended in a copy of the same catch-all: log the
traceback, return {status, "An error occurred; see logs for details",
details: describe_exception(e)} with a 500. They are replaced by one
errorhandler on the api_v3 blueprint that returns exactly that body.
It lives on the blueprint rather than falling through to app.py's global
handler because the two answers differ: the global one adds
error_code: UNKNOWN_ERROR, and api_client.js sends a body with an
error_code to the error modal and one without to a plain toast. A
blueprint handler also gives tests that mount api_v3 on a bare Flask app
the same answer the real app gives.
Only handlers that were byte-for-byte that shape were removed (matched on
the AST, and each rewritten function re-parsed and compared). Handlers
with their own message, extra keys, operation-history records or cleanup
stay, as does execute_plugin_action's step-1 handler, which sits inside
an `except subprocess.TimeoutExpired` arm that would otherwise turn a
plugin's timeout into a 408.
HTTPExceptions raised inside a route go back as themselves in the global
handler's 4xx shape. Where a removed catch-all used to swallow one (only
delete_plugin_asset's non-silent get_json() is reachable), a malformed
request now gets its 415/400 instead of a 500.
Most of the diff is re-indentation from unwrapping the try blocks;
`git diff -w` shows the real change.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin action errors name the real failure, not UnboundLocalError
execute_plugin_action bound a local `logger` in its JSON-parsing arm,
which made `logger` local to the whole function. Every other
`logger.error` in it then raised UnboundLocalError, so a failing OAuth
step-1 script was reported as "UnboundLocalError: cannot access local
variable 'logger'" -- from the step-1 handler, and before the previous
commit from the route's outer catch-all too. Use the module logger.
Found by comparing every api_v3 route's forced-failure response before
and after the catch-all consolidation.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop the error category and exception-name code guessing
WebInterfaceError derived an ErrorCategory from every error code and put
it in each structured error body as `error_category`. Nothing reads it:
not the web UI (static/ and templates/), not the tests beyond the ones
pinning the mapping itself, and not any plugin in ledmatrix-plugins. The
enum, the inference table and the JSON key go.
from_exception() could also guess an error code from the exception's
class name ("Config" -> CONFIG_LOAD_FAILED, and so on). Every caller
passes a code, so the guess never ran; error_code is now required.
suggested_fixes stays: the error dialog in static/v3/js/utils/
error_handler.js lists them.
The REST reference loses error_category and says what an unanticipated
exception in an /api/v3 route answers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one call for the from_exception error responses
Nine plugin routes built a structured error by hand:
from src.web_interface.errors import WebInterfaceError
error = WebInterfaceError.from_exception(e, ErrorCode.X)
return error_response(error.error_code, error.message,
details=error.details, context=error.context,
status_code=500)
That is now exception_error_response(e, ErrorCode.X) in api_helpers, so
error_response() is the only structured-error entry point the routes
use. The three operation-history routes never passed the context, and
with_context=False keeps their bodies exactly as they were; a test
compares the helper against the hand-written pair for both forms.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): one api_v3 error-response path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
ece416c4e5 |
refactor(plugins): one plugin-directory resolver (#623)
* refactor(plugins): one resolver for plugin id -> directory Five places mapped a plugin id to its directory, each with its own rules and each re-reading manifests per lookup: PluginManager discovery and get_plugin_directory, PluginLoader.find_plugin_directory, PluginStoreManager._find_plugin_path / list_installed_plugins, and state_reconciliation.disk_plugin_ids. They disagreed on backup dirs, on whether the manifest id or the directory name is the id, on duplicate ids and on path safety. src/plugin_system/plugin_dirs.py now holds the rules once: PluginDirectoryIndex scans one directory and reads each manifest once; resolve_plugin_dir() searches directories in order. What legitimately differs per caller is an explicit argument: search dirs (discovery and the loader: configured dir only; the store: configured then sibling plugins/), ledmatrix- prefix (not for the store), case folding (loader only), manifest pass (not for get_plugin_directory, whose discovery map already holds it). Behaviour changes, all for layouts installs do not produce: - a directory whose manifest declares the id beats one merely named for it (discovery already worked this way; the loader and store now agree) - the store searches the configured dir completely before plugins/ - backup and hidden dirs are skipped everywhere (the loader's case and manifest scans and list_installed_plugins used to return them) - duplicate ids resolve deterministically (exact name, then ledmatrix-<id>, then by name) with a one-time warning; discovery no longer lists the id twice - disk_plugin_ids / list_installed_plugins report manifest ids, falling back to the directory name; auto-update looks the directory up - ids that are not one plain path segment resolve to nothing in every caller (the loader used to truncate them, the store to join them) The .standalone-backup- marker is one constant, BACKUP_MARKER, used by store_manager's rename-aside names and every lookup. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): one plugin-directory resolver Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
13bbb537f3 |
refactor(web): one logging setup and one TTL cache for the web process (#621)
* refactor(web): use src.logging_config in the web process; routine requests to DEBUG
The web interface had its own logging setup (web_interface/logging_config.py)
that replaced the root handlers with a plain stdout formatter. The web
service's journal lines therefore never carried a syslog priority, so
`journalctl -p err -u ledmatrix-web` returned nothing while errors were
logged, and the line shape differed from the display's (the log viewer's
prefix stripping only matched the display format). It also ran after the
module-level managers were built, so their INFO lines at import (including
"Re-removed N uninstalled plugin(s)") were dropped.
app.py now calls src.logging_config.setup_logging() first thing, the same as
run.py: journald priorities under systemd, LEDMATRIX_DEBUG honoured,
LEDMATRIX_JSON_LOGGING still selects JSON.
Per-request logging moves to web_interface/request_logging.py. Every request
used to be logged at INFO, so the UI's polling filled the journal
("GET /api/v3/errors/summary - 200" every minute per tab). Now a successful
GET/HEAD/OPTIONS is DEBUG, a successful write is INFO, 4xx WARNING, 5xx
ERROR. Durations use perf_counter and print to 0.1ms.
The duplicate module is deleted; nothing else imported it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one thread-safe TTL cache for the web process
web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.
Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
counted and get_cached defaulted to 60s. An entry now expires after the TTL
it was stored with; a reader's ttl_seconds can only shorten that. Both
current callers pass the same value on both sides (fonts_catalog 300s,
system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
same expired key could raise KeyError (reproduced), which the endpoints
turned into a 500.
app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.
Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): web logging and TTL cache
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): only ask systemctl about known units
Codacy flagged the systemctl argv built from a variable. The unit now has
to be one of two literals, and anything else raises.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): response_time_ms reads the same clock request_logging stamps
request_logging now stamps request.start_time from perf_counter, but
success_response still subtracted it from time.time(), so metadata
reported ~1.8e12 ms. Found testing on ledpi.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
afe9001aed |
refactor(fonts): one BDF loader and one BDF rasterizer (#627)
* refactor(fonts): one BDF loader and one BDF rasterizer BDF faces were loaded three ways (FontManager._load_bdf_font, element_style._load_bdf, DisplayManager._load_fonts) and drawn by two copies of the same per-pixel loop (DisplayManager._draw_bdf_text and the plugin test harness's "replicated" copy), which golden images and check_plugin/dev_server previews rely on matching the panel. src/common/bdf_font.py now owns both: - load_bdf_face(path, size) -> (face, realised_px): native-strike fallback for sizes the file lacks, one bounded LRU cache keyed on path, size and mtime. FontManager, element_style and DisplayManager delegate to it; read_bdf_native_size moves here (the old names delegate). - draw_bdf_text(draw, text, x, y, face, color, clip): builds each glyph as a 1-bit mask and fills it with ImageDraw.bitmap instead of a draw.point per pixel. A blending Draw (RGB image, "RGBA" mode) keeps the point path so translucent colours still blend. Pixel-identical: 220,032 renders (every bundled BDF at native and off-strike sizes, 14 strings, 4 colours, clipped on every edge, through each old loader x rasterizer) match origin/main byte for byte. test/test_bdf_font.py keeps a lightweight version against a frozen copy of the old loop. DisplayManager._draw_bdf_text goes from 1.4-23 ms to about 0.1 ms per string (the old loop re-read FreeType's buffer as a Python list for every pixel). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(testing): harness calendar_font is sized like the panel's VisualTestDisplayManager built its 5x7 calendar_font / bdf_5x7_font as a bare freetype.Face. With no size set its ascender reads 0, so BDF text drawn with it landed 6px above where DisplayManager draws it -- entirely off the canvas at y=0 -- and get_font_height() returned 0. Golden images and check_plugin / dev_server previews showed text the panel does not. Load it through load_bdf_face at the panel's 7px, so it is the very face DisplayManager uses. Across the differential run this changes only the cases drawn with the harness's own calendar_font (968 of 220,032), which now match the panel's output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): one BDF face per thread The shared face cache now hands every loader (FontManager, element_style, DisplayManager, the harness) the same freetype.Face. FreeType does not allow two threads to use one face at once, since load_char rewrites its glyph slot, so key the cache by thread as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
abedc46104 |
refactor(sports): merge the sports_shared/sports_card twins that behave identically (#626)
* refactor(sports): wrap the sports_card twins that behave identically SportsCoreSharedMixin (switch mode, via each scoreboard's sports.py) and sports_card (scroll/Vegas mode, via game_renderer.py) carried the same helpers twice. test/test_sports_twins.py now calls every pair with the same inputs -- the eight scoreboards' harness fixture games in flat, flat+nested and nested-only shapes, plus edge cases (favourites by id and abbreviation, NRL's colliding abbreviations, missing and non-numeric scores, bad zones, out-of-range dates, shared font faces). Identical pairs become thin wrappers over the sports_card function: _card_option, _vs_text, _format_game_time, _coerce_rgb, _crisp_size (with the class's own tables), _unshare_element_fonts (with the class's own element map, via a new optional argument), and the colour/month/weekday/ font-grid tables (dicts copied, not aliased). _format_game_date shares the card's formatting body but keeps its own setting, weekday zone and month table; _schema_font_size shares the parser but keeps its per-class cache, because a reloaded plugin gets new classes and a shared path cache would stop it seeing an edited schema. _resolve_font_size agrees but keeps its body so it still dispatches through the overridable hooks. No behaviour change: old and new mixin/card agree on all 22,994 comparisons over the test corpus, and the pairs that do differ (favourite-result colours on nested payloads and by favourites source, the weekday's timezone, the element-name map, per-mode colours) are left alone and pinned in TestPinnedDivergence for an owner decision. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(sports): pin that an ambiguous NRL abbreviation tints in both modes NRL's resolver passes a shared abbreviation ("NEW") through with an error and its _is_favorite_game matches ids only, but both favourite-colour helpers match on abbreviation as well, so both display modes tint a Knights or Warriors result for a user who typed "NEW". The twins agree; neither consults the _favorite_key seam. Pinned so a fix is deliberate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
1fe7237799 |
refactor(install): generate the web sudoers rules in one place (#622)
* refactor(install): generate the web sudoers rules in one place /etc/sudoers.d/ledmatrix_web was written by two copies of the same allow-list: a heredoc in first_time_install.sh Step 10 and a block of echo lines in scripts/install/configure_web_sudo.sh. They drifted before (safe_pip_install.sh was granted by one only), and a test existed just to catch that. Both now call web_sudoers_rules() from the new scripts/install/lib_sudoers.sh and keep their own validate (visudo -c), install and confirm flows. - first_time_install.sh output is byte-for-byte unchanged, so a device re-running the installer gets "already up to date". If the library is missing, Step 10 keeps the installed file and carries on, the same way it handles rules that fail visudo (an empty file would pass visudo). - configure_web_sudo.sh now writes the installer's layout: same 18 rules, different comments and order. It still leaves out reboot, poweroff and journalctl when they are missing; the library does that for both. The drift test now pins the generator's grants, checks that neither installer writes rules of its own, and runs each installer's call line to check the argument order. Tests that read the rule text now read the library. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(install): detect the web service user in one function first_time_install.sh pasted the same WEB_SERVICE_USER detection block three times (Step 3.1's fallback, the plugin-repos setup and Step 11). The copies were identical apart from comments; they now call detect_web_service_user(), whose body is that block unchanged. Behaviour is the same: the function sets the same global and always returns 0, as the inline if-chain did. Checked on Linux against all three original copies across 13 layouts (installed unit with and without User=, the repo as shipped, each grep branch, template placeholders). The comment notes that the install_web_service.sh / install_service.sh greps no longer match anything, so until Step 8 installs the unit the result is "root". That behaviour is left as it was. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
ddf5f085a5 |
perf(cache): tell a stale record from its header instead of parsing it (#633)
The sports plugins cache whole season schedules: 53MB for MLB, 18MB for NHL, 17MB for NCAA baseball. On a Pi 4, orjson.loads of the MLB file takes ~1.8s with the GIL held, and every thread in the display service waits -- the stall watchdog caught the render thread frozen 0.5-1.3s with the interpreter itself blocked, right on these reads. When a season record expired, DiskCache.get paid that whole parse only to find the timestamp too old and throw the result away. CacheManager.set now writes timestamp and ttl ahead of the data, and DiskCache.get reads them from the first 256 bytes of the file, applying the same rule as before (a per-entry ttl wins over max_age; no limit means never stale). A record that is stale is refused without being parsed. Files in the old layout, and records from other writers, don't match the header and are parsed in full as before. Also: ESPN responses in the background data service and espn_dates are parsed with orjson when it is installed (src/common/json_body.py). The stdlib parser behind response.json() takes 3.1s on the MLB season against orjson's 1.8s, both with the GIL held. espn_dates imports it with a fallback, since plugins bundle copies of that module for older cores. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
5baf983fe0 |
docs(scroll): explain the tear across the middle on fast scrolls (#620)
* docs(scroll): explain the tear across the middle on fast scrolls A 1:32-multiplexed 64-row panel lights row 31 almost a whole refresh after row 32, so fast scrolls show a sideways offset at mid-height of about speed x refresh period. Documents the cause, how to read the real refresh rate (show_refresh_rate prints with a carriage return), what was measured on a single-chain 2x128x64 Pi 4 (pwm_bits, gpio_slowdown and an uncapped refresh barely help; gpio_slowdown 2 glitches), and the fix that does help: fewer pixels per output via parallel chains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(scroll): limit the 1:32 row-pair explanation to panels that scan that way Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
82f3a3a3e4 |
fix(redaction): make credential redaction linear, not quadratic (#631)
* fix(redaction): make URL-userinfo redaction linear, not quadratic _REDACT_URL_USERINFO could start a match at every letter of a run of scheme characters, and each attempt read to the end of the run looking for `://`. On a long unbroken run of letters or digits (a hex digest, an ID, part of a response body) that is quadratic: 1.6s for 20k characters. The display service redacts every message, stack trace and context value it publishes in the error snapshot, holding the aggregator lock, and re.sub holds the GIL for the whole call, so one such exception stalled every thread, render loop included (~0.5s measured for 20k chars of hex). It also made test_snapshot_stays_small the slowest test in the suite by far: 142s of a 383s run, 139s of it in this one regex. A match may now only start where a run of scheme characters starts (negative lookbehind). Leading digits and `+.-` are captured in group 1 so the substitution restores them, and the scheme still has to start with a letter, so what gets redacted is unchanged: old and new output were identical on 300k fuzzed inputs. 20k chars now take ~0.5ms, 200k ~6ms, and test_snapshot_stays_small takes 0.8s. test/test_redaction.py pins the exact output for schemes that begin after digits or `+.-`, and bounds 50k-character runs at 1s; against the old pattern those timing tests fail at 3-11s each. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK * fix(redaction): make Authorization-header redaction linear too _REDACT_AUTH_HEADER matched the value's opening as `\s*["\']?\s*`: two `\s*` separated only by an optional quote. With no quote, a whitespace run could be split between them in every possible way, and when no credential followed (end of text, or `,` `"` `<` ...) the engine tried them all before giving up: quadratic, 8s for `authorization:` and 20k spaces, 17s with `Proxy-Authorization:` (tried again at the inner `authorization`). Same stall as the URL pattern: re.sub holds the GIL, and the display service redacts everything it publishes. The quote and the whitespace after it are now one optional unit, `\s*(?:["\']\s*)?`, which matches the same strings with only one way to split them. Output is identical to the old pattern on 300k fuzzed inputs; 20k spaces now take ~1.6ms. A scan of all three redaction patterns over prefix/run/suffix shapes finds none left that scales superlinearly. test/test_redaction.py pins exact output for quoted, tabbed, multi-line and credential-less headers, and bounds header + 20k whitespace at 1s; against the previous pattern those fail at 8-17s each. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
c1ce0b7b04 |
fix(web): two api_v3 paths called names that no longer exist (#625)
The Pixlet editor stop route restarts the display after a SIGKILL with _run_systemctl_command, which starlark.py never imported (since #554). The Starlark device-location resolver fell back to _ensure_cache_manager, which #609 deleted; the resolver already accepts no cache manager. Both raised NameError on the rare path that reaches them. pyflakes finds no other undefined names in src/ or web_interface/. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
14c38a3189 |
fix(perf): count a stall even when the scroll state went missing across it
On hdpi the stall watchdog logged a 1.9s stall during the hourly sports refresh that the soak report never had: its worst gap was 655ms. The frame that ended the stall was recorded as static, so its interval was dropped. "Scrolling" is DisplayManager's scroll state at the moment a frame is presented, and it goes missing mid-scroll: it expires after 2s without activity, and any thread can clear it. Plugins call set_scrolling_state(False) from their own display() (news, stocks, the odds ticker's fallback), and Vegas captures some of those on the render thread between two of its own frames. Vegas sets the state again only after its next frame, so that frame is recorded as static -- along with the capture or stall it followed. One static frame between two scrolling frames, with the scroll resuming within RESUME_SECONDS (1s), is now a frame of the scroll and both of its intervals count, the first at the scroll's own hold (clearing the state drops the hold to 1 too). Two static frames in a row still end the scroll. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
de54fc879a |
docs(offscreen): the two GIL experiments, and what each can and cannot cover
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b67818d5c5 |
feat(perf): LEDMATRIX_STALL_WATCHDOG_MS lowers the stall watchdog's threshold
250ms catches freezes; the hitches left on hdpi are frames 2-5 refreshes late, which look like the render thread waiting for the GIL. At 30ms the watchdog dumps those too, naming what the other threads were running when the frame missed. It polls at a third of the threshold so a stall one poll long is still seen, which costs some GIL time of its own: a diagnostic setting, not one to soak with. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4f30bc62ca |
experiment(vegas): vegas_scroll.prefetch_gate runs the prefetch only while the render thread waits on vsync
The render thread spends most of each refresh in SwapOnVSync with the GIL released, then needs it back the moment the swap returns. With plugin rendering on the prefetch thread, it often has to wait for it -- behind bytecode for up to the switch interval, behind a GIL-holding C call for as long as that takes -- and hdpi's late frames of 2-5 refreshes went up. src/common/render_gate.py opens a window around each swap, up to just before the refresh the swap will return on, and a profile hook on the prefetch thread parks it outside that window. It is never parked holding a lock the render thread also takes (the Vegas buffer, cache and state locks, logging, threading, importlib, the cache), never when no frame has been swapped for 50ms, and never for more than 50ms at a time. Off by default and ignored on a binding that keeps the GIL in SwapOnVSync, where the window would never let the prefetch run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
55a2760892 |
experiment(vegas): vegas_scroll.switch_interval_ms shortens the GIL switch interval during a run
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f79618d4f7 |
refactor(bench): grade render_bench with the shared frame-timing recorder
render_bench.py (from the parallel perf/render-bench work) had its own grading module, frame_pacing, with its own definition of a missed frame and its own refresh estimate. The soak already had both in frame_timing, so the two could have drifted apart on what "late" means. The bench now gives the display manager a fresh FrameTimingRecorder, drains it synchronously at the start and end of the graded run, and prints frame_soak's report with frame_soak's verdict. Its workload is unchanged: the synthetic strip, --busy load, the shared speed resolver, the per-frame scrolling announcement. frame_pacing, its tests and its src.common exports are removed; measure_refresh_hz moves to frame_timing, where scroll_speeds.py now finds it. Two ideas from frame_pacing carry over. The bench seeds the recorder with the idle refresh it measures, so a loop that free-runs (the 827fps bug the first bench caught) shows as early frames and one stuck at half rate as late frames, where an estimate taken from their own intervals finds both self-consistent. And the soak, which has no idle measurement, now calls a run NOT LOCKED when its refresh estimate beats the configured cap. The report also gives the rate held while rendering. Docs: the bench becomes "Without the service" under "Soaking a rig", keeping its hdpi numbers and the idle-vs-rendering refresh finding. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d56ec2ab3a |
feat(bench): measure a rig against the refresh it actually holds
There was no way to answer "does this hardware present every frame on time?" other than watching the panel. `scripts/render_bench.py` drives the production path -- a real DisplayManager and ScrollHelper, configured through the same `scroll_config` resolver every ticker uses -- and grades the run with a new `src.common.frame_pacing`, exiting non-zero when more than 0.1% of frames slipped a refresh. Exit 2 when the run could not be set up at all, so a rig that was never measured cannot pass by accident. A missed frame is defined exactly: an interval that rounds up to at least one more refresh than its frame hold asked for. The half-refresh rounding boundary keeps a frame that ran 1ms long on a 10ms refresh out of the count, because it still presented on the refresh it was meant to. The verdict that matters more is NOT LOCKED. A loop that never blocked on vsync reports a perfect zero misses while presenting nothing -- 8ms frames on a 100Hz panel all land in the one-refresh bucket while running 25% too fast -- so the report also checks the typical frame is not shorter than the panel could physically present. That is what caught the first version of this benchmark announcing its scrolling state once instead of per frame: the state expires on an inactivity threshold, the dirty-tracking skip then fires mid-scroll, and the loop free-ran at 827fps. And the refresh is read back out of the frames rather than taken from an idle measurement. Driving the matrix is bit-banging on the same machine, so pushing frames slows the refresh: a Pi 4 on 512x64 measures 100.4Hz idle and holds 96.3Hz while scrolling. Both are real, and grading against the idle figure reports a locked loop as 4% slow -- or, once the gap passes half a refresh, as missing every frame. The gap between the two is itself worth watching: a rise in it is a render-cost regression even when nothing is missed. Measured on hdpi (Pi 4, 512x64, pwm_bits 8), two minutes each: plain 95.44 fps, 8 missed of 11,449 (0.070%) PASS --busy 2 95.41 fps, 3 missed of 11,445 (0.026%) PASS Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 |
||
|
|
8163104581 |
style(offscreen): lint fixes for the new code (Codacy)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c883a2fd1e |
feat(perf): a stall watchdog that logs what the render thread is waiting on
The recorder counts freezes; it cannot say why. hdpi showed 1-2s freezes in both the #628 and offscreen builds, one lining up with hockey's 2s NHL fetch on the update thread, and nothing in the logs explained it. StallWatchdog polls every 50ms from its own thread. When a scroll's last frame is more than 250ms old (and a scroll is still running, so the end of a scroll is not a stall), it logs the stack of the thread that presented that frame and the top of every other thread's, then the stall's length when frames resume. It also measures how late its own wake-up was: if it was held up as long as the render thread, the whole interpreter was blocked (C code holding the GIL), not one thread on a lock. One dump per 30s at most; LEDMATRIX_STALL_WATCHDOG=0 disables it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0f68fbcfdc |
docs(offscreen): step 1 status and first hdpi soak
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
13b5264d11 |
feat(vegas): render every plugin's ticker content off the render thread
The plugin-facing canvas (DisplayManager.image, draw, matrix) was one shared object, so any plugin whose Vegas content needed it -- display capture, scroll-content generation, narrowed rendering -- was deferred to the render thread and fetched there one at a time. On hdpi that is most plugins, and each fetch stalled the scroll: news ~320ms, hockey ~660ms, in bursts whenever the strip extended. DisplayManager.offscreen() gives the calling thread a canvas of its own. image, draw and matrix are now properties that resolve to the thread's surface while it is inside the block and to the shared canvas otherwise, so the ~100 existing uses become thread-correct unchanged. Inside, update_display(), the hardware half of clear(), and set_scrolling_state()/ set_frame_hold() are inert, so a plugin drawn for Vegas can neither reach the panel nor re-pace the live scroll. render_size() is rebuilt on it. capture_mode() now restores the previous state instead of clearing it, so it cannot end suppression inside an offscreen block. The adapter draws every path on its own canvas (_isolated_canvas) and drops the copy-and-restore of the shared image, which from a background thread would have written a stale frame back over the render loop's. Background fetches take the plugin's update/display lock, waiting up to 2s for a running update() and skipping the plugin that round otherwise; Vegas never took that lock, so render-thread captures already raced update(). A background fetch that comes back empty is no longer queued for the render thread. vegas_scroll.offscreen_prefetch (default true) restores the old deferred path when false. See docs/OFFSCREEN_RENDERING.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
430e2312f8 |
docs(offscreen): redraw on real updates with a 10s floor; sync by operation replay
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
1af5d5fe45 |
docs(offscreen): keep live content fresh: refresh at the gate, replace ahead, patch on screen
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
5edc9195ca |
docs: propose per-thread offscreen rendering for Vegas content
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
ac841f4583 |
docs(perf): hdpi soak results, main vs #628
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
98728d3b81 |
fix(perf): count 1-2s stalls, flag early frames, and keep the refresh estimate honest
Three gaps found by the first hdpi soaks: - Intervals of 1s or more between two scrolling frames were dropped as "gaps between scrolls". But the scrolling state lapses only after 2s, so every 1-2s stall inside a scroll vanished from the report. Those are now freezes (the gap bound is a 5s sanity limit), with a breakdown by length. - A frame a whole refresh early means the swap did not wait for the panel. Those are counted, and a soak with more than the threshold of them fails as NOT LOCKED instead of reporting a flattering late rate. - The refresh estimate took the lowest window it had seen, so one window of non-blocking swaps halved it and made every early frame look on time. A window may now lower it by at most 20%. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3eb7a2e349 |
feat(perf): time every presented frame, and a soak script to judge a rig
Each scroller already logs its own stats line, but in different formats, per source, and Vegas logs a healthy window only at DEBUG. None of it answers the question a release has to answer on each rig: over a long run, how often did a moving frame reach the panel late? Every frame reaches the panel through DisplayManager.update_display, so it is timed there once, whoever drew it: the blit (SetImage), the vsync wait, and the interval since the previous frame. The render thread only appends a tuple. A worker thread aggregates cumulative counters and histograms and rewrites /dev/shm/ledmatrix_frame_stats.json every 10s (RAM, so no SD wear). A frame due after `hold` refreshes that lands one or more refreshes later is "late": the panel repeated the previous frame, a visible hitch. Gaps of 250ms+ inside a scroll are "freezes" (recomposes, handovers, blocking calls), counted separately so one handover does not read as 40 missed refreshes. Static frames, the first frame of a scroll and gaps between scrolls are not timed. The refresh period is estimated from the frames themselves. scripts/frame_soak.py runs next to the service as any user, diffs two snapshots over a run (default 10 minutes), optionally keeps the web preview's viewer marker fresh, and exits non-zero above 0.1% late frames. It also reports whether the loaded rgbmatrix binding releases the GIL. Documented under "Soaking a rig" in docs/SCROLL_PERFORMANCE.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6c01c2b493 |
docs(scroll): Vegas no longer opts into sub-pixel blending
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0a2ce58026 |
fix(vegas): lock the scroll to the panel refresh; encode the preview off the render thread
Vegas advanced by elapsed time, blended neighbouring columns every frame, and paced itself with a sleep to target_fps. On hdpi (4x128x64 on one chain, a 120Hz cap the chain cannot reach, ~95-100Hz real) that ran at 73fps with target 90 and ~89fps with target 125: the sleep drifted against the refresh and missed a vsync every few frames, and the blend read as shimmer on the panel (and as "anti-aliased" text in the preview). smooth_scroll now means the crisp pacing the plugin tickers already use: a whole number of pixels per presented frame, each held for frame_hold refreshes, with SwapOnVSync as the clock. The speed is solved against the panel's measured refresh, timed from our own swaps once scrolling starts, because the configured limit is only a cap -- at "120Hz" 90px/s solves to 3px every 4 refreshes, at the real ~97Hz to 1px every refresh. The old blend stays available as sub_pixel_blend (default off). With the web preview open, the render thread also PNG-encoded the whole 512x64 frame five times a second, 12-14ms each -- longer than a refresh. Mid-scroll that encode now runs on a single-slot writer thread (Pillow releases the GIL while compressing); static frames still write inline. Measured on hdpi, 3-minute soak with the preview open: 3 of 17,280 frames held an extra refresh (0.02%), down from ~6-20%. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f3894916a9 |
feat(web): show which plugins use each font; warn before deleting one (#619)
* feat(web): show which plugins use each font, warn before deleting one
The Fonts tab lists font files from the web process's own scan, and the
plugins that register fonts run in the display process, so the tab had no
way to say whether a font was in use before deleting it.
The display service now publishes {catalog key: [plugin ids]} to the
shared cache (font_usage_snapshot, src/font_usage.py), built from the
loaded plugins' FontManager.register_manager_font() registrations. A
daemon thread checks every 10 s and writes only when the usage changed
(plus a daily refresh so cache cleanup cannot expire it); it never raises.
Families, aliases (press_start, four_by_six, ...) and paths are resolved
through FontManager's catalog to the file stem the Fonts tab keys rows by;
fonts outside assets/fonts are left out. Unloading a plugin drops its
registrations (new FontManager.forget_manager_fonts).
GET /api/v3/fonts/catalog merges used_by into each row per request (the
5-minute scan cache is copied, never edited): a list of plugin ids, or
null when the display service has not reported. The tab shows a Used by
column ("unknown" / "-" / ids, rendered as text) and deleting an in-use
font names the plugins in the confirmation, from a fresh read. The server
still refuses only system fonts. Catalog fetches bypass the browser's
5-second API cache, which otherwise served the pre-delete list.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix: call forget_manager_fonts through a hasattr check pylint can follow
getattr(..., None) then callable() is fine at runtime, but pylint's E1102
("not callable") can't see through it, and Codacy fails the check on it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|