mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
224847cebc23a0ac4e8eb18829c84d79cf59caaf
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
116abb0daa |
fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service BackgroundDataService always sent a season range first and, on a 400, fell back to chunks without recording the rejection, so every background season fetch spent a doomed request and live scoreboards learned nothing from it (or it from them). The worker now consults and sets the same 6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight to month/day chunks, and if every chunk fails the range is asked once for a real error without re-spending the chunks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep plugin asset and action routes inside their directories POST /plugins/assets/upload, GET /plugins/assets/list and POST /plugins/assets/delete joined the request's plugin_id onto assets/plugins unchecked, so '../../config' created, wrote, listed and deleted outside it. #561 guarded only the route that serves the files. All three now go through path_safety.resolve_under and answer 400 for anything but a plain name, and delete only unlinks a metadata path that resolves into that plugin's uploads directory. PluginManager.get_plugin_directory refuses ids that are not one plain path segment, so /plugins/action (which runs a manifest script from the returned directory) and every other caller get the guard; the action route also rejects such ids up front, covering its no-manager fallback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): report a no-op plugin update as already up to date update_plugin() returns True both for a real update and for "nothing to do" (a ZIP-installed monorepo plugin already at the registry version, a bundled plugin). With no git commit to compare, POST /plugins/update called every such success "updated successfully", so Check & Update All counted most official plugins as updated on every run. The route now reads what changed off the plugin itself (commit, else manifest version, else last_updated) and returns data.update_status (updated / up_to_date / local_only). The update-all toast is summarised by PluginInstallManager.summarizeUpdateResults from that status, falling back to the message for older servers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): scoreboard scroll speed no longer follows target_fps sports_scroll computed the crisp speed ladder against the global target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since frame-locked presentation (#545) the helper steps a fixed number of whole pixels per presented frame and the panel presents at its real refresh, so the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a 50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s. The ladder now uses the display manager's refresh_hz, then display.hardware.limit_refresh_rate_hz, then the default. target_fps is not consulted. Docstrings now say scroll_delay is ignored for pacing (no behaviour change there) and describe the fixed-step model. Tests: replace the tests that pinned target_fps as the ladder refresh and described time-based stepping; assert speed independence from target_fps (unit and end-to-end presented px/s against the real helper), that the fixed per-frame step is applied, and that scroll_delay does not change speed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): escape registry and upload values in plugin manager inline handlers The store, saved-repository and custom-registry buttons built onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone, so a custom registry entry whose id contained ' closed the attribute and added its own handler. One helper, jsStringAttr(), now HTML-escapes the JSON literal for every one of those handlers, and the store View button opens only http(s) repo links. The live window.updateImageList (plugins_manager.js loads last, so its copy wins over the file-upload widget's) wrote the uploaded file's original name, path and ids into markup raw; they are escaped now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note plugin asset, action and inline handler guards Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): let the root pip wrapper install web_interface/requirements.txt Update Code, the automatic update's health check and Install Base Requirements install web_interface/requirements.txt through safe_pip_install.sh, which only allowed the root requirements.txt. The first commit changing that file would fail its dependency install, and the automatic updater rolls back any update whose dependencies did not install -- on every device, for every newer commit. The wrapper now lists both core requirement files. Only their folders are resolved, so a requirements.txt symlinked out of the project is compared by its target and refused (previously the root file's own symlink target was what got allowed). The updater's file list is a named constant, and a test runs the real wrapper (pip stubbed) on every file Update Code and the rollback install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): do not retry plugin requests that got an HTTP answer PluginAPI.request wrapped everything that was not a structured error as NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a JSON error without error_code included. Check & Update All retries NETWORK_ERROR, so those updates were re-sent five more times with backoff, contrary to the #587 contract that an HTTP error response is the server's answer. NETWORK_ERROR now means only that fetch() rejected. Any HTTP response without an error_code, or with a body that is not JSON, is API_ERROR with the HTTP status attached. Tested against the shipped api_client.js. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): restart the stats window when an idle gap is dropped by size #582 dropped an idle gap from the frame stats two ways: the reset_scroll() sentinel, which also restarts the 5s window timer, and a size guard for scrollers that never call reset_scroll(), which did not. On that path the first real frame after the gap found the boundary overdue and logged a stats line for a one-frame window. Both paths now share one seeding helper. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): leave plugins alone when update_core's own rollback fails update_core returns rollback_failed directly when a partial pull or an update whose health check never started cannot be rolled back. run() only held plugins back for 'verifying', so those devices still got new plugin versions and a display restart on top of a core in an unknown state -- the opposite of what the health-check path does, and of the 3.4.0 changelog (plugins are left alone if the rollback fails). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(api): make the REST reference match the api_v3 package Every documented request body, query parameter and response shape was re-checked against the handlers in web_interface/blueprints/api_v3/. Fixes calls that failed as documented (repo_url, action_id/params, files/image_id, font_file+font_family, ?font=, cache key, auto_enable_ap_mode, plugin limit keys), removes the font-override endpoints dropped in #566, corrects response shapes (plugins/config, plugins/schema, health, metrics, operation history, github-status, fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds the 26 routes it omitted (backup, system auto-update/git, wifi radio, starlark editor, MQTT bridge, status endpoints, skins). Documents the merge semantics of partial JSON saves to /config/main and /plugins/config and the dim-schedule POST accepting GET's days shape, which land in the same change set. Replaces app.py line numbers and the removed api_v3.py path with file and function names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): remove the General-tab plugin system toggles that did nothing plugin_system.auto_discover, auto_load_enabled and development_mode had General-tab toggles whose help tips promised dormant plugins and verbose logging, but nothing reads them: every enabled plugin is discovered and loaded regardless. Remove the three toggles. The keys stay tolerated in stored configs. The save handler now stores a flag only when a client sends it; treating a missing key as an unchecked box would otherwise rewrite all three to false on every General-tab save, which still posts plugins_directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(scroll): remove dead code left by #523/#570 - Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read them since the numpy blend replaced the scipy path. - Drop ScrollHelper._last_integer_position and frame_time_target, which were written but never read. - Keep target_fps and set_target_fps() but document them as informational: nothing paces off them, yet ledmatrix-elections' test_scroll_pacing.py reads helper.target_fps back and third-party plugins may call the setter. - Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to throttle", set_sub_pixel_scrolling's "default: True", and set_frame_based_scrolling's claim that it steps. The plugins monorepo was grepped for every removed name; none is used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI FONT_MANAGER.md told plugins to read display_manager.font_manager, which does not exist, so a plugin following it failed to load with AttributeError. The shared FontManager lives on the PluginManager and BasePlugin._get_font_manager() returns it (with a fallback for harnesses). Also removes the Fonts-tab override workflow and element-override panels that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls /plugins/store/search does not exist (404) and the list endpoint reads query, not q. The registry guide's curl examples omitted the JSON Content-Type, so the handlers saw an empty body and answered 400. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(config): use the shared core-key list in the last three private copies StartupValidator warned "Plugin 'auto_update' is enabled but not found" on every display start with auto-update or a dim schedule on; the reserved plugin-id check missed auto_update, sync, location and the rest; and ConfigManager's (uncalled) orphan cleanup would have deleted display, schedule and auto_update. All three now read src/core_config_keys.py, which also gains CORE_SECRETS_KEYS for the github/youtube secrets sections. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): partial JSON saves to /config/main change only what they send A JSON body with one field reset every checkbox in the sections it touched: the MQTT bridge's brightness slider turned off disable_hardware_pulsing, inverse_colors, show_refresh_rate and use_short_date_format, and a timezone-only save turned off web-UI autostart and weekly auto-updates. Missing-means-unchecked now applies only to form posts: form-encoded bodies and the v3 forms, which mark themselves with a hidden __form_section input. Also on the config routes: - vegas_min/max_cycle_duration no longer match the generic *_duration rule, so they stop landing in display_durations and a blank one no longer rejects the whole Display save; - saving from the Raw JSON editor calls start_setup_if_needed like the General form, so enabling auto-update there finishes its setup; - the schedule and dim-schedule POSTs accept the per-day days.<day> shape their GETs return, as well as the flat form keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): install plugin dependencies from the configured plugins directory install_plugin_dependencies.sh scanned only plugins/, but the Plugin Store installs into plugin_system.plugins_directory (default plugin-repos), so the documented "Recommended" fix found 0 plugins on every store install. It now reads plugins_directory from config/config.json (relative to the project root or absolute, default plugin-repos) and also scans plugins/ for dev symlinks, installing a plugin reached through both only once. With set -e alone, `pip ... | tee` took tee's exit status, so a failed pip install was reported as success; set -o pipefail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: replace stale API names, line numbers and the api_v3.py path - ADVANCED_FEATURES: StreamManager methods that exist (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the real on-demand status envelope ({status, data: {state, service}}) - app.py:199 / :144 / :607-619 line citations and web_interface/blueprints/api_v3.py (now a package) replaced with file and function names in ADVANCED_FEATURES, CONFIG_DEBUGGING, PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE, PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README - CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use /config/raw/main to replace the file; describe where validation runs - TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints usage) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): verify the web interface that actually ships, on port 5000 verify_installation.sh failed every healthy install: it required the long-removed web_interface_v2.py and looked for a listener on port 5001, while the web interface binds 5000 (web_interface/start.py). It now checks the files ledmatrix-web.service runs (start_web_conditionally.py, web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the same 5001 port in its listen check, HTTP probe and printed URLs. Port matches are anchored so :50001 no longer counts as :5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(plugins): one display-size contract: display_manager.width/height CLAUDE.md (#580) says to read display_manager.width/height because matrix is None when hardware init fails; the development guide, the safety-harness doc and two DisplayManager docstrings still recommended matrix.width/height. The bundled starlark-apps plugin read matrix.width unguarded, so its magnify recommendation and frame scaling raised in fallback mode (e.g. after the Pi 5 hardware refusal). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(install): make install_service.sh --help print usage instead of installing install_service.sh parsed no arguments, so `sudo ./scripts/install/ install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md) rewrote ledmatrix.service, ledmatrix-web.service and both update-verify units and enabled/started them. It now handles -h/--help (usage, exit 0, no changes) and rejects any other argument with exit 2 before doing anything. Running it with no arguments, as first_time_install.sh does, is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scroll): describe the fixed-step model and document frame_hold Since #545 a crisp speed from scroll_config.configure() makes the helper advance a fixed whole-pixel step per presented frame with no clock, and the display manager's frame hold is part of the speed. The docs still described the removed wall-clock model: - scroll_config's module and configure() docstrings said speed is applied in time-based mode and that omitting the hold "falls back to fractional pixels"; omitting it actually runs the scroll frame_hold times too fast. - SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both modes, and read a 20 ms stats median as missed refreshes although that is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step, the hold-dependent healthy median, that target_fps plays no part, and that a hand-added scroll_pixels_per_second loses to a schema-default pair. - PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling) without frame_hold; it now documents the parameter (core 3.4.0) with a configure() + set_scrolling_state example. - update_scroll_position/set_scroll_speed and set_scrolling_state docstrings say the same. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark target_fps legacy; describe what Vegas scroll_delay does - General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after the sports_scroll fix nothing in core scrolling reads it. The field and its API validation stay so saved configs and plugins that read global_config['target_fps'] keep working. CONFIG_REFERENCE says the same. - Vegas frame_based_scrolling/scroll_delay were described as frame-count stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and still advances by elapsed time, so the applied speed is clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The config comments, render_pipeline comment and CONFIG_REFERENCE rows now say so. No behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(deps): describe how plugin dependencies are really installed The guides said the web service runs as root, that installs pick --user from os.geteuid(), and quoted a warning and a PluginManager._install_plugin_dependencies() method that don't exist. The web unit runs as the installing user; store installs go through install_requirements_file() and sudo safe_pip_install.sh (root), with a user-level fallback that says so, and load-time installs run in the display service's own (root) interpreter. Manual paths now use the configured plugins directory (plugin-repos/ by default) instead of plugins/, which store installs no longer use, and install_plugin_dependencies.sh is described as scanning that directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): count local changes one way for the preflight and the pull The automatic update's preflight ignored mode-only changes and anything whose status line contained plugins/ or plugin-repos/, then promised "Automatic updates will not stash your changes". perform_core_update used plain git status (modes count) and ignored only 'plugins/', then ran 'git stash push -- :!plugins', which nothing ever pops. So an edit to a bundled plugin under plugin-repos/, or the installer's chmods on tracked scripts, passed the preflight and was stashed away for good. - auto_update.local_changes() is the one predicate both use: core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/ excluded by leading folder rather than substring (a core file under web_interface/static/v3/js/plugins/ now counts). - Update Code's explicit stash leaves out both plugin folders; the pull's --autostash carries their edits and mode changes across and reapplies them. - The automatic updater calls perform_core_update(stash_local_changes= False), which refuses instead of stashing edits that appeared after the preflight; update_core reports that as 'blocked'. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): diagnostics follow the web autostart default and api_v3 package #556 made a missing web_display_autostart mean "start" (only an explicit false/off keeps the web interface down), but the diagnostics still said otherwise: diagnose_web_ui.sh reported a missing key as "defaults to false", diagnose_web_interface.sh said the web interface "will not start unless this is set to true" and recommended enabling it, and debug_web_manual.py printed False. Troubleshooting a down web UI pointed users at a non-cause. Both shell scripts now evaluate the setting with the launcher's own autostart_enabled() (inline fallback if it cannot be imported) and report on / off / not set (on) / unparseable config; debug_web_manual.py uses the same function. They also check web_interface/blueprints/api_v3/ __init__.py: api_v3.py became a package in #553, so every healthy checkout was reported as missing a file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(install): what install_service.sh installs; verify script port; no sudo for --help install_service.sh installs and starts ledmatrix, ledmatrix-web and the update-verify units, not only ledmatrix.service (systemd/README.md, README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help' as a harmless check; it now shows --help without sudo and warns what a real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh checks the web interface on port 5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note update-all, plugin system settings and script fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): size the preview after orientation and pixel mappers display_geometry.physical_size claimed to give DisplayManager's answer but only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured after the library's pixel mappers, so a Rotate:90 / orientation 90 chain previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for 128x64. Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown names leave it alone), and move the orientation composition here so DisplayManager and the preview share it. The module docstring no longer claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH. Audit finding F18. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): refuse settings the rgbmatrix library aborts on, on every board The library answers several settings with a NULL matrix or abort() rather than an error, so the display service crash-looped (Restart=on-failure) instead of reaching fallback mode: rows above 64, chain_length above 255 (uint8_t binding setter, documented as "no upper limit"), a misspelled hardware_mapping, and parallel 2-3 on a single-output mapping, reachable from the Display form on the default adafruit-hat(-pwm) mapping. #586 only guarded the Pi 5 subset. - src/matrix_support.py holds the rules for every board (Options::Validate ranges, binding integer types, mapping names and outputs from lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the API's numeric ranges. - DisplayManager checks them before building options and raises MatrixSettingsRefused, so a hand-edited config falls back with a logged, reported reason. Emulator mode only warns. - The config API refuses them with a 400 naming the setting; combinations are checked against stored values but reported only when the request sets a field involved. - The hardware status file gains "cause" (settings/library/forced). The fallback log and Display banner give the Pi 5 rebuild hint only for a library failure instead of rebuild + gpio_slowdown advice for every failure; one Pi 5 slowdown recommendation (1-3, start at 1). - The Display form offers classic/classic-pi1 and orientation 90/270 and renders any other stored mapping selected with a warning, so an unrelated save no longer rewrites them; the API accepts 90/270. Audit findings F03, F16, F19, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(display): library limits, template defaults and Pi 5 slowdown - rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs, classic/classic-pi1 mappings and orientation 90/270 documented. - Defaults are the config.template.json values: config migration adds missing keys from the template, so the listed "code defaults" never applied. - One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode, starting at 1. - Troubleshooting describes the refused-settings fallback, and CHANGELOG corrects the Unreleased "no upper limit" entry. Audit findings F19, F20, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py opens the panel with the service's options --measure and --demo built RGBMatrixOptions from a private copy of the display service's builder that had drifted: gpio_slowdown came from display.hardware (default 2) instead of display.runtime (default 3), and rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors, pixel_mapper_config and orientation were skipped, with different defaults (hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown was measured -- or garbled -- in a setup the service never drives. The option filling in DisplayManager._setup_matrix moves, unchanged, into DisplayManager.apply_matrix_options(options, config), which _setup_matrix calls and the script reuses (overriding only limit_refresh_rate_hz for --measure). The script now loads the whole config rather than the hardware block. Tests pin the script's options to the service's attribute for attribute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py recommends keys the resolver honours The ladder ended by telling users to set display_options.scroll_pixels_per_second. scroll_config ranks that key below the scroll_speed + scroll_delay pair, deliberately, and several plugin schemas default the pair into config, so the advised key was silently ignored (a schema-default 1/0.02 pair plus an advised 66 still resolved to 50 px/s). The advice is now the pair that selects the crisp speed exactly (pixels_per_frame every frame_hold/refresh seconds), explains that the pair outranks scroll_pixels_per_second, and gives the scoreboards' per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the printed pair over a schema-default pair and check it lands on the advertised speed and hold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula - SPORTS_UNIFICATION.md still presented honouring global target_fps as sports_scroll's added behaviour and its one user-visible gain; note that it was withdrawn because it had become a speed multiplier. - ADVANCED_FEATURES.md gave Vegas scrolling as (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s by elapsed time, through a 0.1-5 px per scroll_delay clamp when frame_based_scrolling is on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): scroll model fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(dev): link-github links plugins from the ledmatrix-plugins monorepo link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git, and those per-plugin repositories no longer exist: official plugins are directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the monorepo once into the dev directory, finds plugins/<name>, plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and links it under its manifest id. With an explicit repo URL it still links a single-repository plugin as before. dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a fork), plus plugins_repo and plugins_branch; github_pattern, which was documented but never read, is dropped and warned about. Ships dev_plugins.json.example and git-ignores dev_plugins.json, both of which the guide promised. Reading JSON falls back to python3 when jq is missing (get_plugin_id silently returned nothing without jq). update/status/list find the git checkout above a monorepo plugin directory (its .git is not in the plugin dir), and update pulls a shared checkout once. status no longer exits 1 when nothing is broken. Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration, workflow, store integration, hello-world link, submission), and the nonexistent scripts/git-hooks/pre-push-plugin-version and scripts/bump_plugin_version.py replaced with the real rule: bump the manifest version and run update_registry.py. scripts/dev/README.md and CLAUDE.md updated to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scripts): monorepo workspace layout; fix_perms and install READMEs MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin; setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the workspace file opens LEDMatrix plus ../ledmatrix-plugins. scripts/fix_perms/README.md listed cache directories fix_cache_permissions.sh never touches and a 'ledmatrix' service user that doesn't exist (also in scripts/install/README.md); adds safe_pip_install.sh. install/README: install_service.sh installs the web and update-verify units too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): keep the rollback's pip retries inside the unit time limit The health check reinstalled the previous requirements by trying the next bash path after any failure, including a 600 s pip timeout. Two files, two paths: up to 40 minutes of pip alone, while systemd stops ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the rollback half-way and leaving the update 'verifying' until the web UI calls it lost. - Like permission_utils.install_requirements_file, only a sudo refusal moves on to the next bash; a pip that ran and failed or timed out is not repeated. The refusal wording is one list (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only verifier and pinned equal by a test. - All reinstalls in one rollback share a 600 s budget. - WORST_CASE_SECONDS adds up every timeout on the longest path (27.5 min); a test holds it under the unit's TimeoutStartSec and that under the web UI's VERIFY_LOST_SECONDS. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools Plugin config was prepared differently depending on how it arrived: - JSON POST /plugins/config built a partial body on schema defaults, so {"enabled": true} reset every other setting of the plugin. It now merges onto the stored section first, as the form path already did. - Legacy-boolean normalization (#588) ran only at load: GET /plugins/config returned the raw boolean, posting it back failed validation, and hot reload handed plugins the raw section (a legacy dynamic_duration: true came back as a boolean). schema_manager.prepare_plugin_config (normalize, then defaults) is now used by PluginManager.load_plugin, both save paths, GET, the save notifications and DisplayController's hot-reload callback. - The JSON save's filter kept only enabled/display_duration/live_priority and dropped a submitted skin, skin_options or vegas_* tuning key. There is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES, used by validation and by the save filter; PluginManager's CORE_OWNED_CONFIG_KEYS is its vegas subset. - Plugin sections posted to /config/main were stored verbatim, including values /plugins/config rejects. They now go through the same preparation (_prepare_plugin_config_for_save, extracted from save_plugin_config), and a failing section rejects the whole save before anything is written. - dev_server read only top-level defaults and let a schema enabled:false win; build_full_config shallow-merged overrides, dropping sibling defaults; the harness extracted defaults differently from the device. loading.build_config now uses the device's extraction and preparation, and dev_server, check_plugin, render_plugin and the harness all use it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt-bridge): brightness changes apply live and touch nothing else The display service's hot reload applies a saved brightness within a few seconds, and /config/main no longer resets other display settings on a brightness-only JSON body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): automatic update hardening Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI It described web_interface_v2.py and index_v2.html (both gone), client-side form generation, one POST per field with {key, value}, and 'no nested objects'. The v3 UI renders plugin forms server-side from the schema (pages_v3 partial + plugin_config.html macros, nested sections and x-widgets), posts the whole form once, and save_plugin_config() merges onto the stored section, validates, splits x-secret fields and notifies the plugin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt): brightness saves apply via hot reload and leave other settings alone The bridge README said brightness is applied on the display's next restart; the display controller's config hot reload applies it within seconds. It also now states that the bridge's partial JSON save changes only brightness (the /config/main merge fix in this change set). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): don't log pip's output from the health check's reinstall pip can echo a private index URL with embedded credentials; permission_utils redacts it, the stdlib-only verifier cannot, so it logs the exit code only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark the plugin_system toggles as unused legacy keys auto_discover, auto_load_enabled and development_mode are read by nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and the REST reference listed them as live settings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): docs and developer tools group Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): legacy plugin-system toggles no longer count as a General save auto_discover, auto_load_enabled and development_mode have left the General form, so a post carrying only one of them is not a general-settings save and must not treat web_display_autostart and auto_update as unchecked. The plugin_system block itself is left as on main for the branch that reworks it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): config-save and plugin-config preparation fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(claude): re-check matrix_support.py rules when the library submodule is bumped Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address Codacy findings on the core audit PR - plugin_manager.prepare_plugin_config: when the fallback legacy-boolean pass also fails, log a warning instead of a bare except/pass. - api_client.js: request() refuses any endpoint that is not a plain path under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace, control characters) with INVALID_ENDPOINT before calling fetch(), and plugin ids are URL-encoded wherever they are put into a URL (also in the app-shell batch load). - test_update_all.js: pins both against the shipped client. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): check endpoint control characters without a control-character regex Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags the \x00-\x1f range in checkEndpoint's regex. Test the char codes instead; the endpoints refused are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(auto-update): make the seed script executable on disk, not only in the index On Linux Repo.publish() commits with -a, which recorded scripts/run.sh as 100644 upstream because the seed file was never chmod +x. The pull then brought in the same mode the installer chmod had made locally, so installer_chmod saw no mode change left to check. The updater was fine: with the upstream commit at 100755 the --autostash carries the device's chmod across. Verified under Linux (WSL, git 2.43): the old helper fails exactly as CI did, the fixed one passes all 63 tests in the file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d01da3bd9f |
fix(scroll): stop timing the idle gap between scrolls as a frame (#582)
ScrollHelper.last_frame_time was set once in __init__ and thereafter only
at the end of log_frame_rate(). Nothing re-armed it when a scroll began, so
the first frame of every scroll was timed against the last frame of the
*previous* one and the whole idle period between them was recorded as a
single frame.
Measured over 3 hours on a 256x64 Pi 4, that produced 31 windows reading
Scroll frame stats - 0.0 fps over 1 frames | median 136776.02ms
p95 136776.02ms max 136776.02ms min 136776.02ms | stalls 0 (0.0%)
and -- worse, because it is not obviously wrong -- put the same gap in the
max field of otherwise healthy windows, where the worst values were 537s
and 604s. It also counted as one stall per scroll start: at ~500 frames to
a window that is ~0.2%, against measured stall rates of 0.07-0.16%. The
stall rate is the number used to judge whether a scroll change worked, and
it was the same order of magnitude as its own artefact.
The first frame of a scroll has no predecessor, so it has no frame time.
last_frame_time is now None until one is rendered, and reset_scroll() puts
it back -- the same treatment last_update_time already gets three lines
above, for the same reason. reset_scroll() alone is not enough, because the
scrollers actually emitting these lines never call it, so a sample at or
past the 5s log interval is dropped as well: nothing that renders a scroll
takes that long over one frame. Seeding also restarts the window timer, or
the boundary is already overdue when the second frame arrives and every
scroll opens by reporting a window of exactly one frame. A window whose
samples were all dropped now logs nothing rather than reporting the gap.
docs/SCROLL_PERFORMANCE.md documented the diagnostic in terms of a
"Frame time: N ms" line that
|
||
|
|
d12323e7f1 |
perf(scroll): pace frames to the panel — 44→100 fps, stalls 14% → 0.02% (#523)
* perf(scroll): pace frames to the panel, not to a fixed sleep Scrolling ran at 44-46 fps on a 2x128x64 chain and 14-17% of frames took 41-53ms, which reads as judder. Four independent causes, each measured on the hardware; details and the diagnostic recipe are in docs/SCROLL_PERFORMANCE.md. The high-FPS loop slept a flat 8ms after every render. display() has already blocked on the panel's vsync by then, so that sleep was added to a wait that had happened: ~4ms of render plus 8ms put each iteration at ~12ms against a 10ms refresh grid, so every swap missed a refresh and the loop settled at 50fps while asking for 125 -- with no headroom, so a further 14% of frames slipped again. It now sleeps only the remainder, with a 1ms floor so plugin threads still get the GIL. ScrollHelper stepped position on a wall clock at 1/scroll_delay steps per second. Plugins set scroll_delay to the frame period, so that comparison sat exactly on its own threshold: a frame arriving a hair early moved zero pixels and rendered an identical frame, dirty-tracking skipped the swap, it returned in ~2ms, and the beat repeated. No scroll_delay value tunes that out -- a shorter delay trades stalled frames for periodic double-steps. Both modes now accumulate elapsed time at the same configured speed, so position stays proportional to real time. Sub-pixel blending goes back to off by default. It renders a half-step by mixing two adjacent columns, which on a coarse panel showing pixel-font text alternates crisp and smeared frames and reads as shimmer -- visibly worse than integer stepping on the hardware. Vegas mode still opts in. disk_cache uses orjson when importable, falling back to the stdlib. Encoding a ~1MB record drops from 14.8ms to 5.4ms end-to-end, and that work holds the GIL while a marquee is on screen. display_manager also checksummed the whole framebuffer twice per frame (dirty tracking, then the preview snapshot); the snapshot now takes the checksum the caller already computed. New src/common/scroll_config.py resolves scroll settings in one place. Five ticker plugins each hand-rolled this and disagreed: odds-ticker ranked the deprecated scroll_pixels_per_second above the documented scroll_speed/delay pair, and because that key carries a schema default the documented settings were dead for every user (ChuckBuilds/ledmatrix-plugins#408), while ledmatrix-leaderboard read the same key only as a fallback. The resolver also warns when a speed will not advance a whole number of pixels per refresh, which is the property that actually determines whether a scroll looks smooth. scripts/build_rgbmatrix_nogil.sh rebuilds the rgbmatrix binding so it releases the GIL. Upstream declares SwapOnVSync without nogil, unlike SetPixel/Clear/Fill beside it, so the render thread held the GIL for the whole vsync wait and starved background threads into long uninterruptible bursts. The script patches, builds and self-verifies into a scratch tree; --install backs up the original and rolls back if the service does not come back healthy. Measured after: 100 fps locked, no stalls observed, render thread down from 51% to 19% of one core. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): keep the panel swap locked to vsync while scrolling Dirty tracking skipped SwapOnVSync for byte-identical frames. That is the right call for static content, but SwapOnVSync is also what paces the render loop, so skipping it skips the wait for the panel: a duplicate frame returns in ~8ms instead of ~10ms on a 100Hz panel, advances the strip only 0.8px instead of 1.0px, and so makes the next frame more likely to repeat as well. The effect sustains itself once it starts. Measured over 20 minutes on a 2x128x64 chain, both scrollers configured identically at 100 px/s: leaderboard 10ms x35, 11ms x3 (clean) odds-ticker 10ms x26, 8ms x7, 15ms x5 (~20% duplicates mid-scroll) The duplicates were not end-of-cycle idling -- 38% of fast frames fell within 90s of a scroll completion against 35% of normal frames, a null result. The trigger is per-frame work: odds does more of it, and more variably, so it is first to land a frame that advances less than a whole pixel. Pushing an identical frame costs one canvas copy. Falling out of vsync lock costs smooth motion. Static content is untouched, because is_currently_scrolling() expires on its own inactivity threshold -- covered by test_stale_scrolling_state_stops_forcing_pushes so a plugin that stops scrolling without saying so cannot pin the panel into always-push. Also de-flakes test_snapshot_still_written_on_skip, which asserted a strict mtime increase between two writes that can land in the same filesystem tick; it failed about two runs in three on Windows regardless of the code under test. The file is now backdated before the check. 156 tests pass on the Pi. Not yet confirmed by eye on the panel. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): report the frame-time tail, and stop the row-major blit Two problems, both found by looking at the panel rather than the metric. The frame-stats line reported ONE instantaneous frame every 5 seconds -- about 1 frame in 500 -- printed beside a 100-frame average. Both hide exactly the fault they are used to chase: a 2ms duplicate and a 21ms double-wait average to precisely 10ms, so a ticker stalling on half its frames still reports a healthy "Avg FPS: 100.0". That reading cost several rounds of chasing the wrong layer. The line now aggregates every frame since the last log and reports median, p95, max, min, and explicit stall and skip rates (past 1.5x the median missed a refresh; under half never reached the panel, because dirty tracking skipped the swap so the frame never waited on vsync). On the hardware this now reads: leaderboard 100.0 fps over 501 frames | median 10.00ms p95 10.05ms max 10.34ms | stalls 0 (0.0%) skips 0 (0.0%) The binding rebuild's blit patch becomes opt-in (RGB_PATCH_BLIT=1, default off). Reordering that loop to row-major changes what a torn frame looks like: column-major tearing shows as a vertical seam, row-major as a horizontal split between the panel's upper and lower halves. On a 1/32 scan panel that reads as a one-pixel fold across the middle of every panel, which is what was reported on hardware and what went away when the blit was reverted. All of the measured gain comes from the SwapOnVSync change, so the risky half is simply not worth taking; the header says so. Also fixes --install resolving its paths against $HOME, which is /root under sudo, so it looked in /root/rgbmatrix-nogil-build and died with "no built module found" on a machine where the build had just succeeded. It now resolves SUDO_USER's home. Both build paths are verified on the Pi: default yields one GIL-release site, RGB_PATCH_BLIT=1 yields two. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(scroll): let users pick a crisp speed for their own panel Whole-pixel motion was previously only available at multiples of the refresh rate -- 100, 200, 300 px/s on a 100Hz panel. 100 px/s crosses a 256px panel in 2.6s, which is brisk for reading, and everything slower had to blend (blur) or repeat frames unevenly (judder). There was no way to ask for 50 px/s and get clean motion. SwapOnVSync takes a framerate_fraction the display manager never passed. It holds each frame for N panel refreshes; the panel keeps refreshing at its full rate throughout, so holding costs nothing in flicker and only changes how often a NEW image is presented. That turns 50 px/s into one whole pixel every second refresh instead of half a pixel every refresh. The crisp speeds are therefore refresh_hz / hold * pixels_per_frame, and that ladder depends on the panel: a Pi Zero on a long chain has a different set of good speeds from a Pi 4 on a short one. crisp_ladder() enumerates them and solve_crisp() picks the best match for a requested speed. solve_crisp weights motion quality rather than picking the numerically nearest entry, which matters more than it sounds. Asked for 30 px/s, nearest-by-value answers 28.6 -- 2px jumps at 14fps -- over 33.3, which is single-pixel motion at 33fps and obviously better on the panel. The target is also clamped into the ladder's range first, because relative error saturates near 1.0 for a target far outside it and the quality penalty would otherwise answer "10000 px/s" with the slowest entry. configure() snaps to the ladder and applies the hold when given a display manager. Without one the hold silently cannot happen and motion falls back to fractional pixels, so it warns rather than failing quietly. set_frame_hold() resets to 1 when scrolling stops, so one plugin's pacing cannot leak into whatever is on screen next. scripts/scroll_speeds.py is the user-facing part: it prints the ladder for the configured rate, measures what the panel ACTUALLY manages (--measure, for hardware that cannot reach its configured limit), highlights the nearest option to a wanted speed, and demos one live. It never starts or stops the display service itself -- doing that inside a script stranded the panel twice today. Speeds below ~20 px/s remain stepped regardless. That is the pixel pitch, not a software limit. Also fixes the dirty-tracking test spy, which stubbed SwapOnVSync with a single-argument function and would have masked the new call as a failed push, and rewrites a configure() test that had started passing for the wrong reason: it asserted a judder warning, which snapping now prevents, and was matching the unrelated "hold could not be applied" warning instead. 183 tests pass on the Pi. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): tie the frame hold to the scroll, not the plugin The hold applied in configure() never reached the panel. Plugins share one display manager, and set_scrolling_state(False) -- fired whenever ANY other plugin finishes its scroll -- reset the hold to 1. A hold set once at plugin construction was therefore always gone by the time that plugin rendered. The symptom was a log line that lied. ledmatrix-stocks reported Scroll configured: 50.0 px/s (1px every 2 refreshes = 50.0 fps, smooth) while the panel measured 100.0 fps, median 10.00ms. Config, resolution and snapping were all correct; only the pacing silently was not applied. set_scrolling_state(is_scrolling, frame_hold=1) now carries it, so the hold lives exactly as long as the scroll that asked for it. configure() reports the value as ScrollSettings.frame_hold instead of applying it -- applying it behind the caller's back could never have been right on a shared display manager. Existing callers are unaffected; the default keeps one frame per refresh. Verified on hardware: stocks at 50 px/s now measures 50.0 fps over 251 frames | median 20.00ms p95 20.09ms | stalls 0 skips 0 20.00ms being exactly two refreshes, with the panel still refreshing at 100Hz underneath so flicker is unchanged. test_another_plugin_stopping_does_not_strand_a_hold pins the interaction that broke this. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll,cache): resolve CodeRabbit review on #523 Eight findings, all reproduced before fixing. scroll_config.configure() read the refresh rate *after* resolve() had already used it. resolve() fills in target_fps, pixels_per_frame and the judder warning from that rate, so on a 60Hz panel every one of them described 100Hz -- and with snap_to_crisp=False nothing downstream corrected it, so set_target_fps() paced the helper to 100 FPS. The rate is now settled first, and falls back to the global config rather than straight to the default. refresh_hz_from_config() used `(cfg.get("display") or {}).get(...)`, which raises AttributeError when either level is truthy but not a mapping -- out of a function whose whole contract is a rate or a default. The frame-stats line reported the upper-middle sample as the median and the 96th sorted sample as p95 of 100. Both are also thresholds (stalls at 1.5x the median, skips at 0.5x), so the counts were biased too. The arithmetic is now in frame_stats()/format_frame_stats(), testable without a clock. configure()'s docstring and docs/SCROLL_PERFORMANCE.md still said it applies the frame hold and warns when it cannot. It deliberately does neither since "tie the frame hold to the scroll, not the plugin"; a caller following the old text would omit set_scrolling_state() and slow snapped speeds would still present every refresh. disk_cache had no policy for non-finite floats: orjson writes null, the stdlib writes NaN/Infinity, and orjson then rejects those legacy files so DiskCache.get deleted them as corrupt. One behaviour on both paths now -- write null, keep legacy records readable. allow_nan=False detects the values; the replacement walk runs only when there is one, so the ordinary write path is byte-identical and pays nothing. build_rgbmatrix_nogil.sh picked the build artifact with a glob piped to `head -1`, which sorts cpython-311 ahead of cpython-313, so a stale .so staged in from the source tree was installed as core.so while the GIL check -- which reads the generated core.cpp, not the .so -- still passed. It now requires the current interpreter's exact ABI name and fails closed. Its systemctl calls were also unchecked under `set -uo pipefail`: a failed stop left the old service running, the following start succeeded as a no-op, and the health check reported SUCCESS for a binding that was never loaded. orjson floor raised to 3.11.6 for CVE-2025-67221 (unbounded recursion in dumps); it covers the project's Python 3.10-3.13 range. Adds test/test_cache_nonfinite_floats.py (14) plus regression tests in test_scroll_config.py and test_scroll_helper.py. 9 of the cache tests and 9 of the scroll_config tests fail against the pre-fix code. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * test(harness): keep the visual double's signature tied to production Moves set_scrolling_state's frame_hold into the test double here, where DisplayManager gains it, rather than in #534 where it arrived a PR early. CodeRabbit flagged the #534 version correctly: a double that accepts an argument production does not lets the call pass every harness run and raise TypeError on the panel, which is the one failure a safety harness exists to prevent. The drift has now gone both ways across two branches -- double behind production on this branch, double ahead of it on #534 -- so it is pinned instead of remembered. test_display_double_parity.py compares the two signatures and fails with the direction of the drift named. It reads the files with ast rather than importing them, because display_manager imports rgbmatrix at module scope and this check should hold on a laptop and in CI as well as on a Pi. Plugins begin passing frame_hold in ledmatrix-plugins#462, which is why production and the double both need it before that lands. Full suite: 3889 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9930bd33b1 |
test: add 306 new tests covering previously untested modules (#347)
* test: add 306 new tests covering previously untested modules Adds test coverage for six major untested areas: - src/base_classes/api_extractors.py — ESPN football, baseball, hockey, soccer extractors - src/base_classes/data_sources.py — ESPN, MLB, and soccer API data sources (HTTP mocked) - src/common/game_helper.py — game extraction, filtering, sorting, and summaries - src/common/utils.py — all utility functions (normalise, format, validate, parse) - src/common/scroll_helper.py — ScrollHelper init, create, update, visible portion, duration - src/background_data_service.py — cache hit/miss paths, retry, cancel, cleanup, singleton - src/vegas_mode/config.py — VegasModeConfig from_config, validate, update, ordering - src/logo_downloader.py — normalize_abbreviation, filename variations, directory helpers - src/plugin_system/health_monitor.py — HealthStatus determination, metrics, suggestions, lifecycle https://claude.ai/code/session_015792DiGo27JbgH5mk3KBjk * fix(tests): thread cleanup on assertion failure, reduce oversized image - test_health_monitor.py: wrap start_monitoring calls in try/finally so the background thread is always stopped even when an assertion fails - test_scroll_helper.py: reduce 50,000px test image to 5,000px to avoid unnecessary memory pressure on Raspberry Pi Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Chuck <chuck@example.com> |