mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
9d024f24efd1f891282b34ef61b3913238f96e3c
210
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9d024f24ef |
refactor(cache): remove the cache layer's duplicate cleanup and dead lookups (#613)
* refactor(cache): collapse CacheStrategy's all-60 defaults table and twin soccer branch get_sport_live_interval() without a config manager looked the sport up in a table where every value was 60, with 60 as the fallback; it now returns 60. get_data_type_from_key() had an `if 'soccer'` branch returning the same 'sports_live' as its else. test_cache_strategy_intervals pins the returned strategy for every data type x sport key x config-manager shape; it passes unchanged on the old code. A 2,544-entry dump of every CacheStrategy method over a wider grid is identical before and after. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(cache): drop CacheStrategy's `<sport>_scoreboard` config lookup get_sport_live_interval() and get_cache_strategy() read live/recent/ upcoming intervals from config[f"{sport}_scoreboard"]. Those sections belonged to the built-in scoreboards the plugin system replaced; plugin config is keyed by plugin id ("football-scoreboard"), so on a current config the lookup always fell through to the defaults (60 live, 1800 recent, 10800 upcoming), which are now returned directly. The one input where this differs: a config.json upgraded from the pre-plugin era that still carries e.g. an "nfl_scoreboard" section (no code removes them), queried with an explicit sport key. No caller in core or the plugin monorepo passes a sport key here -- get_with_auto_strategy only derives one for keys classed sports_live/live_scores, and its callers (odds managers, odds-ticker) use odds keys -- so the stale section was unreachable in practice. A dump of every CacheStrategy method over 2,544 inputs differs from the previous commit only in those 45 legacy-config entries; the test grid now includes that shape. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * perf(cache): list cache files without holding the memory-tier lock CacheManager.list_cache_files() held the in-memory cache's lock while it listed and stat'd the whole cache directory -- 8,864 files on a real rig -- so every get()/set() from the display loop and plugins waited out the scan. The lock never protected the disk: DiskCache writes and deletes under their own lock, and a file vanishing between listdir and stat was already handled (logged and skipped). The body is unchanged apart from the dedent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(cache): delegate memory-tier cleanup and stats to MemoryCache CacheManager._cleanup_memory_cache() was a line-for-line copy of MemoryCache.cleanup(), and get_memory_cache_stats() a copy of MemoryCache.get_stats(), both reaching into the component's private _cache/_timestamps/_lock through "backward compatibility" aliases bound in __init__. So the component's own cleanup and stats only ever ran in tests, and the aliases went stale whenever the component was swapped (test_cache_ttl_honoured does). Both now delegate, and the aliases are gone: nothing in core, the tests, or the ledmatrix-plugins monorepo reads them. Behaviour is the same. Compared line by line, the two cleanups differ only in the sort key's fallback (0 vs 0.0, which orders identically), range+bounds check vs slice for the eviction, and the logger name on the DEBUG summary line (src.cache_manager -> src.cache.memory_cache). A differential run over 20,000 random memory states (str/None/garbage/ future timestamps, orphan keys, sizes 0-12, forced and throttled runs) gives identical removed counts, resulting dicts and last-cleanup times; the same harness catches each of three seeded mutations of MemoryCache.cleanup. The throttle clock also moves with it: CacheManager kept its own copy of last-cleanup, the component's is used now, and they started equal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(background): inline the sport cache key and drop the unused request queue get_sport_cache_key() constructed a whole CacheManager -- ConfigManager, config parse, cache-dir probing with test-file writes -- to return f"{sport}_{date}". It now builds the key itself in the same format as CacheManager.generate_sport_cache_key() (UTC date, %Y%m%d); tests check the two agree for explicit dates and, with a frozen clock at 03:30 UTC, for the default date. Median per call on Windows: ~0.6 ms -> ~2 us (alternating runs); on a Pi the old path also wrote a probe file per call. request_queue was a PriorityQueue nothing ever put into: requests go straight to the executor, so `priority` never did anything. The queue is gone; the `priority` parameter and FetchRequest field stay (every monorepo scoreboard passes priority=) and are documented as ignored, and get_statistics() keeps reporting queue_size, now a literal 0 as it always was in practice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
269385c97c |
fix(config): write config.json through one durable atomic writer (#611)
save_config() opened config.json with 'w' and streamed json.dump into it, so a power cut or an unencodable value left the file truncated. save_config_atomic() renamed a temp file into place but never fsynced it, rewrote the unchanged secrets file on every save, and re-parsed every backup to rotate them. save_raw_file_content() had its own third copy. All of them, plus rollback and config creation from the template, now go through atomic_write_text(): temp file in the same directory, fsync, final mode set before the rename, rename (retried on Windows while a reader holds the file), directory fsync. A root save copies the previous owner onto the new file so a rename by the display service no longer hands config.json to root; the shared-group fix-up is unchanged. The mode is chosen from the file name, so a "secrets" directory in the install path no longer makes config.json 0640. The secrets file is rewritten only when its content changes, and backup rotation works from filenames alone. Backups keep their names (config/backups/config.json.backup.<version>, paired secrets backup) and the five newest are kept, as before. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
604f58ff07 |
feat: deprecate unused plugin-facing methods for removal in 3.7.0 (#610)
35 methods on CacheManager, DisplayManager, FontManager and PluginManager have no caller in core, the ledmatrix-plugins monorepo or the registry's third-party plugins, but plugins live elsewhere, so they stay for one release. src.deprecation.deprecated logs a warning (and emits a DeprecationWarning) the first time each is called in a process, naming the release that removes it. The list and replacements are in CHANGELOG and PLUGIN_API_REFERENCE's new Deprecated APIs section; a test pins the set. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a8b3e86775 |
refactor(web): delete dead routes, JS files and duplicate definitions (#609)
* refactor(web): drop validators nothing calls escape_html, validate_image_url, validate_font_awesome_class, validate_mime_type, validate_numeric_range, validate_string_length and sanitize_plugin_config had no callers outside their own tests. Only validate_file_upload (fonts upload) is imported by the web interface. dedup_unique_arrays is kept: its one caller in save_plugin_config was removed by the unrelated sync PR (#330), which looks accidental. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(api): remove the music-auth and of-the-day JSON routes POST /plugins/authenticate/spotify and /plugins/authenticate/ytm had no caller but their tests: the music plugin authenticates through its web_ui_actions (authenticate_spotify.py / authenticate_ytm.py) via /plugins/action. POST /plugins/of-the-day/json/upload and /json/delete looked the plugin up by the id ledmatrix-of-the-day (its manifest id is of-the-day), were reachable only from a file_type "json" upload field that no schema declares, and put the plugin directory on sys.path per request to import scripts.update_config. of-the-day manages its files through plugin-file-manager and its own web_ui_actions. The of-the-day branch of GET /plugins/config stays: it matches the real manifest id and still merges the on-disk category files into the form. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(api): read managers only from the blueprints api_v3/__init__.py and pages_v3.py declared module globals (plugin_store_manager, saved_repositories_manager, schema_manager, operation_queue, plugin_state_manager, operation_history, sync_manager, config_manager, plugin_manager) that nothing assigns: app.py sets the managers as attributes on the Blueprint objects, and every route reads them there. The one reader, backup restore's fallback to the module plugin_store_manager, could only ever fall back to None. _ensure_cache_manager() built a second CacheManager in the web process instead of using the one app.py puts on api_v3. The display routes now read api_v3.cache_manager, creating it on the blueprint only when nothing set it (the same None handling as the /cache routes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(web): drop run.sh and the unused log_config_change web_interface/run.sh was referenced only by web_interface/README.md; the service starts the UI through scripts/utils/start_web_conditionally.py and the README already documents `python3 web_interface/start.py`. log_config_change() in web_interface/logging_config.py was never called. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): delete unreferenced store_manager.js, diff_viewer.js, htmx-sse.js - js/plugins/store_manager.js (window.PluginStoreManager) and js/config/diff_viewer.js (window.ConfigDiffViewer) were loaded on every page but nothing reads either global. - htmx-sse.js (plus its CDN fallback) was loaded after HTMX, but no template or plugin page uses sse-connect / hx-ext="sse": the live streams run through LEDStreams in app-shell.js. js/plugins/state_manager.js stays: install_manager.js's updateAll() reads and refreshes window.PluginStateManager. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove app.js helpers nothing calls - hexToRgb, rgbToHex, validateForm, uploadFont and switchTab (whose 'switch-tab' event had no listener) have no caller in the templates, static JS or the plugin monorepo. - installPlugin: plugins_manager.js (loaded last) assigns window.installPlugin, and its own store cards are the only callers. - The showNotification fallback could never install: app-shell.js is deferred ahead of app.js and defines the same fallback at top level. - performanceMonitor only logged with ?debug=perf and read an unset this.measures; the marks it took on every load had no reader. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): drop app-shell.js refreshPlugin A top-level function in app-shell.js, so a window global, but nothing calls it (no inline handler, no window lookup, no string-built name). The other plugin actions in that block stay. updatePlugin is the live window.updatePlugin: plugins_manager.js only installs its own copy when none exists. uninstallPlugin/pollUninstallOperation, updateAllPlugins, executePluginAction and toggleNestedSection are replaced by later deferred scripts, but a click that lands while those scripts are still downloading reaches the app-shell copies, so removing them is not a pure no-op. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): remove definitions plugins_manager.js always overrides All of these are replaced before anything can call them, checked against the load order in base.html and the live window.* values: - openOnDemandModal/requestOnDemandStop stubs: the IIFE later in the same script assigns the real functions synchronously. - updatePlugin and uninstallPlugin stubs (`window.X || stub`): app-shell.js already defined both, so the fallback never installed. Same for the later updatePlugin override, gated on the live function containing '[UPDATE]', which app-shell.js's never does. - The first addArrayObjectItem/removeArrayObjectItem: reassigned by the top-level copies after the IIFE. - The first `function formatDate` in the IIFE: a later declaration of the same name in the same scope wins. - deleteUploadedImage, getCurrentImages, showUploadProgress, formatFileSize and getScheduleSummary: character-for-character copies of js/widgets/file-upload.js, which stays the owner. - `typeof X === 'undefined'` fallbacks and `typeof X !== 'undefined'` re-exports after the IIFE: always false, or a self-assignment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(web): render the shell directly and delete index.html index.html extended base.html with {% block content %}, but base.html defines no blocks, so none of index.html ever rendered: rendering both with jinja2 gives byte-identical output. index() still loaded the config, read config.json and config_secrets.json raw and json.dumps'd them on every page load for variables base.html never reads, and flashed errors that base.html never shows. It now renders base.html with no context. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): stop htmx-config.js replacing console.error and console.warn It swapped both globals for filters that dropped any error mentioning insertBefore / "Cannot read properties of null" when "htmx" appeared in the message or stack, and a list of Permissions-Policy warnings. That hid real errors from every script on the page, and made every logged error and warning report htmx-config.js as its source. The beforeSwap target validation above it, which prevents the insertBefore errors in the first place, stays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(web): quiet the widget load announcements and debug logs About 30 lines hit the console on every page load: one "... widget registered" per widget file, one "[WidgetRegistry] Registered widget: X" per registration, plus the registry, base widget and plugin loader announcing themselves. The load-time announcements are removed; the per-call ones (registry register, plugin widget loads, "Render called") now go through the page's debugLog switch (localStorage.pluginDebug), guarded because the widgets also load in node tests without it. fonts.html and wifi.html debug logging goes through debugLog as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(api): drop the removed music-auth and of-the-day JSON routes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
84afa9d64f |
refactor: delete dead Python code in the core (and stop storing Wi-Fi passwords) (#608)
* refactor(plugins): remove the no-op PluginHealthMonitor Its monitor loop did nothing (`if callbacks: pass`), register_health_check had no callers and api_v3.health_monitor was never read by any route. The live health data comes from PluginHealthTracker, which is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(store): drop the never-set uninstall tombstones Nothing in production called mark_recently_uninstalled, so the reconciler's was_recently_uninstalled check was always False. The persistent uninstall registry is what actually stops resurrection; the reconciler test now exercises that gate instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(common): delete unused config/display/game helpers, utils and error_handler Nothing in core, the web UI, scripts or the plugin monorepo imports config_helper, display_helper, game_helper, utils or error_handler; only their own tests did. The error_handler re-exports leave src.common's __all__; APIHelper, TextHelper, ScrollHelper, LogoHelper and the adaptive layout exports are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(config): drop ConfigService's unused versioning and save API ConfigVersion, get_version/get_version_history/get_version_config, rollback, save_config, reload, get_plugin_config and the backward-compat load_config/get_config_path/get_secrets_path had no callers. The display controller only uses get_config, subscribe, unsubscribe and shutdown, plus the file watcher. Change detection now compares against the current checksum instead of the last history entry. The subscriber tests asserted `callback.called or True`; they now reload the way the watcher does and assert the notification. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): drop unread plugin state history and callbacks plugin_state.PluginStateManager kept a bounded per-plugin transition history that only get_state_history (tests only) read; get_state_info reports a separate lifetime count, which stays. set_error_info and record_display had no callers, and set_state_with_error's `error` argument only fed the history. The web-side state_manager.PluginStateManager loses subscribe_to_state_changes, _notify_callbacks, set_plugin_error and get_state_version, none of which had callers; with no subscribers the old-state copy in update_plugin_state went with them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): remove unused PluginManager methods and attribute guards update_all_plugins was only called by a test (the display loop uses run_scheduled_updates); get_plugin_health_metrics, get_plugin_resource_metrics and get_plugin_state had no callers; and plugin_modules was written but never read. plugin_directories is now initialised in __init__, so the hasattr() guards around it go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): remove unused executor, loader, store and package helpers - PluginExecutor.execute_safe: no callers. - PluginLoader._parse_semver: only its own tests; compatibility.parse_semver is the live copy and test_compatibility.py already covers it. - PluginStoreManager.get_installed_plugin_info: no callers. - PluginResourceMonitor._local: never read. - src.plugin_system.get_store_manager and __api_version__: no importers in core, scripts or the plugin monorepo. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): stop storing Wi-Fi passwords in wifi_config.json WiFiManager appended every joined network's SSID and password, in plaintext, to saved_networks in config/wifi_config.json, and nothing (web UI, backup restore, scripts) ever read them back: NetworkManager keeps its own credentials. The writes are gone, and loading the config now drops any saved_networks key and rewrites the file, so passwords already on disk are scrubbed. Also removes _check_dnsmasq_conflict (never called) and _detect_trixie, whose result only reached one log line, along with the NM_CONNECTIONS_PATHS constant only it used. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): remove unreachable and unused DisplayController code - _follower_rebuild_scroll_image: never called. - mode_duration (never read) and last_mode_change (write-only). - The `chosen_cap <= 0` branch: chosen_cap is either the minimum of caps already filtered to > 0 or DEFAULT_DYNAMIC_DURATION_CAP (180). - The `max_duration < min_duration` branch directly after `max_duration = max(min_duration, max_duration)`. - The circuit-breaker branch's `display_result = False` and `manager_to_display = None`: the first is overwritten a few lines later, the second is already None there. - The bool-to-bool conversion of execute_display's result, which is always a bool. - The `loaded_plugins` lookup in _update_modules: PluginManager has no such attribute, so it always fell through to `plugins`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(vegas): remove unused config update, boundary finder and refresh VegasModeConfig.update had no callers outside its own tests (the coordinator rebuilds the config with from_config on a change); geometry.find_item_boundary and StreamManager._refresh_plugin_content had no callers at all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(run): drop the debug block that pretended to import the plugin system In debug mode run.py put src/plugin_system itself on sys.path and printed "Plugin system import successful" without importing anything. Nothing imports plugin_system modules by bare name, so the path entry did nothing either. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: delete tests that test nothing - test/plugins/test_{basketball_scoreboard,calendar,clock_simple, odds_ticker,soccer_scoreboard,text_display}.py skip everywhere the named plugins are not installed, including CI (LEDMATRIX_PLUGINS_DIR holds only the fixture plugin); test_plugin_matrix.py already covers every discovered plugin. Their PluginTestBase and the fixtures only it used (plugins_dir, mock_display_manager, mock_cache_manager, mock_plugin_manager, base_plugin_config in test/plugins/conftest.py) go with them. - test_plugin_system.py: test_discover_plugins (body was `pass`) and test_dependency_check (a comment), plus the test_plugin_manager fixture only the former requested. - test_display_manager.py: test_draw_image asserted that an image it had just assigned was not None. - test_display_controller.py: the rotation and schedule-override tests re-implemented the run-loop arithmetic inline and asserted on their own result without calling the controller. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: expect one plugin_last_update success stamp after update_all_plugins EveryStampRecordsACompletion required at least two success-path stamps; the second was update_all_plugins, removed as test-only. The worker and synchronous paths share the remaining stamp in _execute_update_now, and the check that every stamp calls _note_update_completed is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
e1ce7189f1 |
fix(install): make one-shot retry() retry, and drop root grants on user files (#606)
retry() in one-shot-install.sh used `if ! "$@"; then status=$?`, where $? is the status of the negation -- always 0. A failed command was never retried and retry() reported success, so a failed `git clone` carried on until a later check noticed the missing checkout. It now retries (3 attempts) and returns the command's status. The two apt steps stay non-fatal: warning and continuing is what they effectively did before, and making them fatal would stop installs that work today. A clone that keeps failing stops the install, as it already did, just sooner and with the one-shot's own error message. Both installers granted the web user NOPASSWD root on display_controller.py, start_display.sh and stop_display.sh. Those files are owned by the user after Step 11's chown, so the grant let the web user rewrite them and run them as root, and nothing ever ran them through sudo. Removed from both installers, with a test that every project file granted as root is a root-owned fix_perms helper. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
342e9164b8 |
fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound Each repeat of an error pattern appended every plugin in the time window to the pattern's list again, so a plugin failing in a loop grew the display process's memory without limit: 3,000 errors from three plugins reached 2.5 million entries. Keep the list unique. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): load a BDF font at its native size instead of PIL's default FreeType rejects any size but a BDF strike's own, and FontManager answered that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf requested at 8 or 10px rendered as PIL's default font. Retry at the native strike, as element_style already does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin toggle failures no longer claim "operation in progress" Every exception in POST /plugins/toggle was mapped to PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation is already in progress". Report the failure as what it is, and record the plugin id in the operation history for form posts too (it read a `data` variable that only the JSON path set). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): route plugin card clicks through handlePluginAction The document-level delegation checked `typeof handlePluginAction`, which is scoped inside the plugin-manager IIFE and so never visible to it. Every card click took a copied fallback that stopped propagation (the grid's own listener never ran), confirmed an uninstall twice, and sent Starlark app uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>. Expose the handler on window and delegate to it. Also run every test/js/unit suite under pytest: they need only node, but CI ran one of the eight. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): apply Rotation durations, WiFi messages and Vegas settings Three settings the web UI saves never reached the display: - Rotation & Durations: display.display_durations was never read. Every plugin inherits get_display_duration() and the plugin was asked first. A saved value now wins. The page shows unsaved screens blank with the plugin's own duration as a placeholder, and saving a blank removes the override, so one save no longer pins every screen. - WiFi status overlay: the controller looked for wifi_status.json one directory above the repo. Both sides now use wifi_manager.get_wifi_status_path(). The message is written by rename so the display never reads it half-written, and the resumed plugin redraws the whole panel afterwards. - Vegas: nothing called coordinator.update_config(), so saved Vegas settings never reached a running scroll. They are now queued when display.vegas_scroll changes, and applied while Vegas is stopped too, so a disable then re-enable works. The follower's scroll-speed default (75) now matches VegasModeConfig's (50). Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows with each plugin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep affected_plugins order when serialized; guard non-Element targets ErrorPattern.to_dict() ran the now-ordered list through set(), so get_error_summary() listed plugins in an unstable order. The document-level card-action listener called event.target.closest() without checking the target is an Element. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
967f3a0567 |
fix(install): parse the sudoers rules before installing them (#602)
Both installers generated the ledmatrix_web rules and copied them straight into /etc/sudoers.d without ever parsing them. Every rule is built from `which` lookups, so an empty or surprising path produces a malformed drop-in -- and a malformed file in /etc/sudoers.d makes sudo refuse every command for every user. On a headless Pi that is unrecoverable over SSH. first_time_install.sh now runs `visudo -c` on the generated file and, if it does not parse, prints what visudo said and leaves the installed file untouched rather than replacing it with a broken one. configure_web_sudo.sh does the same before it offers the rules for confirmation. first_time_install.sh also built the file at a fixed /tmp path as root; mktemp now picks the name. test/test_sudoers_is_validated.py renders the installer's own sudoers heredoc and checks the result with visudo -- the check neither installer had -- and asserts the install stays gated on it. Claude-Session: https://claude.ai/code/session_01Dby94z9PV3zVM25fqGNXTt Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
cf02538d2e |
fix(sports): ask the endpoint the league actually publishes for standings (#600)
* fix(sports): ask the endpoint the league actually publishes for standings ESPNDataSource.fetch_standings tried /standings first regardless of league and fell back to /rankings only on a 404. College leagues answer /standings with a 200 that carries no poll, so the fallback never fired and the poll came back empty every time. Nothing failed; the rank badge simply never appeared, and anything keyed off rankings quietly did nothing. Endpoints are now ordered by whether the league publishes a poll, a 200 that lacks the key counts as a miss so a league answering both still ends up with whichever one carries the poll, and only a 404 is treated as routine -- it is how a league says it has none. A connection error, a timeout or an unparseable body is logged as an error again. This is the implementation the football, baseball and hockey boards already ship; core was the last copy still on the old one. Verified against live ESPN: mens-college-basketball returns a populated rankings key where it previously returned nothing, and nba still resolves from /standings alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(standings): stop the endpoint handler from swallowing its own bugs Addresses both CodeRabbit findings on #600. The handler caught `Exception`, so an AttributeError or TypeError raised while *inspecting* the payload was indistinguishable from an endpoint that failed. The loop would move on and, if the other endpoint had nothing either, return {} -- silently dropping rankings for a league that has them. That is the precise failure this function was written to fix, so the handler was able to reintroduce it. Only the request is guarded now. `requests.RequestException` covers the transport failures and `ValueError` covers a body that will not parse; payload inspection happens after the handler, where a bug surfaces instead of being logged as a missing poll. A non-dict payload is treated as a miss explicitly rather than by tripping over `.get`. Tests: the fallback paths had no coverage -- the old single-endpoint code would have passed the suite unchanged. Added order assertions for both league kinds, a 200-without-a-poll fall-through, 404 and non-404 recovery, a non-object payload, and a guard proving a bug is no longer swallowed. `test_fetch_standings_returns_empty_on_error` faked a transport failure with a bare `Exception`, which only passed because the handler caught everything. It now raises ConnectionError, which is what actually happens. Verified by mutation: restoring standings-first fails 5 tests, restoring the catch-all fails the bug-not-swallowed guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
81e1bc596f |
fix(sports): stop the idle back-off sleeping through a kickoff (#599)
* fix(sports): stop the idle back-off sleeping through a kickoff A league with no live games backs its poll off as empty checks mount, capped by live_idle_max_interval. The escalation counts empty looks and nothing else, so a league three hours before kickoff is indistinguishable from one three months out of season. Both reach the ceiling -- and the ceiling then *is* the blind spot. Measured on two rigs on 2026-09-19: gaps of up to 928s between looks, ten of them at or above 900s. Reproduced in the wild on 2026-09-20, where an unpatched rig sat for fifteen minutes with eight NFL games in progress and had not noticed any of them. That is the "it doesn't pick up new live games until I restart it" report -- restarting being the one thing that forces an immediate look. The clamp costs no extra request: the live fetch already downloads the whole day's scoreboard, upcoming games included, so the earliest start still ahead of us falls out of the payload the manager already has. Before a kickoff the wait is shortened so it cannot run past it; just after one, the live cadence is held for _KICKOFF_GRACE_SECONDS, because a provider that has not yet flipped the status would otherwise look like another empty check and escalate the back-off again, right when the game is starting. The grace window needed a second pass. A soak caught it as dead code: the just-passed kickoff was replaced by the next fixture on the card the instant it passed, `now < start` went true again, and the back-off returned to its ceiling. Observed live -- the rig polled at 13:00:45, found nothing because ESPN had not flipped the status, then went quiet for a quarter of an hour. A kickoff inside the grace window is now kept. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * test(sports): pin absolute tolerances and correct a wrong grace expectation pytest.approx defaults to a relative tolerance. On a unix timestamp that is roughly 1790 seconds, so every kickoff assertion here was effectively vacuous -- it called a kickoff half an hour away "equal". All seven now pin abs=1. That hid a wrong expectation. test_an_earlier_kickoff_still_wins_during_the_grace asserted a game ten minutes out should displace one that kicked off moments ago. It should not, and the code does not: while the grace holds, the wait is the live cadence (30s), which is strictly tighter than clamping to the nearer kickoff would give (~600s). Letting the candidate win would set a ten-minute wait at the exact moment games are starting -- the dead grace window this branch exists to fix. The test now pins the real behaviour plus the safety property that makes it correct, and is renamed to say what it checks. Reported by CodeRabbit on the PR. The finding was right that code and test disagreed; the suggested fix was the wrong way to resolve it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
19686ab761 |
fix(web): plugin settings form shows schema defaults for unsaved keys (#597)
The server-rendered plugin settings partial rendered straight from the saved config, so an option added in a plugin update (geochron 1.2.0's show_date / show_date_line, default true) drew as an unchecked box, and the save route's missing-checkbox handling then stored it as false. Enum dropdowns likewise showed their first option instead of the default. - _load_plugin_config_partial runs the stored section through prepare_plugin_config (as GET /plugins/config does) before masking secrets, so a secret's schema default is masked too. - render_field falls back to the field's own default, covering children of objects that declare a default of their own (where the defaults extraction stops). - The legacy-boolean parity test now compares against the config the plugin actually runs with (defaults included). Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
92f1960d00 |
perf(sports): fetch ESPN date chunks concurrently (#596)
* perf(sports): fetch ESPN date chunks concurrently Since ESPN started rejecting `dates=YYYYMMDD-YYYYMMDD` on 2026-09-15, one season request became a chunk per month -- and a month over the 500-event cap becomes a request per day. A cold college-baseball season is about 130 requests, and they went out one at a time. That is slower than the 20s budget `_update_plugins()` shares across every plugin at startup, so scoreboards were logging `update() timed out` on first run and being deferred to the scheduled tick with nothing on the panel. Measured on a Pi 4 against live ESPN, March+April college baseball (63 requests, 3101 events): 11.2s sequential, 1.6s concurrent. Over a whole boot that moved football-scoreboard, ledmatrix-flights and birdnet-go inside the budget -- 13 plugins deferred before, 10 after. Chunks now go out six at a time, in two passes: months and edge days first, then the days of any month that came back capped. Six keeps the shared Session under requests' default pool_maxsize of 10, so no connection is discarded. Merged events still follow `espn_date_chunks` order -- a capped month's days are spliced back into its own slot -- so the payload does not depend on which request won the race. Request order is no longer significant, so the three tests that pinned it compare the chunks as a set and keep asserting the merged event order, which is the part callers actually see. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): drop capped month payloads before fetching their days Review of the concurrent chunk fetch found it raised the worst-case peak memory more than the concurrency explains. The old loop discarded a month that came back at the 500-event cap the moment it saw it; the rewrite kept every capped month alive in `results`/`slots` until all of their day requests had finished. Measured on a Pi 4 fetching 20260201-20260531 college baseball (four capped months, 5462 events), peak RSS growth over the call: sequential (main) 83 MB concurrent, months retained 121 MB (+43) concurrent, one worker 108 MB -- the retention alone was +25 concurrent, months dropped 98-100 MB (+16) docs/LOW_MEMORY_BOARDS.md puts a 1 GB Pi 3B+ at under 200 MB of headroom, where running out makes the board unreachable until a power cycle, so the difference matters. The remaining +16 MB is six responses parsing at once; three workers saved about 6 MB more, within run-to-run noise, so the worker count stays at six. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(sports): state what ESPN_CHUNK_WORKERS was measured to do, not more The comment claimed the sequential fetch made scoreboards blow the 20s startup update() timeout. A boot on this branch still deferred 12 plugins and timed out baseball-scoreboard while its season fetches took 0.74s and 1.12s: the startup budget is spent on other per-plugin work. Say what was measured -- 17.7s sequential, 2.6-3.3s concurrent -- and nothing else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
116abb0daa |
fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service BackgroundDataService always sent a season range first and, on a 400, fell back to chunks without recording the rejection, so every background season fetch spent a doomed request and live scoreboards learned nothing from it (or it from them). The worker now consults and sets the same 6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight to month/day chunks, and if every chunk fails the range is asked once for a real error without re-spending the chunks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep plugin asset and action routes inside their directories POST /plugins/assets/upload, GET /plugins/assets/list and POST /plugins/assets/delete joined the request's plugin_id onto assets/plugins unchecked, so '../../config' created, wrote, listed and deleted outside it. #561 guarded only the route that serves the files. All three now go through path_safety.resolve_under and answer 400 for anything but a plain name, and delete only unlinks a metadata path that resolves into that plugin's uploads directory. PluginManager.get_plugin_directory refuses ids that are not one plain path segment, so /plugins/action (which runs a manifest script from the returned directory) and every other caller get the guard; the action route also rejects such ids up front, covering its no-manager fallback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): report a no-op plugin update as already up to date update_plugin() returns True both for a real update and for "nothing to do" (a ZIP-installed monorepo plugin already at the registry version, a bundled plugin). With no git commit to compare, POST /plugins/update called every such success "updated successfully", so Check & Update All counted most official plugins as updated on every run. The route now reads what changed off the plugin itself (commit, else manifest version, else last_updated) and returns data.update_status (updated / up_to_date / local_only). The update-all toast is summarised by PluginInstallManager.summarizeUpdateResults from that status, falling back to the message for older servers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): scoreboard scroll speed no longer follows target_fps sports_scroll computed the crisp speed ladder against the global target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since frame-locked presentation (#545) the helper steps a fixed number of whole pixels per presented frame and the panel presents at its real refresh, so the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a 50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s. The ladder now uses the display manager's refresh_hz, then display.hardware.limit_refresh_rate_hz, then the default. target_fps is not consulted. Docstrings now say scroll_delay is ignored for pacing (no behaviour change there) and describe the fixed-step model. Tests: replace the tests that pinned target_fps as the ladder refresh and described time-based stepping; assert speed independence from target_fps (unit and end-to-end presented px/s against the real helper), that the fixed per-frame step is applied, and that scroll_delay does not change speed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): escape registry and upload values in plugin manager inline handlers The store, saved-repository and custom-registry buttons built onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone, so a custom registry entry whose id contained ' closed the attribute and added its own handler. One helper, jsStringAttr(), now HTML-escapes the JSON literal for every one of those handlers, and the store View button opens only http(s) repo links. The live window.updateImageList (plugins_manager.js loads last, so its copy wins over the file-upload widget's) wrote the uploaded file's original name, path and ids into markup raw; they are escaped now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note plugin asset, action and inline handler guards Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): let the root pip wrapper install web_interface/requirements.txt Update Code, the automatic update's health check and Install Base Requirements install web_interface/requirements.txt through safe_pip_install.sh, which only allowed the root requirements.txt. The first commit changing that file would fail its dependency install, and the automatic updater rolls back any update whose dependencies did not install -- on every device, for every newer commit. The wrapper now lists both core requirement files. Only their folders are resolved, so a requirements.txt symlinked out of the project is compared by its target and refused (previously the root file's own symlink target was what got allowed). The updater's file list is a named constant, and a test runs the real wrapper (pip stubbed) on every file Update Code and the rollback install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): do not retry plugin requests that got an HTTP answer PluginAPI.request wrapped everything that was not a structured error as NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a JSON error without error_code included. Check & Update All retries NETWORK_ERROR, so those updates were re-sent five more times with backoff, contrary to the #587 contract that an HTTP error response is the server's answer. NETWORK_ERROR now means only that fetch() rejected. Any HTTP response without an error_code, or with a body that is not JSON, is API_ERROR with the HTTP status attached. Tested against the shipped api_client.js. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): restart the stats window when an idle gap is dropped by size #582 dropped an idle gap from the frame stats two ways: the reset_scroll() sentinel, which also restarts the 5s window timer, and a size guard for scrollers that never call reset_scroll(), which did not. On that path the first real frame after the gap found the boundary overdue and logged a stats line for a one-frame window. Both paths now share one seeding helper. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): leave plugins alone when update_core's own rollback fails update_core returns rollback_failed directly when a partial pull or an update whose health check never started cannot be rolled back. run() only held plugins back for 'verifying', so those devices still got new plugin versions and a display restart on top of a core in an unknown state -- the opposite of what the health-check path does, and of the 3.4.0 changelog (plugins are left alone if the rollback fails). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(api): make the REST reference match the api_v3 package Every documented request body, query parameter and response shape was re-checked against the handlers in web_interface/blueprints/api_v3/. Fixes calls that failed as documented (repo_url, action_id/params, files/image_id, font_file+font_family, ?font=, cache key, auto_enable_ap_mode, plugin limit keys), removes the font-override endpoints dropped in #566, corrects response shapes (plugins/config, plugins/schema, health, metrics, operation history, github-status, fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds the 26 routes it omitted (backup, system auto-update/git, wifi radio, starlark editor, MQTT bridge, status endpoints, skins). Documents the merge semantics of partial JSON saves to /config/main and /plugins/config and the dim-schedule POST accepting GET's days shape, which land in the same change set. Replaces app.py line numbers and the removed api_v3.py path with file and function names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): remove the General-tab plugin system toggles that did nothing plugin_system.auto_discover, auto_load_enabled and development_mode had General-tab toggles whose help tips promised dormant plugins and verbose logging, but nothing reads them: every enabled plugin is discovered and loaded regardless. Remove the three toggles. The keys stay tolerated in stored configs. The save handler now stores a flag only when a client sends it; treating a missing key as an unchecked box would otherwise rewrite all three to false on every General-tab save, which still posts plugins_directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(scroll): remove dead code left by #523/#570 - Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read them since the numpy blend replaced the scipy path. - Drop ScrollHelper._last_integer_position and frame_time_target, which were written but never read. - Keep target_fps and set_target_fps() but document them as informational: nothing paces off them, yet ledmatrix-elections' test_scroll_pacing.py reads helper.target_fps back and third-party plugins may call the setter. - Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to throttle", set_sub_pixel_scrolling's "default: True", and set_frame_based_scrolling's claim that it steps. The plugins monorepo was grepped for every removed name; none is used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI FONT_MANAGER.md told plugins to read display_manager.font_manager, which does not exist, so a plugin following it failed to load with AttributeError. The shared FontManager lives on the PluginManager and BasePlugin._get_font_manager() returns it (with a fallback for harnesses). Also removes the Fonts-tab override workflow and element-override panels that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls /plugins/store/search does not exist (404) and the list endpoint reads query, not q. The registry guide's curl examples omitted the JSON Content-Type, so the handlers saw an empty body and answered 400. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(config): use the shared core-key list in the last three private copies StartupValidator warned "Plugin 'auto_update' is enabled but not found" on every display start with auto-update or a dim schedule on; the reserved plugin-id check missed auto_update, sync, location and the rest; and ConfigManager's (uncalled) orphan cleanup would have deleted display, schedule and auto_update. All three now read src/core_config_keys.py, which also gains CORE_SECRETS_KEYS for the github/youtube secrets sections. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): partial JSON saves to /config/main change only what they send A JSON body with one field reset every checkbox in the sections it touched: the MQTT bridge's brightness slider turned off disable_hardware_pulsing, inverse_colors, show_refresh_rate and use_short_date_format, and a timezone-only save turned off web-UI autostart and weekly auto-updates. Missing-means-unchecked now applies only to form posts: form-encoded bodies and the v3 forms, which mark themselves with a hidden __form_section input. Also on the config routes: - vegas_min/max_cycle_duration no longer match the generic *_duration rule, so they stop landing in display_durations and a blank one no longer rejects the whole Display save; - saving from the Raw JSON editor calls start_setup_if_needed like the General form, so enabling auto-update there finishes its setup; - the schedule and dim-schedule POSTs accept the per-day days.<day> shape their GETs return, as well as the flat form keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): install plugin dependencies from the configured plugins directory install_plugin_dependencies.sh scanned only plugins/, but the Plugin Store installs into plugin_system.plugins_directory (default plugin-repos), so the documented "Recommended" fix found 0 plugins on every store install. It now reads plugins_directory from config/config.json (relative to the project root or absolute, default plugin-repos) and also scans plugins/ for dev symlinks, installing a plugin reached through both only once. With set -e alone, `pip ... | tee` took tee's exit status, so a failed pip install was reported as success; set -o pipefail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: replace stale API names, line numbers and the api_v3.py path - ADVANCED_FEATURES: StreamManager methods that exist (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the real on-demand status envelope ({status, data: {state, service}}) - app.py:199 / :144 / :607-619 line citations and web_interface/blueprints/api_v3.py (now a package) replaced with file and function names in ADVANCED_FEATURES, CONFIG_DEBUGGING, PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE, PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README - CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use /config/raw/main to replace the file; describe where validation runs - TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints usage) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): verify the web interface that actually ships, on port 5000 verify_installation.sh failed every healthy install: it required the long-removed web_interface_v2.py and looked for a listener on port 5001, while the web interface binds 5000 (web_interface/start.py). It now checks the files ledmatrix-web.service runs (start_web_conditionally.py, web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the same 5001 port in its listen check, HTTP probe and printed URLs. Port matches are anchored so :50001 no longer counts as :5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(plugins): one display-size contract: display_manager.width/height CLAUDE.md (#580) says to read display_manager.width/height because matrix is None when hardware init fails; the development guide, the safety-harness doc and two DisplayManager docstrings still recommended matrix.width/height. The bundled starlark-apps plugin read matrix.width unguarded, so its magnify recommendation and frame scaling raised in fallback mode (e.g. after the Pi 5 hardware refusal). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(install): make install_service.sh --help print usage instead of installing install_service.sh parsed no arguments, so `sudo ./scripts/install/ install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md) rewrote ledmatrix.service, ledmatrix-web.service and both update-verify units and enabled/started them. It now handles -h/--help (usage, exit 0, no changes) and rejects any other argument with exit 2 before doing anything. Running it with no arguments, as first_time_install.sh does, is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scroll): describe the fixed-step model and document frame_hold Since #545 a crisp speed from scroll_config.configure() makes the helper advance a fixed whole-pixel step per presented frame with no clock, and the display manager's frame hold is part of the speed. The docs still described the removed wall-clock model: - scroll_config's module and configure() docstrings said speed is applied in time-based mode and that omitting the hold "falls back to fractional pixels"; omitting it actually runs the scroll frame_hold times too fast. - SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both modes, and read a 20 ms stats median as missed refreshes although that is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step, the hold-dependent healthy median, that target_fps plays no part, and that a hand-added scroll_pixels_per_second loses to a schema-default pair. - PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling) without frame_hold; it now documents the parameter (core 3.4.0) with a configure() + set_scrolling_state example. - update_scroll_position/set_scroll_speed and set_scrolling_state docstrings say the same. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark target_fps legacy; describe what Vegas scroll_delay does - General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after the sports_scroll fix nothing in core scrolling reads it. The field and its API validation stay so saved configs and plugins that read global_config['target_fps'] keep working. CONFIG_REFERENCE says the same. - Vegas frame_based_scrolling/scroll_delay were described as frame-count stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and still advances by elapsed time, so the applied speed is clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The config comments, render_pipeline comment and CONFIG_REFERENCE rows now say so. No behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(deps): describe how plugin dependencies are really installed The guides said the web service runs as root, that installs pick --user from os.geteuid(), and quoted a warning and a PluginManager._install_plugin_dependencies() method that don't exist. The web unit runs as the installing user; store installs go through install_requirements_file() and sudo safe_pip_install.sh (root), with a user-level fallback that says so, and load-time installs run in the display service's own (root) interpreter. Manual paths now use the configured plugins directory (plugin-repos/ by default) instead of plugins/, which store installs no longer use, and install_plugin_dependencies.sh is described as scanning that directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): count local changes one way for the preflight and the pull The automatic update's preflight ignored mode-only changes and anything whose status line contained plugins/ or plugin-repos/, then promised "Automatic updates will not stash your changes". perform_core_update used plain git status (modes count) and ignored only 'plugins/', then ran 'git stash push -- :!plugins', which nothing ever pops. So an edit to a bundled plugin under plugin-repos/, or the installer's chmods on tracked scripts, passed the preflight and was stashed away for good. - auto_update.local_changes() is the one predicate both use: core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/ excluded by leading folder rather than substring (a core file under web_interface/static/v3/js/plugins/ now counts). - Update Code's explicit stash leaves out both plugin folders; the pull's --autostash carries their edits and mode changes across and reapplies them. - The automatic updater calls perform_core_update(stash_local_changes= False), which refuses instead of stashing edits that appeared after the preflight; update_core reports that as 'blocked'. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): diagnostics follow the web autostart default and api_v3 package #556 made a missing web_display_autostart mean "start" (only an explicit false/off keeps the web interface down), but the diagnostics still said otherwise: diagnose_web_ui.sh reported a missing key as "defaults to false", diagnose_web_interface.sh said the web interface "will not start unless this is set to true" and recommended enabling it, and debug_web_manual.py printed False. Troubleshooting a down web UI pointed users at a non-cause. Both shell scripts now evaluate the setting with the launcher's own autostart_enabled() (inline fallback if it cannot be imported) and report on / off / not set (on) / unparseable config; debug_web_manual.py uses the same function. They also check web_interface/blueprints/api_v3/ __init__.py: api_v3.py became a package in #553, so every healthy checkout was reported as missing a file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(install): what install_service.sh installs; verify script port; no sudo for --help install_service.sh installs and starts ledmatrix, ledmatrix-web and the update-verify units, not only ledmatrix.service (systemd/README.md, README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help' as a harmless check; it now shows --help without sudo and warns what a real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh checks the web interface on port 5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note update-all, plugin system settings and script fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): size the preview after orientation and pixel mappers display_geometry.physical_size claimed to give DisplayManager's answer but only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured after the library's pixel mappers, so a Rotate:90 / orientation 90 chain previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for 128x64. Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown names leave it alone), and move the orientation composition here so DisplayManager and the preview share it. The module docstring no longer claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH. Audit finding F18. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): refuse settings the rgbmatrix library aborts on, on every board The library answers several settings with a NULL matrix or abort() rather than an error, so the display service crash-looped (Restart=on-failure) instead of reaching fallback mode: rows above 64, chain_length above 255 (uint8_t binding setter, documented as "no upper limit"), a misspelled hardware_mapping, and parallel 2-3 on a single-output mapping, reachable from the Display form on the default adafruit-hat(-pwm) mapping. #586 only guarded the Pi 5 subset. - src/matrix_support.py holds the rules for every board (Options::Validate ranges, binding integer types, mapping names and outputs from lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the API's numeric ranges. - DisplayManager checks them before building options and raises MatrixSettingsRefused, so a hand-edited config falls back with a logged, reported reason. Emulator mode only warns. - The config API refuses them with a 400 naming the setting; combinations are checked against stored values but reported only when the request sets a field involved. - The hardware status file gains "cause" (settings/library/forced). The fallback log and Display banner give the Pi 5 rebuild hint only for a library failure instead of rebuild + gpio_slowdown advice for every failure; one Pi 5 slowdown recommendation (1-3, start at 1). - The Display form offers classic/classic-pi1 and orientation 90/270 and renders any other stored mapping selected with a warning, so an unrelated save no longer rewrites them; the API accepts 90/270. Audit findings F03, F16, F19, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(display): library limits, template defaults and Pi 5 slowdown - rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs, classic/classic-pi1 mappings and orientation 90/270 documented. - Defaults are the config.template.json values: config migration adds missing keys from the template, so the listed "code defaults" never applied. - One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode, starting at 1. - Troubleshooting describes the refused-settings fallback, and CHANGELOG corrects the Unreleased "no upper limit" entry. Audit findings F19, F20, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py opens the panel with the service's options --measure and --demo built RGBMatrixOptions from a private copy of the display service's builder that had drifted: gpio_slowdown came from display.hardware (default 2) instead of display.runtime (default 3), and rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors, pixel_mapper_config and orientation were skipped, with different defaults (hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown was measured -- or garbled -- in a setup the service never drives. The option filling in DisplayManager._setup_matrix moves, unchanged, into DisplayManager.apply_matrix_options(options, config), which _setup_matrix calls and the script reuses (overriding only limit_refresh_rate_hz for --measure). The script now loads the whole config rather than the hardware block. Tests pin the script's options to the service's attribute for attribute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py recommends keys the resolver honours The ladder ended by telling users to set display_options.scroll_pixels_per_second. scroll_config ranks that key below the scroll_speed + scroll_delay pair, deliberately, and several plugin schemas default the pair into config, so the advised key was silently ignored (a schema-default 1/0.02 pair plus an advised 66 still resolved to 50 px/s). The advice is now the pair that selects the crisp speed exactly (pixels_per_frame every frame_hold/refresh seconds), explains that the pair outranks scroll_pixels_per_second, and gives the scoreboards' per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the printed pair over a schema-default pair and check it lands on the advertised speed and hold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula - SPORTS_UNIFICATION.md still presented honouring global target_fps as sports_scroll's added behaviour and its one user-visible gain; note that it was withdrawn because it had become a speed multiplier. - ADVANCED_FEATURES.md gave Vegas scrolling as (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s by elapsed time, through a 0.1-5 px per scroll_delay clamp when frame_based_scrolling is on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): scroll model fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(dev): link-github links plugins from the ledmatrix-plugins monorepo link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git, and those per-plugin repositories no longer exist: official plugins are directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the monorepo once into the dev directory, finds plugins/<name>, plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and links it under its manifest id. With an explicit repo URL it still links a single-repository plugin as before. dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a fork), plus plugins_repo and plugins_branch; github_pattern, which was documented but never read, is dropped and warned about. Ships dev_plugins.json.example and git-ignores dev_plugins.json, both of which the guide promised. Reading JSON falls back to python3 when jq is missing (get_plugin_id silently returned nothing without jq). update/status/list find the git checkout above a monorepo plugin directory (its .git is not in the plugin dir), and update pulls a shared checkout once. status no longer exits 1 when nothing is broken. Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration, workflow, store integration, hello-world link, submission), and the nonexistent scripts/git-hooks/pre-push-plugin-version and scripts/bump_plugin_version.py replaced with the real rule: bump the manifest version and run update_registry.py. scripts/dev/README.md and CLAUDE.md updated to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scripts): monorepo workspace layout; fix_perms and install READMEs MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin; setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the workspace file opens LEDMatrix plus ../ledmatrix-plugins. scripts/fix_perms/README.md listed cache directories fix_cache_permissions.sh never touches and a 'ledmatrix' service user that doesn't exist (also in scripts/install/README.md); adds safe_pip_install.sh. install/README: install_service.sh installs the web and update-verify units too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): keep the rollback's pip retries inside the unit time limit The health check reinstalled the previous requirements by trying the next bash path after any failure, including a 600 s pip timeout. Two files, two paths: up to 40 minutes of pip alone, while systemd stops ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the rollback half-way and leaving the update 'verifying' until the web UI calls it lost. - Like permission_utils.install_requirements_file, only a sudo refusal moves on to the next bash; a pip that ran and failed or timed out is not repeated. The refusal wording is one list (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only verifier and pinned equal by a test. - All reinstalls in one rollback share a 600 s budget. - WORST_CASE_SECONDS adds up every timeout on the longest path (27.5 min); a test holds it under the unit's TimeoutStartSec and that under the web UI's VERIFY_LOST_SECONDS. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools Plugin config was prepared differently depending on how it arrived: - JSON POST /plugins/config built a partial body on schema defaults, so {"enabled": true} reset every other setting of the plugin. It now merges onto the stored section first, as the form path already did. - Legacy-boolean normalization (#588) ran only at load: GET /plugins/config returned the raw boolean, posting it back failed validation, and hot reload handed plugins the raw section (a legacy dynamic_duration: true came back as a boolean). schema_manager.prepare_plugin_config (normalize, then defaults) is now used by PluginManager.load_plugin, both save paths, GET, the save notifications and DisplayController's hot-reload callback. - The JSON save's filter kept only enabled/display_duration/live_priority and dropped a submitted skin, skin_options or vegas_* tuning key. There is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES, used by validation and by the save filter; PluginManager's CORE_OWNED_CONFIG_KEYS is its vegas subset. - Plugin sections posted to /config/main were stored verbatim, including values /plugins/config rejects. They now go through the same preparation (_prepare_plugin_config_for_save, extracted from save_plugin_config), and a failing section rejects the whole save before anything is written. - dev_server read only top-level defaults and let a schema enabled:false win; build_full_config shallow-merged overrides, dropping sibling defaults; the harness extracted defaults differently from the device. loading.build_config now uses the device's extraction and preparation, and dev_server, check_plugin, render_plugin and the harness all use it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt-bridge): brightness changes apply live and touch nothing else The display service's hot reload applies a saved brightness within a few seconds, and /config/main no longer resets other display settings on a brightness-only JSON body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): automatic update hardening Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI It described web_interface_v2.py and index_v2.html (both gone), client-side form generation, one POST per field with {key, value}, and 'no nested objects'. The v3 UI renders plugin forms server-side from the schema (pages_v3 partial + plugin_config.html macros, nested sections and x-widgets), posts the whole form once, and save_plugin_config() merges onto the stored section, validates, splits x-secret fields and notifies the plugin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt): brightness saves apply via hot reload and leave other settings alone The bridge README said brightness is applied on the display's next restart; the display controller's config hot reload applies it within seconds. It also now states that the bridge's partial JSON save changes only brightness (the /config/main merge fix in this change set). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): don't log pip's output from the health check's reinstall pip can echo a private index URL with embedded credentials; permission_utils redacts it, the stdlib-only verifier cannot, so it logs the exit code only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark the plugin_system toggles as unused legacy keys auto_discover, auto_load_enabled and development_mode are read by nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and the REST reference listed them as live settings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): docs and developer tools group Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): legacy plugin-system toggles no longer count as a General save auto_discover, auto_load_enabled and development_mode have left the General form, so a post carrying only one of them is not a general-settings save and must not treat web_display_autostart and auto_update as unchecked. The plugin_system block itself is left as on main for the branch that reworks it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): config-save and plugin-config preparation fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(claude): re-check matrix_support.py rules when the library submodule is bumped Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address Codacy findings on the core audit PR - plugin_manager.prepare_plugin_config: when the fallback legacy-boolean pass also fails, log a warning instead of a bare except/pass. - api_client.js: request() refuses any endpoint that is not a plain path under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace, control characters) with INVALID_ENDPOINT before calling fetch(), and plugin ids are URL-encoded wherever they are put into a URL (also in the app-shell batch load). - test_update_all.js: pins both against the shipped client. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): check endpoint control characters without a control-character regex Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags the \x00-\x1f range in checkEndpoint's regex. Test the char codes instead; the endpoints refused are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(auto-update): make the seed script executable on disk, not only in the index On Linux Repo.publish() commits with -a, which recorded scripts/run.sh as 100644 upstream because the seed file was never chmod +x. The pull then brought in the same mode the installer chmod had made locally, so installer_chmod saw no mode change left to check. The updater was fine: with the upstream commit at 100755 the --autostash carries the device's chmod across. Verified under Linux (WSL, git 2.43): the old helper fails exactly as CI did, the fixed one passes all 63 tests in the file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7e5967e160 |
fix(web): find installed plugins before anything has discovered them (#594)
The web process discovers plugins lazily: plugin_manifests is empty until some endpoint calls discover_plugins(). Three routes consulted it without discovering, so they misbehaved for as long as nothing else had run -- which, after every ledmatrix-web restart, is until someone opens the dashboard: - POST /display/on-demand/start answered 404 "Plugin <id> not found" (or "Mode <mode> not found"). Measured on a rig: 404 for over three minutes after a web restart, until GET /plugins/installed ran. The browser UI loads the plugin list first, so API-only callers (the Home Assistant MQTT bridge, scripts) are the ones who hit it. - POST /plugins/toggle answered 404 "Plugin not found". - POST /config/main did not recognise a plugin section, so it skipped secret separation and merged the section as-is: the plugin's API key was written to config.json in plain text instead of config_secrets.json. Add _discovered_plugin_manifests(), which discovers when nothing has been yet, and rescans once when a specific plugin id (or, for on-demand by mode, a mode) is not found, so a plugin installed since the last scan is found too. _installed_plugin_ids() now uses it. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f475038895 |
fix(cache): web UI can read what the display service caches again (#593)
* fix(cache): web UI can read what the display service caches again ledmatrix-web.service carried CacheDirectory=ledmatrix. With User= set to the installing user, systemd re-owns /var/cache/ledmatrix and everything in it to that user and its primary group whenever the directory's owner differs -- for a directory root created, on the first start. That erased the root:ledmatrix setgid layout the installers set up, so every file the display service (root) wrote afterwards was root:root 0660 and unreadable by the web interface: WARNING - Permission denied loading cache for display_current_state ... Since #547 install_service.sh renders the web unit from the template, so every fresh install hit this. Measured on one rig: 392 unreadable files, and the web UI's display status, on-demand state and plugin health empty. Existing installs only receive `git pull`, never a reinstalled unit, so the fix for them is in the code the root display service runs: - DiskCache.set gives each file the directory's group (when the directory is group-writable) and 0660 on the open descriptor before the rename, independent of setgid. This also closes a window where a fresh file was visible as mkstemp's 0600. - DiskCache.share_existing_files repairs files an older version left behind, once per process from the cleanup thread. It works through O_NOFOLLOW descriptors and skips hard links and other users' files: the directory is writable by the web user, and root must not be steered into changing a file outside it. For new installs, the web unit drops CacheDirectory=/CacheDirectoryMode=, and install_web_service.sh stops replacing an existing directory's ledmatrix group with the user's group. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): on-demand and current-display status read the display's latest state Found testing the cache-permission fix on a rig: once the web interface could read display_on_demand_state at all, /display/on-demand/status kept answering "active" for over 100 seconds while the file on disk said "idle". Both status routes read the display service's keys through the web process's memory tier, which serves the first copy it read for the full max_age (120s). Read them with memory_ttl=0, as every other cross-process reader (plugin health/metrics, the on-demand mailbox) already does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(install): re-group the cache dir whenever the web user is outside its group install_web_service.sh replaced an existing cache directory's group only when it was root's. A directory in any other group the web user is not a member of -- root:ledmatrix, for a user who is not in ledmatrix -- was left alone, and every file root wrote there stayed unreadable to the web interface. Replace the group whenever the installing user is not in it. A directory whose group the user is already in (ledmatrix, or the user's own group where CacheDirectory= left it) is still left as it is: re-grouping a working directory strands the files already in it on the old group. When the group does change and root-owned JSON files carrying the old group are present, try-restart ledmatrix.service so DiskCache.share_existing_files re-groups them through its symlink- and hard-link-safe path, rather than a recursive chgrp. Verified under WSL's systemd for seven directory states (user group, ledmatrix member, ledmatrix non-member with and without root files, root:root, missing, unnamed gid); the previous version left the non-member case unchanged. Addresses CodeRabbit review on #593. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1d51efe4c7 |
fix(backup): restore over existing files on hosts without os.chown (#592)
_copy_file() replaces each restored file and then carries the previous owner across with os.chown. On Windows os.chown does not exist and st_uid/st_gid are 0 rather than absent, so the ownership branch always ran and raised AttributeError. That is not an OSError, so it escaped every per-section handler in restore_backup(): a restore over any existing config aborted at config.json and restored nothing. Skip the ownership step where os.chown is missing, as auto_update_setup.py already does. No change on POSIX. test_restore_over_a_file_the_user_cannot_write simulates root-owned files with chmod 0o444; on Windows that sets the read-only attribute, which blocks any rename over the file, so it is skipped there. The modes the app writes (0o644/0o640/0o600) replace fine on Windows. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7ae614aa35 |
fix(sports): recover from ESPN rejecting scoreboard date ranges (#591)
* fix(sports): recover from ESPN rejecting scoreboard date ranges Since 2026-09-15 ESPN's site API answers `dates=YYYYMMDD-YYYYMMDD` with 400 "Failed to get events endpoint." for every sport. Single days, months (`YYYYMM`) and season years still work. Every season and weeks-window fetch in core failed, including the background service the scoreboards submit their season schedules to. src/common/espn_dates.py re-asks a rejected range as whole-month chunks plus the leftover edge days, which tile the window exactly (a season is 8 requests, not 213). A month that comes back with exactly 500 events is truncated (college baseball's March) and is re-asked day by day. It also clamps `limit` to 500: above that ESPN truncates silently, e.g. college football returns 25 of 68 games for one Saturday at limit=1000. BackgroundDataService recovers rejected ranges on the worker thread and advertises `handles_espn_date_ranges` so plugins can tell whether to hand it a range. SportsCore, sports_shared, ESPNDataSource and APIHelper route through the helper or the clamped limit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): ESPN date-range fallback and limit clamp Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): stop re-sending ESPN date ranges once one is rejected Live scoreboards refresh every 30 seconds, and each refresh sent the range first, got the 400, then fetched the chunks: three requests where one used to do. After a rejection, ranges now go straight to chunks for six hours, then the range is tried again so the workaround retires itself if ESPN reverts. A single-day 400 does not set the memo, and when every chunk fails the range request supplies the error without the chunks being fetched a second time. Per-fetch chunk logging drops to debug; the rejection itself stays a warning. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): clamp limit only on ESPN scoreboard submissions The background service is generic, and limit above 500 only truncates scoreboards. /teams needs limit=1000 (college football has 762 teams and limit=500 returns 500), so a teams submission must keep its limit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2082665252 |
fix(config): load config_secrets.json on hosts without os.geteuid (#590)
ensure_shared_group_ownership() - the chgrp self-heal ConfigManager runs before reading config_secrets.json (#416) - looked up os.geteuid unguarded. That name does not exist on Windows, and the AttributeError is not an OSError, so it escaped the helper's best-effort handling and every except clause in load_config(). Any Windows checkout with a config/config_secrets.json got a ConfigError from every config load and could not import web_interface.app. That is what made test_update_all_plugins.py error at setup: its client fixture imports web_interface.app. It was not state leaked between test files - the trigger is whether the checkout has a secrets file. Return early when os.geteuid or os.chown is missing. No change on POSIX. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f9b3d6ae52 |
fix(web): accept every panel size and row address type the rgbmatrix library does (#586)
* fix(web): accept every panel size and row address type the rgbmatrix library does The Display form capped columns at 128 and chain length at 24, and its submit handler (fixInvalidNumberInputs) rewrote anything larger to the cap, so wide panels and long chains silently saved as the wrong size. The config API checked none of the hardware numbers, so values the library rejects (odd rows, parallel 4, PWM dither bits 3) saved and the matrix then refused to start. - Form limits now match the pinned library: rows even 8-64, cols >= 16 and chain_length >= 1 with no upper bound, parallel 1-3, PWM dither bits 0-2, PWM LSB nanoseconds 50-3000. - save_main_config rejects out-of-range rows, cols, chain_length, parallel, brightness, scan_mode, pwm_bits, pwm_dither_bits, pwm_lsb_nanoseconds and gpio_slowdown with a 400. - A stored gpio_slowdown or pwm_dither_bits of 0 renders as 0 instead of the default, so saving the tab no longer overwrites it. - Row Address Type offers 5 (SM5368 / B707 row shift register). Verified on a Waveshare 96x48 V2 (24S-A1) on a Pi 4 with the Adafruit Triple LED Matrix Bonnet: rows 48, cols 96, row address type 5, BGR, GPIO slowdown 8. - Help text and docs: FM6124-family panels use Panel Type Standard; on a Pi 5 the library supports only row address types 0 and 2. No change to the rpi-rgb-led-matrix submodule. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): drop the rows cap and document every display setting accurately Rows: no upper limit in the form or the API. Still even and at least 8. The current rgbmatrix library rejects more than 64 per panel, so a larger value saves but the matrix won't start; the help tip, README, config reference and troubleshooting section all say so, and nothing here needs changing if the library lifts the limit. limit_refresh_rate_hz: the form accepts 0 (the library's "no cap"), a stored 0 no longer renders and re-saves as 120, and the API rejects negatives. pwm_dither_bits stays 0-2: the library rejects 3 and 4, so the old form's 0-4 only ever let users save a config the display couldn't start with. Docs and help tips, checked against the pinned library and its README: - panel_type and rp1_rio get README entries - show_refresh_rate prints to stdout; it never drew on the panel - dither bits raise the refresh rate; the tip said they lowered it - scan_mode is about interlacing at low refresh, not wrong colours - disable_hardware_pulsing: hardware pulsing needs OE on GPIO 18 and the onboard sound driver off; software timing makes rows flash brighter - gpio_slowdown guidance agrees between the README and the UI - all 22 multiplexing values listed; every numeric setting states its range - troubleshooting for a blank panel after a settings change, jumping rows and brightness flashes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): reject true and 5.5 for row_address_type and multiplexing Both still went straight through int(), so a JSON true saved as 1 and 5.5 as 5. They now use the shared hardware range check like the other panel fields. Review feedback on #586. Also: the RP1 Backend tooltip said it is ignored on Pi 3/4 (it is ignored on every model but the Pi 5), and the README gave the dynamic-duration default cap as 90s; the code default is 180s. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: refuse matrix settings a Raspberry Pi 5 can't drive On a Pi 5 the pinned rgbmatrix library drives the panel through the RP1 chip, and that path supports only row address types 0 and 2, parallel 1-3 and the regular / regular-pi1 / classic / adafruit-hat(-pwm) mappings (Rp1PioConfigSupported in lib/rp1/rp1_pio_backend.cc). For anything else CreateFromOptions returns NULL; the Python binding doesn't check, so the display process crashed on its first call into the matrix and systemd restarted it into the same crash every 10 seconds. - src/pi5_matrix_support.py: the rule and Pi 5 detection, matching the library's /proc/device-tree/model check - DisplayManager raises before creating the matrix, so it is a logged init failure (reported by /api/v3/hardware/status) and fallback mode - the config API rejects those settings on a Pi 5 when a request sets row_address_type, parallel or hardware_mapping - the Display form offers only row address types 0 and 2 on a Pi 5, and warns when a stored value can't be used - CLAUDE.md: re-check the rule whenever the submodule is bumped Review feedback on #586. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9616a5a054 |
fix(web): auto_update and other core settings are not orphaned plugins (#589)
v3.4.0 shows "Plugin Config Warning - In config but not installed: auto_update. Reinstall via the Plugin Store, or remove these entries from config.json." auto_update is the core weekly-update setting from #581. Reconciliation treated every top-level dict not in its private _SYSTEM_CONFIG_KEYS list as a plugin id, and #581 could not know to extend that list. - Move core top-level keys into src/core_config_keys.py (CORE_CONFIG_KEYS) and use it in reconciliation. Tests fail if a config.template.json key or a key written by the general-settings save is missing from it. - A secrets-file key only counts as a non-plugin when no installed plugin has that id. Plugin secrets are namespaced by id, so installed plugins with secrets were reported as missing from config on every run. - still_unresolved() drops "not on disk" findings whose id is no longer a plugin entry in config, so a stored verdict clears without a restart. - A plugin whose id is a core key is skipped with a warning, and the fix never writes a plugin stub over or in place of a core setting. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c200b5837d |
fix(plugins): normalize legacy boolean settings before schema validation (#588)
The news plugin's schema turned global.dynamic_duration from a boolean
into an {enabled, min_duration_seconds, ...} object. Installs that have
not saved the news settings since still hold `true`, so every start
logged "Plugin news config does not match its schema (loading anyway):
Field 'global.dynamic_duration': Expected type object, got bool" and
flagged news degraded.
The settings form already reads such a boolean as {"enabled": <bool>}
(render_nested_section in plugin_config.html) and the next save writes
the object. The loader did not. It now applies the same rule before
merging schema defaults and validating, so the defaults fill in the rest
of the object and the plugin receives it in the new shape.
The rule lives in schema_manager.legacy_bool_as_object /
normalize_legacy_booleans. It applies at any depth of nested objects
but not inside arrays, matching the form, and only to a real bool under
an object-typed property with an `enabled` child. Every other mismatch
still warns. A parity test renders the template macro against the helper
so the two cannot drift.
Nothing is written to config.json at load: the normalization is in
memory, and the next save of the plugin's settings persists the object.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
9f2743471c |
fix(web): update-all skips Starlark apps and no longer misses plugins (#587)
* fix(web): update-all skips Starlark apps and no longer misses plugins Check & Update All posted every entry from /plugins/installed to POST /plugins/update, including the virtual starlark:<app_id> entries that list installed Starlark apps. The store manager cannot find those, so each answered 500 "plugin not found". Update-all now sends only plugin ids (install_manager.js, and the older app-shell.js copy), and the route answers a starlark: id with a 400 saying it is a Starlark app. A request that got no HTTP answer was recorded as failed and never sent again. On a device, a web-service restart mid-run killed the in-flight request and refused the next one, stock-news, which was left on 2.6.2 with 2.8.0 available. Such requests are now re-sent with backoff (about 30s) before being reported as failed. HTTP error answers are not retried. Tests: test/js/unit/test_update_all.js (run from pytest via test/web_interface/test_update_all_plugins.py so CI covers it) and the route contract for starlark: ids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): walk update-all retry delays without indexed lookup Codacy's ESLint security/detect-object-injection rule flagged retryDelays[attempt] as a High issue. The index was a bounded loop counter over a fixed array, but shifting a per-plugin copy of the schedule gives the same backoff without the pattern. No behaviour change: test/js/unit/test_update_all.js (21) and test/web_interface/test_update_all_plugins.py (7) pass unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fddb0e06db |
feat(web): honour x-display: hidden in plugin settings (#585)
* feat(web): honour x-display: hidden in plugin settings Plugins keep deprecated and internal keys declared so stored configs keep validating (weather api_key/radar_zoom, countdown's auto-generated row id), but the settings form drew them as live controls. A property marked "x-display": "hidden" -- or an object whose children are all hidden -- now gets no control at any depth: top level, nested sections, Advanced Settings (not counted either), array-table columns and the row editor. A hidden top-level key is not reported in __rendered_section. Saving never changes a hidden value. Plain and nested fields aren't posted, so the save's deep merge keeps them; _set_missing_booleans_to_false skips hidden booleans at every depth. A posted array row replaces the stored item, so hidden row properties are carried as JSON-encoded hidden inputs and decoded exactly on save (an id "1" stays a string). New rows get none. JSON API saves are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): read hidden row keys without dynamic property access Build the set of x-display: hidden item properties once and look values up through Object.entries, instead of indexing objects by a variable key on the lines this branch added (Codacy: object injection sink, 6 warnings). Behaviour is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8360220809 |
feat(common): sports_helpers — the helpers all nine scoreboards carry identical copies of (#583)
* feat(common): sports_helpers, the helpers all nine scoreboards copy verbatim Add src/common/sports_helpers.py: the helpers the scoreboard plugins' sports.py carry byte-identical copies of (docstring-stripped AST, checked at ledmatrix-plugins f09bff2), so a later plugins PR can delete its copies once it floors on the core release that ships this. - Free functions: clamp_window, clamp_seconds, logo_needs_refresh (lazy src.logo_downloader import, as in the plugins), spread_weighted_order, MIN_WINDOW_DAYS / MAX_WINDOW_DAYS. All nine plugins. - SportsHelpersMixin (no __init__, stateless): _mode_customization, _setting_int, _reset_dwell_on_reentry, _next_switch_index, _spread_weighted_order (all nine), _odds_color and _upcoming_date_and_time_text (all but ufc), plus the _favorite_key seam from base_classes core.py for later phases. A new module rather than more methods on sports_shared: a plugin that deletes a copy and relies on an existing module having grown the method fails at runtime with AttributeError on an older core, which neither the loader nor check_min_core_version.py can see; a missing module fails at load. Tests: behaviour for every helper, a derived host contract, and a parity test that AST-compares every body against every plugin copy when LEDMATRIX_PLUGINS points at a checkout (skipped otherwise). test_common_is_hardware_free.py imports src.common and every sports_* module with rgbmatrix blocked and scans src/common for module-level imports of src.base_classes, src.display_manager and src.plugin_system (no existing violations). Nothing in core imports the new module; no behaviour change. CHANGELOG Unreleased entry and a converging note in docs/SPORTS_UNIFICATION.md. __version__ is not bumped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(common): address review on sports_helpers and the hardware-free test - SportsHelpersMixin docstring and CHANGELOG: constructor-free, but it keeps lazy state on its host (_reset_dwell_on_reentry, _next_switch_index). - test_common_is_hardware_free: the runtime check now filters every FORBIDDEN package, src.plugin_system included; the AST scan resolves relative imports against src.common, so `from .. import plugin_system` and `from ..plugin_system import x` are caught. Guard tests for both. - Parity skip reason names the CI guard that runs the same comparison: ledmatrix-plugins scripts/check_sports_helpers_parity.py (#495). - _odds_color: line-level pylint disable for a not-callable false positive (getter is None-checked); the AST is unchanged, parity still passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
869e36fb2f |
feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback A General-tab toggle (off by default) checks for and installs LEDMatrix and plugin updates once a week, overnight in the configured timezone. - Pre-update checks skip (and report) instead of forcing: local edits or commits, merge/live rebase, no upstream, low disk, missing health check, or a version that was already rolled back. An abandoned rebase (HEAD back on a branch) is cleared, since it would otherwise block every pull. - The pull reuses the Update Code path (now perform_core_update(), which reports dependency install failures as data). - ledmatrix-update-verify.service, started via a .path unit from a request file, restarts the services from its own cgroup, requires them to come up and stay up, and otherwise resets to the previous commit and reinstalls the previous requirements. It runs a copy of the checker taken before the pull. - No SSH needed: switching the toggle on restarts the display service, which (as root) installs the two units from the repo templates for the web user. first_time_install.sh installs them too and takes --enable-auto-update / LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh). - Plugins update after the code passes its check; failures, blocks and rollbacks raise an Overview banner and show under the toggle. Tested end to end on a Pi: web-UI setup, a good update, a broken web service and a broken display (both rolled back), a blocked local edit, and an abandoned rebase found on the device. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(auto-update): address static-analysis findings - Replace the subprocess.CompletedProcess the verifier fabricated for a command that could not start with a plain namedtuple; nothing is executed there, but the scanner flags any CompletedProcess built from variables. - Mark the subprocess imports with the repo's standard B404 annotation (all calls are list-form argv, no shell). - Mark the rollback-failed message as not SQL (B608 matched its wording). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): CI failures on Linux - Keep the setup result when chown fails. CI runs as a non-root user, where chown to the web user raises; that discarded the result file, so the General tab would never learn whether setup worked. Regression test added. - Register the two new /api/v3/system/auto-update routes in the URL map snapshot. - Use utility classes app.css defines (space-y-1, hover:text-red-600). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): address review feedback - Health check: a failed restart command no longer lets the check run against the still-running old process; it counts as a failure (and after a rollback, as a failed rollback). An unreadable restart count is never treated as stable, since a crash loop looks healthy between attempts. - Installer writes the auto_update setting to a temp file and swaps it in, keeping mode and owner, so a running config watcher never reads a truncated config.json. - Verify unit quotes its command-line paths (install folders with spaces); setup refuses folder names systemd would reinterpret (%, quotes, backslashes, control characters) and says so on the General tab. - The auto-update status route no longer returns exception text. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): keep error detail in the status route's 500 test_web_error_detail requires every 5xx handler to log the traceback and return describe_exception(e), which redacts credentials, so failures are diagnosable from the web UI. Dropping it for CodeQL broke that policy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): dismiss route rejects non-object JSON with 400 A JSON array or scalar body made `.get('alert_id')` raise, returning 500. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): let the app-wide handler answer status-route errors CodeQL (py/stack-trace-exposure, #709) flagged the route's own except, which returned describe_exception(e). web_interface/app.py's error handler already logs the traceback and returns the same redacted detail for any unhandled exception, so the local copy is removed: same response, no new exception-to-response flow, and test_web_error_detail's policy still holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d01da3bd9f |
fix(scroll): stop timing the idle gap between scrolls as a frame (#582)
ScrollHelper.last_frame_time was set once in __init__ and thereafter only
at the end of log_frame_rate(). Nothing re-armed it when a scroll began, so
the first frame of every scroll was timed against the last frame of the
*previous* one and the whole idle period between them was recorded as a
single frame.
Measured over 3 hours on a 256x64 Pi 4, that produced 31 windows reading
Scroll frame stats - 0.0 fps over 1 frames | median 136776.02ms
p95 136776.02ms max 136776.02ms min 136776.02ms | stalls 0 (0.0%)
and -- worse, because it is not obviously wrong -- put the same gap in the
max field of otherwise healthy windows, where the worst values were 537s
and 604s. It also counted as one stall per scroll start: at ~500 frames to
a window that is ~0.2%, against measured stall rates of 0.07-0.16%. The
stall rate is the number used to judge whether a scroll change worked, and
it was the same order of magnitude as its own artefact.
The first frame of a scroll has no predecessor, so it has no frame time.
last_frame_time is now None until one is rendered, and reset_scroll() puts
it back -- the same treatment last_update_time already gets three lines
above, for the same reason. reset_scroll() alone is not enough, because the
scrollers actually emitting these lines never call it, so a sample at or
past the 5s log interval is dropped as well: nothing that renders a scroll
takes that long over one frame. Seeding also restarts the window timer, or
the boundary is already overdue when the second frame arrives and every
scroll opens by reporting a window of exactly one frame. A window whose
samples were all dropped now logs nothing rather than reporting the gap.
docs/SCROLL_PERFORMANCE.md documented the diagnostic in terms of a
"Frame time: N ms" line that
|
||
|
|
814c21de1c |
chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0 (#580)
* chore: mark skins unsupported, fix stale docs and preview size, prepare 3.4.0
Skins: no current scoreboard plugin builds on src.base_classes, so the only
skin hook (SportsCore._render_game) never runs. The plugin schema endpoint no
longer injects the Visual Skin dropdown, the store hides and refuses
"type": "skin" registry entries, and GET /api/v3/skins reports
supported: false with a message. Stored skin config still loads and saves.
src/skin_system/ and its tests are unchanged apart from the support flag.
Docs: check_plugin.py/render_plugin.py examples use --plugin; document
BasePlugin.get_update_interval() and its interaction with the manifest
update_interval; CLAUDE.md drops the stale template line number and
recommends display_manager.width/height.
Preview size: new src/display_geometry.py holds the size computation and
defaults DisplayManager uses (double-sided applied, chain_length default 2).
The web preview, /display/current, Starlark magnify default, sync handshake
and two dev scripts use it.
Release: __version__ 3.4.0, CHANGELOG 3.4.0 section plus a 3.3.0 tag note.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: address CodeRabbit review on #580
- Preview fallbacks (SSE stream and /display/current) use logical_size({})
(128x32, the shared default) instead of a hard-coded 128x64.
- display_geometry treats a non-mapping display/hardware block as missing,
so a malformed config.json falls back to defaults instead of raising
AttributeError (which turned the Starlark render into an HTTP 500).
- Docs: the static update interval falls back manifest -> plugin config
-> 60s, in both the API reference and the architecture spec.
Not taken: validating double_sided copies against chain_length/parallel.
An orientation Rotate: or U-mapper pixel mapper decides which axis panels
lie on, so the counts would reject working setups (the existing
vertical-split test is one).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(display_geometry): a non-finite hardware size raises ValueError, not OverflowError
CodeRabbit flagged the Starlark magnify default in
_standalone_render_starlark_app for truthy non-mapping display values. That
case was already handled by
|
||
|
|
914bf2002f |
fix(install): grant and harden safe_pip_install.sh in first_time_install.sh (#579)
first_time_install.sh granted the web user safe_plugin_rm.sh but not safe_pip_install.sh, unlike scripts/install/configure_web_sudo.sh. On devices set up only by the first-time installer, install_requirements_file could not use the root wrapper and fell back to a user-level install that root-run ledmatrix.service may not see. Also harden both sudo-granted helpers to root:root 755. first_time_install.sh never did this, and Step 11's project-wide chown to the user would undo it if placed in Step 10, so it runs at the end of Step 11.1. Add a test that parses the ledmatrix_web sudoers rules from both installers and asserts they grant the same commands, and that every granted helper is hardened (after the chown, in first_time_install.sh). Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
47afaaac2b |
fix(install): build rgbmatrix on ARMv6 Pi Zero / Pi 1, and stop locking users out of the submodule (#577)
The pinned rpi-rgb-led-matrix commit emits the ARMv7-only `dmb ishst`
instruction in lib/rp1/rp1_rio_backend.cc, guarded only by __arm__, so the
build fails on every ARMv6 board ("selected processor does not support
`dmb ishst' in ARM mode"). Bump the pin to upstream 1ee4f76, which merges
12d839f (guard on __ARM_ARCH >= 7) plus docs only.
The installer also needed two changes for that bump to reach anyone:
- git pull never moves an existing submodule checkout, so a device that
already failed would keep building the broken commit. The build step now
moves the checkout forward to the pin — never backward or sideways (a
`git submodule update --remote` checkout is left alone), and never fatal.
- The submodule git commands ran as root on the user's clone (git's SUDO_UID
exemption allows it), leaving .git/modules/rpi-rgb-led-matrix-master
root-owned and the user unable to run git in it. They now run as the
project directory's owner, and root-owned leftovers are handed back.
Root-owned installs keep running as root.
test/test_install_rgb_checkout.py covers the non-root sync scenarios under
the installer's strict mode, checks that every called _helper is defined
before use, and pins the one-shot-install.sh -> first_time_install.sh
contract. Verified the tests fail on five deliberate mutations. Root/owner
scenarios were exercised manually under WSL Ubuntu, and the library was
cross-compiled for arm1176jzf-s at both pins (old: rp1_rio_backend.cc fails
at line 120; new: 16/16 sources compile).
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
6d1cbfb70b |
fix(web): plugin config page survives stored values the schema outgrew (#578)
Two stored shapes broke the config form: * A scalar under a field that is now an object. News' dynamic_duration was a boolean and is becoming an object; render_nested_section did `key in true` and the whole page failed to render. Look into dicts only, and carry a legacy boolean over as the object's `enabled`, so the next save upgrades it without switching the feature off. * A custom feed logo with a path but no id. The template always emitted an empty `logo.id` input, which the save route parsed to null, failing the id's string type on every save. Emit it only when there is an id, as custom-feeds.js already does. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8e6d7c280f |
test(element-style): cover the stateless element_color clamp path (#576)
#569 fixed _normalize_color and #572 covered the resolver path. That test's own docstring notes the resolver "normalizes colour separately from element_color", and the other path had no test: the stateless element_color(), which src.common.sports_card delegates to and which every one of the nine scoreboard plugins takes for each per-element colour it draws. That is the path that regressed. element_color() moved here with the per-element customization framework, the coercion rejected out-of-range components where the reader it replaced clamped them, and a rejection reads as "not configured" -- so one component over 255 painted the element white while the user's colour sat in their config. Every scoreboard's test_element_text_colors.py failed on it, and it took two plugin PRs red on CI to surface. Six cases: clamping, in-range untouched, hex, unparseable fallback, missing element, and agreement with sports_card.coerce_rgb. The last is the point -- the two shared readers disagreed about the same value, so this asserts against coerce_rgb directly rather than restating the arithmetic, and any future move of element_color has to keep them consistent. Verified by mutation: restoring the rejecting coercion fails two of the six, alongside the resolver test from #572. Tests only; no source change. Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
11bf39cd66 |
fix(web): two plugin config saves that always returned 400 (geochron, news) (#575)
* fix(web): render widget-less arrays of objects as a table, not comma text An array of objects with no x-widget (geochron's `cities`) fell through to the comma-separated text input. Jinja joined each item as a Python dict repr, the save route read them back as a list of strings, and the schema rejected them -- so every save of the plugin returned 400 "Configuration validation failed", whatever setting was changed. Default such arrays to the existing array-table widget, which already edits arrays of objects and posts `field.N.key` inputs the save route rebuilds into a list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): don't leave an empty object stub in array items on save The unchecked-checkbox pass walked into every nested object of an array item looking for booleans, creating it when absent. A news custom feed with no logo came out with `logo: {}`, which fails the logo's `required: [id, path]`, so every save of the news plugin returned 400. Recurse into a scratch dict instead and attach it only if a boolean was actually set in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bc60b41445 |
test(element-style): cover ElementStyleResolver's colour clamp path (#572)
* fix(colour): clamp out-of-range text_color components instead of dropping them
_normalize_color returned None for a triple with a component outside 0..255,
and None means "not configured" to element_color -- so configuring
[300, 0, 20] silently handed the element its *default* colour rather than red.
Every scoreboard reads its per-element colours through this path, so the bug
reached all eight.
It is also the odd one out: sports_card.coerce_rgb and
SportsShared._coerce_rgb both clamp, and core's own test is named
test_coerce_rgb_clamps_rather_than_rejecting. The rejecting normaliser arrived
with the shared readers in
|
||
|
|
f9b1f87e8d |
fix: clamp colour components, and let the style editor actually take over (#569)
* fix(element-style): clamp out-of-range colour components instead of rejecting
A regression this framework shipped. The eight scoreboards used to read their
colours through sports_card.coerce_rgb, which clamps; routing them through the
shared element_color sent them through _normalize_color, which rejected any
component outside 0..255 and fell back to the default. So a configured
[999, -5, 20] -- a typo'd bright red -- rendered white instead of (255, 0, 20).
Their own test_element_text_colors.py caught it: one case of nineteen, in all
eight plugins, failing only once the core change reached main.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): the style editor takes over its own blocks -- and gets to at all
Two defects, both found by rendering the real partial in a browser rather than
by reading the code.
It was losing a race to its own fields. The hand-off guard asked "do any
fallback controls differ from their server-rendered defaults?" as a proxy for
"is someone editing this?". But the fallback holds this block's own font
fields, and the font-selector widget populates them on the same 50ms timer --
so a plain page load, with nobody touching anything, raced into "dirty" and the
editor removed itself, leaving the 701-line accordion form it exists to
replace. Measured: seven customization.*.font selects dirty ~60ms after
injection, clean again by 400ms. The question is whether a *person* typed, and
event.isTrusted answers exactly that; the listeners now go on synchronously,
because the edit worth protecting can happen before initWidget runs.
It took over too much. Taking over removed the whole fallback section, but a
customization block can hold more than styling -- football keeps
favorite_result_colors there -- so that removed the only UI those fields had,
and the editor also rendered them as an element, giving every row an "enabled"
and three colour columns. Core now marks the blocks it recognises as styling
(the compact declaration already did; hand-written adoption did not), the
widget renders only those, and the template drops only the children the widget
reports owning.
Verified on football's real schema: 28 rows across four mode tabs, columns
Element/Font/Size/Colour/X/Y, favorite_result_colors still editable with its
ten inputs, no duplicated field names, no console errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): style editor no longer drops layout-only fields it never rendered
CodeRabbit flagged elementKeys() in style-editor.js: render() claims the
whole customization.layout child as the widget's own (removing it from the
generic fallback renderer, since posting the same offset twice is worse),
but elementKeys() only listed keys that also have their own top-level style
block. A hand-written schema can put a key under layout that never got one
-- a logo, a timeout indicator, a possession arrow with a position but no
font or colour -- and that key's only control silently disappeared: no row
in the style editor's table (elementKeys never listed it) and no fallback
section either (layout was removed wholesale).
elementKeys() now appends any layout-declared key not already covered by a
style element, so table() renders a row for it (layout columns only, no
style columns) and the wholesale layout ownership claim stays truthful.
Verified against current code before fixing. New regression test
(test/js/unit/test_style_editor_element_keys.js, following this repo's
existing eval-extraction pattern for testing widget JS without a browser)
fails against the reverted function and passes with the fix; added to
run_all.js and the suite table in test/js/README.md.
Full pytest suite: 4887 passed, 62 skipped, 2 failed -- both the
pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias gap on this sandbox,
identical on origin/main, unrelated to this change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Dpg3HLWohdCUdzz2QNHanm
* fix(web): style editor no longer strands leaf-valued layout fields
A prior fix on this PR made elementKeys() append any layout-only key with
no style block of its own (a logo, a timeout indicator, a possession
arrow), so table() draws a row for it instead of losing it when the
wholesale `layout` claim removes the generic fallback. That covers a
layout-only key shaped like an object (x_offset/y_offset, ...), because
columnsFor() only ever produced columns from a key's *sub-fields*.
It missed the case where the layout-only key's own value is itself a
leaf -- a plain "show_logo" boolean directly under layout, no x/y object
underneath. elementKeys() still lists it (any row: no matching column),
so it renders as an uneditable blank row and its only control -- the
generic fallback checkbox -- is still gone. Confirmed by executing the
real widget's render() against a synthetic schema in Node (a DOM-stub
harness, not committed): the field's name never appeared as an <input>.
columnsFor() now gives such a leaf key a column keyed to itself
('layout-leaf'), and elementRow() binds it to the leaf's own path
(customization.layout.<key>, matching the name the fallback would have
used) instead of leaving every cell blank.
New regression test (test/js/unit/test_style_editor_layout_leaf_columns.js,
following this PR's existing eval-extraction pattern) checks the leaf
column is produced, is self-keyed, doesn't duplicate, and that a schema
with no leaf-valued layout key is unaffected; wired into run_all.js and
the suite table in test/js/README.md.
test/js/run_all.js: 84 + 6 + 6 = all suites passed (jsdom unavailable
here, DOM suites skip as before). Python suite untouched by this change;
test_style_editor_extra_fields.py, test_style_editor_save_roundtrip.py
and the one PIL-dependent style_editor_takeover.py case fail identically
before this commit -- missing flask/PIL in this sandbox, not this PR.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): keep layout-leaf style-editor columns distinct from name collisions
columnsFor() keyed a layout-only leaf field's column by its bare field
name. If an unrelated element's style block or another element's layout
axis block happened to declare a sub-field with that same name, the
`!seen.has(key)` guard skipped creating the leaf's column, silently
dropping its only control again -- the same failure the leaf-column fix
was meant to close, just reached through a name collision (CodeRabbit
review on
|
||
|
|
7e580dc005 |
fix(wifi): make Connect work from the setup AP (#571)
* fix(wifi): make Connect work from the setup AP Joining a network from LEDMatrix-Setup has to take the AP down first, which drops the phone that sent the request. The connect endpoint answered only after the attempt finished, so the browser never got a reply and the WiFi tab's Connect button appeared to do nothing. - /wifi/connect answers 202 immediately while the AP is active and connects in a background thread; the result (never the password) is reported via /wifi/status as last_connect_attempt. A second connect while one is pending gets 409. - connect_to_network holds a /tmp flag for the attempt; the monitor daemon skips AP management while it is fresh. Previously the daemon's disconnected counter, accumulated over the whole AP session, re-enabled the AP on its next tick in the middle of the connect. - The WiFi tab and captive setup page explain the handoff up front, and on reopening show why the last attempt failed. The wrong-password message now works: the route sets the error_type the captive page checks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(wifi): serialize connect attempts on both paths Addresses CodeRabbit review on #571: - Check for a pending attempt before branching on AP state. A background attempt takes the AP down long before it finishes, so a second click used to bypass the 409 and start a competing synchronous connect. - Record pending for the synchronous (non-AP) path too, so two requests can't overlap and have the first clear the daemon's in-progress flag while the second is still connecting. - Clear the pending state if the background thread fails to start, rather than refusing every later request until restart. - Say the setup network returns "within a few minutes": a stale flag plus the daemon's grace period can take longer than one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d1e821c625 |
fix(web): harden, polish and optimize the web UI per the Sept 2026 audit (#568)
* fix(web): harden, polish and optimize the web UI per the September 2026 audit Works through docs/archive/WEB_UI_AUDIT_2026-09.md (health 8/20). Implementation integrity (P0) - app.css now defines every utility class the templates and JS use, including .hidden, so the ~145 JS show/hide toggles work. Button reset, and base component rules (.btn, .form-control) wrapped in :where() so utility classes on the same element win. New static-audit test fails when a used utility class has no rule. Accessibility - Focus rings render (the old ring rule referenced undefined variables); one :focus-visible outline everywhere; skip link; labelled nav landmarks. - Shared dialog helper (js/utils/dialog.js): role/aria-modal, focus trap, Escape, focus return, applied to every modal. - Named icon-only buttons and labelled ~70 form fields. - Toasts announced once; errors persist >= 10s; one showNotification. - Captive WiFi page: live region, timeouts, dark mode, 16px inputs. Performance (Pi Zero 2 W) - SSE streams and tab timers pause when hidden or off-tab; the display stream only runs while a preview is visible. app-shell.js deferred. - Widget scripts served as one versioned bundle (/assets/widgets.js): 52 -> 21 script tags, 66 -> 35 requests on first load. - Stdlib gzip fallback when flask-compress is missing: first-load JS/CSS 1358 KB -> 291 KB on the wire. SSE untouched. Theming and responsive - File managers, form fields and Fonts upload on theme tokens; bare inputs themed in dark mode; no more white surfaces. - No horizontal overflow at 375px on any tab; 44px touch targets on coarse pointers; reduced-motion respected; header title truncates. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): clear Codacy findings on #568 - json-file-manager: focus-trap releases kept in a Map (no dynamic property access or delete; no value-returning forEach callback) - notification / schedule-picker: style and day-label lookups via Map - app.js: move the pending-queue assignment out of the expression - diff_viewer / error_handler: named function declarations instead of arrow consts No behavior change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: check the OAuth widget ships in the widget bundle base.html no longer tags widget scripts one by one; they load through /assets/widgets.js. Assert the page requests the bundle and the bundle contains google-oauth.js, which is what the test was protecting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): address review feedback on #568 - widget bundle version fingerprints every file (name, mtime_ns, size) - gzip fallback appends Accept-Encoding to an existing Vary header - dialog helper: releasing a non-top dialog no longer moves focus out of the dialog the user is in - labels: file-upload targets its file input; fallback config fields get label for/id pairs; native color input has a fallback name - utility audit also reads class names inside bound :class expressions Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): give the native color-picker input an accessible name CodeRabbit flagged this on PR #568 as an outside-diff finding (never posted inline, so it was missed in the round of fixes that addressed the other 6 review comments). The <input type="color"> only carried a title attribute; screen readers don't reliably announce title, and there's no other label naming the control when showHexInput is false. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): clear Codacy findings in app-shell.js - drop the unused catch binding on the SSE JSON parse - move the pending-notification queue assignment out of the expression No behavior change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): contain plugin widgets/ dir and bound style-editor retries From CodeRabbit review on #568 (code that arrived with the main merge): - serve_plugin_widget resolves widgets/ with resolve_under before resolving the manifest script under it, so a symlinked widgets directory can't become the containment base (CWE-22). New test. - style-editor init stops polling after ~10s when the widget never registers and leaves the plain fallback fields in place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
69d408b321 |
feat(core): one per-element display-customization framework, wired into the web UI (#566)
* fix(sports): rebuild un-shared faces through the pinned layout engine unshare_element_fonts re-instantiates a duplicate font face so two elements can be told apart by id(). It did so through bare ImageFont.truetype, which takes PIL's default layout engine rather than the one src/common/font_layout.py pins. Raqm and Basic disagree on fractional advances -- that disagreement is the reason the pin exists, having broken golden images across machines -- so a rebuilt face could measure differently from the shared face it replaced, on any host where Raqm is installed. These were the only two call sites in src/ bypassing the pin. The guard asserts that the rebuild goes through the pinned loader rather than comparing engine values: where Raqm is absent, bare truetype returns BASIC anyway, so an engine comparison passes whether or not the pin is honoured. The first draft of this test did exactly that and passed with the bug reintroduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): drop the two dead client-side config-form renderers generateConfigForm and generateSimpleConfigForm (580 lines) were defined on the Alpine component and never called: server-side Jinja replaced them, as pages_v3.py:641 records. Nothing in any template invokes them -- there is no x-html in the templates and no bracket access on the component. They carried their own x-widget dispatch, which made them an active trap: the next person adding a widget would reasonably think both renderers needed updating. plugins/config_manager.js (PluginConfigManager, 133 lines) goes for the same reason -- loaded on every page from base.html, referenced only by itself and by an archived doc. Kept, having checked them: widgets/example-color-picker.js is the worked example docs/widget-guide.md points plugin authors at, and widgets/plugin-loader.js is the client half of a documented feature (manifest-declared plugin widgets) whose server route is missing -- soccer-scoreboard already ships a widgets/custom-leagues.js that this loader is meant to fetch. That is an unfinished feature to complete, not dead code to delete. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(web): serve plugin-declared widgets, and actually ask for them LEDMatrixWidgets.loadPluginWidget has always fetched /static/plugin-widgets/<plugin>/<widget>.js, and docs/widget-guide.md has always documented that path, but nothing served it. soccer-scoreboard has shipped a 17KB widgets/custom-leagues.js since August that could never load. Both halves were missing, not just the route: - serve_plugin_widget serves the script from the plugin's widgets/ directory as text/javascript. The manifest is the allowlist -- only a widget the plugin declares is reachable -- so installing a plugin does not publish everything it ships. Path handling mirrors the sibling serve_plugin_web_ui: allowlist regexes, os.path.basename, resolve() + relative_to() containment, and the ledmatrix- prefix fallback. The declared script name is guarded too, since it comes from the plugin rather than the request. - The config form never requested one. Its x-widget dispatch is a hardcoded list of core widget names, so a plugin's own widget fell through to a plain text input. An unrecognised x-widget on a string field now asks ensureWidget() for it. The text input stays as the fallback and is removed only once the widget has actually rendered, so a missing or broken widget costs the user an editor rather than their configured value on the next save. - manifest_schema.json gains "widgets", so the declaration is validated rather than merely tolerated by additionalProperties. Verified in a browser against the real partial: a declared widget loads, registers and renders, and its field posts exactly one value; a field whose widget 404s keeps its text input and still posts its value. Not addressed: loadPluginWidgetsFromManifest still has no caller. The per-field ensureWidget path is lazier and is what the form now uses, so that bulk helper is dead weight -- worth removing, but left alone here rather than inventing a call site for it. Known limitation, documented: only string-typed fields take this path. object/array/boolean/number fields and enums are dispatched by the template's own branches, which still only know core widgets. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(element-style): a wrong-size BDF now keeps its font, not its size BDF fonts are fixed-size bitmap strikes: FreeType accepts only the pixel size baked into the file and raises for anything else. 32 of the 35 shipped fonts are BDF, so a size picked in the web UI usually is not a valid strike -- and load_font caught that failure with its generic "unloadable font" handler, which substitutes PressStart2P. Asking for 5x7.bdf at size 10 therefore rendered a completely different typeface, silently. It now falls back to the file's own native size instead, which is what SportsCore._load_custom_font_from_element_config has always done. The native size is read via FontManager._read_bdf_native_size rather than a fourth copy of that parser, matching how core.py already delegates. Also here, because they are the same code path: - native_bdf_size() is exposed for the web UI, which needs to know when a size field can take effect at all. None means "free choice". - ElementStyle.font_size now reports the size actually realised rather than the one requested. Callers lay out from it, and reserving space for a size nothing was drawn at is how this surfaces. - The module font cache is a bounded LRU (256) instead of an unbounded dict. The display process runs for weeks and every config save can add a (font, size) pair; every other hot cache in the codebase is bounded this way. Untouched configs are unaffected: the shipped classic fonts are the three TTFs, so nothing was hitting the substitution path by default. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): per-mode style and offset overrides Lets one element be styled differently per situation -- a scoreboard's live / upcoming / recent cards, weather's current / hourly / daily screens -- under customization.modes.<mode>. The mode is bound at construction rather than passed per call. That is what makes this cheap to adopt: SportsUpcoming and SportsRecent are already separate instances with distinct SKIN_MODE values, so binding once makes every existing style()/offset_value() call site mode-aware without editing any of them. A per-call mode argument exists for the rare host that renders more than one mode. The two layers answer different questions, deliberately: - The base layer keeps the existing "differs from the schema default" rule, because the save flow writes the full default object into config.json whether or not the user touched it. - A mode layer is pure override -- its fields default to None, so presence is intent. Nothing writes into it unasked, so there is nothing for the stricter rule to protect against. None therefore means inherit, and has to stay distinct from 0: a mode y_offset of 0 means "sit at the base position", not "no preference". This is the same distinction scroll_card.switch_* draws with "inherit". A malformed mode value falls back to the resolved base value rather than to the caller's default -- caught by the degradation tests, which is what they are for: resolving the mode first let one bad string in a mode block silently discard a good base offset. With no modes block, and for every existing caller, resolution is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): declare per-mode overrides in config_schema.json A plugin adds "x-style-modes": ["live", "upcoming", "recent"] alongside its x-style-elements declaration and gets a customization.modes.<mode> group per mode, with every field of every declared element repeated as an override. Those override fields are typed nullable and default to null, which is the whole trick. The save flow writes schema defaults into config.json wholesale, so giving a mode field the base element's default would make every mode a frozen copy of the base the first time a user pressed Save, and the base would stop reaching them. Null means inherit. The mutation test for this is explicit: with concrete defaults, a base font_size of 14 resolves as 10 with user_forced set. min/max from the declaration carry into the mode blocks, so an out-of-range override is rejected by validation rather than clamped silently at render time. Also: the emitted font field now carries "x-widget": "font-selector". The widget already shipped and the config form already allowlisted it -- the hint was simply never emitted, so the field rendered as a bare text box that the user had to type a font filename into. Verified through the real SchemaManager path -- load_schema, defaults extraction, merge_with_defaults, validation, then resolution -- rather than against a hand-built dict, since the thing at risk is what that pipeline does to a null. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): render the config form from the schema the save route validates The form read config_schema.json with a raw json.load while api_v3.save_plugin_config went through SchemaManager. Those are not the same schema: SchemaManager applies expand_style_elements, which turns a compact customization.x-style-elements declaration into the per-element blocks the form knows how to render. Without it, that customization object has an x-style-elements key and no "properties", so the template's object branch matched nothing and the section rendered as empty space -- while saving still validated against the expanded shape. of-the-day ships the compact form, so its customization section has been invisible in the web UI. pages_v3 gains a schema_manager the way it already has config_manager and plugin_manager. use_cache=False matches the save route, so an edited schema is not served stale during plugin development. The raw read stays as a fallback for callers that register this blueprint without one. Checked before making the change: load_schema does nothing here except read, validate and expand -- inject_skin_selector is a separate method it does not call -- so this is not a behaviour change for schemas without the declaration. The test pair renders the same compact schema with and without a SchemaManager, so it documents exactly what was broken as well as what is fixed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(web): style-editor widget -- a row per element instead of 65 accordions Rendered element by element, a realistic scoreboard's customization block is 65 nested sections, and reaching one per-mode font size takes five levels of expanding. The widget collapses that to one compact row per element -- font, size, colour, X, Y -- with a tab per declared mode. It emits ordinary inputs under the same dotted names the generic renderer would produce, so the save/validate/merge pipeline is untouched: no hidden JSON blob and no new server-side parsing. It is driven entirely by the schema block it is handed, so fields added to the schema later appear without editing the widget. If it fails to load or throws, the generic nested rendering it replaces is left in place. Fixing two things the save path got wrong for nullable fields, found by posting what the widget actually emits: - The indexed-array recombiner (text_color.0/.1/.2 -> one list) compared the declared type to the string 'array', so a per-mode colour, typed ["array", "null"], was never reassembled and failed validation on save. _parse_form_value_with_schema had the same comparison. - A blank nullable field became [] rather than None, which then failed the minItems the colour array declares. Null is the inherit sentinel, so it has to survive. And two things the widget itself got wrong, found by looking at it: - An unset base control fell back to the select's first option, so an untouched scoreboard claimed every element used 10x20.bdf -- and the size box then locked itself to that bitmap font's fixed size. Base controls now show the schema default; mode controls stay blank, because blank there means inherit. - Elements arrived alphabetised (Detail and Odds above Score). Flask's JSON provider sorts keys, so declaration order has to be stated explicitly; expand_style_elements now emits x-propertyOrder, which the generic renderer already honoured too. Size is disabled and shown as fixed for a bitmap font, using the scalable/native_size the font catalog now reports. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): visibility, alignment and scale per element Completes the customization vocabulary: hide an element, align it, and resize a logo, alongside the font/size/colour/offset that already existed. All three per mode. They resolve to "change nothing" until the user asks for something -- True, None and 1.0 -- rather than to whatever the schema declares. That is the same invariant the font fields keep: a caller that honours them still renders an untouched config exactly as it did before they existed. A schema default therefore does not count as a choice, which matters because the save flow writes that default into config either way. scale sits in the layout block with the offsets rather than in the element block, because it is geometry: a logo has a scale and no font. The widget's columns come from the schema, so a logo row shows visibility, offsets and scale and no empty font cell. Two bugs found by the tests rather than by reading: - A nullable enum needs null in its enum list, not just in its type. The mode copy of `align` defaulted to null and then failed its own schema, so a plugin declaring any enum field with modes could not save at all. Six tests failed on this before any of them reached what they were testing. - defaults_from_schema only ever extracted font/font_size/text_color, so the schema defaults for the new fields were invisible to the resolver and a declared default read as a user choice. Widget: the table scrolls horizontally and pins the element-name column. Nine columns do not fit the config panel, and clipping them hid the offsets entirely while scrolling them made every row anonymous. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): resolve elements under the names plugins actually use Two naming conventions collided as the scoreboards grew. Counted across the published schemas: the style block names elements with a _text suffix (score_text, status_text, detail_text), while the layout block mostly uses the bare noun (score, date, time, odds) -- except status_text, which kept the suffix in seven plugins and lost it in two. records vs record splits seven to two the same way. A lookup now tries the exact name first and then the spellings that mean the same thing. Exact-first is what makes this inert for any config that already matches; the aliases only decide cases that resolved to nothing before. This is also what makes migrating to the compact declaration form safe. That form uses one key for both blocks, so a scoreboard adopting it asks for layout.score_text while its users have layout.score saved -- without the aliases, every offset they had dialled in would silently become 0. Applies to the style block, the layout block, the schema defaults and the per-mode overrides, since the drift shows up in all four. Not attempting to canonicalise on write: renaming keys in config.json would break the plugins still reading the old spelling from their own bundled code, and the drift costs a dict miss rather than correctness. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(plugins): BasePlugin.styles -- per-element styling every plugin inherits Adopting the element-style system meant repeating three things in every plugin: a guarded import, finding its own config_schema.json, and rebuilding the resolver when on_config_change swapped the config dict. This is those three things once, on the class all 45 plugins already inherit from. title = self.styles.style('title_text', classic_font='PressStart2P-Regular.ttf', classic_size=8, classic_color=(255, 255, 255)) The classic_* arguments are the adoption contract: with nothing configured they come back verbatim, so a plugin that switches to this renders exactly as before until a user changes something. A plugin with one instance per display mode sets STYLE_MODE on the class and every existing lookup becomes mode-aware without a call site changing -- which is the point of binding the mode to the resolver rather than passing it per call. styles_for() covers a plugin that renders several modes from one instance. Schema discovery reads the concrete class's own module rather than this file, because this file lives in src/plugin_system where no plugin schema exists -- the same trap SportsCore._config_schema_path documents. The first mutation test for that passed anyway: an installed plugin's module directory and its entry under plugins_dir are the same path, so the test could not tell the two apart. The case where they diverge is a plugin symlinked in for development, and the test now forces that shape. Getting discovery wrong is silent rather than loud: with no schema the resolver has no defaults to compare against, so every configured value reads as a deliberate override and the plugin quietly stops honouring its own shipped styling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(element-style): adopt hand-written customization blocks, and widen the font list Nineteen plugins spell their style elements out longhand instead of declaring them -- football's block is 701 lines for seven elements -- and predate this system entirely. Core now recognises that shape, so they pick up the row-per-element editor and the real font picker on a core update rather than on a plugin release. Checked against every published schema: 21 plugins adopt, and the defaults of each still validate against the schema generated for it. Detection requires *every* field in a block to be one this system understands. A looser "has at least one style field" rule sweeps in baseball's `count`, which carries a text_color beside geometry that means nothing here. That distinction took three attempts to test: the first two assertions passed under both rules, because an over-eager rule leaves a fontless block looking untouched and only surfaces as an extra row in the editor. The hardcoded font enum is replaced rather than extended. Football lists five of the thirty-five installed fonts, which is why a font a user uploads can never appear in one. It is not a curated safe set -- it omits some twenty other faces that fit the declared size cap just as well -- it is the fonts that happened to exist when it was written. Widening it does need a guard, though, and not the one the schema already has: a bitmap font ignores font_size and renders at its size baked into the file, so `maximum: 16` cannot stop a 27px face. The picker now filters out fixed-size fonts taller than the element's own declared ceiling, which drops exactly the four that would overflow a 32px panel and keeps the other thirty. Per-mode overrides stay opt-in: core cannot invent a plugin's display modes, so `x-style-modes` remains the one line that unlocks them. Their layout half covers every positionable element rather than only those with a style block -- the two namespaces do not line up in a hand-written schema, and football positions six things (logos, timeouts, possession) that have no style block at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(web): remove the two Fonts-tab panels that reported invented data "Element Font Overrides" let a user configure an override, showed a success toast, and changed nothing. All three endpoints behind it were stubs -- GET returned a hardcoded {}, POST and DELETE returned success without calling anything -- each marked "This would integrate with the actual font system". Wiring them to FontManager would not have fixed it. The machinery there is real (_load_overrides/_save_overrides persist config/font_overrides.json, resolve_font applies them, and the countdown plugin genuinely consumes it), but the panel's element dropdown offered eleven invented keys -- nfl.live.score, clock.time, weather.current -- that no plugin has ever read. An override saved against one of those would have persisted correctly and still done nothing. "Detected Manager Fonts" goes for the same reason. It claimed to show "fonts currently in use by managers (auto-detected)"; its own comment said "we'll simulate this", and it listed every font in the catalog with a hardcoded usage_count of 1 -- the panel beside it, with fabricated numbers attached. Per-element font choice now lives in each plugin's own config editor, against the elements that plugin actually has, and covers size, colour, offsets, visibility, alignment and scale rather than family and size. Kept: the font library (upload, preview, delete), which works, and /fonts/tokens, which is a stub but genuinely feeds the preview's size dropdown. FontManager's override methods are untouched -- countdown uses them. Verified in a browser with the tab's JS running: no console errors, 35 fonts listed, upload and preview intact. Removing the panel meant unwiring it from populateFontSelects too, which would otherwise have bailed out early on the missing select and left the preview dropdown empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(sports): one reader for element colours and layout offsets There were two copies of the per-element colour read and three of the layout-offset read. They had already drifted -- the scroll-card renderer carries a comment about having ignored offsets its own schema advertised -- and each new capability had to be added to all of them or silently work in some places and not others. All of them now go through src.element_style, which is what carries the alias handling and the per-mode lookup. That lands immediately for the nine plugins importing these modules: a scoreboard asking for `score_text` offsets finds the `layout.score` its users configured, and a Live instance resolves its own colours through SKIN_MODE without any call site passing a mode. _normalize_color learned "#RRGGBB" in the process. The scoreboards' own readers have always accepted it, so the shared one had to, or consolidating would have quietly dropped a form users' configs may hold. _coerce_offset picked up the non-finite guard the scroll-card reader had and the other two did not. _get_layout_offset is promoted onto SportsCoreSharedMixin. Each plugin still carries its own copy in its bundled sports.py, which wins by MRO -- so adopting this is a deletion in the plugin, and until that deletion nothing changes for it. Note for whoever runs the suite next: test_display_dirty_tracking.py is order-dependent. Fifteen of its tests failed in one full run and passed in the next with no change in between, and pass in isolation. Pre-existing, unrelated to this, but it makes a full-run diff untrustworthy until it is fixed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): record the element-style work under Unreleased This file's own preamble asks for it: a plugin may delete its bundled fallback copy of a core module only when its manifest floors on the first release that shipped that module, which requires the additions to be recorded here against a version. Names a plugin can now import and floor on -- the stateless layout_offset and element_color readers, alias_keys, native_bdf_size, the resolver's mode binding, BasePlugin.styles, and the promoted SportsCoreSharedMixin._get_layout_offset -- plus the schema and web-UI changes, the four fixes and the three removals. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(fonts): log the BDF native-size read failure instead of swallowing it The bdf-native-size lookup in get_fonts_catalog() caught any exception and silently discarded it. Every other guarded read added in this PR (the manifest parse in _declared_widget_script, the SchemaManager fallback in _load_plugin_config_partial) logs before falling through to the same degraded behavior. This one didn't, which is the shape a silent-exception-swallow lint rule flags. Behavior is unchanged -- native_size still comes back None -- but a corrupt or unreadable BDF file now leaves a trace. Verified: font-related tests (140) and the full suite still pass, with only the 2 pre-existing Europe/Kiev/Asia/Calcutta tzdata-alias failures already present on origin/main. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address CodeRabbit findings on the style-editor/font-selector PR - Fix _load_font_sized double-wrapping the (font, size) tuple on the missing-font path, which handed callers a tuple instead of a font. - Fix _set_nested_value skipping an explicit None when the key already existed, which silently kept stale overrides when a user cleared a nullable per-mode field or blanked all channels of an indexed color. - Preserve BDF scalable/native_size metadata through fetchFontCatalog's catalog-format mapping so maxFixedSize filtering actually applies. - Stop caching an empty array on a failed font-catalog fetch so a later call can retry instead of being stuck with the failed result. - Keep a saved font selected in the style editor even when it no longer fits a newly declared maxFixedSize, instead of silently deselecting it. - Don't drop in-progress user edits to fallback fields when a plugin widget finishes loading asynchronously and takes over the form. - Tighten the removed font-override endpoint test to assert 405, not just != 200. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): a partial save no longer switches off checkboxes it never showed An HTML checkbox posts nothing when unchecked, so the save route walked the schema and forced every boolean missing from the form to False. That is right for the rendered form and wrong for every other caller: a script, the MQTT bridge or a curl against the documented endpoint never rendered a checkbox, and reading its silence as "all off" turns a one-field save into a mass disable. Found on hardware. Posting four customization.* keys to a live device switched off nfl.enabled, ncaa_fb.enabled and every display-mode toggle in one request. The form now reports the top-level sections it drew (__rendered_section), and inside those an absent checkbox still means unchecked -- including a section whose only fields are checkboxes that are all off, which no heuristic could recover. A post with no marker only touches objects it actually posted a field from. Meta fields are dropped before form keys are treated as config paths, because unknown keys are otherwise written straight into config.json. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(sports): resolve element colour by name, and honour visible/align/scale Two of the three gaps this framework shipped with. Colour by name. A draw resolved its colour by comparing the *identity* of the font object it was handed, which cannot tell two elements apart when they share a face -- so those draws went out white. Every bitmap font is in that case, because a freetype.Face cannot be re-instantiated to un-share it, which is how an element rendered in any of the 32 shipped BDF fonts silently lost a colour its picker had offered all along. _draw_text_with_outline now takes element="score_text" and reads the colour by name; the identity path remains for un-annotated callers, but narrows before giving up -- one configured colour among the sharers is the only thing the user can have meant. Visible, align and scale. The resolver has understood these since the framework landed and nothing consumed them: an element could be marked hidden in the web UI and still render. Adds the stateless readers, the mixin accessors, and a scale parameter on the one shared logo-sizing seam (keyed into the cache, so two elements scaled differently cannot be served each other's image). Naming an element in a draw also honours its visibility. Untouched configs are unaffected: every new parameter defaults to today's behaviour, and all ten affected plugins render pixel-identically to main across every harness size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(plugins): how to declare styleable elements; harden the widget's lookups The plugin-author guide for the compact x-style-elements declaration -- what each key does, how to read values back without breaking the "user-forced only when it differs from the default" rule, and why a hand-written block needs no changes to be adopted. Also clears the static-analysis findings on style-editor.js. Every lookup in that file is keyed by something out of a schema or a saved config, so a key of __proto__ or constructor would walk the prototype chain and hand back a function instead of a schema; reads now go through an own-property helper. The panel registry became a list, and the flagged vars moved to their function roots. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): clear the remaining static-analysis findings Five, all on lines this branch touched. The Python one is not a new defect: _set_missing_booleans_to_false's first parameter was always named `config`, which shadows the `config` submodule imported for its side effects at the bottom of this module. Editing the signature simply put the existing warning on a changed line. The parameter is the plugin's config dict, so `plugin_config` is what it should have been called anyway; callers pass it positionally and are unaffected. The JavaScript ones are the object-injection rule firing on reads keyed by data. own() now goes through a property descriptor, so the one unavoidable data-keyed read is no longer a computed member access; at() consumes its path instead of indexing it; and the column set is a Map, which has no prototype to pollute and needs no guarded reads at all. Verified the widget still renders identically against football's real schema: 29 element rows, all four mode tabs, values populated, no console errors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): drop the hasOwnProperty alias the descriptor read made redundant own() now reads through Object.getOwnPropertyDescriptor, so the alias it used to call has no remaining reference. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
92ac231138 |
fix(fonts): load 4x6 on its pixel grid, from any working directory (#565)
* fix(fonts): load 4x6 on its pixel grid, from any working directory `extra_small_font` loaded 4x6-font.ttf at 6, off the face's 7px grid. Under `draw.fontmode = "1"` the mono rasteriser thresholds each glyph at 50% coverage, so every glyph lost its fourth column and deformed: christmas-countdown rendered "UNTIL" as "VM1JL". The advance is 5px at both sizes, so snapping to 7 reflows nothing. - Sizes in DisplayManager._load_fonts go through crisp_size() instead of literals. crisp_size / FONT_PIXEL_GRID / FONT_NAME_ALIASES move to src/common/font_layout.py; sports_card re-exports them. - Mirror the fix in VisualTestDisplayManager, the harness's fork of _load_fonts. Without it every golden is blessed at the old size. - Resolve bundled font paths against the install root, not the cwd. FontManager._resolve_asset_path now delegates to font_layout.resolve_asset_path (kept by name; plugins probe for it). - The startup banner's middle rung snaps to 7; the 5 rung stays off-grid on purpose (the only size that fits a dotted quad on 64px). - loading.py reads all plugin JSON as UTF-8 (cp1252 on Windows aborted check_plugin.py on a 0x9d byte). - check_plugin.py reports in ASCII and never dies on an unencodable char. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(fonts): resolve relative asset paths from the install root, not the cwd resolve_asset_path checked os.path.exists(relative_path) unconditionally, so a relative asset path was still resolved against the process cwd first -- exactly the dependency this module exists to remove. An unrelated working directory that happens to contain assets/fonts/4x6-font.ttf (a stale checkout, a copied assets folder, another project) would shadow the real bundled font instead of the install root ever being consulted. Only an absolute path is now returned as-is; a relative path always resolves against _INSTALL_ROOT first, matching the docstring's stated contract. FontManager._resolve_asset_path delegates to this function, so it's covered by the same fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9ad7528c9b |
fix(config): stop same-second backups overwriting each other (#564)
* fix(config): stop same-second backups overwriting each other
A backup's version is its identity. save_config_atomic() hands the path
back, rollback_config(backup_version=...) looks that version up, and the
paired secrets backup is found by reusing the same string.
The version was stamped at second granularity, so two saves inside the
same second produced the same filename and the second shutil.copy2()
silently overwrote the first backup. The path a caller was still holding
then pointed at different content, and rolling back to it restored the
wrong config. A user saving twice in quick succession lost a restore
point with no error.
list_backups() made it worse. It parsed the version off Path.stem, which
drops only the last dot-component, so for config.json.backup.20240101_120000
parts was ['config', 'json', 'backup'] and parts[-2] was 'json' -- never
'backup'. The filename branch was unreachable: every backup fell through
to the mtime fallback and reported a second-granularity restamp of its
mtime rather than the name on disk, so a unique filename alone would not
have been enough for rollback to find the right version.
Stamp microseconds, and never overwrite an existing backup -- on a
collision bump a -N suffix rather than lose a restore point. Parse the
version off the exact glob prefix so it round-trips with the filename,
still reading the legacy second-granularity format so restore points that
predate this keep working.
Two tests had encoded the bug:
- test_multiple_config_changes asserted a rollback produced plugin1=45
with plugin2=15, a state no single backup ever held -- 45 was only in
the second backup, 15 only in the first. It passed because the two
saves collided onto one file, so the first version resolved to the
second's content. Corrected to the state that backup actually holds.
- test_backup_rotation asserted against a hardcoded max of 3 while
setUp configured 5, and still passed: every save in its loop collapsed
onto a single filename, so there was only ever one backup to count and
rotation was never exercised. It now asks the manager for its limit
and overshoots it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(config): fold collision suffix into ordering, close backup-path race
_parse_backup_version() stripped any trailing "-segment" unconditionally,
so a collision-suffixed backup parsed to the exact same timestamp as its
sibling and list_backups() had no deterministic way to order them. Only
strip the suffix when it's numeric, and fold it back in as extra
microseconds so same-tick collisions sort newest-first reliably.
_create_backup() also checked backup_path.exists() before shutil.copy2(),
which two concurrent callers can both pass for the same path -- the second
copy2() then silently destroys the first call's restore point. Reserve
each path (config and, when configured, secrets) with exclusive file
creation instead of a check-then-copy, retrying on a real conflict.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
6b3028ad58 |
test: isolate DisplayManager globals across modules, and name the failure (#563)
Follow-up to #562. That commit fixed the actual cause of the intermittent 15-test failure in test_display_dirty_tracking.py -- the emulator's fixed TCP port 8888, a machine-wide singleton that a concurrent pytest process takes away. This adds the two things that would have made it a five-minute diagnosis instead of a long one, and closes the other door into the same failure. Confirmed the module is order-independent as it stands, on this checkout: pytest test/ -q, three times 115 failed / 4464 passed / 63 skipped, byte-identical failure sets, the module 21/21 passed each time module forced last (197 files first) identical failure set module forced first identical failure set module after each of test_display_manager, test_display_controller, test_display_controller_vegas_tick, test_skin_system, test_sports_scroll, test_initial_update_budget, test_display_double_parity, test_initializing_screen all pass four concurrent processes on the file 21/21 each And reproduced the original, to be sure the diagnosis in #562 is the whole story. Holding 0.0.0.0:8888 from a separate process: HEAD's test/conftest.py 21 passed pre-#562 test/conftest.py 15 failed, 6 passed The 15/6 split is not arbitrary: the six survivors are the only tests in the file that never touch dm.matrix. conftest.py: DisplayManager is a process-wide singleton and the RGBMatrix / RGBMatrixOptions names it constructs through are module globals, bound once at import. All three are shared by every test module in the run, so a module that leaves an instance in _instance -- or leaves patch('src.display_manager. RGBMatrix') standing -- changes what the NEXT module builds, invisibly, and only in a full run. A module-scoped autouse fixture now resets the singleton and restores either binding if a patch outlived its module. Module-scoped rather than per-test so that files sharing one manager across their own tests keep doing so; only the leak across the module boundary is cut. Autouse fixtures are set up ahead of requested ones, so this is finalised after a module's own DisplayManager fixture. Verified with a throwaway pair of probe modules -- one leaks a patch and a singleton, the next asserts both are clean -- which passed and were then removed. test_display_dirty_tracking.py: _setup_matrix() swallows every construction failure and falls back to matrix=None, so a broken environment arrived as fifteen identical "'NoneType' object has no attribute 'SwapOnVSync'" errors naming neither the fixture nor the cause. The fixture now fails once, and says where to look; under a held port it reads DisplayManager fell back to matrix=None: RGBMatrix construction raised... Known causes: the emulator adapter losing a fixed TCP port to another process -- see pytest_configure in test/conftest.py -- or a patch('src.display_manager.RGBMatrix') leaked from an earlier test module. with WinError 10048 in the captured log directly above it. No regressions: full suite with both changes is 115 failed / 4464 passed / 63 skipped, failure set identical to the pre-change baseline. The 115 is the pre-existing Windows-environment baseline (os.geteuid, POSIX modes, fcntl); CI on Linux remains authoritative. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
59997594ac |
test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths Two shared-state problems, both of which show up as a permanently dirty checkout or an unreproducible test failure. test_display_dirty_tracking.py builds a real DisplayManager, whose _snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web UI reads. Every pytest process on the machine shares that one file, so two concurrent runs -- CI shards, a second worktree, an agent running the suite alongside -- overwrite each other's snapshot and the mtime assertions stop meaning anything. The module fixture now points it at a session-unique temp path; the individual tests that care still override it further. To be clear about what this does and does not fix: this is a real shared-path hazard, but it is NOT the cause of the intermittent 15-test failure in that module. That turned out to be the emulator's fixed TCP port, fixed in the follow-up commit. This change stands on its own merits. web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json and data/operation_history.json as the web interface runs, into a directory that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig that ever opened the web UI -- and every test run that constructs the app -- left three untracked files behind and a permanently dirty `git status`. Only data/.gitkeep is tracked under data/, so the negation keeps it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: stop the emulator binding a fixed port, so concurrent runs can't collide This is the cause of the intermittent full-suite failures we have been chasing: runs of identical code landing anywhere between 100 and 130 failures, while every implicated test passed in isolation. Six test modules set EMULATOR=true and build a real DisplayManager. The repo's emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to serve the dev preview. That port is a machine-wide singleton, so a second pytest process -- a CI shard, another worktree, an agent running the suite alongside -- loses the bind. RGBMatrix construction then raises, DisplayManager catches it and falls back to `self.matrix = None`, and every test that subsequently touches the matrix dies with AttributeError: 'NoneType' object has no attribute 'SwapOnVSync' which names neither a port nor a socket, and points at the wrong file entirely. Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its matrix-touching tests fail together or not at all -- the 15-test swing that made the totals look random. Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process and running test_display_dirty_tracking.py: without this change 15 failed, 6 passed with this change 21 passed The "raw" adapter renders in memory and binds nothing. Only display_adapter is overridden, in a throwaway config written per pytest process; the repo's emulator_config.json is untouched and `run.py -e` still opens the browser preview on 8888. Nothing in the suite referenced the adapter, and the tests wrap SwapOnVSync on the matrix object itself, so they are indifferent to what sits underneath. allow_adapter_fallback is forced off -- falling back would land us on the browser adapter and its fixed port, which is the whole problem. CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set to an absolute path: the previous behaviour depended on where pytest was invoked from, and silently wrote a default config into whatever directory that was. Verified no regressions: full suite on this branch and with origin/main's versions of the touched files, same machine, back to back -- 115 failed / 4347 passed on both sides, zero failures unique to either. That 115 is the pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell scripts, Linux-only binaries); CI on Linux remains authoritative. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: mark the shell entry points executable Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh` fails with "Permission denied" and only works if you know to prefix `bash`. That one matters most: the web UI's own error hint, added in #560, tells users to run exactly that path when a system action fails for want of passwordless sudo, and following that instruction verbatim did not work. All eleven carry a shebang and are invoked directly, never sourced. The two sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately left non-executable, which is what distinguishes a library from an entry point. Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied with `git update-index --chmod=+x` because this checkout is on Windows, where core.fileMode is off and the working-tree bit is not tracked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f6367d63ae |
security: triage the CodeQL backlog — 129 alerts, three of them live (#561)
* fix(web): escape quotes in every HTML escaper, not just & < >
The escapers are all `div.textContent = x; return div.innerHTML`. That
round-trip escapes &, < and > -- the only characters the HTML serializer
must escape in a text node -- and leaves quotes alone. Every widget then
interpolates the result into a quoted attribute value:
value="${escapeHtml(v)}" title="${escapeHtml(v)}"
so a value of `x" onmouseover="alert(1)` closes the attribute and adds an
event handler of its own. CodeQL reported this 83 times
(js/incomplete-html-attribute-sanitization) across the widget files.
It is one bug, not 83: the widgets each carry a standalone fallback that
did escape quotes, but they all prefer BaseWidget.escapeHtml when
window.BaseWidget exists -- which it always does in the shipped page -- so
the correct fallbacks were dead code and the incomplete shared one ran.
Fixed at each source instead of at the call sites.
app-shell.js already documented this exact gap in a comment and worked
around it by building DOM nodes by hand; that workaround stays (setting a
property cannot be got wrong), the comment is now accurate.
cache.html's delete button interpolated the cache key into
`onclick="deleteCacheFile('...')"`. Escaping cannot help there -- the
browser HTML-decodes the attribute before parsing it as JS, so `'`
becomes a real `'` again -- so the key moves to a data-cache-key
attribute that the handler reads back.
url-input.js additionally wrote a value straight into an <a href> after
validating it against a schema-supplied protocol list, and that list
accepted any RFC 3986 scheme -- "javascript" included. Scriptable schemes
(javascript, data, vbscript, blob, filesystem) are now refused both when
the list is normalised and when a URL is checked against it, and the
render path routes its href through the same check instead of emitting
whatever was stored (js/xss-through-dom).
test/js/unit/test_html_escaping.js reads each escaper out of the shipped
file and runs it, so losing the quote handling again fails a test rather
than a scan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): stop request-supplied names from reaching paths outside their base
Three of the py/path-injection alerts were live, not lint:
* GET /api/v3/plugins/<plugin_id>/static/<path:file_path> read any file
whose resolved path *string-prefixed* the plugin directory. Flask's
default converter forbids a slash but not dots, and
get_plugin_directory('..') returned the parent of the plugins directory
because it exists -- so every file under the project root then prefixed
that directory, config/config_secrets.json included. The prefix check
was also wrong on its own terms: with plugin dir "plugin-repos/foo",
"../foo-evil/x" resolves to "plugin-repos/foo-evil/x", whose string does
start with "plugin-repos/foo".
* POST /api/v3/plugins/of-the-day/json/delete interpolated the request
body's file_id into f"{file_id}.json" and unlinked it, unvalidated. A
file_id of "../../../../etc/something" deleted that file. This is the
one finding in the batch that destroyed data rather than exposing it.
* POST /api/v3/cache/delete passed the body's key through
CacheManager.clear_cache to DiskCache, which joined it as a filename and
called os.remove. Same shape, same result. The guard goes in
DiskCache.get_cache_path, the single choke point get/set/clear share, so
every caller is covered rather than just this route. Real keys are the
stems of files already flat in the cache directory -- that is how
list_cache_files derives them -- so nothing legitimate is turned away.
The rest of the cluster (web_interface/app.py's asset route, the plugin
update handler, _get_plugin_version, the plugin-schema read in config.py)
was guarded in ways that held, but each had grown its own version of the
check. They now go through one helper, src/common/path_safety.py, which
returns the *sanitised value* rather than a verdict -- so a caller cannot
validate one string and open another, which is how the two real bugs
above were shaped.
Also: WiFiManager.connect_to_network took the SSID and password straight
from POST /api/v3/wifi/connect into nmcli's argv. There is no shell there,
so CodeQL's py/command-line-injection alert overstates the risk -- but
nmcli reads a leading "-" as an option, so an SSID of "--ask" asks nmcli
to run differently rather than to join a network. Both values are now
checked for shape (802.11's 32-octet SSID limit, WPA's 8-63 char
passphrase or 64-char hex key, no control characters, no leading dash)
before any subprocess runs.
test/test_path_traversal_guards.py asserts on the filesystem, not just
the status code: a handler that returns 403 and deletes the file anyway
would pass the weaker check. Twelve of its cases fail against the
unpatched code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): refuse a plugin id that is not a plain name, don't truncate it
pages_v3 and scripts/dev_server.py ran request ids through
os.path.basename and carried on with what came out, so "../weather"
rendered the config form for "weather". Nothing escaped the plugins
directory -- the relative_to guards held -- but the handler answered a
request nobody made, and validating one string while the filesystem sees
another is the shape both live traversals earlier in this branch had.
Same treatment as the rest: safe_path_component rejects rather than
truncates, resolve_under returns the path it checked, and the call sites
use what those return. The three handlers that had hand-rolled
resolve-and-relative_to blocks lose about twenty lines to the shared one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(web): say what the plugin web_ui iframe actually is
The docstring claimed the fragment runs "in a sandboxed iframe". The
iframe in plugin_config.html carries no sandbox attribute, so the
fragment runs with the interface's own origin. That is fine -- the file
belongs to an installed plugin, and an installed plugin already runs
Python on the device, so the trust boundary is install rather than this
route -- but a comment promising containment that is not there is worse
than no comment. This is the context for the py/reflective-xss alert on
this handler.
Also drops the now-unused os/os.path imports.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): inline url-input's scheme guard at the previewLink.href sink
CodeQL flagged this line as a new high-severity js/xss-through-dom alert
on this PR even though it is already covered by SCRIPTABLE_SCHEMES: the
guard reached the sink through safeHref -> isValidUrl, two function calls
away, which its DOM-based-XSS sanitizer recognition does not trace.
Behavior is unchanged -- same scheme check, same SCRIPTABLE_SCHEMES list,
same allowedProtocols gate -- just inlined directly above the
previewLink.href assignment it guards, so the barrier is visible in the
same scope as the sink.
Added a regression test that runs the shipped onInput handler (not just
the extracted helpers) against a mocked DOM, so a future change that
reintroduces an unguarded previewLink.href assignment fails here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): address CodeRabbit findings on the CodeQL triage PR
- src/wifi_manager.py: reject non-ASCII WPA-PSK passphrases before any
credential-saving or connect flow runs. NetworkManager only accepts
printable ASCII passphrases (or a 64-char hex key); a non-ASCII value
was previously saved/attempted before nmcli itself rejected it.
- web_interface/blueprints/api_v3/config.py: fail closed when the
plugin config schema path can't be resolved under the plugins
directory (e.g. a symlinked plugin dir). Previously this fell
through with secret_fields left empty, so submitted credentials for
that plugin were saved as ordinary, unencrypted configuration.
- web_interface/static/v3/js/widgets/plugin-file-manager.js: stop
splicing the JSON day/column key into an inline oninput="..." handler
string. escHtml() escapes quotes for a normal HTML attribute, but the
browser HTML-decodes the attribute before running it as script, which
undoes that escaping and lets a crafted column name (e.g. from an
uploaded JSON file) break out of the JS string and execute. Cell
edits now travel through data-day/data-col attributes read by one
delegated 'input' listener instead.
While in this file: fixed 6 pre-existing missing-')' typos on
multi-line safeSetHTML(...) calls (already flagged by Biome in this
PR's own CodeRabbit run as syntax errors blocking its lint pass).
These predate this PR (present on main too) but made the whole file
fail to parse in any JS engine, which is a bigger problem than the
XSS finding itself and directly touches the same lines.
Added/extended regression tests for each fix; full suites pass
(pytest: 4580 passed, 62 skipped; JS: 84 assertions).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5137e86d16 |
feat(tools): MQTT bridge and Pixlet editor, ported onto the api_v3 split (#554)
* feat(tools): manage the MQTT bridge and Pixlet editor from the Tools tab PR #544's change, ported onto the api_v3 package split (#553). Identical behaviour; only the placement of the new code differs. The original added 508 lines to web_interface/blueprints/api_v3.py, which #553 deletes, so every hunk of it would conflict irreconcilably. Ported by AST: 26 new top-level items sorted to where the split puts each kind -- __init__.py 2 imports, 11 constants, 7 helpers starlark.py 4 routes (/starlark/editor/{apps,status,start,stop}) misc.py 2 routes (/integrations/mqtt-bridge{,/config}) Everything outside api_v3.py -- the Tools partial, the installer scripts, the JS tests -- applied unchanged. Routes: 111 from the split plus these 6 = 117, and the url-map snapshot is regenerated to match, which is exactly what test_api_v3_url_map.py is designed to make you do when routes are added. Full Python suite: 4,278 passed, 68 skipped, 0 failed. The JS tests this PR ships could not be run here -- node is not installed on this machine -- so test/js/dom/test_tools_sections.js is unverified. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(starlark): don't crash the pixlet editor's start/stop routes, and honor an operator-set PIXLET_EDITOR_HOST The AST-based port of #544 onto the api_v3 package split dropped `time` from starlark.py's import list. start_pixlet_editor() and stop_pixlet_editor() both call time.time()/time.sleep() directly, so every start (NameError building `state['started_at']`) and every stop that has to wait out the EXIT trap crashed with a 500. No test caught it because the route's own tests mock subprocess.Popen but never actually invoked it before now. Also carries over #544's later fix that this port branched before: env['PIXLET_EDITOR_HOST'] = '0.0.0.0' unconditionally overrode an operator who had already pinned PIXLET_EDITOR_HOST to loopback, forcing the unauthenticated `pixlet serve` process onto the LAN regardless (CodeQL CWE-1188). Switched to env.setdefault(...), same as api_v3.starlark.py's siblings already do for _pkg-owned names. Both fixes route the shared _pkg.time reference the rest of the package's route modules already use for anything a test might need to patch, rather than a bare `import time` local to this file. Ported the existing regression test from #544 (TestPixletEditorHostDefaultsButDoesNotOverride) onto this branch's module layout (web_interface.blueprints.api_v3.starlark instead of the old monolithic api_v3 module), which is what caught the NameError. Full suite: 4330 passed, 62 skipped, 2 failed -- identical on this branch and on origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): clear the six lint errors this rebase introduced All six were introduced by rebasing this branch onto the merged blueprint split, not by the split itself. Confirmed by diffing pyflakes output against main with line numbers normalised -- everything else it reports is present on main too and is the package's deliberate re-export pattern. starlark.py used _STARLARK_APPS_DIR three times without importing it (F821). The rebase resolved an import-list conflict as a union of both sides, and that symbol was on neither side of the conflict hunk, so it was silently lost. It is defined in __init__.py and is now imported like its neighbours. This was the only one of the six that would fail at runtime rather than merely lint. __init__.py imported contextlib twice (F811): the cherry-pick added one next to the existing import. Removed the duplicate; the original at line 19 is used. __init__.py imported signal purely to re-export it to starlark.py, so pyflakes saw it as unused (F401). signal is stdlib and does not need routing through the blueprint package, so starlark.py imports it directly and __init__.py no longer does. contextlib stays re-exported because this module genuinely uses it. _read_mqtt_bridge_config()'s local `config` shadowed the `config` submodule this module imports at the bottom for its route side effects (F811). Renamed to `settings`, with a comment saying why, since the name is otherwise the obvious one to reach for. Verified: pyflakes now reports nothing on this branch that main does not, the package imports, all nine route modules load, and 117 routes register, matching the pinned URL-map snapshot. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): reject MQTT bridge bodies the endpoint cannot apply Two CodeRabbit findings on the bridge settings endpoint, both of which returned 200 while doing something other than what the caller asked. `request.get_json(silent=True) or {}` turned a missing or unparseable body -- and the JSON literals null, [] and false -- into an empty dict, which then satisfied the isinstance(data, dict) guard on the very next line. The guard was there to reject exactly those bodies. Dropping the `or {}` lets None fail it. The same `or {}` on /errors/clear is left alone: its docstring documents the body as optional, so an absent body legitimately means "use the defaults". The difference is that saving settings has nothing sensible to do with no body. `if data.get('clear_password'):` accepted any truthy value, and the string "false" is truthy in Python -- so a client echoing the field back as a string wiped a password it meant to keep. Now coerced through the package's existing _coerce_to_bool, which already maps 'true'/'on'/'1'/'yes' and nothing else. test_mqtt_bridge_config_endpoint.py covers both: five unusable body shapes plus a missing body, and clear_password across truthy and falsy spellings. Verified against the unfixed code -- reverting the body guard fails 5, reverting the coercion fails 3. Not changed here: CodeRabbit also asks this endpoint to reject MQTT credentials when TLS is off (CWE-319). That is a policy decision about the feature rather than a defect -- unencrypted MQTT on a trusted LAN is common and often deliberate -- so it is raised on the PR for a maintainer call instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: work through the remaining review findings on the editor and bridge allow_insecure_mqtt (CWE-319, requested): a password with TLS disabled crosses the network in cleartext. Refused now rather than merely warned about -- but refused, not forbidden, because unencrypted MQTT on a trusted LAN is a normal deliberate setup. allow_insecure_mqtt is the explicit acknowledgement, defaults false, and is coerced like the other booleans so the string "false" cannot switch the guard off. starlark.py:796 -- the supported service runs Flask threaded, so two start requests could each see running=False, each launch an editor, and the second state write replace the first PID, orphaning a process that holds the display down with nothing recording it. The check-launch-write sequence now takes a module-level lock. starlark.py:848 -- if the state write failed the route returned success with an editor running and no PID recorded: status and stop both reported no session while the display stayed down until the timeout expired. It now terminates the process group and returns an error. starlark.py:890 -- SIGKILL gives the script's EXIT trap no chance to run, so nothing hands the display back, yet the response said "the display is restarting". After an escalation the display is now restarted explicitly, and a failure to do so returns an error naming the manual step instead of a success. pixlet_config_editor.sh:184 -- find_pixlet supports Darwin but macOS ships no timeout(1); GNU coreutils installs it as gtimeout. Resolved up front so the failure lands before the display is stopped rather than after. pixlet_config_editor.sh:154 -- wildcard, loopback and an explicit interface address are three cases, not two. Collapsing the last two printed a URL saying "localhost" whenever PIXLET_EDITOR_HOST named a LAN address. tools.html:1254 -- escHtml does not encode single quotes, and the app id was interpolated into an inline onclick="startPixletEditor('...')", so a directory containing an apostrophe could break out of the JS string and run script. The handler binds with addEventListener and reads the id from dataset, where it is only ever parsed as an HTML attribute. Tests: test_mqtt_bridge_config_endpoint.py grows to 23 cases covering the opt-in in both directions. The tools DOM suite gains three guards asserting the edit buttons carry no inline onclick and pass the id via dataset -- those need jsdom and did not run here, so CI verifies them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): log the traceback on the editor state-write failure The 848 fix answers 500 when the session state cannot be written, and logged that at error level -- but without exc_info, so the traceback never reached the log. test_web_error_detail.py guards exactly this: a handler returning 5xx must write an error-level record *with* the traceback and return the sanitized detail, because checking that merely something was logged is too weak. Caught by Core unit tests on the previous commit, not locally: the guard parses every module under web_interface/blueprints/api_v3 as one source, so it only fires once the whole package is read together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bdb9a94033 |
refactor(api-v3): split the 10,469-line blueprint into a package (#553)
* refactor(api-v3): split the 10,469-line blueprint into a package web_interface/blueprints/api_v3.py held 111 routes, 56 helpers and 181 functions in one module -- 9% of the core by line count and three times the next largest file. It becomes a package of nine route modules grouped by path segment, plus __init__.py for the shared imports, constants, Blueprint and helpers. Every route module decorates the SAME api_v3 Blueprint object, so endpoint names stay api_v3.<function>, the URL map is unchanged and app.py is untouched. Verified: 111 routes before, 111 after, byte-identical rules, endpoints and methods, and every endpoint still on the one blueprint. plugins 3,867 config 1,178 starlark 692 system 619 fonts 452 misc 398 wifi 361 display 326 backup 212 __init__ 1,787 (imports, constants, Blueprint, 56 helpers) Two things the URL-map check could not catch, both found by running the suite: 1. PROJECT_ROOT = Path(__file__).parent.parent.parent. Moving the code one directory deeper made that resolve to web_interface/ instead of the project root. Nothing failed at import; it surfaced as ~110 tests failing with 404s and "installation script not found", because every path built from it was one level too shallow. Now parents[3], and test_api_v3_url_map.py asserts PROJECT_ROOT/run.py exists so the next move cannot repeat it. 2. Module-attribute patching. Tests do monkeypatch.setattr(api_v3_module, "_BACKUP_EXPORT_DIR", ...) and a route module that binds such a name by value never sees the patch. The shared code therefore stays in __init__.py rather than moving to a _common submodule -- it has to live on the module the tests patch -- and the eleven names tests patch are read back through the package (_pkg.X) instead of bound by value. Those eleven were found by AST-scanning every setattr in the test tree, not by guessing; "time" is among them, used to drive a fake clock through the second-resolution credential-backup filenames. Test changes are confined to what genuinely moved: patch targets that now name the owning route module, imports of helpers, and six tests that scan the api_v3 source as a file and now read the package directory. Full suite: 4,278 passed, 68 skipped, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(api-v3): address CodeRabbit findings from the blueprint-split review Fixes to the api_v3 package split (PR #553), one per finding verified against the actual code: - __init__.py: _redact_credentials only blanked scalar values under a credential-named key; a bare list of secrets under such a key (e.g. tokens: ["a", "b"]) passed through untouched, since the list branch recursed with no memory that its key looked like a credential. Nested dicts still walk normally (a documented, tested behaviour -- a container like secrets: {api_key: ..., note: ...} is a section name, not a value to blank outright), but any value reached under a credential-shaped key is now actually blanked. - __init__.py: the OAuth helper script's raw stderr/stdout went to logger.error unredacted (CWE-532) right next to a comment claiming this was deliberate; the HTTP response already used the existing redact_text helper. Routed the log line through the same helper. - __init__.py / starlark.py: the standalone Starlark manifest fallback (used when the plugin instance isn't loaded) read-modified-wrote manifest.json with no lock, unlike StarlarkAppsPlugin._update_manifest_safe (plugin-repos/starlark-apps/manager.py), which already holds an flock for the same file when the plugin is loaded. Added _starlark_manifest_lock, mirroring that pattern, and wrapped every standalone read-modify-write call site in it. The app-config update route also wrote config.json and the manifest as two separate, non-transactional writes (a second, distinct finding at the same call site); config.json is now rolled back if the manifest write that follows it fails. - backup.py: restore options used bare bool() on values from the request, so {"restore_secrets": "false"} restored secrets anyway (bool("false") is True). Switched to the existing _coerce_to_bool helper already used for this exact purpose elsewhere in the package. - config.py: an automated import-rewrite mangled four user-facing validation strings and their neighbouring comments -- "Invalid start time" had become "Invalid start _pkg.time" (and likewise for "end time") in both the schedule and dim-schedule per-day validation paths. - display.py: `import _pkg.time as time_module` -- _pkg is a local alias for the package, not a real importable module, so this raised ModuleNotFoundError whenever a caller restarted an already-running display service via /display/on-demand/start, after the on-demand request was already written to cache. Fixed to `import time`. Audited the rest of the package for the same `_pkg.<module>` import mistake; every other `_pkg.` reference is a legitimate attribute read-through (`_pkg.time.time()`, `_pkg._get_starlark_plugin()`, ...), not a broken import statement. - fonts.py: validate_file_upload's max_size_mb parameter is silently unused by that helper (it only checks filename/extension) -- the font upload route saved arbitrarily large files as a result. Added the same seek-and-check pattern already used for the sibling .star upload. - wifi.py: two ad hoc, inconsistent bool coercions. POST /wifi/ap/auto-enable used bare bool(), so a JSON string "false" enabled it. POST /wifi/radio's enabled/force parsing recognized real bool and some strings but not int 1/0 (1 is True is False in Python). Factored one small _parse_bool_ish helper local to this file and used it at all three sites. Not changed: the "unknown/misspelled restore option keys default to True" half of the backup.py finding -- the file's own comment documents that a missing key deliberately means "restore everything," matching the already-existing JSON-parse-failure guard a few lines above it; only the bool-coercion defect was a real bug. Added or extended regression tests for every fix, following each area's existing test conventions. Full suite: 4328 passed, 62 skipped, 2 failed on both this branch and origin/main (missing tzdata package breaks two timezone-alias tests in test_onboarding_checklist.py, unrelated to this change) -- no new failures. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S3bPMESe2TfrGvbs1ef9c5 * fix(api-v3): reject unknown restore option keys CodeRabbit's review of the blueprint split (#553) asked that POST /backup/restore reject option keys outside RestoreOptions' known set. The follow-up commit fixed the bool("false")-is-True bug with _coerce_to_bool but never added the key check: a typo'd or renamed key (e.g. "restoreSecrets") is silently ignored by opts_dict.get(key, True), so the flag stays at its True default and secrets get restored despite the caller's request saying otherwise -- with no indication anything was wrong. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vmcwf5vMgYqdt8bJTZtiwb * fix(api-v3): address CodeRabbit findings on the blueprint split - _redact_credentials: blank scalar descendants of objects reached through a credential-owned list (e.g. tokens: [{"value": "secret"}]) regardless of field name -- the existing name-based walk only protected direct dict values under a credential key, not list items. - wifi.py: reject enabled/force/auto_enable_ap_mode values _parse_bool_ish can't recognize (400) instead of silently treating them as False, which could disable Wi-Fi or the radio itself. - Starlark manifest locking: lock a stable manifest.json.lock sidecar instead of manifest.json itself, in both the standalone route path (_starlark_manifest_lock) and the plugin path (StarlarkAppsPlugin._save_manifest / _update_manifest_safe). manifest.json is replaced by an atomic rename on every write, which swaps in a fresh inode; a lock held on the old inode does not exclude a second locker that opens the path afresh right after the rename and gets the new inode, so two writers could race despite each holding "a lock". A sidecar that no write ever touches always resolves to the same inode for every locker. Skipped as stale: the "serialize the complete manifest read-modify-write" finding at api_v3/__init__.py -- every standalone handler that calls _write_starlark_manifest is already wrapped in _starlark_manifest_lock() on this branch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(api-v3): re-check reconciliation findings by the reconciler's own rules Both CodeRabbit findings on the merge commit, verified against the code first. Major, plugins.py: the stale-findings filter derived its own notion of "in config" and "on disk", and both were looser than the reconciliation module's. set(load_config()) also contains system keys, the secrets-file keys load_config() merges in, and non-dict values; and any directory holding a manifest.json counted as installed even when that manifest does not parse. Either looseness clears a finding that is still true -- and a secrets key read as a plugin is the precise bug the filter exists to stop reporting, so reintroducing that asymmetry while re-checking was the wrong way round. The two extractions now live in state_reconciliation.py as config_plugin_ids() and disk_plugin_ids(), with ignored_config_keys() and secrets_top_level_keys() alongside. _get_config_state() and _get_disk_state() use them too, so there is one definition rather than two that can drift. _get_disk_state() re-reads each manifest for version/name after taking membership from the shared extractor; that costs one extra small read per plugin on a path that runs once per boot. Minor, the new test: the fixture assigned api_v3.config_manager and api_v3.plugin_manager directly. Those live on a module-level blueprint singleton, so the mocks leaked into every later test that imports api_v3 -- pointing at a tmp_path already deleted. Both now go through monkeypatch.setattr, which restores them. This is the same pollution class that made an earlier test in this session break seven unrelated ones, so it is worth getting right. Five cases added for the parity itself: a secrets key, a system key and a non-dict value must not clear an "installed but missing from config" finding, and neither an unparseable manifest nor a .standalone-backup- directory may count as installed. All five fail against the looser version. Linux CI on the preceding commit: Core unit tests, plugin harness, CodeQL and CodeRabbit all pass. Codacy reads action_required on every commit of this branch including the first, so it is pre-existing and not from this work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0ab95586fb |
fix(web): say when a system action failed for want of passwordless sudo (#560)
* fix(web): say when a system action failed for want of passwordless sudo
POSTing reboot_system to a Pi returns, in full:
{"message": "Action failed; see logs for details", "status": "error"}
The cause is that the web interface runs unprivileged, and its
systemctl/reboot/journalctl calls only work once
scripts/install/configure_web_sudo.sh has granted NOPASSWD. first_time_install.sh
never invokes that script and no user-facing doc mentions it, so on a fresh
device every privileged action fails -- start_display, stop_display, the
autostart toggles, reboot, and the log viewer.
That last one closes the loop: "see logs for details" is unreachable advice
when journalctl is refused for the same reason. This is exactly the failure
src/web_interface/error_handler.py's describe_exception() was written to break,
and /system/action's exception handler was still discarding the cause instead
of using the helper the module already imports.
Two changes, no behaviour change when things work:
- The exception path now returns 'details': describe_exception(e), matching how
the other handlers in this blueprint already report.
- A failure whose stderr or exception text is sudo refusing to prompt ("a
password is required", "no tty present", "a terminal is required") reports
what to do about it, naming configure_web_sudo.sh. Unrelated failures keep
the generic message and their stderr, so a missing unit is not blamed on
sudo.
Granting the sudo rights is left alone deliberately: auto-running a script that
hands out NOPASSWD is a security decision for the maintainer, not something to
slip into an installer. Making the refusal legible is the part that is
unambiguously an improvement.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(web): apply the sudo hint on the on-demand start_display path too
start_display with a mode builds its own response and returns before the shared
nonzero-result path, so a recognized sudo refusal there reported only "Failed to
start display" and said nothing about the passwordless sudo that refused it --
the exact gap the rest of this PR closes everywhere else.
Raised by CodeRabbit on #560 and verified against the code before fixing: the
branch at api_v3.py:2058 does return early past the shared handler.
Three regression cases: the on-demand branch reports the sudo cause, keeps its
"Display started" message on success, and does not blame an unrelated failure on
sudo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
39f27d285d |
fix(plugins): stop reconciliation inventing plugins and telling users to delete real config (#557)
On a device running four installed, configured, working plugins, the overview
banner read:
Stale plugin config entries found: football-scoreboard, odds-ticker, data,
ledmatrix-weather, starlark-apps. Remove them from config.json or reinstall
via the Plugin Store.
Every claim in that sentence was wrong, and following its advice would have
deleted 4.9KB of working league settings. Four separate defects combined.
1. Secrets keys became phantom plugins. load_config() merges
config_secrets.json into the config it returns, and the ignore list named
only 'github' and 'youtube'. A 'data' key in that file therefore read as a
plugin id and was reported as "in config but not on disk" forever. Read the
secrets file's own top-level keys instead of hardcoding two of them.
2. The auto-fix clobbered real config. The handler for "on disk but not in
config" assigned `config[plugin_id] = {'enabled': False}` unconditionally,
so whenever detection was wrong it replaced a plugin's entire configuration
with a stub. On the reported device it only failed to do so because the
write hit EACCES. Now it refuses to overwrite an entry that already exists.
3. The banner gave backwards advice. plugin_missing_in_config ("on disk, not in
config") and plugin_missing_on_disk ("in config, not on disk") are opposite
problems, and both were rendered as "stale config entries ... remove them
from config.json" -- which is correct for the second and destructive for the
first. They are now reported separately, each with the advice that fits.
4. A stale verdict was served indefinitely. The result is a snapshot written
once per run to a status file, and a run that fails to apply a fix also
declares it will not retry. A condition that had since resolved kept being
reported for hours. The status endpoint now re-checks stored findings
against current state, dropping only what it can prove stale and keeping
any kind it cannot re-verify.
The secrets-key lookup is deliberately fail-safe: an unreadable, absent,
malformed or non-path secrets location narrows the ignore set rather than
raising. An earlier revision let TypeError escape, which the broad handler in
_get_config_state() swallowed as "Error reading config state" -- emptying the
config state and making every downstream detection wrong. The existing
reconciliation tests caught it; there is now a regression test for it too.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
dcd6e39c96 |
fix(web): report real disk usage and MemAvailable on the live status stream (#558)
The SSE status stream sent 'disk_used_percent': 0 as a literal, so every
consumer of the live view showed 0% disk no matter how full the card was.
/api/v3/system/status computed it correctly; the stream that the dashboard
actually watches did not. On a Pi with a modest SD card that is the warning a
user most needs, and it was guaranteed to never appear.
The stream also omitted memory_available_mb. /api/v3/system/status carries it
with a comment spelling out why it matters: MemAvailable accounts for
reclaimable page cache, so it is what separates a board reading 70% "used" that
is fine from one reading 70% that is about to fail fork(). A 1GB Pi 3B+ can sit
at either. The number that predicts the failure was missing from the live view.
An unreadable disk now reports None rather than 0. The UI already renders null
as '--'; a confident 0 reads as "plenty of room", which is worse than a blank.
Metric collection moves to web_interface/system_metrics.py, with no Flask or app
imports. That is not cosmetic: importing web_interface.app constructs the Flask
application and a CacheManager, and the latter claims the cache directory with a
cleanup thread. The first version of these tests imported the generator directly
and broke test_cache_cleanup_thread_ownership ("one thread per directory") plus
four starlark route tests through that side effect. Reading a CPU percentage
should not boot a web application, and testing it should not either.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
ad5bc4b819 |
perf(sports): LRU-bound the decoded logo cache (#559)
SportsCore._logo_cache was a plain dict keyed by team abbreviation with no eviction. Its entries are not file bytes but decoded RGBA thumbnails sized to display*1.5 -- roughly 36KB on a 256x64 panel, more for wide wordmarks -- and assets/sports/ncaa_logos ships 307 of them. A plugin that walked a full league held the whole league resident: about 11-18MB per manager instance, and a league runs three (live/recent/upcoming) that each keep their own cache, so the same logos were duplicated across them. On the 1GB Pi 3B+ this was measured on, one board was sitting at 439MB resident with ~290MB available, so tens of megabytes of duplicated league logos is real money. Bounded to 64 entries, which holds a full "other games" cycle (on the order of 20 games, 40 teams) without thrashing while capping the cache well below a 307-team league. Eviction is LRU rather than clear-when-full, using the OrderedDict/popitem pattern the neighbouring caches in this codebase already use (_IMAGE_CACHE_MAX, _FIT_CACHE_MAX, _TEXT_WIDTH_CACHE_MAX). That ordering matters: the logos on screen right now are precisely the ones that must not be discarded, so a cache hit moves the entry to the end. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fb3b293ace |
fix(plugins): let a plugin ask to be polled faster while it has live content (#555)
* fix(plugins): let a plugin ask to be polled faster while it has live content
Reported: "the football plugin with live games only updates the live game in
progress if I restart the display."
The data path was never the problem. NFLLiveManager fetches ESPN with no cache,
SportsLive.update() refreshes current_game in place when the game IDs are
unchanged, and the scorebug redraws from the game dict every frame -- which is
why the reporter's logs look healthy.
The problem is cadence. _get_plugin_update_interval() read only the manifest's
static update_interval, football's manifest pins that to 60, and the plugin's
own live_update_interval (15s) was invisible to the scheduler. Measured on a rig
during the fourth quarter of the game in the report:
23:21:49 23:22:50 23:23:50 23:24:50 23:25:50 <- exactly 60s apart
A clock and score up to a minute stale during a two-minute drill reads as a
frozen panel, and a restart is the one moment it is ever current.
A single static number cannot say "every 15 seconds while a game is on, every 15
minutes in July", and only the plugin knows which is true. get_update_interval()
lets it say so per tick; returning None means "no opinion" and the existing
manifest/config resolution applies, so every plugin that predates this is
unaffected.
Requests are clamped to MIN_DYNAMIC_UPDATE_INTERVAL (5s): a plugin returning 0
would otherwise be re-entered on every tick of the render loop, busy-waiting
against its own API. A hook that raises or returns a non-number is ignored
rather than propagated -- a scheduler that fails on one plugin's bug stops
updating all the others.
Deliberately NOT changed: the manifest still beats config in the static path.
That looked like the obvious fix -- user config being silently ignored -- until
checking a real rig, where football and baseball both carry update_interval 3600
in config against a manifest 60, and weather 1800 against 60. Those values are
stale precisely because nothing has been honouring them; making config win would
have slowed three plugins by 60x, turning a one-minute lag into an hour. The
dynamic hook makes the flip unnecessary. There is a test pinning the current
precedence with that reasoning attached.
Full suite: 4,283 passed, 68 skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* test(plugins): drive the real scheduler, not just the interval resolver
test_plugin_dynamic_update_interval.py asserts that
_get_plugin_update_interval() returns the number the plugin asked for. That is
not the same claim as "the plugin gets updated more often", and the gap between
those two is exactly where the original bug lived: the plugin knew it wanted
15s, said so in live_update_interval, and nothing downstream acted on it.
So this ticks the real run_scheduled_updates() through a simulated hour and
counts dispatches. Against pre-fix core it reports "10 updates in 10 minutes of
a live game" -- the 60s manifest cadence, matching what was measured on a rig
during the reported game. Against the fix it reports ~40.
Also pins the regression that would be worse than the bug: an idle hour must
still be ~60 updates, not 240. Asking for the live interval year-round would
poll ESPN four times a minute all summer.
Scope note, since it is easy to over-read this fix: the *switch* display path
already refreshed the manager immediately before drawing, via
_try_manager_display() -> _ensure_manager_updated(), which honours the manager's
own 15s interval. So a switch-mode card was already <=15s stale at draw time
before this change. What this fixes is the background cadence, which is what
live-priority detection, Vegas content and scroll preparation all read.
Full suite: 4,288 passed, 68 skipped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(plugins): reject bool and -inf hook results in dynamic interval
get_update_interval() ran bool through float() (bool is an int subclass,
so True/False became 1.0/0.0) and only checked for +inf, not -inf. Both
cases landed on the MIN_DYNAMIC_UPDATE_INTERVAL floor by coincidence
instead of falling back to the static/manifest interval as invalid
input should.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Tst9cied2ri9bH4QRWa6H
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
28bc79566f |
fix(logo): remember a missing logo instead of re-warning every rotation (#548)
* fix(logo): remember a missing logo instead of re-warning every rotation load_logo() stat'd the path and logged a WARNING on every call, and the positive cache never covered it because a miss returns None and caches nothing. A file that is simply not there therefore produced one warning per rotation for as long as the process ran -- measured on a live rig at 114 lines in 24 hours for a single missing ticker icon, for a file nobody was going to add. Misses are now remembered for 10 minutes: warn once, then return None without touching the disk. Bounded rather than permanent because logo_downloader writes logos at runtime, so a file that appears later must still be picked up without a restart. Downloads through load_logo_with_download() clear the entry outright -- load_logo() consults the miss record before it stats the disk, so without that a freshly downloaded logo would stay invisible for the whole window. This is in the core rather than in ledmatrix-stocks, where it was found, so every plugin that goes through LogoHelper gets it. _cache_order stays a list. Swapping the pair for an OrderedDict would shave an O(n) scan per cache hit, but n is capped at cache_size (100 by default) and test_logo_helper.py pins the current structure; not worth the churn. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * fix(logo): make the miss TTL longer than the rotation it is meant to outlast Deployed the previous commit to a live rig and measured it: no change at all. "Logo not found for VOO" stayed at ~6 lines an hour, exactly the baseline. The TTL was 600s and the display rotation is ~618s, so every recheck expired just as the plugin came round again and the negative cache never once got to suppress a warning. The fix was correct in shape and useless in practice, which only measuring on the rig would show. An hour instead. That is safe because the TTL is not the main way an entry clears: load_logo_with_download() drops it the moment a download succeeds and clear_cache() drops all of them. The TTL only covers a file that appeared some other way -- someone copying one in by hand -- and waiting up to an hour for that, or restarting, is a fair trade for not re-warning about a file nobody is going to add. The general lesson is in the comment: a TTL has to be long relative to the loop that does the asking, not merely "a while". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |