mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-10 17:16:36 +00:00
2e7cab7d5358471b2c21c94a977328489ead5197
254
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2e7cab7d53 |
fix(display): on_demand_request_id has a class default for controllers built without __init__
_on_demand_state now publishes it, and test_state_stream_readers builds controllers with __new__. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3d79d7d1cd |
fix(web): a delivered on-demand start reads as starting until the display acts on it
On ledpi (three cold starts) the display acknowledged the start as its socket opened, then took ~5 s to act on it while Vegas built its first strip; the status routes meanwhile showed the display's own idle state, so a UI polling every 700 ms flashed idle. The display's on-demand state now names the request it answers (request_id). A delivered start keeps reading as status "starting" with delivered: true, in /display/on-demand/status and as on_demand_pending in /display/current-status, until the display publishes state for that request id (a display without the field: any state newer than the delivery), for at most DELIVERED_SHOWN_SECONDS (30 s). The display's startup state, which can be published after the acknowledgement, names no request and does not end it. Tests: stays starting against the startup idle state (no id, an older id); the matching active state and the matching error take over; an older display's newer state takes over; the 30 s cap; the display's state names its request. Mutation check: 11 mutants, 11 killed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
62ff14b972 |
feat(ipc)!: remove the cache-key mailboxes (control socket stage 5)
The control socket is now the only way the web interface sends the display a command. The display stops reading display_on_demand_request and plugin_error_clear_request, and the web interface stops writing them. - Display: no mailbox poll (MailboxWatch, the 1 s / 0.25 s cadence, _consume_on_demand_request, the deprecation log) and no persisted display_on_demand_processed_id guard; the error publisher reads no clear request. CacheManager.file_signature and MailboxWatch are removed. - A write to either retired key is dropped by CacheManager.save_cache and logged once per writer, naming the plugin from the call stack (or the request's plugin_id), with the API to move to. - Web: on-demand start with no display listening starts the service (when start_service) and sends the request again once the socket answers (45 s, 10 s for a running service without a socket yet); every other failure is a 503 (400 for invalid_args). Stop answers 503 when no display listens, unless stop_service. errors/clear answers 503 with a reason-specific message instead of writing a request; clear_pending is always false. src.ipc.client.should_fall_back is replaced by display_not_listening. - Kept: display_current_state, display_on_demand_state, plugin_runtime_snapshot and the heartbeat (read whenever the socket cannot answer), and display_on_demand_config (the display's resume record). Tests: mailbox-only tests removed (test_on_demand_mailbox.py, the mailbox cadence, file_signature and MailboxWatch tests); tests that injected requests through the mailbox now use the socket queue or a plugin's in-process request. The run-loop harness sends on-demand requests over its fake control socket, so four golden traces change: on-demand starts and stops land at the request instant instead of the next 0.25 s mailbox look (one frame fewer on the screen they end), and in vegas.json within one frame instead of 263 ms, which shifts the later 1 s-throttled WiFi-notice check by under a second. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
20bae8b609 |
fix(display): a failed on-demand request drops the session it ended (#779)
_set_on_demand_error ends any running session (_reset_on_demand_fields) but left its saved copy, display_on_demand_config, in the cache. A failed request that replaced a running session therefore made the next restart resume the session that had already ended. Clear the saved copy where every error path goes through, and drop the two restore-failed callers' own clears, which this now covers. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6c533d62af |
fix(plugins): web mode lookups use the modes the display registered (#668) (#769)
* fix(plugins): web mode lookups use the modes the display registered (#668) A plugin may compute its display modes from its config: soccer-scoreboard registers soccer_<league>_live/recent/upcoming for every custom_leagues entry, which no manifest can list ahead of time. The display always rotated them (_register_loaded_plugin prefers plugin.modes), but the web process reads plugins as files, so /display/modes, the on-demand dialog and on-demand/start with a mode and no plugin_id saw only manifests -- a custom league's mode was missing from every list and 404'd on lookup. - PluginStateManager.record_modes(): the controller records what it registered, on the loaded record (an unload or reload forgets it) - the runtime snapshot carries it per plugin as "modes" (bounded), and PluginRuntimeView.display_modes() reports it only while live - PluginCatalog takes a runtime_source; get_plugin_display_modes and find_plugin_for_mode prefer the live modes, falling back to the manifest when the display is stopped or has not loaded the plugin. The view is read at most once a second, so a listing is one read, not one per plugin. No manifest or plugin change needed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): call the runtime view's display_modes directly Codacy flagged the getattr/callable indirection as 'lookup is not callable'. The view is a PluginRuntimeView or None; anything else raises inside the existing try and falls back to the manifest. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(plugins): address review -- no manifest fallback for live plugins, keep mode names whole, send registered spelling Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
cb06124b42 |
fix: a Vegas static pause survives a non-numeric display duration; aliased store installs ask for a restart (#753)
* fix(vegas): a display duration that is not a number no longer cancels a static pause The Vegas static pause compared plugin.get_display_duration() with the clock. clock-simple, calendar and countdown return their display_duration setting straight from config.json, so a value saved as "20" or null reached that comparison as a string or None. The TypeError went to the pause's broad except, which ended the pause: the plugin flashed up and the scroll went straight on, at every one of its turns. inf held the pause until something interrupted it, and NaN, False, 0 or a negative number ended it at once. The pause now reads the duration the way the rotation has since #739, with the same helper, then the rotation's fallbacks: 30 s for anything that is not a number or a get_display_duration() that raises, 15 s for a number at or below zero. Logged once per plugin. test_vegas_static_mode.py's pauses used 0 to mean "no wait"; they now use 0.01. The helper moves from display_controller (_finite_seconds) to base_plugin (finite_seconds), unchanged: the coordinator cannot import from display_controller, which imports src.vegas_mode at module level, and a new src module would turn ledmatrix-plugins' min-core table check red until it was listed. base_plugin is already loaded whenever either one is. Tests: test/test_vegas_static_pause_duration.py, on a fake clock, including TestSameAsTheRotation, which runs every value through both the pause and the rotation's _get_display_duration/_resolve_durations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): a store install asks for a restart by the id it installed as POST /plugins/install decides restart_required from whether config.json already enables the plugin: the display loads a plugin when its enabled flag changes, so one already enabled (a reinstall, or a config carried over) keeps running the copy it loaded until a restart. The route read that flag under the registry id. Weather, Music, Stocks and Leaderboard install under the id their manifests declare (weather -> ledmatrix-weather), which is the config section's id, so reinstalling an enabled one never reported that a restart was needed. Both the queued and the direct path now look up the installed id once (#746's _installed_plugin_id) and use it for the plugin_id they answer with and for the enabled check. Tests: test/test_api_v3_install_restart_installed_id.py, through the Flask test client, both paths. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): type the static pause fallback on its own (mypy ratchet) _static_pause_duration assigned the fallback to `seconds`, which the except branch typed as float before finite_seconds() reassigned it as float | None. A separate `fallback` keeps both types exact; behaviour is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
e745ae8060 |
feat(display): cap malloc arenas in-process and malloc_trim between screens (#774)
* feat(display): cap malloc arenas in-process and malloc_trim between screens Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: malloc_tuning with ctypes mocked; add to the mypy ratchet Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): malloc arena cap and malloc_trim between screens Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): spacing Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0577c807eb |
feat(plugins): request_on_demand() / end_on_demand() -- plugins ask for the screen in-process (#768)
* feat(plugins): request_on_demand() / end_on_demand() -- plugins ask for the screen in-process Four plugins (birdnet-go, mqtt-notifications, on-air, pomodoro-timer) take the screen by writing the display_on_demand_request mailbox, which the display reads once a second while the control socket is up and which stage 5 removes. This is the in-process way in that stage needed. - BasePlugin.request_on_demand(mode=None, duration=None, pinned=False) and end_on_demand(), safe from any thread, go through PluginManager to DisplayController.submit_plugin_on_demand, which only queues (at most 32) and wakes the render thread through ControlServer.wake(). The render thread applies them in _drain_control_commands, after socket commands, through _handle_on_demand_request, so they land within a frame; without a socket, on the next pending-changes pass. - A plugin's stop ends only its own session; a mailbox stop still ends any. - Both answer the request id, or None with no display in the process (web interface, check_plugin.py), a full queue, or a mock manager -- a plugin's cue to write the mailbox, which the display still reads. - docs/PLUGIN_API_REFERENCE.md documents the hasattr pattern for plugins that must keep working on older cores; IPC_CONTROL_SOCKET.md and the CHANGELOG are updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): wire the on-demand handler only on a manager that has it Tests and the golden traces stand in simpler plugin managers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3866aa4519 |
feat(ipc): the control socket carries every web command; mailboxes are a fallback (stage 4) (#765)
- client: ControlError.sent says whether the display had the request; should_fall_back() allows a mailbox write only when it did not, or when the display is too old to know the command (upgrade case) - on-demand start/stop: a display that had the request and failed it is answered 503 (400 for invalid_args), no mailbox copy - errors.clear: new socket command, answered on the connection thread by a handler the display registers; applied and republished before the answer; plugin_error_clear_request only on fallback - display: on-demand mailbox looked at once a second while the socket is up (0.25 s without), read only when its file changed (one stat via CacheManager.file_signature / MailboxWatch); socket commands no longer touch the mailbox; a processed duplicate is consumed; writers logged once - error publisher: mailbox read only when changed; snapshot carries applied_clear_cutoff so an older mailbox request is not shown pending - docs and CHANGELOG (mailboxes kept for one release) Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
26cae3e5d6 |
refactor(display): run loop stage 3 -- ScreenRunner, PREEMPTED, OnDemand/Live/Rotation Sources (#762)
* refactor(display): run each screen through a ScreenRunner (run loop stage 3) The two frame loops, the make-up dwell and the dynamic-duration exit move out of DisplayController.run() into src/screen_runner.py. ScreenRunner paces with an injected FrameClock (production: this module's time, looked up per call so the golden harness's fake clock still drives it) and returns one Outcome whose ExitReason is DURATION, CYCLE_COMPLETE, EMPTY, ERROR, DISPLAY_FALSE, RELOAD or PREEMPTED. PREEMPTED replaces the re-checks that used to follow each frame loop and the make-up dwell (current_display_mode != active_mode, the schedule, a pending WiFi notice): the runner asks its host at named service points (FRAME, AFTER_LOOP, after_dwell, FINAL), and on PREEMPTED run() goes to the next pass without advancing the rotation, as each `continue` did. RELOAD is the one early end that still advances, as a reload always did. Each service point reads the WiFi notice file exactly when the loop did (NoticeRead), because the read is throttled and deletes expired files. The host answers still use the old checks; the following commits move them to the Arbiter. _screen_preempted is gone (folded into the FRAME check); _wait_frame_interval returns the preempting plan instead of a bool. The frame pacing (8 ms deadline, 1 ms minimum yield, 1 Hz wait with socket wake) is the same code, moved. Golden traces unchanged. A capture of every harness run (all 67, with every sleep, display() call, wifi read, live scan, publish and dwell logged) is identical to origin/main except for throttled WiFi reads that returned the cached answer (no side effect) after a notice preempted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): on-demand is an Arbiter Source (run loop stage 3) Arbiter.decide() now answers for an active on-demand session itself (Source.ON_DEMAND) instead of returning LEGACY: - ArbiterState gains the session: its mode list, index, expiry and pin, plus current_mode, snapshotted from the controller's fields by _arbiter_state(). The controller's attributes stay the record that the web UI, the control socket and the cache read. - The OnDemand plan is the session's current mode (an index past the end of a shortened list starts again at 0), with what is left of a timed session at `now` as max_duration and the expiry as deadline. A session with no modes left is a plan with no mode; the controller ends it and shows the rotation's mode, as _resolve_active_mode did. - on_demand_bound() is _clamp_to_on_demand made pure. It is still applied after the first frame, with the clock read there. - ArbiterState.next_on_demand() is the step _advance_on_demand takes. run() asks decide() for the screen at the point it used to call _resolve_active_mode (after any Vegas iteration, so a session that started mid-iteration still shows next), and _take_plan() applies it. Golden traces and the 67-run capture identical to origin/main. Adds TestOnDemand and the bound table to test_display_arbiter.py; the stage-2 table's on-demand rows now name ON_DEMAND instead of LEGACY. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): live priority is an Arbiter Source (run loop stage 3) The live-priority step of run() (step 7) and the live checks around the Vegas iteration become the Arbiter's Live Source: - ArbiterInputs gains live_modes (the scan, None where run() made none), vegas_enabled, vegas_live_in_ticker and vegas_yielded. ArbiterState gains the rotation and its index, the live resume point and the "takeover not shown yet" flag. - Live picks the next live mode round-robin (live_pick, now also what _check_live_priority returns), not advancing past a mid-screen takeover that has not shown. It outranks Vegas unless the ticker keeps live content, in which case it has no say at all, as before. With nothing live, a plan below it carries ends_live and the interrupted rotation resumes. - ArbiterState.claim_live/release_live are _apply_live_priority's bookkeeping made pure; _apply_live_priority applies them. run() reads the inputs below the WiFi notice where it always did (_arbiter_inputs_below_wifi: the Vegas check, then the scan), asks decide() once more, and _take_plan applies the claim or the resume. The Vegas iteration moves to _run_vegas_iteration, which re-decides with vegas_yielded after a yield, so a game that stopped the ticker or an on-demand session that started mid-iteration still shows next. LEGACY now means Vegas or the rotation. One redundant call is gone: a Vegas pass scanned the live plugins twice at the same instant (step 7, then step 8's "is anything live?"); it scans once. Golden traces unchanged. The 67-run capture is identical to origin/main once that duplicate scan and _apply_live_priority(None) calls that changed nothing are left out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): the rotation is an Arbiter Source; LEGACY means Vegas (run loop stage 3) decide() now names the screen for every pass: the Rotation Source (Source.ROTATION) answers with the rotation's current mode, after the resume when live priority just ended. LEGACY is left meaning only Vegas, whose iteration is still run()'s own code until stage 4. Once this pass's iteration has yielded (vegas_yielded), Vegas passes and the screen it fell through to is decided like any other. The rotation's mode is state.current_mode rather than rotation[rotation_index]: they agree except where something moved the panel off the list and the rotation carries on from there (a live mode no entry names, or None after a session ended with nothing to resume to), and run() always showed current_display_mode. ArbiterState.after(outcome) is _advance_after_screen's step: an on-demand session moves to its next mode; otherwise the rotation advances unless the mode just shown is still live. The Outcome carries the two facts only the controller can see at the end of the screen (on_demand_active, the live hold from _still_live). Ending a session with no modes left stays in the controller, because it is not pure. Golden traces unchanged; the 67-run capture is identical to origin/main with the same two exclusions as the previous commit. Adds the Vegas / Rotation table and TestAfter; the stage-2 rows that said LEGACY for "live, Vegas or rotation" now say ROTATION. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): one decide() call at each of the runner's service points (run loop stage 3) The three mid-screen checks the frame loops made one after another -- _check_live_takeover, then _screen_preempted with _wifi_notice_pending in it -- become one call: Arbiter.decide(state, inputs, now, running=plan). It returns `running` itself while the screen holds, else the plan that ends it, from these rules in the order the loops checked them: 1. Live: a game went live while a non-live screen runs. First because it is the one preemption that changes the state (the rotation moves to the live mode and remembers where it was), and it is still claimed when a WiFi notice is pending too; the next pass shows the notice, then the game, as before. 2. The panel's mode moved under the screen (on-demand started, ended or changed; the rotation was rebuilt). 3. The schedule turned the panel off. 4. A WiFi notice arrived (on-demand outranks it; compared with expiry). 5. A plugin reload is waiting (between frames only). Each rule is gated by plan.preemptible_by: every screen may be preempted by the gate, OnDemand, Wifi, Live, Rotation and a reload, except that a live screen leaves Live out. A follower and Vegas never preempt mid-screen. The pure helper live_takeover() is the Live rule, shared with the dwell sleep's _check_live_takeover. The controller only gathers and applies. _screen_service applies pending changes and makes the live scan when one is due (_scan_for_takeover: the same throttle and gates as before); _screen_check reads the WiFi notice exactly where the loop did (the read is throttled and deletes an expired file, so an extra read would move both), calls decide() once, and claims a live takeover. _screen_preempted is gone; _check_live_takeover and _wifi_notice_pending remain for the dwell sleep and the Vegas yield path, built on the same rules. Golden traces unchanged. The 67-run capture is identical to origin/main (with the earlier two exclusions) except for one event: in the 125 Hz loop the live scan still runs before the frame's sleep, but the claim is now made by the service point after it, so the "live" state change is logged 8 ms later (test_live_game_cuts_a_scrolling_screen_short: 8.064 -> 8.072). The screen still ends at the same frame (8.072) and every frame, sleep and pass is unchanged. Adds the mid-screen table (24 rows), live_takeover's table, and test/test_screen_runner.py (the runner on a scripted host, plus the controller's service point: which reads it makes at which checkpoint). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs, tests: run loop stage 3 -- doc, changelog, mutation survivors docs/RUN_LOOP_REDESIGN.md describes run() as it now is (two decide() calls per pass, the runner and its service points, the state snapshot and its transitions), records what stage 3 shipped and how it was checked, and adds one open "may be wrong" behaviour the mutation run surfaced: a Vegas iteration stopped for a sync follower falls through to a full rotation screen before the follower gets the panel (pinned by test_vegas_yielding_to_a_follower_shows_a_rotation_screen_first; passes on origin/main too). docs/IPC_CONTROL_SOCKET.md no longer names _screen_preempted. CHANGELOG entry under Unreleased. A mutation run broke 46 moved or new pieces once each (OnDemand, Live, Vegas/Rotation, after(), each mid-screen rule, the runner's pacing, exits and service points, the controller's gathering and claims). Three survived and get a test here: - the after-loop service point not reading the WiFi notice: the completed-loop checkpoint gets its own name, and a run-loop test has a notice pending when a later frame comes back empty; - the Vegas yield path not marking vegas_yielded: the follower test above; - _take_plan not writing back a reset on-demand index: a controller test with an index past a shortened list. All 46 now fail at least one test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
caec9f5bf5 |
feat(display): report a scrolling screen held by its plugin's update() (#758)
While a plugin's update() runs it holds the plugin's lock and its frames are skipped -- on a scroller, a frozen strip -- with nothing logged. The high-FPS loop now times each run of skipped frames (report_hold=True); one of 250 ms or more logs 'Display of X held N ms by its update()' (rate-limited per plugin) and is recorded as a 'display hold' busy skip, which never touches the circuit breaker. The 1 Hz loop is left out: one skipped frame there measures the loop interval on a screen that did not visibly freeze (seen on ledpi as ~1000 ms reports on clock-simple and switch-mode football). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
e09e251553 |
fix(display): a frame the preview throttle skipped still reaches the snapshot (#752)
A screen that draws its card once and holds it no longer leaves the web preview black: DisplayManager remembers a changed frame the snapshot throttle skipped, and the render loop writes it (write_owed_snapshot(), called from _display_once) once the interval has passed. A failed owed write stays owed and is retried. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7026eeb156 |
fix(on-demand): show a named live mode; end a session that cannot resume (#748)
On-demand: a mode requested by name is shown first (even a quiet live mode); a session that can't resume after a restart, or whose plugin system failed to start, ends with status restore-failed instead of staying dead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
064b9c9912 |
fix(display): a non-numeric plugin duration no longer stops the display; narrow scroll strips no longer raise (#739)
* fix(display): a plugin duration that is not a number no longer stops the display DisplayController._get_display_duration returned whatever the plugin's get_display_duration() gave back. clock-simple, calendar and countdown return their display_duration setting straight from config.json, so a value saved as "20" or null reached _resolve_durations as a string or None, and its `<= 0` check raised a TypeError. Nothing in the loop caught it: run()'s outer handler logged "Unexpected error in display controller" and cleanup() ended the service when that plugin's screen came up, and systemd restarted it into the same crash. The plugin's answer is now read as seconds: a finite number or a numeric string is used (as BasePlugin.get_display_duration already accepts), a number at or below zero still goes to _resolve_durations' 15 s rule, and anything else -- None, a non-numeric string, a bool, NaN, infinity, or a get_display_duration() that raises -- gets the 30 s a mode without a plugin gets. The warning is logged once per plugin, not at every screen. Tests: test/test_display_duration_not_a_number.py, including the real run() on the run-loop harness, which returned at t=30 before the fix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(scroll): a strip narrower than the panel no longer raises on every frame ScrollHelper._get_visible_portion_integer handled a frame that runs off the end of the strip by copying the strip's tail and then the rest of the frame from its head, which assumed the head was at least that wide. For a strip narrower than the panel that raised "could not broadcast input array" at every position, so get_visible_portion() never returned a frame and the caller logged a traceback each frame. Vegas composes such a strip (lead_in_width defaults to 0) when its content is narrower than the chain. A wrapping frame is now taken column by column modulo the strip's width (np.take, mode='wrap', into the reused frame buffer): the tail then the head, as before, and a narrow strip repeated across the panel. The same path takes a position before the start of the strip, whose [-n:m] slice was empty and made frombytes raise; the integer and sub-pixel fast paths now leave a negative start to it. A zero-width strip is still a black frame. Tests: test/test_scroll_helper_narrow_strip.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
0d179fdf12 |
perf(display): throttle the per-frame update tick; check strips without building them (#731)
The frame loops and the dwell sleep run PluginManager.run_scheduled_updates() at most every 0.25 s instead of after every frame (the top of each loop pass still ticks unthrottled), and SportsScrollDisplay and the sync follower ask ScrollHelper.has_strip() instead of building cached_image just to see whether a strip exists. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
515248b34e |
feat(ipc): control socket stage 3 - a state stream replaces polled cache keys (#735)
Adds state.get / state.subscribe to the display's control socket (StateHub in src/ipc/server.py). The web interface holds one subscription per process (web_interface/display_state.py) and reads current-status, on-demand status, plugin runtime and /health's display_loop from it, falling back to the cache keys and heartbeat file. While the socket serves readers, display_current_state and plugin_runtime_snapshot are written less often (about 1.5 instead of 5 cache writes a minute for 15 s screens). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a5ec645d25 |
refactor(display): run() stage 2 - an Arbiter decides the scheduled-off blank, follower and WiFi notice, no behaviour change (#733)
run() stage 2: a pure Arbiter.decide() (src/display_arbiter.py) chooses scheduled-off, follower and WiFi notices; everything else takes the existing path. Golden traces byte-identical. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c14002edc3 |
fix(status): runtime status agrees with the heartbeat; current-status republishes on wake (#726)
The runtime status snapshot now agrees with the display heartbeat, and current-status is republished when the display wakes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8136a2d525 |
fix(ipc): a plugin reload no longer freezes the panel during Vegas (#723)
A plugin.reload no longer freezes the panel during Vegas: the old instance is torn down and the new one loaded off the render thread (frame gap 3017 ms -> 9 ms in the ledpi reproduction). A failed or timed-out teardown stops the reload with a restart hint instead of loading over stale modules or tearing down an instance still in use. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
85be4bf25d |
fix(display): end the scroll before the schedule-off blank and WiFi notice (#721)
The schedule-off blank and the WiFi notice are drawn by the display controller, not dispatched to a plugin, so #716's handover never reached them. Drawn while the last scroll's state was still set, the blank went out with the ticker's lagging rows on a scan-compensated panel and stayed up for its 60 s dwell, and the notice's redraws (which #712 now shows over a running scroller or Vegas) were timed as 0.5-1 s freezes and logged as a mid-scroll Render stall. The controller now calls set_scrolling_state(False) before drawing either; a scroller that resumes sets the state again on its next frame. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
746dcfcadb |
feat(ipc): control socket stage 2 - wake the render thread, brightness.set, plugin.reload (#720)
Control socket stage 2: the render thread wakes for queued commands (static screens ~1 ms, Vegas within one frame), brightness.set, and plugin.reload after a store update, with mailbox/restart fallbacks. Rig checks listed in the PR body are still to run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
1ecf3aba03 |
fix(display): narrow the WiFi status before reading expires_at (#715)
Narrow the WiFi status before reading expires_at (mypy Optional index; behaviour unchanged). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
34be83d595 |
fix(display): end the scroll state at scroller-to-static handovers (#716)
A static plugin screen that follows a scroller no longer starts with the ticker's lagging rows on scan-compensated panels, and the 1 Hz loop's second frame is no longer recorded as a ~1 s mid-scroll freeze / Render stall. The display controller calls DisplayManager.end_scroll_for_static_screen() before a static screen's first display() (clears the scan history; _scan_segments passes its frames through in one swap) and set_scrolling_state(False) after it; the scroller's hold stays until then, so late-frame counts are unchanged. A screen's first frame is tagged 'handover': gaps of 250 ms or more before it go to the additive handover_freezes (frame_soak prints 'Handover gaps'), not freezes. The display thread is named display-<plugin id>. The WiFi notice and the schedule-off blank are not covered yet (docs list them as a follow-up). ledpi A B B A soak (20 min each, --preview): main 0.118% / 0.113% late with 6 / 3 freezes; with this and #717 0.107% / 0.104% late, 0 freezes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
41db488c73 |
fix(display): schedule windows end at the end time; on-demand ending in off hours blanks at once (#714)
Schedule and dim windows are half-open [start, end): on from the start time, off at exactly the end time, whatever second the check runs. An on-demand session that ends in scheduled-off hours (expiry or stop) clears the once-a-minute schedule gate, so the panel blanks within about a second. Golden: schedule; two test_display_pending_changes.py tests now say end_time 23:00. Merged with #712 and #713: with all three in, docs/RUN_LOOP_REDESIGN.md's 'may be wrong' list is empty, so that section now records that all six items are fixed and by which PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
77e0ea91ae |
fix(display): live games take over within a second, and straight from Vegas (#713)
A game that goes live takes over within about a second (_check_live_takeover in the frame loops and the dwell sleep, throttled to 1 s, never during on-demand, scheduled-off, live_in_ticker or an already-live screen); an interrupted Vegas iteration switches straight to the game; has_live_content() is asked once per plugin per scan. Goldens: live_priority, vegas. Merged with #712: after an interrupted Vegas iteration the WiFi-notice check runs before the live switch (WiFi outranks live). Adds test/test_run_loop_wifi_and_live.py, pinning that a notice and a game arriving during the same screen (1 Hz, 125 Hz, Vegas) show the notice first, then the game, and neither while scheduled off. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6c6394a1a7 |
fix(display): a WiFi notice preempts the current screen and Vegas yields to it (#712)
A WiFi notice preempts the current screen within about a second (frame loops, post-loop check, make-up dwell), and an interrupted Vegas iteration that yielded for a notice ends the pass so the notice shows next. Goldens: wifi_notice, vegas. First of three run-loop fixes (#712, #713, #714), pre-tested together. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
4be53d048b |
feat(fetch): shared fetch service, stage 1 (pooling, merging, host budgets, counters) (#702)
Core's own HTTP fetch paths (APIHelper, fetch_espn_scoreboard and its date chunks, BackgroundDataService, BaseOddsManager.get_odds) go through one service in src/common/fetch_service.py: shared connection pools per retry policy, merged identical in-flight GETs, per-host token-bucket budgets (fetch_service.rate_limits), and per-plugin request counters published to GET /api/v3/plugins/fetch-stats. Return values, exceptions, cache keys, TTLs and retry policies are unchanged. Core-internal in this release; plugins should not import it directly yet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9edeb6da14 |
fix(display): a raising display() counts as a circuit-breaker failure (#707)
The first-frame dispatch (_dispatch_first_frame) now asks PluginExecutor.execute_display() to re-raise (raise_errors=True) and records a raise inside the executor as a breaker failure, with the original exception as last_error, instead of a success. The screen is still an empty pass and rotation is unchanged; a hung display() is still recorded once, as a hang. The run-loop golden trace plugin_error.json is regenerated (crashy now records health failures and is skipped by the breaker), and behaviour 7 is dropped from docs/RUN_LOOP_REDESIGN.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
a21650e746 |
refactor(display): run() stage 1 - golden traces and extracted helpers, no behaviour change (#704)
Adds golden trace tests for DisplayController.run() (test/test_run_loop_golden.py on a fake clock with fake plugins, 15 scenarios, fixtures in test/fixtures/run_loop_golden/) and moves twelve blocks of run() into named helpers (_dispatch_first_frame, _resolve_durations, _resolve_active_mode, _needs_high_fps, _advance_after_screen and others) with the traces identical before and after. docs/RUN_LOOP_REDESIGN.md describes the target structure. Hardware-checked on hdpi: Vegas late-frame rate unchanged in an ABBA A/B. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
695ff92009 |
feat(ipc): display control socket, stage 1 - on-demand with acks (#706)
The display serves a control socket (/run/ledmatrix/control.sock) carrying versioned JSON commands, one per line, each answered. Stage 1 covers on-demand start, stop and status; commands are queued on the socket thread and applied on the render thread through the mailbox's own handler, and the web interface falls back to the file mailbox when the socket is unavailable. Protocol and security model: docs/IPC_CONTROL_SOCKET.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
74696d2108 |
fix(display): routine rotation log lines are DEBUG (30% fewer journal lines) (#693)
"Processing mode", "display() returned False" and "No content to display" repeated what "Switching to mode" already logs on every rotation; they are now DEBUG. On ledpi this cut the display's journal lines by about 30%; the measured SD-write saving is small (within noise). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7f06cc9c3b |
feat(vegas): keep live games in the ticker by default (#699)
* perf(timing): say which render-thread work a late frame followed The soak already says how often a moving frame reached the panel late, but not what the render thread was doing just before it. Vegas does two kinds of work there between frames -- building its strip (compose, extend) and, with live elements, patching changed pixels into it -- and deciding whether either is affordable needs their own numbers. - FrameTimingRecorder.note_op(kind, nbytes) tags the next presented frame. Totals gain op_frames, late_op_frames, op_freezes and op_bytes per kind; aggregate() still takes frames without ops. The file schema is unchanged. - Vegas tags compose and every strip extension (with the bytes it copied). - frame_soak prints an "after work" table: frames, late %, freezes and MB moved per kind, only when something tagged its work. - render_bench gains --strip-screens (Vegas-sized strips), --patch-bytes / --patch-every / --patch-where (in-place column writes, as a live element update does) and --extend-every-screens / --extend-width (append + trim on a fixed cadence that holds the strip's width). No runtime behaviour changes: this is the measurement gate for live Vegas elements. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): note the frame-op attribution and bench modes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * perf(scroll): build the strip's PIL image only when something reads it Every Vegas strip extension rebuilt ScrollHelper.cached_image from cached_array in full, twice (append, then trim), on the render thread: Image.fromarray is 1.7ms for an 8,000px strip and 3.8ms for 20,000px on a Pi 4 (measured on ledpi), about two thirds of an extension's render-thread cost. Nothing on the frame path reads the image's pixels; every frame is cut from the array. cached_image is now a property. append_content and drop_scrolled_prefix defer it; the first read builds it from the array it started with and keeps it only if the strip has not changed meanwhile, so a sync push racing an extension cannot leave a stale image cached. Assigning cached_image stores exactly what was assigned, as before. has_strip() says whether there is a strip without building its image; the helper's frame path, Vegas and the adapter's scroll-cache invalidation use it. The strip is also no longer held in memory twice. In Vegas the image is now built only by a multi-display sync push. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): live elements -- a plugin API for content that changes while it scrolls Vegas bakes each plugin's pictures into one strip, so a card already on its way across the panel keeps what it showed when it was drawn. This adds the API and bookkeeping for content that can be updated in place; the worker that redraws and swaps it follows separately. No shipped plugin implements the hook yet, so nothing changes for users. Plugin API (core 3.8.0), all no-ops by default: - BasePlugin.get_vegas_elements() -> [VegasElement(key, image, version, live, refresh_hz)]: named, fixed-width pieces of Vegas content. - BasePlugin.redraw_vegas_element(key, width, height, at): a lock-free redraw for content that changes with time. - BasePlugin.notify_vegas_data_changed(): data that lands outside update(). - src/plugin_system/vegas_elements.py (VegasElement, re-exported from base_plugin). Core: - PluginAdapter asks a plugin that implements the hook for elements on the background fetch only (under its lock, on its own canvas); every other path keeps get_vegas_content(). Live elements are pinned (padded with content_padding, never trimmed), tagged with their key, digest and data epoch in Image.info so the existing cache and group plumbing carry them unchanged, and untagged if a width budget crops them. - RenderPipeline records where each live element lands (ElementRecord), in absolute strip columns a trim does not move; the block-start arithmetic is shared with the STATIC markers. - PluginManager update listeners (add/remove_update_listener, notify_data_changed): told the moment update() completes, not at the next ~4s Vegas poll. The coordinator uses one to move each plugin's data epoch on. - vegas_scroll.live_refresh (kill switch), live_max_hz, live_min_interval, live_lead_screens; per-plugin core-owned vegas_live. Live elements are off under multi-display sync, in swap mode and with offscreen_prefetch off. - scripts/check_plugin.py checks the element contract (src/plugin_system/testing/vegas.py); test/fixtures/plugins/vegas-live-stub is a working example. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): live elements update in place while they scroll One background worker (src/vegas_mode/live_worker.py) redraws a plugin's live elements when its data epoch moves on (update listener) or on their refresh_hz, nearest the screen first, and hands changed pixels lock-free to the render thread, which copies them into the strip between frames (RenderPipeline.apply_live_patches, ScrollHelper.patch_columns): at most four patches or two screens of bytes a frame, no drawing or locks there. The worker takes over group prefetch once a live element is placed, runs inside the render gate, and is supervised. Update tick 1s while live elements exist. Web UI switch for live_refresh. OFFSCREEN_RENDERING.md describes what was built and why SegmentStrip was not needed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(sports): live Vegas cards for the scoreboards (shared layer) One live element per game, drawn only when what the card shows changes, so a score changes on a card already crossing the panel. The shared part, so each scoreboard adopts it in a few lines: - src/common/sports_vegas.py: game_key, game_fingerprint (the whole game dict, frozen: no drawn field can be missed), dedupe_games, VegasCardCache, StickyOdds (odds a live poll left out stay drawn), finished_games / with_finished_games (a game that just went final keeps its card, after its league's live games; one a heuristic only judged over keeps its live state, so a tied end of regulation never shows FINAL early). - SportsScrollDisplay.make_vegas_renderer() is the override point; build_vegas_elements() and SportsScrollDisplayManager .get_vegas_elements_for() do the rest. A card's version includes its teams' ranks, which the renderer draws from the rankings cache. - SportsLiveSharedMixin._record_finished_game() / finished_games_snapshot(): held for FINISHED_GAME_TTL after it leaves the live list. A sport that does not implement make_vegas_renderer keeps its ordinary Vegas content, so no scoreboard changes until it opts in. scripts/render_plugin.py --vegas renders a plugin's Vegas block as the ticker lays it out, and --timeline stacks it at successive moments as the ticker would update it in place; the join is now render_pipeline.join_plugin_rows(). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(vegas): keep live games in the ticker by default display.vegas_scroll.live_in_ticker now defaults to true: through a live game the marquee keeps running and the live scoreboard takes extra turns in it -- its cards updating in place while they scroll -- instead of the ticker giving way to the full-screen scoreboard. The new default would reach nobody on its own: every existing config holds an explicit false copied from the template (there was no control for it), and the template merge only adds missing keys. ConfigManager therefore turns a stored false on once, with a backup, and records live_in_ticker_migrated so a false chosen afterwards stays. The marker is never in the template. A "Keep live games in the ticker" checkbox under Vegas mode sets it. Tests that pin the full-screen takeover now say live_in_ticker=false. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(sports): a default _determine_game_type on SportsScrollDisplay render_vegas_card looked the method up with getattr and a None default, which static analysis (Codacy) reports as calling something that may not be callable. The base class now has the default -- the card type from the game's state -- and the plugins that define their own override it as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: review follow-ups on the shared live-card layer - The reused Vegas renderer always gets the current rankings, empty included, so ranks cleared since are not kept drawn. - render_plugin.py: --timeline refuses --no-live (a timeline shows live elements changing), --timeline/--no-live need --vegas, and the Vegas paths create the output's directory like the display path does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
64c7289593 |
feat(display): systemd watchdog and heartbeat for a frozen render loop (#687)
If the render loop gets stuck inside a plugin's display(), ledmatrix.service stays active and the panel stays frozen. This adds a way to detect that. - src/display_watchdog.py (standard library only) sends sd_notify over $NOTIFY_SOCKET and writes /run/ledmatrix/display-heartbeat.json. Only the render thread counts: beats from other threads are ignored. - ledmatrix.service: WatchdogSec=120, NotifyAccess=main, RuntimeDirectory=ledmatrix (0755), RestartSteps=4 and RestartMaxDelaySec=2min. It stays Type=simple. run.py widens the watchdog to 15 min for start-up, and load_plugin() does the same on the render thread. The loop arms after its first frame. - /api/v3/health adds checks.display_loop: running, stalled (no heartbeat for over 60s, which makes the status degraded) or not_reported. With web login on, a caller who is not logged in still gets only healthy/degraded, and a stall degrades that answer. - The update verifier requires a fresh heartbeat from the restarted display when the display it replaced was writing one. A frozen panel is rolled back. - Existing installs get the systemd watchdog only after install_service.sh is re-run. The heartbeat works right away. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b09434a418 |
refactor(plugins): the display publishes plugin runtime state; retire plugin_state.json (#690)
Stage 2 of the web plugin catalog, after #688. - The display publishes a plugin runtime snapshot (plugin_runtime.py) to the shared cache: per plugin loaded, lifecycle state, a short redacted error summary, the version it loaded and when, plus published_at / stale_after / running. Written on change (throttled to 10 s; the RUNNING/ENABLED flip of an ordinary update is not a change) and once a minute otherwise; cleanup() publishes running: false. - The web reads it back and restores loaded / state / error_info in /api/v3/plugins/installed (plus loaded_version, loaded_at and data.runtime). Only a live snapshot counts; stale, stopped or missing answers null and says which. - data/plugin_state.json is retired: every reader and writer moved to config + disk (desired) or the snapshot (observed). Nothing in it was non-derivable, so nothing is migrated and an existing file is left unread. The web-side PluginStateManager (state_manager.py) is removed; the display's plugin_state.PluginStateManager is the only state machine. - StateReconciliation compares config + disk with the snapshot, reporting enabled-but-not-loaded and older-version-loaded as no_action findings. - Backups list installed manifests with enabled from config.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c0d97e4867 |
fix(plugins): one hung plugin no longer stops every plugin from updating (#677)
The update worker no longer blocks forever on a plugin whose display() never returns. It waits at most PLUGIN_LOCK_TIMEOUT (5s) for a plugin's lock, then skips that plugin's update (a report-only "busy skip" in health) and keeps updating every other plugin. display() frames are timed (slow calls logged and counted; calls past the executor timeout recorded as hangs), a hung update() is recorded, and on_config_change() now runs under the plugin lock or is deferred to the worker. The plugin-facing API is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c3a7a110c4 |
fix(display): on-demand loads a disabled plugin live instead of failing (#678)
* fix(web): on-demand no longer restarts a running display service POST /display/on-demand/start treated start_service (default true, sent by "Preview on display", the on-demand dialog and the MQTT bridge) as "restart": with the service running it ran systemctl stop, slept 1.5s and started it again. Every request cold-started the display process -- every plugin reloaded, panel blank -- to deliver a request the running process already reads from the cache mailbox every ON_DEMAND_POLL_INTERVAL (0.25s), including mid-dwell, mid-screen and mid-Vegas. The restart bought nothing: startup only restores a session the display saved itself (display_on_demand_config), so the new request arrived through the same mailbox either way. start_service now means "start it if it is not running". The stop route coerces stop_service to a boolean so "false" no longer stops the service. test_api_v3_on_demand_restart.py pinned the old restart path; it now pins the replacement. Docs updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): on-demand loads a disabled plugin live instead of failing The display process only loads enabled plugins, so an on-demand request for a disabled one -- "Preview on display" offers it on every config page, with a note that the plugin will be enabled for the preview -- failed with invalid-mode. Nothing enabled it short of a restart, and the on-demand route no longer restarts the service. _activate_on_demand now loads an installed-but-not-running plugin through the live-enable path (load_plugin + _register_loaded_plugin), with a new load_plugin(force_enabled=True) so the instance runs enabled while config.json keeps saying disabled. The plugin is tracked in _on_demand_loaded_plugins, and the main loop unloads it through _unregister_plugin once on-demand moves off it (stop, expiry, another request, or a failed request that ends the session) -- right after its own poll, where no display() is on the stack. A failed load publishes status error with load-failed. A plugin enabled during the session stays loaded. A session restored after a restart uses the same tracking instead of setting enabled in the config dict config_manager caches, so its plugin is unloaded when the session ends rather than staying loaded until the next restart. Ending a session no longer resumes the rotation onto a plugin that is about to be unloaded, which a restored session did. Also: a stop sent while on-demand is inactive clears a failed request's error, instead of /display/on-demand/status reporting status: error until the state aged out. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6f45ff5e63 |
fix(display): thread-safety for deferred updates, BDF faces and follower image; one refresh default (#652)
- DisplayManager.defer_update()/process_deferred_updates(): one lock around every queue mutation (appends from the update thread were lost to the render thread's filter/slice reassignments); callables run outside it. - FontManager and element_style no longer cache BDF freetype.Face objects process-wide (load_bdf_face caches them per thread); element_style's LRU is locked against get/move_to_end vs eviction races. - limit_refresh_rate_hz default is one constant, DEFAULT_REFRESH_LIMIT_HZ = 100 (the template's), for the library options, refresh_hz, the matrix guard, Vegas and scroll_config. Previously a missing key capped the panel at 90 while pacing assumed 100. - Sync follower: the TCP thread queues the leader's scroll image; the render thread swaps image/array/width in between frames. - update_display() error log rate-limited (traceback first, then once a minute with a count); swallowed DisplayController exceptions log at DEBUG. - Root display_controller.py runs run.py via runpy. - stream_manager: correct the RLock release comments; merge duplicate if. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
224847cebc |
fix(display): stop the run loop spinning when no mode has anything to show (#649)
* fix(display): stop the run loop spinning when no mode has anything to show A mode whose display() reports nothing rotates to the next at once, with no dwell. With every enabled mode empty (only a sports plugin in its off-season, say) the loop went round with no sleep: on ledpi, 169% CPU and ~1,800 "No content" log lines every 10 seconds. After one full rotation of empty passes it now pauses EMPTY_ROTATION_PAUSE (1s) per pass, servicing plugin updates and returning early on on-demand or schedule changes; live priority is still checked at the top of every pass, and the streak resets as soon as any mode shows something. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): restart the loop if the empty-rotation pause starts on-demand; per-rotation streak - an on-demand request serviced during the pause returned early into the on-demand branch, which advanced past the mode just requested; the loop now restarts when the pause changed the mode, on-demand state or schedule - the streak is reset when the rotation changes (on-demand start/stop, a plugin enabled or disabled), so a streak from one rotation can't make another pause before its own modes are tried - docstring: live content is picked up within the pause, not "at once" Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7eb7a58d0c |
fix: web UI and src.common bugs (wifi wrong-password, plugin icon, starlark toggle, API caching, scroll/logo/font helpers) (#646)
- wifi: keep the "wrong_password:" prefix through the restore/AP fallback so the UI's incorrect-password prompt fires again. - /plugins/installed returns the manifest's icon (string only). - /starlark/apps/<id>/toggle coerces `enabled` and delegates to _toggle_starlark_app (disk before memory, no KeyError, "false" is false). - /api/v3/ JSON GETs are sent Cache-Control: no-store; non-JSON keeps 5s. - ScrollHelper.set_scrolling_image converts non-RGB input (alpha onto black); create/set_scrolling_image reset last_update_time like reset_scroll. - LogoHelper backs off a failed download per path for MISSING_LOGO_RECHECK_SECONDS; cleared on invalidate/clear_cache. - refresh_placeholder_timestamp saves atomically. - FontManager.clear_cache / _clear_plugin_font_cache bump cache_generation. - Odds manager: per-game logs to DEBUG; JSON decode error caught before RequestException (same cooldown). - element_style mangled continuations; startup validator skips null plugin blocks and reuses the controller's discovery. - src/common/README lists frame_timing, json_body, render_gate. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6bc13a8934 |
fix(display): Vegas resumes after live priority, and six smaller runtime fixes (#644)
- Vegas: a live-priority pause was only lifted from inside run_frame(), which returns before that check while paused, so the ticker never came back until a restart. run_iteration() now resumes it (the controller only calls it when nothing preempts Vegas); start()/stop() clear the pause state. Iteration length is timed with the monotonic clock. - Dim schedule: a per-day disabled day now updates the minute-gate cache, so brightness no longer flips back to dim within each minute. - On-demand: a second request no longer overwrites the rotation resume index with the first request's mode. - Render pipeline: reset() drops the prepared group and deferred queue, and a prefetch in flight across a reset discards its result. - Sync: stop() removes the status file (and the controller's cleanup now calls it), standalone removes a stale one at startup, and writes use a unique mkstemp temp file. - render_gate.swap_releases_gil() delegates to frame_timing. - Stale docstrings/comments corrected. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
da5937da3d |
fix: six bugs found testing main on a real Pi (#641)
* fix: six bugs found testing main on a real Pi (ledpi) - Stopping the service now runs cleanup. systemd stops ledmatrix.service with SIGTERM, whose default action ended Python before run()'s finally block, so the update worker, Vegas and the panel were never torn down. main() now turns SIGTERM into KeyboardInterrupt, the Ctrl-C path. - "Now showing" no longer turns into "unknown". display_current_state was only written on a mode change and the web UI reads it with max_age=120, so a live game or a single plugin on screen for longer read as unknown. It is republished every 30 s while unchanged. - Switching Vegas on in the web UI works when it was off at startup. The coordinator was only created at startup; the config watcher now flags it and the render thread creates it. - configure_web_sudo.sh finds reboot and poweroff in /usr/sbin. Run as the web user it could not, silently dropped their rules and still said it granted them, so the web UI's Reboot/Shutdown stopped working. - check_system_compatibility.sh reports installed packages as installed. `dpkg -l | grep -q` under pipefail failed when grep exited early. - A network failure fetching GitHub repo info logs a WARNING, not ERROR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): fixes found testing on a Pi Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: run the Linux-only script tests correctly The sbin-lookup test set PATH=/nonexistent and then could not find bash itself; call it by absolute path. The dpkg-query stub read $4, but the package name is the third (last) argument. Both now pass on a Pi. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): cover the follower, long-render and startup cases Review follow-ups on the ledpi fixes: - The pending Vegas start is applied before the sync-follower branch too (_apply_pending_vegas_init), which skips _is_vegas_mode_active() while a follower is connected but needs the coordinator for the leader's image. - _service_pending_changes(), which runs inside Vegas iterations and long screens, republishes a stale display_current_state as well; the main loop alone could be away for a 240 s Vegas iteration. - The SIGTERM handler is installed after DisplayController() is built, so a stop during parallel plugin loading keeps the default immediate exit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b9416ef803 |
fix(display): Vegas teardown and 240s default; remove dead Vegas buffer code (#637)
* fix(display): tear down Vegas mode on controller cleanup DisplayController.cleanup() never called VegasModeCoordinator.cleanup(), so the Vegas teardown (stop, pipeline/stream reset, adapter cache drop) was unreachable. Call it before the display manager is cleaned up, and skip it when Vegas was never created. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): default max_cycle_duration to the documented 240s The template, the web UI help, CONFIG_REFERENCE and the controller all say 240, but the code defaulted to 600 in two places, so a config without the key ran Vegas iterations 2.5x longer than documented. from_config now falls back to the dataclass field defaults instead of repeating each one, so the two copies can no longer drift, and the controller's follower scroll-speed default reads VegasModeConfig's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): let run.py -d show display_manager's DEBUG output display_manager pinned its logger to INFO at import, overriding the root level, so debug mode never showed its DEBUG lines. Use get_logger() from src.logging_config like the rest of the core and leave the level to the logging setup. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): run each startup validation check once StartupValidator.validate_all() ran twice at boot, before and after the plugin manager was created, so every config, cache, display and systemd-unit warning was logged twice. The second pass now runs only the plugin checks. Drop the commented-out raise_on_errors line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): one INFO line per plugin-list refresh StreamManager logged "=" * 60 banners and a line per plugin (INCLUDED, SKIPPED, FETCHING CONTENT, SEGMENT CREATED) at INFO on every refresh and fetch, i.e. at each cycle start and every 30s. Log one INFO summary of the rotation per refresh and move the per-plugin detail, the weighting breakdown and "no content this cycle" to DEBUG (the adapter still warns when every content path fails). Also drop the check/cross marks from the controller's log messages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): drop the per-iteration static-mode plugin scan run_iteration() rebuilt _static_mode_plugins on every iteration, asking every plugin for its display mode and logging the set at INFO, but nothing ever read it: static pauses are triggered by _check_static_plugin_trigger() from the next segment. Delete it, the coordinator's get_ordered_plugins() that only it used, and the write-only _static_pause_plugin / _static_pause_start. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): remove the staging buffer that was never filled StreamManager and RenderPipeline carried a double-buffer design that nothing used: _staging_buffer was only ever cleared or swapped, so swap_buffers() never did anything and should_recompose()'s staging_count > 0 branch was dead, and _active_scroll_image, _staging_scroll_image, _is_rendering, _last_frame_time and _frame_interval were written but never read. Delete the machinery and rewrite the docstrings around what actually carries updates: _pending_updates, consumed by process_updates() in swap mode and invalidate_pending_updates() in continuous mode. should_recompose() no longer builds a buffer-status dict every frame. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): tidy the display controller without changing behaviour - Import VegasModeCoordinator locally instead of through module globals (there is no circular import to avoid). - Drop hasattr() checks on attributes PluginManager.__init__ always sets (plugin_executor, plugin_last_update, get_plugin_lock, run_scheduled_updates*, stop_update_worker) and the dead "older manager" fallbacks; keep the health_tracker None checks, now via _health_tracker(). - Extract _display_once() for the per-frame display call both render loops copied, _advance_on_demand() for the two on-demand rotations, _reset_on_demand_fields() for the error and clear paths, and _timezone() / _in_window() for the two schedule checks. - Remove always-true conditions and the unreachable non-plugin else branch in run(), and read _was_display_active / _last_published_mode / vegas_coordinator directly now that __init__ declares them. - Declare the follower render state in __init__, name its tuning constants, add _follower_sign(), and share the 90/s sync send interval with the render pipeline (SYNC_SEND_INTERVAL). - Delete history narration and the "Opt #N" labels, fix the comment that called _scroll_speed constant (hot reload updates it), and drop a startup timing log that measured nothing. - render_pipeline / plugin_adapter: read display_manager.width/height as the properties they are, drop an empty TYPE_CHECKING block, an aliased threading import and a duplicated `if result and self.sync_manager:`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): trim dead code from display_manager - Add _new_canvas() for the image/draw/fontmode="1" setup that was copied six times. - Call resolve_double_sided() and compose_pixel_mapper_config() directly instead of through a module alias and a passthrough method, and replace the comment that said the passthrough read class attributes. - Delete the unused _initialized flag and _ORIENTATION_ROTATE_DEGREES alias (no core or monorepo reader; tests stop resetting the flag), the test pattern's unreachable no-matrix branch (it only runs once the matrix exists), `del old_image # help GC` (a no-op on a local), a duplicated early return in process_deferred_updates, and stale comments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(vegas): remove unread fields and test-only helpers, fix docstrings - ContentSegment: drop total_width, fetched_at, is_stale, image_count and is_static, none of which is read. - StreamManager: drop _current_index (never advanced) and the test-only get_all_content_for_composition() and has_pending_updates(); VegasModeConfig: drop the test-only is_plugin_included(). - geometry.find_blank_cut() has had no production caller since the crop moved to item boundaries; delete it and its tests. - PluginAdapter: the _finalize docstring described separator_width between every image, and _crop_to_budget's said cuts snap to the nearest blank column; both now describe what the code does. - Coordinator: the static-pause interrupt log no longer blames follower mode for every interrupt, and set_update_callback names the callback the controller actually wires. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(scroll): correct ScrollHelper comments and drop dead branches - Four comments said the strip always starts with display_width of blank; it does only when lead_gap is None (Vegas passes its own). - Delete the "Width calculation mismatch" warning: the image is created at the calculated width, so the two can never differ. - Remove the two scroll_delay <= 0 fallbacks (which disagreed with each other): set_scroll_delay clamps it to at least 0.001 and nothing in core or the plugin monorepo assigns it directly. - Trim the scipy history from the blend docstring. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(run): drop a redundant comment Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): log set_scrolling_state only when it changes Vegas and scrolling plugins set the scrolling state every frame, so once display_manager's DEBUG output became visible in debug mode it printed "Scrolling state set to: True" about 120 times a second. Log only when the value differs from the previous one; the state, activity timestamp and frame hold still update on every call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): display-vegas Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f3894916a9 |
feat(web): show which plugins use each font; warn before deleting one (#619)
* feat(web): show which plugins use each font, warn before deleting one
The Fonts tab lists font files from the web process's own scan, and the
plugins that register fonts run in the display process, so the tab had no
way to say whether a font was in use before deleting it.
The display service now publishes {catalog key: [plugin ids]} to the
shared cache (font_usage_snapshot, src/font_usage.py), built from the
loaded plugins' FontManager.register_manager_font() registrations. A
daemon thread checks every 10 s and writes only when the usage changed
(plus a daily refresh so cache cleanup cannot expire it); it never raises.
Families, aliases (press_start, four_by_six, ...) and paths are resolved
through FontManager's catalog to the file stem the Fonts tab keys rows by;
fonts outside assets/fonts are left out. Unloading a plugin drops its
registrations (new FontManager.forget_manager_fonts).
GET /api/v3/fonts/catalog merges used_by into each row per request (the
5-minute scan cache is copied, never edited): a list of plugin ids, or
null when the display service has not reported. The tab shows a Used by
column ("unknown" / "-" / ids, rendered as text) and deleting an in-use
font names the plugins in the confirmation, from a fresh read. The server
still refuses only system fonts. Catalog fetches bypass the browser's
5-second API cache, which otherwise served the pre-delete list.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix: call forget_manager_fonts through a hasattr check pylint can follow
getattr(..., None) then callable() is fine at runtime, but pylint's E1102
("not callable") can't see through it, and Codacy fails the check on it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
4a1fd7464a |
fix(errors): serve /api/v3/errors/* from the display service; add a Plugin errors panel (#614)
* fix(errors): serve /api/v3/errors/* from the display service's aggregator The error aggregator is a per-process singleton and only the display service runs plugins, so only its aggregator records anything. The routes read the web process's own, empty one and always reported no errors. The display service now publishes a bounded snapshot of its aggregator to the shared cache (plugin_error_snapshot) from a daemon thread: at most once every 10 s and only when something changed, never raising into the caller. The routes read it and keep their response shapes, adding snapshot_available, generated_at and clear_pending; exception text has credentials redacted. POST /errors/clear writes a clear request (plugin_error_clear_request) that the display applies on its next 5 s tick via the new clear_before(), which keeps errors recorded after the cutoff and rebuilds the counts. Until the snapshot acknowledges the request, reads hide everything before the cutoff, so a snapshot written just before the click cannot bring errors back. Adds "all": true; cleared_count is null when only the display can know it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(web): show plugin errors in the Logs tab A compact panel under the log viewer: per-plugin error counts, repeating errors (type, count, affected plugins, a sample message, last seen) and a Clear button, with empty states for "no errors" and "display service hasn't reported yet". Polls every 15 s while the tab is active; all text goes through escapeHtml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: describe where plugin error reports come from and how clear works Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(errors): redact the published snapshot before clipping it Keeping only a traceback's tail (or clipping a message) could cut an `api_key=` marker off while keeping the secret after it, and the web side's redaction would then have nothing to match. The display now redacts every free-text field of the snapshot first. The patterns move to a Flask-free src/redaction.py so the display service can use them; redact_text in the web error handler uses the same function, unchanged in behaviour. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
cd5a4e2251 |
fix(display): apply on-demand, brightness and schedule changes mid-screen (#618)
* fix(display): apply on-demand, brightness and schedule changes mid-screen The main loop read the on-demand mailbox, the on/off schedule and the brightness target once per pass -- once per screen. A dwell can be a minute and a Vegas iteration runs for max_cycle_duration (240s), so on a Pi an on-demand request posted at 10:54:27 was activated at 10:57:24, and two brightness saves 12s apart inside one 30s screen never reached the panel. During Vegas nothing read the mailbox at all: _check_vegas_interrupt only checked on_demand_active, which only the main-loop read sets. _service_pending_changes does the main loop's on-demand poll, expiry, schedule and brightness steps, throttled to PENDING_CHANGES_INTERVAL (the existing 0.25s mailbox floor), on the display thread. It runs from the Vegas interrupt checker, the high-FPS and once-a-second render loops (replacing their direct on-demand poll) and _sleep_with_plugin_updates; between passes it costs one monotonic compare. A brightness change re-pushes the current frame, since the panel only shows it from the next push. Callers act on what it leaves behind: Vegas yields on an on-demand start or the display being scheduled off (and the main loop then blanks instead of rendering a screen), the render loops break on a schedule-off as they already did on a mode change, and the dwell sleep returns early on an on-demand start/stop or a schedule flip -- so the 60s scheduled-off sleep now wakes for an on-demand request. The main loop no longer rotates after a dwell that ended that way, which advanced a new on-demand session past the mode that was asked for. A brightness set_brightness() refuses is not retried until the target changes, so the 4Hz pass doesn't log the same failure (fallback mode) four times a second. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(display): a screen scheduled off midway stops rendering Covers the schedule-off break added to the high-FPS and once-a-second render loops: with the display scheduled off halfway through a 120s screen, neither loop renders for more than one redraw plus one service interval past the boundary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
84afa9d64f |
refactor: delete dead Python code in the core (and stop storing Wi-Fi passwords) (#608)
* refactor(plugins): remove the no-op PluginHealthMonitor Its monitor loop did nothing (`if callbacks: pass`), register_health_check had no callers and api_v3.health_monitor was never read by any route. The live health data comes from PluginHealthTracker, which is untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(store): drop the never-set uninstall tombstones Nothing in production called mark_recently_uninstalled, so the reconciler's was_recently_uninstalled check was always False. The persistent uninstall registry is what actually stops resurrection; the reconciler test now exercises that gate instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(common): delete unused config/display/game helpers, utils and error_handler Nothing in core, the web UI, scripts or the plugin monorepo imports config_helper, display_helper, game_helper, utils or error_handler; only their own tests did. The error_handler re-exports leave src.common's __all__; APIHelper, TextHelper, ScrollHelper, LogoHelper and the adaptive layout exports are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(config): drop ConfigService's unused versioning and save API ConfigVersion, get_version/get_version_history/get_version_config, rollback, save_config, reload, get_plugin_config and the backward-compat load_config/get_config_path/get_secrets_path had no callers. The display controller only uses get_config, subscribe, unsubscribe and shutdown, plus the file watcher. Change detection now compares against the current checksum instead of the last history entry. The subscriber tests asserted `callback.called or True`; they now reload the way the watcher does and assert the notification. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): drop unread plugin state history and callbacks plugin_state.PluginStateManager kept a bounded per-plugin transition history that only get_state_history (tests only) read; get_state_info reports a separate lifetime count, which stays. set_error_info and record_display had no callers, and set_state_with_error's `error` argument only fed the history. The web-side state_manager.PluginStateManager loses subscribe_to_state_changes, _notify_callbacks, set_plugin_error and get_state_version, none of which had callers; with no subscribers the old-state copy in update_plugin_state went with them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): remove unused PluginManager methods and attribute guards update_all_plugins was only called by a test (the display loop uses run_scheduled_updates); get_plugin_health_metrics, get_plugin_resource_metrics and get_plugin_state had no callers; and plugin_modules was written but never read. plugin_directories is now initialised in __init__, so the hasattr() guards around it go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(plugins): remove unused executor, loader, store and package helpers - PluginExecutor.execute_safe: no callers. - PluginLoader._parse_semver: only its own tests; compatibility.parse_semver is the live copy and test_compatibility.py already covers it. - PluginStoreManager.get_installed_plugin_info: no callers. - PluginResourceMonitor._local: never read. - src.plugin_system.get_store_manager and __api_version__: no importers in core, scripts or the plugin monorepo. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(wifi): stop storing Wi-Fi passwords in wifi_config.json WiFiManager appended every joined network's SSID and password, in plaintext, to saved_networks in config/wifi_config.json, and nothing (web UI, backup restore, scripts) ever read them back: NetworkManager keeps its own credentials. The writes are gone, and loading the config now drops any saved_networks key and rewrites the file, so passwords already on disk are scrubbed. Also removes _check_dnsmasq_conflict (never called) and _detect_trixie, whose result only reached one log line, along with the NM_CONNECTIONS_PATHS constant only it used. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): remove unreachable and unused DisplayController code - _follower_rebuild_scroll_image: never called. - mode_duration (never read) and last_mode_change (write-only). - The `chosen_cap <= 0` branch: chosen_cap is either the minimum of caps already filtered to > 0 or DEFAULT_DYNAMIC_DURATION_CAP (180). - The `max_duration < min_duration` branch directly after `max_duration = max(min_duration, max_duration)`. - The circuit-breaker branch's `display_result = False` and `manager_to_display = None`: the first is overwritten a few lines later, the second is already None there. - The bool-to-bool conversion of execute_display's result, which is always a bool. - The `loaded_plugins` lookup in _update_modules: PluginManager has no such attribute, so it always fell through to `plugins`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(vegas): remove unused config update, boundary finder and refresh VegasModeConfig.update had no callers outside its own tests (the coordinator rebuilds the config with from_config on a change); geometry.find_item_boundary and StreamManager._refresh_plugin_content had no callers at all. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(run): drop the debug block that pretended to import the plugin system In debug mode run.py put src/plugin_system itself on sys.path and printed "Plugin system import successful" without importing anything. Nothing imports plugin_system modules by bare name, so the path entry did nothing either. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: delete tests that test nothing - test/plugins/test_{basketball_scoreboard,calendar,clock_simple, odds_ticker,soccer_scoreboard,text_display}.py skip everywhere the named plugins are not installed, including CI (LEDMATRIX_PLUGINS_DIR holds only the fixture plugin); test_plugin_matrix.py already covers every discovered plugin. Their PluginTestBase and the fixtures only it used (plugins_dir, mock_display_manager, mock_cache_manager, mock_plugin_manager, base_plugin_config in test/plugins/conftest.py) go with them. - test_plugin_system.py: test_discover_plugins (body was `pass`) and test_dependency_check (a comment), plus the test_plugin_manager fixture only the former requested. - test_display_manager.py: test_draw_image asserted that an image it had just assigned was not None. - test_display_controller.py: the rotation and schedule-override tests re-implemented the run-loop arithmetic inline and asserted on their own result without calling the controller. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test: expect one plugin_last_update success stamp after update_all_plugins EveryStampRecordsACompletion required at least two success-path stamps; the second was update_all_plugins, removed as test-only. The worker and synchronous paths share the remaining stamp in _execute_update_now, and the check that every stamp calls _note_update_completed is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
342e9164b8 |
fix: settings the display ignored, a memory leak, and the plugin card handler (#605)
* fix(errors): stop affected_plugins growing without bound Each repeat of an error pattern appended every plugin in the time window to the pattern's list again, so a plugin failing in a loop grew the display process's memory without limit: 3,000 errors from three plugins reached 2.5 million entries. Keep the list unique. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(fonts): load a BDF font at its native size instead of PIL's default FreeType rejects any size but a BDF strike's own, and FontManager answered that with ImageFont.load_default() -- a different typeface -- so 5x7.bdf requested at 8 or 10px rendered as PIL's default font. Retry at the native strike, as element_style already does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): plugin toggle failures no longer claim "operation in progress" Every exception in POST /plugins/toggle was mapped to PLUGIN_OPERATION_CONFLICT, so any failure told the user "A plugin operation is already in progress". Report the failure as what it is, and record the plugin id in the operation history for form posts too (it read a `data` variable that only the JSON path set). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(web): route plugin card clicks through handlePluginAction The document-level delegation checked `typeof handlePluginAction`, which is scoped inside the plugin-manager IIFE and so never visible to it. Every card click took a copied fallback that stopped propagation (the grid's own listener never ran), confirmed an uninstall twice, and sent Starlark app uninstalls to POST /plugins/uninstall instead of DELETE /starlark/apps/<id>. Expose the handler on window and delegate to it. Also run every test/js/unit suite under pytest: they need only node, but CI ran one of the eight. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): apply Rotation durations, WiFi messages and Vegas settings Three settings the web UI saves never reached the display: - Rotation & Durations: display.display_durations was never read. Every plugin inherits get_display_duration() and the plugin was asked first. A saved value now wins. The page shows unsaved screens blank with the plugin's own duration as a placeholder, and saving a blank removes the override, so one save no longer pins every screen. - WiFi status overlay: the controller looked for wifi_status.json one directory above the repo. Both sides now use wifi_manager.get_wifi_status_path(). The message is written by rename so the display never reads it half-written, and the resumed plugin redraws the whole panel afterwards. - Vegas: nothing called coordinator.update_config(), so saved Vegas settings never reached a running scroll. They are now queued when display.vegas_scroll changes, and applied while Vegas is stopped too, so a disable then re-enable works. The follower's scroll-speed default (75) now matches VegasModeConfig's (50). Also throttles Vegas's per-frame live-priority scan to 4Hz. It cost 139us per frame on a Pi 4 with two scoreboards (1.7% of a 125fps frame) and grows with each plugin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep affected_plugins order when serialized; guard non-Element targets ErrorPattern.to_dict() ran the now-ordered list through set(), so get_error_summary() listed plugins in an unstable order. The document-level card-action listener called event.target.closest() without checking the target is an Element. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
116abb0daa |
fix: September 16 core audit — partial saves, asset path safety, auto-update, display settings the library refuses, scroll speed (#595)
* fix(sports): share the ESPN rejected-range memo with the background service BackgroundDataService always sent a season range first and, on a 400, fell back to chunks without recording the rejection, so every background season fetch spent a doomed request and live scoreboards learned nothing from it (or it from them). The worker now consults and sets the same 6-hour memo fetch_espn_scoreboard() uses: a known rejection goes straight to month/day chunks, and if every chunk fails the range is asked once for a real error without re-spending the chunks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): keep plugin asset and action routes inside their directories POST /plugins/assets/upload, GET /plugins/assets/list and POST /plugins/assets/delete joined the request's plugin_id onto assets/plugins unchecked, so '../../config' created, wrote, listed and deleted outside it. #561 guarded only the route that serves the files. All three now go through path_safety.resolve_under and answer 400 for anything but a plain name, and delete only unlinks a metadata path that resolves into that plugin's uploads directory. PluginManager.get_plugin_directory refuses ids that are not one plain path segment, so /plugins/action (which runs a manifest script from the returned directory) and every other caller get the guard; the action route also rejects such ids up front, covering its no-manager fallback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): report a no-op plugin update as already up to date update_plugin() returns True both for a real update and for "nothing to do" (a ZIP-installed monorepo plugin already at the registry version, a bundled plugin). With no git commit to compare, POST /plugins/update called every such success "updated successfully", so Check & Update All counted most official plugins as updated on every run. The route now reads what changed off the plugin itself (commit, else manifest version, else last_updated) and returns data.update_status (updated / up_to_date / local_only). The update-all toast is summarised by PluginInstallManager.summarizeUpdateResults from that status, falling back to the message for older servers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(sports): scoreboard scroll speed no longer follows target_fps sports_scroll computed the crisp speed ladder against the global target_fps whenever limit_refresh_rate_hz was the 100 Hz default. Since frame-locked presentation (#545) the helper steps a fixed number of whole pixels per presented frame and the panel presents at its real refresh, so the General tab's "Scroll Frame Rate" became a speed multiplier: 60 ran a 50 px/s scoreboard at 100 px/s, 200 ran it at 25 px/s. The ladder now uses the display manager's refresh_hz, then display.hardware.limit_refresh_rate_hz, then the default. target_fps is not consulted. Docstrings now say scroll_delay is ignored for pacing (no behaviour change there) and describe the fixed-step model. Tests: replace the tests that pinned target_fps as the ladder refresh and described time-based stepping; assert speed independence from target_fps (unit and end-to-end presented px/s against the real helper), that the fixed per-frame step is applied, and that scroll_delay does not change speed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): escape registry and upload values in plugin manager inline handlers The store, saved-repository and custom-registry buttons built onclick='...(${JSON.stringify(id)})...'. JSON.stringify leaves ' alone, so a custom registry entry whose id contained ' closed the attribute and added its own handler. One helper, jsStringAttr(), now HTML-escapes the JSON literal for every one of those handlers, and the store View button opens only http(s) repo links. The live window.updateImageList (plugins_manager.js loads last, so its copy wins over the file-upload widget's) wrote the uploaded file's original name, path and ids into markup raw; they are escaped now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note plugin asset, action and inline handler guards Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): let the root pip wrapper install web_interface/requirements.txt Update Code, the automatic update's health check and Install Base Requirements install web_interface/requirements.txt through safe_pip_install.sh, which only allowed the root requirements.txt. The first commit changing that file would fail its dependency install, and the automatic updater rolls back any update whose dependencies did not install -- on every device, for every newer commit. The wrapper now lists both core requirement files. Only their folders are resolved, so a requirements.txt symlinked out of the project is compared by its target and refused (previously the root file's own symlink target was what got allowed). The updater's file list is a named constant, and a test runs the real wrapper (pip stubbed) on every file Update Code and the rollback install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): do not retry plugin requests that got an HTTP answer PluginAPI.request wrapped everything that was not a structured error as NETWORK_ERROR: a proxy's 502 HTML page (response.json() throws) and a JSON error without error_code included. Check & Update All retries NETWORK_ERROR, so those updates were re-sent five more times with backoff, contrary to the #587 contract that an HTTP error response is the server's answer. NETWORK_ERROR now means only that fetch() rejected. Any HTTP response without an error_code, or with a body that is not JSON, is API_ERROR with the HTTP status attached. Tested against the shipped api_client.js. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): restart the stats window when an idle gap is dropped by size #582 dropped an idle gap from the frame stats two ways: the reset_scroll() sentinel, which also restarts the 5s window timer, and a size guard for scrollers that never call reset_scroll(), which did not. On that path the first real frame after the gap found the boundary overdue and logged a stats line for a one-frame window. Both paths now share one seeding helper. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): leave plugins alone when update_core's own rollback fails update_core returns rollback_failed directly when a partial pull or an update whose health check never started cannot be rolled back. run() only held plugins back for 'verifying', so those devices still got new plugin versions and a display restart on top of a core in an unknown state -- the opposite of what the health-check path does, and of the 3.4.0 changelog (plugins are left alone if the rollback fails). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(api): make the REST reference match the api_v3 package Every documented request body, query parameter and response shape was re-checked against the handlers in web_interface/blueprints/api_v3/. Fixes calls that failed as documented (repo_url, action_id/params, files/image_id, font_file+font_family, ?font=, cache key, auto_enable_ap_mode, plugin limit keys), removes the font-override endpoints dropped in #566, corrects response shapes (plugins/config, plugins/schema, health, metrics, operation history, github-status, fonts/catalog, cache/list, logs, wifi, on-demand, SSE streams), and adds the 26 routes it omitted (backup, system auto-update/git, wifi radio, starlark editor, MQTT bridge, status endpoints, skins). Documents the merge semantics of partial JSON saves to /config/main and /plugins/config and the dim-schedule POST accepting GET's days shape, which land in the same change set. Replaces app.py line numbers and the removed api_v3.py path with file and function names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): remove the General-tab plugin system toggles that did nothing plugin_system.auto_discover, auto_load_enabled and development_mode had General-tab toggles whose help tips promised dormant plugins and verbose logging, but nothing reads them: every enabled plugin is discovered and loaded regardless. Remove the three toggles. The keys stay tolerated in stored configs. The save handler now stores a flag only when a client sends it; treating a missing key as an unchecked box would otherwise rewrite all three to false on every General-tab save, which still posts plugins_directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(scroll): remove dead code left by #523/#570 - Drop the optional scipy.ndimage import and HAS_SCIPY; nothing read them since the numpy blend replaced the scipy path. - Drop ScrollHelper._last_integer_position and frame_time_target, which were written but never read. - Keep target_fps and set_target_fps() but document them as informational: nothing paces off them, yet ledmatrix-elections' test_scroll_pacing.py reads helper.target_fps back and third-party plugins may call the setter. - Fix stale comments: fixed_pixels_per_frame's "use scroll_delay to throttle", set_sub_pixel_scrolling's "default: True", and set_frame_based_scrolling's claim that it steps. The plugins monorepo was grepped for every removed name; none is used. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(fonts): point plugins at plugin_manager.font_manager; drop removed overrides UI FONT_MANAGER.md told plugins to read display_manager.font_manager, which does not exist, so a plugin following it failed to load with AttributeError. The shared FontManager lives on the PluginManager and BasePlugin._get_font_manager() returns it (with a fallback for harnesses). Also removes the Fonts-tab override workflow and element-override panels that #566 deleted, from FONT_MANAGER.md and WEB_INTERFACE_GUIDE.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(store): search via /plugins/store/list?query=; send Content-Type on registry curls /plugins/store/search does not exist (404) and the list endpoint reads query, not q. The registry guide's curl examples omitted the JSON Content-Type, so the handlers saw an empty body and answered 400. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(config): use the shared core-key list in the last three private copies StartupValidator warned "Plugin 'auto_update' is enabled but not found" on every display start with auto-update or a dim schedule on; the reserved plugin-id check missed auto_update, sync, location and the rest; and ConfigManager's (uncalled) orphan cleanup would have deleted display, schedule and auto_update. All three now read src/core_config_keys.py, which also gains CORE_SECRETS_KEYS for the github/youtube secrets sections. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): partial JSON saves to /config/main change only what they send A JSON body with one field reset every checkbox in the sections it touched: the MQTT bridge's brightness slider turned off disable_hardware_pulsing, inverse_colors, show_refresh_rate and use_short_date_format, and a timezone-only save turned off web-UI autostart and weekly auto-updates. Missing-means-unchecked now applies only to form posts: form-encoded bodies and the v3 forms, which mark themselves with a hidden __form_section input. Also on the config routes: - vegas_min/max_cycle_duration no longer match the generic *_duration rule, so they stop landing in display_durations and a blank one no longer rejects the whole Display save; - saving from the Raw JSON editor calls start_setup_if_needed like the General form, so enabling auto-update there finishes its setup; - the schedule and dim-schedule POSTs accept the per-day days.<day> shape their GETs return, as well as the flat form keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): install plugin dependencies from the configured plugins directory install_plugin_dependencies.sh scanned only plugins/, but the Plugin Store installs into plugin_system.plugins_directory (default plugin-repos), so the documented "Recommended" fix found 0 plugins on every store install. It now reads plugins_directory from config/config.json (relative to the project root or absolute, default plugin-repos) and also scans plugins/ for dev symlinks, installing a plugin reached through both only once. With set -e alone, `pip ... | tee` took tee's exit status, so a failed pip install was reported as success; set -o pipefail. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: replace stale API names, line numbers and the api_v3.py path - ADVANCED_FEATURES: StreamManager methods that exist (get_next_segment, take_next_group, refresh, advance_cycle, ...), and the real on-demand status envelope ({status, data: {state, service}}) - app.py:199 / :144 / :607-619 line citations and web_interface/blueprints/api_v3.py (now a package) replaced with file and function names in ADVANCED_FEATURES, CONFIG_DEBUGGING, PLUGIN_ARCHITECTURE_SPEC, PLUGIN_QUICK_REFERENCE, PLUGIN_CONFIGURATION_TABS, TROUBLESHOOTING and web_interface/README - CONFIG_DEBUGGING: partial /config/main saves change only sent keys; use /config/raw/main to replace the file; describe where validation runs - TROUBLESHOOTING: clear_cache.py needs --clear-all (no args only prints usage) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): verify the web interface that actually ships, on port 5000 verify_installation.sh failed every healthy install: it required the long-removed web_interface_v2.py and looked for a listener on port 5001, while the web interface binds 5000 (web_interface/start.py). It now checks the files ledmatrix-web.service runs (start_web_conditionally.py, web_interface/start.py, app.py) and port 5000. verify_web_ui.sh had the same 5001 port in its listen check, HTTP probe and printed URLs. Port matches are anchored so :50001 no longer counts as :5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(plugins): one display-size contract: display_manager.width/height CLAUDE.md (#580) says to read display_manager.width/height because matrix is None when hardware init fails; the development guide, the safety-harness doc and two DisplayManager docstrings still recommended matrix.width/height. The bundled starlark-apps plugin read matrix.width unguarded, so its magnify recommendation and frame scaling raised in fallback mode (e.g. after the Pi 5 hardware refusal). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(install): make install_service.sh --help print usage instead of installing install_service.sh parsed no arguments, so `sudo ./scripts/install/ install_service.sh --help` (presented as harmless in MIGRATION_GUIDE.md) rewrote ledmatrix.service, ledmatrix-web.service and both update-verify units and enabled/started them. It now handles -h/--help (usage, exit 0, no changes) and rejects any other argument with exit 2 before doing anything. Running it with no arguments, as first_time_install.sh does, is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scroll): describe the fixed-step model and document frame_hold Since #545 a crisp speed from scroll_config.configure() makes the helper advance a fixed whole-pixel step per presented frame with no clock, and the display manager's frame hold is part of the speed. The docs still described the removed wall-clock model: - scroll_config's module and configure() docstrings said speed is applied in time-based mode and that omitting the hold "falls back to fractional pixels"; omitting it actually runs the scroll frame_hold times too fast. - SCROLL_PERFORMANCE.md said ScrollHelper accumulates elapsed time in both modes, and read a 20 ms stats median as missed refreshes although that is a healthy 50 px/s (hold 2) scroll. It now explains the fixed step, the hold-dependent healthy median, that target_fps plays no part, and that a hand-added scroll_pixels_per_second loses to a schema-default pair. - PLUGIN_API_REFERENCE.md documented set_scrolling_state(is_scrolling) without frame_hold; it now documents the parameter (core 3.4.0) with a configure() + set_scrolling_state example. - update_scroll_position/set_scroll_speed and set_scrolling_state docstrings say the same. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark target_fps legacy; describe what Vegas scroll_delay does - General tab "Scroll Frame Rate" (target_fps) is labelled legacy: after the sports_scroll fix nothing in core scrolling reads it. The field and its API validation stay so saved configs and plugins that read global_config['target_fps'] keep working. CONFIG_REFERENCE says the same. - Vegas frame_based_scrolling/scroll_delay were described as frame-count stepping at ~50 FPS. Neither steps nor sets a frame rate: frame-based mode converts the speed to px per scroll_delay, clamps it to 0.1-5, and still advances by elapsed time, so the applied speed is clamp(scroll_speed * scroll_delay, 0.1, 5) / scroll_delay px/s. The config comments, render_pipeline comment and CONFIG_REFERENCE rows now say so. No behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(deps): describe how plugin dependencies are really installed The guides said the web service runs as root, that installs pick --user from os.geteuid(), and quoted a warning and a PluginManager._install_plugin_dependencies() method that don't exist. The web unit runs as the installing user; store installs go through install_requirements_file() and sudo safe_pip_install.sh (root), with a user-level fallback that says so, and load-time installs run in the display service's own (root) interpreter. Manual paths now use the configured plugins directory (plugin-repos/ by default) instead of plugins/, which store installs no longer use, and install_plugin_dependencies.sh is described as scanning that directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): count local changes one way for the preflight and the pull The automatic update's preflight ignored mode-only changes and anything whose status line contained plugins/ or plugin-repos/, then promised "Automatic updates will not stash your changes". perform_core_update used plain git status (modes count) and ignored only 'plugins/', then ran 'git stash push -- :!plugins', which nothing ever pops. So an edit to a bundled plugin under plugin-repos/, or the installer's chmods on tracked scripts, passed the preflight and was stashed away for good. - auto_update.local_changes() is the one predicate both use: core.fileMode=false, porcelain -z, and plugins/ and plugin-repos/ excluded by leading folder rather than substring (a core file under web_interface/static/v3/js/plugins/ now counts). - Update Code's explicit stash leaves out both plugin folders; the pull's --autostash carries their edits and mode changes across and reapplies them. - The automatic updater calls perform_core_update(stash_local_changes= False), which refuses instead of stashing edits that appeared after the preflight; update_core reports that as 'blocked'. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): diagnostics follow the web autostart default and api_v3 package #556 made a missing web_display_autostart mean "start" (only an explicit false/off keeps the web interface down), but the diagnostics still said otherwise: diagnose_web_ui.sh reported a missing key as "defaults to false", diagnose_web_interface.sh said the web interface "will not start unless this is set to true" and recommended enabling it, and debug_web_manual.py printed False. Troubleshooting a down web UI pointed users at a non-cause. Both shell scripts now evaluate the setting with the launcher's own autostart_enabled() (inline fallback if it cannot be imported) and report on / off / not set (on) / unparseable config; debug_web_manual.py uses the same function. They also check web_interface/blueprints/api_v3/ __init__.py: api_v3.py became a package in #553, so every healthy checkout was reported as missing a file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(install): what install_service.sh installs; verify script port; no sudo for --help install_service.sh installs and starts ledmatrix, ledmatrix-web and the update-verify units, not only ledmatrix.service (systemd/README.md, README.md). MIGRATION_GUIDE presented 'sudo install_service.sh --help' as a harmless check; it now shows --help without sudo and warns what a real run does. SSH_UNAVAILABLE_AFTER_INSTALL: verify_installation.sh checks the web interface on port 5000. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): note update-all, plugin system settings and script fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): size the preview after orientation and pixel mappers display_geometry.physical_size claimed to give DisplayManager's answer but only computed cols*chain x rows*parallel. RGBMatrix.width/height are measured after the library's pixel mappers, so a Rotate:90 / orientation 90 chain previewed 128x32 for a 32x128 panel and a U-mapper chain of four 256x32 for 128x64. Model the built-in mappers' size effect as the pinned lib/pixel-mapper.cc does (Rotate, U-mapper, V-mapper, StackToRow, Remap; Mirror and unknown names leave it alone), and move the orientation composition here so DisplayManager and the preview share it. The module docstring no longer claims the sync handshake uses it; that imports only DEFAULT_CHAIN_LENGTH. Audit finding F18. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): refuse settings the rgbmatrix library aborts on, on every board The library answers several settings with a NULL matrix or abort() rather than an error, so the display service crash-looped (Restart=on-failure) instead of reaching fallback mode: rows above 64, chain_length above 255 (uint8_t binding setter, documented as "no upper limit"), a misspelled hardware_mapping, and parallel 2-3 on a single-output mapping, reachable from the Display form on the default adafruit-hat(-pwm) mapping. #586 only guarded the Pi 5 subset. - src/matrix_support.py holds the rules for every board (Options::Validate ranges, binding integer types, mapping names and outputs from lib/hardware-mapping.c) plus the Pi 5 ones, and is the one source of the API's numeric ranges. - DisplayManager checks them before building options and raises MatrixSettingsRefused, so a hand-edited config falls back with a logged, reported reason. Emulator mode only warns. - The config API refuses them with a 400 naming the setting; combinations are checked against stored values but reported only when the request sets a field involved. - The hardware status file gains "cause" (settings/library/forced). The fallback log and Display banner give the Pi 5 rebuild hint only for a library failure instead of rebuild + gpio_slowdown advice for every failure; one Pi 5 slowdown recommendation (1-3, start at 1). - The Display form offers classic/classic-pi1 and orientation 90/270 and renders any other stored mapping selected with a warning, so an unrelated save no longer rewrites them; the API accepts 90/270. Audit findings F03, F16, F19, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(display): library limits, template defaults and Pi 5 slowdown - rows 8-64, chain_length 1-255, parallel limited by the mapping's outputs, classic/classic-pi1 mappings and orientation 90/270 documented. - Defaults are the config.template.json values: config migration adds missing keys from the template, so the listed "code defaults" never applied. - One Raspberry Pi 5 gpio_slowdown recommendation: 1-3 in PIO mode, starting at 1. - Troubleshooting describes the refused-settings fallback, and CHANGELOG corrects the Unreleased "no upper limit" entry. Audit findings F19, F20, F21. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py opens the panel with the service's options --measure and --demo built RGBMatrixOptions from a private copy of the display service's builder that had drifted: gpio_slowdown came from display.hardware (default 2) instead of display.runtime (default 3), and rp1_rio, panel_type, disable_hardware_pulsing, inverse_colors, pixel_mapper_config and orientation were skipped, with different defaults (hardware_mapping "regular", pwm_bits 11). A panel needing a high slowdown was measured -- or garbled -- in a setup the service never drives. The option filling in DisplayManager._setup_matrix moves, unchanged, into DisplayManager.apply_matrix_options(options, config), which _setup_matrix calls and the script reuses (overriding only limit_refresh_rate_hz for --measure). The script now loads the whole config rather than the hardware block. Tests pin the script's options to the service's attribute for attribute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): scroll_speeds.py recommends keys the resolver honours The ladder ended by telling users to set display_options.scroll_pixels_per_second. scroll_config ranks that key below the scroll_speed + scroll_delay pair, deliberately, and several plugin schemas default the pair into config, so the advised key was silently ignored (a schema-default 1/0.02 pair plus an advised 66 still resolved to 50 px/s). The advice is now the pair that selects the crisp speed exactly (pixels_per_frame every frame_hold/refresh seconds), explains that the pair outranks scroll_pixels_per_second, and gives the scoreboards' per-league scroll_settings.scroll_speed (px/s) form. Tests resolve the printed pair over a schema-default pair and check it lands on the advertised speed and hold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: withdraw the target_fps claim for sports_scroll; fix the Vegas speed formula - SPORTS_UNIFICATION.md still presented honouring global target_fps as sports_scroll's added behaviour and its one user-visible gain; note that it was withdrawn because it had become a speed multiplier. - ADVANCED_FEATURES.md gave Vegas scrolling as (scroll_speed / target_fps) * elapsed; the real rule is scroll_speed px/s by elapsed time, through a 0.1-5 px per scroll_delay clamp when frame_based_scrolling is on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): scroll model fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(dev): link-github links plugins from the ledmatrix-plugins monorepo link-github <name> cloned https://github.com/ChuckBuilds/ledmatrix-<name>.git, and those per-plugin repositories no longer exist: official plugins are directories in the ledmatrix-plugins monorepo. It now clones (or pulls) the monorepo once into the dev directory, finds plugins/<name>, plugins/ledmatrix-<name> or the plugin whose manifest id is <name>, and links it under its manifest id. With an explicit repo URL it still links a single-repository plugin as before. dev_plugins.json: github_user is honoured again (monorepo owner, e.g. a fork), plus plugins_repo and plugins_branch; github_pattern, which was documented but never read, is dropped and warned about. Ships dev_plugins.json.example and git-ignores dev_plugins.json, both of which the guide promised. Reading JSON falls back to python3 when jq is missing (get_plugin_id silently returned nothing without jq). update/status/list find the git checkout above a monorepo plugin directory (its .git is not in the plugin dir), and update pulls a shared checkout once. status no longer exits 1 when nothing is broken. Docs: PLUGIN_DEVELOPMENT_GUIDE (quick start, link-github, configuration, workflow, store integration, hello-world link, submission), and the nonexistent scripts/git-hooks/pre-push-plugin-version and scripts/bump_plugin_version.py replaced with the real rule: bump the manifest version and run update_registry.py. scripts/dev/README.md and CLAUDE.md updated to match. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(scripts): monorepo workspace layout; fix_perms and install READMEs MULTI_ROOT_WORKSPACE_SETUP described one sibling repository per plugin; setup_plugin_repos.py links ../ledmatrix-plugins/plugins/* into plugin-repos/ and update_plugin_repos.py pulls only the monorepo, and the workspace file opens LEDMatrix plus ../ledmatrix-plugins. scripts/fix_perms/README.md listed cache directories fix_cache_permissions.sh never touches and a 'ledmatrix' service user that doesn't exist (also in scripts/install/README.md); adds safe_pip_install.sh. install/README: install_service.sh installs the web and update-verify units too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): keep the rollback's pip retries inside the unit time limit The health check reinstalled the previous requirements by trying the next bash path after any failure, including a 600 s pip timeout. Two files, two paths: up to 40 minutes of pip alone, while systemd stops ledmatrix-update-verify.service at TimeoutStartSec=30min -- killing the rollback half-way and leaving the update 'verifying' until the web UI calls it lost. - Like permission_utils.install_requirements_file, only a sudo refusal moves on to the next bash; a pip that ran and failed or timed out is not repeated. The refusal wording is one list (permission_utils.SUDO_REFUSAL_PHRASES), mirrored in the stdlib-only verifier and pinned equal by a test. - All reinstalls in one rollback share a 600 s budget. - WORST_CASE_SECONDS adds up every timeout on the longest path (27.5 min); a test holds it under the unit's TimeoutStartSec and that under the web UI's VERIFY_LOST_SECONDS. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(plugins): prepare plugin configs one way for load, saves, GET, hot reload and dev tools Plugin config was prepared differently depending on how it arrived: - JSON POST /plugins/config built a partial body on schema defaults, so {"enabled": true} reset every other setting of the plugin. It now merges onto the stored section first, as the form path already did. - Legacy-boolean normalization (#588) ran only at load: GET /plugins/config returned the raw boolean, posting it back failed validation, and hot reload handed plugins the raw section (a legacy dynamic_duration: true came back as a boolean). schema_manager.prepare_plugin_config (normalize, then defaults) is now used by PluginManager.load_plugin, both save paths, GET, the save notifications and DisplayController's hot-reload callback. - The JSON save's filter kept only enabled/display_duration/live_priority and dropped a submitted skin, skin_options or vegas_* tuning key. There is now one core-owned per-plugin list, schema_manager.CORE_PLUGIN_PROPERTIES, used by validation and by the save filter; PluginManager's CORE_OWNED_CONFIG_KEYS is its vegas subset. - Plugin sections posted to /config/main were stored verbatim, including values /plugins/config rejects. They now go through the same preparation (_prepare_plugin_config_for_save, extracted from save_plugin_config), and a failing section rejects the whole save before anything is written. - dev_server read only top-level defaults and let a schema enabled:false win; build_full_config shallow-merged overrides, dropping sibling defaults; the harness extracted defaults differently from the device. loading.build_config now uses the device's extraction and preparation, and dev_server, check_plugin, render_plugin and the harness all use it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt-bridge): brightness changes apply live and touch nothing else The display service's hot reload applies a saved brightness within a few seconds, and /config/main no longer resets other display settings on a brightness-only JSON body. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): automatic update hardening Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): rewrite PLUGIN_CONFIG_ARCHITECTURE for the v3 web UI It described web_interface_v2.py and index_v2.html (both gone), client-side form generation, one POST per field with {key, value}, and 'no nested objects'. The v3 UI renders plugin forms server-side from the schema (pages_v3 partial + plugin_config.html macros, nested sections and x-widgets), posts the whole form once, and save_plugin_config() merges onto the stored section, validates, splits x-secret fields and notifies the plugin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(mqtt): brightness saves apply via hot reload and leave other settings alone The bridge README said brightness is applied on the display's next restart; the display controller's config hot reload applies it within seconds. It also now states that the bridge's partial JSON save changes only brightness (the /config/main merge fix in this change set). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(update): don't log pip's output from the health check's reinstall pip can echo a private index URL with embedded credentials; permission_utils redacts it, the stdlib-only verifier cannot, so it logs the exit code only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(config): mark the plugin_system toggles as unused legacy keys auto_discover, auto_load_enabled and development_mode are read by nothing and leave the General tab in this change set (F40). CONFIG_REFERENCE said they were read by the plugin loader; PLUGIN_CONFIGURATION_GUIDE and the REST reference listed them as live settings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): docs and developer tools group Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): legacy plugin-system toggles no longer count as a General save auto_discover, auto_load_enabled and development_mode have left the General form, so a post carrying only one of them is not a general-settings save and must not treat web_display_autostart and auto_update as unchecked. The plugin_system block itself is left as on main for the branch that reworks it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(changelog): config-save and plugin-config preparation fixes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(claude): re-check matrix_support.py rules when the library submodule is bumped Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address Codacy findings on the core audit PR - plugin_manager.prepare_plugin_config: when the fallback legacy-boolean pass also fails, log a warning instead of a bare except/pass. - api_client.js: request() refuses any endpoint that is not a plain path under /api/v3 ("//host", backslashes, ".." or "." segments, whitespace, control characters) with INVALID_ENDPOINT before calling fetch(), and plugin ids are URL-encoded wherever they are put into a URL (also in the app-shell batch load). - test_update_all.js: pins both against the shipped client. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(web): check endpoint control characters without a control-character regex Codacy (ESLint no-control-regex, Biome noControlCharactersInRegex) flags the \x00-\x1f range in checkEndpoint's regex. Test the char codes instead; the endpoints refused are unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(auto-update): make the seed script executable on disk, not only in the index On Linux Repo.publish() commits with -a, which recorded scripts/run.sh as 100644 upstream because the seed file was never chmod +x. The pull then brought in the same mode the installer chmod had made locally, so installer_chmod saw no mode change left to check. The updater was fine: with the upstream commit at 100755 the --autostash carries the device's chmod across. Verified under Linux (WSL, git 2.43): the old helper fails exactly as CI did, the fixed one passes all 63 tests in the file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
869e36fb2f |
feat(web): weekly automatic updates with health check and rollback (#581)
* feat(web): weekly automatic updates with health check and rollback A General-tab toggle (off by default) checks for and installs LEDMatrix and plugin updates once a week, overnight in the configured timezone. - Pre-update checks skip (and report) instead of forcing: local edits or commits, merge/live rebase, no upstream, low disk, missing health check, or a version that was already rolled back. An abandoned rebase (HEAD back on a branch) is cleared, since it would otherwise block every pull. - The pull reuses the Update Code path (now perform_core_update(), which reports dependency install failures as data). - ledmatrix-update-verify.service, started via a .path unit from a request file, restarts the services from its own cgroup, requires them to come up and stay up, and otherwise resets to the previous commit and reinstalls the previous requirements. It runs a copy of the checker taken before the pull. - No SSH needed: switching the toggle on restarts the display service, which (as root) installs the two units from the repo templates for the web user. first_time_install.sh installs them too and takes --enable-auto-update / LEDMATRIX_AUTO_UPDATE (passed through by one-shot-install.sh). - Plugins update after the code passes its check; failures, blocks and rollbacks raise an Overview banner and show under the toggle. Tested end to end on a Pi: web-UI setup, a good update, a broken web service and a broken display (both rolled back), a blocked local edit, and an abandoned rebase found on the device. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(auto-update): address static-analysis findings - Replace the subprocess.CompletedProcess the verifier fabricated for a command that could not start with a plain namedtuple; nothing is executed there, but the scanner flags any CompletedProcess built from variables. - Mark the subprocess imports with the repo's standard B404 annotation (all calls are list-form argv, no shell). - Mark the rollback-failed message as not SQL (B608 matched its wording). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): CI failures on Linux - Keep the setup result when chown fails. CI runs as a non-root user, where chown to the web user raises; that discarded the result file, so the General tab would never learn whether setup worked. Regression test added. - Register the two new /api/v3/system/auto-update routes in the URL map snapshot. - Use utility classes app.css defines (space-y-1, hover:text-red-600). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): address review feedback - Health check: a failed restart command no longer lets the check run against the still-running old process; it counts as a failure (and after a rollback, as a failed rollback). An unreadable restart count is never treated as stable, since a crash loop looks healthy between attempts. - Installer writes the auto_update setting to a temp file and swaps it in, keeping mode and owner, so a running config watcher never reads a truncated config.json. - Verify unit quotes its command-line paths (install folders with spaces); setup refuses folder names systemd would reinterpret (%, quotes, backslashes, control characters) and says so on the General tab. - The auto-update status route no longer returns exception text. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): keep error detail in the status route's 500 test_web_error_detail requires every 5xx handler to log the traceback and return describe_exception(e), which redacts credentials, so failures are diagnosable from the web UI. Dropping it for CodeQL broke that policy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): dismiss route rejects non-object JSON with 400 A JSON array or scalar body made `.get('alert_id')` raise, returning 500. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(auto-update): let the app-wide handler answer status-route errors CodeQL (py/stack-trace-exposure, #709) flagged the route's own except, which returned describe_exception(e). web_interface/app.py's error handler already logs the traceback and returns the same redacted detail for any unhandled exception, so the local copy is removed: same response, no new exception-to-response flow, and test_web_error_detail's policy still holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
577f5501a6 |
perf(plugins): stop re-deriving a display() signature the caller already cached (#549)
display_controller resolves once, and caches, whether a plugin's display() takes a display_mode keyword -- self._plugin_accepts_display_mode, populated right before the dispatch. It then handed the executor a types.SimpleNamespace wrapping a closure, and execute_display() ran inspect.signature() on that to work out the same thing. Because the SimpleNamespace is rebuilt per call, the callable was new every time, so nothing inside the executor could ever cache it either. Measured at ~39us per dispatch on a Pi 4, for a value the caller had a line earlier. execute_display() now takes accepts_display_mode, falling back to inspecting only when a caller does not pass it, so existing callers are unaffected. Also documents two things that read as bugs and are not: - execute_with_timeout()'s timeout is advisory. Nothing cancels the thread -- Python cannot -- so on expiry the operation runs to completion in the background and only the caller gives up. A permanently hung plugin leaks a daemon thread per attempt. This is why callers holding a lock across the call must release it from inside the wrapped callable, as run()'s _release_display_lock already does. - Only the first display() of each mode goes through the executor; the per-frame loops call display() directly. That is deliberate: a thread per frame would cost more than an advisory timeout buys. Both loops now say so, so the asymmetry does not read as an oversight. Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |