mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
dfd67c7c8b8b01ae776616e17d2e103e01e8c33a
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7f9c73e9aa |
fix(vegas): smooth Vegas scroll pacing -- whole pixels per refresh, measured refresh, off-thread preview writes (#628)
Vegas scrolls a whole number of pixels per panel refresh, locked to SwapOnVSync, against the refresh the panel really holds (measured from swap gaps), instead of blending sub-pixel positions against the refresh cap. The web preview PNG is encoded off the render thread while scrolling, with writes ordered and retried. On hdpi, late frames fell from 6.3% to 0.7%. See docs/SCROLL_PERFORMANCE.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b9416ef803 |
fix(display): Vegas teardown and 240s default; remove dead Vegas buffer code (#637)
* fix(display): tear down Vegas mode on controller cleanup DisplayController.cleanup() never called VegasModeCoordinator.cleanup(), so the Vegas teardown (stop, pipeline/stream reset, adapter cache drop) was unreachable. Call it before the display manager is cleaned up, and skip it when Vegas was never created. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): default max_cycle_duration to the documented 240s The template, the web UI help, CONFIG_REFERENCE and the controller all say 240, but the code defaulted to 600 in two places, so a config without the key ran Vegas iterations 2.5x longer than documented. from_config now falls back to the dataclass field defaults instead of repeating each one, so the two copies can no longer drift, and the controller's follower scroll-speed default reads VegasModeConfig's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): let run.py -d show display_manager's DEBUG output display_manager pinned its logger to INFO at import, overriding the root level, so debug mode never showed its DEBUG lines. Use get_logger() from src.logging_config like the rest of the core and leave the level to the logging setup. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): run each startup validation check once StartupValidator.validate_all() ran twice at boot, before and after the plugin manager was created, so every config, cache, display and systemd-unit warning was logged twice. The second pass now runs only the plugin checks. Drop the commented-out raise_on_errors line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): one INFO line per plugin-list refresh StreamManager logged "=" * 60 banners and a line per plugin (INCLUDED, SKIPPED, FETCHING CONTENT, SEGMENT CREATED) at INFO on every refresh and fetch, i.e. at each cycle start and every 30s. Log one INFO summary of the rotation per refresh and move the per-plugin detail, the weighting breakdown and "no content this cycle" to DEBUG (the adapter still warns when every content path fails). Also drop the check/cross marks from the controller's log messages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): drop the per-iteration static-mode plugin scan run_iteration() rebuilt _static_mode_plugins on every iteration, asking every plugin for its display mode and logging the set at INFO, but nothing ever read it: static pauses are triggered by _check_static_plugin_trigger() from the next segment. Delete it, the coordinator's get_ordered_plugins() that only it used, and the write-only _static_pause_plugin / _static_pause_start. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(vegas): remove the staging buffer that was never filled StreamManager and RenderPipeline carried a double-buffer design that nothing used: _staging_buffer was only ever cleared or swapped, so swap_buffers() never did anything and should_recompose()'s staging_count > 0 branch was dead, and _active_scroll_image, _staging_scroll_image, _is_rendering, _last_frame_time and _frame_interval were written but never read. Delete the machinery and rewrite the docstrings around what actually carries updates: _pending_updates, consumed by process_updates() in swap mode and invalidate_pending_updates() in continuous mode. should_recompose() no longer builds a buffer-status dict every frame. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): tidy the display controller without changing behaviour - Import VegasModeCoordinator locally instead of through module globals (there is no circular import to avoid). - Drop hasattr() checks on attributes PluginManager.__init__ always sets (plugin_executor, plugin_last_update, get_plugin_lock, run_scheduled_updates*, stop_update_worker) and the dead "older manager" fallbacks; keep the health_tracker None checks, now via _health_tracker(). - Extract _display_once() for the per-frame display call both render loops copied, _advance_on_demand() for the two on-demand rotations, _reset_on_demand_fields() for the error and clear paths, and _timezone() / _in_window() for the two schedule checks. - Remove always-true conditions and the unreachable non-plugin else branch in run(), and read _was_display_active / _last_published_mode / vegas_coordinator directly now that __init__ declares them. - Declare the follower render state in __init__, name its tuning constants, add _follower_sign(), and share the 90/s sync send interval with the render pipeline (SYNC_SEND_INTERVAL). - Delete history narration and the "Opt #N" labels, fix the comment that called _scroll_speed constant (hot reload updates it), and drop a startup timing log that measured nothing. - render_pipeline / plugin_adapter: read display_manager.width/height as the properties they are, drop an empty TYPE_CHECKING block, an aliased threading import and a duplicated `if result and self.sync_manager:`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(display): trim dead code from display_manager - Add _new_canvas() for the image/draw/fontmode="1" setup that was copied six times. - Call resolve_double_sided() and compose_pixel_mapper_config() directly instead of through a module alias and a passthrough method, and replace the comment that said the passthrough read class attributes. - Delete the unused _initialized flag and _ORIENTATION_ROTATE_DEGREES alias (no core or monorepo reader; tests stop resetting the flag), the test pattern's unreachable no-matrix branch (it only runs once the matrix exists), `del old_image # help GC` (a no-op on a local), a duplicated early return in process_deferred_updates, and stale comments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(vegas): remove unread fields and test-only helpers, fix docstrings - ContentSegment: drop total_width, fetched_at, is_stale, image_count and is_static, none of which is read. - StreamManager: drop _current_index (never advanced) and the test-only get_all_content_for_composition() and has_pending_updates(); VegasModeConfig: drop the test-only is_plugin_included(). - geometry.find_blank_cut() has had no production caller since the crop moved to item boundaries; delete it and its tests. - PluginAdapter: the _finalize docstring described separator_width between every image, and _crop_to_budget's said cuts snap to the nearest blank column; both now describe what the code does. - Coordinator: the static-pause interrupt log no longer blames follower mode for every interrupt, and set_update_callback names the callback the controller actually wires. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(scroll): correct ScrollHelper comments and drop dead branches - Four comments said the strip always starts with display_width of blank; it does only when lead_gap is None (Vegas passes its own). - Delete the "Width calculation mismatch" warning: the image is created at the calculated width, so the two can never differ. - Remove the two scroll_delay <= 0 fallbacks (which disagreed with each other): set_scroll_delay clamps it to at least 0.001 and nothing in core or the plugin monorepo assigns it directly. - Trim the scipy history from the blend docstring. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * refactor(run): drop a redundant comment Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(display): log set_scrolling_state only when it changes Vegas and scrolling plugins set the scrolling state every frame, so once display_manager's DEBUG output became visible in debug mode it printed "Scrolling state set to: True" about 120 times a second. Log only when the value differs from the previous one; the state, activity timestamp and frame hold still update on every call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(changelog): display-vegas Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
6b3028ad58 |
test: isolate DisplayManager globals across modules, and name the failure (#563)
Follow-up to #562. That commit fixed the actual cause of the intermittent 15-test failure in test_display_dirty_tracking.py -- the emulator's fixed TCP port 8888, a machine-wide singleton that a concurrent pytest process takes away. This adds the two things that would have made it a five-minute diagnosis instead of a long one, and closes the other door into the same failure. Confirmed the module is order-independent as it stands, on this checkout: pytest test/ -q, three times 115 failed / 4464 passed / 63 skipped, byte-identical failure sets, the module 21/21 passed each time module forced last (197 files first) identical failure set module forced first identical failure set module after each of test_display_manager, test_display_controller, test_display_controller_vegas_tick, test_skin_system, test_sports_scroll, test_initial_update_budget, test_display_double_parity, test_initializing_screen all pass four concurrent processes on the file 21/21 each And reproduced the original, to be sure the diagnosis in #562 is the whole story. Holding 0.0.0.0:8888 from a separate process: HEAD's test/conftest.py 21 passed pre-#562 test/conftest.py 15 failed, 6 passed The 15/6 split is not arbitrary: the six survivors are the only tests in the file that never touch dm.matrix. conftest.py: DisplayManager is a process-wide singleton and the RGBMatrix / RGBMatrixOptions names it constructs through are module globals, bound once at import. All three are shared by every test module in the run, so a module that leaves an instance in _instance -- or leaves patch('src.display_manager. RGBMatrix') standing -- changes what the NEXT module builds, invisibly, and only in a full run. A module-scoped autouse fixture now resets the singleton and restores either binding if a patch outlived its module. Module-scoped rather than per-test so that files sharing one manager across their own tests keep doing so; only the leak across the module boundary is cut. Autouse fixtures are set up ahead of requested ones, so this is finalised after a module's own DisplayManager fixture. Verified with a throwaway pair of probe modules -- one leaks a patch and a singleton, the next asserts both are clean -- which passed and were then removed. test_display_dirty_tracking.py: _setup_matrix() swallows every construction failure and falls back to matrix=None, so a broken environment arrived as fifteen identical "'NoneType' object has no attribute 'SwapOnVSync'" errors naming neither the fixture nor the cause. The fixture now fails once, and says where to look; under a held port it reads DisplayManager fell back to matrix=None: RGBMatrix construction raised... Known causes: the emulator adapter losing a fixed TCP port to another process -- see pytest_configure in test/conftest.py -- or a patch('src.display_manager.RGBMatrix') leaked from an earlier test module. with WinError 10048 in the captured log directly above it. No regressions: full suite with both changes is 115 failed / 4464 passed / 63 skipped, failure set identical to the pre-change baseline. The 115 is the pre-existing Windows-environment baseline (os.geteuid, POSIX modes, fcntl); CI on Linux remains authoritative. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
59997594ac |
test: fix the emulator port collision behind the intermittent suite failures (#562)
* chore: stop tests and rigs writing to shared paths Two shared-state problems, both of which show up as a permanently dirty checkout or an unreproducible test failure. test_display_dirty_tracking.py builds a real DisplayManager, whose _snapshot_path defaults to the fixed /tmp/led_matrix_preview.png that the web UI reads. Every pytest process on the machine shares that one file, so two concurrent runs -- CI shards, a second worktree, an agent running the suite alongside -- overwrite each other's snapshot and the mtime assertions stop meaning anything. The module fixture now points it at a session-unique temp path; the individual tests that care still override it further. To be clear about what this does and does not fix: this is a real shared-path hazard, but it is NOT the cause of the intermittent 15-test failure in that module. That turned out to be the emulator's fixed TCP port, fixed in the follow-up commit. This change stands on its own merits. web_interface/app.py writes data/plugin_operations.json, data/plugin_state.json and data/operation_history.json as the web interface runs, into a directory that ships tracked (data/.gitkeep) and was otherwise unignored. So every rig that ever opened the web UI -- and every test run that constructs the app -- left three untracked files behind and a permanently dirty `git status`. Only data/.gitkeep is tracked under data/, so the negation keeps it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: stop the emulator binding a fixed port, so concurrent runs can't collide This is the cause of the intermittent full-suite failures we have been chasing: runs of identical code landing anywhere between 100 and 130 failures, while every implicated test passed in isolation. Six test modules set EMULATOR=true and build a real DisplayManager. The repo's emulator_config.json selects the "browser" adapter, which binds TCP port 8888 to serve the dev preview. That port is a machine-wide singleton, so a second pytest process -- a CI shard, another worktree, an agent running the suite alongside -- loses the bind. RGBMatrix construction then raises, DisplayManager catches it and falls back to `self.matrix = None`, and every test that subsequently touches the matrix dies with AttributeError: 'NoneType' object has no attribute 'SwapOnVSync' which names neither a port nor a socket, and points at the wrong file entirely. Because test_display_dirty_tracking's fixture is module-scoped, all 15 of its matrix-touching tests fail together or not at all -- the 15-test swing that made the totals look random. Demonstrated rather than assumed. Holding 0.0.0.0:8888 from a separate process and running test_display_dirty_tracking.py: without this change 15 failed, 6 passed with this change 21 passed The "raw" adapter renders in memory and binds nothing. Only display_adapter is overridden, in a throwaway config written per pytest process; the repo's emulator_config.json is untouched and `run.py -e` still opens the browser preview on 8888. Nothing in the suite referenced the adapter, and the tests wrap SwapOnVSync on the matrix object itself, so they are indifferent to what sits underneath. allow_adapter_fallback is forced off -- falling back would land us on the browser adapter and its fixed port, which is the whole problem. CONFIG_PATH is a bare relative filename resolved against the CWD, so it is set to an absolute path: the previous behaviour depended on where pytest was invoked from, and silently wrote a default config into whatever directory that was. Verified no regressions: full suite on this branch and with origin/main's versions of the touched files, same machine, back to back -- 115 failed / 4347 passed on both sides, zero failures unique to either. That 115 is the pre-existing Windows-environment baseline (POSIX file modes, fcntl, shell scripts, Linux-only binaries); CI on Linux remains authoritative. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: mark the shell entry points executable Eleven scripts shipped as 100644, so `./scripts/install/configure_web_sudo.sh` fails with "Permission denied" and only works if you know to prefix `bash`. That one matters most: the web UI's own error hint, added in #560, tells users to run exactly that path when a system action fails for want of passwordless sudo, and following that instruction verbatim did not work. All eleven carry a shebang and are invoked directly, never sourced. The two sourced libraries -- lib_lowmem.sh and lib_systemd_render.sh -- are deliberately left non-executable, which is what distinguishes a library from an entry point. Mode bits only, no content: 11 files changed, 0 insertions, 0 deletions. Applied with `git update-index --chmod=+x` because this checkout is on Windows, where core.fileMode is off and the working-tree bit is not tracked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f8e2e89edc |
refactor(sports): put the scoreboards on the shared scroll resolver (#542)
* refactor(sports): put the scoreboards on the shared scroll resolver Eight sports scoreboards -- afl, baseball, basketball, football, hockey, lacrosse, nrl, soccer -- scrolled through this module's own pacing while the other eleven scrolling plugins went through src/common/scroll_config. Two implementations of the same job, and this one was on the losing side of every difference. It never called set_scrolling_state. Two consequences, both of which this release's work was about: - The frame hold is applied through that call, so a speed the crisp ladder could render in whole pixels still presented a new frame every refresh. - Core only runs deferred updates while nothing is scrolling. Believing nothing was, it ran blocking work in the middle of these scrolls. The default is non-crisp today: scroll_speed 50.0 with scroll_delay 0.01 is 50 px/s, which on a 100Hz panel is half a pixel per refresh. That cannot render as motion -- it alternates 0px and 1px steps and judders at a 50Hz beat, on every scoreboard, out of the box. Resolved through the ladder it stays 50 px/s and holds each frame for two refreshes: same speed, whole-pixel motion. The stepping disagreement that used to justify a separate module is gone. scroll_config avoided frame-based mode because it stepped on a wall clock at 1/scroll_delay with scroll_delay set to the frame period, so the decision sat on its own threshold and flipped on sub-millisecond jitter. That branch now accumulates elapsed time, identical arithmetic to the time-based one, so the two differ only in the units the speed arrives in. What is NOT shared, and must not be: the two modules read identically-named keys with different meanings. Here scroll_speed is px/SECOND and scroll_delay only converts to px/frame; in scroll_config scroll_speed is px per STEP, so px/s is speed/delay. Passing this module's settings dict to the resolver turns 50 px/s into 5000, clamped to 500 -- a tenfold speed-up everywhere. So _get_scroll_settings keeps sole ownership of reading sports config, including the league merging, and hands the resolver a plain px/s. A test pins that specific number, because it is the mistake the refactor invites. MIN/MAX_PIXELS_PER_FRAME are gone; the resolver bounds speed and the helper clamps FPS. _resolve_target_fps stays, re-purposed: under the old model that key was the rate frames were presented at, so it is the faithful translation into the refresh the ladder is computed against, used when no hardware refresh is configured. Speed changes for panels that are not 100Hz: 50 px/s becomes 60 at 60Hz (+20%) and 48 at 120Hz (-4%). At 100Hz it is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): drop the frame hold when a scroll times out, not just when it says so set_scrolling_state(False) clears the hold. The other way a scroll ends is is_currently_scrolling() deciding, after scroll_inactivity_threshold of silence, that it is over -- which is what happens when the rotation moves on mid-scroll or a plugin is torn down. That path cleared the flag and kept the hold, so every later plugin, scrolling or static, was presented at refresh/N by whoever scrolled last, until something called the explicit stop. The method's own docstring already states the rule this breaks: the hold "must not outlive the scroll that asked for it". The timeout was the exception it did not cover. Pre-existing, but reachable by three plugins before and eleven after the sports scoreboards moved onto the shared resolver, so it belongs with that change. The test ages the activity timestamp past the threshold rather than sleeping. Also adds scripts/sports_scroll_check.py. The sports scroll path is per-league opt-in, so a rig showing static game cards never constructs a SportsScrollDisplay and none of its pacing can be observed from a normal run -- which is exactly what happened when this change was first put on hardware: 26 minutes, zero sports scroll lines. The script drives the path directly with synthetic games and asserts the three things the resolver is meant to buy: the speed lands on whole pixels, the hold is published, and it is released after. It never starts or stops the display service, matching scroll_speeds.py, so a crash here cannot leave the panel dark. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scripts): refuse to grab the panel while the display service has it The module docstring already said to stop ledmatrix first. Nothing enforced it, and running the script against a live service is not a harmless mistake: rpi-rgb-led-matrix configures GPIO directions and the hardware PWM inside RGBMatrix(), and when the root check fails it calls exit() from C with no cleanup. The service keeps rendering and swapping onto pins that have been reconfigured underneath it, so the panel goes black while every diagnostic says the display is healthy -- fresh framebuffer, every pixel lit, "RGB Matrix initialized successfully", nothing in the log. A restart fixes it, once you work out that is what happened. Found the hard way: this is what took the panel down on the test rig, not the change the script was written to verify. --fallback skips the check, since it never opens the matrix. --force is there for anyone who means it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(scripts): annotate the subprocess call the way this repo already does Codacy fails a PR on one new issue, and bandit B404 fires on any subprocess import. scripts/run_plugin_tests.py carries the same suppression with the same justification -- list-form argv, no shell -- so this follows it rather than inventing a second convention. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d12323e7f1 |
perf(scroll): pace frames to the panel — 44→100 fps, stalls 14% → 0.02% (#523)
* perf(scroll): pace frames to the panel, not to a fixed sleep Scrolling ran at 44-46 fps on a 2x128x64 chain and 14-17% of frames took 41-53ms, which reads as judder. Four independent causes, each measured on the hardware; details and the diagnostic recipe are in docs/SCROLL_PERFORMANCE.md. The high-FPS loop slept a flat 8ms after every render. display() has already blocked on the panel's vsync by then, so that sleep was added to a wait that had happened: ~4ms of render plus 8ms put each iteration at ~12ms against a 10ms refresh grid, so every swap missed a refresh and the loop settled at 50fps while asking for 125 -- with no headroom, so a further 14% of frames slipped again. It now sleeps only the remainder, with a 1ms floor so plugin threads still get the GIL. ScrollHelper stepped position on a wall clock at 1/scroll_delay steps per second. Plugins set scroll_delay to the frame period, so that comparison sat exactly on its own threshold: a frame arriving a hair early moved zero pixels and rendered an identical frame, dirty-tracking skipped the swap, it returned in ~2ms, and the beat repeated. No scroll_delay value tunes that out -- a shorter delay trades stalled frames for periodic double-steps. Both modes now accumulate elapsed time at the same configured speed, so position stays proportional to real time. Sub-pixel blending goes back to off by default. It renders a half-step by mixing two adjacent columns, which on a coarse panel showing pixel-font text alternates crisp and smeared frames and reads as shimmer -- visibly worse than integer stepping on the hardware. Vegas mode still opts in. disk_cache uses orjson when importable, falling back to the stdlib. Encoding a ~1MB record drops from 14.8ms to 5.4ms end-to-end, and that work holds the GIL while a marquee is on screen. display_manager also checksummed the whole framebuffer twice per frame (dirty tracking, then the preview snapshot); the snapshot now takes the checksum the caller already computed. New src/common/scroll_config.py resolves scroll settings in one place. Five ticker plugins each hand-rolled this and disagreed: odds-ticker ranked the deprecated scroll_pixels_per_second above the documented scroll_speed/delay pair, and because that key carries a schema default the documented settings were dead for every user (ChuckBuilds/ledmatrix-plugins#408), while ledmatrix-leaderboard read the same key only as a fallback. The resolver also warns when a speed will not advance a whole number of pixels per refresh, which is the property that actually determines whether a scroll looks smooth. scripts/build_rgbmatrix_nogil.sh rebuilds the rgbmatrix binding so it releases the GIL. Upstream declares SwapOnVSync without nogil, unlike SetPixel/Clear/Fill beside it, so the render thread held the GIL for the whole vsync wait and starved background threads into long uninterruptible bursts. The script patches, builds and self-verifies into a scratch tree; --install backs up the original and rolls back if the service does not come back healthy. Measured after: 100 fps locked, no stalls observed, render thread down from 51% to 19% of one core. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(display): keep the panel swap locked to vsync while scrolling Dirty tracking skipped SwapOnVSync for byte-identical frames. That is the right call for static content, but SwapOnVSync is also what paces the render loop, so skipping it skips the wait for the panel: a duplicate frame returns in ~8ms instead of ~10ms on a 100Hz panel, advances the strip only 0.8px instead of 1.0px, and so makes the next frame more likely to repeat as well. The effect sustains itself once it starts. Measured over 20 minutes on a 2x128x64 chain, both scrollers configured identically at 100 px/s: leaderboard 10ms x35, 11ms x3 (clean) odds-ticker 10ms x26, 8ms x7, 15ms x5 (~20% duplicates mid-scroll) The duplicates were not end-of-cycle idling -- 38% of fast frames fell within 90s of a scroll completion against 35% of normal frames, a null result. The trigger is per-frame work: odds does more of it, and more variably, so it is first to land a frame that advances less than a whole pixel. Pushing an identical frame costs one canvas copy. Falling out of vsync lock costs smooth motion. Static content is untouched, because is_currently_scrolling() expires on its own inactivity threshold -- covered by test_stale_scrolling_state_stops_forcing_pushes so a plugin that stops scrolling without saying so cannot pin the panel into always-push. Also de-flakes test_snapshot_still_written_on_skip, which asserted a strict mtime increase between two writes that can land in the same filesystem tick; it failed about two runs in three on Windows regardless of the code under test. The file is now backdated before the check. 156 tests pass on the Pi. Not yet confirmed by eye on the panel. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): report the frame-time tail, and stop the row-major blit Two problems, both found by looking at the panel rather than the metric. The frame-stats line reported ONE instantaneous frame every 5 seconds -- about 1 frame in 500 -- printed beside a 100-frame average. Both hide exactly the fault they are used to chase: a 2ms duplicate and a 21ms double-wait average to precisely 10ms, so a ticker stalling on half its frames still reports a healthy "Avg FPS: 100.0". That reading cost several rounds of chasing the wrong layer. The line now aggregates every frame since the last log and reports median, p95, max, min, and explicit stall and skip rates (past 1.5x the median missed a refresh; under half never reached the panel, because dirty tracking skipped the swap so the frame never waited on vsync). On the hardware this now reads: leaderboard 100.0 fps over 501 frames | median 10.00ms p95 10.05ms max 10.34ms | stalls 0 (0.0%) skips 0 (0.0%) The binding rebuild's blit patch becomes opt-in (RGB_PATCH_BLIT=1, default off). Reordering that loop to row-major changes what a torn frame looks like: column-major tearing shows as a vertical seam, row-major as a horizontal split between the panel's upper and lower halves. On a 1/32 scan panel that reads as a one-pixel fold across the middle of every panel, which is what was reported on hardware and what went away when the blit was reverted. All of the measured gain comes from the SwapOnVSync change, so the risky half is simply not worth taking; the header says so. Also fixes --install resolving its paths against $HOME, which is /root under sudo, so it looked in /root/rgbmatrix-nogil-build and died with "no built module found" on a machine where the build had just succeeded. It now resolves SUDO_USER's home. Both build paths are verified on the Pi: default yields one GIL-release site, RGB_PATCH_BLIT=1 yields two. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(scroll): let users pick a crisp speed for their own panel Whole-pixel motion was previously only available at multiples of the refresh rate -- 100, 200, 300 px/s on a 100Hz panel. 100 px/s crosses a 256px panel in 2.6s, which is brisk for reading, and everything slower had to blend (blur) or repeat frames unevenly (judder). There was no way to ask for 50 px/s and get clean motion. SwapOnVSync takes a framerate_fraction the display manager never passed. It holds each frame for N panel refreshes; the panel keeps refreshing at its full rate throughout, so holding costs nothing in flicker and only changes how often a NEW image is presented. That turns 50 px/s into one whole pixel every second refresh instead of half a pixel every refresh. The crisp speeds are therefore refresh_hz / hold * pixels_per_frame, and that ladder depends on the panel: a Pi Zero on a long chain has a different set of good speeds from a Pi 4 on a short one. crisp_ladder() enumerates them and solve_crisp() picks the best match for a requested speed. solve_crisp weights motion quality rather than picking the numerically nearest entry, which matters more than it sounds. Asked for 30 px/s, nearest-by-value answers 28.6 -- 2px jumps at 14fps -- over 33.3, which is single-pixel motion at 33fps and obviously better on the panel. The target is also clamped into the ladder's range first, because relative error saturates near 1.0 for a target far outside it and the quality penalty would otherwise answer "10000 px/s" with the slowest entry. configure() snaps to the ladder and applies the hold when given a display manager. Without one the hold silently cannot happen and motion falls back to fractional pixels, so it warns rather than failing quietly. set_frame_hold() resets to 1 when scrolling stops, so one plugin's pacing cannot leak into whatever is on screen next. scripts/scroll_speeds.py is the user-facing part: it prints the ladder for the configured rate, measures what the panel ACTUALLY manages (--measure, for hardware that cannot reach its configured limit), highlights the nearest option to a wanted speed, and demos one live. It never starts or stops the display service itself -- doing that inside a script stranded the panel twice today. Speeds below ~20 px/s remain stepped regardless. That is the pixel pitch, not a software limit. Also fixes the dirty-tracking test spy, which stubbed SwapOnVSync with a single-argument function and would have masked the new call as a failed push, and rewrites a configure() test that had started passing for the wrong reason: it asserted a judder warning, which snapping now prevents, and was matching the unrelated "hold could not be applied" warning instead. 183 tests pass on the Pi. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll): tie the frame hold to the scroll, not the plugin The hold applied in configure() never reached the panel. Plugins share one display manager, and set_scrolling_state(False) -- fired whenever ANY other plugin finishes its scroll -- reset the hold to 1. A hold set once at plugin construction was therefore always gone by the time that plugin rendered. The symptom was a log line that lied. ledmatrix-stocks reported Scroll configured: 50.0 px/s (1px every 2 refreshes = 50.0 fps, smooth) while the panel measured 100.0 fps, median 10.00ms. Config, resolution and snapping were all correct; only the pacing silently was not applied. set_scrolling_state(is_scrolling, frame_hold=1) now carries it, so the hold lives exactly as long as the scroll that asked for it. configure() reports the value as ScrollSettings.frame_hold instead of applying it -- applying it behind the caller's back could never have been right on a shared display manager. Existing callers are unaffected; the default keeps one frame per refresh. Verified on hardware: stocks at 50 px/s now measures 50.0 fps over 251 frames | median 20.00ms p95 20.09ms | stalls 0 skips 0 20.00ms being exactly two refreshes, with the panel still refreshing at 100Hz underneath so flicker is unchanged. test_another_plugin_stopping_does_not_strand_a_hold pins the interaction that broke this. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(scroll,cache): resolve CodeRabbit review on #523 Eight findings, all reproduced before fixing. scroll_config.configure() read the refresh rate *after* resolve() had already used it. resolve() fills in target_fps, pixels_per_frame and the judder warning from that rate, so on a 60Hz panel every one of them described 100Hz -- and with snap_to_crisp=False nothing downstream corrected it, so set_target_fps() paced the helper to 100 FPS. The rate is now settled first, and falls back to the global config rather than straight to the default. refresh_hz_from_config() used `(cfg.get("display") or {}).get(...)`, which raises AttributeError when either level is truthy but not a mapping -- out of a function whose whole contract is a rate or a default. The frame-stats line reported the upper-middle sample as the median and the 96th sorted sample as p95 of 100. Both are also thresholds (stalls at 1.5x the median, skips at 0.5x), so the counts were biased too. The arithmetic is now in frame_stats()/format_frame_stats(), testable without a clock. configure()'s docstring and docs/SCROLL_PERFORMANCE.md still said it applies the frame hold and warns when it cannot. It deliberately does neither since "tie the frame hold to the scroll, not the plugin"; a caller following the old text would omit set_scrolling_state() and slow snapped speeds would still present every refresh. disk_cache had no policy for non-finite floats: orjson writes null, the stdlib writes NaN/Infinity, and orjson then rejects those legacy files so DiskCache.get deleted them as corrupt. One behaviour on both paths now -- write null, keep legacy records readable. allow_nan=False detects the values; the replacement walk runs only when there is one, so the ordinary write path is byte-identical and pays nothing. build_rgbmatrix_nogil.sh picked the build artifact with a glob piped to `head -1`, which sorts cpython-311 ahead of cpython-313, so a stale .so staged in from the source tree was installed as core.so while the GIL check -- which reads the generated core.cpp, not the .so -- still passed. It now requires the current interpreter's exact ABI name and fails closed. Its systemctl calls were also unchecked under `set -uo pipefail`: a failed stop left the old service running, the following start succeeded as a no-op, and the health check reported SUCCESS for a binding that was never loaded. orjson floor raised to 3.11.6 for CVE-2025-67221 (unbounded recursion in dumps); it covers the project's Python 3.10-3.13 range. Adds test/test_cache_nonfinite_floats.py (14) plus regression tests in test_scroll_config.py and test_scroll_helper.py. 9 of the cache tests and 9 of the scroll_config tests fail against the pre-fix code. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * test(harness): keep the visual double's signature tied to production Moves set_scrolling_state's frame_hold into the test double here, where DisplayManager gains it, rather than in #534 where it arrived a PR early. CodeRabbit flagged the #534 version correctly: a double that accepts an argument production does not lets the call pass every harness run and raise TypeError on the panel, which is the one failure a safety harness exists to prevent. The drift has now gone both ways across two branches -- double behind production on this branch, double ahead of it on #534 -- so it is pinned instead of remembered. test_display_double_parity.py compares the two signatures and fails with the direction of the drift named. It reads the files with ast rather than importing them, because display_manager imports rgbmatrix at module scope and this check should hold on a laptop and in CI as well as on a Pi. Plugins begin passing frame_hold in ledmatrix-plugins#462, which is why production and the double both need it before that lands. Full suite: 3889 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9837315308 |
perf(display): dirty tracking in update_display + plugin FPS declaration (#406)
* perf(display): dirty tracking in update_display + plugin FPS declaration update_display now skips SetImage+SwapOnVSync when the frame is byte-identical to the last pushed one (adler32 digest) AND brightness is unchanged — brightness is part of the digest, and set_brightness additionally resets it, so a dim-schedule change can never be skipped. clear() resets the digest (it writes to the matrix directly). Skipping a swap is hardware-safe: the panel refreshes the current frame from the driver's own thread; swaps only change content. Kill switch: display.dirty_tracking: false restores always-push. display_controller's high-FPS decision gains a precedence step: a plugin exposing needs_high_fps is honored first (so static-image can declare False for still PNGs and stop burning a 125fps loop on them); static-image without the attribute keeps its historical forced high-FPS (GIF back-compat); scrolling logic is otherwise unchanged. Verified with 7 tests against the real DisplayManager on RGBMatrixEmulator (identical-frame skip, pixel-change push, clear and brightness invalidation, snapshot-through-skip, kill switch) plus the 202-test display/controller/vegas suites, and a clean devpi deploy. Audit: every SetImage/SwapOnVSync/Clear/brightness call site is inside display_manager — no external writer can bypass the digest. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FqzC1nzTWL4kaqgMaQZFam * fix(display): serialize update_display, narrow brightness exception, log fixes CodeRabbit review on #406, verified against current code: - update_display() can genuinely be called from background threads (some sports base classes call it directly from inside update() for an immediate "live" refresh), not just the render loop — confirmed via the existing follower-mode gating wrapper in display_controller.py, which exists specifically because "background plugin threads" can reach it. Without a lock, two callers could both pass the digest check before either writes _last_pushed_digest back, causing a redundant push, or interleave the offscreen/current canvas swap. Added self._update_lock (RLock, in case of re-entrant callers) around the full method body so every call site is automatically covered — no caller changes needed. (No prior lock existed to reuse on DisplayManager; this adds one.) - Narrowed the brightness-read exception handler to AttributeError, matching the established pattern in get_brightness()/set_brightness() — a getattr() with a default already swallows AttributeError, so the only case this guards is the property getter itself raising, and the established pattern treats that as an expected, specific failure mode rather than something to blanket-catch. - FPS-check debug log now includes the plugin_id already in scope (previously only active_mode) and a "[DisplayController]" prefix for grep-ability, matching the sibling log two lines below it. - test_display_dirty_tracking.py: dm fixture and test_config_flag_wires_through now reset the DisplayManager singleton on teardown, matching the pattern test_display_manager.py already uses elsewhere in the same file family. - test_snapshot_still_written_on_skip previously only exercised the non-skip (push) path despite its name; now performs a second update that meets the skip conditions (identical frame) and asserts the snapshot is still written even though the panel push itself is skipped. All 7 dirty-tracking tests pass, plus the full display_manager/ display_controller/vegas suite (140 passed). Full repo suite has only the 5 known pre-existing failures (double-sided config x2, state_reconciliation x2, and test_circuit_breaker's conftest.py mock signature drift — the latter fixed in #400, which this branch's base predates). --------- Co-authored-by: Chuck <chuck@example.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |