250ms catches freezes; the hitches left on hdpi are frames 2-5 refreshes
late, which look like the render thread waiting for the GIL. At 30ms the
watchdog dumps those too, naming what the other threads were running when
the frame missed. It polls at a third of the threshold so a stall one poll
long is still seen, which costs some GIL time of its own: a diagnostic
setting, not one to soak with.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
render_bench.py (from the parallel perf/render-bench work) had its own
grading module, frame_pacing, with its own definition of a missed frame
and its own refresh estimate. The soak already had both in frame_timing,
so the two could have drifted apart on what "late" means.
The bench now gives the display manager a fresh FrameTimingRecorder,
drains it synchronously at the start and end of the graded run, and prints
frame_soak's report with frame_soak's verdict. Its workload is unchanged:
the synthetic strip, --busy load, the shared speed resolver, the
per-frame scrolling announcement. frame_pacing, its tests and its
src.common exports are removed; measure_refresh_hz moves to frame_timing,
where scroll_speeds.py now finds it.
Two ideas from frame_pacing carry over. The bench seeds the recorder with
the idle refresh it measures, so a loop that free-runs (the 827fps bug
the first bench caught) shows as early frames and one stuck at half rate
as late frames, where an estimate taken from their own intervals finds
both self-consistent. And the soak, which has no idle measurement, now
calls a run NOT LOCKED when its refresh estimate beats the configured cap.
The report also gives the rate held while rendering.
Docs: the bench becomes "Without the service" under "Soaking a rig",
keeping its hdpi numbers and the idle-vs-rendering refresh finding.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The recorder counts freezes; it cannot say why. hdpi showed 1-2s freezes
in both the #628 and offscreen builds, one lining up with hockey's 2s
NHL fetch on the update thread, and nothing in the logs explained it.
StallWatchdog polls every 50ms from its own thread. When a scroll's last
frame is more than 250ms old (and a scroll is still running, so the end
of a scroll is not a stall), it logs the stack of the thread that
presented that frame and the top of every other thread's, then the
stall's length when frames resume. It also measures how late its own
wake-up was: if it was held up as long as the render thread, the whole
interpreter was blocked (C code holding the GIL), not one thread on a
lock. One dump per 30s at most; LEDMATRIX_STALL_WATCHDOG=0 disables it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three gaps found by the first hdpi soaks:
- Intervals of 1s or more between two scrolling frames were dropped as
"gaps between scrolls". But the scrolling state lapses only after 2s, so
every 1-2s stall inside a scroll vanished from the report. Those are now
freezes (the gap bound is a 5s sanity limit), with a breakdown by length.
- A frame a whole refresh early means the swap did not wait for the panel.
Those are counted, and a soak with more than the threshold of them fails
as NOT LOCKED instead of reporting a flattering late rate.
- The refresh estimate took the lowest window it had seen, so one window of
non-blocking swaps halved it and made every early frame look on time. A
window may now lower it by at most 20%.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each scroller already logs its own stats line, but in different formats,
per source, and Vegas logs a healthy window only at DEBUG. None of it
answers the question a release has to answer on each rig: over a long
run, how often did a moving frame reach the panel late?
Every frame reaches the panel through DisplayManager.update_display, so
it is timed there once, whoever drew it: the blit (SetImage), the vsync
wait, and the interval since the previous frame. The render thread only
appends a tuple. A worker thread aggregates cumulative counters and
histograms and rewrites /dev/shm/ledmatrix_frame_stats.json every 10s
(RAM, so no SD wear).
A frame due after `hold` refreshes that lands one or more refreshes
later is "late": the panel repeated the previous frame, a visible hitch.
Gaps of 250ms+ inside a scroll are "freezes" (recomposes, handovers,
blocking calls), counted separately so one handover does not read as 40
missed refreshes. Static frames, the first frame of a scroll and gaps
between scrolls are not timed. The refresh period is estimated from the
frames themselves.
scripts/frame_soak.py runs next to the service as any user, diffs two
snapshots over a run (default 10 minutes), optionally keeps the web
preview's viewer marker fresh, and exits non-zero above 0.1% late
frames. It also reports whether the loaded rgbmatrix binding releases
the GIL. Documented under "Soaking a rig" in docs/SCROLL_PERFORMANCE.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>