mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-10 09:06:36 +00:00
feat(perf): time every presented frame, and a soak script to judge a rig
Each scroller already logs its own stats line, but in different formats, per source, and Vegas logs a healthy window only at DEBUG. None of it answers the question a release has to answer on each rig: over a long run, how often did a moving frame reach the panel late? Every frame reaches the panel through DisplayManager.update_display, so it is timed there once, whoever drew it: the blit (SetImage), the vsync wait, and the interval since the previous frame. The render thread only appends a tuple. A worker thread aggregates cumulative counters and histograms and rewrites /dev/shm/ledmatrix_frame_stats.json every 10s (RAM, so no SD wear). A frame due after `hold` refreshes that lands one or more refreshes later is "late": the panel repeated the previous frame, a visible hitch. Gaps of 250ms+ inside a scroll are "freezes" (recomposes, handovers, blocking calls), counted separately so one handover does not read as 40 missed refreshes. Static frames, the first frame of a scroll and gaps between scrolls are not timed. The refresh period is estimated from the frames themselves. scripts/frame_soak.py runs next to the service as any user, diffs two snapshots over a run (default 10 minutes), optionally keeps the web preview's viewer marker fresh, and exits non-zero above 0.1% late frames. It also reports whether the loaded rgbmatrix binding releases the GIL. Documented under "Soaking a rig" in docs/SCROLL_PERFORMANCE.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -231,6 +231,9 @@ advances by elapsed time at `scroll_speed / scroll_delay` px/s.
|
||||
|
||||
## Diagnosing a juddery scroller
|
||||
|
||||
To check a whole rig rather than one scroller, soak it -- see *Soaking a rig*
|
||||
below.
|
||||
|
||||
**An average will lie to you.** A 2 ms duplicate frame and a 21 ms double-wait
|
||||
mean exactly 10 ms, so a ticker stalling on half its frames still averages to a
|
||||
healthy 100 fps. The stats line reports the tail for that reason — read the
|
||||
@@ -303,6 +306,49 @@ journalctl -u ledmatrix --since "-5min" --no-pager | grep -iE "px/s|px/frame"
|
||||
If a plugin logs its scroll config **twice** with different modes, the second
|
||||
line is what is running.
|
||||
|
||||
## Soaking a rig
|
||||
|
||||
The per-scroller lines above tell you *which* scroller misbehaves. The soak
|
||||
answers the question a release has to answer for each rig: **over a long run,
|
||||
how often did a moving frame reach the panel late?**
|
||||
|
||||
Every frame reaches the panel through `DisplayManager.update_display`, so it is
|
||||
timed there once, whoever drew it -- Vegas, a ticker plugin, anything. The
|
||||
render thread only appends a tuple; a worker thread aggregates and rewrites
|
||||
`/dev/shm/ledmatrix_frame_stats.json` every 10 seconds (RAM, so no SD-card
|
||||
wear). `src/common/frame_timing.py` has the details.
|
||||
|
||||
```bash
|
||||
python3 scripts/frame_soak.py # 10 minutes, as the display is now
|
||||
python3 scripts/frame_soak.py --preview # with the web preview open
|
||||
python3 scripts/frame_soak.py --show # totals since the service started
|
||||
python3 scripts/frame_soak.py --json a.json # keep the report to compare later
|
||||
```
|
||||
|
||||
It runs as any user next to the display service and stops nothing. It needs
|
||||
something to *scroll* during the run: a live game holding a static scoreboard
|
||||
on screen gives no verdict. `--preview` keeps the web preview's viewer marker
|
||||
fresh, which puts the preview's PNG encoding at full rate -- run it as the web
|
||||
service's user.
|
||||
|
||||
| line | what it tells you |
|
||||
|---|---|
|
||||
| **Late frames** | Frames presented one or more refreshes after they were due: the panel showed the previous frame again, a visible hitch. **The pass/fail number**, 0.1% by default (`--max-late-pct`). Only intervals between two scrolling frames count, and a frame held for `frame_hold` refreshes is due `frame_hold` refreshes after the last. |
|
||||
| **Freezes** | Gaps of 250 ms or more inside a scroll: recomposes, plugin handovers, blocking calls on the render thread. Reported but not failed on, because some are handovers between plugins rather than faults. |
|
||||
| **blit** | Copying the frame into the matrix canvas (`SetImage`). It grows with width × height × `pwm_bits`: ~5.5 ms at 512×64 with 8 bits on a Pi 4. It is the biggest fixed cost, and it sets the refresh rates a rig can hold one pixel per refresh at. |
|
||||
| **wait** | Time blocked in `SwapOnVSync`, i.e. the slack left in each refresh. A p50 near zero means the rig has no headroom and anything extra lands a frame late. |
|
||||
| **work** | Everything else between two frames: drawing, scrolling, and waiting for the GIL. A wide gap between its p50 and p99 is another thread getting in the way. |
|
||||
| **Binding** | `STOCK` means the rgbmatrix binding holds the GIL through the vsync wait, which starves every other thread. See *Rebuilding the binding*. |
|
||||
|
||||
The refresh rate is estimated from the frames themselves (swaps that block on
|
||||
vsync can only land on refresh boundaries). Cross-check it with
|
||||
`scroll_speeds.py --measure` if it looks wrong. It can read high on a rig where
|
||||
nothing ever presented at the full refresh rate.
|
||||
|
||||
A soak is only meaningful against a fixed workload. Compare runs with the same
|
||||
content and `--preview` setting, and alternate which build goes first when you
|
||||
A/B two of them. A live-API workload drifts over time.
|
||||
|
||||
## Rebuilding the binding
|
||||
|
||||
```bash
|
||||
|
||||
Reference in New Issue
Block a user