Files
LEDMatrix/test
a51fb7ce11 feat(vegas): make scroll stutter visible, and catch it in the act (#454)
* feat(vegas): make scroll stutter visible, and catch it in the act

The loop reported only a mean FPS over a five-second window. At 120fps
that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee --
moves the average from 120.0 to 115.4 and reads as healthy. Stutter was
literally unmeasurable.

The FPS line now carries p99, the worst frame, and a hitch count. On the
dev rig that immediately turned "it sometimes stutters" into a number:
two freezes of 3.2s and 0.7s in twenty minutes, with every other frame
under 81ms.

Statistics say a stall happened but not what caused it, and by the time
they are logged the stack is gone. So there is also a watchdog that dumps
every thread's stack while the loop is still wedged. It is off unless
LEDMATRIX_STALL_WATCHDOG is set to a threshold in seconds, since it
prints a lot. Pointed at the 3.2s freeze it named the culprit on the
first try: a plugin generating a 17,000px scroll image, logo PNG decode
and all, synchronously on the render thread.

The hitch threshold is relative to what frames actually cost, not to the
configured target. The target is routinely set above what the panel can
hold so vsync does the pacing; measured against that budget every
ordinary frame counts as a hitch, and the first version of this counter
duly reported 250 per window on a display running perfectly smoothly.

The watchdog is owned by the coordinator, not created per iteration --
run_iteration is called repeatedly, so building one there would leak a
thread each time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): let the watchdog see stalls that hold the GIL

The watchdog only noticed a late heartbeat, which a whole class of
freeze can never produce: if the loop is inside one long C call that
holds the GIL, this thread cannot run during the stall, and by the time
it does the loop has already checked in. On the dev rig that hid a
recurring 3.2s freeze completely -- twenty minutes of watching produced
one dump, for an unrelated 0.4s stall.

What it can still observe is that its own sleep ran long. A badly
overshot wait is now reported as a stall in its own right. The stacks
are stale by then and the message says so, but knowing the freeze is
GIL-holding is most of the diagnosis: it rules out lock contention and
scheduling, and points at a single long C call.

This also explains why lowering sys.setswitchinterval changed nothing --
the switch interval cannot preempt a C call that never releases the GIL.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* feat(vegas): report the worst frame, not just the mean

The loop logged only a mean FPS over a five-second window. At 120fps
that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee
-- moves the average from 120.0 to 115.4 and reads as perfectly healthy.
Stutter was unmeasurable, which is why "it sometimes freezes" went
unpinned for so long.

Adding p99 and the worst frame turned that into a number immediately: on
the dev rig, two freezes of 3.2s and 0.7s in twenty minutes with every
other frame under 81ms. Not general slowness -- two rare, total stalls,
which is a different problem with a different fix.

Costs 0.96us per frame, about 0.012% of an 8.3ms frame.

This replaces an earlier version that also shipped a stall watchdog and
a hitch counter. The watchdog never found anything -- one dump in
forty-five minutes, for an unrelated stall -- because it can only notice
a late heartbeat, and the freeze happens in coordinator.start() before
the frame loop begins beating. py-spy found the cause in one recording
by sampling the process externally, which needs no code here. The hitch
counter went with it: it needed a rolling median every frame, which was
most of the cost, to produce a number the worst frame already tells you.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

* fix(vegas): use the nearest-rank index for p99

int(n * 0.99) is off by one, and at exactly 100 samples it selects the
maximum -- which is the number logged immediately beside it as the worst
frame. The two columns exist to say different things, p99 the
bad-but-ordinary frame and worst the outlier, so they agreed precisely
when the sample was smallest and least informative.

Nearest rank is ceil(n * fraction) - 1. Extracted so it can be tested
directly rather than only through a five-second logging interval.

Reported by CodeRabbit on the PR.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 14:06:32 -04:00
..