mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-08-12 22:28:06 +00:00
* feat(vegas): make scroll stutter visible, and catch it in the act The loop reported only a mean FPS over a five-second window. At 120fps that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee -- moves the average from 120.0 to 115.4 and reads as healthy. Stutter was literally unmeasurable. The FPS line now carries p99, the worst frame, and a hitch count. On the dev rig that immediately turned "it sometimes stutters" into a number: two freezes of 3.2s and 0.7s in twenty minutes, with every other frame under 81ms. Statistics say a stall happened but not what caused it, and by the time they are logged the stack is gone. So there is also a watchdog that dumps every thread's stack while the loop is still wedged. It is off unless LEDMATRIX_STALL_WATCHDOG is set to a threshold in seconds, since it prints a lot. Pointed at the 3.2s freeze it named the culprit on the first try: a plugin generating a 17,000px scroll image, logo PNG decode and all, synchronously on the render thread. The hitch threshold is relative to what frames actually cost, not to the configured target. The target is routinely set above what the panel can hold so vsync does the pacing; measured against that budget every ordinary frame counts as a hitch, and the first version of this counter duly reported 250 per window on a display running perfectly smoothly. The watchdog is owned by the coordinator, not created per iteration -- run_iteration is called repeatedly, so building one there would leak a thread each time. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(vegas): let the watchdog see stalls that hold the GIL The watchdog only noticed a late heartbeat, which a whole class of freeze can never produce: if the loop is inside one long C call that holds the GIL, this thread cannot run during the stall, and by the time it does the loop has already checked in. On the dev rig that hid a recurring 3.2s freeze completely -- twenty minutes of watching produced one dump, for an unrelated 0.4s stall. What it can still observe is that its own sleep ran long. A badly overshot wait is now reported as a stall in its own right. The stacks are stale by then and the message says so, but knowing the freeze is GIL-holding is most of the diagnosis: it rules out lock contention and scheduling, and points at a single long C call. This also explains why lowering sys.setswitchinterval changed nothing -- the switch interval cannot preempt a C call that never releases the GIL. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * feat(vegas): report the worst frame, not just the mean The loop logged only a mean FPS over a five-second window. At 120fps that is ~600 frames, so a 200ms freeze -- plainly visible on a marquee -- moves the average from 120.0 to 115.4 and reads as perfectly healthy. Stutter was unmeasurable, which is why "it sometimes freezes" went unpinned for so long. Adding p99 and the worst frame turned that into a number immediately: on the dev rig, two freezes of 3.2s and 0.7s in twenty minutes with every other frame under 81ms. Not general slowness -- two rare, total stalls, which is a different problem with a different fix. Costs 0.96us per frame, about 0.012% of an 8.3ms frame. This replaces an earlier version that also shipped a stall watchdog and a hitch counter. The watchdog never found anything -- one dump in forty-five minutes, for an unrelated stall -- because it can only notice a late heartbeat, and the freeze happens in coordinator.start() before the frame loop begins beating. py-spy found the cause in one recording by sampling the process externally, which needs no code here. The hitch counter went with it: it needed a rolling median every frame, which was most of the cost, to produce a number the worst frame already tells you. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 * fix(vegas): use the nearest-rank index for p99 int(n * 0.99) is off by one, and at exactly 100 samples it selects the maximum -- which is the number logged immediately beside it as the worst frame. The two columns exist to say different things, p99 the bad-but-ordinary frame and worst the outlier, so they agreed precisely when the sample was smallest and least informative. Nearest rank is ceil(n * fraction) - 1. Extracted so it can be tested directly rather than only through a five-second logging interval. Reported by CodeRabbit on the PR. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Udr6MfaFLUPhX5Fgo67Jf5 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>