mirror of
https://github.com/ChuckBuilds/LEDMatrix.git
synced 2026-10-04 14:25:08 +00:00
* fix(core): register tom_thumb, accept frame_hold in the test double, wire api_v3's managers Three independent fixes found while validating every plugin on a 256x64 rig. FontManager never registered tom_thumb even though assets/fonts/tom-thumb.bdf ships with the core, so every plugin offering it logged "Font family 'tom_thumb' not found" (16 warnings per countdown render) and had to carry a private loader to use a bundled font. Closes #524. VisualTestDisplayManager.set_scrolling_state() lacked the frame_hold parameter that DisplayManager gained, so any plugin passing it died with TypeError at render time and failed every size. Nine plugins now make that call; ledmatrix-stocks and ledmatrix-leaderboard were failing outright and the other seven only passed because their scroll path was unreachable without data. Closes #525. api_v3 declared module-level config_manager/plugin_manager = None that nothing ever assigned -- app.py sets the blueprint attributes, which the other 150+ call sites use. Three sites read the decoys, so /health reported the config unreadable and the plugin system uninitialised (making "degraded" permanent and unreachable-by-design) and /display/current fell back to a hardcoded 128x64 on every rig. The decoys are removed rather than assigned, so a bare name is now a NameError at test time instead of a silent None. The same function's first-call uptime was computed from two separate clock reads and came out negative. Closes #529. Verified on the rig: both previously-failing plugins render, the tom_thumb warnings are gone, /health reports "healthy" with all three checks passing, and /display/current reports the real 256x64. * fix(core): unique snapshot temp name, honour on-demand requests, skip empty starlark The preview snapshot wrote through a fixed "<snapshot>.tmp". /tmp is world-writable and sticky, and the display service runs as a different user from the tooling, so a leftover temp owned by anyone else became unopenable even by root -- fs.protected_regular refuses O_CREAT on a foreign file in a sticky directory. The preview and the health check's liveness proxy then froze until someone deleted the file by hand; on the test rig that meant 23 hours of a healthy display reporting "hardware: stale". Now uses tempfile.mkstemp with cleanup on failure, matching the hardware-status write a few hundred lines above. Closes #528. _poll_on_demand_requests read its mailbox with max_age=3600, and get() defaults the in-memory TTL to max_age -- so the first request was pinned in memory for an hour and every later poll returned that stale copy. No second on-demand request was honoured until the service restarted, while the API kept returning 200. get() already documents memory_ttl=0 for exactly this cross-process case. The consumed request is also now deleted: leaving it on disk meant a restart replayed the previous request, activated it, and ignored the one the caller had just made. Closes #530. starlark-apps returned None from display() when it has no app to show, which is the state of every install without Pixlet and of a fresh one before any app is added. The controller only skips on a boolean False, so that held a black panel for the full display_duration instead of rotating on. Closes #456 (core side). Verified on the rig: two consecutive on-demand requests with no restart between them are both activated, where the second was previously dropped in silence. * perf(harness): share one cache across a plugin's renders _instantiate built a fresh MockCacheManager for every (size, mode), and that mock is a per-instance in-memory dict, so each render was a cold start. A plugin that fetches per game or per player re-fetched everything N times over -- baseball-scoreboard at one size took 840s for nine renders where the arithmetic said ~72s, and at eight sizes it exceeded a 900s timeout. The second and later renders also never exercised the cache-hit path, which is what a running rig executes almost all of the time, so a caching regression could not be caught here. The cache is now built once per render_plugin_matrix call and threaded down. The display manager stays per-render -- the bounds checking depends on that -- so only fetched data is shared. Measured on the rig, same render counts and same goldens: tide-display 2s -> 1s (32 renders) cricket-scoreboard 10s -> 3s (24 renders) No pass/fail change across tide-display, cricket-scoreboard, clock-simple, geochron, christmas-countdown, of-the-day, web-ui-info and incoming-packages. Closes #533. * fix(scripts): run standalone plugin tests instead of collecting nothing run_plugin_tests.py discovered every plugin test file and handed the lot to pytest. Most plugin tests are standalone scripts -- module-level main() plus an `if __name__ == "__main__"` guard, signalling through an exit code -- and pytest collects zero items from those. The run printed how many files it had *found*, then "no tests ran", and exited without executing any of them. On a rig with all 44 first-party plugins that is 151 of 248 files. Files are now classified and each kind runs under the right runner: pytest for real test modules, subprocess for scripts, honouring the 0 pass / 2 skip / 1 fail convention ledmatrix-plugins' own runner established (a script that wants a tty or an LED matrix is a skip, not a regression). Before: $ python3 scripts/run_plugin_tests.py -p countdown -d ~/LEDMatrix/plugin-repos Found 1 test file(s) collected 0 items no tests ran in 0.31s rc=0 After: Found 1 test file(s) -- 0 collectable, 1 standalone script(s) 1 passed, 0 skipped, 0 failed (scripts) rc=0 Verified across three shapes: countdown (1 script), jellyfin-now-playing and pomodoro-timer (pytest only, 16 and 42 tests), and ledmatrix-flights (11 files split 4 collectable / 7 scripts, all seven of which had never run). Closes #532. Running the flights scripts for the first time also surfaced four genuinely failing tests there, hidden by the mirror-image bug in the plugins repo's own runner -- filed as ChuckBuilds/ledmatrix-plugins#464 and #465. * fix(harness): give an empty-looking mode a few frames before warning about it check_plugin's "drew nothing but display() returned X" warning fired on a single frame, rendered with force_clear=True, under a frozen clock. All three defeat a scrolling plugin, whose first frame is legitimately its blank scroll-in buffer. Across 44 first-party plugins, 60 of 76 warnings were false -- the rate at which people stop reading a warning, which matters because the true positives are real: a mode that draws nothing and does not return False holds a blank panel for its whole display duration. An apparently-empty frame is now re-driven for up to 48 more frames with force_clear=False (force_clear means "reset the scroll", so repeating it would redraw frame 1 for ever) and with the clock advancing -- freezegun's factory where time is frozen, a real sleep where it is not, since scroll position is usually a function of elapsed time. The first frame that draws content replaces the result. The clock is moved back afterwards. It is shared by every render in the matrix, so time borrowed by the probe leaked into later modes and drifted their goldens -- f1_upcoming picked up 5 spurious drifts before this was restored. Measured on the rig: empty warns check before after f1-scoreboard 42 0 48 PASS / 0 FAIL, goldens intact ledmatrix-elections 16 0 16 PASS / 0 FAIL on-air 8 8 true positive, kept nfl-draft 8 8 true positive, kept clock-simple/geochron/ 0 0 unchanged christmas-countdown 58 false positives gone, both true positives kept, no golden regressions. Cost is confined to modes that really are blank: plugins that draw immediately are unchanged (clock-simple and tide-display still 2s), while on-air -- eight deliberately blank modes -- goes to 21s. Closes #527. * fix(harness): load nested schema defaults, and merge caller config at leaf level load_config_defaults read only top-level properties. An object property carries its defaults on its children, not on itself, so everything nested was dropped -- 2,386 defaults across 37 of 44 plugins, soccer-scoreboard alone losing 539 of 565. render_plugin_matrix's comment says the plugin then "behaves like a real install", which for most of the fleet it did not. _defaults_from_properties now recurses. merge_config deep-merges the caller's config onto the result so an override lands at the leaf: a shallow merge would let -c '{"nhl": {"enabled": true}}' replace the whole nhl subtree and discard every other nhl default, which is the same class of bug being fixed here. Measured before/after across all 49 installed plugins on the rig: **no render changed** -- identical PASS/FAIL counts, byte-identical output, goldens intact. Plugins already fall back to the same values internally via config.get(key, default), so supplying them explicitly agrees with what they were doing. The defaults really are arriving now: ufc-scoreboard 9 -> 87 defaults ledmatrix-flights 51 -> 95 masters-tournament 10 -> 51 cricket-scoreboard 22 -> 50 tide-display 12 -> 18 and hockey-scoreboard, which used to load nhl.enabled=None, now gets nhl.enabled=True with its full display_modes block. Caveat worth carrying: the eight plugins with the most nested config (soccer, baseball, basketball, hockey, lacrosse, football, afl, nrl -- 1,634 of the 2,386 dropped defaults, 68%) could not be measured. They import src.common.sports_shared, which the test rig's core branch predates, so they fail to load there identically before and after. Re-run this comparison against a core that has that module before trusting the "nothing changed" result for them; those are exactly the plugins whose renders should change most. Closes #531. * refactor: narrow the exception handlers this branch introduced Codacy flagged the new code; it passes on other recent PRs, so the finding is mine. Four of the five broad `except Exception` clauses I added were catching far more than they needed to, which is the same shape as several bugs this branch fixes -- hello-world's TypeError sat invisible for exactly this reason. freezer() / move_to() / tick() -> (AttributeError, TypeError, ValueError) cache_manager.delete() -> (OSError, AttributeError, KeyError) The fifth stays broad and now says why: it wraps a call into a plugin's own display(), which can raise anything, and the first frame has already rendered -- so a failure there must not turn a good result into an error. Verified against a checkout of main: f1-scoreboard 48 PASS / 0 FAIL with 0 empty warnings, on-air keeps its 8 true positives, clock-simple 8 PASS. geochron shows 7 golden drifts both before and after this branch, so it is not from these changes -- its committed goldens predate #521's 1-bit text rendering. * fix: resolve CodeRabbit review and Codacy findings on #534 CodeRabbit raised six; all six were real. The test double had drifted ahead of production. VisualTestDisplayManager accepted set_scrolling_state(frame_hold=...) while DisplayManager did not, so such a call passed every harness run and would raise TypeError on the panel -- the one failure a safety harness exists to prevent. frame_hold belongs to the change that adds it to DisplayManager (#523), so it moves there and the double matches main again. The harness swallowed exceptions from re-rendered frames. _settle_loop re-renders a mode that came back blank, to give a scroll time to draw; returning silently on a crash meant a mode that renders one good frame and then explodes was reported as passing. Recorded on result.error now, keeping the captured frame so the failure stays inspectable. starlark-apps display() returned True after _display_frame() failed, so the controller held a dead frame for the whole display_duration instead of rotating on. _display_frame now returns bool on all three paths. run_plugin_tests.py used env.setdefault for PYTHONPATH and LEDMATRIX_CORE, so an inherited value won and the subprocess imported a different core than the one under test -- ledmatrix-plugins#467 exactly. Prepends PROJECT_ROOT and sets LEDMATRIX_CORE unconditionally. The on-demand mailbox is polled after every frame, ~125x/second on a scrolling mode, and the read is deliberately uncached, so it was that many disk reads per second to find nothing. Floored at 250ms, which is imperceptible for a web-UI click. Consuming it also deleted whatever was present rather than what had just been processed, so a request posted while the previous one was in flight was thrown away and never ran; the delete is now keyed by request_id. That narrows the window rather than closing it -- a true atomic claim needs a primitive the cache layer does not offer, and the code says so rather than implying otherwise. Codacy's 2 criticals were bandit B404/B603 on the subprocess call added to run_plugin_tests.py. Fixed interpreter, argument list, no shell; annotated with the repo's existing nosec convention. Bandit is clean on the file. Adds test/test_on_demand_mailbox.py (8), test_starlark_display_contract.py (4) and two settle cases in test_harness_empty_claimed.py. 4, 4 and 2 of those fail against the pre-fix code. Full suite: 3961 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * chore: satisfy Codacy's subprocess checks on the new test runner Codacy runs Bandit and Opengrep (its Semgrep fork). The new subprocess.run in scripts/run_plugin_tests.py trips three patterns, on two different lines: Bandit B404 on the import, B603 on the call Opengrep dangerous-subprocess-use-audit on the run( line dangerous-subprocess-use-tainted-env-args on the argv line A nosemgrep applies only to its own line, so the call line and the argv line each need one; a single comment on the call covered neither rule fully. Suppression is the right answer here rather than a rewrite: the interpreter is sys.executable, the arguments are a list, and no shell is involved, so there is nothing to word-split or expand. Matches the pair the rest of the repo already uses for this shape -- permission_utils.py, plugin_loader.py, install_dependencies_apt.py. Codacy: 0 new issues, up to standards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * chore: leave visual_display_manager untouched so #523 can merge The only change this branch made to that file was a docstring, and it collided with #523's rewrite of the same method -- so #534 and #523 each merged cleanly against main but conflicted with each other. Reverted to main's text; #523 owns this method and adds frame_hold to it. The note the docstring carried ('frame_hold arrives in #523') would have been stale the moment #523 landed anyway. The parity test in #523 is what actually keeps the two signatures honest. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 * chore: add the Ruff suppression nosec/nosemgrep do not cover Ruff reports S603 on the same call Bandit and Opengrep do, and none of the three suppressions covers the others. Confirmed the precondition first: path comes from discover_plugin_tests(), which globs test files inside the repo, and the call is a fixed interpreter with a list argv and no shell. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
566 lines
24 KiB
Python
566 lines
24 KiB
Python
"""
|
|
Plugin safety harness.
|
|
|
|
Renders a plugin across every declared screen (mode) and every supported matrix
|
|
size, capturing crashes and overflow. Used by scripts/check_plugin.py and the
|
|
pytest matrix test to guarantee a plugin change doesn't break a screen at a size
|
|
the author didn't try.
|
|
|
|
The render flow mirrors scripts/render_plugin.py (same PluginLoader call), but
|
|
this module adds: multi-size iteration, per-mode rendering, overflow detection
|
|
via BoundsCheckingDisplayManager, and golden-image comparison.
|
|
"""
|
|
|
|
import contextlib
|
|
import http.client
|
|
import inspect
|
|
import time
|
|
from datetime import timedelta
|
|
import socket
|
|
import ssl
|
|
import urllib.error
|
|
from dataclasses import dataclass
|
|
from pathlib import Path
|
|
from typing import Any, Dict, List, Optional, Tuple
|
|
|
|
from PIL import Image, ImageChops
|
|
|
|
from src.logging_config import get_logger
|
|
from .bounds_display_manager import BoundsCheckingDisplayManager
|
|
from .loading import load_config_defaults, load_manifest, merge_config
|
|
from .sizes import DEFAULT_TEST_SIZES, safe_mode_filename, size_label
|
|
|
|
logger = get_logger("[Plugin Harness]")
|
|
|
|
|
|
def _tolerated_update_errors() -> Tuple[type, ...]:
|
|
"""Exception types from update() we treat as a tolerated no-connectivity
|
|
failure (expected in CI / headless dev) rather than a real plugin bug.
|
|
|
|
Anything NOT in this set is a genuine regression — a plugin that lets a
|
|
non-network exception escape update() should fail the harness, not pass
|
|
green because display() happened to survive.
|
|
"""
|
|
types: List[type] = [
|
|
ConnectionError, TimeoutError, # builtins
|
|
socket.gaierror, socket.timeout, # DNS / socket timeouts
|
|
ssl.SSLError,
|
|
urllib.error.URLError,
|
|
http.client.HTTPException,
|
|
]
|
|
try: # requests is optional; cover its whole error tree when present
|
|
import requests
|
|
types.append(requests.exceptions.RequestException)
|
|
except ImportError: # pragma: no cover - requests not installed
|
|
logger.debug("requests not installed; its connectivity errors won't be specifically tolerated")
|
|
return tuple(types)
|
|
|
|
|
|
_TOLERATED_UPDATE_ERRORS = _tolerated_update_errors()
|
|
|
|
|
|
@dataclass
|
|
class RenderResult:
|
|
"""Outcome of rendering one (size, mode) of a plugin."""
|
|
plugin_id: str
|
|
width: int
|
|
height: int
|
|
mode: str
|
|
image: Optional[Image.Image] = None
|
|
error: Optional[str] = None # fatal: load/display crash, or a non-network update() error
|
|
update_error: Optional[str] = None # tolerated: connectivity error from update() (no network in CI)
|
|
overflow: Optional[Tuple[int, int, int, int]] = None # bbox past the panel
|
|
# golden comparison (populated only when a golden was provided)
|
|
golden_checked: bool = False
|
|
golden_ok: Optional[bool] = None
|
|
golden_diff_pixels: int = 0
|
|
golden_max_delta: int = 0
|
|
# what display() handed back; the controller skips a mode only on False
|
|
display_returned: Any = None
|
|
# empty-frame check: rendered nothing while not reporting "no content"
|
|
empty_claimed: Optional[bool] = None # True when that happened
|
|
empty_ok: Optional[bool] = None # False only in strict mode
|
|
# fill / scale-up check (populated only for sizes >= 2x the design size)
|
|
fill_checked: bool = False
|
|
fill_ok: Optional[bool] = None # False only in strict mode
|
|
fill_extent: Optional[Tuple[float, float]] = None # (extent_x, extent_y)
|
|
|
|
@property
|
|
def size_label(self) -> str:
|
|
return size_label(self.width, self.height)
|
|
|
|
@property
|
|
def ok(self) -> bool:
|
|
"""Phase-1 pass: rendered without crashing and without overflow, and if a
|
|
golden was checked it matched."""
|
|
if self.error is not None or self.overflow is not None:
|
|
return False
|
|
if self.golden_checked and self.golden_ok is False:
|
|
return False
|
|
if self.fill_ok is False:
|
|
return False
|
|
if self.empty_ok is False:
|
|
return False
|
|
return True
|
|
|
|
|
|
def list_modes(plugin_instance: Any, manifest: Dict[str, Any], plugin_id: str) -> List[str]:
|
|
"""Enumerate a plugin's screens: instance.modes wins, then manifest
|
|
display_modes, then the plugin id as a single mode."""
|
|
modes = getattr(plugin_instance, "modes", None)
|
|
if modes:
|
|
return [str(m) for m in modes]
|
|
declared = manifest.get("display_modes")
|
|
if declared:
|
|
return [str(m) for m in declared]
|
|
return [plugin_id]
|
|
|
|
|
|
def _instantiate(plugin_id: str, manifest: Dict[str, Any], plugin_dir: Path,
|
|
config: Dict[str, Any], mock_data: Dict[str, Any],
|
|
display_manager: Any, cache_manager: Any = None) -> Any:
|
|
"""Load and construct a plugin instance with mocked managers.
|
|
|
|
Pass ``cache_manager`` to share one cache across the renders of a plugin.
|
|
Building a fresh one per (size, mode) made every render a cold start, so a
|
|
plugin that fetches per game or per player re-fetched everything N times --
|
|
baseball-scoreboard took 840s for nine renders where ~72s was the arithmetic
|
|
-- and the cache-hit path, which is what a running rig executes almost
|
|
always, was never exercised.
|
|
"""
|
|
from src.plugin_system.plugin_loader import PluginLoader
|
|
from src.plugin_system.testing import MockCacheManager, MockPluginManager
|
|
|
|
if cache_manager is None:
|
|
cache_manager = MockCacheManager()
|
|
for key, value in (mock_data or {}).items():
|
|
cache_manager.set(key, value)
|
|
|
|
loader = PluginLoader()
|
|
plugin_instance, _module = loader.load_plugin(
|
|
plugin_id=plugin_id,
|
|
manifest=manifest,
|
|
plugin_dir=plugin_dir,
|
|
config=config,
|
|
display_manager=display_manager,
|
|
cache_manager=cache_manager,
|
|
plugin_manager=MockPluginManager(),
|
|
install_deps=False,
|
|
)
|
|
return plugin_instance
|
|
|
|
|
|
def _render_mode(plugin_instance: Any, mode: str) -> Any:
|
|
"""Render a specific screen. Prefer an explicit display_mode kwarg; otherwise
|
|
drive the plugin's internal mode state machine (first display() call renders
|
|
modes[current_mode_index] when current_display_mode is None).
|
|
|
|
Returns whatever display() returned. The display controller skips a mode
|
|
whose display() returns False, so that value decides whether an empty mode
|
|
is rotated past or sat on -- which makes it worth reporting rather than
|
|
discarding."""
|
|
sig = inspect.signature(plugin_instance.display)
|
|
if "display_mode" in sig.parameters:
|
|
return plugin_instance.display(force_clear=True, display_mode=mode)
|
|
|
|
modes = getattr(plugin_instance, "modes", None)
|
|
if modes and mode in modes:
|
|
plugin_instance.current_mode_index = list(modes).index(mode)
|
|
if hasattr(plugin_instance, "current_display_mode"):
|
|
plugin_instance.current_display_mode = None
|
|
return plugin_instance.display(force_clear=False)
|
|
|
|
|
|
# How many extra frames to drive before believing a mode really draws nothing.
|
|
# A scroll starts with its content off-panel, so frame 1 is legitimately blank;
|
|
# measured across the fleet, content appears by frame 2-4 (f1-scoreboard),
|
|
# frame 4 (ledmatrix-elections) and frame 38 at 64px (ledmatrix-leaderboard).
|
|
EMPTY_RECHECK_FRAMES = 48
|
|
# Seconds to advance the clock between those frames. Scroll position is usually
|
|
# driven by elapsed time, which a frozen clock never provides.
|
|
EMPTY_RECHECK_STEP = 0.05
|
|
|
|
|
|
def _has_content(image) -> bool:
|
|
"""True when any pixel is lit above the threshold."""
|
|
if image is None:
|
|
return False
|
|
return image.convert("L").point(
|
|
lambda p: 255 if p > _LIT_THRESHOLD else 0).getbbox() is not None
|
|
|
|
|
|
def _render_mode_again(plugin_instance: Any, mode: str) -> Any:
|
|
"""Draw one more frame WITHOUT force_clear.
|
|
|
|
_render_mode passes force_clear=True, which for a scrolling plugin means
|
|
"reset the scroll to the start" -- so repeating it would redraw frame 1 for
|
|
ever. The re-check needs the plugin to advance.
|
|
"""
|
|
sig = inspect.signature(plugin_instance.display)
|
|
if "display_mode" in sig.parameters:
|
|
return plugin_instance.display(force_clear=False, display_mode=mode)
|
|
return plugin_instance.display(force_clear=False)
|
|
|
|
|
|
def _settle_empty_frame(inst, mode, dm, result, freezer) -> None:
|
|
"""Give an apparently-empty mode a few frames to draw before believing it.
|
|
|
|
One frame is not evidence: a scroll's first frame is its blank scroll-in
|
|
buffer. Without this, every scrolling plugin was warned about -- 60 of 76
|
|
warnings on a 44-plugin rig were false, which is the rate at which people
|
|
stop reading a warning.
|
|
"""
|
|
if result.error is not None or result.display_returned is False:
|
|
return
|
|
if _has_content(result.image):
|
|
return
|
|
# The frozen clock is shared by every render in the matrix, so any time this
|
|
# probe borrows has to be given back -- otherwise a mode that scrolls in
|
|
# leaves the clock advanced and every later mode renders at the wrong
|
|
# instant, drifting its golden. Seen as 5 spurious f1_upcoming drifts.
|
|
resume_at = None
|
|
if freezer is not None:
|
|
try:
|
|
resume_at = freezer()
|
|
except (AttributeError, TypeError, ValueError):
|
|
# Not a freezegun factory, or a version whose factory is not
|
|
# callable. Only used to restore the clock, never load-bearing.
|
|
resume_at = None
|
|
try:
|
|
_settle_loop(inst, mode, dm, result, freezer)
|
|
finally:
|
|
if resume_at is not None:
|
|
try:
|
|
freezer.move_to(resume_at)
|
|
except (AttributeError, TypeError, ValueError):
|
|
pass
|
|
|
|
|
|
def _settle_loop(inst, mode, dm, result, freezer) -> None:
|
|
tick = getattr(freezer, "tick", None) if freezer is not None else None
|
|
for _ in range(EMPTY_RECHECK_FRAMES):
|
|
if tick is not None:
|
|
# timedelta rather than a bare float: freezegun has accepted a
|
|
# number only since 1.x, and a stale pin would raise here.
|
|
try:
|
|
tick(timedelta(seconds=EMPTY_RECHECK_STEP))
|
|
except (AttributeError, TypeError, ValueError):
|
|
# Pacing is best-effort; a freezegun that will not take a
|
|
# timedelta just means this probe runs without advancing time.
|
|
pass
|
|
else:
|
|
# No frozen clock, so the real one has to do the advancing. Without
|
|
# this the 48 frames run in microseconds, elapsed time stays ~0, and
|
|
# a scroll driven by elapsed time never moves -- which is exactly
|
|
# the plugin this check is trying not to slander.
|
|
time.sleep(EMPTY_RECHECK_STEP)
|
|
try:
|
|
result.display_returned = _render_mode_again(inst, mode)
|
|
except Exception as e: # noqa: BLE001
|
|
# Deliberately broad: this calls a plugin's display(), which can
|
|
# raise anything. Recorded rather than swallowed -- a mode that
|
|
# renders one good frame and then crashes on the next is broken,
|
|
# and returning silently here reported it as passing. The frame
|
|
# already captured stays on the result so the failure is still
|
|
# inspectable.
|
|
result.error = repr(e)
|
|
return
|
|
image = dm.get_image()
|
|
if _has_content(image):
|
|
result.image = image
|
|
result.overflow = dm.check_overflow()
|
|
return
|
|
|
|
|
|
def _freeze(freeze_time: Optional[str]):
|
|
"""Context manager that freezes wall-clock time when freeze_time is given,
|
|
so time-dependent plugins (clocks, countdowns) render deterministic goldens."""
|
|
if not freeze_time:
|
|
return contextlib.nullcontext()
|
|
try:
|
|
from freezegun import freeze_time as _ft
|
|
except ImportError as e: # pragma: no cover - only hit without the dep
|
|
raise RuntimeError(
|
|
"freeze_time requires the 'freezegun' package (pip install freezegun)"
|
|
) from e
|
|
return _ft(freeze_time)
|
|
|
|
|
|
def render_plugin_matrix(
|
|
plugin_id: str,
|
|
plugin_dir: Path,
|
|
config: Optional[Dict[str, Any]] = None,
|
|
mock_data: Optional[Dict[str, Any]] = None,
|
|
sizes: Optional[List[Tuple[int, int]]] = None,
|
|
run_update: bool = True,
|
|
freeze_time: Optional[str] = None,
|
|
) -> List[RenderResult]:
|
|
"""Render every (size, mode) combination for a plugin.
|
|
|
|
Returns a flat list of RenderResult. A fresh plugin instance is built per
|
|
(size, mode) so state never leaks between screens. Pass freeze_time (e.g.
|
|
"2025-08-01 15:25:00") to make time-dependent plugins reproducible.
|
|
"""
|
|
plugin_dir = Path(plugin_dir)
|
|
manifest = load_manifest(plugin_dir)
|
|
# Start from config_schema.json defaults so the plugin behaves like a real
|
|
# install; explicit caller config still wins over a schema default.
|
|
config = merge_config(
|
|
merge_config({"enabled": True}, load_config_defaults(plugin_dir)),
|
|
config or {})
|
|
sizes = sizes or DEFAULT_TEST_SIZES
|
|
results: List[RenderResult] = []
|
|
|
|
# The largest panel in this run. Every (smaller) canvas is padded out to it
|
|
# so a coordinate meant for the biggest configuration is still caught when
|
|
# rendering a smaller one, instead of being clipped into a false pass.
|
|
extent = (max(w for w, _ in sizes), max(h for _, h in sizes))
|
|
|
|
# One cache for the whole matrix: see _instantiate. The display manager
|
|
# stays per-render (the bounds checking depends on that); only fetched data
|
|
# is shared.
|
|
from src.plugin_system.testing import MockCacheManager
|
|
cache_manager = MockCacheManager()
|
|
for key, value in (mock_data or {}).items():
|
|
cache_manager.set(key, value)
|
|
|
|
with _freeze(freeze_time) as freezer:
|
|
for width, height in sizes:
|
|
results.extend(_render_size(
|
|
plugin_id, manifest, plugin_dir, config, mock_data or {},
|
|
width, height, run_update, extent, cache_manager, freezer,
|
|
))
|
|
|
|
return results
|
|
|
|
|
|
def _render_size(plugin_id, manifest, plugin_dir, config, mock_data,
|
|
width, height, run_update, extent,
|
|
cache_manager=None, freezer=None) -> List[RenderResult]:
|
|
"""Render every mode at one size. A fresh instance per mode avoids state leaks."""
|
|
results: List[RenderResult] = []
|
|
|
|
# Discover modes once per size (instance build can depend on config).
|
|
try:
|
|
probe_dm = BoundsCheckingDisplayManager(width=width, height=height, overflow_extent=extent)
|
|
probe = _instantiate(plugin_id, manifest, plugin_dir, config, mock_data, probe_dm,
|
|
cache_manager)
|
|
modes = list_modes(probe, manifest, plugin_id)
|
|
except Exception as e: # noqa: BLE001 — surface any load failure as a result
|
|
return [RenderResult(plugin_id, width, height, "<load>", error=repr(e))]
|
|
|
|
for mode in modes:
|
|
result = RenderResult(plugin_id, width, height, mode)
|
|
dm = BoundsCheckingDisplayManager(width=width, height=height, overflow_extent=extent)
|
|
try:
|
|
inst = _instantiate(plugin_id, manifest, plugin_dir, config, mock_data, dm,
|
|
cache_manager)
|
|
if run_update:
|
|
try:
|
|
inst.update()
|
|
except _TOLERATED_UPDATE_ERRORS as e:
|
|
# Expected when CI / headless dev has no network: record it
|
|
# (surfaced in the report) but don't fail the run.
|
|
result.update_error = repr(e)
|
|
logger.debug("update() connectivity error for %s [%s]: %s", plugin_id, mode, e)
|
|
except Exception as e: # noqa: BLE001 — a non-network update() failure is a real bug
|
|
# A regression in update() must not pass green just because
|
|
# display() survives, so treat it as a failure of this render.
|
|
result.error = repr(e)
|
|
logger.warning("update() raised a non-connectivity error for %s [%s]: %s",
|
|
plugin_id, mode, e)
|
|
if result.error is None:
|
|
result.display_returned = _render_mode(inst, mode)
|
|
result.image = dm.get_image()
|
|
result.overflow = dm.check_overflow()
|
|
# A blank first frame is not proof of a blank mode; see
|
|
# _settle_empty_frame.
|
|
_settle_empty_frame(inst, mode, dm, result, freezer)
|
|
except Exception as e: # noqa: BLE001 — a display crash is a real failure
|
|
result.error = repr(e)
|
|
results.append(result)
|
|
|
|
return results
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Golden-image comparison
|
|
# ---------------------------------------------------------------------------
|
|
|
|
def compare_images(rendered: Image.Image, golden: Image.Image,
|
|
max_delta: int = 0, max_diff_pixels: int = 0) -> Tuple[bool, int, int]:
|
|
"""Compare two images. Returns (ok, diff_pixel_count, max_per_channel_delta).
|
|
|
|
Tolerances default to exact match; bump them only to absorb known platform
|
|
anti-aliasing noise (requires a pinned Pillow + bundled fonts for stability).
|
|
"""
|
|
if rendered.size != golden.size:
|
|
return False, rendered.size[0] * rendered.size[1], 255
|
|
a = rendered.convert("RGB")
|
|
b = golden.convert("RGB")
|
|
diff = ImageChops.difference(a, b)
|
|
bbox = diff.getbbox()
|
|
if bbox is None:
|
|
return True, 0, 0
|
|
# Count pixels whose largest per-channel delta exceeds the allowed tolerance,
|
|
# and track the worst delta seen (for reporting).
|
|
diff_pixels = 0
|
|
observed_max = 0
|
|
for px in diff.crop(bbox).getdata():
|
|
m = max(px) if isinstance(px, tuple) else px
|
|
if m > observed_max:
|
|
observed_max = m
|
|
if m > max_delta:
|
|
diff_pixels += 1
|
|
# Pass when the number of out-of-tolerance pixels is within budget.
|
|
ok = diff_pixels <= max_diff_pixels
|
|
return ok, diff_pixels, observed_max
|
|
|
|
|
|
def golden_path(golden_dir: Path, width: int, height: int, mode: str) -> Path:
|
|
"""Location of a golden image: <golden_dir>/<WxH>/<mode>.png.
|
|
|
|
The mode is sanitized to a safe basename so a mode name with '/' or '..'
|
|
can't read or write outside the golden directory.
|
|
"""
|
|
return Path(golden_dir) / size_label(width, height) / f"{safe_mode_filename(mode)}.png"
|
|
|
|
|
|
def compare_to_goldens(results: List[RenderResult], golden_dir: Path,
|
|
max_delta: int = 0, max_diff_pixels: int = 0) -> List[RenderResult]:
|
|
"""Compare rendered results against committed goldens, mutating each result's
|
|
golden_* fields. Results with no golden file on disk are left unchecked."""
|
|
for r in results:
|
|
if r.image is None:
|
|
continue
|
|
gp = golden_path(golden_dir, r.width, r.height, r.mode)
|
|
if not gp.exists():
|
|
continue
|
|
r.golden_checked = True
|
|
with Image.open(gp) as g:
|
|
ok, diff_pixels, observed_max = compare_images(
|
|
r.image, g, max_delta=max_delta, max_diff_pixels=max_diff_pixels)
|
|
r.golden_ok = ok
|
|
r.golden_diff_pixels = diff_pixels
|
|
r.golden_max_delta = observed_max
|
|
return results
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Fill / scale-up check
|
|
# ---------------------------------------------------------------------------
|
|
#
|
|
# Overflow catches content that is too BIG for a panel; nothing catches
|
|
# content that stays tiny on a panel much larger than the plugin's design
|
|
# size (e.g. 128x32 content in the corner of a 256x128 renders "green").
|
|
# These helpers measure how much of the panel the lit content spans so the
|
|
# harness can flag plugins that don't scale up.
|
|
|
|
# A pixel counts as "lit" above this luminance — low enough to catch dim
|
|
# content, high enough to ignore near-black noise.
|
|
_LIT_THRESHOLD = 16
|
|
# Content must span at least this fraction of an axis that is >= 2x the
|
|
# design size. Lenient on purpose: margins are fine, a tiny corner is not.
|
|
_MIN_FILL_EXTENT = 0.5
|
|
|
|
|
|
def fill_metrics(image: Image.Image) -> Tuple[float, float, float]:
|
|
"""Measure lit-content coverage: (extent_x, extent_y, ink_ratio).
|
|
|
|
extent_* are the lit bounding box's spans as fractions of the panel;
|
|
ink_ratio is the fraction of pixels lit (reporting only — sparse pixel
|
|
fonts legitimately have low ink ratios)."""
|
|
lit = image.convert("L").point(lambda p: 255 if p > _LIT_THRESHOLD else 0)
|
|
bbox = lit.getbbox()
|
|
if bbox is None:
|
|
return (0.0, 0.0, 0.0)
|
|
extent_x = (bbox[2] - bbox[0]) / image.width
|
|
extent_y = (bbox[3] - bbox[1]) / image.height
|
|
ink = sum(1 for p in lit.getdata() if p) / (image.width * image.height)
|
|
return (extent_x, extent_y, ink)
|
|
|
|
|
|
def check_empty_claimed(results: List[RenderResult],
|
|
strict: bool = False) -> List[RenderResult]:
|
|
"""Flag a mode that rendered nothing without reporting "no content".
|
|
|
|
The display controller skips a mode whose ``display()`` returns False, and
|
|
treats anything else -- including None -- as "content was shown". A mode
|
|
that draws nothing and does not return False therefore holds whatever is on
|
|
the panel for its whole display duration. Since a mode switch clears first,
|
|
that is a blank screen. Two sports plugins shipped exactly this: their
|
|
``display()`` returned None on every path, so an out-of-season league sat
|
|
blank for its full duration rather than being rotated past.
|
|
|
|
Warn-only by default, because a blank frame is not automatically wrong: a
|
|
scroll mode whose first frame is its blank scroll-in buffer renders empty
|
|
and is behaving correctly. ``strict=True`` sets ``empty_claimed`` such that
|
|
``RenderResult.ok`` fails -- opt in per plugin via harness.json
|
|
``{"empty_check": "strict"}`` once its modes are known to draw on the
|
|
fixture data.
|
|
|
|
Note this can only catch what the fixtures actually render. A plugin whose
|
|
harness fixture seeds content never exercises its empty path here; the
|
|
source-level gate in the plugins repo covers that case.
|
|
"""
|
|
for r in results:
|
|
if r.image is None or r.error is not None:
|
|
continue
|
|
# An explicit False is the plugin correctly saying "nothing to show".
|
|
if r.display_returned is False:
|
|
continue
|
|
if r.image.convert("L").point(
|
|
lambda p: 255 if p > _LIT_THRESHOLD else 0).getbbox() is not None:
|
|
continue
|
|
r.empty_claimed = True
|
|
if strict:
|
|
r.empty_ok = False
|
|
return results
|
|
|
|
|
|
def check_scale_up(results: List[RenderResult],
|
|
design_size: Tuple[int, int] = (128, 32),
|
|
min_extent: float = _MIN_FILL_EXTENT,
|
|
strict: bool = False) -> List[RenderResult]:
|
|
"""Flag renders that leave a big panel mostly empty.
|
|
|
|
For each result whose panel is at least 2x the design size on an axis,
|
|
require the lit content to span >= min_extent of that axis. Mutates the
|
|
results' fill_* fields. In the default warn-only mode fill_ok is left
|
|
None (reported, never failing); strict=True sets fill_ok=False, which
|
|
fails RenderResult.ok — opt in per plugin via harness.json
|
|
{"fill_check": "strict"} once its adaptive layout is in place.
|
|
"""
|
|
design_w, design_h = design_size
|
|
for r in results:
|
|
if r.image is None or r.error is not None:
|
|
continue
|
|
check_x = r.width >= 2 * design_w
|
|
check_y = r.height >= 2 * design_h
|
|
if not (check_x or check_y):
|
|
continue
|
|
extent_x, extent_y, _ink = fill_metrics(r.image)
|
|
r.fill_checked = True
|
|
r.fill_extent = (round(extent_x, 3), round(extent_y, 3))
|
|
underfilled = ((check_x and extent_x < min_extent)
|
|
or (check_y and extent_y < min_extent))
|
|
if underfilled and strict:
|
|
r.fill_ok = False
|
|
elif not underfilled:
|
|
r.fill_ok = True
|
|
# warn-only underfill: fill_ok stays None; fill_extent tells the story
|
|
return results
|
|
|
|
|
|
def write_goldens(results: List[RenderResult], golden_dir: Path) -> int:
|
|
"""Write each successfully-rendered result to its golden path. Returns count."""
|
|
written = 0
|
|
for r in results:
|
|
if r.image is None or r.error is not None:
|
|
continue
|
|
gp = golden_path(golden_dir, r.width, r.height, r.mode)
|
|
gp.parent.mkdir(parents=True, exist_ok=True)
|
|
r.image.save(gp, format="PNG")
|
|
written += 1
|
|
return written
|