84043468 also gated the ESPN chunk fetches and the background data
service's workers, for the hourly sports refresh. A burst test on hdpi
(baseball and football refreshing every 5 minutes, 10-minute soaks, G F F G):
G prefetch gated only 0.87%, 0.83% late; 6+ late 20, 16; fetches 0.3-1.8s
F + fetch threads gated 1.23%, 1.05% late; 6+ late 12, 16; fetches 1.6-4.6s
Every parked fetch thread wakes at each swap and has to take the GIL again
just to park at the end of the window, so twenty of them cost more than
they saved, and the fetches ran two to three times as long. The plugins'
own copies of espn_dates (half the burst) were never gated anyway.
espn_dates and the background data service go back to main's versions and
the module-level active gate goes. Kept from that commit: the render thread
is never gated, a live refresh from another thread can't take its place,
and nested blocks keep the outer boundary.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codacy (Bandit B110, B108). The bench's silent except now prints why it read
config.json directly. The stats file's fixed name in /dev/shm is safe:
write() goes through mkstemp and os.replace, which replaces a planted
symlink instead of following it; the comment says so and marks it nosec.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The gate never parks a thread inside logging, threading, importlib or the
cache, and it matched those as substrings of each frame's file path. On
GitHub's runners Python lives under /opt/hostedtoolcache, so every stdlib
frame said "cache" and the gate never parked anything -- three tests failed
there and passed here. A virtualenv under ~/.cache would have done the same
on a Pi. Match the frame's module name (f_globals['__name__']) instead.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The hourly sports refresh froze hdpi's Vegas scroll for 1.9s: about twenty
espn-chunk threads fetching and parsing at once, and the render thread
(and the stall watchdog) queued behind all of them for the GIL. The
prefetch gate only covered the prefetch thread.
Vegas now makes its gate the active one (render_gate.set_active), and
render_gate.yielding() gives way through it when there is one and does
nothing otherwise. espn_dates wraps each chunk fetch in it (behind the
same import fallback as json_body, for the copies plugins bundle), and the
background data service wraps each worker. The render thread is never
gated -- the first thread to swap is exempt, and a plugin pushing a live
refresh from its update thread cannot take its place -- and nested blocks
keep the outermost frame as the boundary for the lock checks.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two runs per arm, about 81,000 frames each, order A B C C B A:
A step 1 as is 0.90% late, 20.1 per 10k two+ refreshes late
B switch_interval_ms 1 0.78% late, 15.8 per 10k
C prefetch_gate 0.60% late, 2.5 per 10k
No freezes in any arm, and the next group was ready at every strip
extension, so parking the prefetch thread (3-6s per 8-minute run) cost
nothing visible. The gate is now on unless vegas_scroll.prefetch_gate is
false; on a stock binding it cannot work and says so at INFO once a run
rather than warning on every install. switch_interval_ms stays off.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): answer unhandled api_v3 errors from one blueprint handler
Fifty-three api_v3 routes ended in a copy of the same catch-all: log the
traceback, return {status, "An error occurred; see logs for details",
details: describe_exception(e)} with a 500. They are replaced by one
errorhandler on the api_v3 blueprint that returns exactly that body.
It lives on the blueprint rather than falling through to app.py's global
handler because the two answers differ: the global one adds
error_code: UNKNOWN_ERROR, and api_client.js sends a body with an
error_code to the error modal and one without to a plain toast. A
blueprint handler also gives tests that mount api_v3 on a bare Flask app
the same answer the real app gives.
Only handlers that were byte-for-byte that shape were removed (matched on
the AST, and each rewritten function re-parsed and compared). Handlers
with their own message, extra keys, operation-history records or cleanup
stay, as does execute_plugin_action's step-1 handler, which sits inside
an `except subprocess.TimeoutExpired` arm that would otherwise turn a
plugin's timeout into a 408.
HTTPExceptions raised inside a route go back as themselves in the global
handler's 4xx shape. Where a removed catch-all used to swallow one (only
delete_plugin_asset's non-silent get_json() is reachable), a malformed
request now gets its 415/400 instead of a 500.
Most of the diff is re-indentation from unwrapping the try blocks;
`git diff -w` shows the real change.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): plugin action errors name the real failure, not UnboundLocalError
execute_plugin_action bound a local `logger` in its JSON-parsing arm,
which made `logger` local to the whole function. Every other
`logger.error` in it then raised UnboundLocalError, so a failing OAuth
step-1 script was reported as "UnboundLocalError: cannot access local
variable 'logger'" -- from the step-1 handler, and before the previous
commit from the route's outer catch-all too. Use the module logger.
Found by comparing every api_v3 route's forced-failure response before
and after the catch-all consolidation.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): drop the error category and exception-name code guessing
WebInterfaceError derived an ErrorCategory from every error code and put
it in each structured error body as `error_category`. Nothing reads it:
not the web UI (static/ and templates/), not the tests beyond the ones
pinning the mapping itself, and not any plugin in ledmatrix-plugins. The
enum, the inference table and the JSON key go.
from_exception() could also guess an error code from the exception's
class name ("Config" -> CONFIG_LOAD_FAILED, and so on). Every caller
passes a code, so the guess never ran; error_code is now required.
suggested_fixes stays: the error dialog in static/v3/js/utils/
error_handler.js lists them.
The REST reference loses error_category and says what an unanticipated
exception in an /api/v3 route answers.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one call for the from_exception error responses
Nine plugin routes built a structured error by hand:
from src.web_interface.errors import WebInterfaceError
error = WebInterfaceError.from_exception(e, ErrorCode.X)
return error_response(error.error_code, error.message,
details=error.details, context=error.context,
status_code=500)
That is now exception_error_response(e, ErrorCode.X) in api_helpers, so
error_response() is the only structured-error entry point the routes
use. The three operation-history routes never passed the context, and
with_context=False keeps their bodies exactly as they were; a test
compares the helper against the hand-written pair for both forms.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): one api_v3 error-response path
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(plugins): one resolver for plugin id -> directory
Five places mapped a plugin id to its directory, each with its own rules
and each re-reading manifests per lookup: PluginManager discovery and
get_plugin_directory, PluginLoader.find_plugin_directory,
PluginStoreManager._find_plugin_path / list_installed_plugins, and
state_reconciliation.disk_plugin_ids. They disagreed on backup dirs,
on whether the manifest id or the directory name is the id, on duplicate
ids and on path safety.
src/plugin_system/plugin_dirs.py now holds the rules once:
PluginDirectoryIndex scans one directory and reads each manifest once;
resolve_plugin_dir() searches directories in order. What legitimately
differs per caller is an explicit argument: search dirs (discovery and
the loader: configured dir only; the store: configured then sibling
plugins/), ledmatrix- prefix (not for the store), case folding (loader
only), manifest pass (not for get_plugin_directory, whose discovery map
already holds it).
Behaviour changes, all for layouts installs do not produce:
- a directory whose manifest declares the id beats one merely named for
it (discovery already worked this way; the loader and store now agree)
- the store searches the configured dir completely before plugins/
- backup and hidden dirs are skipped everywhere (the loader's case and
manifest scans and list_installed_plugins used to return them)
- duplicate ids resolve deterministically (exact name, then
ledmatrix-<id>, then by name) with a one-time warning; discovery no
longer lists the id twice
- disk_plugin_ids / list_installed_plugins report manifest ids, falling
back to the directory name; auto-update looks the directory up
- ids that are not one plain path segment resolve to nothing in every
caller (the loader used to truncate them, the store to join them)
The .standalone-backup- marker is one constant, BACKUP_MARKER, used by
store_manager's rename-aside names and every lookup.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): one plugin-directory resolver
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): use src.logging_config in the web process; routine requests to DEBUG
The web interface had its own logging setup (web_interface/logging_config.py)
that replaced the root handlers with a plain stdout formatter. The web
service's journal lines therefore never carried a syslog priority, so
`journalctl -p err -u ledmatrix-web` returned nothing while errors were
logged, and the line shape differed from the display's (the log viewer's
prefix stripping only matched the display format). It also ran after the
module-level managers were built, so their INFO lines at import (including
"Re-removed N uninstalled plugin(s)") were dropped.
app.py now calls src.logging_config.setup_logging() first thing, the same as
run.py: journald priorities under systemd, LEDMATRIX_DEBUG honoured,
LEDMATRIX_JSON_LOGGING still selects JSON.
Per-request logging moves to web_interface/request_logging.py. Every request
used to be logged at INFO, so the UI's polling filled the journal
("GET /api/v3/errors/summary - 200" every minute per tab). Now a successful
GET/HEAD/OPTIONS is DEBUG, a successful write is INFO, 4xx WARNING, 5xx
ERROR. Durations use perf_counter and print to 0.1ms.
The duplicate module is deleted; nothing else imported it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(web): one thread-safe TTL cache for the web process
web_interface/cache.py becomes a small TTLCache class (lock-guarded,
monotonic clock) with the existing get_cached/set_cached/delete_cached/
invalidate_cache helpers kept on top of a shared instance, so the api_v3
callers are unchanged.
Bugs fixed:
- set_cached(ttl_seconds=...) ignored its TTL; only the reader's value
counted and get_cached defaulted to 60s. An entry now expires after the TTL
it was stored with; a reader's ttl_seconds can only shorten that. Both
current callers pass the same value on both sides (fonts_catalog 300s,
system_status 10s), so their observable TTLs are unchanged.
- get_cached deleted expired keys without a lock; two threads reading the
same expired key could raise KeyError (reproduced), which the endpoints
turned into a 500.
app.py's two hand-rolled systemctl caches (_ap_mode_cache, 30s, and
_ledmatrix_service_cache, 15s) now share one helper over a private
TTLCache, with the same TTLs. The AP-mode check used to retry on every
request after a failure (and log an ERROR each time); a failure now keeps the
last known answer for the TTL, as the display-service check already did. With
no systemctl at all (a dev machine) it answers False without forking.
Left alone as not TTL memoisation: the gzip cache (size-bounded, keyed by URL
and version), the settings search index (keyed by installed-plugin set), the
widget bundle (keyed by file fingerprint) and CacheManager (cross-process).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): web logging and TTL cache
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): only ask systemctl about known units
Codacy flagged the systemctl argv built from a variable. The unit now has
to be one of two literals, and anything else raises.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): response_time_ms reads the same clock request_logging stamps
request_logging now stamps request.start_time from perf_counter, but
success_response still subtracted it from time.time(), so metadata
reported ~1.8e12 ms. Found testing on ledpi.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(fonts): one BDF loader and one BDF rasterizer
BDF faces were loaded three ways (FontManager._load_bdf_font,
element_style._load_bdf, DisplayManager._load_fonts) and drawn by two
copies of the same per-pixel loop (DisplayManager._draw_bdf_text and the
plugin test harness's "replicated" copy), which golden images and
check_plugin/dev_server previews rely on matching the panel.
src/common/bdf_font.py now owns both:
- load_bdf_face(path, size) -> (face, realised_px): native-strike fallback
for sizes the file lacks, one bounded LRU cache keyed on path, size and
mtime. FontManager, element_style and DisplayManager delegate to it;
read_bdf_native_size moves here (the old names delegate).
- draw_bdf_text(draw, text, x, y, face, color, clip): builds each glyph as
a 1-bit mask and fills it with ImageDraw.bitmap instead of a draw.point
per pixel. A blending Draw (RGB image, "RGBA" mode) keeps the point path
so translucent colours still blend.
Pixel-identical: 220,032 renders (every bundled BDF at native and
off-strike sizes, 14 strings, 4 colours, clipped on every edge, through
each old loader x rasterizer) match origin/main byte for byte.
test/test_bdf_font.py keeps a lightweight version against a frozen copy of
the old loop. DisplayManager._draw_bdf_text goes from 1.4-23 ms to about
0.1 ms per string (the old loop re-read FreeType's buffer as a Python list
for every pixel).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(testing): harness calendar_font is sized like the panel's
VisualTestDisplayManager built its 5x7 calendar_font / bdf_5x7_font as a
bare freetype.Face. With no size set its ascender reads 0, so BDF text
drawn with it landed 6px above where DisplayManager draws it -- entirely
off the canvas at y=0 -- and get_font_height() returned 0. Golden images
and check_plugin / dev_server previews showed text the panel does not.
Load it through load_bdf_face at the panel's 7px, so it is the very face
DisplayManager uses. Across the differential run this changes only the
cases drawn with the harness's own calendar_font (968 of 220,032), which
now match the panel's output.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(fonts): one BDF face per thread
The shared face cache now hands every loader (FontManager, element_style,
DisplayManager, the harness) the same freetype.Face. FreeType does not allow
two threads to use one face at once, since load_char rewrites its glyph
slot, so key the cache by thread as well.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(sports): wrap the sports_card twins that behave identically
SportsCoreSharedMixin (switch mode, via each scoreboard's sports.py) and
sports_card (scroll/Vegas mode, via game_renderer.py) carried the same
helpers twice. test/test_sports_twins.py now calls every pair with the
same inputs -- the eight scoreboards' harness fixture games in flat,
flat+nested and nested-only shapes, plus edge cases (favourites by id and
abbreviation, NRL's colliding abbreviations, missing and non-numeric
scores, bad zones, out-of-range dates, shared font faces).
Identical pairs become thin wrappers over the sports_card function:
_card_option, _vs_text, _format_game_time, _coerce_rgb, _crisp_size (with
the class's own tables), _unshare_element_fonts (with the class's own
element map, via a new optional argument), and the colour/month/weekday/
font-grid tables (dicts copied, not aliased). _format_game_date shares the
card's formatting body but keeps its own setting, weekday zone and month
table; _schema_font_size shares the parser but keeps its per-class cache,
because a reloaded plugin gets new classes and a shared path cache would
stop it seeing an edited schema. _resolve_font_size agrees but keeps its
body so it still dispatches through the overridable hooks.
No behaviour change: old and new mixin/card agree on all 22,994
comparisons over the test corpus, and the pairs that do differ
(favourite-result colours on nested payloads and by favourites source,
the weekday's timezone, the element-name map, per-mode colours) are left
alone and pinned in TestPinnedDivergence for an owner decision.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(sports): pin that an ambiguous NRL abbreviation tints in both modes
NRL's resolver passes a shared abbreviation ("NEW") through with an error
and its _is_favorite_game matches ids only, but both favourite-colour
helpers match on abbreviation as well, so both display modes tint a
Knights or Warriors result for a user who typed "NEW". The twins agree;
neither consults the _favorite_key seam. Pinned so a fix is deliberate.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(install): generate the web sudoers rules in one place
/etc/sudoers.d/ledmatrix_web was written by two copies of the same
allow-list: a heredoc in first_time_install.sh Step 10 and a block of
echo lines in scripts/install/configure_web_sudo.sh. They drifted before
(safe_pip_install.sh was granted by one only), and a test existed just
to catch that.
Both now call web_sudoers_rules() from the new
scripts/install/lib_sudoers.sh and keep their own validate (visudo -c),
install and confirm flows.
- first_time_install.sh output is byte-for-byte unchanged, so a device
re-running the installer gets "already up to date". If the library is
missing, Step 10 keeps the installed file and carries on, the same way
it handles rules that fail visudo (an empty file would pass visudo).
- configure_web_sudo.sh now writes the installer's layout: same 18 rules,
different comments and order. It still leaves out reboot, poweroff and
journalctl when they are missing; the library does that for both.
The drift test now pins the generator's grants, checks that neither
installer writes rules of its own, and runs each installer's call line
to check the argument order. Tests that read the rule text now read the
library.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(install): detect the web service user in one function
first_time_install.sh pasted the same WEB_SERVICE_USER detection block
three times (Step 3.1's fallback, the plugin-repos setup and Step 11).
The copies were identical apart from comments; they now call
detect_web_service_user(), whose body is that block unchanged.
Behaviour is the same: the function sets the same global and always
returns 0, as the inline if-chain did. Checked on Linux against all
three original copies across 13 layouts (installed unit with and without
User=, the repo as shipped, each grep branch, template placeholders).
The comment notes that the install_web_service.sh / install_service.sh
greps no longer match anything, so until Step 8 installs the unit the
result is "root". That behaviour is left as it was.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The sports plugins cache whole season schedules: 53MB for MLB, 18MB for
NHL, 17MB for NCAA baseball. On a Pi 4, orjson.loads of the MLB file
takes ~1.8s with the GIL held, and every thread in the display service
waits -- the stall watchdog caught the render thread frozen 0.5-1.3s with
the interpreter itself blocked, right on these reads. When a season record
expired, DiskCache.get paid that whole parse only to find the timestamp
too old and throw the result away.
CacheManager.set now writes timestamp and ttl ahead of the data, and
DiskCache.get reads them from the first 256 bytes of the file, applying
the same rule as before (a per-entry ttl wins over max_age; no limit
means never stale). A record that is stale is refused without being
parsed. Files in the old layout, and records from other writers, don't
match the header and are parsed in full as before.
Also: ESPN responses in the background data service and espn_dates are
parsed with orjson when it is installed (src/common/json_body.py). The
stdlib parser behind response.json() takes 3.1s on the MLB season
against orjson's 1.8s, both with the GIL held. espn_dates imports it with
a fallback, since plugins bundle copies of that module for older cores.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs(scroll): explain the tear across the middle on fast scrolls
A 1:32-multiplexed 64-row panel lights row 31 almost a whole refresh after
row 32, so fast scrolls show a sideways offset at mid-height of about
speed x refresh period. Documents the cause, how to read the real refresh
rate (show_refresh_rate prints with a carriage return), what was measured on
a single-chain 2x128x64 Pi 4 (pwm_bits, gpio_slowdown and an uncapped
refresh barely help; gpio_slowdown 2 glitches), and the fix that does help:
fewer pixels per output via parallel chains.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(scroll): limit the 1:32 row-pair explanation to panels that scan that way
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(redaction): make URL-userinfo redaction linear, not quadratic
_REDACT_URL_USERINFO could start a match at every letter of a run of
scheme characters, and each attempt read to the end of the run looking
for `://`. On a long unbroken run of letters or digits (a hex digest, an
ID, part of a response body) that is quadratic: 1.6s for 20k characters.
The display service redacts every message, stack trace and context value
it publishes in the error snapshot, holding the aggregator lock, and
re.sub holds the GIL for the whole call, so one such exception stalled
every thread, render loop included (~0.5s measured for 20k chars of hex).
It also made test_snapshot_stays_small the slowest test in the suite by
far: 142s of a 383s run, 139s of it in this one regex.
A match may now only start where a run of scheme characters starts
(negative lookbehind). Leading digits and `+.-` are captured in group 1
so the substitution restores them, and the scheme still has to start
with a letter, so what gets redacted is unchanged: old and new output
were identical on 300k fuzzed inputs. 20k chars now take ~0.5ms, 200k
~6ms, and test_snapshot_stays_small takes 0.8s.
test/test_redaction.py pins the exact output for schemes that begin after
digits or `+.-`, and bounds 50k-character runs at 1s; against the old
pattern those timing tests fail at 3-11s each.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK
* fix(redaction): make Authorization-header redaction linear too
_REDACT_AUTH_HEADER matched the value's opening as `\s*["\']?\s*`: two
`\s*` separated only by an optional quote. With no quote, a whitespace
run could be split between them in every possible way, and when no
credential followed (end of text, or `,` `"` `<` ...) the engine tried
them all before giving up: quadratic, 8s for `authorization:` and 20k
spaces, 17s with `Proxy-Authorization:` (tried again at the inner
`authorization`). Same stall as the URL pattern: re.sub holds the GIL,
and the display service redacts everything it publishes.
The quote and the whitespace after it are now one optional unit,
`\s*(?:["\']\s*)?`, which matches the same strings with only one way to
split them. Output is identical to the old pattern on 300k fuzzed
inputs; 20k spaces now take ~1.6ms. A scan of all three redaction
patterns over prefix/run/suffix shapes finds none left that scales
superlinearly.
test/test_redaction.py pins exact output for quoted, tabbed, multi-line
and credential-less headers, and bounds header + 20k whitespace at 1s;
against the previous pattern those fail at 8-17s each.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KMXdS2S4NXTJ8ET96GymhK
---------
Co-authored-by: Claude <noreply@anthropic.com>
The Pixlet editor stop route restarts the display after a SIGKILL with
_run_systemctl_command, which starlark.py never imported (since #554). The
Starlark device-location resolver fell back to _ensure_cache_manager, which
#609 deleted; the resolver already accepts no cache manager. Both raised
NameError on the rare path that reaches them. pyflakes finds no other
undefined names in src/ or web_interface/.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
On hdpi the stall watchdog logged a 1.9s stall during the hourly sports
refresh that the soak report never had: its worst gap was 655ms. The frame
that ended the stall was recorded as static, so its interval was dropped.
"Scrolling" is DisplayManager's scroll state at the moment a frame is
presented, and it goes missing mid-scroll: it expires after 2s without
activity, and any thread can clear it. Plugins call
set_scrolling_state(False) from their own display() (news, stocks, the odds
ticker's fallback), and Vegas captures some of those on the render thread
between two of its own frames. Vegas sets the state again only after its
next frame, so that frame is recorded as static -- along with the capture
or stall it followed.
One static frame between two scrolling frames, with the scroll resuming
within RESUME_SECONDS (1s), is now a frame of the scroll and both of its
intervals count, the first at the scroll's own hold (clearing the state
drops the hold to 1 too). Two static frames in a row still end the scroll.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
250ms catches freezes; the hitches left on hdpi are frames 2-5 refreshes
late, which look like the render thread waiting for the GIL. At 30ms the
watchdog dumps those too, naming what the other threads were running when
the frame missed. It polls at a third of the threshold so a stall one poll
long is still seen, which costs some GIL time of its own: a diagnostic
setting, not one to soak with.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The render thread spends most of each refresh in SwapOnVSync with the GIL
released, then needs it back the moment the swap returns. With plugin
rendering on the prefetch thread, it often has to wait for it -- behind
bytecode for up to the switch interval, behind a GIL-holding C call for as
long as that takes -- and hdpi's late frames of 2-5 refreshes went up.
src/common/render_gate.py opens a window around each swap, up to just
before the refresh the swap will return on, and a profile hook on the
prefetch thread parks it outside that window. It is never parked holding a
lock the render thread also takes (the Vegas buffer, cache and state
locks, logging, threading, importlib, the cache), never when no frame has
been swapped for 50ms, and never for more than 50ms at a time.
Off by default and ignored on a binding that keeps the GIL in
SwapOnVSync, where the window would never let the prefetch run.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
render_bench.py (from the parallel perf/render-bench work) had its own
grading module, frame_pacing, with its own definition of a missed frame
and its own refresh estimate. The soak already had both in frame_timing,
so the two could have drifted apart on what "late" means.
The bench now gives the display manager a fresh FrameTimingRecorder,
drains it synchronously at the start and end of the graded run, and prints
frame_soak's report with frame_soak's verdict. Its workload is unchanged:
the synthetic strip, --busy load, the shared speed resolver, the
per-frame scrolling announcement. frame_pacing, its tests and its
src.common exports are removed; measure_refresh_hz moves to frame_timing,
where scroll_speeds.py now finds it.
Two ideas from frame_pacing carry over. The bench seeds the recorder with
the idle refresh it measures, so a loop that free-runs (the 827fps bug
the first bench caught) shows as early frames and one stuck at half rate
as late frames, where an estimate taken from their own intervals finds
both self-consistent. And the soak, which has no idle measurement, now
calls a run NOT LOCKED when its refresh estimate beats the configured cap.
The report also gives the rate held while rendering.
Docs: the bench becomes "Without the service" under "Soaking a rig",
keeping its hdpi numbers and the idle-vs-rendering refresh finding.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was no way to answer "does this hardware present every frame on
time?" other than watching the panel. `scripts/render_bench.py` drives the
production path -- a real DisplayManager and ScrollHelper, configured
through the same `scroll_config` resolver every ticker uses -- and grades
the run with a new `src.common.frame_pacing`, exiting non-zero when more
than 0.1% of frames slipped a refresh. Exit 2 when the run could not be set
up at all, so a rig that was never measured cannot pass by accident.
A missed frame is defined exactly: an interval that rounds up to at least
one more refresh than its frame hold asked for. The half-refresh rounding
boundary keeps a frame that ran 1ms long on a 10ms refresh out of the
count, because it still presented on the refresh it was meant to.
The verdict that matters more is NOT LOCKED. A loop that never blocked on
vsync reports a perfect zero misses while presenting nothing -- 8ms frames
on a 100Hz panel all land in the one-refresh bucket while running 25% too
fast -- so the report also checks the typical frame is not shorter than the
panel could physically present. That is what caught the first version of
this benchmark announcing its scrolling state once instead of per frame:
the state expires on an inactivity threshold, the dirty-tracking skip then
fires mid-scroll, and the loop free-ran at 827fps.
And the refresh is read back out of the frames rather than taken from an
idle measurement. Driving the matrix is bit-banging on the same machine, so
pushing frames slows the refresh: a Pi 4 on 512x64 measures 100.4Hz idle
and holds 96.3Hz while scrolling. Both are real, and grading against the
idle figure reports a locked loop as 4% slow -- or, once the gap passes
half a refresh, as missing every frame. The gap between the two is itself
worth watching: a rise in it is a render-cost regression even when nothing
is missed.
Measured on hdpi (Pi 4, 512x64, pwm_bits 8), two minutes each:
plain 95.44 fps, 8 missed of 11,449 (0.070%) PASS
--busy 2 95.41 fps, 3 missed of 11,445 (0.026%) PASS
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
The recorder counts freezes; it cannot say why. hdpi showed 1-2s freezes
in both the #628 and offscreen builds, one lining up with hockey's 2s
NHL fetch on the update thread, and nothing in the logs explained it.
StallWatchdog polls every 50ms from its own thread. When a scroll's last
frame is more than 250ms old (and a scroll is still running, so the end
of a scroll is not a stall), it logs the stack of the thread that
presented that frame and the top of every other thread's, then the
stall's length when frames resume. It also measures how late its own
wake-up was: if it was held up as long as the render thread, the whole
interpreter was blocked (C code holding the GIL), not one thread on a
lock. One dump per 30s at most; LEDMATRIX_STALL_WATCHDOG=0 disables it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The plugin-facing canvas (DisplayManager.image, draw, matrix) was one
shared object, so any plugin whose Vegas content needed it -- display
capture, scroll-content generation, narrowed rendering -- was deferred to
the render thread and fetched there one at a time. On hdpi that is most
plugins, and each fetch stalled the scroll: news ~320ms, hockey ~660ms,
in bursts whenever the strip extended.
DisplayManager.offscreen() gives the calling thread a canvas of its own.
image, draw and matrix are now properties that resolve to the thread's
surface while it is inside the block and to the shared canvas otherwise,
so the ~100 existing uses become thread-correct unchanged. Inside,
update_display(), the hardware half of clear(), and set_scrolling_state()/
set_frame_hold() are inert, so a plugin drawn for Vegas can neither reach
the panel nor re-pace the live scroll. render_size() is rebuilt on it.
capture_mode() now restores the previous state instead of clearing it,
so it cannot end suppression inside an offscreen block.
The adapter draws every path on its own canvas (_isolated_canvas) and
drops the copy-and-restore of the shared image, which from a background
thread would have written a stale frame back over the render loop's.
Background fetches take the plugin's update/display lock, waiting up to
2s for a running update() and skipping the plugin that round otherwise;
Vegas never took that lock, so render-thread captures already raced
update(). A background fetch that comes back empty is no longer queued
for the render thread.
vegas_scroll.offscreen_prefetch (default true) restores the old deferred
path when false. See docs/OFFSCREEN_RENDERING.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three gaps found by the first hdpi soaks:
- Intervals of 1s or more between two scrolling frames were dropped as
"gaps between scrolls". But the scrolling state lapses only after 2s, so
every 1-2s stall inside a scroll vanished from the report. Those are now
freezes (the gap bound is a 5s sanity limit), with a breakdown by length.
- A frame a whole refresh early means the swap did not wait for the panel.
Those are counted, and a soak with more than the threshold of them fails
as NOT LOCKED instead of reporting a flattering late rate.
- The refresh estimate took the lowest window it had seen, so one window of
non-blocking swaps halved it and made every early frame look on time. A
window may now lower it by at most 20%.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each scroller already logs its own stats line, but in different formats,
per source, and Vegas logs a healthy window only at DEBUG. None of it
answers the question a release has to answer on each rig: over a long
run, how often did a moving frame reach the panel late?
Every frame reaches the panel through DisplayManager.update_display, so
it is timed there once, whoever drew it: the blit (SetImage), the vsync
wait, and the interval since the previous frame. The render thread only
appends a tuple. A worker thread aggregates cumulative counters and
histograms and rewrites /dev/shm/ledmatrix_frame_stats.json every 10s
(RAM, so no SD wear).
A frame due after `hold` refreshes that lands one or more refreshes
later is "late": the panel repeated the previous frame, a visible hitch.
Gaps of 250ms+ inside a scroll are "freezes" (recomposes, handovers,
blocking calls), counted separately so one handover does not read as 40
missed refreshes. Static frames, the first frame of a scroll and gaps
between scrolls are not timed. The refresh period is estimated from the
frames themselves.
scripts/frame_soak.py runs next to the service as any user, diffs two
snapshots over a run (default 10 minutes), optionally keeps the web
preview's viewer marker fresh, and exits non-zero above 0.1% late
frames. It also reports whether the loaded rgbmatrix binding releases
the GIL. Documented under "Soaking a rig" in docs/SCROLL_PERFORMANCE.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Vegas advanced by elapsed time, blended neighbouring columns every frame,
and paced itself with a sleep to target_fps. On hdpi (4x128x64 on one
chain, a 120Hz cap the chain cannot reach, ~95-100Hz real) that ran at
73fps with target 90 and ~89fps with target 125: the sleep drifted
against the refresh and missed a vsync every few frames, and the blend
read as shimmer on the panel (and as "anti-aliased" text in the preview).
smooth_scroll now means the crisp pacing the plugin tickers already use:
a whole number of pixels per presented frame, each held for frame_hold
refreshes, with SwapOnVSync as the clock. The speed is solved against the
panel's measured refresh, timed from our own swaps once scrolling starts,
because the configured limit is only a cap -- at "120Hz" 90px/s solves to
3px every 4 refreshes, at the real ~97Hz to 1px every refresh. The old
blend stays available as sub_pixel_blend (default off).
With the web preview open, the render thread also PNG-encoded the whole
512x64 frame five times a second, 12-14ms each -- longer than a refresh.
Mid-scroll that encode now runs on a single-slot writer thread (Pillow
releases the GIL while compressing); static frames still write inline.
Measured on hdpi, 3-minute soak with the preview open: 3 of 17,280
frames held an extra refresh (0.02%), down from ~6-20%.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(web): show which plugins use each font, warn before deleting one
The Fonts tab lists font files from the web process's own scan, and the
plugins that register fonts run in the display process, so the tab had no
way to say whether a font was in use before deleting it.
The display service now publishes {catalog key: [plugin ids]} to the
shared cache (font_usage_snapshot, src/font_usage.py), built from the
loaded plugins' FontManager.register_manager_font() registrations. A
daemon thread checks every 10 s and writes only when the usage changed
(plus a daily refresh so cache cleanup cannot expire it); it never raises.
Families, aliases (press_start, four_by_six, ...) and paths are resolved
through FontManager's catalog to the file stem the Fonts tab keys rows by;
fonts outside assets/fonts are left out. Unloading a plugin drops its
registrations (new FontManager.forget_manager_fonts).
GET /api/v3/fonts/catalog merges used_by into each row per request (the
5-minute scan cache is copied, never edited): a list of plugin ids, or
null when the display service has not reported. The tab shows a Used by
column ("unknown" / "-" / ids, rendered as text) and deleting an in-use
font names the plugins in the confirmation, from a fresh read. The server
still refuses only system fonts. Catalog fetches bypass the browser's
5-second API cache, which otherwise served the pre-delete list.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix: call forget_manager_fonts through a hasattr check pylint can follow
getattr(..., None) then callable() is fine at runtime, but pylint's E1102
("not callable") can't see through it, and Codacy fails the check on it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(starlark): blank app locations use the device location, not San Francisco
A Starlark (Tidbyt) app whose Location field is blank rendered at its
author's hard-coded DEFAULT_LOCATION -- usually San Francisco -- even with
the device city set under General settings. A user in Charlotte, NC got San
Francisco weather and radar with nothing in config.json to explain it.
src/device_location.py fills unset location fields at render time (display
plugin and the web standalone render): the device city is geocoded once via
Open-Meteo, preferring a match in the configured state/country, and cached
permanently. A saved location always wins; if the lookup fails the field is
dropped so the app uses its own default, and the failure is not retried for
30 minutes.
Also fixes the config form: clearing a location omitted the key, and the
save merges, so the old value could never be removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(starlark): say what happens when the device location can't be used
A blank app Location only renders at the device's city when one is set and
the Open-Meteo lookup finds it. With no city, no match, or the geocoder
unreachable (retried after 30 minutes), the app gets no location and keeps
its author's default. The guide, the config page hint, CONFIG_REFERENCE and
the CHANGELOG entry now say so.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor: remove the skin system
Skins never rendered with the current scoreboard plugins: the only hook was
SportsCore._render_game in src/base_classes, which no plugin builds on, so
the UI and store already treated them as unsupported. The owner decided on
2026-09-23 to remove them outright.
Removed src/skin_system/ (runtime, base class, fixtures), skins/,
scripts/validate_skin.py and their tests; the store's "type": "skin"
installer, uninstaller and hide/refuse filters (the official registry lists
no skins); SchemaManager.inject_skin_selector; and GET /api/v3/skins.
Stored skin/skin_options config values are handled in the next commit.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(config): drop retired skin/skin_options keys instead of validating them
A config.json written while the skin system existed can carry skin and
skin_options in any plugin section, and most plugin schemas set
additionalProperties: false. They are no longer core plugin properties;
RETIRED_PLUGIN_KEYS in schema_manager lists them and
drop_retired_plugin_keys removes them (unless the plugin's own schema
declares the name) in prepare_plugin_config, which loading, hot reload,
GET /plugins/config and both web saves already share, and in
validate_config_against_schema for callers that validate a raw section.
POST /plugins/config and /config/main also drop them from the stored
section they merge into, so they leave config.json on the next save.
Tests cover the load path (real PluginManager.load_plugin: no schema
warning, not degraded), raw and prepared validation,
validate_all_plugin_configs, and the JSON, form and /config/main saves.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor: remove the unused src/base_classes package
No scoreboard plugin builds on src.base_classes: the nine monorepo
scoreboards ship their own sports.py and share code through src/common
(docs/SPORTS_UNIFICATION.md), and none of the third-party registry plugins
imports it. The one import anywhere, baseball-scoreboard's
rankings_manager.py, is a lazy import of ESPNDataSource in a class nothing
instantiates.
Removed the package and the eight test files that only tested it
(test_api_extractors, test_data_sources, test_sports_base_characterization,
test_sports_capabilities, test_sports_core_promotions,
test_sports_logo_cache_bounded, test_sports_modes_promotions,
test_sports_odds_fanout). test_common_is_hardware_free no longer lists
src.base_classes as a forbidden import, and comments in sports_helpers.py
and base_odds_manager.py stop pointing at it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: drop the skin system and src/base_classes from the docs
Deletes docs/SKIN_SYSTEM.md and docs/CREATING_SKINS.md and every link to
them (docs/README.md, README.md, PLUGIN_DEVELOPMENT_GUIDE.md, the /skins
section of REST_API_REFERENCE.md), the skin section of CLAUDE.md and the
term in PRODUCT.md. SPORTS_UNIFICATION.md now says src/base_classes was
removed and shared code lives in src/common, in the Layering section and
the view-model-contract rule. Other docs stop pointing at the removed
package. CHANGELOG records both removals under Unreleased.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(store): hide and refuse registry entries that aren't plugins
The skin filters went with the skin system, but a custom registry can still
list "type": "skin" entries, and installing one as a plugin would unpack it
into the plugins directory. PluginStoreManager.is_plugin_entry() (a missing
type means plugin) now hides non-plugin entries from the store and
custom-registry listings, and install refuses them, in the route with a
clear 400 and in _install_plugin_impl for any other caller.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(starlark): stop the root display service locking the web UI out
Reported after a fresh install: installing an app from the Starlark tab
failed with "install failed: Failed to install from repository", and so did
uploading a .star file and installing from a GitHub directory. The reporter
found the cause only by reading service logs, and fixed it with
sudo chown -R ledpi:ledpi /home/ledpi/LEDMatrix/starlark-apps
starlark-apps is gitignored, so it is never checked out -- it is created
lazily by whichever process reaches it first. Those processes run as
different users. systemd/ledmatrix.service is User=root and constructs this
plugin at startup, which is where _get_apps_directory() is called from;
systemd/ledmatrix-web.service runs as the login user and is what actually
installs apps.
The documented first step is to install pixlet and reboot, so on a fresh
machine the display service usually wins that race and mkdir() leaves the
directory root-owned. The web process then fails in _install_star_file() on
app_dir.mkdir(), which catches nothing, so PermissionError reaches the
route's outer `except Exception` and becomes the generic message the user
saw. All three install paths write to the same directory, which is why all
three failed.
The web user cannot repair this -- chown needs root. So root does it, on
every startup, which also heals machines already broken by this without the
owner having to find the chown themselves. It is a no-op when not root, when
the platform has no POSIX ownership, and when the checkout genuinely belongs
to root; a chown that fails warns rather than killing startup.
Also made the failure legible if the handover is ever prevented: a
PermissionError now names the directory, the automatic repair, and the
manual chown, instead of a message that names neither path nor cause.
Verified by mutation: dropping the handover call, chowning a genuinely
root-owned checkout, and letting a non-root process chown each fail their
own test. 121 starlark tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
* fix(starlark): address the review on the ownership repair
Findings from the automated review of #604.
Symlinks (CWE-59, the serious one). A root chown that follows links is a
privilege-escalation primitive: anyone able to write in starlark-apps could
point a link at a root-owned file and have the repair hand it over. Entries
are now read with os.lstat, symlinks are skipped outright, and the chown
passes follow_symlinks=False. Descendants are processed before the directory
itself, so the container does not change hands while its contents are still
being walked.
install_app() caught PermissionError in its broad handler and returned
False, which both routes report as a generic install failure -- the exact
shape of the bug this PR exists to fix, since the caller could not tell
"this app is broken" from "this process cannot write here". PermissionError
is now re-raised; every other failure still returns False.
The test fixtures skipped on bare Exception, which would have turned a
syntax error or NameError in the plugin into a green run. They now skip only
for a named absent dependency and re-raise anything else.
Also fixed the _Stat stub that failed in CI but passed locally: it carried
only st_uid/st_gid, and pathlib reads st_mode while walking. It now wraps
the real stat result and overrides ownership alone.
NOT taken: the CodeQL "information exposure through an exception" finding on
the hint response. Dropping `details` would contradict this package's
documented rule -- "if it returns 5xx, it says why" -- which
test_no_api_v3_handler_discards_its_exception enforces with an allowance
that may shrink and never grow. The Starlark routes are the ones that policy
was written for: they answered 500 with no detail for three releases.
describe_exception already redacts credentials and truncates. Keeping the
detail is the deliberate trade-off, so the finding is declined rather than
silently worked around.
Verified on hdpi with the updated code: a symlink to /etc/shadow planted in
starlark-apps was skipped while the directory was handed back, and
/etc/shadow stayed root:shadow.
Mutation-checked all three behaviours. The symlink test was vacuous on the
first attempt -- the link already had the target owner, so it was skipped
for the wrong reason and the mutation passed. It now forces the link to look
like it needs handing over, and fails when the check is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RRtqXDCnvnY6EQwhT5CV9
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(errors): serve /api/v3/errors/* from the display service's aggregator
The error aggregator is a per-process singleton and only the display
service runs plugins, so only its aggregator records anything. The routes
read the web process's own, empty one and always reported no errors.
The display service now publishes a bounded snapshot of its aggregator to
the shared cache (plugin_error_snapshot) from a daemon thread: at most once
every 10 s and only when something changed, never raising into the caller.
The routes read it and keep their response shapes, adding
snapshot_available, generated_at and clear_pending; exception text has
credentials redacted.
POST /errors/clear writes a clear request (plugin_error_clear_request) that
the display applies on its next 5 s tick via the new clear_before(), which
keeps errors recorded after the cutoff and rebuilds the counts. Until the
snapshot acknowledges the request, reads hide everything before the cutoff,
so a snapshot written just before the click cannot bring errors back. Adds
"all": true; cleared_count is null when only the display can know it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* feat(web): show plugin errors in the Logs tab
A compact panel under the log viewer: per-plugin error counts, repeating
errors (type, count, affected plugins, a sample message, last seen) and a
Clear button, with empty states for "no errors" and "display service
hasn't reported yet". Polls every 15 s while the tab is active; all text
goes through escapeHtml.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs: describe where plugin error reports come from and how clear works
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(errors): redact the published snapshot before clipping it
Keeping only a traceback's tail (or clipping a message) could cut an
`api_key=` marker off while keeping the secret after it, and the web side's
redaction would then have nothing to match. The display now redacts every
free-text field of the snapshot first. The patterns move to a Flask-free
src/redaction.py so the display service can use them; redact_text in the web
error handler uses the same function, unchanged in behaviour.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): apply on-demand, brightness and schedule changes mid-screen
The main loop read the on-demand mailbox, the on/off schedule and the
brightness target once per pass -- once per screen. A dwell can be a minute
and a Vegas iteration runs for max_cycle_duration (240s), so on a Pi an
on-demand request posted at 10:54:27 was activated at 10:57:24, and two
brightness saves 12s apart inside one 30s screen never reached the panel.
During Vegas nothing read the mailbox at all: _check_vegas_interrupt only
checked on_demand_active, which only the main-loop read sets.
_service_pending_changes does the main loop's on-demand poll, expiry,
schedule and brightness steps, throttled to PENDING_CHANGES_INTERVAL (the
existing 0.25s mailbox floor), on the display thread. It runs from the Vegas
interrupt checker, the high-FPS and once-a-second render loops (replacing
their direct on-demand poll) and _sleep_with_plugin_updates; between passes
it costs one monotonic compare. A brightness change re-pushes the current
frame, since the panel only shows it from the next push.
Callers act on what it leaves behind: Vegas yields on an on-demand start or
the display being scheduled off (and the main loop then blanks instead of
rendering a screen), the render loops break on a schedule-off as they
already did on a mode change, and the dwell sleep returns early on an
on-demand start/stop or a schedule flip -- so the 60s scheduled-off sleep
now wakes for an on-demand request. The main loop no longer rotates after
a dwell that ended that way, which advanced a new on-demand session past
the mode that was asked for.
A brightness set_brightness() refuses is not retried until the target
changes, so the 4Hz pass doesn't log the same failure (fallback mode)
four times a second.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(display): a screen scheduled off midway stops rendering
Covers the schedule-off break added to the high-FPS and once-a-second
render loops: with the display scheduled off halfway through a 120s screen,
neither loop renders for more than one redraw plus one service interval
past the boundary.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(logos): harden the plugin logo download and share core HTTP headers
download_missing_logo / LogoDownloader.download_logo, the path the
scoreboard plugins use, read response.content with no size cap and wrote
straight to the final path, so a failed or corrupt download could be left
in place and cached as the logo. It now goes through fetch_logo: streamed
with a 10 MB cap, image/* only, decoded by Pillow, converted to RGBA once,
and moved into place atomically. A failure leaves no partial or temp file
and keeps any logo already on disk. LogoHelper._download_logo delegates to
the same code. Public signatures and return values are unchanged; saved
files are pixel-identical to before (RGBA, palette+tRNS, L+tRNS, LA, JPEG).
download_missing_logo reuses one downloader per thread instead of a new
Session per logo. Per thread rather than behind a lock: Session is not
documented thread-safe, and a lock would serialise every plugin's
downloads behind the slowest one.
Placeholders are written atomically, without the test_write.tmp probe.
The logo downloader and background data service now send the real
ChuckBuilds User-Agent from src.common.api_helper (USER_AGENT,
DEFAULT_HTTP_HEADERS) instead of a yourusername/contact@example.com
placeholder, and no longer hand-set Accept-Encoding: br (brotli is not
installed).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(http): drop APIHelper's hand-set brotli encoding; LogoHelper sends the real UA
APIHelper advertised `br` though brotli isn't installed, so a server that
honoured it would send a body requests can't decode. LogoHelper sent a bare
`LEDMatrix-Common/1.0`, the kind of User-Agent ESPN has been rejecting.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
project_root was only assigned in the relative-path branch, so an absolute
plugin_system.plugins_directory made web_interface/app.py raise NameError
at import (first use: the SchemaManager construction). Define it before the
if/else; plugins_dir resolution is unchanged.
Adds a regression test that imports the real module in a fresh interpreter
with an absolute and a relative plugins_directory.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(cache): collapse CacheStrategy's all-60 defaults table and twin soccer branch
get_sport_live_interval() without a config manager looked the sport up in
a table where every value was 60, with 60 as the fallback; it now returns
60. get_data_type_from_key() had an `if 'soccer'` branch returning the
same 'sports_live' as its else.
test_cache_strategy_intervals pins the returned strategy for every data
type x sport key x config-manager shape; it passes unchanged on the old
code. A 2,544-entry dump of every CacheStrategy method over a wider grid
is identical before and after.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(cache): drop CacheStrategy's `<sport>_scoreboard` config lookup
get_sport_live_interval() and get_cache_strategy() read live/recent/
upcoming intervals from config[f"{sport}_scoreboard"]. Those sections
belonged to the built-in scoreboards the plugin system replaced; plugin
config is keyed by plugin id ("football-scoreboard"), so on a current
config the lookup always fell through to the defaults (60 live, 1800
recent, 10800 upcoming), which are now returned directly.
The one input where this differs: a config.json upgraded from the
pre-plugin era that still carries e.g. an "nfl_scoreboard" section (no
code removes them), queried with an explicit sport key. No caller in core
or the plugin monorepo passes a sport key here -- get_with_auto_strategy
only derives one for keys classed sports_live/live_scores, and its callers
(odds managers, odds-ticker) use odds keys -- so the stale section was
unreachable in practice. A dump of every CacheStrategy method over 2,544
inputs differs from the previous commit only in those 45 legacy-config
entries; the test grid now includes that shape.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* perf(cache): list cache files without holding the memory-tier lock
CacheManager.list_cache_files() held the in-memory cache's lock while it
listed and stat'd the whole cache directory -- 8,864 files on a real rig
-- so every get()/set() from the display loop and plugins waited out the
scan. The lock never protected the disk: DiskCache writes and deletes
under their own lock, and a file vanishing between listdir and stat was
already handled (logged and skipped). The body is unchanged apart from
the dedent.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(cache): delegate memory-tier cleanup and stats to MemoryCache
CacheManager._cleanup_memory_cache() was a line-for-line copy of
MemoryCache.cleanup(), and get_memory_cache_stats() a copy of
MemoryCache.get_stats(), both reaching into the component's private
_cache/_timestamps/_lock through "backward compatibility" aliases bound
in __init__. So the component's own cleanup and stats only ever ran in
tests, and the aliases went stale whenever the component was swapped
(test_cache_ttl_honoured does). Both now delegate, and the aliases are
gone: nothing in core, the tests, or the ledmatrix-plugins monorepo reads
them.
Behaviour is the same. Compared line by line, the two cleanups differ
only in the sort key's fallback (0 vs 0.0, which orders identically),
range+bounds check vs slice for the eviction, and the logger name on the
DEBUG summary line (src.cache_manager -> src.cache.memory_cache). A
differential run over 20,000 random memory states (str/None/garbage/
future timestamps, orphan keys, sizes 0-12, forced and throttled runs)
gives identical removed counts, resulting dicts and last-cleanup times;
the same harness catches each of three seeded mutations of
MemoryCache.cleanup. The throttle clock also moves with it:
CacheManager kept its own copy of last-cleanup, the component's is used
now, and they started equal.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(background): inline the sport cache key and drop the unused request queue
get_sport_cache_key() constructed a whole CacheManager -- ConfigManager,
config parse, cache-dir probing with test-file writes -- to return
f"{sport}_{date}". It now builds the key itself in the same format as
CacheManager.generate_sport_cache_key() (UTC date, %Y%m%d); tests check
the two agree for explicit dates and, with a frozen clock at 03:30 UTC,
for the default date. Median per call on Windows: ~0.6 ms -> ~2 us
(alternating runs); on a Pi the old path also wrote a probe file per call.
request_queue was a PriorityQueue nothing ever put into: requests go
straight to the executor, so `priority` never did anything. The queue is
gone; the `priority` parameter and FetchRequest field stay (every
monorepo scoreboard passes priority=) and are documented as ignored, and
get_statistics() keeps reporting queue_size, now a literal 0 as it
always was in practice.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
save_config() opened config.json with 'w' and streamed json.dump into it,
so a power cut or an unencodable value left the file truncated.
save_config_atomic() renamed a temp file into place but never fsynced it,
rewrote the unchanged secrets file on every save, and re-parsed every
backup to rotate them. save_raw_file_content() had its own third copy.
All of them, plus rollback and config creation from the template, now go
through atomic_write_text(): temp file in the same directory, fsync,
final mode set before the rename, rename (retried on Windows while a
reader holds the file), directory fsync. A root save copies the previous
owner onto the new file so a rename by the display service no longer
hands config.json to root; the shared-group fix-up is unchanged. The
mode is chosen from the file name, so a "secrets" directory in the
install path no longer makes config.json 0640.
The secrets file is rewritten only when its content changes, and backup
rotation works from filenames alone. Backups keep their names
(config/backups/config.json.backup.<version>, paired secrets backup) and
the five newest are kept, as before.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
35 methods on CacheManager, DisplayManager, FontManager and PluginManager
have no caller in core, the ledmatrix-plugins monorepo or the registry's
third-party plugins, but plugins live elsewhere, so they stay for one
release. src.deprecation.deprecated logs a warning (and emits a
DeprecationWarning) the first time each is called in a process, naming the
release that removes it. The list and replacements are in CHANGELOG and
PLUGIN_API_REFERENCE's new Deprecated APIs section; a test pins the set.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>