- client: ControlError.sent says whether the display had the request;
should_fall_back() allows a mailbox write only when it did not, or when
the display is too old to know the command (upgrade case)
- on-demand start/stop: a display that had the request and failed it is
answered 503 (400 for invalid_args), no mailbox copy
- errors.clear: new socket command, answered on the connection thread by
a handler the display registers; applied and republished before the
answer; plugin_error_clear_request only on fallback
- display: on-demand mailbox looked at once a second while the socket is
up (0.25 s without), read only when its file changed (one stat via
CacheManager.file_signature / MailboxWatch); socket commands no longer
touch the mailbox; a processed duplicate is consumed; writers logged once
- error publisher: mailbox read only when changed; snapshot carries
applied_clear_cutoff so an older mailbox request is not shown pending
- docs and CHANGELOG (mailboxes kept for one release)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): run each screen through a ScreenRunner (run loop stage 3)
The two frame loops, the make-up dwell and the dynamic-duration exit move
out of DisplayController.run() into src/screen_runner.py. ScreenRunner
paces with an injected FrameClock (production: this module's time, looked
up per call so the golden harness's fake clock still drives it) and
returns one Outcome whose ExitReason is DURATION, CYCLE_COMPLETE, EMPTY,
ERROR, DISPLAY_FALSE, RELOAD or PREEMPTED.
PREEMPTED replaces the re-checks that used to follow each frame loop and
the make-up dwell (current_display_mode != active_mode, the schedule, a
pending WiFi notice): the runner asks its host at named service points
(FRAME, AFTER_LOOP, after_dwell, FINAL), and on PREEMPTED run() goes to
the next pass without advancing the rotation, as each `continue` did.
RELOAD is the one early end that still advances, as a reload always did.
Each service point reads the WiFi notice file exactly when the loop did
(NoticeRead), because the read is throttled and deletes expired files.
The host answers still use the old checks; the following commits move
them to the Arbiter. _screen_preempted is gone (folded into the FRAME
check); _wait_frame_interval returns the preempting plan instead of a
bool. The frame pacing (8 ms deadline, 1 ms minimum yield, 1 Hz wait with
socket wake) is the same code, moved.
Golden traces unchanged. A capture of every harness run (all 67, with
every sleep, display() call, wifi read, live scan, publish and dwell
logged) is identical to origin/main except for throttled WiFi reads that
returned the cached answer (no side effect) after a notice preempted.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): on-demand is an Arbiter Source (run loop stage 3)
Arbiter.decide() now answers for an active on-demand session itself
(Source.ON_DEMAND) instead of returning LEGACY:
- ArbiterState gains the session: its mode list, index, expiry and pin,
plus current_mode, snapshotted from the controller's fields by
_arbiter_state(). The controller's attributes stay the record that the
web UI, the control socket and the cache read.
- The OnDemand plan is the session's current mode (an index past the end
of a shortened list starts again at 0), with what is left of a timed
session at `now` as max_duration and the expiry as deadline. A session
with no modes left is a plan with no mode; the controller ends it and
shows the rotation's mode, as _resolve_active_mode did.
- on_demand_bound() is _clamp_to_on_demand made pure. It is still applied
after the first frame, with the clock read there.
- ArbiterState.next_on_demand() is the step _advance_on_demand takes.
run() asks decide() for the screen at the point it used to call
_resolve_active_mode (after any Vegas iteration, so a session that
started mid-iteration still shows next), and _take_plan() applies it.
Golden traces and the 67-run capture identical to origin/main. Adds
TestOnDemand and the bound table to test_display_arbiter.py; the stage-2
table's on-demand rows now name ON_DEMAND instead of LEGACY.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): live priority is an Arbiter Source (run loop stage 3)
The live-priority step of run() (step 7) and the live checks around the
Vegas iteration become the Arbiter's Live Source:
- ArbiterInputs gains live_modes (the scan, None where run() made none),
vegas_enabled, vegas_live_in_ticker and vegas_yielded. ArbiterState
gains the rotation and its index, the live resume point and the
"takeover not shown yet" flag.
- Live picks the next live mode round-robin (live_pick, now also what
_check_live_priority returns), not advancing past a mid-screen takeover
that has not shown. It outranks Vegas unless the ticker keeps live
content, in which case it has no say at all, as before. With nothing
live, a plan below it carries ends_live and the interrupted rotation
resumes.
- ArbiterState.claim_live/release_live are _apply_live_priority's
bookkeeping made pure; _apply_live_priority applies them.
run() reads the inputs below the WiFi notice where it always did
(_arbiter_inputs_below_wifi: the Vegas check, then the scan), asks
decide() once more, and _take_plan applies the claim or the resume. The
Vegas iteration moves to _run_vegas_iteration, which re-decides with
vegas_yielded after a yield, so a game that stopped the ticker or an
on-demand session that started mid-iteration still shows next. LEGACY
now means Vegas or the rotation.
One redundant call is gone: a Vegas pass scanned the live plugins twice
at the same instant (step 7, then step 8's "is anything live?"); it scans
once. Golden traces unchanged. The 67-run capture is identical to
origin/main once that duplicate scan and _apply_live_priority(None) calls
that changed nothing are left out.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): the rotation is an Arbiter Source; LEGACY means Vegas (run loop stage 3)
decide() now names the screen for every pass: the Rotation Source
(Source.ROTATION) answers with the rotation's current mode, after the
resume when live priority just ended. LEGACY is left meaning only Vegas,
whose iteration is still run()'s own code until stage 4. Once this pass's
iteration has yielded (vegas_yielded), Vegas passes and the screen it fell
through to is decided like any other.
The rotation's mode is state.current_mode rather than
rotation[rotation_index]: they agree except where something moved the
panel off the list and the rotation carries on from there (a live mode no
entry names, or None after a session ended with nothing to resume to),
and run() always showed current_display_mode.
ArbiterState.after(outcome) is _advance_after_screen's step: an on-demand
session moves to its next mode; otherwise the rotation advances unless the
mode just shown is still live. The Outcome carries the two facts only the
controller can see at the end of the screen (on_demand_active, the live
hold from _still_live). Ending a session with no modes left stays in the
controller, because it is not pure.
Golden traces unchanged; the 67-run capture is identical to origin/main
with the same two exclusions as the previous commit. Adds the Vegas /
Rotation table and TestAfter; the stage-2 rows that said LEGACY for "live,
Vegas or rotation" now say ROTATION.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* refactor(display): one decide() call at each of the runner's service points (run loop stage 3)
The three mid-screen checks the frame loops made one after another --
_check_live_takeover, then _screen_preempted with _wifi_notice_pending in
it -- become one call: Arbiter.decide(state, inputs, now, running=plan).
It returns `running` itself while the screen holds, else the plan that
ends it, from these rules in the order the loops checked them:
1. Live: a game went live while a non-live screen runs. First because it
is the one preemption that changes the state (the rotation moves to
the live mode and remembers where it was), and it is still claimed
when a WiFi notice is pending too; the next pass shows the notice,
then the game, as before.
2. The panel's mode moved under the screen (on-demand started, ended or
changed; the rotation was rebuilt).
3. The schedule turned the panel off.
4. A WiFi notice arrived (on-demand outranks it; compared with expiry).
5. A plugin reload is waiting (between frames only).
Each rule is gated by plan.preemptible_by: every screen may be preempted
by the gate, OnDemand, Wifi, Live, Rotation and a reload, except that a
live screen leaves Live out. A follower and Vegas never preempt
mid-screen. The pure helper live_takeover() is the Live rule, shared with
the dwell sleep's _check_live_takeover.
The controller only gathers and applies. _screen_service applies pending
changes and makes the live scan when one is due (_scan_for_takeover: the
same throttle and gates as before); _screen_check reads the WiFi notice
exactly where the loop did (the read is throttled and deletes an expired
file, so an extra read would move both), calls decide() once, and claims
a live takeover. _screen_preempted is gone; _check_live_takeover and
_wifi_notice_pending remain for the dwell sleep and the Vegas yield path,
built on the same rules.
Golden traces unchanged. The 67-run capture is identical to origin/main
(with the earlier two exclusions) except for one event: in the 125 Hz loop
the live scan still runs before the frame's sleep, but the claim is now
made by the service point after it, so the "live" state change is logged
8 ms later (test_live_game_cuts_a_scrolling_screen_short: 8.064 -> 8.072).
The screen still ends at the same frame (8.072) and every frame, sleep
and pass is unchanged.
Adds the mid-screen table (24 rows), live_takeover's table, and
test/test_screen_runner.py (the runner on a scripted host, plus the
controller's service point: which reads it makes at which checkpoint).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs, tests: run loop stage 3 -- doc, changelog, mutation survivors
docs/RUN_LOOP_REDESIGN.md describes run() as it now is (two decide()
calls per pass, the runner and its service points, the state snapshot and
its transitions), records what stage 3 shipped and how it was checked,
and adds one open "may be wrong" behaviour the mutation run surfaced: a
Vegas iteration stopped for a sync follower falls through to a full
rotation screen before the follower gets the panel (pinned by
test_vegas_yielding_to_a_follower_shows_a_rotation_screen_first; passes on
origin/main too). docs/IPC_CONTROL_SOCKET.md no longer names
_screen_preempted. CHANGELOG entry under Unreleased.
A mutation run broke 46 moved or new pieces once each (OnDemand, Live,
Vegas/Rotation, after(), each mid-screen rule, the runner's pacing, exits
and service points, the controller's gathering and claims). Three
survived and get a test here:
- the after-loop service point not reading the WiFi notice: the
completed-loop checkpoint gets its own name, and a run-loop test has a
notice pending when a later frame comes back empty;
- the Vegas yield path not marking vegas_yielded: the follower test above;
- _take_plan not writing back a reset on-demand index: a controller test
with an index past a shortened list.
All 46 now fail at least one test.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
GET /plugins/installed called get_registry_info() per plugin. Despite the
"no network call" comment, a cold or expired cache made that download
plugins.json (10 s timeout, three attempts), and with no cached copy each
plugin's lookup repeated it -- offline, every load waited out the timeouts.
The route now reads the registry copy already in memory, however old, via
get_cached_registry_info(). A missing or expired copy starts a single
background refresh (backing off after an offline failure), so a later load
gets update and verified badges. Store, install and update paths still fetch.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
fetch-stats now reports wire_bytes (urllib3's raw socket byte count) beside the decoded bytes in every counter set. ESPN gzips its scoreboards, so the decoded count overstated real traffic ~10-14x: ledpi measured football at 33.2 MB/h decoded vs 3.15 MB/h on the wire.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
While a plugin's update() runs it holds the plugin's lock and its frames are skipped -- on a scroller, a frozen strip -- with nothing logged. The high-FPS loop now times each run of skipped frames (report_hold=True); one of 250 ms or more logs 'Display of X held N ms by its update()' (rate-limited per plugin) and is recorded as a 'display hold' busy skip, which never touches the circuit breaker. The 1 Hz loop is left out: one skipped frame there measures the loop interval on a screen that did not visibly freeze (seen on ledpi as ~1000 ms reports on clock-simple and switch-mode football).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- fetch_espn_date_chunks() asks for a window's partial edge month whole when the window covers ESPN_MONTH_COVER_MIN_DAYS (7) or more of its days, trimmed to the window by US Eastern start date. New espn_request_chunks().
- Chunk requests share one process-wide cap of ESPN_CHUNK_WORKERS (6) in flight.
- A new process starts as if a range had just been rejected, so it no longer spends a doomed 400 per window at start.
- Also: _eastern_zone() without try/except/pass (Codacy), and test_on_demand_live_and_restore reads the last on-demand state write rather than the last cache write (the font-usage publisher raced it; main CI had failed on it since #748).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A screen that draws its card once and holds it no longer leaves the web preview black: DisplayManager remembers a changed frame the snapshot throttle skipped, and the render loop writes it (write_owed_snapshot(), called from _display_once) once the interval has passed. A failed owed write stays owed and is retried.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On-demand: a mode requested by name is shown first (even a quiet live mode); a session that can't resume after a restart, or whose plugin system failed to start, ends with status restore-failed instead of staying dead.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): pass plugin action params to the wrapper on stdin
POST /api/v3/plugins/action runs a plugin's script through a generated
Python wrapper, and the params went into that wrapper's source as
`params = <json.dumps(params)>`. JSON true, false and null are undefined
names in Python, so any params holding one made the wrapper die with a
NameError before the script ran, and the route answered "Action failed".
The plugin file manager's category toggle sends {"category_name": ...,
"enabled": true}, so of-the-day's category toggle failed every time.
The wrapper now reads the params from its own stdin (json.loads) and the
route passes them there; nothing taken from the request is written into
the generated source any more. The script's side is unchanged: the same
json.dumps(params) on its stdin, LEDMATRIX_ROOT set, stdout parsed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a refused on-demand start leaves no request in the mailbox
POST /api/v3/display/on-demand/start delivered the request (control
socket, else the file mailbox) before it checked the display service.
With the service stopped the socket is absent, so the request went to the
mailbox; the route then answered 400 "Display service is not running"
when start_service was off, or 500 "Failed to start display service" when
the start failed. The display reads that mailbox with max_age=3600 and
never checks a request's timestamp, so the next time it was started it
ran the refused request, pinned if asked.
The service is now checked before anything is delivered, and nothing is
posted when start_service is off and the service is down. When the start
itself fails, the request is withdrawn from the mailbox, but only while
the mailbox still holds this request_id (the compare-before-delete the
display's _consume_on_demand_request uses), so a newer request posted in
the meantime is left for the display.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a pending plugin operation's status no longer answers 500
PluginOperationQueue.enqueue_operation stores the operation's callback in
operation.parameters['_callback'], and the worker pops it only when it
runs the operation. PluginOperation.to_dict() returned parameters as they
were, so GET /api/v3/plugins/operation/<id> for an operation still
waiting in the queue (an install queued behind another plugin's) handed
jsonify a function and answered 500 "A system error occurred" on every
poll until the worker reached it.
to_dict() now leaves out parameters whose name starts with "_". The
operation itself keeps its callback for the worker; every other field of
the answer, and the operation-history records (a different class), are
unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a second install or uninstall of a busy plugin is a 409
PluginOperationQueue.enqueue_operation raises ValueError when the plugin
already has an operation waiting or running. /plugins/install did not
catch it, so a double-clicked Install (the button is never disabled)
answered 500 "An error occurred; see logs for details" from the
blueprint's catch-all while the first install carried on.
/plugins/uninstall caught it in its own catch-all: a 500 "Failed to
uninstall plugin", plus an "uninstall failed" operation-history record
for an uninstall that never started.
Both routes now enqueue through _enqueue_or_conflict, which turns the
queue's refusal into a 409 PLUGIN_OPERATION_CONFLICT naming the plugin,
and records nothing.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): serve binary plugin static files instead of a 500
GET /api/v3/plugins/<plugin_id>/static/<path> read every file with
open(..., 'r', encoding='utf-8') and returned the decoded text, so any
binary file -- a plugin icon or preview image, which is what the REST API
reference says the route is for -- raised UnicodeDecodeError and answered
500.
The file is now sent with send_file, as bytes. HTML, JavaScript, CSS and
JSON keep the content types the route always set, and other text keeps
text/plain; anything else gets the type mimetypes knows it by (image/png
for a .png). The plugin id and path validation and the resolve_under
containment check are untouched.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a socket-acknowledged on-demand start is a success
cdaeb385 checked the systemd unit before delivering the on-demand
request, so a display run by hand or in the emulator (no active unit)
with start_service off now got nothing, where before the request went
over the control socket and took effect behind a 400. A socket
acknowledgement is the display itself saying it is running and has the
request queued, so it is the better witness than systemd.
The request is delivered first again. When the display acknowledged it
over the socket, the route answers success without consulting systemd for
the "not running" 400 and without starting the unit (with start_service
on it tried to start a second display beside the one that answered); the
service is still reported the way _ensure_display_service_running reports
a running one. When it went to the mailbox, the 400 (service down,
start_service off) and the failed-start 500 both withdraw this request_id
from the mailbox, leaving a newer request alone, so neither refusal runs
later.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): mask the Config Editor's secrets like GET /config/secrets
The Config Editor tab (/partials/raw-json) filled its config_secrets.json
editor with the file as it is on disk. GET /api/v3/config/secrets masks every
value because the interface is reachable without a login by default, but
this page handed the same credentials (GitHub token, Home Assistant token,
plugin API keys) to anyone who loaded it. The masked-save path in
save_raw_secrets_config was written for a masked editor and never got one.
_load_raw_json_partial now masks the section with mask_all_secret_values
after strip_auth_section, exactly as the GET does. Saving it back is safe:
save_raw_secrets_config drops the masks (strip_masked_values) and merges the
rest onto the stored file (deep_merge), so an untouched secret stays as it
is and a replaced mask is the only value that changes.
The config.json editor is left as it is. Its save (save_raw_main_config)
writes the posted object verbatim, with no mask stripping or merge, so a
masked main editor would write the bullets over any credential it holds.
Masking it needs a merge-on-save of its own first.
Tests: TestConfigEditorRoundTrip renders the partial over a real
ConfigManager, checks no real value is in the editor, and posts the editor
back unchanged (the file is identical) and with one mask replaced (only that
value changes).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): keep disabled plugins in the saved rotation order and Vegas exclusions
PluginOrderList draws one row per enabled plugin and, once drawn, rewrites
its hidden inputs (plugin_rotation_order, vegas_plugin_order,
vegas_excluded_plugins) from those rows. A disabled plugin has no row, so
merely opening the Display or Rotation & Durations tab took it out of the
inputs, and the next save of that form stored the lists without it. Exclude
Clock from Vegas, disable it, change the brightness, re-enable it: Clock was
scrolling in Vegas again and had moved to the end of the rotation.
syncInputs now keeps the saved ids that have no row. In the order, each one
keeps its saved slot and the rows fill the other slots in their current
order, with rows not in the saved order last, as before. In the exclusions
they follow the unchecked rows. Only string ids are carried over, once each:
/config/main refuses a list holding anything else, which would block every
later save of the tab.
Tests: test/js/unit/test_plugin_order_list.js runs the shipped widget in a vm
with a fake DOM (draw, reorder, include/exclude, the rotation list, junk ids)
and is in run_all.js and the README. The durations DOM suite now reads only
its own rows' ids from the input, since a rig's saved order can hold others.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a restore reinstalls only the plugins that are missing
POST /backup/restore with reinstall_plugins (the "Reinstall missing plugins"
box) passed every plugin in the backup's plugins.json to
install_plugin(). That replaces an installed copy with a fresh download, so
a restore onto the same device re-downloaded every plugin inside the
request. A plugin installed from its own URL is not in the registry, so its
install returned False, plugins_failed set success to False, and the restore
answered 500 "Restore incomplete ... plugins not reinstalled: <id>" (shown
as "Restore failed") with the plugin still installed and the config
restored.
Each plugin is now looked up first with the store's _existing_install, the
same lookup install_plugin makes to decide a copy exists: the id, or an id
the registry proves is the same plugin (aliases, the plugin_path name), and
never a bare ledmatrix-<id> folder (#686). One that is installed is recorded
in result.skipped as "plugin:<id> (installed)", which the page lists under
Skipped; a missing one is installed as before. The list_installed_plugins
docstring said every listed plugin is reinstalled and now says otherwise.
Tests: TestInstalledPluginsAreNotReinstalled, with a mocked store (installed
skipped, missing installed; an installed plugin the store can't install is
not a failure) and with a real PluginStoreManager (a registry alias and a
third-party install are skipped, a missing plugin installed).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): /config/main answers malformed JSON with a 400
save_main_config read a JSON body with request.get_json(), which raises
Werkzeug's BadRequest for a body that does not parse (or an empty one sent as
application/json). That happened inside the handler's try, so the
catch-all answered 500 CONFIG_SAVE_FAILED with "Check file permissions on
config directory" among its suggested fixes and logged a traceback at
ERROR, for what was the caller's mistake.
It now reads with get_json(silent=True), as save_raw_main_config does, and
answers a sent-but-unparseable body with the same 400
{"status": "error", "message": "Invalid JSON in request body"}. An empty
JSON body falls through to the existing 400 "No data provided". The change
is limited to the lines that read the body.
Tests: TestMalformedBody in test_api_v3_partial_main_save.py (the 400 and its
shape, identical to /config/raw/main's, and nothing saved; the empty body).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a restore that brings back fonts clears the font catalog cache
GET /api/v3/fonts/catalog caches its answer as fonts_catalog for five
minutes. Font upload and delete clear that entry (fonts.py), but
POST /backup/restore copies user fonts into assets/fonts without touching
it, so restored fonts were missing from the Fonts tab and every font picker
until the cache expired.
backup_restore now clears fonts_catalog when the result lists restored fonts
(restore_backup records them as "fonts (<count>)"). A restore that restored
no fonts leaves the cache alone.
Tests: TestFontsCatalogCache in test_api_v3_backup_restore.py.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): drop uninstalled plugins from the carried-over order and exclusions
2b34f254 made the plugin order list keep every saved id that has no row,
so a disabled plugin keeps its rotation slot and Vegas exclusion. That
also kept the ids of plugins that have since been uninstalled: they stayed
in plugin_rotation_order and vegas_excluded_plugins for good, where before
the next save of the tab dropped them.
The widget already fetches /api/v3/plugins/installed, every installed plugin
with its enabled flag, and draws only the enabled ones. It now keeps that
response's full id set and carries over only saved ids that are installed
but have no row (disabled). An id outside the set is dropped, as before.
With no list, nothing is dropped: a failed request draws no rows and leaves
the inputs as saved, and the carry-over keeps everything if the set was
never filled.
Tests: test/js/unit/test_plugin_order_list.js adds a disabled plugin kept
while an uninstalled one is dropped (order and exclusions; fails on
2b34f254), and a failed plugin list leaving both inputs as saved. The
CHANGELOG bullet and the README row say so.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* test(js): register the order-list suite apart from other branches' suites
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(config): load_config hands each caller a private copy
ConfigManager.load_config() returned its cached self.config itself (the
mtime fast path from #410 kept the full path's aliasing). Web handlers
edit what they load and then validate: the plugin form save applies the
posted fields to the loaded section (a shallow .copy(), so nested dicts
were the cache's own), and save_main_config sets its checkboxes before
it checks auto_update_channel. When the save was refused, the edit
stayed in the cache the fast path serves, and the next save of any
other setting wrote it to config.json: the refused value, and a nested
secret typed into the same form (mqtt.password, league.espn_s2,
flightaware.api_key) in plain text, since it never reached
config_secrets.json to be stripped. The form also reloaded showing the
refused values.
load_config() now returns a private copy on both paths, and
save_config/save_config_atomic keep a copy of what they were given, so
nothing a caller edits reaches the cache unless it is saved. Fixing it
here rather than in each handler covers every route that edits before it
validates. No caller relies on editing the cache without saving: every
src/ and web_interface/ caller either reads, or saves the dict it
edited. get_config() still returns the live dict for the display
process's readers.
The copy is a pickle round trip: on a Pi 4 with its real 64 KiB config,
2.0 ms against 6.9 ms for copy.deepcopy (json round trip 3.4 ms). Two
tests asserted the aliasing itself and now assert a copy.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): GET /plugins/config masks secrets and refuses core sections
The route returned the plugin's section as load_config() has it, with
config_secrets.json merged in: API keys and tokens went out in plain
text. #276 masked them here; #330's rewrite of the route dropped it,
while the settings page and GET /config/secrets kept masking. It also
took any plugin_id, so ?plugin_id=web_auth returned the login's
cookie-signing key and password hash, and ?plugin_id=github the Plugin
Store token, which GET /config/main strips and redacts.
The route now refuses what _non_plugin_id_error refuses for reset and
uninstall (core sections, malformed ids) with a 400, and blanks x-secret
fields with mask_secret_fields after the defaults merge, as the page
does. A plugin with no schema has its credential-named fields blanked by
_redact_credentials, as GET /config/main does. Blank rather than the
bullets of GET /config/secrets: the save drops a blank secret as
"unchanged" (remove_empty_secrets) but would store the bullets, so the
response must post back as it came. Tested: GET, then POST the response
unchanged, keeps every stored secret.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): parse a table row's cells against the list's item schema
An array of objects drawn as a table posts each cell as
"cities.0.timezone". _get_schema_property stopped at "cities" (an array,
not an object with properties), so _parse_form_value_with_schema got no
schema for the cell and guessed: a blank optional text cell became None
and a text cell holding digits became an int. Validation refused both,
so every save of the page failed for as long as such a row existed --
geochron's city without a timezone, a countdown named "2027". A secret
cell is always drawn blank, so a plugin with secrets in its rows could
not be saved from the form at all.
The lookup now steps from an index segment into the array's items: to
the item schema itself for "color.2", into its properties for a row
cell. Number, boolean and required cells convert as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a blank secret field saves as "unchanged", required or not
The settings page draws a stored secret blank (mask_secret_fields) and
posts the blank back. _parse_form_value_with_schema turned a blank
optional string into "" -- which the save drops as unchanged
(remove_empty_secrets) -- but a blank required one into None. For a
secret that is required with no default (youtube-stats' api_key) that
None failed validation, so every save of the page was refused until the
key was typed in again.
A blank text secret (x-secret, type string) now parses to "", whatever
its required list says; a list or object secret keeps getting [] or {},
which the save drops the same way. Not _SKIP_FIELD: skipping keeps the
value load_config() merged in, and the save would then write it back to
config_secrets.json -- after a secret change the cached section can
still hold the old one, so that write reverted it. A test covers that
sequence.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): POST /plugins/config refuses core sections and malformed ids
Reset and uninstall check the plugin id with _non_plugin_id_error; the
save did not. {"plugin_id": "display", "config": {...}} found no schema,
so nothing was validated or filtered, and the body was merged into the
core display section along with "enabled": true -- rows: "banana"
included. A plugin_id that was not a string (a list, an object, a number)
reached config.get() or the schema lookup, raised TypeError, and came
back as a 500.
Both the JSON and the form path now call _non_plugin_id_error first and
answer its 400.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(web): a text field keeps "true", "[1, 2]" and "{}" as typed
_parse_form_value_with_schema guessed before it consulted the schema:
"true"/"false" became booleans, and a value starting with "[" or "{"
that parsed as JSON became a list or object, whatever the field's type.
A text setting holding "true", "False", "[1, 2]" or "{}" was then
refused by validation ("Expected type string, got bool"), and the save
with it.
A field whose schema type is string, or string-or-null, now returns the
posted text as it came. Every other type goes through the conversions as
before; numbers in text fields were already left alone.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(config): copy the cached config without pickle
_private_copy was a pickle round trip. It only ever unpickled bytes it had
just made from our own dict, so nothing untrusted reached it, but it put
pickle in the config path and Codacy failed the PR for it (B301/B403).
The config is JSON data, so copying its dicts and lists is a full copy;
every other value is immutable. Measured on ledpi (Pi 4) with its real
60 KiB config: 2.11 ms, against 1.92 ms for pickle and 6.75 ms for
copy.deepcopy.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* docs(changelog): describe the config copy without pickle
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
With scroll_card.date_format "weekday", a Friday 8 PM ET game read
"Sat Oct 2" on the scroll and Vegas cards.
Cause: the extractor prints the "M/D" in the plugin's resolved zone (its
own setting, then the global one, then the system zone), but the card is
handed only the plugin's config. Its timezone ships as "", so
card_tzinfo fell back to UTC and the weekday belonged to the UTC date:
the next day for evening games in the Americas, the previous day for
morning games east of UTC (Auckland, Kiritimati).
Fix: every zone is within a day of UTC, so the printed date is the
start's UTC date or a neighbour of it. _format_date_as now takes the game
and names the weekday of whichever of those days has the printed month
and day, falling back to the zone-based weekday only when the start
cannot place the date (no offset, unparseable, or more than a day away).
The switch-mode scorebug shares the formatter and passes the game too, so
the twins stay identical; it already used the resolved zone and draws
what it drew before. Public signatures are unchanged.
Tests cover US DST end, New Year's Eve, both sides of the date line, NZ
DST start and UTC+14. The twins test's weekday pin is updated: the drawn
date now agrees, and only the bare weekday helpers still differ.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugins): drop a plugin's package modules when it unloads
A plugin that keeps helpers in a package (providers/feed.py, imported as
`from providers.feed import ...`) leaves dotted entries in sys.modules.
PluginLoader only tracked bare names: `providers` was namespaced and
dropped on unload, `providers.feed` stayed. A reload after a store update
imported a fresh `providers`, then got the old `feed` back from the module
cache, so the new manager.py ran against the old helpers until the display
restarted. A load that failed part-way left them behind the same way.
Elections (providers/), flights (enrichment/) and olympics (data/,
renderers/) ship packages.
The loader now records the dotted modules whose file (or, for a namespace
package, every __path__ entry) lies inside the plugin directory. They keep
their names while the plugin runs, as before, and unregister_plugin_modules()
drops them, only while sys.modules still holds that plugin's module. The
failed-load cleanup in load_module() drops them too.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugins): remove a symlinked dev plugin as a link
PluginStoreManager._safe_remove_directory, behind uninstall and behind
discarding the set-aside copy after an install or update, handed a
symlinked dev plugin (scripts/dev/dev_plugin_setup.sh) to shutil.rmtree,
which refuses a symlink. The chmod fallback then walked through the link
and set every directory and file in the linked checkout to 0700, and the
sudo stage refused the resolved path as outside the plugins directory. The
removal failed, the link stayed, and the developer's checkout lost its
group/other permissions. A dangling link read as already removed, because
exists() follows it, and was left behind.
A symlink is now unlinked before any other stage runs, and before the
exists() check.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugins): load a dev plugin linked in under a different name
contained_plugin_dir(), the containment check before a plugin's
dependencies are installed, resolved the plugin directory and looked for
the resolved folder's name among the plugins directory's entries. A dev
plugin symlinked in under its id by a name its checkout does not share --
`dev_plugin_setup.sh link-github foo <url>` clones ledmatrix-foo, the
repository naming convention, and links it as plugins/foo -- has no such
entry, so install_dependencies() returned False and the load failed with
"Dependency installation failed", even with no requirements.txt.
When the path sits directly in the plugins directory, the entry it names
(the link) is looked up first; anything else is resolved and matched by
name as before. The answer is still always rebuilt from a name os.scandir()
returned for the plugins directory, so a path outside it is still refused.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(plugins): release a plugin whose update() raises a BaseException
On the async update worker, the wrapped update() finished its bookkeeping
(_finish: release the plugin lock, drop the pending slot, state back to
ENABLED) only for an Exception. asyncio.CancelledError and SystemExit
derive from BaseException, so one raised from update() skipped _finish:
the plugin kept its lock and stayed RUNNING for the life of the process,
never rescheduled, with every display() skipped as busy. PluginExecutor
caught only Exception as well, so its thread died with the call never
marked complete and an immediate failure was logged and recorded as a
timeout.
_target_update now runs _finish for any BaseException and re-raises it,
and the executor's thread stores it like any other exception, so it is
reported as the operation's failure (PluginError) on both the async and
the synchronous path. _finish and _record_update_failure take a
BaseException.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(config): notify config subscribers outside the service lock
ConfigService._load_config ran every subscriber while holding _lock. The
display's per-plugin subscriber calls PluginManager.apply_config_change,
which waits up to PLUGIN_LOCK_TIMEOUT (5 s) for a plugin busy in update().
A save that enables or disables a plugin also flags a reconcile, which the
render thread runs: its get_config(), and the unsubscribe() of a plugin it
disables, both take _lock, so the panel froze behind every slow callback,
up to 5 s per busy plugin.
The config is now swapped under _lock and the subscribers are called after
it is released, from a copy of the subscriber lists. A separate
_notify_lock is held across a whole reload (read, swap, notify), so one
reload's notifications still finish before the next one's start. Each
callback is checked against the live lists just before it runs, and
unsubscribe() waits only for a call of that same callback already in
progress (unless it is that callback's own thread), so a callback it
removed is not running and will not run once it returns, as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
StateSubscription._run reset its backoff only when _follow() returned
normally, which happens only on stop(). Every real disconnect raises
ControlError, so the wait kept doubling across connections: after
successive display restarts the web resubscribed 1, 2, 4, 8, 16 and then
30 s later for good, answering from one-shot state.get connections in the
meantime. The docs promise "1 s up to 30 s" per outage.
The wait now goes back to the minimum once a connection got as far as
storing a snapshot, whatever ended it. A display without the stream
(unknown_command) is still retried at the slow interval.
The frozen-timestamp bug found in the same review is fixed by #737.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(display): a plugin duration that is not a number no longer stops the display
DisplayController._get_display_duration returned whatever the plugin's
get_display_duration() gave back. clock-simple, calendar and countdown
return their display_duration setting straight from config.json, so a
value saved as "20" or null reached _resolve_durations as a string or
None, and its `<= 0` check raised a TypeError. Nothing in the loop caught
it: run()'s outer handler logged "Unexpected error in display controller"
and cleanup() ended the service when that plugin's screen came up, and
systemd restarted it into the same crash.
The plugin's answer is now read as seconds: a finite number or a numeric
string is used (as BasePlugin.get_display_duration already accepts), a
number at or below zero still goes to _resolve_durations' 15 s rule, and
anything else -- None, a non-numeric string, a bool, NaN, infinity, or a
get_display_duration() that raises -- gets the 30 s a mode without a
plugin gets. The warning is logged once per plugin, not at every screen.
Tests: test/test_display_duration_not_a_number.py, including the real
run() on the run-loop harness, which returned at t=30 before the fix.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(scroll): a strip narrower than the panel no longer raises on every frame
ScrollHelper._get_visible_portion_integer handled a frame that runs off
the end of the strip by copying the strip's tail and then the rest of the
frame from its head, which assumed the head was at least that wide. For a
strip narrower than the panel that raised "could not broadcast input
array" at every position, so get_visible_portion() never returned a frame
and the caller logged a traceback each frame. Vegas composes such a strip
(lead_in_width defaults to 0) when its content is narrower than the chain.
A wrapping frame is now taken column by column modulo the strip's width
(np.take, mode='wrap', into the reused frame buffer): the tail then the
head, as before, and a narrow strip repeated across the panel. The same
path takes a position before the start of the strip, whose [-n:m] slice
was empty and made frombytes raise; the integer and sub-pixel fast paths
now leave a negative start to it. A zero-width strip is still a black
frame.
Tests: test/test_scroll_helper_narrow_strip.py.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(cache): store keys too long to be a filename
The calendar plugin's cache key joins every calendar id the user picked.
On hdpi it passed 300 bytes; ext4 caps a filename at 255, so every write
(the temp file, the direct-write fallback and the home-directory fallback)
failed with ENAMETOOLONG, once an hour, and the final warning said
"(permission denied)" whatever the error was.
DiskCache.get_cache_path keeps a key of up to 200 UTF-8 bytes as its
filename, exactly as before, and turns a longer one into its first 183
bytes (cut on a character boundary) plus a 16-hex-digit hash of the whole
key. The temp file adds 15 bytes, so the longest name is 215. The
shortened stem is itself short, so the web UI's cache list, which names a
key by its filename, deletes the same file. The give-up warning now names
the real error.
Validated on ledpi's ext4: the old module drops the hdpi-shaped key, the
new one writes a 205-byte filename and reads it back.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(cache): judge a memory hit by the record's own timestamp
A record loaded from disk went into the memory tier timed from the load,
so get(key, max_age=300) could return data close to 600 s old: after a
restart, after the memory sweep, or in a second process. A stored ttl was
stretched the same way. #728's _fresh_cached works around it for the
scoreboard; every other caller was exposed.
get_cached_data and load_cache now also check a memory hit against the
record's embedded timestamp, with DiskCache.get's rule that a stored ttl
wins over max_age. A stale copy is dropped and the read falls through to
disk, which returns the other process's newer write if there is one.
Records without a timestamp keep the memory tier's own clock.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
State stream ticks carry the volatile timestamps (display.last_updated, plugins.published_at), so current-status and the plugin runtime stay fresh while one mode stays on screen.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
src/common/__init__.py and src/plugin_system/__init__.py resolve their re-exports lazily (PEP 562 __getattr__, __all__ and __dir__ unchanged, TYPE_CHECKING imports for mypy), and sync_manager imports numpy only where send_frame uses it. The web process no longer loads numpy, freetype helpers and PluginManager just to import path_safety, store_manager or schema_manager (~67 MB to ~54 MB RSS on a Pi 4). from src.common import X and from src.plugin_system import X keep working, including submodule imports.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Six small savings with no behaviour change: the odds fetch no longer pretty-prints every response for a debug line; the scroll integer-slice path drops a redundant full-frame np.ascontiguousarray; ledmatrix-web.service gets MALLOC_ARENA_MAX=2 like the display unit; core ESPN responses are parsed via response_json (orjson when installed); and the scroll frame stats go to INFO only for degraded windows plus a 5-minute heartbeat.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The disk cache's unchanged-payload skip now ignores a CacheManager.set() record's timestamp, so unchanged re-saves are skipped; a skip moves the file's mtime to the new timestamp instead, and readers take a record's age from the newer of the two (never more than an hour past the embedded timestamp). Per-plugin plugin_metrics:<id> records become one plugin_metrics_snapshot written at most once a minute, and CacheManager builds its ConfigManager on first use. On hdpi, cache file writes went from ~37 to 8.6 a minute.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The frame loops and the dwell sleep run PluginManager.run_scheduled_updates() at most every 0.25 s instead of after every frame (the top of each loop pass still ticks unthrottled), and SportsScrollDisplay and the sync follower ask ScrollHelper.has_strip() instead of building cached_image just to see whether a strip exists.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Updates that move HEAD now install changed systemd units through a root-owned helper (/usr/local/sbin/ledmatrix-refresh-units, two literal sudo lines), with a backup restored on rollback; a refresh that fails part-way puts the old units back. Devices without the new sudo rule keep updating and are told to re-run the installer once. The one-shot installer now checks out the newest vX.Y.Z release (LEDMATRIX_CHANNEL=beta keeps main) and never moves an existing checkout backwards.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds state.get / state.subscribe to the display's control socket (StateHub in src/ipc/server.py). The web interface holds one subscription per process (web_interface/display_state.py) and reads current-status, on-demand status, plugin runtime and /health's display_loop from it, falling back to the cache keys and heartbeat file. While the socket serves readers, display_current_state and plugin_runtime_snapshot are written less often (about 1.5 instead of 5 cache writes a minute for 15 s screens).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
LayoutContext.fit_image keyed images without a cache_key by id() and held
a strong reference to the source so the id could not be recycled. A
plugin following the documented one-liner -- draw_image(Image.open(path),
box) each frame -- never hit that cache and kept the last 64 sources
alive: ~64MB for 500x500 RGBA team logos (median size under
assets/sports), up to ~600MB for the largest.
The entry now holds a weak reference whose callback drops it when the
source is freed, and a hit re-checks that the referent is the same
image. Sources that cannot be weak-referenced are still pinned. Keyed
entries (the only kind any plugin on ledmatrix-plugins main uses today:
football-scoreboard's logo fit) are unchanged.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
A collection during interpreter shutdown called GcMonitor after the
module's `time` global was torn down, printing "Exception ignored while
calling GC callback ... 'NoneType' object has no attribute
'perf_counter'" at the end of service and test runs.
- GcMonitor binds its clock and sys.is_finalizing at construction and
does nothing once the interpreter is finalizing.
- install_gc_monitor() unregisters it with atexit; new
uninstall_gc_monitor().
- DisplayManager.cleanup() (reached from SIGTERM via run()'s finally)
unregisters it alongside the frame recorder.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
run() stage 2: a pure Arbiter.decide() (src/display_arbiter.py) chooses scheduled-off, follower and WiFi notices; everything else takes the existing path. Golden traces byte-identical.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fetch service stage 2: one ESPN scoreboard cache key shared across the sports base classes (legacy keys still read), and a max-age response cache in the fetch service.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The runtime status snapshot now agrees with the display heartbeat, and current-status is republished when the display wakes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A plugin.reload no longer freezes the panel during Vegas: the old instance is torn down and the new one loaded off the render thread (frame gap 3017 ms -> 9 ms in the ledpi reproduction). A failed or timed-out teardown stops the reload with a restart hint instead of loading over stale modules or tearing down an instance still in use.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A GcMonitor in src/common/frame_timing.py, installed once per process from gc.callbacks by the display manager (and render_bench), counts collections and seconds per generation, the longest, and those of 20 ms or more. A long one tags the next presented frame 'gc' in record(), so frame_soak shows its late rate under 'after work'; the stats file gains an additive 'gc' block printed as a 'Garbage collection' line; and a Render stall dump says when a long collection ran inside the stall. Diagnostic only: nothing tunes, freezes or disables the collector.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The schedule-off blank and the WiFi notice are drawn by the display controller, not dispatched to a plugin, so #716's handover never reached them. Drawn while the last scroll's state was still set, the blank went out with the ticker's lagging rows on a scan-compensated panel and stayed up for its 60 s dwell, and the notice's redraws (which #712 now shows over a running scroller or Vegas) were timed as 0.5-1 s freezes and logged as a mid-scroll Render stall. The controller now calls set_scrolling_state(False) before drawing either; a scroller that resumes sets the state again on its next frame.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A scoreboard in scroll mode no longer redraws every card at the start of a recent/upcoming turn whose games have not changed (~1.4 s for seven football cards at 192x48 on a Pi 4, with the render thread waiting). SportsScrollDisplayManager keeps one display per slate (game type + leagues; up to 4 per game type, at most 6 MB per plugin of strips not on screen) and rewinds the strip it built last time when its games, rankings, config, panel size and date are unchanged and it is under 10 minutes old. First turns, changed slates, live strips and empty turns are drawn as before; get_scroll_display() still answers with the strip on screen.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Control socket stage 2: the render thread wakes for queued commands (static screens ~1 ms, Vegas within one frame), brightness.set, and plugin.reload after a store update, with mailbox/restart fallbacks. Rig checks listed in the PR body are still to run.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Mid-scroll, update_display() no longer checksums every frame: it asks is_currently_scrolling() once per frame and reuses the answer, hashes only when dirty tracking can skip a static frame, and the preview snapshot asks its policy first and hashes only when a write or touch could follow (decide() is monotone, pinned by a property test). With the preview open the snapshot is written at most once a second (VIEWER_INTERVAL 1.0 s, was 0.2 s; the SSE stream re-read it once a second, so four encodes in five went unread); the stream now polls its mtime every 0.25 s (VIEWER_POLL_INTERVAL), so the preview stays about as fresh. The PNG is written at compress_level=1. --preview soaks are not comparable across this change.
Merged with #716: _scan_segments takes the frame's one scrolling answer and #716's static-handover pass-through, as soaked on ledpi (A B B A, 20 min each: main 0.118% / 0.113% late, with #716 and this 0.107% / 0.104%).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
New draw_text_outlined(draw, xy, text, font, fill, outline_color=(0, 0, 0), offsets=OUTLINE_SQUARE) in src/common/text_helper.py, with OUTLINE_SQUARE and OUTLINE_CROSS. It rasterizes the string once and stamps the mask at each outline offset instead of one draw.text per offset: the same pixels as the nine-draw loop (an equivalence sweep across the bundled fonts, image and font modes, colours and positions pins it, on Windows and Linux), about 8x faster per outlined string. Fractional coordinates, multiline text, fonts other than a plain FreeTypeFont, other image modes and a replaced draw.text take the old loop. SportsCoreSharedMixin._draw_text_with_outline and TextHelper.draw_text_with_outline draw through it; the scoreboards' own game_renderer loops adopt it in a ledmatrix-plugins change after a core release.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A static plugin screen that follows a scroller no longer starts with the ticker's lagging rows on scan-compensated panels, and the 1 Hz loop's second frame is no longer recorded as a ~1 s mid-scroll freeze / Render stall. The display controller calls DisplayManager.end_scroll_for_static_screen() before a static screen's first display() (clears the scan history; _scan_segments passes its frames through in one swap) and set_scrolling_state(False) after it; the scroller's hold stays until then, so late-frame counts are unchanged. A screen's first frame is tagged 'handover': gaps of 250 ms or more before it go to the additive handover_freezes (frame_soak prints 'Handover gaps'), not freezes. The display thread is named display-<plugin id>. The WiFi notice and the schedule-off blank are not covered yet (docs list them as a follow-up).
ledpi A B B A soak (20 min each, --preview): main 0.118% / 0.113% late with 6 / 3 freezes; with this and #717 0.107% / 0.104% late, 0 freezes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Schedule and dim windows are half-open [start, end): on from the start time, off at exactly the end time, whatever second the check runs. An on-demand session that ends in scheduled-off hours (expiry or stop) clears the once-a-minute schedule gate, so the panel blanks within about a second. Golden: schedule; two test_display_pending_changes.py tests now say end_time 23:00.
Merged with #712 and #713: with all three in, docs/RUN_LOOP_REDESIGN.md's 'may be wrong' list is empty, so that section now records that all six items are fixed and by which PR.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A game that goes live takes over within about a second (_check_live_takeover in the frame loops and the dwell sleep, throttled to 1 s, never during on-demand, scheduled-off, live_in_ticker or an already-live screen); an interrupted Vegas iteration switches straight to the game; has_live_content() is asked once per plugin per scan. Goldens: live_priority, vegas.
Merged with #712: after an interrupted Vegas iteration the WiFi-notice check runs before the live switch (WiFi outranks live). Adds test/test_run_loop_wifi_and_live.py, pinning that a notice and a game arriving during the same screen (1 Hz, 125 Hz, Vegas) show the notice first, then the game, and neither while scheduled off.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A WiFi notice preempts the current screen within about a second (frame loops, post-loop check, make-up dwell), and an interrupted Vegas iteration that yielded for a notice ends the pass so the notice shows next. Goldens: wifi_notice, vegas. First of three run-loop fixes (#712, #713, #714), pre-tested together.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Core's own HTTP fetch paths (APIHelper, fetch_espn_scoreboard and its date chunks, BackgroundDataService, BaseOddsManager.get_odds) go through one service in src/common/fetch_service.py: shared connection pools per retry policy, merged identical in-flight GETs, per-host token-bucket budgets (fetch_service.rate_limits), and per-plugin request counters published to GET /api/v3/plugins/fetch-stats. Return values, exceptions, cache keys, TTLs and retry policies are unchanged. Core-internal in this release; plugins should not import it directly yet.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The first-frame dispatch (_dispatch_first_frame) now asks PluginExecutor.execute_display() to re-raise (raise_errors=True) and records a raise inside the executor as a breaker failure, with the original exception as last_error, instead of a success. The screen is still an empty pass and rotation is unchanged; a hung display() is still recorded once, as a hang. The run-loop golden trace plugin_error.json is regenerated (crashy now records health failures and is skipped by the breaker), and behaviour 7 is dropped from docs/RUN_LOOP_REDESIGN.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds golden trace tests for DisplayController.run() (test/test_run_loop_golden.py on a fake clock with fake plugins, 15 scenarios, fixtures in test/fixtures/run_loop_golden/) and moves twelve blocks of run() into named helpers (_dispatch_first_frame, _resolve_durations, _resolve_active_mode, _needs_high_fps, _advance_after_screen and others) with the traces identical before and after. docs/RUN_LOOP_REDESIGN.md describes the target structure. Hardware-checked on hdpi: Vegas late-frame rate unchanged in an ABBA A/B.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scan-order compensation only ran at one frame per refresh, so a crisp scroll
like 60 px/s on a 120 Hz panel (1px every 2 refreshes) showed a half-pixel
step across the middle of the panel. A held frame is now presented as a
sequence of swaps (scan_order.refresh_plan): the lagging half shows the
previous frame for its first refresh and the new one for the rest, so it
steps one refresh after the rest. Skipped when a blit takes over half a
refresh, since the second blit has to land before the next vsync.
Soaked on ledpi (60 px/s, 120 Hz): 0.16% late frames, as before the change.
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* feat(scroll): show which scroll speeds are smooth on this panel
The Vegas Scroll Speed slider now says what the panel will do with the
chosen speed and offers the nearest smooth ones to click. Backed by
scroll_config.speed_advice() and GET /api/v3/config/scroll-speed-advice,
which uses the refresh the display measured rather than the cap.
Also stops the default 50 px/s snapping to a stepped 48 px/s (2px every 5
refreshes, 24fps) on a 120Hz panel: the low-fps penalty in solve_crisp()
now loses to 60 or 40 px/s. 100Hz panels are unchanged.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
* fix(scroll): hint threw before its timer variables existed; count 25-30fps as stepped
The Vegas speed hint called refreshScrollSpeedHint() before the let
declarations it uses, so it never rendered (found on ledpi). And the
solver's low-fps penalty stopped at 25fps, which let a measured 125.7Hz
panel keep a 25.1fps 2px-every-5-refreshes scroll.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
* test: add the scroll-speed-advice route to the /api/v3 URL map snapshot
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* chore(deprecation): remove the 35 APIs deprecated for 3.8.0
The usage scan (docs/DEPRECATIONS_3.8.md, regenerated 2026-10-01 and
committed here) finds no call or override of any of them in the 46
monorepo plugins or the 8 third-party plugins plugins.json lists; the
only core callers were other deprecated methods removed alongside.
- CacheManager: 13 methods, plus the private helpers only
has_data_changed used (_has_*_changed, _is_market_open).
- DisplayManager: 7 methods, plus WEATHER_COLORS and the private
_draw_sun/_cloud/_rain/_snow/_storm helpers only the icon methods used.
- FontManager: 14 methods, plus size_tokens, _save_overrides and
_clear_plugin_font_cache. font_overrides and _load_overrides stay:
resolve_font() still applies config/font_overrides.json.
performance_stats stays: get_font() keeps it and tests read it.
- PluginManager.get_enabled_plugins.
test_deprecation.py pins only the two 3.9.0 markers now; the scanner
tests run against a stand-in core instead of the real markers. The
memory-tier tests read stats through log_memory_cache_stats() and the
component, and the test of the removed _clear_plugin_font_cache goes.
Docs drop the removed methods' reference entries; the Deprecated APIs
table becomes "Removed in 3.8.0". CHANGELOG gains a Removed section.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(deprecation): drop the test harness's copies of the removed icon methods
VisualTestDisplayManager still drew weather icons that DisplayManager no
longer has, so a plugin's visual tests could pass on calls that raise
AttributeError on the real display. Its draw_sun/draw_cloud/draw_rain/
draw_snow/draw_weather_icon/draw_text_with_icons, WEATHER_COLORS and the
private helpers go, with the tests that exercised them. The CHANGELOG's
Deprecations entries no longer say nothing is removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: prepare the 3.8.0 release
Bumps src.__version__ to 3.8.0 and turns Unreleased into ## 3.8.0, with a
summary and a New modules list (vegas_elements, testing.vegas, sports_vegas;
display_watchdog, plugin_catalog, plugin_runtime) for plugins flooring on
3.8.0. Adds the CHANGELOG line #701's second commit lacked (blocks laid out
off the render thread).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore: 3.8.0 also ships what landed on main since the prep
#703, #705, #706 and #693 merged after this branch was cut; their CHANGELOG
entries now sit under 3.8.0. The summary and New modules list name them
(sports consolidation stage 4's four modules, src/ipc, field_model), stage
4's section says to floor on 3.8.0, and src/common/README.md marks its four
modules 3.8.0 instead of Unreleased.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* chore(deprecation): remove the 35 APIs deprecated for 3.8.0
The usage scan (docs/DEPRECATIONS_3.8.md, regenerated 2026-10-01 and
committed here) finds no call or override of any of them in the 46
monorepo plugins or the 8 third-party plugins plugins.json lists; the
only core callers were other deprecated methods removed alongside.
- CacheManager: 13 methods, plus the private helpers only
has_data_changed used (_has_*_changed, _is_market_open).
- DisplayManager: 7 methods, plus WEATHER_COLORS and the private
_draw_sun/_cloud/_rain/_snow/_storm helpers only the icon methods used.
- FontManager: 14 methods, plus size_tokens, _save_overrides and
_clear_plugin_font_cache. font_overrides and _load_overrides stay:
resolve_font() still applies config/font_overrides.json.
performance_stats stays: get_font() keeps it and tests read it.
- PluginManager.get_enabled_plugins.
test_deprecation.py pins only the two 3.9.0 markers now; the scanner
tests run against a stand-in core instead of the real markers. The
memory-tier tests read stats through log_memory_cache_stats() and the
component, and the test of the removed _clear_plugin_font_cache goes.
Docs drop the removed methods' reference entries; the Deprecated APIs
table becomes "Removed in 3.8.0". CHANGELOG gains a Removed section.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* chore(deprecation): drop the test harness's copies of the removed icon methods
VisualTestDisplayManager still drew weather icons that DisplayManager no
longer has, so a plugin's visual tests could pass on calls that raise
AttributeError on the real display. Its draw_sun/draw_cloud/draw_rain/
draw_snow/draw_weather_icon/draw_text_with_icons, WEATHER_COLORS and the
private helpers go, with the tests that exercised them. The CHANGELOG's
Deprecations entries no longer say nothing is removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The display serves a control socket (/run/ledmatrix/control.sock) carrying versioned JSON commands, one per line, each answered. Stage 1 covers on-demand start, stop and status; commands are queued on the socket thread and applied on the render thread through the mailbox's own handler, and the web interface falls back to the file mailbox when the socket is unavailable. Protocol and security model: docs/IPC_CONTROL_SOCKET.md.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"Processing mode", "display() returned False" and "No content to display" repeated what "Switching to mode" already logs on every rotation; they are now DEBUG. On ledpi this cut the display's journal lines by about 30%; the measured SD-write saving is small (within noise).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>