Compare commits

...
Author SHA1 Message Date
ChuckandClaude Opus 5.5 e39dcb18b4 feat(scroll): report a panel that cannot reach its refresh cap, and suggest one it can hold
Scroll speeds are solved against display.hardware.limit_refresh_rate_hz,
which is only a ceiling. A panel that cannot reach it still moves whole
pixels per frame, but every scroll runs slow by the shortfall and the
"smooth" ladder is the cap's, not the panel's. A user rig (Pi 4, 2x128x64,
adafruit-hat-pwm, pwm_bits 9, gpio_slowdown 5) measured 107.6-113.1 Hz under
a 120 Hz cap: 60 px/s ran at 55, and nothing said why.

- scroll_config: refresh_shortfall() (more than 3% under the planned rate),
  holdable_cap() (a multiple of 10, 5% under the measurement, since the
  measurement is the fast end of an uncapped panel's drift), and
  describe_refresh_shortfall().
- FrameTimingRecorder.plan_refresh(): once the measured period has held for
  three trusted windows, a shortfall is logged once as a warning naming the
  cap to use. DisplayManager calls it only for a real panel, not the
  emulator or the fallback canvas. The stats file records
  planned_refresh_hz (additive).
- GET /api/v3/config/refresh-rate, plus a hint under the Display tab's
  Limit Refresh Rate field with a button that fills in the suggested cap.
- _panel_refresh_hz (behind the Vegas slider's advice) ignores a measurement
  written under a different cap, so a changed cap stops being advised from
  the old rate before the display restarts.

Verified on ledpi with a temporary 200 Hz cap: the warning logged about a
minute after the restart ("about 132 Hz ... Set Limit Refresh Rate to
120 Hz"), the endpoint returned the same shortfall, and the Display tab
showed the hint; its button filled in 120. ledpi was restored afterwards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 17:11:09 -04:00
ChuckandClaude Opus 5.5 a74b5a2f0f fix(espn): fetch a window's edge months whole, cap chunk requests per process (#751)
- fetch_espn_date_chunks() asks for a window's partial edge month whole when the window covers ESPN_MONTH_COVER_MIN_DAYS (7) or more of its days, trimmed to the window by US Eastern start date. New espn_request_chunks().
- Chunk requests share one process-wide cap of ESPN_CHUNK_WORKERS (6) in flight.
- A new process starts as if a range had just been rejected, so it no longer spends a doomed 400 per window at start.
- Also: _eastern_zone() without try/except/pass (Codacy), and test_on_demand_live_and_restore reads the last on-demand state write rather than the last cache write (the font-usage publisher raced it; main CI had failed on it since #748).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 15:46:15 -04:00
ChuckandClaude Opus 5.5 57d7df6705 fix(web): schedule saves, restart_required, health and current-status agree with the rig (#750)
- POST /api/v3/config/schedule and /config/dim-schedule accept a disabled per-day schedule with every day off, and keep an off day's times.
- POST /api/v3/config/main answers restart_required only when the save changed a setting the running display does not apply live.
- GET /api/v3/health reports degraded with checks.display_loop.status stopped when the display service is stopped.
- GET /api/v3/display/current-status answers unknown (null fields) after the display stops instead of the cached last state. New web_interface.display_state.display_gone().

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 15:07:19 -04:00
ChuckandClaude Opus 5.5 e09e251553 fix(display): a frame the preview throttle skipped still reaches the snapshot (#752)
A screen that draws its card once and holds it no longer leaves the web preview black: DisplayManager remembers a changed frame the snapshot throttle skipped, and the render loop writes it (write_owed_snapshot(), called from _display_once) once the interval has passed. A failed owed write stays owed and is retried.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 14:41:47 -04:00
ChuckandClaude Opus 5.5 7026eeb156 fix(on-demand): show a named live mode; end a session that cannot resume (#748)
On-demand: a mode requested by name is shown first (even a quiet live mode); a session that can't resume after a restart, or whose plugin system failed to start, ends with status restore-failed instead of staying dead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 13:25:12 -04:00
ChuckandClaude Opus 5.5 2236ff3081 fix(web-ui): MQTT password without TLS, Overview poll that never stopped, brightness slider error, token form left dirty (#745)
* fix(web-ui): let the MQTT bridge form save a password without TLS

PUT /api/v3/integrations/mqtt-bridge/config refuses a stored password
while mqtt_tls is off unless allow_insecure_mqtt is set (the CWE-319
guard in api_v3/misc.py). The Tools tab form neither rendered a control
for that flag nor sent it, so a password-protected broker on a LAN
without TLS could never be saved from the UI, and once such a password
was in bridge_config.json every later save from the form was refused.

The form now shows "Allow without TLS (trusted network)" while "Use
TLS" is unchecked, prefilled from the GET's config.allow_insecure_mqtt,
and mqttBody() sends its state as allow_insecure_mqtt. The box is off
until the user ticks it, so the server's guard still refuses a
cleartext password by default.

Tests: the Tools DOM suite checks the control, its show/hide with the
TLS box, the prefill and the value saved; a Flask test pins that the
GET reports the opt-in (false until saved on).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): stop the Overview reconciliation poll from running forever

The reconciliation banner script in partials/overview.html re-asked
/api/v3/plugins/reconciliation-status every 2 s until the answer said
done, with no limit. The route answers done: false whenever
ledmatrix_reconciliation.json is missing or unreadable, which happens
when _run_startup_reconciliation raises before writing it or when /tmp
is cleaned under a long-running web service (reconciliation runs once
per process). The browser then sent that request every 2 s for as long
as the page stayed open, on every tab, since the poll was never tied to
the Overview being visible.

The poll now gives up after 30 tries (a minute) and runs only while the
Overview is the active, visible tab, registered with LEDVisibility under
its own key like the other partials' pollers. Dismissing the banner
ends it too.

Test: test/js/unit/test_overview_reconciliation_poll.js runs the shipped
script in a vm with fake timers and fetch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): drop the Display tab's lookup of a removed brightness label

The brightness slider's input handler in partials/display.html set the
text of both #brightness-value and #brightness-display. #387
(978a03b42) removed the "LED brightness: N%" line that carried
#brightness-display, so getElementById returned null and every step of
the slider threw "Cannot set properties of null" into the console. The
visible label still updated, because it is written first.

The dead lookup is removed.

Test: test/js/unit/test_display_partial_ids.js checks every literal
getElementById() in the partial's inline scripts against the ids its
markup renders, and runs the shipped script in a vm to move the slider.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): a created API token leaves the General tab's form clean

app.js marks a form data-dirty on any input inside it and removes the
mark only after a successful htmx request; its beforeunload handler
asks "Leave site?" while a visible form is still dirty. The API token
form in partials/general.html posts through window.webLogin.createToken
with fetch, so the mark survived the token being created and a reload
of the page with the General tab open prompted about a change that had
already been saved.

createToken now removes data-dirty after a successful create, next to
the form.reset() it already did. A refused request keeps the mark.

Test: test/js/unit/test_general_web_login_token.js runs the shipped
script in a vm with a fake fetch and DOM.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(js): match <script> tags the way CodeQL's tag-filter rule expects

The three new suites pull the inline scripts out of their partials with
/<script>([\s\S]*?)<\/script>/g. CodeQL flags that shape as a bad HTML
filtering regexp (js/bad-tag-filter: misses upper case and tags with
attributes or whitespace), four high alerts that blocked the PR. These are
our own templates read by tests, not user input, but the stricter pattern
costs nothing: /<script\b[^>]*>(...)<\/script[^>]*>/gi, as
test_html_escaping.js already uses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(js): slice the Display partial's markup around its scripts

CodeQL read the script-stripping replace() as an incomplete HTML sanitizer
(js/incomplete-multi-character-sanitization). The test only reads our own
template, but slicing between the matched blocks gives the same markup
without the pattern.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:30:51 -04:00
ChuckandClaude Opus 5.5 5ad5e9aa59 fix(web): plugin action params, refused on-demand starts, pending-operation 500, double-click 409, binary static files (#744)
* fix(web): pass plugin action params to the wrapper on stdin

POST /api/v3/plugins/action runs a plugin's script through a generated
Python wrapper, and the params went into that wrapper's source as
`params = <json.dumps(params)>`. JSON true, false and null are undefined
names in Python, so any params holding one made the wrapper die with a
NameError before the script ran, and the route answered "Action failed".
The plugin file manager's category toggle sends {"category_name": ...,
"enabled": true}, so of-the-day's category toggle failed every time.

The wrapper now reads the params from its own stdin (json.loads) and the
route passes them there; nothing taken from the request is written into
the generated source any more. The script's side is unchanged: the same
json.dumps(params) on its stdin, LEDMATRIX_ROOT set, stdout parsed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a refused on-demand start leaves no request in the mailbox

POST /api/v3/display/on-demand/start delivered the request (control
socket, else the file mailbox) before it checked the display service.
With the service stopped the socket is absent, so the request went to the
mailbox; the route then answered 400 "Display service is not running"
when start_service was off, or 500 "Failed to start display service" when
the start failed. The display reads that mailbox with max_age=3600 and
never checks a request's timestamp, so the next time it was started it
ran the refused request, pinned if asked.

The service is now checked before anything is delivered, and nothing is
posted when start_service is off and the service is down. When the start
itself fails, the request is withdrawn from the mailbox, but only while
the mailbox still holds this request_id (the compare-before-delete the
display's _consume_on_demand_request uses), so a newer request posted in
the meantime is left for the display.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a pending plugin operation's status no longer answers 500

PluginOperationQueue.enqueue_operation stores the operation's callback in
operation.parameters['_callback'], and the worker pops it only when it
runs the operation. PluginOperation.to_dict() returned parameters as they
were, so GET /api/v3/plugins/operation/<id> for an operation still
waiting in the queue (an install queued behind another plugin's) handed
jsonify a function and answered 500 "A system error occurred" on every
poll until the worker reached it.

to_dict() now leaves out parameters whose name starts with "_". The
operation itself keeps its callback for the worker; every other field of
the answer, and the operation-history records (a different class), are
unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a second install or uninstall of a busy plugin is a 409

PluginOperationQueue.enqueue_operation raises ValueError when the plugin
already has an operation waiting or running. /plugins/install did not
catch it, so a double-clicked Install (the button is never disabled)
answered 500 "An error occurred; see logs for details" from the
blueprint's catch-all while the first install carried on.
/plugins/uninstall caught it in its own catch-all: a 500 "Failed to
uninstall plugin", plus an "uninstall failed" operation-history record
for an uninstall that never started.

Both routes now enqueue through _enqueue_or_conflict, which turns the
queue's refusal into a 409 PLUGIN_OPERATION_CONFLICT naming the plugin,
and records nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): serve binary plugin static files instead of a 500

GET /api/v3/plugins/<plugin_id>/static/<path> read every file with
open(..., 'r', encoding='utf-8') and returned the decoded text, so any
binary file -- a plugin icon or preview image, which is what the REST API
reference says the route is for -- raised UnicodeDecodeError and answered
500.

The file is now sent with send_file, as bytes. HTML, JavaScript, CSS and
JSON keep the content types the route always set, and other text keeps
text/plain; anything else gets the type mimetypes knows it by (image/png
for a .png). The plugin id and path validation and the resolve_under
containment check are untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a socket-acknowledged on-demand start is a success

cdaeb385 checked the systemd unit before delivering the on-demand
request, so a display run by hand or in the emulator (no active unit)
with start_service off now got nothing, where before the request went
over the control socket and took effect behind a 400. A socket
acknowledgement is the display itself saying it is running and has the
request queued, so it is the better witness than systemd.

The request is delivered first again. When the display acknowledged it
over the socket, the route answers success without consulting systemd for
the "not running" 400 and without starting the unit (with start_service
on it tried to start a second display beside the one that answered); the
service is still reported the way _ensure_display_service_running reports
a running one. When it went to the mailbox, the 400 (service down,
start_service off) and the failed-start 500 both withdraw this request_id
from the mailbox, leaving a newer request alone, so neither refusal runs
later.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:30:39 -04:00
ChuckandClaude Opus 5.5 8a0cce1aaf fix(web): mask the Config Editor's secrets; keep disabled plugins' rotation slot and Vegas exclusion; restore only missing plugins (#743)
* fix(web): mask the Config Editor's secrets like GET /config/secrets

The Config Editor tab (/partials/raw-json) filled its config_secrets.json
editor with the file as it is on disk. GET /api/v3/config/secrets masks every
value because the interface is reachable without a login by default, but
this page handed the same credentials (GitHub token, Home Assistant token,
plugin API keys) to anyone who loaded it. The masked-save path in
save_raw_secrets_config was written for a masked editor and never got one.

_load_raw_json_partial now masks the section with mask_all_secret_values
after strip_auth_section, exactly as the GET does. Saving it back is safe:
save_raw_secrets_config drops the masks (strip_masked_values) and merges the
rest onto the stored file (deep_merge), so an untouched secret stays as it
is and a replaced mask is the only value that changes.

The config.json editor is left as it is. Its save (save_raw_main_config)
writes the posted object verbatim, with no mask stripping or merge, so a
masked main editor would write the bullets over any credential it holds.
Masking it needs a merge-on-save of its own first.

Tests: TestConfigEditorRoundTrip renders the partial over a real
ConfigManager, checks no real value is in the editor, and posts the editor
back unchanged (the file is identical) and with one mask replaced (only that
value changes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): keep disabled plugins in the saved rotation order and Vegas exclusions

PluginOrderList draws one row per enabled plugin and, once drawn, rewrites
its hidden inputs (plugin_rotation_order, vegas_plugin_order,
vegas_excluded_plugins) from those rows. A disabled plugin has no row, so
merely opening the Display or Rotation & Durations tab took it out of the
inputs, and the next save of that form stored the lists without it. Exclude
Clock from Vegas, disable it, change the brightness, re-enable it: Clock was
scrolling in Vegas again and had moved to the end of the rotation.

syncInputs now keeps the saved ids that have no row. In the order, each one
keeps its saved slot and the rows fill the other slots in their current
order, with rows not in the saved order last, as before. In the exclusions
they follow the unchecked rows. Only string ids are carried over, once each:
/config/main refuses a list holding anything else, which would block every
later save of the tab.

Tests: test/js/unit/test_plugin_order_list.js runs the shipped widget in a vm
with a fake DOM (draw, reorder, include/exclude, the rotation list, junk ids)
and is in run_all.js and the README. The durations DOM suite now reads only
its own rows' ids from the input, since a rig's saved order can hold others.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a restore reinstalls only the plugins that are missing

POST /backup/restore with reinstall_plugins (the "Reinstall missing plugins"
box) passed every plugin in the backup's plugins.json to
install_plugin(). That replaces an installed copy with a fresh download, so
a restore onto the same device re-downloaded every plugin inside the
request. A plugin installed from its own URL is not in the registry, so its
install returned False, plugins_failed set success to False, and the restore
answered 500 "Restore incomplete ... plugins not reinstalled: <id>" (shown
as "Restore failed") with the plugin still installed and the config
restored.

Each plugin is now looked up first with the store's _existing_install, the
same lookup install_plugin makes to decide a copy exists: the id, or an id
the registry proves is the same plugin (aliases, the plugin_path name), and
never a bare ledmatrix-<id> folder (#686). One that is installed is recorded
in result.skipped as "plugin:<id> (installed)", which the page lists under
Skipped; a missing one is installed as before. The list_installed_plugins
docstring said every listed plugin is reinstalled and now says otherwise.

Tests: TestInstalledPluginsAreNotReinstalled, with a mocked store (installed
skipped, missing installed; an installed plugin the store can't install is
not a failure) and with a real PluginStoreManager (a registry alias and a
third-party install are skipped, a missing plugin installed).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): /config/main answers malformed JSON with a 400

save_main_config read a JSON body with request.get_json(), which raises
Werkzeug's BadRequest for a body that does not parse (or an empty one sent as
application/json). That happened inside the handler's try, so the
catch-all answered 500 CONFIG_SAVE_FAILED with "Check file permissions on
config directory" among its suggested fixes and logged a traceback at
ERROR, for what was the caller's mistake.

It now reads with get_json(silent=True), as save_raw_main_config does, and
answers a sent-but-unparseable body with the same 400
{"status": "error", "message": "Invalid JSON in request body"}. An empty
JSON body falls through to the existing 400 "No data provided". The change
is limited to the lines that read the body.

Tests: TestMalformedBody in test_api_v3_partial_main_save.py (the 400 and its
shape, identical to /config/raw/main's, and nothing saved; the empty body).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a restore that brings back fonts clears the font catalog cache

GET /api/v3/fonts/catalog caches its answer as fonts_catalog for five
minutes. Font upload and delete clear that entry (fonts.py), but
POST /backup/restore copies user fonts into assets/fonts without touching
it, so restored fonts were missing from the Fonts tab and every font picker
until the cache expired.

backup_restore now clears fonts_catalog when the result lists restored fonts
(restore_backup records them as "fonts (<count>)"). A restore that restored
no fonts leaves the cache alone.

Tests: TestFontsCatalogCache in test_api_v3_backup_restore.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): drop uninstalled plugins from the carried-over order and exclusions

2b34f254 made the plugin order list keep every saved id that has no row,
so a disabled plugin keeps its rotation slot and Vegas exclusion. That
also kept the ids of plugins that have since been uninstalled: they stayed
in plugin_rotation_order and vegas_excluded_plugins for good, where before
the next save of the tab dropped them.

The widget already fetches /api/v3/plugins/installed, every installed plugin
with its enabled flag, and draws only the enabled ones. It now keeps that
response's full id set and carries over only saved ids that are installed
but have no row (disabled). An id outside the set is dropped, as before.
With no list, nothing is dropped: a failed request draws no rows and leaves
the inputs as saved, and the carry-over keeps everything if the set was
never filled.

Tests: test/js/unit/test_plugin_order_list.js adds a disabled plugin kept
while an uninstalled one is dropped (order and exclusions; fails on
2b34f254), and a failed plugin list leaving both inputs as saved. The
CHANGELOG bullet and the README row say so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(js): register the order-list suite apart from other branches' suites

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:30:28 -04:00
ChuckandClaude Opus 5.5 0b039c875f fix(web): plugin settings endpoints - a refused save no longer leaks into config.json; GET masks secrets (#742)
* fix(config): load_config hands each caller a private copy

ConfigManager.load_config() returned its cached self.config itself (the
mtime fast path from #410 kept the full path's aliasing). Web handlers
edit what they load and then validate: the plugin form save applies the
posted fields to the loaded section (a shallow .copy(), so nested dicts
were the cache's own), and save_main_config sets its checkboxes before
it checks auto_update_channel. When the save was refused, the edit
stayed in the cache the fast path serves, and the next save of any
other setting wrote it to config.json: the refused value, and a nested
secret typed into the same form (mqtt.password, league.espn_s2,
flightaware.api_key) in plain text, since it never reached
config_secrets.json to be stripped. The form also reloaded showing the
refused values.

load_config() now returns a private copy on both paths, and
save_config/save_config_atomic keep a copy of what they were given, so
nothing a caller edits reaches the cache unless it is saved. Fixing it
here rather than in each handler covers every route that edits before it
validates. No caller relies on editing the cache without saving: every
src/ and web_interface/ caller either reads, or saves the dict it
edited. get_config() still returns the live dict for the display
process's readers.

The copy is a pickle round trip: on a Pi 4 with its real 64 KiB config,
2.0 ms against 6.9 ms for copy.deepcopy (json round trip 3.4 ms). Two
tests asserted the aliasing itself and now assert a copy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): GET /plugins/config masks secrets and refuses core sections

The route returned the plugin's section as load_config() has it, with
config_secrets.json merged in: API keys and tokens went out in plain
text. #276 masked them here; #330's rewrite of the route dropped it,
while the settings page and GET /config/secrets kept masking. It also
took any plugin_id, so ?plugin_id=web_auth returned the login's
cookie-signing key and password hash, and ?plugin_id=github the Plugin
Store token, which GET /config/main strips and redacts.

The route now refuses what _non_plugin_id_error refuses for reset and
uninstall (core sections, malformed ids) with a 400, and blanks x-secret
fields with mask_secret_fields after the defaults merge, as the page
does. A plugin with no schema has its credential-named fields blanked by
_redact_credentials, as GET /config/main does. Blank rather than the
bullets of GET /config/secrets: the save drops a blank secret as
"unchanged" (remove_empty_secrets) but would store the bullets, so the
response must post back as it came. Tested: GET, then POST the response
unchanged, keeps every stored secret.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): parse a table row's cells against the list's item schema

An array of objects drawn as a table posts each cell as
"cities.0.timezone". _get_schema_property stopped at "cities" (an array,
not an object with properties), so _parse_form_value_with_schema got no
schema for the cell and guessed: a blank optional text cell became None
and a text cell holding digits became an int. Validation refused both,
so every save of the page failed for as long as such a row existed --
geochron's city without a timezone, a countdown named "2027". A secret
cell is always drawn blank, so a plugin with secrets in its rows could
not be saved from the form at all.

The lookup now steps from an index segment into the array's items: to
the item schema itself for "color.2", into its properties for a row
cell. Number, boolean and required cells convert as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a blank secret field saves as "unchanged", required or not

The settings page draws a stored secret blank (mask_secret_fields) and
posts the blank back. _parse_form_value_with_schema turned a blank
optional string into "" -- which the save drops as unchanged
(remove_empty_secrets) -- but a blank required one into None. For a
secret that is required with no default (youtube-stats' api_key) that
None failed validation, so every save of the page was refused until the
key was typed in again.

A blank text secret (x-secret, type string) now parses to "", whatever
its required list says; a list or object secret keeps getting [] or {},
which the save drops the same way. Not _SKIP_FIELD: skipping keeps the
value load_config() merged in, and the save would then write it back to
config_secrets.json -- after a secret change the cached section can
still hold the old one, so that write reverted it. A test covers that
sequence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): POST /plugins/config refuses core sections and malformed ids

Reset and uninstall check the plugin id with _non_plugin_id_error; the
save did not. {"plugin_id": "display", "config": {...}} found no schema,
so nothing was validated or filtered, and the body was merged into the
core display section along with "enabled": true -- rows: "banana"
included. A plugin_id that was not a string (a list, an object, a number)
reached config.get() or the schema lookup, raised TypeError, and came
back as a 500.

Both the JSON and the form path now call _non_plugin_id_error first and
answer its 400.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web): a text field keeps "true", "[1, 2]" and "{}" as typed

_parse_form_value_with_schema guessed before it consulted the schema:
"true"/"false" became booleans, and a value starting with "[" or "{"
that parsed as JSON became a list or object, whatever the field's type.
A text setting holding "true", "False", "[1, 2]" or "{}" was then
refused by validation ("Expected type string, got bool"), and the save
with it.

A field whose schema type is string, or string-or-null, now returns the
posted text as it came. Every other type goes through the conversions as
before; numbers in text fields were already left alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(config): copy the cached config without pickle

_private_copy was a pickle round trip. It only ever unpickled bytes it had
just made from our own dict, so nothing untrusted reached it, but it put
pickle in the config path and Codacy failed the PR for it (B301/B403).
The config is JSON data, so copying its dicts and lists is a full copy;
every other value is immutable. Measured on ledpi (Pi 4) with its real
60 KiB config: 2.11 ms, against 1.92 ms for pickle and 6.75 ms for
copy.deepcopy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): describe the config copy without pickle

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:30:16 -04:00
ChuckandClaude Opus 5.5 07abd87d5e fix(sports): scroll and Vegas cards name the printed date's own weekday (#747)
With scroll_card.date_format "weekday", a Friday 8 PM ET game read
"Sat Oct 2" on the scroll and Vegas cards.

Cause: the extractor prints the "M/D" in the plugin's resolved zone (its
own setting, then the global one, then the system zone), but the card is
handed only the plugin's config. Its timezone ships as "", so
card_tzinfo fell back to UTC and the weekday belonged to the UTC date:
the next day for evening games in the Americas, the previous day for
morning games east of UTC (Auckland, Kiritimati).

Fix: every zone is within a day of UTC, so the printed date is the
start's UTC date or a neighbour of it. _format_date_as now takes the game
and names the weekday of whichever of those days has the printed month
and day, falling back to the zone-based weekday only when the start
cannot place the date (no offset, unparseable, or more than a day away).
The switch-mode scorebug shares the formatter and passes the game too, so
the twins stay identical; it already used the resolved zone and draws
what it drew before. Public signatures are unchanged.

Tests cover US DST end, New Year's Eve, both sides of the date line, NZ
DST start and UTC+14. The twins test's weekday pin is updated: the drawn
date now agrees, and only the bare weekday helpers still differ.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:19:05 -04:00
ChuckandClaude Opus 5.5 e32d177cbd fix(web-ui): Plugin Manager - enable aliased installs, Update All, on-demand modes, long installs, categories, GitHub-URL install (#746)
* fix(web-ui): Update All sends the live installed list and redraws the grid

updateAll() preferred PluginStateManager.installedPlugins over
window.installedPlugins. Only updateAll's own end-of-run refresh ever
fills PluginStateManager, so from the second run on it sent the first
run's plugins: one uninstalled since failed with "plugin not found" and
one installed since was never updated. That refresh also only replaced
window.installedPlugins, so the installed cards and the Updates badge
kept offering "Update to vX" for what had just been updated.

Read window.installedPlugins, the list plugins_manager.js republishes
after every install, uninstall and refresh, keeping PluginStateManager
as the fallback for a page without it, and refresh through
pluginManager.loadInstalledPlugins(true), which redraws the grid.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): list each plugin's display modes in /plugins/installed

The on-demand modal fills its Display Mode select from
plugin.display_modes, but /plugins/installed never sent the field. Every
plugin offered one option, its own id, under "This plugin exposes a
single display mode"; the display resolved that id to the plugin's first
mode, so a multi-mode plugin could only be started, or pinned, there.

Add display_modes to each entry, read from the plugin catalog
(get_plugin_display_modes), the same declared list /display/modes and
on-demand/start use, keeping only strings. Single-mode plugins still get
one option and the same hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): enable a store install by its installed id, and not on reinstall

The store's Install button enabled the new plugin by the registry id it
installed. Weather, Music, Stocks and Leaderboard install under the id
their manifests declare (weather -> ledmatrix-weather); the plugin list,
the config section and /plugins/toggle know only that id, so the toggle
answered 404 "Plugin not found" and the plugin stayed disabled behind
"installed, but enabling it failed". The same button on an installed
plugin (Reinstall) enabled it too, switching a plugin the user had
turned off back on.

POST /plugins/install now names the installed plugin: plugin_id in the
direct answer and in the queued operation's result, read from the
installed manifest found the way the store's update and uninstall find
it (_find_plugin_path: id, aliases, plugin_path name), else the
requested id. The client reloads the list, then enables that id; from
an answer without it, the installed entry the store entry matches
(findInstalledStorePlugin, which isStorePluginInstalled now uses). A
reinstall, decided by the same match that labelled the button, reloads
the list and leaves the enabled state alone.

test/js/plugins_manager_sandbox.js runs the whole of
plugins_manager.js in a vm context against a fake DOM and API, for
suites that drive its real flows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): wait for long store installs; on timeout reload, not fail

pollOperationStatus gave a queued install 60 polls, a second apart,
then reported "Install operation timed out" as an error and stopped.
The server allows the plugin's dependency install 300 s on its own
(install_requirements_file in store_install.py), after a download that
fetches the plugin a file at a time, so installs that went on to
succeed were reported as failed, never enabled, and left out of the
installed list until the page was reloaded.

Give installs INSTALL_POLL_MAX_ATTEMPTS (600, ten minutes). When even
that runs out, reload the installed list and the store badges and warn
that the install may still be running; nothing is enabled without the
operation's answer. Uninstall keeps the default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): build the store's category filter from the store's plugins

The #plugin-category select listed seven fixed categories while the
registry uses about twenty (productivity, utility, transit, finance,
...), so roughly a third of the store could not be filtered to, and
"Financial" missed the plugin filed under "finance".

The template now ships only "All Categories"; syncStoreCategoryOptions,
run by applyStoreFiltersAndSort, adds one option per category the cached
store plugins have (case folded, as the filter compares), keeps the
current choice, and rebuilds only when the set changes or the partial
was swapped in afresh -- the way the Starlark section builds its own.

The test sandbox gains window.addEventListener (initPluginsPage needs
it) and quiets the script's "element not found" warnings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(web-ui): one handler for the GitHub-URL Install button

#install-plugin-from-url had an inline onclick calling
window.handleGitHubPluginInstall, and attachInstallButtonHandler also
gave it a click listener that installs, so both ran on every click
(and on Enter, which clicks it). The inline handler threw a
ReferenceError -- it called isGithubUrl, which is local to the
plugin-manager IIFE, from outside it -- so only the listener's request
went out; correcting that scope alone would have sent every install
twice.

Remove the inline onclick and the window.handleGitHubPluginInstall it
called, which nothing else uses. The listener, which already sent the
only request, is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:54 -04:00
ChuckandClaude Opus 5.5 d18e4d3c9d fix(plugins): sub-package reload, symlinked dev plugins, BaseException in update(), config callbacks outside the lock (#741)
* fix(plugins): drop a plugin's package modules when it unloads

A plugin that keeps helpers in a package (providers/feed.py, imported as
`from providers.feed import ...`) leaves dotted entries in sys.modules.
PluginLoader only tracked bare names: `providers` was namespaced and
dropped on unload, `providers.feed` stayed. A reload after a store update
imported a fresh `providers`, then got the old `feed` back from the module
cache, so the new manager.py ran against the old helpers until the display
restarted. A load that failed part-way left them behind the same way.
Elections (providers/), flights (enrichment/) and olympics (data/,
renderers/) ship packages.

The loader now records the dotted modules whose file (or, for a namespace
package, every __path__ entry) lies inside the plugin directory. They keep
their names while the plugin runs, as before, and unregister_plugin_modules()
drops them, only while sys.modules still holds that plugin's module. The
failed-load cleanup in load_module() drops them too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): remove a symlinked dev plugin as a link

PluginStoreManager._safe_remove_directory, behind uninstall and behind
discarding the set-aside copy after an install or update, handed a
symlinked dev plugin (scripts/dev/dev_plugin_setup.sh) to shutil.rmtree,
which refuses a symlink. The chmod fallback then walked through the link
and set every directory and file in the linked checkout to 0700, and the
sudo stage refused the resolved path as outside the plugins directory. The
removal failed, the link stayed, and the developer's checkout lost its
group/other permissions. A dangling link read as already removed, because
exists() follows it, and was left behind.

A symlink is now unlinked before any other stage runs, and before the
exists() check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): load a dev plugin linked in under a different name

contained_plugin_dir(), the containment check before a plugin's
dependencies are installed, resolved the plugin directory and looked for
the resolved folder's name among the plugins directory's entries. A dev
plugin symlinked in under its id by a name its checkout does not share --
`dev_plugin_setup.sh link-github foo <url>` clones ledmatrix-foo, the
repository naming convention, and links it as plugins/foo -- has no such
entry, so install_dependencies() returned False and the load failed with
"Dependency installation failed", even with no requirements.txt.

When the path sits directly in the plugins directory, the entry it names
(the link) is looked up first; anything else is resolved and matched by
name as before. The answer is still always rebuilt from a name os.scandir()
returned for the plugins directory, so a path outside it is still refused.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(plugins): release a plugin whose update() raises a BaseException

On the async update worker, the wrapped update() finished its bookkeeping
(_finish: release the plugin lock, drop the pending slot, state back to
ENABLED) only for an Exception. asyncio.CancelledError and SystemExit
derive from BaseException, so one raised from update() skipped _finish:
the plugin kept its lock and stayed RUNNING for the life of the process,
never rescheduled, with every display() skipped as busy. PluginExecutor
caught only Exception as well, so its thread died with the call never
marked complete and an immediate failure was logged and recorded as a
timeout.

_target_update now runs _finish for any BaseException and re-raises it,
and the executor's thread stores it like any other exception, so it is
reported as the operation's failure (PluginError) on both the async and
the synchronous path. _finish and _record_update_failure take a
BaseException.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(config): notify config subscribers outside the service lock

ConfigService._load_config ran every subscriber while holding _lock. The
display's per-plugin subscriber calls PluginManager.apply_config_change,
which waits up to PLUGIN_LOCK_TIMEOUT (5 s) for a plugin busy in update().
A save that enables or disables a plugin also flags a reconcile, which the
render thread runs: its get_config(), and the unsubscribe() of a plugin it
disables, both take _lock, so the panel froze behind every slow callback,
up to 5 s per busy plugin.

The config is now swapped under _lock and the subscribers are called after
it is released, from a copy of the subscriber lists. A separate
_notify_lock is held across a whole reload (read, swap, notify), so one
reload's notifications still finish before the next one's start. Each
callback is checked against the live lists just before it runs, and
unsubscribe() waits only for a call of that same callback already in
progress (unless it is that callback's own thread), so a callback it
removed is not running and will not run once it returns, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:42 -04:00
ChuckandClaude Opus 5.5 04f0d8d134 fix(ipc): reset the state subscription's reconnect wait after a good connection (#740)
StateSubscription._run reset its backoff only when _follow() returned
normally, which happens only on stop(). Every real disconnect raises
ControlError, so the wait kept doubling across connections: after
successive display restarts the web resubscribed 1, 2, 4, 8, 16 and then
30 s later for good, answering from one-shot state.get connections in the
meantime. The docs promise "1 s up to 30 s" per outage.

The wait now goes back to the minimum once a connection got as far as
storing a snapshot, whatever ended it. A display without the stream
(unknown_command) is still retried at the slow interval.

The frozen-timestamp bug found in the same review is fixed by #737.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:31 -04:00
ChuckandClaude Opus 5.5 064b9c9912 fix(display): a non-numeric plugin duration no longer stops the display; narrow scroll strips no longer raise (#739)
* fix(display): a plugin duration that is not a number no longer stops the display

DisplayController._get_display_duration returned whatever the plugin's
get_display_duration() gave back. clock-simple, calendar and countdown
return their display_duration setting straight from config.json, so a
value saved as "20" or null reached _resolve_durations as a string or
None, and its `<= 0` check raised a TypeError. Nothing in the loop caught
it: run()'s outer handler logged "Unexpected error in display controller"
and cleanup() ended the service when that plugin's screen came up, and
systemd restarted it into the same crash.

The plugin's answer is now read as seconds: a finite number or a numeric
string is used (as BasePlugin.get_display_duration already accepts), a
number at or below zero still goes to _resolve_durations' 15 s rule, and
anything else -- None, a non-numeric string, a bool, NaN, infinity, or a
get_display_duration() that raises -- gets the 30 s a mode without a
plugin gets. The warning is logged once per plugin, not at every screen.

Tests: test/test_display_duration_not_a_number.py, including the real
run() on the run-loop harness, which returned at t=30 before the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(scroll): a strip narrower than the panel no longer raises on every frame

ScrollHelper._get_visible_portion_integer handled a frame that runs off
the end of the strip by copying the strip's tail and then the rest of the
frame from its head, which assumed the head was at least that wide. For a
strip narrower than the panel that raised "could not broadcast input
array" at every position, so get_visible_portion() never returned a frame
and the caller logged a traceback each frame. Vegas composes such a strip
(lead_in_width defaults to 0) when its content is narrower than the chain.

A wrapping frame is now taken column by column modulo the strip's width
(np.take, mode='wrap', into the reused frame buffer): the tail then the
head, as before, and a narrow strip repeated across the panel. The same
path takes a position before the start of the strip, whose [-n:m] slice
was empty and made frombytes raise; the integer and sub-pixel fast paths
now leave a negative start to it. A zero-width strip is still a black
frame.

Tests: test/test_scroll_helper_narrow_strip.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:19 -04:00
ChuckandClaude Opus 5.5 a6e9e3ef1c fix(cache): cache keys too long to be a filename; memory hits judged by the record's own age (#738)
* fix(cache): store keys too long to be a filename

The calendar plugin's cache key joins every calendar id the user picked.
On hdpi it passed 300 bytes; ext4 caps a filename at 255, so every write
(the temp file, the direct-write fallback and the home-directory fallback)
failed with ENAMETOOLONG, once an hour, and the final warning said
"(permission denied)" whatever the error was.

DiskCache.get_cache_path keeps a key of up to 200 UTF-8 bytes as its
filename, exactly as before, and turns a longer one into its first 183
bytes (cut on a character boundary) plus a 16-hex-digit hash of the whole
key. The temp file adds 15 bytes, so the longest name is 215. The
shortened stem is itself short, so the web UI's cache list, which names a
key by its filename, deletes the same file. The give-up warning now names
the real error.

Validated on ledpi's ext4: the old module drops the hdpi-shaped key, the
new one writes a 205-byte filename and reads it back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cache): judge a memory hit by the record's own timestamp

A record loaded from disk went into the memory tier timed from the load,
so get(key, max_age=300) could return data close to 600 s old: after a
restart, after the memory sweep, or in a second process. A stored ttl was
stretched the same way. #728's _fresh_cached works around it for the
scoreboard; every other caller was exposed.

get_cached_data and load_cache now also check a memory hit against the
record's embedded timestamp, with DiskCache.get's rule that a stored ttl
wins over max_age. A stale copy is dropped and the read falls through to
disk, which returns the other process's newer write if there is one.
Records without a timestamp keep the memory tier's own clock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:18:00 -04:00
ChuckandClaude Opus 5.5 41192b9588 fix(ipc): ticks carry the volatile timestamps, so current-status stays known (#737)
State stream ticks carry the volatile timestamps (display.last_updated, plugins.published_at), so current-status and the plugin runtime stay fresh while one mode stays on screen.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 22:00:07 -04:00
ChuckandClaude Opus 5.5 ef69201770 perf: import package re-exports on first use (#724)
src/common/__init__.py and src/plugin_system/__init__.py resolve their re-exports lazily (PEP 562 __getattr__, __all__ and __dir__ unchanged, TYPE_CHECKING imports for mypy), and sync_manager imports numpy only where send_frame uses it. The web process no longer loads numpy, freetype helpers and PluginManager just to import path_safety, store_manager or schema_manager (~67 MB to ~54 MB RSS on a Pi 4). from src.common import X and from src.plugin_system import X keep working, including submodule imports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:38:27 -04:00
ChuckandClaude Opus 5.5 ed0753c1ca perf: cheap per-frame and per-fetch savings (#725)
Six small savings with no behaviour change: the odds fetch no longer pretty-prints every response for a debug line; the scroll integer-slice path drops a redundant full-frame np.ascontiguousarray; ledmatrix-web.service gets MALLOC_ARENA_MAX=2 like the display unit; core ESPN responses are parsed via response_json (orjson when installed); and the scroll frame stats go to INFO only for degraded windows plus a 5-minute heartbeat.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:26:03 -04:00
ChuckandClaude Opus 5.5 7bb85c0356 perf(cache): skip rewriting unchanged CacheManager.set() records (#730)
The disk cache's unchanged-payload skip now ignores a CacheManager.set() record's timestamp, so unchanged re-saves are skipped; a skip moves the file's mtime to the new timestamp instead, and readers take a record's age from the newer of the two (never more than an hour past the embedded timestamp). Per-plugin plugin_metrics:<id> records become one plugin_metrics_snapshot written at most once a minute, and CacheManager builds its ConfigManager on first use. On hdpi, cache file writes went from ~37 to 8.6 a minute.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:14:08 -04:00
ChuckandClaude Opus 5.5 0d179fdf12 perf(display): throttle the per-frame update tick; check strips without building them (#731)
The frame loops and the dwell sleep run PluginManager.run_scheduled_updates() at most every 0.25 s instead of after every frame (the top of each loop pass still ticks unthrottled), and SportsScrollDisplay and the sync follower ask ScrollHelper.has_strip() instead of building cached_image just to see whether a strip exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 14:01:23 -04:00
ChuckandClaude Opus 5.5 ee789775b4 perf(install): build rpi-rgb-led-matrix with a faster SetImage (#736)
first_time_install.sh applies patches/rpi-rgb-led-matrix/0001-bulk-setimage.patch just before building the Python binding and reverts it afterwards (and from the EXIT trap), so the submodule stays at its pinned commit. The patch copies each image row with one bulk FrameCanvas::SetPixels call and writes the bit planes branch-free: on a 512x64 Pi 4 the frame copy went from 6.57 ms to 2.21 ms. A patch that no longer applies is reported and skipped. Existing installs get it on a rebuild (RPI_RGB_FORCE_REBUILD=1).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:48:27 -04:00
ChuckandClaude Opus 5.5 05deb1ee7d feat(install): support Raspberry Pi OS Bookworm (Python 3.11) alongside Trixie (3.13) (#689)
The installer and scripts/check_system_compatibility.sh share one set of OS rules (scripts/install/lib_os.sh): Bookworm (Debian 12, Python 3.11) and Trixie (Debian 13, Python 3.13) are supported, python3 older than 3.11 stops the install before anything changes, and dhcpcd gets a warning with directions. setcap targets /usr/bin/python3, the apt fallback honours the requirement floors, and the desktop check no longer misreads under pipefail. CI runs the unit and plugin-safety suites on 3.11 and 3.13 (tooling jobs on 3.13); mypy targets 3.11.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:35:13 -04:00
ChuckandClaude Opus 5.5 f841fa36b6 feat(install): updates refresh systemd units; new installs run the newest release (#729)
Updates that move HEAD now install changed systemd units through a root-owned helper (/usr/local/sbin/ledmatrix-refresh-units, two literal sudo lines), with a backup restored on rollback; a refresh that fails part-way puts the old units back. Devices without the new sudo rule keep updating and are told to re-run the installer once. The one-shot installer now checks out the newest vX.Y.Z release (LEDMATRIX_CHANNEL=beta keeps main) and never moves an existing checkout backwards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 13:09:53 -04:00
ChuckandClaude Opus 5.5 515248b34e feat(ipc): control socket stage 3 - a state stream replaces polled cache keys (#735)
Adds state.get / state.subscribe to the display's control socket (StateHub in src/ipc/server.py). The web interface holds one subscription per process (web_interface/display_state.py) and reads current-status, on-demand status, plugin runtime and /health's display_loop from it, falling back to the cache keys and heartbeat file. While the socket serves readers, display_current_state and plugin_runtime_snapshot are written less often (about 1.5 instead of 5 cache writes a minute for 15 s screens).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 12:56:47 -04:00
ChuckandClaude Opus 5.5 6bf7c3fa51 perf(layout): id-keyed fit_image cache entries no longer pin the source (#732)
LayoutContext.fit_image keyed images without a cache_key by id() and held
a strong reference to the source so the id could not be recycled. A
plugin following the documented one-liner -- draw_image(Image.open(path),
box) each frame -- never hit that cache and kept the last 64 sources
alive: ~64MB for 500x500 RGBA team logos (median size under
assets/sports), up to ~600MB for the largest.

The entry now holds a weak reference whose callback drops it when the
source is freed, and a hit re-checks that the referent is the same
image. Sources that cannot be weak-referenced are still pinned. Keyed
entries (the only kind any plugin on ledmatrix-plugins main uses today:
football-scoreboard's logo fit) are unchanged.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 12:11:27 -04:00
ChuckandClaude Opus 5.5 76ad71d1c5 fix(frame-timing): keep the GC monitor quiet at interpreter shutdown (#734)
A collection during interpreter shutdown called GcMonitor after the
module's `time` global was torn down, printing "Exception ignored while
calling GC callback ... 'NoneType' object has no attribute
'perf_counter'" at the end of service and test runs.

- GcMonitor binds its clock and sys.is_finalizing at construction and
  does nothing once the interpreter is finalizing.
- install_gc_monitor() unregisters it with atexit; new
  uninstall_gc_monitor().
- DisplayManager.cleanup() (reached from SIGTERM via run()'s finally)
  unregisters it alongside the frame recorder.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 11:51:06 -04:00
170 changed files with 14652 additions and 803 deletions
+3
View File
@@ -10,3 +10,6 @@
# Generated by scripts/build_css.py; collapsed in diffs, not hand-edited. # Generated by scripts/build_css.py; collapsed in diffs, not hand-edited.
web_interface/static/v3/tailwind.css linguist-generated=true web_interface/static/v3/tailwind.css linguist-generated=true
web_interface/static/v3/plugin-frame.css linguist-generated=true web_interface/static/v3/plugin-frame.css linguist-generated=true
# Installed as an executable (its shebang runs it) by install_service.sh.
scripts/install/ledmatrix_refresh_units.py text eol=lf
+1 -1
View File
@@ -31,7 +31,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
# No dependencies: the script reads src/__init__.py and CHANGELOG.md only. # No dependencies: the script reads src/__init__.py and CHANGELOG.md only.
- name: Assert the tag, CHANGELOG and src.__version__ agree - name: Assert the tag, CHANGELOG and src.__version__ agree
+19 -8
View File
@@ -14,8 +14,14 @@ permissions:
jobs: jobs:
plugin-safety: plugin-safety:
name: Plugin safety harness + unit tests name: Plugin safety harness + unit tests (Python ${{ matrix.python-version }})
runs-on: ubuntu-latest runs-on: ubuntu-latest
# The two Pythons the installer supports: Raspberry Pi OS Bookworm ships
# 3.11 and Trixie 3.13.
strategy:
fail-fast: false
matrix:
python-version: ["3.11", "3.13"]
env: env:
# The bundled fixture plugin gives the harness at least one real plugin # The bundled fixture plugin gives the harness at least one real plugin
# to render, and REQUIRE_PLUGINS turns "discovered zero plugins" into a # to render, and REQUIRE_PLUGINS turns "discovered zero plugins" into a
@@ -29,7 +35,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: ${{ matrix.python-version }}
cache: pip cache: pip
- name: Install dependencies - name: Install dependencies
@@ -43,8 +49,13 @@ jobs:
pytest --no-cov test/plugins/ pytest --no-cov test/plugins/
unit-tests: unit-tests:
name: Core unit tests name: Core unit tests (Python ${{ matrix.python-version }})
runs-on: ubuntu-latest runs-on: ubuntu-latest
# Bookworm's Python (3.11) and Trixie's (3.13); see plugin-safety.
strategy:
fail-fast: false
matrix:
python-version: ["3.11", "3.13"]
steps: steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2 - uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with: with:
@@ -52,7 +63,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: ${{ matrix.python-version }}
cache: pip cache: pip
- name: Install dependencies - name: Install dependencies
@@ -84,7 +95,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
cache: pip cache: pip
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0 - uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0
@@ -123,7 +134,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
# Downloads the pinned standalone Tailwind CLI (SHA-256 checked; no # Downloads the pinned standalone Tailwind CLI (SHA-256 checked; no
# Node), rebuilds static/v3/tailwind.css and plugin-frame.css from the # Node), rebuilds static/v3/tailwind.css and plugin-frame.css from the
@@ -142,7 +153,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
cache: pip cache: pip
# The runtime requirements are installed so mypy sees the real types of # The runtime requirements are installed so mypy sees the real types of
@@ -181,7 +192,7 @@ jobs:
- uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0 - uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
with: with:
python-version: "3.12" python-version: "3.13"
# Stdlib only; exits 0 whatever it finds. # Stdlib only; exits 0 whatever it finds.
- name: Report method-family drift across the nine scoreboards - name: Report method-family drift across the nine scoreboards
+670
View File
@@ -19,6 +19,288 @@ accepts both, but the store flags the old spelling as deprecated
## Unreleased ## Unreleased
### Scroll speed: a panel slower than its refresh cap is reported
- Scroll speeds are solved against `limit_refresh_rate_hz`, so a panel that
cannot reach its cap ran every scroll slow by the shortfall, with no sign
why (one Pi 4 on a 120 Hz cap refreshed at ~110 Hz: 60 px/s ran at 55).
Once the display has measured the real rate over three windows of
scrolling, a panel more than 3% short of the cap is logged once, as a
warning from `src.common.frame_timing` that names a cap it can hold (a
multiple of 10, 5% under the measurement). The Display tab shows the same
under Limit Refresh Rate, with a button that fills it in, from the new
`GET /api/v3/config/refresh-rate`. Not checked in the emulator or on the
fallback canvas.
- The frame-stats file records `planned_refresh_hz` (additive), and the
scroll-speed advice behind the Vegas slider ignores a measurement written
under a different cap. Until now, after the cap changed, the slider kept
advising from the old rate until the display restarted.
- New in `src.common.scroll_config`: `refresh_shortfall()`, `holdable_cap()`
and `describe_refresh_shortfall()`.
### Fixed
- The web preview and `/api/v3/display/current` no longer stay black for a
whole screen that draws its card once and then holds it. The snapshot is
written from `update_display()` at most once per write interval, so a frame
pushed inside that interval was skipped and left for the next
`update_display()` -- which such a screen never makes. Soccer's
recent/upcoming cards skip redundant redraws, and the first one after an
on-demand start lands a few milliseconds after the start's clear wrote a
black frame: on ledpi the preview showed 0 lit pixels for the whole 15 s
while the panel showed the card. `DisplayManager` now remembers a skipped
changed frame, and the render loop writes it (`write_owed_snapshot()`)
once the interval has passed. The cadence is unchanged, and nothing extra
runs when no frame is owed.
### ESPN date-range fetches: fewer requests, fewer at once
A soccer board (8 leagues, ESPN rejecting `dates=` ranges) logged ~90
`NameResolutionError` lines and an `update() timed out` at every start on a
Pi: each league's fortnight-either-side window was 29 day requests, fetched
by several managers at once, ~40 in flight. Measured against live ESPN with
soccer-scoreboard 2.39.2, alternating runs: **~450 requests per start, peak
~45 in flight, ~75 DNS lookups -> 46 requests, peak 13, ~30 lookups**.
- `fetch_espn_date_chunks()` asks for a window's partial edge month whole
when the window covers `ESPN_MONTH_COVER_MIN_DAYS` (7) or more of its days,
and trims the answer to the window's days by each event's US Eastern start
date -- the day ESPN's `dates=YYYYMMDD` means (417 of 417 live soccer
events matched). A 29-day window spanning two months is 2 requests instead
of 29. Short windows (a live poll's 1-2 days) stay day by day. A trimmed
month that comes back at the 500-event cap re-asks only the window's days.
An event with no readable date is kept. New: `espn_request_chunks()`.
- Chunk requests share one process-wide cap of `ESPN_CHUNK_WORKERS` (6) in
flight, across every window being fetched, instead of six per window.
- A new process starts as if a range had just been rejected, so it no longer
spends one doomed 400 per window at every start (eleven at once from a
soccer board); the range is still retried `RANGE_RETRY_SECONDS` in.
### Cheap per-frame and per-fetch savings
- `BaseOddsManager.get_odds()` no longer pretty-prints every odds response
for a debug line: the `json.dumps(..., indent=2)` calls in the fetch path
and `_extract_espn_data` are guarded with `isEnabledFor(DEBUG)`, and the
other debug f-strings there take %-style arguments. Same messages at DEBUG.
- `ScrollHelper`'s integer frame path (`_get_visible_portion_integer`) takes
`tobytes()` straight from the strip's column slice instead of copying it
with `np.ascontiguousarray()` first; the bytes are identical (a test pins
them). At 512x64 on a Pi 4 the bytes step went from ~45 us to ~21 us a
frame.
- `systemd/ledmatrix-web.service` sets `MALLOC_ARENA_MAX=2`, as
`ledmatrix.service` has since #476. Existing installs pick it up when
`scripts/install/install_service.sh` or `install_web_service.sh` is re-run;
until then the startup drift check reports the web unit as changed.
- `APIHelper.get()`/`post()`, `BaseOddsManager.get_odds()`, the two
`LogoDownloader` team fetches and `DynamicTeamResolver`'s rankings fetch
parse with `src.common.json_body.response_json` (orjson when installed),
like `background_data_service` already did. A body orjson rejects falls
back to `response.json()`, so a bad body raises the same
`requests.exceptions.JSONDecodeError` these call sites already catch.
- The `Scroll frame stats` line is logged at INFO only for a degraded window
(fps under 0.9 of the rate the window was locked to, or more than 1% of
frames stalled), the window after one, and a 5-minute heartbeat per
scroller, as the `Vegas FPS` line already was; every window is still logged
at DEBUG. `docs/SCROLL_PERFORMANCE.md` says how to see them all.
### Fewer SD-card writes from the cache
- **An unchanged `CacheManager.set()` no longer rewrites the file.**
`DiskCache` already skipped a payload identical to the last one it wrote,
but `set()` stamps every record with the current time, so for `set()` the
payload never matched and every unchanged re-save was a full rewrite. The
comparison now leaves out a header-first record's timestamp (the `ttl` and
the data still count), and the newer timestamp is kept in the file's mtime
instead: a skipped save touches the file to the record's timestamp, and a
real write pins mtime to the record's own timestamp. Every reader ages a
record from the newer of the two -- `DiskCache.get`, its header-only
staleness check, and the record it returns, whose `timestamp` is the newer
value, so `CacheManager.get`, the memory tier and plugins reading
`record['timestamp']` all agree; the retention sweep and the web UI's cache
list already used mtime. The mtime is trusted at most an hour past the
record's own timestamp, and unchanged data is rewritten once an hour, so a
file copied without its mtime reads at most an hour fresher than its
contents. 100 identical `set()` calls of a 32 KB record: 100 writes before,
1 after.
- **Plugin metrics are one record, written at most once a minute.** The
resource monitor wrote a `plugin_metrics:<id>` record per plugin, each at
most every 30 s: two writes a minute per plugin, 28 on a fourteen-plugin
rig. Every plugin's metrics now go in one `plugin_metrics_snapshot` record
(`{"schema": 1, "plugins": {id: record}}`, each record shaped as before),
written at most once a minute. `GET /api/v3/plugins/metrics` and
`/plugins/metrics/<id>` return the same fields; the numbers can be up to a
minute old instead of 30 s. A plugin the snapshot does not have yet is
still read from its old `plugin_metrics:<id>` record, which nothing writes
any more and the cache's retention removes. Each write starts from the
snapshot on disk, so plugins the display has not run since a restart keep
their numbers, and a reset from the web UI sticks for a plugin the display
is not running, as it did. A plugin with no call for 30 days is dropped from
the snapshot, as its record used to age out.
- **`CacheManager` no longer loads the config when it is built.** Every
manager built a `ConfigManager` and loaded the whole config for a cache
strategy that stopped reading it. `cache_manager.config_manager` is still
there -- the sports plugins resolve the global timezone through it -- and
is now built and loaded on first access; assigning it still replaces it.
`CacheStrategy` is given no config manager (it reads none).
### Plugin update tick: a few times a second, not every frame
- The frame loops and the dwell sleep ran
`PluginManager.run_scheduled_updates()` after every frame, about 125 times
a second on a scroller. Each pass copies the plugin dict and takes several
locks per plugin, almost always to find nothing due: about 100 us with 20
plugins on a Pi 4, 1.2% of the render thread. They now call
`DisplayController._tick_plugin_updates_if_due()`, which runs the pass at
most every `PLUGIN_UPDATE_TICK_INTERVAL` (0.25 s), so a 4 s scroll runs 16
passes instead of 500. No update interval is shorter than 5 s
(`MIN_DYNAMIC_UPDATE_INTERVAL`), and the 1 Hz frame loop already ticked
once a second, so an update starts at most a quarter second later.
- The top of each loop pass still runs it unthrottled, so a plugin just
loaded, reloaded or enabled for on-demand is updated at once. Vegas's own
update thread (`_tick_plugin_updates_for_vegas`) is unchanged.
`test/test_plugin_update_tick_throttle.py` covers both, on the real
`run()` through the golden-trace harness.
### Strip checks no longer build the PIL image
- `SportsScrollDisplay.display_scroll_frame` (every frame) and
`has_cached_content`, and the sync follower's per-frame check of the Vegas
strip, asked whether there was a strip by reading
`ScrollHelper.cached_image`. After the helper deferred the image (an
append, trim or patch), that read built it from `cached_array` and kept
it: 3.5 ms and about 1 MiB more held for a 4288x64 strip, on top of the
array's 0.8 MiB. They now ask `has_strip()`, which gives the same answer
from the helper's bookkeeping. A scoreboard whose `scroll_helper` has no
`has_strip` (its own helper, a test double) is still asked
`cached_image`.
- A strip built with `create_scrolling_image` or `set_scrolling_image`
still keeps both the image and the array, as before.
### Faster frame copy into the panel (library patch, applied at build time)
- Copying each frame into the panel buffer (`SetImage`) was the biggest CPU
cost LEDMatrix owns on large panels: 6-7.5 ms per frame on a 512x64 Pi 4 at
~85 fps, about 60% of a core. The library's binding walked the image column
by column and set one pixel at a time, and each pixel rewrote a word in
every PWM bit plane, 2KB apart, so nearly every write missed the cache.
`patches/rpi-rgb-led-matrix/0001-bulk-setimage.patch` copies row by row in
one bulk call per row, with the colour lookup done once and branch-free
bit-plane writes. The panel buffer is byte-identical to before (882 checks
across image types, offsets, PWM bits, brightness, inverse colours and a
pixel mapper).
- Measured on hdpi (Pi 4, 4x128x64): frame copy 6.57 -> 2.21 ms, the display
process 139% -> 103% of a core, late frames 7.8 -> 5.4 per 1,000.
- `first_time_install.sh` applies the patch to `rpi-rgb-led-matrix-master`
just before building the binding and takes it back out straight after (and
on any exit), so the submodule stays at its pinned commit with no local
changes. A patch that no longer applies after a submodule bump is reported
and skipped; the unpatched library still builds.
- Existing installs keep the library they have until it is rebuilt:
`sudo RPI_RGB_FORCE_REBUILD=1 ./first_time_install.sh`.
`scripts/build_rgbmatrix_nogil.sh` builds from an unpatched copy and is
unchanged.
### Install
- Raspberry Pi OS **Bookworm** (Debian 12, Python 3.11) is supported,
alongside **Trixie** (Debian 13, Python 3.13). The installer used to stop
on anything but Trixie. Which releases and Pythons are accepted now lives
in one place, `scripts/install/lib_os.sh`, which `first_time_install.sh`
and `scripts/check_system_compatibility.sh` both read, so the two can no
longer disagree (the compatibility check called Bookworm an error, and
still accepted Python 3.10, which the rgbmatrix bindings refuse). An
unsupported system gets plain directions to the right image; a `python3`
older than 3.11 stops the install before anything changes.
- The installer says up front when the Pi runs dhcpcd instead of
NetworkManager, and how to switch back: the web page's WiFi tab and the
`LEDMatrix-Setup` hotspot need NetworkManager. Not fatal, and it does not
switch the network stack itself, since that can cut the SSH session.
- The desktop check no longer misses a desktop install: `dpkg -l | grep -q`
under `pipefail` read a match as "not found".
- `cap_sys_nice` is set on the interpreter the services run
(`/usr/bin/python3`); it preferred `/usr/bin/python3.13` whenever it
existed.
- The Step 7 dependency fallback (`scripts/install_dependencies_apt.py`) no
longer accepts apt packages older than the pins -- Bookworm's Flask 2.2.2
and Pillow 9.4, Trixie's Flask 3.1.1 and Pillow 11.1. The floors are read
from `web_interface/requirements.txt`, and pip is asked for `Pillow`, not
`PIL`.
- CI runs the unit and plugin-safety suites on Python 3.11 and 3.13 (was
3.12); mypy targets 3.11.
### Updates refresh the systemd units; new installs run the newest release
- **Updates now install changed systemd units.** An update (Update Code, or
the weekly automatic update) moved the checkout's `systemd/*.service`
templates but never the units systemd runs, so settings added after a
device was installed -- #687's render-loop watchdog, for one -- only ever
arrived with a reinstall. After an update that moves HEAD, the web
interface compares the installed `ledmatrix.service`,
`ledmatrix-web.service` and `ledmatrix-update-verify.{service,path}` with
the new templates (rendered exactly as `install_service.sh` does, comments
ignored as the startup drift warning does) and, when they differ, runs the
new root-owned helper `/usr/local/sbin/ledmatrix-refresh-units`
(`scripts/install/ledmatrix_refresh_units.py`) through sudo: it installs
the changed units and runs `systemctl daemon-reload`, so the restart that
follows the update runs under them. Update Code's message says so.
- **Rollback restores them.** The helper keeps the units it replaced
(`/var/lib/ledmatrix/unit-backup`, root only); when the automatic update's
health check rolls an update back, it runs `ledmatrix-refresh-units
--restore` before restarting the services onto the old code.
- **The sudo rule needs a reinstall.** `install_service.sh` installs the
helper and `lib_sudoers.sh` grants it with exactly two command lines (no
arguments, and `--restore`). A device installed before this has neither;
its updates keep working, log that the new unit settings need a reinstall
and say so in Update Code's message, the same remedy as the startup
"unit drift" warning. Re-run `sudo ./first_time_install.sh` once (or
`sudo ./scripts/install/install_service.sh` then
`./scripts/install/configure_web_sudo.sh`).
- `install_service.sh` now leaves the units it installs mode `0644`, as
`first_time_install.sh` already did; run on its own it left them `0600`.
- **New installs run the newest release.** The one-shot installer cloned
`main`'s tip, so a new device ran unreleased code until the next release.
It now checks out the newest `vX.Y.Z` tag after cloning (the same semver
rules as `web_interface/update_channel.py`), and that release's own
`first_time_install.sh` runs. `LEDMATRIX_CHANNEL=beta` installs `main`
instead and records the beta channel; `first_time_install.sh --beta` (or
`LEDMATRIX_CHANNEL=beta|stable`) records a channel for a manual install.
- **Re-running the one-shot never moves backwards.** On an existing stable
checkout it moves to the newest release only when that release contains
the current commit; a checkout newer than every release keeps its
fast-forward pull (on a branch) or stays put (detached), and beta keeps the
pull it always had. It used to fast-forward a detached release checkout to
`main`'s tip.
### Control socket stage 3: the display's state over the socket
- Two new commands, still protocol version 1. `state.get` returns a
versioned snapshot of what the display is doing: the current mode and
plugin, the on-demand session, the brightness, the plugin runtime
snapshot and the render loop's heartbeat age. With `since`/`epoch` it
returns a short "unchanged" answer. `state.subscribe` returns the same
snapshot, then pushes a `state` event on every change (always the latest
version) and a `tick` at least every 5 s. The display serves all of it
from memory (`StateHub` in `src/ipc/server.py`), and publishing never
waits for a reader. Subscribers have their own bound (4), separate from
the 8 request slots, and one that stops reading is dropped after the 2 s
IO timeout. See `docs/IPC_CONTROL_SOCKET.md`, "The state stream".
- The web interface holds one subscription per process
(`web_interface/display_state.py`). `/display/current-status`,
`/display/on-demand/status`, the plugin runtime fields of
`/plugins/installed` and `/plugins/state`, the reconciliations and
`/health`'s `display_loop` read it first. When the socket is missing (a
stopped or older display, Windows), they fall back to the cache keys and
the heartbeat file. Each answer has a `source` (`socket`, `cache` or
`heartbeat_file`). The stale and stalled rules from #726 apply the same
way to both.
- Fewer SD-card writes while the socket serves those readers.
`display_current_state` is written once a minute and on a flag change,
not on every mode change. The `plugin_runtime_snapshot` refresh goes from
60 s to 120 s. For a rotation of 15 s screens, that is 1.5 cache writes a
minute instead of 5. Both keys keep being written for one release.
- `RenderWatchdog.liveness()` reports the heartbeat age from memory.
`PluginRuntimeView` has a `source`, and `describe()` includes it.
### Web UI: four more tabs are ES-module pages (stage 2) ### Web UI: four more tabs are ES-module pages (stage 2)
- Rotation, Operation History, Config Editor and Backup & Restore follow the - Rotation, Operation History, Config Editor and Backup & Restore follow the
@@ -56,6 +338,27 @@ accepts both, but the store flags the old spelling as deprecated
`scripts/render_bench.py` records the same. Diagnostic only: nothing tunes, `scripts/render_bench.py` records the same. Diagnostic only: nothing tunes,
freezes or disables the collector. freezes or disables the collector.
### Web interface: lighter package imports
- `src.common` and `src.plugin_system` now import their re-exported names on
first use (PEP 562 module `__getattr__`) instead of in `__init__.py`.
`from src.common import ScrollHelper`, `src.plugin_system.PluginManager`,
`from src.common import *` and every submodule import work as before and
return the same objects. What changes is that importing a submodule --
the web interface's `src.common.path_safety`, `src.plugin_system.store_manager`
and the like -- no longer loads `ScrollHelper`, `LogoHelper`, `APIHelper`,
the adaptive layout helpers and `PluginManager` with it. `sync_manager`
imports numpy inside `send_frame`, the one place it uses it, since the API
blueprint imports that module only for its constants.
- The web process no longer loads numpy at all. On a Pi 4 (Python 3.13),
importing `web_interface.app` went from ~67 MB to ~54 MB RSS and from
~2.8 s to ~1.3 s (`-X importtime`, median of five). A bare
`import src.common` went from ~50 MB / ~0.85 s to ~10 MB / ~30 ms. The
display process loads the same modules as before, only later.
- A misspelt name in `from src.common import ...` still raises `ImportError`.
`test/test_lazy_package_imports.py` checks that the packages import nothing
heavy and that every name in `__all__` resolves to its home module's object.
### Outlined text: one rasterization ### Outlined text: one rasterization
- New `draw_text_outlined(draw, xy, text, font, fill, outline_color=(0, 0, - New `draw_text_outlined(draw, xy, text, font, fill, outline_color=(0, 0,
@@ -242,6 +545,91 @@ policies are unchanged.
### Fixes ### Fixes
- A cache key too long to be a filename is now cached. The calendar
plugin's key joins every calendar id the user picked; on a real install
it passed 300 bytes, ext4 refuses names over 255, and every write failed
with `File name too long` — logged as "(permission denied)", so it read
like a cache-directory ownership problem. `DiskCache.get_cache_path` now
keeps a key of up to 200 UTF-8 bytes as its filename, as before, and
turns a longer one into its first bytes plus a hash of the whole key. The
web UI's cache list and delete keep working, because the shortened name
maps back to the same file. A failed write now names the real error.
- The cache's memory tier no longer serves data older than the reader asked
for. A record loaded from disk was timed in memory from the load, not
from when it was written, so `get(key, max_age=300)` could return data
close to 600 s old (after a restart, after the hourly memory sweep, or in
the other process, which only ever loads the record from disk), and a
stored `ttl` was stretched the same way. A memory hit is now also checked against
the record's own timestamp, and a stale one falls through to disk, which
returns a newer write if there is one.
- An on-demand request that names a `*_live` mode now shows that mode. On
ledpi, `{"plugin_id": "football-scoreboard", "mode": "ncaa_fb_live"}` with
15 college games on answered 200 and showed `nfl_recent`. The session's
mode list kept a live mode only when the plugin's `has_live_content()`
said so. That method answers the live-priority question, and the sports
plugins answer it for favourite teams only. A mode the request names
(not one resolved from a bare plugin id) now leads the session, with the
plugin's other modes after it. If it has nothing to draw, the session
moves on to the next of those modes, like any empty on-demand mode. The
name is saved with the session (`named_mode` in
`display_on_demand_config`), so a restart resumes on it.
- A restart during an on-demand session whose plugin then fails to load no
longer leaves a session with no modes. On ledpi, `clock-simple` failed
config validation after a crash. The display logged `No valid display
modes found for on-demand plugin 'clock-simple' after restoration` and
kept reporting the session as active until its first pass ended it as
`idle`. The cached request stayed behind for the next restart. The session
now ends at startup with status `error` and error `restore-failed`, which
`/display/on-demand/status` reports, and the cached request is dropped. The
same applies when the plugin system itself fails to start.
- `POST /api/v3/config/schedule` and `/config/dim-schedule` accept a
disabled per-day schedule with every day off. That is the shape
`config.template.json` ships, so posting back what GET returned on a fresh
install answered 400 "At least one day must be enabled". An enabled per-day
schedule still needs a day on. A day that is off now keeps the times it
was posted with (the schedule picker sends them). Before, saving dropped
them, so turning the day back on showed the defaults.
- `POST /api/v3/config/main` answers `restart_required: true` only when the
save changed a setting the running display does not apply by itself.
Brightness (`brightness.set` and the config watcher), the per-mode
durations and plugin sections are applied live. A brightness-only save,
such as the MQTT bridge's slider, or a save that changed nothing, no longer
shows the restart banner. Hardware, rotation order, timezone and every
other setting still ask for the restart.
- `GET /api/v3/health` reports `degraded` when the display service is
stopped. Before, only the sub-checks changed, and the overall status stayed
`healthy` for as long as the last preview frame was under 60 s old.
`checks.display_loop.status` is now `stopped` when three things agree:
systemd says the service is not active, the control socket does not
answer, and there is no live heartbeat. Where the platform has no socket
(Windows) or it is switched off, nothing changes.
- `GET /api/v3/display/current-status` no longer reports the stopped
display's last state (`is_display_active: true`) from the cache for up to
120 s. When the control socket does not answer and the render loop's
heartbeat is absent, stale, or from a process that is gone (#726's rules),
the answer is unknown, with every field `null`. A display that still beats
without a socket, Windows and a socket switched off read the cache as
before. New `web_interface.display_state.display_gone()`.
- The garbage-collection timer (`GcMonitor`, above) no longer prints
`Exception ignored while calling GC callback ... 'NoneType' object has no
attribute 'perf_counter'` when the display service or a test run exits.
A collection during interpreter shutdown called it after the module's
`time` global was torn down. The monitor now binds its clock at
construction and does nothing once `sys.is_finalizing()`;
`install_gc_monitor()` unregisters it with `atexit`, and
`DisplayManager.cleanup()` (reached from SIGTERM through `run()`'s
`finally`) unregisters it with the frame recorder. New
`frame_timing.uninstall_gc_monitor()`.
- The web interface's state subscription (`StateSubscription`,
`src/ipc/client.py`) resubscribes about 1 s after a display restart, every
time. Its reconnect wait went back to the minimum only when the
subscription was stopped. A disconnect after a working connection kept
doubling the wait, so successive display restarts were followed by waits
of 1, 2, 4, 8, 16 and then 30 s for good.
During each wait the web answered from one-shot `state.get` connections
instead of its copy. The wait now resets once a connection has stored a
snapshot. A display that does not offer the stream is still retried
slowly.
- A plugin reload after a store update (`plugin.reload`, #720) no longer - A plugin reload after a store update (`plugin.reload`, #720) no longer
freezes the panel during Vegas. On ledpi a football reload froze it for freezes the panel during Vegas. On ledpi a football reload froze it for
3.0 s (`Render stall over: no frame for 3043ms`). The reload ran on the 3.0 s (`Render stall over: no frame for 3043ms`). The reload ran on the
@@ -258,6 +646,25 @@ policies are unchanged.
for the plugin is refused (`plugin-reloading`), and a config reconcile for the plugin is refused (`plugin-reloading`), and a config reconcile
neither loads it twice nor unloads it mid-load. A Vegas fetch that waited neither loads it twice nor unloads it mid-load. A Vegas fetch that waited
out a reload for the lock skips the old instance. out a reload for the lock skips the old instance.
- A plugin display duration that is not a number no longer stops the
display. Several plugins (clock-simple, calendar, countdown) return their
`display_duration` setting as it is in config.json, so a value saved as
`"20"` or `null` (the raw config editor, a hand edit) reached the run loop
as a string or None. Comparing it with 0 raised a TypeError that no
handler in the loop caught: the display service exited when that plugin's
screen came up, and systemd restarted it into the same crash. The
controller now reads the plugin's answer as a number: a numeric string
counts, and anything else (or a `get_display_duration()` that raises)
shows the mode for 30 s, with one warning per plugin.
- A scroll strip narrower than the panel scrolls instead of raising on every
frame. When a frame ran off the end of the strip, `ScrollHelper` copied
the strip's tail and then the rest of the frame from its head, which
assumed the head was that wide; for a narrower strip that raised
`ValueError: could not broadcast` at every position, so nothing was drawn
and each frame logged a traceback. Vegas builds such a strip, with no
lead-in, when its content is narrower than the chain. A frame that runs
off the strip now continues from its head column by column, so a narrow
strip repeats across the panel; a wide strip wraps exactly as before.
- The schedule-off blank and the WiFi notice no longer start with a - The schedule-off blank and the WiFi notice no longer start with a
scroller's leftovers. Both are drawn by the display controller rather than scroller's leftovers. Both are drawn by the display controller rather than
dispatched to a plugin, so #716's handover never reached them: drawn while dispatched to a plugin, so #716's handover never reached them: drawn while
@@ -267,6 +674,46 @@ policies are unchanged.
scroller or Vegas) were counted as 0.5-1 s freezes and logged as a scroller or Vegas) were counted as 0.5-1 s freezes and logged as a
`Render stall ... mid-scroll`. The controller now ends the scroll state `Render stall ... mid-scroll`. The controller now ends the scroll state
before drawing either. before drawing either.
- A plugin that keeps helpers in a package (elections' `providers/`,
flights' `enrichment/`, olympics' `data/` and `renderers/`) now runs its
updated helpers after a reload. Unloading dropped the package itself but
left its modules (`providers.feed`) in `sys.modules`, so the reload after a
store update imported the new `manager.py` and got the old helpers back from
the cache until the display restarted. `PluginLoader` now drops a plugin's
package modules when it unloads, and when a load fails part-way.
- Uninstalling a dev plugin that `scripts/dev/dev_plugin_setup.sh` linked
into the plugins directory now removes the link and leaves the checkout
alone. The store's removal passed the link to `shutil.rmtree`, which
refuses a symlink; its fallback then walked through the link and chmodded
every directory and file of the linked checkout to 0700, and the sudo stage
refused a path outside the plugins directory, so the uninstall failed with
the link still in place. The same removal discards the set-aside copy after
an install or update. A symlink, dangling or not, is now unlinked.
- A dev plugin linked in under a name its checkout does not share now loads.
`dev_plugin_setup.sh link-github foo <url>` clones `ledmatrix-foo` (the
repository naming convention) and links it as `plugins/foo`. The loader's
containment check for dependency installs resolved the link and looked for
`ledmatrix-foo` among the plugins directory's entries, found none, and
refused the plugin, so the load failed with "Dependency installation
failed" even when it had no `requirements.txt`. The check now looks for the
entry the path itself names in the plugins directory, the link, and still
only ever answers with an entry it found there.
- A plugin whose `update()` raises `asyncio.CancelledError` or `SystemExit`
no longer goes dark until a restart. Both derive from `BaseException`, not
`Exception`, and the update worker's bookkeeping caught only `Exception`:
the plugin kept its lock and stayed RUNNING, so it was never updated again
and every `display()` was skipped as busy. It is now recorded as that
update's failure, the same as any other raise. The plugin executor
reported such a call as a timeout; it now reports it as a failure.
- Saving a config change no longer freezes the panel while a plugin is busy.
`ConfigService` told its subscribers about a change while holding its lock,
and the display's per-plugin subscriber waits up to 5 s for a plugin in the
middle of an update. A save that enables or disables a plugin also queues a
reconcile, which the render thread runs, and its `get_config()` and
`unsubscribe()` waited behind every one of those callbacks. Subscribers now
run after the lock is released. One reload's notifications still finish
before the next one's start, and a callback `unsubscribe()` removed is not
running, and will not run, once it returns.
- A plugin whose `display()` raises now opens its circuit breaker. The first - A plugin whose `display()` raises now opens its circuit breaker. The first
frame of each screen goes through the plugin executor, which caught the frame of each screen goes through the plugin executor, which caught the
exception and returned False. The display read that as "no content" and exception and returned False. The display read that as "no content" and
@@ -276,6 +723,57 @@ policies are unchanged.
the plugin leaves rotation until the cooldown ends, the same as a raising the plugin leaves rotation until the cooldown ends, the same as a raising
`update()`. The display still moves straight on to the next mode. A hung `update()`. The display still moves straight on to the next mode. A hung
`display()` is still recorded once, as a hang. `display()` is still recorded once, as a hang.
- A plugin settings save that failed validation no longer leaks into the next
save. `ConfigManager.load_config()` returned its cached config itself (the
fast path from #410), so the form save's edits went into the cache before
validation ran, and a refused save left them there. The next save of any
other setting (another plugin's, a plugin toggle, the schedule) wrote them
to config.json: the refused value, and a nested secret typed into the same
form (`mqtt.password`, `league.espn_s2`, `flightaware.api_key`) in plain
text, because it had never reached config_secrets.json to be stripped.
The form also reloaded showing the refused values. `load_config()` now
returns a private copy, and the saves keep one, so nothing a caller edits
reaches the cache unless it is saved. The copy duplicates only the dicts
and lists (every other JSON value is immutable): 2.1 ms for a real 60 KiB
config on a Pi 4, against 6.8 ms for `copy.deepcopy`.
- `GET /api/v3/plugins/config` no longer returns secrets. It sent back the
plugin's section with config_secrets.json merged in, API keys and tokens
in plain text: the masking #276 added was dropped in #330. It also took
any id, so `?plugin_id=web_auth` returned the login's cookie-signing key
and password hash and `?plugin_id=github` the Plugin Store token. Secret
fields now come back blank, as the settings page renders them, and a
plugin with no schema has its credential-named fields blanked, as
`GET /config/main` does. Blank rather than the `••••••••` of
`GET /config/secrets`, because the save reads a blank secret as
"unchanged", so a client can post the response back without erasing
one. Core sections and malformed ids get a 400, as they already did from
reset and uninstall.
- Plugin settings with a table (a list of rows, such as geochron's cities
or the countdowns) save again when a text cell is blank or holds only
digits. A row posts its cells as `cities.0.timezone`, and the schema
lookup stopped at the list, so each cell was parsed with no schema: a
blank optional text cell became null, and a name like "2027" became a
number. Either failed validation, and every save of the page failed for
as long as the row existed. A plugin with a secret in its rows could not
be saved from the page at all, since the secret cell is drawn blank. The
lookup now steps from the index into the list's item schema.
- A plugin whose API key is required and has no default (youtube-stats)
can be saved from its settings page without typing the key in again. The
page draws a stored secret blank and posts the blank back; for a required
secret the save read that blank as null, failed validation, and refused
every save of the page. A blank secret field now means "unchanged", as it
already did for an optional one.
- `POST /api/v3/plugins/config` refuses a core section or a malformed
plugin id with a 400, as reset and uninstall already did.
`{"plugin_id": "display", ...}` merged unvalidated values into the core
display section (and added `"enabled": true` to it), and an id that was
not a string answered with a 500.
- A plugin text setting saves what was typed when that looks like a
boolean or JSON. The form save tried `true`/`false` and `[...]`/`{...}`
before it looked at the schema, so a text field holding "true", "False",
"[1, 2]" or "{}" was stored as a boolean, list or object, and the save
failed validation. Text fields, nullable ones included, are now taken as
typed; other types convert as before.
- A WiFi notice (such as "Connected to HomeNet" or "AP mode on") now shows - A WiFi notice (such as "Connected to HomeNet" or "AP mode on") now shows
within about a second of being posted. It was only checked between within about a second of being posted. It was only checked between
screens, so a 5 s notice posted during a 20 s screen expired before that screens, so a 5 s notice posted during a 20 s screen expired before that
@@ -284,6 +782,42 @@ policies are unchanged.
notice is what shows next, and Vegas resumes after it; before, a rotation notice is what shows next, and Vegas resumes after it; before, a rotation
screen showed instead and the notice expired behind it. An active screen showed instead and the notice expired behind it. An active
on-demand session still holds the panel until it ends. on-demand session still holds the panel until it ends.
- The Config Editor tab no longer shows API keys and tokens in plain
text. Its `config_secrets.json` editor (`/partials/raw-json`) was filled
with the file as it is on disk, so while the web login is off (the
default) anyone who could reach the port could read every credential,
although `GET /api/v3/config/secrets` masks them. The editor now shows the
same masked values. Saving it unchanged changes nothing, because the save
drops the masks and merges onto the stored file; to change a secret,
replace its mask. A list of secrets still needs every entry's real value
to be changed. The `config.json` editor is unchanged: its save writes the
file as given, so a mask there would be stored.
- A disabled plugin keeps its place in the rotation order and its Vegas
exclusion when the Display or Rotation & Durations tab is saved. The order
lists show enabled plugins only and rewrite their hidden inputs from those
rows as soon as they are drawn, so any save of either tab stored the lists
without the disabled plugin. Once re-enabled, it came back at the end of
the rotation and scrolling in Vegas again. A disabled plugin's saved id
now stays in its saved place (`widgets/plugin-order-list.js`); the id of
a plugin that is no longer installed is still dropped.
- Restoring a backup with "Reinstall missing plugins" installs only the
plugins that are missing. Every plugin the backup listed was sent to the
store's install, which replaces an installed copy with a fresh download,
so a restore onto the same device re-downloaded all of them in one
request. A plugin installed from its own URL is not in the registry, so
its "reinstall" failed and the restore answered "Restore failed" while
the plugin sat there installed. An installed plugin, found by the store's
own lookup (registry aliases included), is now listed under Skipped as
`plugin:<id> (installed)`.
- `POST /api/v3/config/main` answers a JSON body that does not parse with
400 `Invalid JSON in request body`, as `/config/raw/main` does, and an
empty JSON body with 400 `No data provided`. Both were a 500
`CONFIG_SAVE_FAILED` suggesting file permissions and disk space, with a
traceback logged at ERROR: `get_json()` raised inside the handler's
catch-all.
- Fonts restored from a backup show up in the Fonts tab and the font
pickers straight away. The font catalog is cached for five minutes, and
upload and delete cleared it but a restore did not.
- A game that goes live now takes over the panel within about a second. - A game that goes live now takes over the panel within about a second.
Live priority was only checked between screens, so a game that went live Live priority was only checked between screens, so a game that went live
during a 30 s screen waited for that screen to end. The frame loops and the during a 30 s screen waited for that screen to end. The frame loops and the
@@ -293,6 +827,45 @@ policies are unchanged.
screen showed first and the game came after it. Each check also asks each screen showed first and the game came after it. Each check also asks each
plugin `has_live_content()` once, where a plugin registered under several plugin `has_live_content()` once, where a plugin registered under several
modes used to be asked once per mode. modes used to be asked once per mode.
- A plugin action whose params hold `true`, `false` or `null` runs again.
`/api/v3/plugins/action` wrote the params into the source of the wrapper
that runs the plugin's script, and those JSON words are not Python, so the
wrapper stopped with a NameError and the action answered "Action failed".
The plugin file manager's category toggle sends `"enabled": true`, so
turning a category on or off in of-the-day always failed. The params now
reach the wrapper on its stdin; the script still receives them as JSON on
its own stdin, as before.
- An on-demand request that `/api/v3/display/on-demand/start` refuses no
longer runs later. With the display stopped the request goes to the
display's mailbox, and the display reads that mailbox for an hour without
looking at a request's age. So with "Start display service" unticked, the
answer was "Display service is not running", yet the next time the
display was started it ran that plugin, pinned if the request said so.
The same happened after "Failed to start display service". On either
refusal the route now takes its request back out of the mailbox, unless a
newer one has replaced it. A request the display acknowledges over the
control socket is now a success whatever systemd reports: a display run
by hand or in the emulator was told "not running" for a request it had
already taken, and with "Start display service" ticked the route tried to
start the service beside it.
- `/api/v3/plugins/operation/<id>` reports a queued operation as `pending`
instead of answering 500. The queue keeps an operation's callback among
its parameters until it runs, and the status route tried to send that
function as JSON. An install queued behind another plugin's install
failed every status poll until the first one finished. Parameters whose
name starts with `_` are internal and are no longer in the answer.
- A second click on Install while that plugin is still installing, or an
Uninstall during its install, now answers 409 "already has an install,
update or uninstall in progress" instead of 500 "An error occurred". The
first operation carried on either way. The uninstall route also stopped
recording a failed uninstall in the operation history for an uninstall
that never started.
- `/api/v3/plugins/<plugin_id>/static/<path>` serves images and other
binary files. It opened every file as UTF-8 text, so a plugin's icon or
preview image answered 500 `UnicodeDecodeError`. Files are now sent as
they are on disk, an image with its own content type; HTML, JavaScript,
CSS, JSON and other text keep the types they had. The path checks are
unchanged.
- The display schedule turns the panel off at exactly the end time. A window - The display schedule turns the panel off at exactly the end time. A window
now runs from its start time up to, but not including, its end time: with now runs from its start time up to, but not including, its end time: with
07:00-23:00 the panel is on at 07:00 and off at 23:00. Before, the end 07:00-23:00 the panel is on at 07:00 and off at 23:00. Before, the end
@@ -300,10 +873,77 @@ policies are unchanged.
the panel went off at 23:00 or at 23:01 depending on when in the minute the panel went off at 23:00 or at 23:01 depending on when in the minute
that check ran. Windows that cross midnight and per-day schedules follow that check ran. Windows that cross midnight and per-day schedules follow
the same rule, and so does the dim schedule. the same rule, and so does the dim schedule.
- The MQTT bridge settings on the Tools tab can save a broker password with
TLS off. The server refuses that unless `allow_insecure_mqtt` is set, and
the form had no way to set it, so a password-protected broker on a home
network without TLS could not be saved from the web UI, and once such a
password was stored every later save failed too. While "Use TLS" is
unchecked the form now shows "Allow without TLS (trusted network)",
prefilled from the saved settings. It is off until ticked, so the server
still refuses a cleartext password by default.
- The Overview's plugin-config warning check stops polling. It asked
`/api/v3/plugins/reconciliation-status` every 2 s until startup
reconciliation reported done, and the route reports not done whenever its
status file is missing: reconciliation raised before writing it, or /tmp
was cleaned under a long-running web service. The page then sent that
request every 2 s for as long as it stayed open, whichever tab was showing.
It now gives up after a minute and only polls while the Overview is on
screen.
- Moving the Brightness slider on the Display tab no longer throws an error
in the browser console on every step. Its handler also updated a "LED
brightness" line that was removed from the page in #387; the lookup is
gone.
- Creating an API token on the General tab no longer leaves the page asking
"Leave site?" on reload. The unsaved-changes guard marks a form when you
type in it and clears the mark only after an htmx save, and the token form
saves with a plain request, so it stayed marked after the token was
created. It is cleared once the token is saved.
- An on-demand session that ends during scheduled-off hours, by expiring or - An on-demand session that ends during scheduled-off hours, by expiring or
being stopped, blanks the panel within about a second. It used to stay on being stopped, blanks the panel within about a second. It used to stay on
until the next minute, because the once-a-minute schedule check had until the next minute, because the once-a-minute schedule check had
already run that minute and the session had overridden its answer. already run that minute and the session had overridden its answer.
- Check & Update All updates what is installed now. A second run in the
same page sent the plugins the first run had seen, so a plugin uninstalled
since then failed with "plugin not found" and one installed since was
skipped. After a run the installed cards and the Updates badge show the
new versions; they kept offering "Update to vX" for what had just been
updated until the page was reloaded.
- The Run On-Demand dialog lists a plugin's display modes, so a mode other
than the first can be started, and pinned. `/api/v3/plugins/installed`
never sent `display_modes`, which the dialog reads, so every plugin
offered only its own id under "This plugin exposes a single display
mode", and the display started its first mode. Each entry now carries
`display_modes`, the modes its manifest declares.
- Installing Weather, Music, Stocks or Leaderboard from the Plugin Store
enables it, as installing any other plugin does. Each installs under the
id its manifest declares (`ledmatrix-weather` for the store's `weather`),
but the store enabled the store id, which `/api/v3/plugins/toggle`
answered with "Plugin not found": the plugin stayed disabled behind
"installed, but enabling it failed". `POST /api/v3/plugins/install` now
answers with the installed `plugin_id` (in the operation's result when it
is queued), and the store enables that.
- Reinstalling a plugin from the Plugin Store leaves it enabled or disabled
as it was. Reinstall enabled it as a fresh install does, so a plugin the
user had switched off came back on.
- A Plugin Store install that takes more than a minute is no longer
reported as failed. The store stopped waiting after 60 s and showed
"Install operation timed out" while the server, which allows the
plugin's dependency install 300 s on its own, carried on and usually
succeeded; the plugin was then neither enabled nor listed until the page
was reloaded. The store now waits up to 10 minutes, and if it still has
no answer it reloads the installed list and says the install may still
be running.
- The Plugin Store's category filter lists every category its plugins
have. It offered a fixed seven while the registry uses about twenty, so
plugins filed under productivity, utility, transit and the rest could not
be filtered to, and "Financial" missed the plugin filed under "finance".
The choices are now built from the store's plugins, as the Starlark
section's are.
- The Install button under Install Single Plugin (Plugin Manager > Install
from GitHub) runs one handler per click. It also had an inline `onclick`
whose handler threw a `ReferenceError` on every click; only the other
handler's request went out, and making the inline one work would have
sent every install twice. The inline handler is gone.
- `/api/v3/plugins/installed` no longer reports the display's plugins as - `/api/v3/plugins/installed` no longer reports the display's plugins as
`live` while `/api/v3/health` says `display_loop: stalled`. The runtime `live` while `/api/v3/health` says `display_loop: stalled`. The runtime
snapshot is written from its own thread, which kept going while the render snapshot is written from its own thread, which kept going while the render
@@ -315,11 +955,41 @@ policies are unchanged.
heartbeat when the service stops), is `stale` at once instead of `live` heartbeat when the service stops), is `stale` at once instead of `live`
for up to 180 s. No new files or writes: both checks are on the reading for up to 180 s. No new files or writes: both checks are on the reading
side. side.
- A scoreboard's scroll and Vegas cards with `scroll_card.date_format:
"weekday"` now show the printed date's own weekday. A Friday 8 PM ET game
read "Sat Oct 2". The card took the weekday in the plugin's own
`timezone` setting, which ships blank, so it fell back to UTC, while the
"Oct 2" beside it came from the zone the plugin actually resolves (its
setting, then the global one, then the system zone). Every zone is within a
day of UTC, so the card now finds which day near the start's UTC date has
the printed month and day and names that one. Games east of UTC (Auckland,
Kiritimati) were off by a day the other way and are fixed the same way.
The switch-mode scorebug, which already used the plugin's resolved zone,
shares the same formatter and draws what it drew before.
- `/api/v3/display/current-status` reflects a wake from scheduled-off, a - `/api/v3/display/current-status` reflects a wake from scheduled-off, a
schedule-off blank, or an on-demand session starting or ending at once, schedule-off blank, or an on-demand session starting or ending at once,
even when the mode name stays the same. The display republished its even when the mode name stays the same. The display republished its
current state only on a mode change or every 30 s, so `is_display_active` current state only on a mode change or every 30 s, so `is_display_active`
and `on_demand_active` could be up to 30 s out of date. and `on_demand_active` could be up to 30 s out of date.
- `/api/v3/display/current-status` no longer answers `mode: null` over the
control socket (#735) once the same mode has been on screen for more than
two minutes: a live game under live priority, Vegas, or a single plugin.
On ledpi it returned nulls in every sample for 90 minutes while the
display was live. The state stream's version leaves out the timestamps
that move on every publish, and a subscriber's keepalive tick carried only
the loop's heartbeat. So the web interface's copy kept the
`display.last_updated` of the last real change, and the reader's 120 s
rule called it unknown. The plugin runtime section had the same problem:
with no plugin changing state, `/plugins/state` and the `runtime` in
`/plugins/installed` read `stale` after 180 s. A tick (and a `state.get`
answer with `since`) now carries `volatile`: the current values of those
timestamps (`display.last_updated`, `on_demand.last_updated` and
`remaining`, `plugins.published_at`), and the subscription merges them
into its copy. The verdicts are unchanged. A render thread that stops
publishing still reads as `stalled` after 60 s and as unknown after 120 s,
a runtime publisher that stops still goes `stale`, and a subscription that
goes quiet still falls back to the cache. The cache path's 120 s rule is
unchanged.
### Scrolling ### Scrolling
+6 -1
View File
@@ -151,6 +151,11 @@ The system supports live, recent, and upcoming game information for multiple spo
- **1GB models (Pi 3B / 3B+), the 512MB Pi Zero 2 W and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`. Once running, keep an eye on memory: see [docs/LOW_MEMORY_BOARDS.md](docs/LOW_MEMORY_BOARDS.md). - **1GB models (Pi 3B / 3B+), the 512MB Pi Zero 2 W and other low-memory boards**: supported, but the `rpi-rgb-led-matrix` C++ build needs more memory than the Pi has. The installer detects this automatically, compiles with fewer parallel jobs, and adds a temporary swapfile for the build which it removes afterwards. Expect that step to take 15-25 minutes instead of 2-5, and leave at least **3GB free** on the SD card. If you manage swap yourself, opt out with `--skip-swap`. To pin the compiler down further, use `--build-jobs 1`. Once running, keep an eye on memory: see [docs/LOW_MEMORY_BOARDS.md](docs/LOW_MEMORY_BOARDS.md).
### Operating system
- **Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12)**, 64-bit recommended. Trixie is the current release and the one to pick for a new SD card; an existing Bookworm install works as it is, no upgrade needed. The installer checks this first and stops with directions on anything else (Bullseye and older, the desktop edition, other distributions).
- **Python**: whatever the OS ships, 3.13 on Trixie and 3.11 on Bookworm. Don't install a different Python; the installer and the services use the system `python3`.
- **Networking**: NetworkManager, the default on both. Choosing a WiFi network from the web page and the `LEDMatrix-Setup` hotspot need it; if you switched to dhcpcd in `raspi-config`, switch back (Advanced Options → Network Config → NetworkManager).
### RGB Matrix Bonnet / HAT ### RGB Matrix Bonnet / HAT
- [Adafruit RGB Matrix Bonnet/HAT](https://www.adafruit.com/product/3211) – supports one “chain” of horizontally connected displays - [Adafruit RGB Matrix Bonnet/HAT](https://www.adafruit.com/product/3211) – supports one “chain” of horizontally connected displays
- [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular` as hardware mapping)* - [Adafruit Triple LED Matrix Bonnet](https://www.adafruit.com/product/6358) – supports up to 3 vertical “chains” of horizontally connected displays *(use `regular` as hardware mapping)*
@@ -249,7 +254,7 @@ These are not required and you can probably rig up something basic with stuff yo
<img width="512" height="361" alt="Step 2 Other " src="https://github.com/user-attachments/assets/166a22e8-8067-48df-9f80-50c91f573356" /> <img width="512" height="361" alt="Step 2 Other " src="https://github.com/user-attachments/assets/166a22e8-8067-48df-9f80-50c91f573356" />
5. Then choose Raspbian OS (64-bit) Lite (Trixie) 5. Then choose Raspbian OS (64-bit) Lite (Trixie). Bookworm Lite (listed as Legacy) also works; see [Operating system](#operating-system) below
<img width="512" height="361" alt="Step 4 Trixie Lite 64" src="https://github.com/user-attachments/assets/3b8590ce-b810-4dfe-9253-26e0d4f8ed1e" /> <img width="512" height="361" alt="Step 4 Trixie Lite 64" src="https://github.com/user-attachments/assets/3b8590ce-b810-4dfe-9253-26e0d4f8ed1e" />
+2 -2
View File
@@ -17,13 +17,13 @@ The LEDMatrix emulator allows you to run and test LEDMatrix displays on your com
## Prerequisites ## Prerequisites
### System Requirements ### System Requirements
- Python 3.10 or higher - Python 3.11 or higher (3.11 and 3.13 are tested)
- Windows, macOS, or Linux - Windows, macOS, or Linux
- At least 2GB RAM (4GB recommended) - At least 2GB RAM (4GB recommended)
- Internet connection for plugin downloads - Internet connection for plugin downloads
### Required Software ### Required Software
- Python 3.10+ - Python 3.11+
- pip (Python package manager) - pip (Python package manager)
- Git (for plugin management) - Git (for plugin management)
+19 -1
View File
@@ -15,6 +15,12 @@ This guide will help you set up your LEDMatrix display for the first time and ge
- Power supply (5V, 4A minimum recommended) - Power supply (5V, 4A minimum recommended)
- MicroSD card (16GB minimum) - MicroSD card (16GB minimum)
**Software:**
- Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12). Trixie
is the current release; Bookworm is listed as Legacy in Raspberry Pi
Imager. No other system is supported, and the installer says so up front.
- The OS's own Python: 3.13 on Trixie, 3.11 on Bookworm
**Network:** **Network:**
- WiFi network (or Ethernet cable) - WiFi network (or Ethernet cable)
- Computer with web browser on same network - Computer with web browser on same network
@@ -28,7 +34,8 @@ This guide will help you set up your LEDMatrix display for the first time and ge
There is no prebuilt SD card image — you install LEDMatrix onto stock There is no prebuilt SD card image — you install LEDMatrix onto stock
Raspberry Pi OS Lite yourself: Raspberry Pi OS Lite yourself:
1. Flash Raspberry Pi OS Lite to the MicroSD card (Raspberry Pi Imager) 1. Flash Raspberry Pi OS Lite (Trixie, or Bookworm) to the MicroSD card
(Raspberry Pi Imager)
2. Connect the LED matrix to your Raspberry Pi, insert the card, and 2. Connect the LED matrix to your Raspberry Pi, insert the card, and
power on power on
3. SSH into the Pi and run the one-shot installer: 3. SSH into the Pi and run the one-shot installer:
@@ -39,6 +46,17 @@ Raspberry Pi OS Lite yourself:
[README Installation Steps / Quick Install](../README.md#installation-steps) [README Installation Steps / Quick Install](../README.md#installation-steps)
for full details for full details
The one-shot installer installs the newest release (the **stable** update
channel). To run the newest, unreleased code from `main` instead (the
**beta** channel), put `LEDMATRIX_CHANNEL=beta` in front of `bash`:
```bash
curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | LEDMATRIX_CHANNEL=beta bash
```
A manual clone starts on `main`; add `--beta` to `first_time_install.sh`
to stay on it, or leave it off and the first update after the next
release moves the device onto releases. You can switch channels later on
the General tab.
**Expected Behavior after install:** **Expected Behavior after install:**
- LED matrix will light up - LED matrix will light up
- A fresh install ships only the bundled `starlark-apps` and - A fresh install ships only the bundled `starlark-apps` and
+250 -24
View File
@@ -4,14 +4,17 @@ The display process serves a Unix socket that the web interface uses to send
it commands and get an answer back. It replaces the cache-file "mailboxes" on it commands and get an answer back. It replaces the cache-file "mailboxes" on
the SD card one command at a time. Stage 1 carries on-demand start, stop and the SD card one command at a time. Stage 1 carries on-demand start, stop and
status. Stage 2 makes those commands land within a frame on every kind of status. Stage 2 makes those commands land within a frame on every kind of
screen, and adds `brightness.set` and `plugin.reload`. The file mailbox stays screen, and adds `brightness.set` and `plugin.reload`. Stage 3 adds a state
as a fallback for one release. stream (`state.get`, `state.subscribe`), so the web interface reads what the
display is doing from the socket instead of from cache files the display
wrote to the SD card. The file mailbox and the cache keys stay as a fallback
for one release.
| | | | | |
|---|---| |---|---|
| Socket | `/run/ledmatrix/control.sock` (tmpfs) | | Socket | `/run/ledmatrix/control.sock` (tmpfs) |
| Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` | | Served by | the display process ([`src/ipc/server.py`](../src/ipc/server.py)), started by `DisplayController.run()` |
| Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop`, `POST /api/v3/plugins/update` (reload), `POST /api/v3/config/main` (brightness) | | Used by | the web interface ([`src/ipc/client.py`](../src/ipc/client.py)): `POST /api/v3/display/on-demand/start` and `/stop`, `POST /api/v3/plugins/update` (reload), `POST /api/v3/config/main` (brightness); and through [`web_interface/display_state.py`](../web_interface/display_state.py) (the state stream), `GET /api/v3/display/current-status`, `/display/on-demand/status`, `/plugins/installed` (`runtime`), `/plugins/state` and the reconciliations, `/health` (`display_loop`) |
| Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it | | Contract | [`src/ipc/contract.py`](../src/ipc/contract.py): messages, versions, framing and the socket path; both sides import it |
| Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it | | Override | `LEDMATRIX_CONTROL_SOCKET=/some/path.sock` for both processes, or `=off` to disable it |
@@ -36,6 +39,11 @@ running, the socket does not exist, and the web interface knows right away.
## Protocol (version 1) ## Protocol (version 1)
Stage 3 is still version 1: `state.get` and `state.subscribe` are new
commands, and a stage-2 display answers them `unknown_command`, which the
web interface treats as "no socket" and falls back from.
**Framing.** One JSON object per line (newline-delimited JSON), UTF-8, at **Framing.** One JSON object per line (newline-delimited JSON), UTF-8, at
most 64 KiB per line (`MAX_MESSAGE_BYTES`). Senders encode with most 64 KiB per line (`MAX_MESSAGE_BYTES`). Senders encode with
`ensure_ascii`, so a newline never appears inside a message. A connection `ensure_ascii`, so a newline never appears inside a message. A connection
@@ -75,6 +83,8 @@ one. Clients branch on `error.code`, never on the message text.
| `on_demand.status` | — | `{on_demand: {...}, current_mode, display_active}` | answered directly | | `on_demand.status` | — | `{on_demand: {...}, current_mode, display_active}` | answered directly |
| `brightness.set` | `{brightness: int 0-100}` | `{brightness, panel_brightness, dimmed, display_active}` | queued, awaited (2 s) | | `brightness.set` | `{brightness: int 0-100}` | `{brightness, panel_brightness, dimmed, display_active}` | queued, awaited (2 s) |
| `plugin.reload` | `{plugin_id}` | `{plugin_id, reloaded: true, version, modes}` | queued, awaited (10 s) | | `plugin.reload` | `{plugin_id}` | `{plugin_id, reloaded: true, version, modes}` | queued, awaited (10 s) |
| `state.get` | `{since?, epoch?}` | a state snapshot (see "The state stream") | answered directly |
| `state.subscribe` | — | a state snapshot, then pushed `state` / `tick` events | answered directly, then a stream |
`duration` is a number of seconds, or a numeric string. `0`, `null` or `""` `duration` is a number of seconds, or a numeric string. `0`, `null` or `""`
mean "until stopped". `pinned` must be a real boolean: the REST route has mean "until stopped". `pinned` must be a real boolean: the REST route has
@@ -130,6 +140,19 @@ web interface treats like any other socket failure and falls back from, and
`hello` lists the commands a display knows. The version changes only when the `hello` lists the commands a display knows. The version changes only when the
envelope or the meaning of an existing command changes. envelope or the meaning of an existing command changes.
**Events.** `state.subscribe` is the one command with more than one message
in reply. After its response, the display pushes events on the same
connection until either side hangs up:
```json
{"v": 1, "id": "<the subscribe id>", "event": "state", "result": {...a state snapshot...}}
{"v": 1, "id": "<the subscribe id>", "event": "tick", "result": {"version": 7, "epoch": "…", "pid": 812, "served_at": 1790000000.1, "changed": false, "loop": {...}, "volatile": {"display": {"last_updated": 1790000000.0}, "...": "..."}}}
```
An event has `event` where a response has `ok`, which is how a reader tells
them apart. The client sends nothing after the subscribe; anything it does
send is ignored.
**Error codes:** `bad_json`, `bad_request`, `message_too_large`, **Error codes:** `bad_json`, `bad_request`, `message_too_large`,
`unsupported_version`, `unknown_command`, `invalid_args`, `busy` (queue full, `unsupported_version`, `unknown_command`, `invalid_args`, `busy` (queue full,
or too many connections), `forbidden` (peer credentials refused), `internal`. or too many connections), `forbidden` (peer credentials refused), `internal`.
@@ -147,6 +170,177 @@ print(client.brightness_set(60))
EOF EOF
``` ```
## The state stream (stage 3)
Before stage 3 the web interface learned what the display was doing by
reading files the display kept writing:
| What | Written by the display | How often | Medium |
|---|---|---|---|
| current mode, plugin, `is_display_active`, `on_demand_active` | `display_current_state` | every mode change, every flag change, and every 30 s | cache (SD card) |
| on-demand session | `display_on_demand_state` | on each on-demand event | cache (SD card) |
| plugin runtime snapshot (#690) | `plugin_runtime_snapshot` | on a change (at most every 10 s), else every 60 s | cache (SD card) |
| render-loop liveness (#687) | `display-heartbeat.json` | every 5 s | tmpfs |
Now the display also keeps the same state in memory and serves it on the
socket.
**The snapshot.** `state.get` and `state.subscribe` answer with one object:
```json
{"schema": 1, "version": 42, "epoch": "3f9c0d1e2a4b5c6d", "pid": 812,
"served_at": 1790000000.1, "changed": true,
"loop": {"heartbeat_age_seconds": 1.8, "armed": true, "stale_after": 60.0},
"state": {
"display": {"mode": "nfl_live", "plugin_id": "football-scoreboard", "mode_index": 3,
"total_modes": 9, "on_demand_active": false, "is_display_active": true,
"last_updated": 1790000000.0},
"on_demand": {"active": false, "status": "idle", "...": "as display_on_demand_state"},
"brightness": {"brightness": 80, "panel_brightness": 40, "dimmed": true},
"plugins": {"schema": 1, "running": true, "published_at": 1789999998.5, "...": "as plugin_runtime_snapshot"},
"loop": {"heartbeat_age_seconds": 1.8, "armed": true, "stale_after": 60.0}
}}
```
- `display` and `on_demand` are the dicts the cache keys hold, `plugins` is
the runtime snapshot (`build_runtime_snapshot`), and `brightness` is the
configured level, what the panel shows now, and whether the dim schedule
has it dimmed. A section not published yet is `null`.
- `loop` is not published: the display measures it when it answers, from
the render thread's last beat in memory (`RenderWatchdog.liveness()`),
the same beat that writes the heartbeat file. So it keeps ageing while the
render thread is stuck, and the socket's connection threads still answer.
`heartbeat_age_seconds` is `null` until the loop has drawn its first frame.
- `version` goes up whenever a section changes, ignoring the timestamps that
move on every publish (`last_updated`, `remaining`, `published_at`). It
counts within an `epoch`, one run of the display process, so a reader that
sees a new `epoch` has a restarted display.
- `state.get` with `since` and `epoch` from an earlier answer gets just
`{changed: false, version, epoch, pid, served_at, loop, volatile}` while
nothing has changed. `volatile` is `{section: {key: value}}`: the current
values of those ignored timestamps, which the reader merges into the copy
it has. They don't make a new version, but they are still news:
`display.last_updated` is how a reader knows the render thread is still
publishing, and `plugins.published_at` the runtime publisher. Without
them a reader's copy kept the timestamps of the last real change, so a
mode on screen for over 120 s read as unknown.
- A snapshot that would not fit in a message (hundreds of plugins) is sent
without `plugins`, and `truncated: ["plugins"]` says so. Readers then use
the cache for that section only.
**The stream.** `state.subscribe` answers with the snapshot, then:
- a `state` event (a full snapshot) whenever the version changes, and
- a `tick` at least every 5 s (`SUBSCRIBE_KEEPALIVE_SECONDS`) when nothing
changed. It is the short `changed: false` answer, so it carries `loop`
(a stalled render loop shows up within one tick) and `volatile` (the
timestamps stay as fresh as the writers keep them), and it tells the
reader the connection is alive.
A slow reader is never sent a backlog: each event is the latest version, so
one that falls behind skips the versions in between. A reader that has heard
nothing for 15 s (three keepalives) stops trusting its copy.
**Who publishes, and when.** All of it is in memory, with no disk writes:
- the render thread, at the places it already published the cache keys:
`display` and `brightness` on every pass of
`_publish_current_mode_state_if_changed()` (every loop pass, and every
`_service_pending_changes()` in a dwell, a scrolling screen or Vegas), and
`on_demand` in `_publish_on_demand_state()`. Every pass refreshes
`display.last_updated`, so a reader can tell when the render thread has
stopped publishing, just as the cache key's 120 s `max_age` does.
- the plugin runtime publisher's thread, on every 5 s tick: the snapshot is
rebuilt when the state machine changed, otherwise only its `published_at`
moves. A change reaches subscribers within a tick, without the cache's
10 s throttle.
Publishing is a hand-off, as the command queue is in the other direction.
The hub (`StateHub` in [`src/ipc/server.py`](../src/ipc/server.py)) holds a
lock only to swap a dict reference, compare it with the last one and bump the
version. Every socket write happens on the subscriber's own connection
thread. The render thread never waits for a reader.
### Readers in the web interface
[`web_interface/display_state.py`](../web_interface/display_state.py) holds
one `state.subscribe` connection per web process
(`src.ipc.client.StateSubscription`, a daemon thread, started on the first
read and reconnecting with a backoff of 1 s up to 30 s). A route answers
from the latest pushed snapshot in memory. Before the subscription has one,
the route asks once with `state.get` (0.5 s timeout). When neither works, it
reads the cache keys and the heartbeat file as before:
| Route | From the socket | Fallback |
|---|---|---|
| `GET /api/v3/display/current-status` | `state.display` | `display_current_state` |
| `GET /api/v3/display/on-demand/status` | `state.on_demand`, with `remaining` worked out from `expires_at` now | `display_on_demand_state` |
| `GET /api/v3/plugins/installed` (`runtime`), `/plugins/state`, `POST /plugins/state/reconcile` and the startup reconciliation | `state.plugins` + `state.loop` | `plugin_runtime_snapshot` + `display-heartbeat.json` |
| `GET /api/v3/health` (`checks.display_loop`) | `state.loop` | `display-heartbeat.json` |
Each answer says where it came from: `source: "socket" | "cache"` (or
`"heartbeat_file"` for the health check).
The SSE display stream (`/api/v3/stream/display`) reads the preview frame
file, not a cache key, so it does not change.
**The same verdicts either way.** The socket's answers are judged by the
rules the cache readers apply (#726):
- the runtime view is `stalled` when the render loop's heartbeat age is at
least `HEARTBEAT_STALE_SECONDS` (60 s, the health check's threshold), and
then reports no per-plugin facts;
- it is `stale` when the snapshot is older than its `stale_after` (the
publisher thread stopped);
- with no beat yet, the snapshot is judged on its own;
- there is no pid check, because the display that answered is alive;
- a `display` section the render thread has not refreshed for 120 s reads
as unknown, as the cache key does once it ages out.
The age a reader uses is the age the display measured, plus the time since
the snapshot arrived.
### Fewer SD writes
The cache keys are still written, for one release, as the fallback. While
the socket serves the readers, the display writes two of them less often.
"Serves the readers" means a subscriber is connected, or a `state.get` came
within the last 60 s (`StateHub.readers_active()`):
- `display_current_state` is no longer written on every mode change: once
every 60 s (`CURRENT_STATE_RELAXED_REFRESH_SECONDS`, inside the readers'
120 s `max_age`), and at once when `is_display_active` or
`on_demand_active` changes.
- `plugin_runtime_snapshot`'s refresh goes from 60 s to 120 s
(`RELAXED_REFRESH_INTERVAL`), and the snapshot says so in its own
`refresh_interval` and `stale_after` (360 s). Changes are still written at
once, at most every 10 s.
`display_on_demand_state` is written only on events, so it is unchanged.
The heartbeat file is on tmpfs, so it costs no SD writes, and it stays: the
automatic update's health check reads it.
This is safe because the relaxed rate only applies while readers are using
the socket. If they stop (the web interface loses the socket, or is stopped),
the next publish after the reader window writes a changed mode at once, and
the runtime refresh goes back to 60 s. A fallback reader in that window sees
a mode up to 60 s old, never one older than its `max_age`.
Measured with fake clocks (`test_cache_writes_per_minute_with_and_without_socket_readers`
in `test/test_state_stream_readers.py`), for a rotation of 15 s screens:
| Key | Writes/min, no socket readers | Writes/min, socket readers |
|---|---|---|
| `display_current_state` | 4.0 | 1.0 |
| `plugin_runtime_snapshot` | 1.0 | 0.5 |
| Total | 5.0 | 1.5 |
That is 70% fewer writes for these keys: about 2,200 a day instead of 7,200.
Shorter screens save more, because the old rate followed the mode changes.
A display that rarely changes mode (one plugin, a long live game) saves less. Plugin
data caches, the error snapshot and font usage are written by other code
and are not affected.
## How the display applies a command ## How the display applies a command
The server's threads never touch rendering. A connection thread parses the The server's threads never touch rendering. A connection thread parses the
@@ -279,7 +473,16 @@ block the render loop or crash it:
process created. process created.
- **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind - **Never fatal.** If the server cannot start (Windows, no `AF_UNIX`, a bind
failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs failure, `LEDMATRIX_CONTROL_SOCKET=off`), it logs that and the display runs
as before. The web interface then uses the mailbox. as before. The web interface then uses the mailbox, and reads the cache
keys and the heartbeat file.
- **Subscribers (stage 3).** A `state.subscribe` connection gives its request
slot back and takes one of 4 subscriber slots (`MAX_SUBSCRIBERS`). A fifth
gets `busy`. So a few browsers' web processes holding streams can never
use up the 8 slots that commands need. Each subscriber has its own thread.
A send that cannot finish within the 2 s IO timeout (a reader that stopped
reading) drops that subscriber. Nothing else waits for it, and the render
thread only publishes to the hub. `close()` wakes every subscriber, so
they end at once.
## Security model ## Security model
@@ -313,8 +516,10 @@ read its state, set the brightness, and reload a plugin the display is
already running, all of which anyone who can reach the web UI can already do already running, all of which anyone who can reach the web UI can already do
(the last by restarting the display). Nothing on the socket runs a shell, (the last by restarting the display). Nothing on the socket runs a shell,
writes a file, or names a path, and `plugin.reload` cannot make the display writes a file, or names a path, and `plugin.reload` cannot make the display
import a plugin it was not running. Stage 2 changed none of the access rules import a plugin it was not running. Stages 2 and 3 changed none of the
above. access rules above. The state stream carries what the cache keys already
held, and those are readable by the same group. A subscriber goes through
the same connect-time and peer-credential checks as any other connection.
**Development.** A display that is not root and cannot write to **Development.** A display that is not root and cannot write to
`/run/ledmatrix`, such as `python3 run.py -e` from a checkout, serves the `/run/ledmatrix`, such as `python3 run.py -e` from a checkout, serves the
@@ -352,29 +557,34 @@ device never touches the live display.
"which sections changed" ack had no reader: the web interface knows "which sections changed" ack had no reader: the web interface knows
what it saved. A reload from the socket thread would also run every what it saved. A reload from the socket thread would also run every
config subscriber on a second thread beside the watcher's. config subscriber on a second thread beside the watcher's.
3. **A state stream.** A `subscribe` command that keeps the connection open 3. **A state stream (done).** `state.get` (a versioned snapshot) and
and pushes events: mode changes, on-demand state (including the outcome of `state.subscribe` (the snapshot, then pushed changes and keepalive ticks)
an acked on-demand command, which today is only published), plugin carry the current mode, the on-demand state (including the outcome of an
runtime state, the outcome of a reload that answered `pending`, and the acked on-demand command), the brightness, the plugin runtime snapshot and
heartbeat. It replaces the polled `display_current_state`, the render loop's liveness, all served from memory (see "The state
`plugin_runtime_snapshot` (#690) and `display-heartbeat.json` (#687) for stream"). The web interface's readers use it and fall back to the cache
readers that hold a connection. The web interface relays it to its keys and the heartbeat file. `display_current_state` and
existing SSE stream. The files remain for one release for older readers. `plugin_runtime_snapshot` are written less often while it serves them.
- The server's per-connection threads (8 at most) do not suit long-lived The keys remain for one release.
subscribers. A subscriber needs its own bound and a writer that drops - Left for later: the outcome of a `plugin.reload` that answered
events for a slow reader rather than blocking the display. `pending` is visible only as the plugin's new `loaded_version` in
- Events are produced on the render thread, so publishing must be a `state.plugins`, not as an event of its own.
non-blocking hand-off, like the queue in the other direction. - Left for later: the SSE display stream reads the preview frame, not
- The store's install of an already-enabled plugin, and an uninstall that state, so nothing relays the stream to the browser yet. A browser still
keeps its config, still answer `restart_required`. With the stream they polls the REST routes, which now answer from memory.
can use a load/unload command and report the result the same way the - Left for later: the store's install of an already-enabled plugin, and
update route does now. an uninstall that keeps its config, still answer `restart_required`.
They can now use a load/unload command and report the result the same
way the update route does.
4. **Retire the mailboxes.** After a release in which every device has had the 4. **Retire the mailboxes.** After a release in which every device has had the
socket, the web interface stops writing `display_on_demand_request`, and socket, the web interface stops writing `display_on_demand_request`, and
the display stops polling it, logging the plugins that still write it so the display stops polling it, logging the plugins that still write it so
they can move to an in-process `request_display()`. The other cache keys they can move to an in-process `request_display()`. The other cache keys
used as messages (`plugin_error_clear_request` and the remaining used as messages (`plugin_error_clear_request` and the remaining
`display_*` keys) move to the socket or to tmpfs. `display_*` keys) move to the socket or to tmpfs. The display also stops
writing `display_current_state`, `display_on_demand_state` and
`plugin_runtime_snapshot` once the web interface no longer falls back to
them.
## Checking it on a device ## Checking it on a device
@@ -405,3 +615,19 @@ sudo journalctl -u ledmatrix | grep -E "Brightness set|Reload(ing|ed) plugin"
`unknown_command` in `brightness_socket_error` or `reload_error` means the `unknown_command` in `brightness_socket_error` or `reload_error` means the
display runs a stage-1 build: restart it once to pick up this one. display runs a stage-1 build: restart it once to pick up this one.
The state stream:
```bash
curl -s localhost:5000/api/v3/display/current-status # ... "source": "socket"
curl -s localhost:5000/api/v3/health | python3 -m json.tool | grep -A3 display_loop
python3 - <<'EOF'
from src.ipc import client # run from the project directory
snap = client.state_get()
print(snap['version'], snap['epoch'], snap['loop'], snap['state']['display'])
EOF
```
`"source": "cache"` means the web interface could not use the socket: the
display is stopped, predates stage 3, or the web user is not in the
socket's group.
+20 -1
View File
@@ -33,6 +33,9 @@ in again (services pick them up on restart).
| `/run/ledmatrix/control.sock` | `root` : cache directory's group (`ledmatrix`) | `660` | The display's control socket; only root and that group can connect. See [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md#security-model) | | `/run/ledmatrix/control.sock` | `root` : cache directory's group (`ledmatrix`) | `660` | The display's control socket; only root and that group can connect. See [IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md#security-model) |
| `scripts/fix_perms/safe_plugin_rm.sh`, `safe_pip_install.sh` | `root:root` | `755` | Run as root through sudo, so the web user must not be able to edit them | | `scripts/fix_perms/safe_plugin_rm.sh`, `safe_pip_install.sh` | `root:root` | `755` | Run as root through sudo, so the web user must not be able to edit them |
| `/etc/sudoers.d/ledmatrix_web`, `ledmatrix_wifi` | `root` | `440` | | | `/etc/sudoers.d/ledmatrix_web`, `ledmatrix_wifi` | `root` | `440` | |
| `/usr/local/sbin/ledmatrix-refresh-units` | `root:root` | `755` | Copy of `scripts/install/ledmatrix_refresh_units.py`, installed by `install_service.sh`. Outside the project so the web user cannot edit what sudo runs |
| `/var/lib/ledmatrix/unit-backup/` | `root` | `700` | The units the last refresh replaced, for the automatic update's rollback |
| `/etc/systemd/system/ledmatrix*.service`, `.path` | `root:root` | `644` | Readable so the web interface can compare them with the templates after an update |
What keeps it that way at runtime: What keeps it that way at runtime:
@@ -86,6 +89,20 @@ password:
- `journalctl -u ledmatrix.service *`, `-u ledmatrix *`, `-t ledmatrix *`, - `journalctl -u ledmatrix.service *`, `-u ledmatrix *`, `-t ledmatrix *`,
tagged `NOEXEC`: journalctl opens a pager on a terminal, and a shell tagged `NOEXEC`: journalctl opens a pager on a terminal, and a shell
escape from that pager would be a root shell escape from that pager would be a root shell
- `/usr/local/sbin/ledmatrix-refresh-units ""` and
`/usr/local/sbin/ledmatrix-refresh-units --restore` — exactly these two
command lines (`""` means "no arguments"). After an update the first
installs the systemd units whose templates changed and runs
`systemctl daemon-reload`; the automatic update's rollback runs the second
to put the previous units back. The helper takes nothing from the caller:
the project folder and the web user come from the installed, root-owned
`ledmatrix.service` and `ledmatrix-web.service`. It only replaces the four
units `install_service.sh` installs, only if they are already installed,
and refuses a template that would change a unit's `User=` (root for the
display, the web user for the rest) or `WorkingDirectory=`, or that is a
symlink, not a regular file, or over 64 KB. It grants nothing new: the
templates are files the web user can edit, but so is `run.py`, which the
display service already runs as root.
### `/etc/sudoers.d/ledmatrix_wifi` ### `/etc/sudoers.d/ledmatrix_wifi`
@@ -136,7 +153,9 @@ directory.
| `safe_plugin_rm.sh`, `safe_pip_install.sh` | — | Called by the web interface through sudo | Not for manual use | | `safe_plugin_rm.sh`, `safe_pip_install.sh` | — | Called by the web interface through sudo | Not for manual use |
To reinstall the sudoers rules, run To reinstall the sudoers rules, run
`./scripts/install/configure_web_sudo.sh` (web rules) or `./scripts/install/configure_web_sudo.sh` (web rules; the
`ledmatrix-refresh-units` rules also need the helper itself, which
`sudo ./scripts/install/install_service.sh` installs) or
`./scripts/install/configure_wifi_permissions.sh` (WiFi rules and polkit) as `./scripts/install/configure_wifi_permissions.sh` (WiFi rules and polkit) as
the web user, not with `sudo`. the web user, not with `sudo`.
+49 -16
View File
@@ -159,14 +159,16 @@ there an unchecked checkbox — which the browser omits — is saved as
} }
``` ```
`restart_required` is always true here: display hardware, rotation, `restart_required` is true when the save changed a setting that takes
durations and general settings take effect when the display restarts, and effect when the display restarts: display hardware, rotation order,
the web UI shows its restart banner on the flag. (Plugin sections saved timezone, general settings and the rest. The web UI shows its restart banner
through this route reach the running plugin live, like on the flag. It is false when the save changed only what the running display
`POST /plugins/config`.) applies by itself, or nothing: `brightness`, the per-mode durations
(`duration__<mode>`, `display.display_durations`) and plugin sections, which
reach the running plugin live, like `POST /plugins/config`.
A saved `brightness` is the exception: it reaches the panel without a A saved `brightness` reaches the panel without a restart. The route also
restart. The route also sends it to the running display over the control sends it to the running display over the control
socket (`brightness.set`), which puts it on the panel at once, and the socket (`brightness.set`), which puts it on the panel at once, and the
response adds `"brightness_transport": "socket"`. Otherwise it is response adds `"brightness_transport": "socket"`. Otherwise it is
`"config"`, with `brightness_socket_error` giving the reason, and the `"config"`, with `brightness_socket_error` giving the reason, and the
@@ -246,7 +248,10 @@ Replace the schedule configuration.
``` ```
A day whose `<day>_enabled` key is absent counts as enabled, with default A day whose `<day>_enabled` key is absent counts as enabled, with default
times `07:00`-`23:00`. At least one day must be enabled. times `07:00`-`23:00`. An enabled schedule needs at least one day enabled; a
disabled one (`"enabled": false`) may have every day off, as
`config.template.json` ships it. A day that is off keeps the times sent for
it, when they are valid `HH:MM`.
**Response**: **Response**:
```json ```json
@@ -331,12 +336,23 @@ by the display process (stale after 120 seconds).
"data": { "data": {
"mode": "nfl_live", "mode": "nfl_live",
"plugin_id": "football-scoreboard", "plugin_id": "football-scoreboard",
"last_updated": 1234567890.123 "last_updated": 1234567890.123,
"source": "socket"
} }
} }
``` ```
When nothing has been published, every field is `null`. When nothing has been published, every field is `null`. `source` is
`socket` when the answer came from the display's state stream over the
control socket ([IPC_CONTROL_SOCKET.md](IPC_CONTROL_SOCKET.md)), and `cache`
when it came from the `display_current_state` cache key (no socket: the
display is stopped or older, or this is Windows). A display whose render
loop has not refreshed its state for 120 seconds is reported with every
field `null`, either way. So is a stopped display: when the socket does not
answer and the render loop's heartbeat
(`/run/ledmatrix/display-heartbeat.json`) is absent, stale or from a process
that is gone, the cache's last entry is not used. A display still beating
without a socket, Windows, or a socket switched off reads the cache.
### List Display Modes ### List Display Modes
@@ -413,11 +429,16 @@ Get the current on-demand display state.
"returncode": 0, "returncode": 0,
"stdout": "active", "stdout": "active",
"stderr": "" "stderr": ""
} },
"source": "socket"
} }
} }
``` ```
`source` is `socket` (the display's state stream, with `remaining` worked
out at the time of the request) or `cache` (the `display_on_demand_state`
cache key).
With no on-demand request, `state` is With no on-demand request, `state` is
`{"active": false, "status": "idle", "last_updated": null}`. `{"active": false, "status": "idle", "last_updated": null}`.
@@ -552,7 +573,8 @@ List all installed plugins with their status and metadata.
"published_at": 1790000030.0, "published_at": 1790000030.0,
"age_seconds": 12.4, "age_seconds": 12.4,
"stale_after": 180.0, "stale_after": 180.0,
"heartbeat_age_seconds": 2.1 "heartbeat_age_seconds": 2.1,
"source": "socket"
} }
} }
} }
@@ -581,7 +603,10 @@ the display is hung or died), `stopped` (the display shut down) or `unknown`
(nothing published yet). Unless it is `live`, every one of those fields is (nothing published yet). Unless it is `live`, every one of those fields is
`null`. `heartbeat_age_seconds` is the heartbeat's age when it was taken into `null`. `heartbeat_age_seconds` is the heartbeat's age when it was taken into
account, `null` otherwise (no heartbeat, as on the dev server, or one from account, `null` otherwise (no heartbeat, as on the dev server, or one from
another process). Health and metrics are at [`/plugins/health`](#get-plugin-health) another process). `runtime.source` is `socket` when the snapshot and the
heartbeat age came from the display's state stream over the control socket,
and `cache` when they came from the `plugin_runtime_snapshot` cache key and
the heartbeat file; the rules above are the same for both. Health and metrics are at [`/plugins/health`](#get-plugin-health)
and `/plugins/metrics`. and `/plugins/metrics`.
`vegas_participation` is what Vegas mode does with the plugin: `"scroll"`, `vegas_participation` is what Vegas mode does with the plugin: `"scroll"`,
@@ -2265,7 +2290,14 @@ display snapshot. `data.status` is `healthy` or `degraded`, with
(with `heartbeat_age_seconds`), `stalled` (no heartbeat for 60s: the panel is (with `heartbeat_age_seconds`), `stalled` (no heartbeat for 60s: the panel is
frozen even if the service is active; the status turns `degraded`), or frozen even if the service is active; the status turns `degraded`), or
`not_reported` when the display writes none (not started yet, the dev server, `not_reported` when the display writes none (not started yet, the dev server,
Windows), which does not affect the status. Windows), which does not affect the status, or `stopped` (with `source:
"service"`) when the display service is not active, the control socket does
not answer and there is no live heartbeat; the status then turns
`degraded`. A platform with no control socket (Windows) or a socket switched
off never reports `stopped`. Its `source` is `socket` when the
age came from the display's state stream over the control socket (measured
in memory by the display) and `heartbeat_file` when it came from
`/run/ledmatrix/display-heartbeat.json`.
Open even when the web login is on, for uptime monitors; a caller that is not Open even when the web login is on, for uptime monitors; a caller that is not
logged in (and has no token) then gets only `{"status": "success", "data": logged in (and has no token) then gets only `{"status": "success", "data":
@@ -2325,8 +2357,9 @@ Replace the dim schedule. `dim_brightness` is 0-100 (default 30). In
`per-day` mode the days can be sent either as the `days` object that GET `per-day` mode the days can be sent either as the `days` object that GET
returns, or as the web form's flat fields (`monday_enabled`, returns, or as the web form's flat fields (`monday_enabled`,
`monday_start`, `monday_end`, ...). A day that is not sent counts as `monday_start`, `monday_end`, ...). A day that is not sent counts as
enabled with default times `20:00`-`07:00`; at least one day must be enabled with default times `20:00`-`07:00`. As for the schedule above, an
enabled. enabled dim schedule needs at least one day enabled and a disabled one may
have every day off.
--- ---
+3 -2
View File
@@ -181,13 +181,14 @@ that the harness patches in today.
- dynamic duration (cycle complete, plugin cap, global cap) - dynamic duration (cycle complete, plugin cap, global cap)
- live priority taking over and handing back; live round-robin - live priority taking over and handing back; live round-robin
- on-demand start/stop/expiry; pinned on-demand; a session resumed after - on-demand start/stop/expiry; pinned on-demand; a session resumed after
a restart a restart, and one that cannot resume (its plugin did not load); a
request naming a live mode the plugin's live check would drop
- schedule off and dim, with an on-demand override during downtime - schedule off and dim, with an on-demand override during downtime
- WiFi notice; sync follower - WiFi notice; sync follower
- Vegas, with and without `live_in_ticker` - Vegas, with and without `live_in_ticker`
- Each trace row is `[start, mode, duration, exit_reason, frames, - Each trace row is `[start, mode, duration, exit_reason, frames,
force_clear]`. The exit reason is the event that decided what came next. force_clear]`. The exit reason is the event that decided what came next.
- All 16 tests run in under a second. The goldens were generated from - All 18 tests run in under a second. The goldens were generated from
main's `run()` before any code moved. main's `run()` before any code moved.
- Vegas uses `FakeVegas`, which implements only the contract the controller - Vegas uses `FakeVegas`, which implements only the contract the controller
depends on: `run_iteration()` returns True after its duration and False depends on: `run_iteration()` returns True after its duration and False
+39 -2
View File
@@ -77,6 +77,31 @@ The Vegas **Scroll Speed** slider in the web UI shows the same thing live: a
line under it says what your speed will run as on this panel, and links to the line under it says what your speed will run as on this panel, and links to the
nearest smooth speeds. nearest smooth speeds.
### A panel that cannot reach its cap
Speeds are solved against `limit_refresh_rate_hz`, the configured cap, but a
cap is only a ceiling: a long chain, a high `pwm_bits` or a big
`gpio_slowdown` can leave the panel below it. One Pi 4 driving 2×128×64 on
`adafruit-hat-pwm` with `pwm_bits 9` and `gpio_slowdown 5` measured
107.6–113.1 Hz under a 120 Hz cap. Frames still move whole pixels, but
every scroll runs that much slower than configured (60 px/s ran at 55 px/s),
and the smooth speeds are the cap's rather than the panel's.
The display measures the real rate from its own frames. About a minute
into scrolling, a panel more than 3% short of its cap is logged once:
```
WARNING - src.common.frame_timing - The panel refreshes at about 113 Hz, below
the 120 Hz that scroll speeds are planned for ... Set Limit Refresh Rate to
100 Hz (web UI, Display tab), which this panel can hold, and restart.
```
The Display tab says the same under **Limit Refresh Rate**, with a button
that fills in the suggested cap (`GET /api/v3/config/refresh-rate`). The
suggestion is a multiple of 10 at least 5% under the measurement, because
an uncapped panel drifts and the measurement is the fast end of it. A cap the
panel holds also stops the drift.
### How a slow speed stays crisp ### How a slow speed stays crisp
`SwapOnVSync(canvas, framerate_fraction)` holds each frame for N panel `SwapOnVSync(canvas, framerate_fraction)` holds each frame for N panel
@@ -246,8 +271,16 @@ mean exactly 10 ms, so a ticker stalling on half its frames still averages to a
healthy 100 fps. The stats line reports the tail for that reason — read the healthy 100 fps. The stats line reports the tail for that reason — read the
percentiles, not the fps. percentiles, not the fps.
Every scroller emits one line every 5 seconds covering *every* frame in that Every scroller summarises each 5-second window, covering *every* frame in it,
window, tagged with the plugin it came from: in one line tagged with the plugin it came from. At the default log level the
line reaches the journal only when it is worth reading: a **degraded** window
(frame rate below 90% of the rate the window was locked to, i.e. 1 / its own
median -- the same 0.9 Vegas's `Vegas FPS` line uses -- or more than 1% of its
frames stalled), the first window after one (the recovery), and otherwise once
every 5 minutes per scroller as a heartbeat, so silence means stopped rather
than fine. Every window is logged at DEBUG: to see them all, run the display
with `-d` or `LEDMATRIX_DEBUG=true` (see
[CONFIG_DEBUGGING.md](CONFIG_DEBUGGING.md#enable-debug-logging)).
```bash ```bash
journalctl -u ledmatrix --since "-10min" --no-pager | grep "Scroll frame stats" journalctl -u ledmatrix --since "-10min" --no-pager | grep "Scroll frame stats"
@@ -286,6 +319,10 @@ journalctl -u ledmatrix --since "-3h" --no-pager | grep "Scroll frame stats" \
| sort -k7 -rn | sort -k7 -rn
``` ```
At the default log level that ranks the windows the journal kept -- the
degraded ones, recoveries and heartbeats -- so it over-weights bad windows;
rank a debug run for an unbiased average, or soak the rig (below).
The `$2 < 1000` guard drops windows whose median is a whole second or more. The `$2 < 1000` guard drops windows whose median is a whole second or more.
Those are not frames. Until the idle-gap fix in `log_frame_rate()`, the first Those are not frames. Until the idle-gap fix in `log_frame_rate()`, the first
frame of every scroll was timed against the end of the *previous* scroll, so frame of every scroll was timed against the end of the *previous* scroll, so
+91 -3
View File
@@ -84,6 +84,43 @@ python3 web_interface/start.py
### Installation & Build Issues ### Installation & Build Issues
#### "This version of Raspberry Pi OS is not supported"
LEDMatrix installs on Raspberry Pi OS Lite **Trixie** (Debian 13, Python
3.13) or **Bookworm** (Debian 12, Python 3.11). The installer checks
`/etc/os-release` before it changes anything and stops on anything else.
**Check what you have:**
```bash
grep -E '^(PRETTY_NAME|VERSION_ID)=' /etc/os-release
python3 --version
```
**Solutions:**
- `VERSION_ID="11"` (Bullseye) or older: flash a new card with Raspberry Pi
Imager, choosing Raspberry Pi OS Lite (64-bit). Trixie is recommended;
Bookworm (Legacy) also works. An in-place upgrade from Bullseye is not
supported by Raspberry Pi and is not worth the risk.
- "Desktop environment detected": use the Lite image, not the desktop one.
- "python3 is Python 3.x; LEDMatrix needs Python 3.11 or newer": something
has replaced the system `python3`. Point it back at the OS's own Python
(`/usr/bin/python3` should be 3.11 on Bookworm, 3.13 on Trixie).
- `sudo bash scripts/check_system_compatibility.sh` runs the same checks
without installing anything.
#### "This Pi manages its network with dhcpcd, not NetworkManager"
A warning, not an error: the install carries on and the display works. But
choosing a WiFi network from the web page and the `LEDMatrix-Setup` hotspot
both need NetworkManager, the default on Bookworm and Trixie. It appears
when dhcpcd was selected in `raspi-config`. Switch back with a keyboard and
screen attached (or over Ethernet), since the WiFi connection drops briefly:
```bash
sudo raspi-config # Advanced Options -> Network Config -> NetworkManager
sudo reboot
```
#### Step 6 fails: "Failed building wheel for rgbmatrix" #### Step 6 fails: "Failed building wheel for rgbmatrix"
**Symptoms:** **Symptoms:**
@@ -328,6 +365,49 @@ commit, then switches to releases on its own.
3. **Local changes after a channel switch:** edits that no longer fit the new 3. **Local changes after a channel switch:** edits that no longer fit the new
version are kept in the git stash rather than lost; `git stash list` version are kept in the git stash rather than lost; `git stash list`
shows them as "LEDMatrix autostash before update". shows them as "LEDMatrix autostash before update".
4. **A new install is on a release, not `main`.** The one-shot installer
checks out the newest release. For the newest code instead, install with
`LEDMATRIX_CHANNEL=beta`:
```bash
curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | LEDMATRIX_CHANNEL=beta bash
```
---
#### Issue: "service settings ... are not applied yet" after an update
**Symptoms:**
- Update Code's message, or the web interface log, says an update changes
service settings that are not applied yet, and to run the installer
- The display logs `ledmatrix.service differs from systemd/ledmatrix.service`
at startup
**Explanation:** updates install the systemd units a new version changes
through the root helper `/usr/local/sbin/ledmatrix-refresh-units`, which the
installer sets up and grants to the web user in
`/etc/sudoers.d/ledmatrix_web`. A device installed before that has neither,
so the new unit settings (for example the display's watchdog) wait for a
reinstall. The update itself is fine.
**Solution:** re-run the installer once, as root:
```bash
cd ~/LEDMatrix
sudo ./first_time_install.sh
# or, lighter: install the units and helper, then the sudo rules
sudo ./scripts/install/install_service.sh
./scripts/install/configure_web_sudo.sh
```
Check it worked:
```bash
ls -l /usr/local/sbin/ledmatrix-refresh-units # root root, rwxr-xr-x
sudo -l | grep ledmatrix-refresh-units # the two rules
```
A message that the helper **refused** a unit (`refusing to install it`)
means a template in `systemd/` was edited so that it would run as another
account or from another folder. The message names the template. Look at
what changed with `git diff -- systemd/`, save any edit you want to keep,
then restore only that file, for example
`git checkout -- systemd/ledmatrix-web.service`.
--- ---
@@ -364,9 +444,16 @@ commit, then switches to releases on its own.
5. **Check required services:** 5. **Check required services:**
```bash ```bash
systemctl is-active NetworkManager # must say "active"
sudo systemctl status hostapd sudo systemctl status hostapd
sudo systemctl status dnsmasq sudo systemctl status dnsmasq
``` ```
On a fresh install `hostapd` shows as **masked**. That is expected, on
Bookworm and Trixie alike: Debian's hostapd package masks the service
when it is installed without a configuration, so the hotspot is brought
up through NetworkManager instead (look for `nmcli hotspot fallback` in
`journalctl -u ledmatrix-wifi-monitor`). If NetworkManager is not
active, see "This Pi manages its network with dhcpcd" above.
6. **Manually enable AP mode:** 6. **Manually enable AP mode:**
```bash ```bash
@@ -590,9 +677,10 @@ stack into the log, so it says which plugin was stuck.
apart, so a plugin that hangs on every start does not restart the display apart, so a plugin that hangs on every start does not restart the display
hundreds of times an hour. hundreds of times an hour.
4. **Is the watchdog installed?** Installs from before it keep their old unit 4. **Is the watchdog installed?** Updates install new unit settings once the
until the installer is re-run (a startup warning says the unit differs installer has set up `ledmatrix-refresh-units`; installs from before that
from its template): keep their old unit until the installer is re-run (a startup warning says
the unit differs from its template):
```bash ```bash
systemctl show -p WatchdogUSec ledmatrix # 2min once running; 0 = not installed systemctl show -p WatchdogUSec ledmatrix # 2min once running; 0 = not installed
sudo ./scripts/install/install_service.sh sudo ./scripts/install/install_service.sh
+160 -47
View File
@@ -47,38 +47,51 @@ if echo "${DEVICE_MODEL:-}" | grep -qi "Raspberry Pi 5"; then
echo "Raspberry Pi 5 detected — will verify RP1 library support." echo "Raspberry Pi 5 detected — will verify RP1 library support."
fi fi
# Check OS version - must be Raspberry Pi OS Lite (Trixie) # Check OS version - must be Raspberry Pi OS Lite, Bookworm or Trixie.
# The rules live in scripts/install/lib_os.sh, shared with
# scripts/check_system_compatibility.sh.
echo "" echo ""
echo "Checking operating system requirements..." echo "Checking operating system requirements..."
echo "----------------------------------------" echo "----------------------------------------"
OS_CHECK_FAILED=0 OS_CHECK_FAILED=0
OS_RELEASE=""
if [ -f /etc/os-release ]; then OS_LIB="$(cd "$(dirname "$0")" && pwd)/scripts/install/lib_os.sh"
. /etc/os-release if [ ! -f "$OS_LIB" ]; then
echo "Detected OS: $PRETTY_NAME" echo "✗ ERROR: $OS_LIB is missing, so the operating system cannot be checked."
echo "Version ID: ${VERSION_ID:-unknown}" echo " Your LEDMatrix download is incomplete. Download it again and re-run this script:"
echo " git clone https://github.com/ChuckBuilds/LEDMatrix.git"
# Check if it's Raspberry Pi OS or Debian exit 1
if [[ "$ID" != "raspbian" ]] && [[ "$ID" != "debian" ]]; then fi
echo "✗ ERROR: This script requires Raspberry Pi OS (raspbian/debian)" # shellcheck source=scripts/install/lib_os.sh
echo " Detected OS ID: $ID" . "$OS_LIB"
OS_CHECK_FAILED=1
fi if [ -r "$LM_OS_RELEASE_FILE" ]; then
echo "Detected OS: $(lm_os_field PRETTY_NAME)"
# Check if it's Debian 13 (Trixie) OS_VERSION_ID=$(lm_os_field VERSION_ID)
if [ "${VERSION_ID:-0}" != "13" ]; then echo "Version ID: ${OS_VERSION_ID:-unknown}"
echo "✗ ERROR: This script requires Raspberry Pi OS Lite (Trixie) - Debian 13"
echo " Detected version: ${VERSION_ID:-unknown}" if OS_RELEASE=$(lm_os_release); then
echo " Please upgrade to Raspberry Pi OS Lite (Trixie) before continuing" echo "✓ $(lm_release_label "$OS_RELEASE") detected"
OS_CHECK_FAILED=1
else else
echo "✓ Debian 13 (Trixie) detected" OS_ID=$(lm_os_field ID)
if [[ "$OS_ID" != "raspbian" ]] && [[ "$OS_ID" != "debian" ]]; then
echo "✗ ERROR: This script requires Raspberry Pi OS (raspbian/debian)"
echo " Detected OS ID: ${OS_ID:-unknown}"
else
echo "✗ ERROR: This version of Raspberry Pi OS is not supported"
echo " Detected version: ${OS_VERSION_ID:-unknown}"
echo " Supported: Trixie (Debian 13) and Bookworm (Debian 12)"
fi
OS_CHECK_FAILED=1
fi fi
# Check if it's the Lite version (no desktop environment) # Check if it's the Lite version (no desktop environment)
# Check for desktop packages or desktop services # Check for desktop packages or desktop services
DESKTOP_DETECTED=0 DESKTOP_DETECTED=0
if dpkg -l | grep -qE "^ii.*raspberrypi-ui-mods|^ii.*lxde|^ii.*xfce|^ii.*gnome|^ii.*kde"; then # grep without -q: -q exits at the first match, dpkg then dies of SIGPIPE,
# and pipefail turns a found desktop into "not found".
if dpkg -l | grep -E "^ii.*raspberrypi-ui-mods|^ii.*lxde|^ii.*xfce|^ii.*gnome|^ii.*kde" >/dev/null; then
DESKTOP_DETECTED=1 DESKTOP_DETECTED=1
fi fi
if systemctl list-units --type=service --state=running 2>/dev/null | grep -qE "lightdm|gdm3|sddm|lxdm"; then if systemctl list-units --type=service --state=running 2>/dev/null | grep -qE "lightdm|gdm3|sddm|lxdm"; then
@@ -96,23 +109,52 @@ if [ -f /etc/os-release ]; then
echo "✓ Lite version confirmed (no desktop environment)" echo "✓ Lite version confirmed (no desktop environment)"
fi fi
else else
echo "✗ ERROR: Could not detect OS version (/etc/os-release not found)" echo "✗ ERROR: Could not detect OS version ($LM_OS_RELEASE_FILE not found)"
OS_CHECK_FAILED=1 OS_CHECK_FAILED=1
fi fi
# Python: whatever python3 the release ships (3.11 on Bookworm, 3.13 on
# Trixie). Checked only when python3 is already there -- Step 1 installs it
# otherwise, and on a supported release that brings the release's own version.
if [ "$OS_CHECK_FAILED" -eq 0 ]; then
if PYTHON3_VERSION=$(lm_python_version); then
case "$(lm_python_check "$PYTHON3_VERSION")" in
ok)
echo "✓ Python $PYTHON3_VERSION detected"
;;
too-old)
echo "✗ ERROR: python3 is Python $PYTHON3_VERSION; LEDMatrix needs Python 3.$LM_PYTHON_MIN_MINOR or newer"
echo " $(lm_release_label "$OS_RELEASE") ships Python $(lm_release_python "$OS_RELEASE"). Something on this"
echo " system has changed which Python 'python3' runs; point it back at the system Python."
OS_CHECK_FAILED=1
;;
*)
echo "⚠ python3 is Python $PYTHON3_VERSION, which LEDMatrix has not been tested with"
echo " (tested: 3.$LM_PYTHON_MIN_MINOR to 3.$LM_PYTHON_MAX_MINOR). Continuing anyway."
;;
esac
else
echo "python3 not found yet; Step 1 installs it."
fi
fi
if [ "$OS_CHECK_FAILED" -eq 1 ]; then if [ "$OS_CHECK_FAILED" -eq 1 ]; then
echo "" echo ""
echo "Installation cannot continue. Please install Raspberry Pi OS Lite (Trixie) and try again." echo "Installation cannot continue."
echo "" lm_print_supported_os_help
echo "To install Raspberry Pi OS Lite (Trixie):"
echo " 1. Download from: https://www.raspberrypi.com/software/operating-systems/"
echo " 2. Select 'Raspberry Pi OS Lite (64-bit)' with Debian 13 (Trixie)"
echo " 3. Flash to SD card using Raspberry Pi Imager"
echo " 4. Boot and run this script again"
exit 1 exit 1
fi fi
echo "✓ OS requirements met" echo "✓ OS requirements met"
# WiFi setup (the web page's WiFi tab and the LEDMatrix-Setup hotspot) needs
# NetworkManager. Both releases use it by default; say so plainly if this Pi
# does not, but carry on -- the display itself does not depend on it.
case "$(lm_network_stack)" in
networkmanager) echo "✓ NetworkManager is managing the network" ;;
dhcpcd) lm_print_dhcpcd_advice ;;
*) echo "⚠ Could not tell which service manages the network; WiFi setup from the web page needs NetworkManager" ;;
esac
echo "" echo ""
# The user who ran the installer: SUDO_USER once we are running under sudo # The user who ran the installer: SUDO_USER once we are running under sudo
@@ -190,6 +232,51 @@ _sync_rgb_submodule() {
fi fi
return 0 return 0
} }
# LEDMatrix's own changes to the library live in patches/rpi-rgb-led-matrix/ and
# are applied only for the build: _apply_rgb_patches before it, _revert_rgb_patches
# after it, success or not. The checkout is left exactly as it was, so `git pull`
# and _sync_rgb_submodule never meet local modifications in the submodule.
# A patch that no longer applies (a submodule bump, a hand-edited checkout) is
# reported and skipped -- the unpatched library still builds and works, so it is
# never fatal. One that is already applied is left alone and not reverted.
_RGB_APPLIED_PATCHES=()
_apply_rgb_patches() {
local sub="$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master"
local dir="$PROJECT_ROOT_DIR/patches/rpi-rgb-led-matrix" patch name
_RGB_APPLIED_PATCHES=()
[ -d "$dir" ] || return 0
for patch in "$dir"/*.patch; do
[ -f "$patch" ] || continue
name=$(basename "$patch")
if _git_as_repo_owner -C "$sub" apply --check "$patch" >/dev/null 2>&1; then
if _git_as_repo_owner -C "$sub" apply "$patch"; then
_RGB_APPLIED_PATCHES+=("$patch")
echo "Applied library patch $name"
else
echo "⚠ Could not apply library patch $name; building without it"
fi
elif _git_as_repo_owner -C "$sub" apply --reverse --check "$patch" >/dev/null 2>&1; then
echo "Library patch $name is already applied"
else
echo "⚠ Library patch $name does not apply to this checkout; building without it"
fi
done
return 0
}
_revert_rgb_patches() {
local sub="$PROJECT_ROOT_DIR/rpi-rgb-led-matrix-master" i
# Last applied first, in case two patches touch the same file.
for ((i = ${#_RGB_APPLIED_PATCHES[@]} - 1; i >= 0; i--)); do
if ! _git_as_repo_owner -C "$sub" apply --reverse "${_RGB_APPLIED_PATCHES[i]}"; then
echo "⚠ Could not revert $(basename "${_RGB_APPLIED_PATCHES[i]}"); restore the checkout with: git -C $sub checkout -- ."
fi
done
_RGB_APPLIED_PATCHES=()
return 0
}
# --- end rpi-rgb-led-matrix checkout helpers --------------------------------- # --- end rpi-rgb-led-matrix checkout helpers ---------------------------------
# Determine the Project Root Directory (where this script is located) # Determine the Project Root Directory (where this script is located)
@@ -223,6 +310,8 @@ SKIP_SWAP=${LEDMATRIX_SKIP_SWAP:-0}
BUILD_JOBS_OVERRIDE=${LEDMATRIX_BUILD_JOBS:-} BUILD_JOBS_OVERRIDE=${LEDMATRIX_BUILD_JOBS:-}
# Weekly automatic updates: 1 on, 0 off, empty = ask (interactive) or leave as is. # Weekly automatic updates: 1 on, 0 off, empty = ask (interactive) or leave as is.
AUTO_UPDATE=${LEDMATRIX_AUTO_UPDATE:-} AUTO_UPDATE=${LEDMATRIX_AUTO_UPDATE:-}
# Update channel written to config.json: stable, beta, or empty = leave as is.
UPDATE_CHANNEL=$(printf '%s' "${LEDMATRIX_CHANNEL:-}" | tr '[:upper:]' '[:lower:]')
usage() { usage() {
cat <<USAGE cat <<USAGE
@@ -240,12 +329,18 @@ Options:
--enable-auto-update Turn on weekly automatic updates (with health --enable-auto-update Turn on weekly automatic updates (with health
check and automatic rollback) check and automatic rollback)
--no-auto-update Leave weekly automatic updates off --no-auto-update Leave weekly automatic updates off
--beta Follow main, the newest code (the beta update
channel). Without it, updates follow releases
(stable). It sets the channel; it does not move
this checkout -- the one-shot installer picks the
version, and so does the next update.
-h, --help Show this help message and exit -h, --help Show this help message and exit
Environment variables (same effect as flags): Environment variables (same effect as flags):
LEDMATRIX_ASSUME_YES=1, RPI_RGB_FORCE_REBUILD=1, LEDMATRIX_SKIP_SOUND=1, LEDMATRIX_ASSUME_YES=1, RPI_RGB_FORCE_REBUILD=1, LEDMATRIX_SKIP_SOUND=1,
LEDMATRIX_SKIP_PERF=1, LEDMATRIX_SKIP_REBOOT_PROMPT=1, LEDMATRIX_SKIP_PERF=1, LEDMATRIX_SKIP_REBOOT_PROMPT=1,
LEDMATRIX_SKIP_SWAP=1, LEDMATRIX_BUILD_JOBS=N, LEDMATRIX_AUTO_UPDATE=1|0 LEDMATRIX_SKIP_SWAP=1, LEDMATRIX_BUILD_JOBS=N, LEDMATRIX_AUTO_UPDATE=1|0,
LEDMATRIX_CHANNEL=stable|beta
Low-memory devices: Low-memory devices:
On a Pi with under 2GB of RAM the C++ build is limited to fewer parallel On a Pi with under 2GB of RAM the C++ build is limited to fewer parallel
@@ -265,6 +360,7 @@ while [ $# -gt 0 ]; do
--skip-swap) SKIP_SWAP=1 ;; --skip-swap) SKIP_SWAP=1 ;;
--enable-auto-update) AUTO_UPDATE=1 ;; --enable-auto-update) AUTO_UPDATE=1 ;;
--no-auto-update) AUTO_UPDATE=0 ;; --no-auto-update) AUTO_UPDATE=0 ;;
--beta) UPDATE_CHANNEL=beta ;;
--build-jobs) --build-jobs)
shift shift
if [ $# -eq 0 ]; then echo "--build-jobs requires a number"; usage; exit 1; fi if [ $# -eq 0 ]; then echo "--build-jobs requires a number"; usage; exit 1; fi
@@ -291,10 +387,12 @@ else
lm_remove_build_swap() { return 0; } lm_remove_build_swap() { return 0; }
fi fi
# Remove the temporary build swapfile no matter how the script ends. Step 6 # Remove the temporary build swapfile, and take any library patches back out
# tears it down itself; this is the backstop for the error path, since # of the submodule, no matter how the script ends. Step 6 does both itself;
# on_error ends in `exit` and EXIT traps still run. # this is the backstop for the error path (on_error ends in `exit` and EXIT
trap 'lm_remove_build_swap' EXIT # traps still run) and for an interrupted build. _revert_rgb_patches only
# touches patches it applied, so running it twice is harmless.
trap 'lm_remove_build_swap; _revert_rgb_patches' EXIT
# Helpers # Helpers
retry() { retry() {
@@ -873,15 +971,25 @@ if [ -z "$AUTO_UPDATE" ] && [ "$ASSUME_YES" != "1" ] && [ -t 0 ]; then
echo echo
if [[ $REPLY =~ ^[Yy]$ ]]; then AUTO_UPDATE=1; else AUTO_UPDATE=0; fi if [[ $REPLY =~ ^[Yy]$ ]]; then AUTO_UPDATE=1; else AUTO_UPDATE=0; fi
fi fi
if [ "$AUTO_UPDATE" = "1" ] || [ "$AUTO_UPDATE" = "0" ]; then case "$UPDATE_CHANNEL" in
if python3 - "$PROJECT_ROOT_DIR/config/config.json" "$AUTO_UPDATE" <<'PY' stable|beta|"") ;;
*) echo "⚠ LEDMATRIX_CHANNEL=$UPDATE_CHANNEL is not stable or beta; leaving the update channel as it is"
UPDATE_CHANNEL="" ;;
esac
# The update channel, likewise only when asked for (--beta / LEDMATRIX_CHANNEL).
if [ "$AUTO_UPDATE" = "1" ] || [ "$AUTO_UPDATE" = "0" ] || [ -n "$UPDATE_CHANNEL" ]; then
if python3 - "$PROJECT_ROOT_DIR/config/config.json" "$AUTO_UPDATE" "$UPDATE_CHANNEL" <<'PY'
import json, os, sys, tempfile import json, os, sys, tempfile
path, enabled = sys.argv[1], sys.argv[2] == "1" path, enabled = sys.argv[1], sys.argv[2]
channel = sys.argv[3] if len(sys.argv) > 3 else ""
with open(path, encoding="utf-8") as f: with open(path, encoding="utf-8") as f:
config = json.load(f) config = json.load(f)
if not isinstance(config.get("auto_update"), dict): if not isinstance(config.get("auto_update"), dict):
config["auto_update"] = {} config["auto_update"] = {}
config["auto_update"]["enabled"] = enabled if enabled in ("0", "1"):
config["auto_update"]["enabled"] = enabled == "1"
if channel:
config["auto_update"]["channel"] = channel
# Written beside the original and swapped in whole: the display service's # Written beside the original and swapped in whole: the display service's
# config watcher may be running and must never read a half-written file. # config watcher may be running and must never read a half-written file.
original = os.stat(path) original = os.stat(path)
@@ -902,9 +1010,11 @@ except BaseException:
raise raise
PY PY
then then
if [ "$AUTO_UPDATE" = "1" ]; then echo "✓ Weekly automatic updates enabled"; else echo "✓ Weekly automatic updates off"; fi if [ "$AUTO_UPDATE" = "1" ]; then echo "✓ Weekly automatic updates enabled"
elif [ "$AUTO_UPDATE" = "0" ]; then echo "✓ Weekly automatic updates off"; fi
if [ -n "$UPDATE_CHANNEL" ]; then echo "✓ Update channel: $UPDATE_CHANNEL"; fi
else else
echo "⚠ Could not set auto_update in config/config.json; turn it on from the General tab instead" echo "⚠ Could not set auto_update in config/config.json; set it from the General tab instead"
fi fi
fi fi
@@ -1226,9 +1336,11 @@ else
fi fi
BUILD_OUTPUT=$(mktemp) BUILD_OUTPUT=$(mktemp)
BUILD_SUCCESS=false BUILD_SUCCESS=false
_apply_rgb_patches
if run_rgbmatrix_build "$BUILD_JOBS" "$BUILD_OUTPUT"; then if run_rgbmatrix_build "$BUILD_JOBS" "$BUILD_OUTPUT"; then
BUILD_SUCCESS=true BUILD_SUCCESS=true
fi fi
_revert_rgb_patches
cat "$BUILD_OUTPUT" >> "$LOG_FILE" cat "$BUILD_OUTPUT" >> "$LOG_FILE"
if [ "$BUILD_SUCCESS" != true ]; then if [ "$BUILD_SUCCESS" != true ]; then
print_rgbmatrix_build_failure "$BUILD_OUTPUT" print_rgbmatrix_build_failure "$BUILD_OUTPUT"
@@ -1349,15 +1461,16 @@ if ! command -v setcap >/dev/null 2>&1; then
echo "⚠ setcap not found, skipping capability configuration" echo "⚠ setcap not found, skipping capability configuration"
echo " Install libcap2-bin if you need hardware timing capabilities" echo " Install libcap2-bin if you need hardware timing capabilities"
else else
# Find the Python binary and resolve symlinks to get the real binary # The binary the services run (ExecStart=/usr/bin/python3), symlinks
# resolved: python3.11 on Bookworm, python3.13 on Trixie. This used to
# prefer /usr/bin/python3.13 whenever it existed, which would set the
# capability on an interpreter the services never run if python3 pointed
# elsewhere.
PYTHON_BIN="" PYTHON_BIN=""
PYTHON_VER="" PYTHON_VER=""
if [ -f "/usr/bin/python3.13" ]; then if [ -f "/usr/bin/python3" ]; then
PYTHON_BIN=$(readlink -f /usr/bin/python3.13)
PYTHON_VER="3.13"
elif [ -f "/usr/bin/python3" ]; then
PYTHON_BIN=$(readlink -f /usr/bin/python3) PYTHON_BIN=$(readlink -f /usr/bin/python3)
PYTHON_VER=$(python3 --version 2>&1 | grep -oP '(?<=Python )\d+\.\d+' || echo "unknown") PYTHON_VER=$(lm_python_version /usr/bin/python3) || PYTHON_VER="unknown"
fi fi
if [ -n "$PYTHON_BIN" ] && [ -f "$PYTHON_BIN" ]; then if [ -n "$PYTHON_BIN" ] && [ -f "$PYTHON_BIN" ]; then
+5 -4
View File
@@ -6,8 +6,9 @@
files = src files = src
exclude = (^|/)(test|__pycache__)/ exclude = (^|/)(test|__pycache__)/
# Python version # Python version: the oldest the installer supports (Raspberry Pi OS
python_version = 3.10 # Bookworm ships 3.11; Trixie ships 3.13).
python_version = 3.11
# Platform (Linux/Raspberry Pi) # Platform (Linux/Raspberry Pi)
platform = linux platform = linux
@@ -103,8 +104,8 @@ ignore_missing_imports = True
# numpy's own stubs (numpy>=2.3) use Python 3.12 `type` statements, which mypy # numpy's own stubs (numpy>=2.3) use Python 3.12 `type` statements, which mypy
# refuses to parse under python_version = 3.10 -- and 3.10 is the floor this # refuses to parse under python_version = 3.11 -- and 3.11 (Bookworm) is the
# code has to run on, so it stays. Treat numpy as Any instead: skip it, and # floor this code has to run on, so it stays. Treat numpy as Any instead: skip it, and
# follow_imports_for_stubs makes the skip apply to its .pyi files too. # follow_imports_for_stubs makes the skip apply to its .pyi files too.
[mypy-numpy.*] [mypy-numpy.*]
follow_imports = skip follow_imports = skip
@@ -0,0 +1,174 @@
Faster SetImage for rpi-rgb-led-matrix (applied by first_time_install.sh at build
time; the submodule itself stays at its pinned commit).
Copying a frame into the panel buffer was the biggest CPU cost LEDMatrix owns on
large panels: the binding's SetPixelsPillow walked the image column by column
and called SetPixel per pixel, and each SetPixel read-modify-writes one word per
PWM bit plane, 2KB apart, so consecutive pixels were a whole double-row apart and
almost every write missed the cache. This patch:
* FrameCanvas gets its own SetPixelsPillow: row by row, one bulk SetPixels call
per row;
* Framebuffer::SetPixels clips once, looks colours up once per pixel, walks each
row's designators in order and writes the bit planes branch-free;
* the base Canvas.SetPixelsPillow loop (RGBMatrix.SetImage) is row-major.
The bit-plane buffer is byte-identical to the old code's (882 memcmp checks over
noise/gradient/solid/low-value/sparse images, clipped offsets, pwm 7/8/11,
brightness 1/50/90/100, inverse colours, luminance correction off and a pixel
mapper). Measured on a Pi 4 at 512x64: 6.0-6.3 ms -> 1.8 ms per frame through
the Python binding; on hdpi (Pi 4, 4x128x64) frame copy 6.57 -> 2.21 ms and the
display process 139% -> 103% of a core.
LEDMatrix always draws into the canvas that is not on screen and swaps it in
(DisplayManager.update_display), so the write order cannot show as tearing.
Against hzeller/rpi-rgb-led-matrix 1ee4f76.
diff --git a/bindings/python/rgbmatrix/core.pyx b/bindings/python/rgbmatrix/core.pyx
index 230d87f..babc3bb 100644
--- a/bindings/python/rgbmatrix/core.pyx
+++ b/bindings/python/rgbmatrix/core.pyx
@@ -2,6 +2,7 @@
from libcpp cimport bool
from libc.stdint cimport uint8_t, uint32_t, uintptr_t
+from libc.stdlib cimport malloc, free
import cython
cdef extern from "Python.h":
@@ -59,8 +60,9 @@ cdef class Canvas:
buffer = get_pillow_buffer(image_capsule)
- for col in range(max(0, -xstart), min(width, frame_width - xstart)):
- for row in range(max(0, -ystart), min(height, frame_height - ystart)):
+ # Row-major: walks both the image and the bitplane buffer sequentially.
+ for row in range(max(0, -ystart), min(height, frame_height - ystart)):
+ for col in range(max(0, -xstart), min(width, frame_width - xstart)):
pixel = buffer[row][col]
r = (pixel ) & 0xFF
g = (pixel >> 8) & 0xFF
@@ -86,6 +88,41 @@ cdef class FrameCanvas(Canvas):
def SetPixel(self, int x, int y, uint8_t red, uint8_t green, uint8_t blue):
(<cppinc.FrameCanvas*>self._getCanvas()).SetPixel(x, y, red, green, blue)
+ @cython.boundscheck(False)
+ @cython.wraparound(False)
+ def SetPixelsPillow(self, int xstart, int ystart, int width, int height, object image_capsule):
+ # Same result as Canvas.SetPixelsPillow(), but hands each image row
+ # to the C++ bulk FrameCanvas::SetPixels() instead of calling the
+ # virtual SetPixel() once per pixel.
+ cdef cppinc.FrameCanvas* my_canvas = <cppinc.FrameCanvas*>self._getCanvas()
+ cdef int col_start = max(0, -xstart)
+ cdef int col_end = min(width, my_canvas.width() - xstart)
+ cdef int row_start = max(0, -ystart)
+ cdef int row_end = min(height, my_canvas.height() - ystart)
+ cdef int row, col, pixel
+ cdef int *src
+ cdef cppinc.Color *line
+ cdef int **buffer
+
+ if col_end <= col_start or row_end <= row_start:
+ return
+ buffer = get_pillow_buffer(image_capsule)
+ line = <cppinc.Color*>malloc((col_end - col_start) * sizeof(cppinc.Color))
+ if line == NULL:
+ raise MemoryError()
+ try:
+ for row in range(row_start, row_end):
+ src = buffer[row]
+ for col in range(col_start, col_end):
+ pixel = src[col]
+ line[col - col_start].r = pixel & 0xFF
+ line[col - col_start].g = (pixel >> 8) & 0xFF
+ line[col - col_start].b = (pixel >> 16) & 0xFF
+ my_canvas.SetPixels(xstart + col_start, ystart + row,
+ col_end - col_start, 1, line)
+ finally:
+ free(line)
+
property width:
def __get__(self): return (<cppinc.FrameCanvas*>self._getCanvas()).width()
diff --git a/bindings/python/rgbmatrix/cppinc.pxd b/bindings/python/rgbmatrix/cppinc.pxd
index 8bec241..314332d 100644
--- a/bindings/python/rgbmatrix/cppinc.pxd
+++ b/bindings/python/rgbmatrix/cppinc.pxd
@@ -25,6 +25,7 @@ cdef extern from "led-matrix.h" namespace "rgb_matrix":
FrameCanvas *SwapOnVSync(FrameCanvas*, uint8_t)
cdef cppclass FrameCanvas(Canvas):
+ void SetPixels(int, int, int, int, Color*) nogil
bool SetPWMBits(uint8_t)
uint8_t pwmbits()
void SetBrightness(uint8_t)
diff --git a/lib/framebuffer.cc b/lib/framebuffer.cc
index 36d138b..aee62ca 100644
--- a/lib/framebuffer.cc
+++ b/lib/framebuffer.cc
@@ -807,11 +807,60 @@ void Framebuffer::SetPixel(int x, int y, uint8_t r, uint8_t g, uint8_t b) {
}
}
+// Bulk version of SetPixel(); produces exactly the same bitplane content.
+// Faster because it hoists the per-pixel work out of the loop: the color
+// mapping becomes one 256-entry table built per call (each channel maps
+// independently through the same function), the pixel designators of a row
+// are contiguous in the PixelDesignatorMap, and the bit-plane loop is
+// branchless (the color bits are effectively random, so the branches in
+// SetPixel() mispredict a lot).
void Framebuffer::SetPixels(int x, int y, int width, int height, Color *colors) {
- for (int iy = 0; iy < height; ++iy) {
- for (int ix = 0; ix < width; ++ix) {
- SetPixel(x + ix, y + iy, colors->r, colors->g, colors->b);
- ++colors;
+ PixelDesignatorMap *const mapper = *shared_mapper_;
+ const int ix_start = std::max(0, -x);
+ const int ix_end = std::min(width, mapper->width() - x);
+ const int iy_start = std::max(0, -y);
+ const int iy_end = std::min(height, mapper->height() - y);
+ if (ix_start >= ix_end || iy_start >= iy_end) return;
+
+ // Common case (luminance correction, no inversion): use the precomputed
+ // table directly; otherwise build one. Cheap enough to do per call, which
+ // matters for callers that send one row at a time.
+ uint16_t local_map[256];
+ const uint16_t *color_map;
+ if (do_luminance_correct_ && !inverse_color_) {
+ color_map = ColorLookupTable::GetLookup(brightness_).color;
+ } else {
+ for (int c = 0; c < 256; ++c) {
+ uint16_t unused1, unused2;
+ MapColors(c, 0, 0, &local_map[c], &unused1, &unused2);
+ }
+ color_map = local_map;
+ }
+
+ const int min_bit_plane = kBitPlanes - pwm_bits_;
+ gpio_bits_t *const plane_start = bitplane_buffer_ + columns_ * min_bit_plane;
+ for (int iy = iy_start; iy < iy_end; ++iy) {
+ const Color *c = colors + iy * width + ix_start;
+ const PixelDesignator *designator = mapper->get(x + ix_start, y + iy);
+ for (int ix = ix_start; ix < ix_end; ++ix, ++c, ++designator) {
+ const long pos = designator->gpio_word;
+ if (pos < 0) continue; // non-used pixel marker.
+ const uint16_t red = color_map[c->r];
+ const uint16_t green = color_map[c->g];
+ const uint16_t blue = color_map[c->b];
+ const gpio_bits_t r_bits = designator->r_bit;
+ const gpio_bits_t g_bits = designator->g_bit;
+ const gpio_bits_t b_bits = designator->b_bit;
+ const gpio_bits_t designator_mask = designator->mask;
+ gpio_bits_t *bits = plane_start + pos;
+ for (int plane = min_bit_plane; plane < kBitPlanes; ++plane) {
+ const gpio_bits_t color_bits =
+ (r_bits & -(gpio_bits_t)((red >> plane) & 1))
+ | (g_bits & -(gpio_bits_t)((green >> plane) & 1))
+ | (b_bits & -(gpio_bits_t)((blue >> plane) & 1));
+ *bits = (*bits & designator_mask) | color_bits;
+ bits += columns_;
+ }
}
}
}
+1 -1
View File
@@ -1,5 +1,5 @@
# LEDMatrix Core Dependencies # LEDMatrix Core Dependencies
# Compatible with Python 3.10, 3.11, 3.12, and 3.13 # Compatible with Python 3.11, 3.12 and 3.13; CI tests 3.11 and 3.13
# Tested on Raspbian OS 12 (Bookworm) and 13 (Trixie) # Tested on Raspbian OS 12 (Bookworm) and 13 (Trixie)
# Image processing # Image processing
+69 -35
View File
@@ -53,26 +53,35 @@ else
fi fi
echo "" echo ""
# Check OS version # Check OS version. The supported releases come from the same library the
# installer uses, so the two cannot disagree.
echo "2. Checking Operating System Version..." echo "2. Checking Operating System Version..."
echo "---------------------------------------" echo "---------------------------------------"
if [ -f /etc/os-release ]; then OS_LIB="$(cd "$(dirname "$0")" && pwd)/install/lib_os.sh"
. /etc/os-release OS_LIB_LOADED=0
echo "OS: $PRETTY_NAME" OS_RELEASE=""
echo "Version ID: ${VERSION_ID:-unknown}" if [ -f "$OS_LIB" ]; then
# shellcheck source=scripts/install/lib_os.sh
# first_time_install.sh refuses anything but Raspberry Pi OS / Debian 13 . "$OS_LIB"
# (Trixie), so anything else is an error here too, not a warning. OS_LIB_LOADED=1
if [[ "$ID" == "raspbian" ]] || [[ "$ID" == "debian" ]]; then fi
if [ "${VERSION_ID:-0}" = "13" ]; then
print_success "Detected Debian 13 Trixie - supported" if [ "$OS_LIB_LOADED" = "0" ]; then
elif [ "${VERSION_ID:-0}" = "12" ]; then print_error "$OS_LIB is missing - download LEDMatrix again"
print_error "Debian 12 Bookworm is not supported - the installer requires Raspberry Pi OS Lite (Trixie), Debian 13" elif [ -r "$LM_OS_RELEASE_FILE" ]; then
else OS_ID=$(lm_os_field ID)
print_error "Debian/Raspbian ${VERSION_ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite (Trixie), Debian 13" OS_VERSION_ID=$(lm_os_field VERSION_ID)
fi echo "OS: $(lm_os_field PRETTY_NAME)"
echo "Version ID: ${OS_VERSION_ID:-unknown}"
# first_time_install.sh refuses anything else, so this is an error here
# too, not a warning.
if OS_RELEASE=$(lm_os_release); then
print_success "Detected $(lm_release_label "$OS_RELEASE") - supported"
elif [[ "$OS_ID" == "raspbian" ]] || [[ "$OS_ID" == "debian" ]]; then
print_error "Debian/Raspbian ${OS_VERSION_ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12)"
else else
print_error "${ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite (Trixie), Debian 13" print_error "${OS_ID:-unknown} is not supported - the installer requires Raspberry Pi OS Lite, Trixie (Debian 13) or Bookworm (Debian 12)"
fi fi
else else
print_error "Could not detect OS version" print_error "Could not detect OS version"
@@ -92,7 +101,7 @@ if [ "$KERNEL_MAJOR" -ge "6" ]; then
print_success "Kernel version is compatible (6.x or newer)" print_success "Kernel version is compatible (6.x or newer)"
if [ "$KERNEL_MAJOR" -eq "6" ] && [ "$KERNEL_MINOR" -ge "12" ]; then if [ "$KERNEL_MAJOR" -eq "6" ] && [ "$KERNEL_MINOR" -ge "12" ]; then
print_success "Running latest Trixie kernel (6.12 LTS)" print_success "Running a 6.12 LTS or newer kernel"
fi fi
elif [ "$KERNEL_MAJOR" -eq "5" ] && [ "$KERNEL_MINOR" -ge "10" ]; then elif [ "$KERNEL_MAJOR" -eq "5" ] && [ "$KERNEL_MINOR" -ge "10" ]; then
print_success "Kernel version is compatible (5.10+)" print_success "Kernel version is compatible (5.10+)"
@@ -104,25 +113,34 @@ echo ""
# Check Python version # Check Python version
echo "4. Checking Python Version..." echo "4. Checking Python Version..."
echo "-----------------------------" echo "-----------------------------"
if command -v python3 >/dev/null 2>&1; then if [ "$OS_LIB_LOADED" = "1" ] && command -v python3 >/dev/null 2>&1; then
PYTHON_VERSION=$(python3 -c 'import sys; print(f"{sys.version_info.major}.{sys.version_info.minor}.{sys.version_info.micro}")') PYTHON_VERSION=$(python3 -c 'import sys; print("%d.%d.%d" % sys.version_info[:3])')
PYTHON_MAJOR=$(python3 -c 'import sys; print(sys.version_info.major)') PYTHON_MINOR_VERSION=$(lm_python_version) || PYTHON_MINOR_VERSION=""
PYTHON_MINOR=$(python3 -c 'import sys; print(sys.version_info.minor)') PYTHON_RANGE="3.${LM_PYTHON_MIN_MINOR}-3.${LM_PYTHON_MAX_MINOR}"
echo "Python: $PYTHON_VERSION" echo "Python: $PYTHON_VERSION"
if [ "$PYTHON_MAJOR" -eq "3" ]; then case "$(lm_python_check "$PYTHON_MINOR_VERSION")" in
if [ "$PYTHON_MINOR" -ge "10" ] && [ "$PYTHON_MINOR" -le "13" ]; then ok)
print_success "Python version is supported (3.10-3.13)" print_success "Python version is supported ($PYTHON_RANGE)"
elif [ "$PYTHON_MINOR" -ge "14" ]; then ;;
print_warning "Python 3.${PYTHON_MINOR} is very new - some packages may not be compatible yet" too-old)
else # The rgbmatrix bindings declare requires-python >=3.11, so the
# Pillow 12 and the pinned test tools need 3.10+, so this won't install. # display cannot be built on anything older.
print_error "Python 3.${PYTHON_MINOR} is too old - Python 3.10+ is required" print_error "Python $PYTHON_MINOR_VERSION is too old - Python 3.${LM_PYTHON_MIN_MINOR}+ is required"
fi ;;
else too-new)
print_error "Python 2.x detected - Python 3.10+ is required" print_warning "Python $PYTHON_MINOR_VERSION is newer than LEDMatrix has been tested with ($PYTHON_RANGE)"
;;
*)
print_warning "Could not read the Python version"
;;
esac
if [ -n "$OS_RELEASE" ] && [ "$PYTHON_MINOR_VERSION" != "$(lm_release_python "$OS_RELEASE")" ]; then
print_warning "$(lm_release_label "$OS_RELEASE") ships Python $(lm_release_python "$OS_RELEASE"), but python3 runs $PYTHON_MINOR_VERSION"
fi fi
elif command -v python3 >/dev/null 2>&1; then
print_warning "Cannot check the Python version without $OS_LIB"
else else
print_error "Python 3 not found - installation required" print_error "Python 3 not found - installation required"
fi fi
@@ -268,6 +286,22 @@ if command -v ping >/dev/null 2>&1; then
else else
print_warning "Ping command not available - cannot verify network" print_warning "Ping command not available - cannot verify network"
fi fi
# WiFi setup from the web page and the LEDMatrix-Setup hotspot drive
# NetworkManager, the default on both Bookworm and Trixie.
if [ "$OS_LIB_LOADED" = "1" ]; then
case "$(lm_network_stack)" in
networkmanager)
print_success "NetworkManager manages the network (needed for WiFi setup)"
;;
dhcpcd)
print_warning "dhcpcd manages the network - WiFi setup from the web page and the setup hotspot need NetworkManager (sudo raspi-config -> Advanced Options -> Network Config)"
;;
*)
print_warning "Could not tell which service manages the network - WiFi setup from the web page needs NetworkManager"
;;
esac
fi
echo "" echo ""
# Print summary # Print summary
+13 -3
View File
@@ -5,11 +5,18 @@ This directory contains scripts for installing and configuring the LEDMatrix sys
## Scripts ## Scripts
- **`one-shot-install.sh`** - Single-command installer; clones the - **`one-shot-install.sh`** - Single-command installer; clones the
repo, checks prerequisites, then runs `first_time_install.sh`. repo, checks out the newest release (or `main` with
Invoked via `curl ... | bash` from the project root README. `LEDMATRIX_CHANNEL=beta`), checks prerequisites, then runs
`first_time_install.sh`. Invoked via `curl ... | bash` from the project
root README. Re-running it never moves a checkout to an older version.
- **`install_service.sh`** - Installs, enables and starts the display - **`install_service.sh`** - Installs, enables and starts the display
service (`ledmatrix.service`), the web interface service service (`ledmatrix.service`), the web interface service
(`ledmatrix-web.service`) and the update-verify units (systemd) (`ledmatrix-web.service`) and the update-verify units (systemd), and
installs `/usr/local/sbin/ledmatrix-refresh-units`
- **`ledmatrix_refresh_units.py`** - Not run from here: `install_service.sh`
installs a root-owned copy as `/usr/local/sbin/ledmatrix-refresh-units`,
which updates run through sudo to install changed units (and the
automatic update's rollback, with `--restore`, to put them back)
- **`install_web_service.sh`** - Installs only the web interface service - **`install_web_service.sh`** - Installs only the web interface service
and the update-verify units (systemd) and the update-verify units (systemd)
- **`install_wifi_monitor.sh`** - Installs the WiFi monitor daemon service - **`install_wifi_monitor.sh`** - Installs the WiFi monitor daemon service
@@ -34,6 +41,9 @@ Libraries (sourced, not run):
script that renders a unit from `systemd/*.service` script that renders a unit from `systemd/*.service`
- **`lib_lowmem.sh`** - Build-job sizing and temporary swap for the C++ - **`lib_lowmem.sh`** - Build-job sizing and temporary swap for the C++
build on low-memory Pis (`first_time_install.sh` Step 6) build on low-memory Pis (`first_time_install.sh` Step 6)
- **`lib_os.sh`** - Which releases (Bookworm, Trixie) and Python versions
(3.11-3.13) the installer accepts, and which service runs the network;
shared by `first_time_install.sh` and `scripts/check_system_compatibility.sh`
## Usage ## Usage
+2
View File
@@ -138,6 +138,8 @@ echo "- View system logs via journalctl"
echo "- Reboot and shutdown the system" echo "- Reboot and shutdown the system"
echo "- Remove plugin directories (for update/uninstall when root-owned files block deletion)" echo "- Remove plugin directories (for update/uninstall when root-owned files block deletion)"
echo "- Install plugin/base requirements.txt as root (so ledmatrix.service can see them)" echo "- Install plugin/base requirements.txt as root (so ledmatrix.service can see them)"
echo "- Install the LEDMatrix systemd units an update changed, and restore them on rollback"
echo " (/usr/local/sbin/ledmatrix-refresh-units, installed by install_service.sh)"
echo "" echo ""
# Ask for confirmation # Ask for confirmation
+24
View File
@@ -143,6 +143,30 @@ for VERIFY_UNIT in ledmatrix-update-verify.service ledmatrix-update-verify.path;
fi fi
done done
# The helper updates run (through sudo, see lib_sudoers.sh) to install these
# same units when a new version changes their templates, and to put the old
# ones back if the automatic update rolls back. Root-owned and outside the
# checkout, so the web user who owns the checkout cannot change what sudo runs.
# Not fatal: without it, updates leave the units for the next reinstall.
REFRESH_UNITS_SRC="$PROJECT_ROOT_DIR/scripts/install/ledmatrix_refresh_units.py"
REFRESH_UNITS_DEST=/usr/local/sbin/ledmatrix-refresh-units
if [ -f "$REFRESH_UNITS_SRC" ]; then
if sudo install -D -o root -g root -m 0755 "$REFRESH_UNITS_SRC" "$REFRESH_UNITS_DEST"; then
echo "Installed $REFRESH_UNITS_DEST (lets updates refresh these units)"
else
echo "WARNING: could not install $REFRESH_UNITS_DEST; updates will not refresh the systemd units" >&2
fi
fi
# The units above are copied from mktemp files, which are 0600. 0644 is what
# first_time_install.sh (Step 8.1) sets, and lets the web interface compare
# them with the templates after an update without root.
for INSTALLED_UNIT in ledmatrix.service ledmatrix-web.service \
ledmatrix-update-verify.service ledmatrix-update-verify.path; do
if [ -f "/etc/systemd/system/$INSTALLED_UNIT" ]; then
sudo chmod 644 "/etc/systemd/system/$INSTALLED_UNIT" || true
fi
done
echo "Reloading systemd daemon for web service..." echo "Reloading systemd daemon for web service..."
sudo systemctl daemon-reload sudo systemctl daemon-reload
+451
View File
@@ -0,0 +1,451 @@
#!/usr/bin/python3 -I
"""Refresh the installed LEDMatrix systemd units from the checkout's templates.
Installed by scripts/install/install_service.sh as a root-owned copy,
/usr/local/sbin/ledmatrix-refresh-units, and granted to the web interface's
user by /etc/sudoers.d/ledmatrix_web (scripts/install/lib_sudoers.sh) with
exactly two command lines:
ledmatrix-refresh-units (no arguments)
ledmatrix-refresh-units --restore
An update (Update Code, or the weekly automatic update) pulls new unit
templates into systemd/, but the units systemd runs are the copies in
/etc/systemd/system, which only the installer used to write. So a setting
added to a template -- the render-loop watchdog, a memory limit -- never
reached a device that was already installed. After an update the web
interface runs this, and the next restart picks the new units up.
* **No arguments:** render each installed unit from systemd/<unit> exactly as
install_service.sh does (__PROJECT_ROOT_DIR__ and __USER__ replaced
literally), and install the ones whose content differs (comments and blank
lines aside, as src/startup_validator.py compares them), then
``systemctl daemon-reload``. The units replaced are saved first, so the
automatic update's rollback can put them back.
* ``--restore``: put back the units the last refresh replaced, and
daemon-reload. Nothing saved means nothing to do.
* ``--check``: print the units that would change, one per line. Needs no
root and changes nothing.
What it trusts, and why. It takes no other input: the project directory and
the web interface's user come from the installed, root-owned
ledmatrix.service and ledmatrix-web.service, not from the caller, and sudo
strips the caller's environment (``-I`` ignores the PYTHON* variables too).
It only replaces units that are already installed, only the four
install_service.sh installs, and only with a rendering that keeps each unit's
User= (root for the display, the web user for the others) and
WorkingDirectory=. The templates are files the web user can edit -- but so is
run.py, which ledmatrix.service already runs as root, so a template grants
nothing that user did not have; the checks keep a damaged or hostile template
from changing who a unit runs as, and keep this from reading anything but a
regular file under the checkout's systemd/ folder.
Standard library only, and no imports from the checkout: the installed copy
must not run code the web user can change.
"""
import json
import os
import re
import stat
import subprocess # nosec B404 - fixed argv, no shell # nosemgrep
import sys
import tempfile
SYSTEMD_DIR = '/etc/systemd/system'
#: Root-only: the units the last refresh replaced, for --restore.
BACKUP_DIR = '/var/lib/ledmatrix/unit-backup'
MANIFEST = 'manifest.json'
INSTALLED_PATH = '/usr/local/sbin/ledmatrix-refresh-units'
DISPLAY_UNIT = 'ledmatrix.service'
WEB_UNIT = 'ledmatrix-web.service'
VERIFY_SERVICE = 'ledmatrix-update-verify.service'
VERIFY_PATH = 'ledmatrix-update-verify.path'
#: What install_service.sh installs, in its order. Nothing else is touched.
UNITS = (DISPLAY_UNIT, WEB_UNIT, VERIFY_SERVICE, VERIFY_PATH)
MAX_TEMPLATE_BYTES = 64 * 1024
_USER_RE = re.compile(r'^[a-z_][a-z0-9_-]{0,31}$')
#: systemd expands % specifiers, and a quote, backslash or line break would
#: be reinterpreted in a unit file (src/auto_update_setup.py refuses the same).
#: (On Windows, where the tests also run, a backslash is the path separator.)
_UNSAFE_PATH_CHARS = set('%"') | ({'\\'} if os.sep == '/' else set())
EXIT_OK = 0
EXIT_FAILED = 1
EXIT_USAGE = 2
class RefreshError(Exception):
"""Why the units were left alone, in words for the web interface's log."""
class UnitsUnreadable(RefreshError):
"""An installed unit is not readable by this (unprivileged) user.
install_service.sh used to leave units mode 0600 (first_time_install.sh's
Step 8.1 makes them 0644), so ``--check`` as the web user cannot always
tell; the root helper itself can.
"""
def directive_values(text, key):
"""Every value of ``key=`` in a unit's text, in order (systemd allows spaces around ``=``)."""
return [m.group(1).strip() for m in re.finditer(rf'^[ \t]*{key}[ \t]*=(.*)$', text or '', re.M)]
def layout_problem(text, section, keys):
"""What would make ``directive_values`` misread the unit as systemd reads it, or None.
A ``User=`` inside a backslash-continued line is part of the line before,
and one under [Unit] is ignored, so either could pass a check that systemd
then does not apply. Neither appears in the shipped templates.
"""
current = None
for raw in (text or '').splitlines():
line = raw.strip()
if not line or line.startswith(('#', ';')):
continue
if line.endswith('\\'):
return 'continues a line with a backslash'
if line.startswith('[') and line.endswith(']'):
current = line[1:-1]
continue
key = line.split('=', 1)[0].strip()
if key in keys and current != section:
return f'sets {key}= outside [{section}]'
return None
def unit_body(text):
"""A unit's meaningful lines in order: no comments, no blank lines.
The same comparison src/startup_validator.py uses for its drift warning,
so what this refreshes is exactly what that warns about.
"""
lines = []
for line in (text or '').splitlines():
line = line.strip()
if line and not line.startswith('#'):
lines.append(line)
return '\n'.join(lines)
def render(template, project_root, user):
"""install_service.sh's ``sed "s|__PROJECT_ROOT_DIR__|...|g; s|__USER__|...|g"``."""
return template.replace('__PROJECT_ROOT_DIR__', project_root).replace('__USER__', user)
def _read_regular(path, limit=MAX_TEMPLATE_BYTES, dir_fd=None):
"""A regular file's text, never through a symlink, a FIFO or a device."""
flags = os.O_RDONLY | getattr(os, 'O_NOFOLLOW', 0) | getattr(os, 'O_NONBLOCK', 0)
kwargs = {'dir_fd': dir_fd} if dir_fd is not None else {}
fd = os.open(path, flags, **kwargs)
try:
info = os.fstat(fd)
if not stat.S_ISREG(info.st_mode):
raise RefreshError(f'{path} is not a regular file')
if info.st_size > limit:
raise RefreshError(f'{path} is larger than {limit} bytes')
data = b''
while True:
chunk = os.read(fd, limit + 1 - len(data))
if not chunk:
break
data += chunk
if len(data) > limit:
raise RefreshError(f'{path} is larger than {limit} bytes')
finally:
os.close(fd)
if b'\0' in data:
raise RefreshError(f'{path} is not a text file')
try:
return data.decode('utf-8')
except UnicodeDecodeError as e:
raise RefreshError(f'{path} is not UTF-8') from e
def _read_installed(systemd_dir, name):
path = os.path.join(systemd_dir, name)
try:
with open(path, 'r', encoding='utf-8') as f:
return f.read()
except FileNotFoundError:
return None
except PermissionError as e:
raise UnitsUnreadable(f'cannot read the installed {name}: {e}') from e
except (OSError, UnicodeDecodeError) as e:
raise RefreshError(f'cannot read the installed {name}: {e}') from e
def _lookup_user(user):
try:
import pwd
except ImportError: # not a POSIX host (the tests on Windows)
return True
try:
pwd.getpwnam(user)
return True
except KeyError:
return False
class Refresher:
def __init__(self, systemd_dir=SYSTEMD_DIR, backup_dir=BACKUP_DIR, run=subprocess.run,
is_root=None, user_exists=_lookup_user, log=None):
self.systemd_dir = systemd_dir
self.backup_dir = backup_dir
self.run = run
self.is_root = is_root or (lambda: hasattr(os, 'geteuid') and os.geteuid() == 0)
self.user_exists = user_exists
self.log = log or (lambda msg: print(msg, flush=True))
# -- what the installed units say -------------------------------------
def context(self, installed):
"""(project root, web user) from the installed, root-owned units."""
display = installed.get(DISPLAY_UNIT)
if display is None:
raise RefreshError(f'{DISPLAY_UNIT} is not installed; run scripts/install/install_service.sh')
roots = directive_values(display, 'WorkingDirectory')
if len(roots) != 1:
raise RefreshError(f'the installed {DISPLAY_UNIT} does not name one WorkingDirectory')
root = roots[0]
if (not os.path.isabs(root) or any(ch in _UNSAFE_PATH_CHARS or ord(ch) < 32 for ch in root)
or os.path.normpath(root) != root):
raise RefreshError(f'the installed {DISPLAY_UNIT} runs from {root!r}, which cannot be used')
if not os.path.isdir(root):
raise RefreshError(f'{root} (the installed {DISPLAY_UNIT} WorkingDirectory) does not exist')
user = None
web = installed.get(WEB_UNIT)
if web is not None:
users = directive_values(web, 'User')
user = users[0] if len(users) == 1 else ('root' if not users else None)
if user is None or not _USER_RE.match(user) or not self.user_exists(user):
raise RefreshError(f'the installed {WEB_UNIT} runs as an account that cannot be used')
if directive_values(web, 'WorkingDirectory') != [root]:
raise RefreshError(f'the installed {WEB_UNIT} and {DISPLAY_UNIT} run from different folders')
return root, user
@staticmethod
def expected_user(name, web_user):
return 'root' if name == DISPLAY_UNIT else web_user
def _template(self, root, name):
"""systemd/<name> under the checkout, as a regular file, never via a symlink."""
dir_flags = os.O_RDONLY | getattr(os, 'O_DIRECTORY', 0) | getattr(os, 'O_NOFOLLOW', 0)
if os.open in getattr(os, 'supports_dir_fd', set()):
try:
dfd = os.open(os.path.join(root, 'systemd'), dir_flags)
except OSError as e:
raise RefreshError(f'cannot open {root}/systemd: {e}') from e
try:
return _read_regular(name, dir_fd=dfd)
except FileNotFoundError:
return None
except OSError as e:
raise RefreshError(f'cannot read systemd/{name}: {e}') from e
finally:
os.close(dfd)
path = os.path.join(root, 'systemd', name)
if os.path.islink(os.path.join(root, 'systemd')):
raise RefreshError(f'{root}/systemd is a symlink')
try:
return _read_regular(path)
except FileNotFoundError:
return None
except OSError as e:
raise RefreshError(f'cannot read systemd/{name}: {e}') from e
def _validate(self, name, rendered, root, user):
problem = layout_problem(rendered, 'Service', ('User', 'WorkingDirectory'))
if problem:
raise RefreshError(f'systemd/{name} {problem}; refusing to install it')
if directive_values(rendered, 'User') != [user]:
raise RefreshError(f'systemd/{name} would not run as {user}; refusing to install it')
if directive_values(rendered, 'WorkingDirectory') != [root]:
raise RefreshError(f'systemd/{name} would not run from {root}; refusing to install it')
def plan(self):
"""{unit: (installed text, new text)} for every installed unit that would change.
Raises RefreshError, and so changes nothing, if any unit cannot be
rendered safely: four units refreshed as a set or not at all.
"""
installed = {name: _read_installed(self.systemd_dir, name) for name in UNITS}
root, web_user = self.context(installed)
changes = {}
for name in UNITS:
current = installed[name]
if current is None:
continue # never installed here: installing is the installer's job
user = self.expected_user(name, web_user)
if user is None:
continue # the web unit is not installed, so neither is its user
template = self._template(root, name)
if template is None:
continue # a version without this unit leaves the installed one alone
rendered = render(template, root, user)
# A path unit runs nothing itself; what matters is what it starts.
if name.endswith('.service'):
self._validate(name, rendered, root, user)
else:
self._validate_path(name, rendered)
if unit_body(rendered) != unit_body(current):
changes[name] = (current, rendered)
return changes
def _validate_path(self, name, rendered):
problem = layout_problem(rendered, 'Path', ('Unit',))
if problem:
raise RefreshError(f'systemd/{name} {problem}; refusing to install it')
if directive_values(rendered, 'Unit') != [VERIFY_SERVICE]:
raise RefreshError(f'systemd/{name} does not start {VERIFY_SERVICE}; refusing to install it')
if directive_values(rendered, 'User'):
raise RefreshError(f'systemd/{name} sets User=; refusing to install it')
# -- writing ------------------------------------------------------------
def _write_unit(self, name, text):
fd, tmp = tempfile.mkstemp(dir=self.systemd_dir, prefix=f'.{name}.')
try:
with os.fdopen(fd, 'w', encoding='utf-8', newline='\n') as f:
f.write(text)
os.chmod(tmp, 0o644)
os.replace(tmp, os.path.join(self.systemd_dir, name))
except BaseException:
try:
os.unlink(tmp)
except OSError:
pass
raise
def _backup_dir(self):
"""The backup folder, created root-only; refused if it is not a plain folder."""
os.makedirs(os.path.dirname(self.backup_dir), mode=0o755, exist_ok=True)
try:
os.mkdir(self.backup_dir, 0o700)
except FileExistsError:
pass
info = os.lstat(self.backup_dir)
if not stat.S_ISDIR(info.st_mode):
raise RefreshError(f'{self.backup_dir} is not a folder')
if hasattr(os, 'geteuid') and info.st_uid != os.geteuid():
raise RefreshError(f'{self.backup_dir} is not owned by root')
return self.backup_dir
def _clear_backup(self, folder):
for entry in os.listdir(folder):
path = os.path.join(folder, entry)
if os.path.isfile(path) or os.path.islink(path):
os.unlink(path)
def _systemctl(self, *args):
result = self.run(['systemctl', *args], capture_output=True, text=True, timeout=60)
if result.returncode != 0:
raise RefreshError(f'"systemctl {" ".join(args)}" failed: '
f'{(result.stderr or result.stdout or "").strip()}')
def _restart_path_unit_if_active(self, names):
"""A rewritten path unit watches the old path until it is restarted."""
if VERIFY_PATH not in names:
return
state = self.run(['systemctl', 'is-active', VERIFY_PATH], capture_output=True, text=True, timeout=30)
if (state.stdout or '').strip() == 'active':
self._systemctl('restart', VERIFY_PATH)
def refresh(self):
if not self.is_root():
raise RefreshError('must run as root (sudo)')
changes = self.plan()
folder = self._backup_dir()
# Always reset: the backup belongs to this refresh, so a --restore
# after an update that changed nothing restores nothing.
self._clear_backup(folder)
if not changes:
self.log('units: up to date')
return []
for name, (current, _) in changes.items():
with open(os.path.join(folder, name), 'w', encoding='utf-8', newline='\n') as f:
f.write(current)
with open(os.path.join(folder, MANIFEST), 'w', encoding='utf-8') as f:
json.dump({'units': sorted(changes)}, f)
try:
for name, (_, rendered) in changes.items():
self._write_unit(name, rendered)
self._systemctl('daemon-reload')
except BaseException:
# A failed refresh is reported as a failure, so the update records
# no units_refreshed and a rollback would not --restore. Put the
# replaced units back now, rather than leave a half-written set
# under the old code.
self._undo(changes, folder)
raise
self._restart_path_unit_if_active(changes)
self.log('units refreshed: ' + ' '.join(sorted(changes)))
return sorted(changes)
def _undo(self, changes, folder):
"""Best effort: reinstall the units a failed refresh replaced."""
undone = True
for name, (current, _) in changes.items():
try:
self._write_unit(name, current)
except OSError as e:
undone = False
self.log(f'units: could not put back {name}: {e}')
try:
self.run(['systemctl', 'daemon-reload'], capture_output=True, text=True, timeout=60)
except (OSError, subprocess.SubprocessError) as e:
self.log(f'units: daemon-reload after putting units back failed: {e}')
if undone:
# Nothing is left to restore; keep the backup only if a unit could
# not be put back, so a manual --restore still can.
self._clear_backup(folder)
def restore(self):
if not self.is_root():
raise RefreshError('must run as root (sudo)')
folder = self._backup_dir()
try:
manifest = json.loads(_read_regular(os.path.join(folder, MANIFEST)))
except FileNotFoundError:
self.log('units: nothing to restore')
return []
names = [n for n in (manifest or {}).get('units', []) if n in UNITS]
for name in names:
self._write_unit(name, _read_regular(os.path.join(folder, name)))
self._systemctl('daemon-reload')
self._restart_path_unit_if_active(names)
self._clear_backup(folder)
self.log('units restored: ' + ' '.join(names))
return names
def main(argv, refresher=None):
args = argv[1:]
if args not in ([], ['--restore'], ['--check']):
print('usage: ledmatrix-refresh-units [--restore | --check]', file=sys.stderr)
return EXIT_USAGE
refresher = refresher or Refresher()
try:
if args == ['--check']:
for name in sorted(refresher.plan()):
print(name)
elif args == ['--restore']:
refresher.restore()
else:
refresher.refresh()
except (RefreshError, OSError, subprocess.SubprocessError, ValueError) as e:
print(f'ledmatrix-refresh-units: {e}', file=sys.stderr)
return EXIT_FAILED
return EXIT_OK
if __name__ == '__main__':
# Only as the installed program: sudo already sets a secure PATH, and
# this pins the one systemctl comes from. (Not in main(), which the
# tests call in-process.)
os.environ['PATH'] = '/usr/sbin:/usr/bin:/sbin:/bin'
sys.exit(main(sys.argv))
+138
View File
@@ -0,0 +1,138 @@
#!/bin/bash
# Which operating systems and Python versions LEDMatrix installs on.
#
# Sourced by first_time_install.sh and scripts/check_system_compatibility.sh,
# so the installer and the compatibility checker cannot disagree about what
# is supported. Pure functions: nothing here installs, changes or exits --
# the callers decide what to do with the answers.
#
# Supported (Lite, no desktop):
# Raspberry Pi OS / Debian 12 "Bookworm" -- Python 3.11
# Raspberry Pi OS / Debian 13 "Trixie" -- Python 3.13
#
# Everything the installer asks apt for (python3-pip, python3-venv,
# python-dev-is-python3, python3-pil, python3-pil.imagetk, build-essential,
# python3-setuptools, python3-wheel, cmake, ninja-build, git, curl, wget,
# unzip, and hostapd, dnsmasq, network-manager for WiFi setup) has the same
# name on both releases. Both ship a pip (23.0.1 and 25.1.1) that is PEP 668
# "externally managed" and accepts --break-system-packages, and a cmake (3.25
# and 3.31) new enough for the rgbmatrix build (3.22). So no step needs a
# per-release branch today; if one ever does, the release name comes from
# lm_os_release below.
# Test hook: the os-release file to read.
LM_OS_RELEASE_FILE="${LM_OS_RELEASE_FILE:-/etc/os-release}"
# Oldest and newest python3 minor versions the installer accepts. 3.11 is
# Bookworm's, and also the floor of the rgbmatrix bindings (requires-python
# >=3.11 in rpi-rgb-led-matrix-master/pyproject.toml); 3.13 is Trixie's.
LM_PYTHON_MIN_MINOR=11
LM_PYTHON_MAX_MINOR=13
# lm_os_field KEY -- one value from os-release with its quotes removed; empty
# when the key or the file is missing. Parsed rather than sourced so that
# os-release's ID, VERSION and friends do not land in the caller's variables.
lm_os_field() {
[ -r "$LM_OS_RELEASE_FILE" ] || return 0
sed -n "/^$1=/{s/^$1=//;s/^[\"']//;s/[\"']\$//;p;q;}" "$LM_OS_RELEASE_FILE"
}
# lm_os_release -- print "bookworm" or "trixie" and succeed on a supported
# release; print nothing and fail on anything else. VERSION_ID decides; the
# codename is used only when VERSION_ID is missing.
lm_os_release() {
local id version
id=$(lm_os_field ID)
version=$(lm_os_field VERSION_ID)
[ -n "$version" ] || version=$(lm_os_field VERSION_CODENAME)
case "$id" in
raspbian|debian) ;;
*) return 1 ;;
esac
case "$version" in
12|bookworm) echo bookworm ;;
13|trixie) echo trixie ;;
*) return 1 ;;
esac
}
# lm_release_label RELEASE -- how to name a release to a person.
lm_release_label() {
case "$1" in
bookworm) echo "Debian 12 (Bookworm)" ;;
trixie) echo "Debian 13 (Trixie)" ;;
*) echo "$1" ;;
esac
}
# lm_release_python RELEASE -- the python3 version a release ships, e.g. 3.11.
lm_release_python() {
case "$1" in
bookworm) echo 3.11 ;;
trixie) echo 3.13 ;;
*) return 1 ;;
esac
}
# lm_python_version [PYTHON] -- "3.11" and so on for python3 (or PYTHON);
# prints nothing and fails when it cannot be run.
lm_python_version() {
"${1:-python3}" -c 'import sys; print("%d.%d" % sys.version_info[:2])' 2>/dev/null
}
# lm_python_check VERSION -- print "ok", "too-old", "too-new" or "unknown"
# for a version such as 3.11. Always succeeds, so it is safe under set -e.
lm_python_check() {
local major minor
major=${1%%.*}
minor=${1#*.}
minor=${minor%%.*}
case "$major:$minor" in
*[!0-9:]*|:*|*:) echo unknown; return 0 ;;
esac
if [ "$major" -lt 3 ] || { [ "$major" -eq 3 ] && [ "$minor" -lt "$LM_PYTHON_MIN_MINOR" ]; }; then
echo too-old
elif [ "$major" -gt 3 ] || [ "$minor" -gt "$LM_PYTHON_MAX_MINOR" ]; then
echo too-new
else
echo ok
fi
}
# lm_network_stack -- which service runs the network: "networkmanager",
# "dhcpcd" or "unknown". Raspberry Pi OS uses NetworkManager on both Bookworm
# and Trixie; dhcpcd appears when someone switched back to it in raspi-config.
lm_network_stack() {
if systemctl is-active --quiet NetworkManager 2>/dev/null; then
echo networkmanager
elif systemctl is-active --quiet dhcpcd 2>/dev/null; then
echo dhcpcd
else
echo unknown
fi
}
# lm_print_dhcpcd_advice -- the explanation for a Pi running dhcpcd. WiFi
# setup from the web page and the LEDMatrix-Setup hotspot both drive
# NetworkManager (nmcli). The installer does not switch the network stack
# itself: doing that over SSH can cut the connection it is running on.
lm_print_dhcpcd_advice() {
echo "⚠ This Pi manages its network with dhcpcd, not NetworkManager."
echo " LEDMatrix installs and the display works, but choosing a WiFi network"
echo " from the web page and the LEDMatrix-Setup hotspot both need NetworkManager."
echo " To switch (with a keyboard and screen attached, or over Ethernet):"
echo " sudo raspi-config -> Advanced Options -> Network Config -> NetworkManager"
echo " then reboot."
}
# lm_print_supported_os_help -- what to do on an unsupported system.
lm_print_supported_os_help() {
echo "LEDMatrix needs Raspberry Pi OS Lite: Trixie (Debian 13) or Bookworm (Debian 12)."
echo ""
echo "To install Raspberry Pi OS Lite:"
echo " 1. Download Raspberry Pi Imager from: https://www.raspberrypi.com/software/"
echo " 2. Choose 'Raspberry Pi OS Lite (64-bit)'. Trixie is the current version and"
echo " is recommended; Bookworm (listed as Legacy) also works"
echo " 3. Flash it to the SD card"
echo " 4. Boot the Pi and run this script again"
}
+9
View File
@@ -10,6 +10,11 @@
# #
# Add or remove a grant here and nowhere else. # Add or remove a grant here and nowhere else.
# Root-owned copy of scripts/install/ledmatrix_refresh_units.py, installed by
# install_service.sh. Outside the checkout on purpose: the web user owns the
# checkout, so a granted file inside it could be rewritten and run as root.
LEDMATRIX_REFRESH_UNITS_PATH=/usr/local/sbin/ledmatrix-refresh-units
# web_sudoers_rules WEB_USER PROJECT_ROOT SYSTEMCTL_PATH BASH_PATH REBOOT_PATH POWEROFF_PATH JOURNALCTL_PATH # web_sudoers_rules WEB_USER PROJECT_ROOT SYSTEMCTL_PATH BASH_PATH REBOOT_PATH POWEROFF_PATH JOURNALCTL_PATH
# #
# Print the ledmatrix_web sudoers rules to stdout. # Print the ledmatrix_web sudoers rules to stdout.
@@ -58,6 +63,10 @@ $WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/scripts/fix_perms/safe_pl
# Install a requirements.txt as root via vetted helper, so packages are visible # Install a requirements.txt as root via vetted helper, so packages are visible
# to root-run ledmatrix.service (not just the web interface's own user). # to root-run ledmatrix.service (not just the web interface's own user).
$WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/scripts/fix_perms/safe_pip_install.sh * $WEB_USER ALL=(ALL) NOPASSWD: $BASH_PATH $PROJECT_ROOT/scripts/fix_perms/safe_pip_install.sh *
# After an update, install the new systemd units (no arguments: "" allows none)
# and, on the automatic update's rollback, put the previous ones back.
$WEB_USER ALL=(ALL) NOPASSWD: $LEDMATRIX_REFRESH_UNITS_PATH ""
$WEB_USER ALL=(ALL) NOPASSWD: $LEDMATRIX_REFRESH_UNITS_PATH --restore
EOF EOF
if [ -n "$JOURNALCTL_PATH" ]; then if [ -n "$JOURNALCTL_PATH" ]; then
cat << EOF cat << EOF
+119 -1
View File
@@ -3,6 +3,10 @@
# LED Matrix One-Shot Installation Script # LED Matrix One-Shot Installation Script
# This script provides a single-command installation experience # This script provides a single-command installation experience
# Usage: curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | bash # Usage: curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | bash
#
# A new install runs the newest release (the stable update channel). For the
# newest code from main instead (the beta channel), set LEDMATRIX_CHANNEL=beta:
# curl -fsSL https://raw.githubusercontent.com/ChuckBuilds/LEDMatrix/main/scripts/install/one-shot-install.sh | LEDMATRIX_CHANNEL=beta bash
set -Eeuo pipefail set -Eeuo pipefail
@@ -205,6 +209,114 @@ check_sudo() {
print_success "Sudo access confirmed" print_success "Sudo access confirmed"
} }
# --- release checkout helpers ------------------------------------------------
# Which version an install runs. The rules are web_interface/update_channel.py's,
# so the installer and the web interface's updates agree:
# stable (default) the newest vX.Y.Z tag by semantic version; pre-releases
# (v3.8.0-rc1), leading zeros and other tags are ignored
# beta main, the newest code
# Never backwards: an existing checkout moves to a release only when that
# release contains its current commit (git merge-base --is-ancestor).
# Never fatal: whatever goes wrong, the install carries on with the checkout
# as it is.
# Print "stable" or "beta": LEDMATRIX_CHANNEL when it is set, else the
# existing install's auto_update.channel (CONFIG_FILE), else stable.
_lm_channel() {
local config_file="${1:-}" value
value=$(printf '%s' "${LEDMATRIX_CHANNEL:-}" | tr '[:upper:]' '[:lower:]' | tr -d '[:space:]')
case "$value" in
stable|beta) printf '%s\n' "$value"; return 0 ;;
"") ;;
*) print_warning "LEDMATRIX_CHANNEL=${LEDMATRIX_CHANNEL} is not stable or beta; using stable" >&2
printf 'stable\n'; return 0 ;;
esac
if [ -n "$config_file" ] && [ -f "$config_file" ] && command -v python3 >/dev/null 2>&1; then
value=$(python3 - "$config_file" 2>/dev/null <<'PY' || true
import json, sys
try:
with open(sys.argv[1], encoding="utf-8") as f:
section = json.load(f).get("auto_update")
value = section.get("channel") if isinstance(section, dict) else None
print(value.strip().lower() if isinstance(value, str) else "")
except Exception:
print("")
PY
)
if [ "$value" = "beta" ]; then
printf 'beta\n'
return 0
fi
fi
printf 'stable\n'
}
# Print the newest release tag of the repository in the current directory,
# or nothing when it has none.
_lm_newest_release_tag() {
git tag --list 'v*' 2>/dev/null \
| grep -E '^v(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)\.(0|[1-9][0-9]*)$' \
| sort -t. -k1.2,1n -k2,2n -k3,3n \
| tail -n 1 || true
}
# A fresh clone (on main): move to the newest release unless beta was asked for.
_lm_checkout_release_after_clone() {
local channel tag
channel=$(_lm_channel "")
if [ "$channel" = "beta" ]; then
print_success "Beta channel: installing the newest code from main"
return 0
fi
tag=$(_lm_newest_release_tag)
if [ -z "$tag" ]; then
print_warning "No release found; installing the newest code from main"
return 0
fi
if git -c advice.detachedHead=false checkout --quiet --detach "${tag}^{commit}"; then
print_success "Installing release $tag (stable channel)"
else
print_warning "Could not check out release $tag; installing the newest code from main"
fi
return 0
}
# An existing checkout: move it forward along its channel, never backwards.
# Returns 1 when it should be updated the way it always was (a fast-forward
# pull of its branch): beta, or stable on a branch newer than every release.
_lm_update_existing_checkout() {
local channel tag head tag_sha
channel=$(_lm_channel "config/config.json")
if [ "$channel" = "beta" ]; then
return 1
fi
if ! git fetch --quiet --tags --force origin >/dev/null 2>&1; then
print_warning "Could not fetch release tags; keeping the current version"
return 0
fi
tag=$(_lm_newest_release_tag)
head=$(git rev-parse --verify --quiet HEAD 2>/dev/null || true)
if [ -n "$tag" ] && [ -n "$head" ] && git merge-base --is-ancestor "$head" "$tag" 2>/dev/null; then
tag_sha=$(git rev-parse --verify --quiet "${tag}^{commit}" 2>/dev/null || true)
if [ "$head" = "$tag_sha" ]; then
print_success "Already on the newest release, $tag"
elif git -c advice.detachedHead=false checkout --quiet --detach "${tag}^{commit}"; then
print_success "Updated to release $tag (stable channel)"
else
print_warning "Could not move to release $tag (local changes?); keeping the current version"
fi
return 0
fi
if git symbolic-ref --quiet HEAD >/dev/null 2>&1; then
# Newer than the newest release (or no release yet): follow the branch
# until a release includes this version, as updates do.
return 1
fi
print_success "This checkout is newer than the newest release${tag:+ ($tag)}; leaving it as it is"
return 0
}
# --- end release checkout helpers --------------------------------------------
# Main installation function # Main installation function
main() { main() {
print_step "LED Matrix One-Shot Installation" print_step "LED Matrix One-Shot Installation"
@@ -292,7 +404,10 @@ main() {
# Try to safely update current branch first (fast-forward only to avoid unintended merges) # Try to safely update current branch first (fast-forward only to avoid unintended merges)
PULL_SUCCESS=false PULL_SUCCESS=false
if git pull --ff-only origin "$CURRENT_BRANCH" >/dev/null 2>&1; then # Stable: the newest release, if it contains this version.
if _lm_update_existing_checkout; then
PULL_SUCCESS=true
elif git pull --ff-only origin "$CURRENT_BRANCH" >/dev/null 2>&1; then
print_success "Repository updated successfully (branch: $CURRENT_BRANCH)" print_success "Repository updated successfully (branch: $CURRENT_BRANCH)"
PULL_SUCCESS=true PULL_SUCCESS=true
else else
@@ -323,10 +438,12 @@ main() {
rm -rf "$REPO_DIR" rm -rf "$REPO_DIR"
print_success "Cloning repository..." print_success "Cloning repository..."
retry git clone "$REPO_URL" "$REPO_DIR" retry git clone "$REPO_URL" "$REPO_DIR"
(cd "$REPO_DIR" && _lm_checkout_release_after_clone) || print_warning "Could not choose a release; installing the newest code from main"
fi fi
else else
print_success "Cloning repository to $REPO_DIR..." print_success "Cloning repository to $REPO_DIR..."
retry git clone "$REPO_URL" "$REPO_DIR" retry git clone "$REPO_URL" "$REPO_DIR"
(cd "$REPO_DIR" && _lm_checkout_release_after_clone) || print_warning "Could not choose a release; installing the newest code from main"
fi fi
# Verify repository is accessible # Verify repository is accessible
@@ -397,6 +514,7 @@ main() {
sudo -E env TMPDIR=/tmp LEDMATRIX_ASSUME_YES=1 \ sudo -E env TMPDIR=/tmp LEDMATRIX_ASSUME_YES=1 \
LEDMATRIX_APT_UPDATED="${LEDMATRIX_APT_UPDATED:-0}" \ LEDMATRIX_APT_UPDATED="${LEDMATRIX_APT_UPDATED:-0}" \
LEDMATRIX_AUTO_UPDATE="${LEDMATRIX_AUTO_UPDATE:-}" \ LEDMATRIX_AUTO_UPDATE="${LEDMATRIX_AUTO_UPDATE:-}" \
LEDMATRIX_CHANNEL="${LEDMATRIX_CHANNEL:-}" \
bash ./first_time_install.sh -y </dev/null bash ./first_time_install.sh -y </dev/null
fi fi
INSTALL_EXIT_CODE=$? INSTALL_EXIT_CODE=$?
+63 -19
View File
@@ -4,13 +4,14 @@ Alternative dependency installer that tries apt packages first,
then falls back to pip with --break-system-packages then falls back to pip with --break-system-packages
""" """
import re
import subprocess import subprocess
import sys import sys
import tempfile import tempfile
import warnings import warnings
from collections import deque from collections import deque
from pathlib import Path from pathlib import Path
from typing import List, Tuple from typing import Dict, List, Tuple
# How many trailing lines of a failed command's output to keep for the # How many trailing lines of a failed command's output to keep for the
# end-of-run failure summary. Keeps the root cause near the end of the log, # end-of-run failure summary. Keeps the root cause near the end of the log,
@@ -81,6 +82,8 @@ def install_via_pip(package_name: str) -> Tuple[bool, str]:
Returns (success, output). Returns (success, output).
""" """
# pip knows PIL as Pillow; the others are asked for by their own name.
package_name = _dist_name(package_name)
print(f"Installing {package_name} via pip...") print(f"Installing {package_name} via pip...")
success, output = _run([ success, output = _run([
sys.executable, '-m', 'pip', 'install', sys.executable, '-m', 'pip', 'install',
@@ -99,26 +102,66 @@ IMPORT_NAME_MAP = {
'freetype-py': 'freetype', 'freetype-py': 'freetype',
} }
# Minimum versions that must be met for an already-installed package to count # The packages above are keyed by what main() lists; these are the ones whose
# as satisfied. Debian Bookworm's python3-freetype is 2.3.0, below the # pip distribution name differs from that key.
# freetype-py>=2.5.1 pin in requirements.txt, so an import-only check would DIST_NAME_MAP = {
# wrongly skip the pip upgrade. 'PIL': 'Pillow',
MIN_VERSIONS = {
'freetype-py': (2, 5, 1),
} }
REQUIREMENTS_FILE = Path(__file__).resolve().parent.parent / 'web_interface' / 'requirements.txt'
def _version_tuple(text: str) -> tuple:
parts = []
for part in text.split('.'):
digits = ''.join(ch for ch in part if ch.isdigit())
if not digits:
break
parts.append(int(digits))
return tuple(parts)
def _requirement_floors(path: Path = REQUIREMENTS_FILE) -> Dict[str, tuple]:
"""``>=`` floors from a requirements file, keyed by lower-cased name.
The apt copies of these packages are older than the pins on both
supported releases -- Bookworm ships Flask and Werkzeug 2.2.2, Pillow 9.4,
requests 2.28, psutil 5.9, pytz 2022.7 and freetype-py 2.3; Trixie ships
Flask 3.1.1, Werkzeug 3.1.3, Pillow 11.1 and requests 2.32 --
so a package that merely imports is not enough. Read from the file rather
than copied here so the two cannot drift.
"""
floors: Dict[str, tuple] = {}
try:
lines = path.read_text(encoding='utf-8').splitlines()
except OSError:
return floors
for line in lines:
match = re.match(r'\s*([A-Za-z0-9][A-Za-z0-9._-]*)[^#]*?>=\s*([0-9][0-9.]*)', line)
if match:
floors[match.group(1).lower()] = _version_tuple(match.group(2))
return floors
def _dist_name(package_name: str) -> str:
return DIST_NAME_MAP.get(package_name, package_name)
def _minimum_version(package_name: str) -> tuple:
"""The required floor for ``package_name``, or () when there is none."""
return MIN_VERSIONS.get(_dist_name(package_name).lower(), ())
# Minimum versions that must be met for an already-installed package to count
# as satisfied.
MIN_VERSIONS = _requirement_floors()
def _installed_version_tuple(dist_name: str) -> tuple: def _installed_version_tuple(dist_name: str) -> tuple:
"""Return the installed distribution version as an int tuple, or () if unknown.""" """Return the installed distribution version as an int tuple, or () if unknown."""
try: try:
from importlib.metadata import version from importlib.metadata import version
parts = [] return _version_tuple(version(dist_name))
for part in version(dist_name).split('.'):
digits = ''.join(ch for ch in part if ch.isdigit())
if not digits:
break
parts.append(int(digits))
return tuple(parts)
except Exception: except Exception:
return () return ()
@@ -134,9 +177,9 @@ def check_package_installed(package_name: str) -> bool:
__import__(import_name) __import__(import_name)
except ImportError: except ImportError:
return False return False
minimum = MIN_VERSIONS.get(package_name) minimum = _minimum_version(package_name)
if minimum: if minimum:
installed = _installed_version_tuple(package_name) installed = _installed_version_tuple(_dist_name(package_name))
if not installed or installed < minimum: if not installed or installed < minimum:
print(f"{package_name} is installed but below the required " print(f"{package_name} is installed but below the required "
f"{'.'.join(map(str, minimum))}; will upgrade via pip") f"{'.'.join(map(str, minimum))}; will upgrade via pip")
@@ -188,10 +231,11 @@ def main():
continue continue
# Try apt first, then pip. An apt install only counts if it also # Try apt first, then pip. An apt install only counts if it also
# satisfies any minimum version (Debian's python3-freetype can be # satisfies the requirements floor (the apt copies of most of these
# older than the freetype-py pin), otherwise fall through to pip. # are older than the pins on both Bookworm and Trixie), otherwise
# fall through to pip.
ok, apt_output = install_via_apt(package) ok, apt_output = install_via_apt(package)
if ok and package in MIN_VERSIONS and not check_package_installed(package): if ok and _minimum_version(package) and not check_package_installed(package):
ok = False ok = False
apt_output = f"apt version of {package} is below the required minimum" apt_output = f"apt version of {package} is below the required minimum"
if not ok: if not ok:
+25 -5
View File
@@ -15,7 +15,8 @@ The updater leaves data/auto_update_pending.json:
{"status": "pending", "old_head": ..., "new_head": ..., {"status": "pending", "old_head": ..., "new_head": ...,
"old_ref": "main" | "" (detached) | absent (older updaters), "old_ref": "main" | "" (detached) | absent (older updaters),
"display_was_active": bool, "dependency_failures": [...], "created_at": ...} "display_was_active": bool, "dependency_failures": [...],
"units_refreshed": bool (absent from older updaters), "created_at": ...}
This moves its status to "verifying" and then to one of "success", This moves its status to "verifying" and then to one of "success",
"rolled_back" or "rollback_failed", with "reason" and "detail" saying why. "rolled_back" or "rollback_failed", with "reason" and "detail" saying why.
@@ -74,6 +75,10 @@ BASH_CANDIDATES = ('/usr/bin/bash', '/bin/bash')
#: ...and, like it, moves to the next one only when sudo refused the command #: ...and, like it, moves to the next one only when sudo refused the command
#: line (permission_utils.SUDO_REFUSAL_PHRASES), never after pip itself ran. #: line (permission_utils.SUDO_REFUSAL_PHRASES), never after pip itself ran.
SUDO_REFUSAL_PHRASES = ('a password is required', 'is not allowed to run', 'no tty present') SUDO_REFUSAL_PHRASES = ('a password is required', 'is not allowed to run', 'no tty present')
#: The root-owned helper that installed the update's systemd units
#: (web_interface/unit_refresh.py); ``--restore`` puts the previous ones back.
REFRESH_UNITS_PATH = '/usr/local/sbin/ledmatrix-refresh-units'
UNIT_RESTORE_TIMEOUT_SECONDS = 90
#: The longest one health check can take: restart and wait, roll back #: The longest one health check can take: restart and wait, roll back
#: (diff, reset, reinstalls), restart and wait again. A wait's last poll can #: (diff, reset, reinstalls), restart and wait again. A wait's last poll can
@@ -81,7 +86,8 @@ SUDO_REFUSAL_PHRASES = ('a password is required', 'is not allowed to run', 'no t
_WAIT_WORST_SECONDS = (HEALTH_TIMEOUT_SECONDS + STABLE_SECONDS + WEB_CHECK_TIMEOUT_SECONDS _WAIT_WORST_SECONDS = (HEALTH_TIMEOUT_SECONDS + STABLE_SECONDS + WEB_CHECK_TIMEOUT_SECONDS
+ 2 * SYSTEMCTL_QUERY_TIMEOUT_SECONDS + POLL_SECONDS) + 2 * SYSTEMCTL_QUERY_TIMEOUT_SECONDS + POLL_SECONDS)
WORST_CASE_SECONDS = (2 * (2 * RESTART_TIMEOUT_SECONDS + _WAIT_WORST_SECONDS) WORST_CASE_SECONDS = (2 * (2 * RESTART_TIMEOUT_SECONDS + _WAIT_WORST_SECONDS)
+ GIT_TIMEOUT_SECONDS + GIT_RESET_TIMEOUT_SECONDS + PIP_BUDGET_SECONDS) + GIT_TIMEOUT_SECONDS + GIT_RESET_TIMEOUT_SECONDS + UNIT_RESTORE_TIMEOUT_SECONDS
+ PIP_BUDGET_SECONDS)
#: What a command that could not run at all reports: its callers only read #: What a command that could not run at all reports: its callers only read
#: these three fields, the same ones a completed subprocess has. #: these three fields, the same ones a completed subprocess has.
@@ -300,12 +306,26 @@ class Verifier:
if result.returncode != 0: if result.returncode != 0:
return False, (f'"git reset --hard {old}" failed: ' return False, (f'"git reset --hard {old}" failed: '
f'{(result.stderr or result.stdout or "").strip()}') f'{(result.stderr or result.stdout or "").strip()}')
notes = []
# The update also installed its own systemd units: put the previous
# ones back before anything restarts onto the rolled-back code.
if pending.get('units_refreshed') and not self.restore_units():
notes.append('restoring the previous service settings failed; run '
'"sudo ./scripts/install/install_service.sh" in the LEDMatrix folder')
deadline = self.clock() + PIP_BUDGET_SECONDS deadline = self.clock() + PIP_BUDGET_SECONDS
failed = [rel for rel in requirements if not self.install_requirements(rel, deadline)] failed = [rel for rel in requirements if not self.install_requirements(rel, deadline)]
if failed: if failed:
return True, ('reinstalling the previous dependencies from ' + ', '.join(failed) notes.append('reinstalling the previous dependencies from ' + ', '.join(failed)
+ ' failed; run Install Base Requirements from the Tools tab') + ' failed; run Install Base Requirements from the Tools tab')
return True, '' return True, '; '.join(notes)
def restore_units(self):
"""Reinstall the systemd units the update replaced. True on success."""
result = self._run(['sudo', '-n', REFRESH_UNITS_PATH, '--restore'],
timeout=UNIT_RESTORE_TIMEOUT_SECONDS)
if result.returncode != 0:
self.log(f'restoring the previous systemd units failed: {(result.stderr or "").strip()}')
return result.returncode == 0
# -- the check itself ------------------------------------------------- # -- the check itself -------------------------------------------------
+34 -12
View File
@@ -28,6 +28,7 @@ freetype.Face, so it drops straight into DisplayManager.draw_text().
""" """
import logging import logging
import weakref
from collections import OrderedDict from collections import OrderedDict
from dataclasses import dataclass from dataclasses import dataclass
from typing import Any, Dict, List, Optional, Sequence, Tuple, Union from typing import Any, Dict, List, Optional, Sequence, Tuple, Union
@@ -332,9 +333,10 @@ class LayoutContext:
# a plugin fitting changing text (a live game clock, a ticker) on a # a plugin fitting changing text (a live game clock, a ticker) on a
# 24/7 service would otherwise grow this without bound. # 24/7 service would otherwise grow this without bound.
self._fit_cache: "OrderedDict[Any, FitResult]" = OrderedDict() self._fit_cache: "OrderedDict[Any, FitResult]" = OrderedDict()
# LRU-bounded (images are big). Entries hold a strong reference to # LRU-bounded (images are big). An id()-keyed entry watches its
# the source image when keyed by id() so the id can't be recycled # source image through a weak reference and is dropped when the
# out from under the cache. # source is freed (see fit_image), so the id can't be recycled out
# from under the cache and the cache never keeps the source alive.
self._image_cache: "OrderedDict[Any, Tuple[Any, Any]]" = OrderedDict() self._image_cache: "OrderedDict[Any, Tuple[Any, Any]]" = OrderedDict()
_IMAGE_CACHE_MAX = 64 _IMAGE_CACHE_MAX = 64
@@ -536,8 +538,15 @@ class LayoutContext:
cached per (image, box size, options) for this panel size. cached per (image, box size, options) for this panel size.
Prefer a stable ``cache_key`` (e.g. "logo:KC") for images that get Prefer a stable ``cache_key`` (e.g. "logo:KC") for images that get
reloaded — the default id()-based key is safe (the entry pins the reloaded — the default id()-based key misses across reloads of the
source image) but misses across reloads of the same content. same content.
An id()-keyed entry lives only as long as its source image: it holds
a weak reference and is dropped when the source is freed. It used to
pin the source instead, so a plugin passing a freshly loaded image
each frame (``draw_image(Image.open(path), box)``, the documented
one-liner) never hit and kept the last 64 sources alive — ~64MB for
500x500 RGBA team logos, the median size under assets/sports.
""" """
from src.adaptive_images import fit_image as _fit_image from src.adaptive_images import fit_image as _fit_image
@@ -547,18 +556,31 @@ class LayoutContext:
key = ("image", identity, img.size, box_w, box_h, mode, key = ("image", identity, img.size, box_w, box_h, mode,
crop_to_ink, anchor, resample_name, upscale) crop_to_ink, anchor, resample_name, upscale)
cached = self._image_cache.get(key) cache = self._image_cache
if cached is not None: cached = cache.get(key)
self._image_cache.move_to_end(key) # An id()-keyed hit must still be this very image; the callback below
# normally removes a dead source's entry before its id can recur.
if cached is not None and (cache_key is not None or cached[1]() is img):
cache.move_to_end(key)
return cached[0] return cached[0]
result = _fit_image(img, (box_w, box_h), mode=mode, result = _fit_image(img, (box_w, box_h), mode=mode,
crop_to_ink=crop_to_ink, anchor=anchor, crop_to_ink=crop_to_ink, anchor=anchor,
resample=resample, upscale=upscale) resample=resample, upscale=upscale)
# Pin the source only for id()-keyed entries (see docstring). source = None
self._image_cache[key] = (result, img if cache_key is None else None) if cache_key is None:
while len(self._image_cache) > self._IMAGE_CACHE_MAX: def _forget(ref: Any, key: Any = key) -> None:
self._image_cache.popitem(last=False) entry = cache.get(key)
if entry is not None and entry[1] is ref:
cache.pop(key, None)
try:
source = weakref.ref(img, _forget)
except TypeError:
# Not weak-referenceable: pin it, as before.
source = lambda img=img: img # noqa: E731
cache[key] = (result, source)
while len(cache) > self._IMAGE_CACHE_MAX:
cache.popitem(last=False)
return result return result
# ---- text utilities ------------------------------------------------ # ---- text utilities ------------------------------------------------
+3 -2
View File
@@ -213,8 +213,9 @@ def list_installed_plugins(project_root: Path) -> List[Dict[str, Any]]:
The plugins are the ``manifest.json`` files in the configured plugin The plugins are the ``manifest.json`` files in the configured plugin
directory (see :func:`_plugins_directory`), with the manifest's version; directory (see :func:`_plugins_directory`), with the manifest's version;
``enabled`` is config.json's flag by the display's rule (a missing flag ``enabled`` is config.json's flag by the display's rule (a missing flag
is disabled). A restore reinstalls every listed plugin and takes enabled is disabled). A restore installs each listed plugin that is missing and
state from the restored config.json, so ``enabled`` is informational. takes enabled state from the restored config.json, so ``enabled`` is
informational.
``data/plugin_state.json`` is not read: it only ever repeated config's ``data/plugin_state.json`` is not read: it only ever repeated config's
enabled flags and the manifests' versions, and is retired (nothing enabled flags and the manifests' versions, and is retired (nothing
+22 -13
View File
@@ -20,6 +20,7 @@ from typing import Dict, Any, Optional, List, cast
from src.common.api_helper import DEFAULT_HTTP_HEADERS from src.common.api_helper import DEFAULT_HTTP_HEADERS
from src.common.fetch_service import fetch_get, share_connection_pool from src.common.fetch_service import fetch_get, share_connection_pool
from src.common.json_body import response_json
@@ -146,7 +147,7 @@ class BaseOddsManager:
if _is_no_odds_marker(cached_data): if _is_no_odds_marker(cached_data):
self.logger.debug("Cached no-odds marker for %s", cache_key) self.logger.debug("Cached no-odds marker for %s", cache_key)
return None return None
self.logger.debug(f"Using cached odds from ESPN for {cache_key}") self.logger.debug("Using cached odds from ESPN for %s", cache_key)
return cached_data return cached_data
if time.monotonic() < self._skip_network_until: if time.monotonic() < self._skip_network_until:
@@ -159,7 +160,7 @@ class BaseOddsManager:
self._skip_network_until - time.monotonic()) self._skip_network_until - time.monotonic())
return None return None
self.logger.debug(f"Cache miss - fetching fresh odds from ESPN for {cache_key}") self.logger.debug("Cache miss - fetching fresh odds from ESPN for %s", cache_key)
try: try:
# Map league names to ESPN API format # Map league names to ESPN API format
@@ -173,26 +174,30 @@ class BaseOddsManager:
espn_league = league_mapping.get(league, league) espn_league = league_mapping.get(league, league)
url = f"{self.base_url}/{sport}/leagues/{espn_league}/events/{event_id}/competitions/{event_id}/odds" url = f"{self.base_url}/{sport}/leagues/{espn_league}/events/{event_id}/competitions/{event_id}/odds"
self.logger.debug(f"Requesting odds from URL: {url}") self.logger.debug("Requesting odds from URL: %s", url)
# The response cache may answer only inside this caller's own # The response cache may answer only inside this caller's own
# interval, the age at which its cached odds expire anyway. # interval, the age at which its cached odds expire anyway.
response = fetch_get(self.session, url, timeout=self.request_timeout, response = fetch_get(self.session, url, timeout=self.request_timeout,
cache_max_age=interval) cache_max_age=interval)
response.raise_for_status() response.raise_for_status()
raw_data = response.json() raw_data = response_json(response)
self._skip_network_until = 0.0 # reachable again self._skip_network_until = 0.0 # reachable again
self.logger.debug(f"Received raw odds data from ESPN: {json.dumps(raw_data, indent=2)}") # Guarded, not just %-style: the json.dumps argument would still be
# built for every response with DEBUG off.
if self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug("Received raw odds data from ESPN: %s",
json.dumps(raw_data, indent=2))
odds_data = self._extract_espn_data(raw_data) odds_data = self._extract_espn_data(raw_data)
if odds_data: if odds_data:
self.logger.debug(f"Successfully extracted odds data: {odds_data}") self.logger.debug("Successfully extracted odds data: %s", odds_data)
self.cache_manager.set(cache_key, odds_data, ttl=interval) self.cache_manager.set(cache_key, odds_data, ttl=interval)
self.logger.debug(f"Saved odds data to cache for {cache_key} with TTL {interval}s") self.logger.debug("Saved odds data to cache for %s with TTL %ss", cache_key, interval)
else: else:
self.logger.debug(f"No odds data available for {cache_key}") self.logger.debug("No odds data available for %s", cache_key)
# Cache the absence too, so the game is not re-requested # Cache the absence too, so the game is not re-requested
# on every update until the interval passes. # on every update until the interval passes.
self.cache_manager.set(cache_key, {"no_odds": True}, ttl=interval) self.cache_manager.set(cache_key, {"no_odds": True}, ttl=interval)
@@ -226,12 +231,12 @@ class BaseOddsManager:
Returns: Returns:
Formatted odds data dictionary or None Formatted odds data dictionary or None
""" """
self.logger.debug(f"Extracting ESPN odds data. Data keys: {list(data.keys())}") self.logger.debug("Extracting ESPN odds data. Data keys: %s", list(data.keys()))
if "items" in data and data["items"]: if "items" in data and data["items"]:
self.logger.debug(f"Found {len(data['items'])} items in odds data") self.logger.debug("Found %d items in odds data", len(data['items']))
item = data["items"][0] item = data["items"][0]
self.logger.debug(f"First item keys: {list(item.keys())}") self.logger.debug("First item keys: %s", list(item.keys()))
# The ESPN API returns odds data directly in the item, not in a # The ESPN API returns odds data directly in the item, not in a
# providers array. ESPN sends explicit JSON nulls for absent # providers array. ESPN sends explicit JSON nulls for absent
@@ -254,13 +259,17 @@ class BaseOddsManager:
.get("pointSpread") or {}).get("value") .get("pointSpread") or {}).get("value")
} }
} }
self.logger.debug(f"Returning extracted odds data: {json.dumps(extracted_data, indent=2)}") if self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug("Returning extracted odds data: %s",
json.dumps(extracted_data, indent=2))
return extracted_data return extracted_data
# Check if this is a valid empty response or an unexpected structure # Check if this is a valid empty response or an unexpected structure
if "count" in data and data["count"] == 0 and "items" in data and data["items"] == []: if "count" in data and data["count"] == 0 and "items" in data and data["items"] == []:
# This is a valid empty response - no odds available for this game # This is a valid empty response - no odds available for this game
self.logger.debug(f"No odds available for this game. Response: {json.dumps(data, indent=2)}") if self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug("No odds available for this game. Response: %s",
json.dumps(data, indent=2))
return None return None
else: else:
# This is an unexpected response structure # This is an unexpected response structure
+221 -42
View File
@@ -4,6 +4,7 @@ Disk Cache
Handles persistent disk-based caching with atomic writes and error recovery. Handles persistent disk-based caching with atomic writes and error recovery.
""" """
import hashlib
import json import json
import math import math
import os import os
@@ -14,7 +15,7 @@ import tempfile
import logging import logging
import threading import threading
import zlib import zlib
from typing import Dict, Any, Optional, Protocol from typing import Dict, Any, Optional, Protocol, Tuple
from datetime import datetime from datetime import datetime
from src.common.path_safety import safe_path_component from src.common.path_safety import safe_path_component
@@ -31,6 +32,35 @@ except ImportError: # pragma: no cover - exercised on hosts without the wheel
# useful, and a half-written file was never useful. # useful, and a half-written file was never useful.
_ORPHAN_TEMP_MAX_AGE_SECONDS = 3600 _ORPHAN_TEMP_MAX_AGE_SECONDS = 3600
# Longest key, in UTF-8 bytes, used verbatim as a filename stem. ext4 caps a
# name at 255 bytes and set()'s temp file is ".<stem>.json.<8 random>", 15
# bytes longer than the stem, so anything near the cap could never be written:
# the calendar plugin's key joins every calendar id and passed 300 bytes on a
# real install, failing every write with ENAMETOOLONG. Longer keys keep this
# many bytes as a readable prefix and end in a hash of the whole key.
_MAX_KEY_FILENAME_BYTES = 200
_KEY_HASH_CHARS = 16
def _filename_stem(key: str) -> str:
"""The filename stem for a key that is already a safe path component.
Short keys are used as they are, so every file already on disk keeps its
name. A long one becomes its first bytes plus a hash of the full key: the
prefix keeps the stem recognisable (and keeps the data-type words that
cleanup's retention lookup reads from it), the hash keeps two keys that
share a long prefix apart. The result is itself short, so a stem read back
from a filename -- which is how the web UI names a key it deletes -- maps to
the same file.
"""
encoded = key.encode('utf-8')
if len(encoded) <= _MAX_KEY_FILENAME_BYTES:
return key
digest = hashlib.sha256(encoded).hexdigest()[:_KEY_HASH_CHARS]
keep = _MAX_KEY_FILENAME_BYTES - _KEY_HASH_CHARS - 1
prefix = encoded[:keep].decode('utf-8', errors='ignore')
return f"{prefix}-{digest}"
class CacheStrategyProtocol(Protocol): class CacheStrategyProtocol(Protocol):
@@ -111,18 +141,91 @@ _HEAD_RE = re.compile(
) )
def _stale_from_head(head: bytes, max_age: Optional[int], now: float) -> bool: # UNCHANGED RE-SAVES: THE FILE'S MTIME CARRIES THE NEWER TIMESTAMP
# ----------------------------------------------------------------
# Plugins re-save unchanged API data every update cycle, and every one of
# those saves was a full rewrite on the SD card. DiskCache.set skips the write
# when the payload matches the last one it wrote for the key -- but
# CacheManager.set stamps each record with time.time(), so for set() the
# payload never matched and the skip never fired.
#
# The digest now leaves out a header-first record's timestamp, so an unchanged
# set() is skipped. What the skip must not do is make the record look older
# than it is: the timestamp inside the file is from the last real write, and
# a reader in another process (the web interface, with memory_ttl=0) or after
# a restart would call fresh data stale. So the newer timestamp goes where it
# costs no data write -- the file's mtime -- and readers take a record's age
# from the newer of the two. The invariant that makes that safe:
#
# a file's mtime is the timestamp of the newest record saved for its key
#
# real write mtime is set to the record's own timestamp, so a record saved
# with an old timestamp (data as of some earlier time) cannot
# borrow freshness from the moment it hit the disk
# skip mtime is set to the skipped record's timestamp -- exactly what
# a rewrite would have stored, without the rewrite
#
# Readers of the on-disk timestamp, all of which go through _effective_timestamp:
# DiskCache.get (the header check and the full parse; it also returns the
# record with 'timestamp' set to the effective value, so CacheManager.get's
# max_age path, the memory tier hydrated from disk, and any plugin reading
# record['timestamp'] all see it). Readers that use mtime alone already see the
# newer value: the retention sweep below, CacheManager.list_cache_files (the
# web UI's cache list). Nothing else opens cache files: web_interface and
# scripts reach them only through CacheManager.
#
# Something other than this class can also move an mtime forward -- a copy
# without -p, an rsync without -t, a `touch`. (backup_manager.py does not
# back up or restore the cache directory, so the in-tree restore cannot.) That
# must not make old data fresh, so the lift is bounded: a reader never takes
# the mtime as more than _MAX_TIMESTAMP_LIFT past the embedded timestamp, and
# set() rewrites the file for real once a skip would need more than that, so
# an honest lift never reaches the bound. A file copied a day after it was
# written therefore reads at most an hour fresher than its contents say, and a
# 30-second live-score record from yesterday stays stale. CacheManager.set
# records written before this change have mtime == write time == embedded
# timestamp, give or take the write itself, and read exactly as before; a
# file an older version wrote or touched later than its embedded timestamp
# says reads at most the same hour fresher, once, until it is next saved.
#: Longest a skipped write may stand in for a real one, and so the furthest a
#: file's mtime is ever trusted past the record's own timestamp. Unchanged data
#: is rewritten at least this often, at most once an hour per key instead of
#: once per update cycle.
_MAX_TIMESTAMP_LIFT = 3600.0
def _record_timestamp(value: Any) -> Optional[float]:
"""A record's timestamp as a finite float, or None if it has no usable one."""
if isinstance(value, bool) or not isinstance(value, (int, float)):
return None
value = float(value)
return value if math.isfinite(value) else None
def _effective_timestamp(embedded: float, mtime: Optional[float]) -> float:
"""When a record was last saved: its timestamp, or the file's mtime if a
later unchanged save moved that forward -- never by more than
_MAX_TIMESTAMP_LIFT. See "UNCHANGED RE-SAVES" above."""
if mtime is None:
return embedded
return max(embedded, min(mtime, embedded + _MAX_TIMESTAMP_LIFT))
def _stale_from_head(head: bytes, max_age: Optional[int], now: float,
mtime: Optional[float] = None) -> bool:
"""True when a record's header alone shows it has expired. """True when a record's header alone shows it has expired.
Mirrors the expiry rule in DiskCache.get: a per-entry ttl wins over the Mirrors the expiry rule in DiskCache.get: a per-entry ttl wins over the
caller's max_age, and no limit at all means never stale. False whenever the caller's max_age, and no limit at all means never stale. False whenever the
header cannot be read, so the full parse decides as it always did. header cannot be read, so the full parse decides as it always did. ``mtime``
is the file's, which may carry a newer save than the header does.
""" """
match = _HEAD_RE.match(head) match = _HEAD_RE.match(head)
if not match: if not match:
return False return False
try: try:
timestamp = float(match.group(1)) timestamp = _effective_timestamp(float(match.group(1)), mtime)
limit = max_age limit = max_age
if match.group(2) is not None: if match.group(2) is not None:
ttl = float(match.group(2)) ttl = float(match.group(2))
@@ -179,7 +282,7 @@ else:
# -------------------------------------------- # --------------------------------------------
# The display service runs as root and the web interface as the installing # The display service runs as root and the web interface as the installing
# user, and the web interface reads records only the display writes # user, and the web interface reads records only the display writes
# (display_current_state, display_on_demand_state, plugin_metrics:*). Files are # (display_current_state, display_on_demand_state, plugin_metrics_snapshot). Files are
# written 0660, so the web interface can read one only through its group. # written 0660, so the web interface can read one only through its group.
# #
# The installers rely on the directory's setgid bit to set that group. That is # The installers rely on the directory's setgid bit to set that group. That is
@@ -248,11 +351,14 @@ class DiskCache:
self.cache_dir = cache_dir self.cache_dir = cache_dir
self.logger = logger or logging.getLogger(__name__) self.logger = logger or logging.getLogger(__name__)
self._lock = threading.Lock() self._lock = threading.Lock()
# key -> adler32 of the last payload successfully written to the # key -> what set() last put at the primary cache path: the adler32 of
# primary cache path; lets set() skip rewriting identical data # the payload (less a header-first timestamp), the timestamp the file
# (per-process only — worst case another process rewrites, never # holds (None for records without one), and the file's inode and size.
# a missed write). Guarded by _lock. # Lets set() skip rewriting identical data. Per-process only, and the
self._write_digests: Dict[str, int] = {} # inode/size check means another process's write is never mistaken
# for ours -- worst case a redundant write, never a missed one.
# Guarded by _lock.
self._write_digests: Dict[str, Tuple[int, Optional[float], int, int]] = {}
def get_cache_path(self, key: str) -> Optional[str]: def get_cache_path(self, key: str) -> Optional[str]:
""" """
@@ -267,6 +373,8 @@ class DiskCache:
derives them), so rejecting anything with a path component turns derives them), so rejecting anything with a path component turns
away only inputs that could never have been written here. away only inputs that could never have been written here.
A key too long to be a filename is shortened by _filename_stem.
Args: Args:
key: Cache key key: Cache key
@@ -280,7 +388,7 @@ class DiskCache:
if safe_key is None: if safe_key is None:
self.logger.warning("Rejected unsafe cache key %r", key) self.logger.warning("Rejected unsafe cache key %r", key)
return None return None
return os.path.join(self.cache_dir, f"{safe_key}.json") return os.path.join(self.cache_dir, f"{_filename_stem(safe_key)}.json")
def get(self, key: str, max_age: Optional[int] = 300) -> Optional[Dict[str, Any]]: def get(self, key: str, max_age: Optional[int] = 300) -> Optional[Dict[str, Any]]:
""" """
@@ -301,32 +409,41 @@ class DiskCache:
try: try:
with self._lock: with self._lock:
with open(cache_path, 'rb') as f: with open(cache_path, 'rb') as f:
# The open file's mtime, not the path's: the file a skipped
# write touched is the one being read.
mtime = os.fstat(f.fileno()).st_mtime
# Decide staleness from the header before paying for the # Decide staleness from the header before paying for the
# parse. A stale read is the common case for the biggest # parse. A stale read is the common case for the biggest
# records (a season schedule is re-fetched when its cache # records (a season schedule is re-fetched when its cache
# expires), and parsing 53MB to throw it away held the GIL # expires), and parsing 53MB to throw it away held the GIL
# for ~1.8s -- a visible freeze on the panel. # for ~1.8s -- a visible freeze on the panel.
if _stale_from_head(f.read(_HEAD_BYTES), max_age, time.time()): if _stale_from_head(f.read(_HEAD_BYTES), max_age, time.time(), mtime):
return None return None
f.seek(0) f.seek(0)
record = _loads(f.read()) record = _loads(f.read())
# Determine record timestamp (prefer embedded, else file mtime) # Determine record timestamp: the embedded one, moved forward by a
# later unchanged save if there was one (see "UNCHANGED RE-SAVES"),
# else the file mtime.
record_ts = None record_ts = None
if isinstance(record, dict): if isinstance(record, dict):
record_ts = record.get('timestamp') record_ts = record.get('timestamp')
if record_ts is None: if record_ts is None:
try: record_ts = mtime
record_ts = os.path.getmtime(cache_path) else:
except OSError: embedded_ts = _record_timestamp(record_ts)
record_ts = None if embedded_ts is None:
try:
if record_ts is not None: record_ts = float(record_ts)
try: except (TypeError, ValueError):
record_ts = float(record_ts) record_ts = None
except (TypeError, ValueError): else:
record_ts = None record_ts = _effective_timestamp(embedded_ts, mtime)
if record_ts != embedded_ts:
# Hand the record back as a rewrite would have left it,
# so callers that age it themselves agree with us.
record['timestamp'] = record_ts
now = time.time() now = time.time()
# An explicit per-entry ttl wins over the caller's max_age. The # An explicit per-entry ttl wins over the caller's max_age. The
@@ -403,7 +520,12 @@ class DiskCache:
self.logger.warning("Cache data for key '%s' not serializable: %s", key, e) self.logger.warning("Cache data for key '%s' not serializable: %s", key, e)
return return
digest = zlib.adler32(payload) timestamp = _record_timestamp(data.get('timestamp')) if isinstance(data, dict) else None
# A header-first record (CacheManager.set's layout) is compared without
# its timestamp, which differs on every save; see "UNCHANGED RE-SAVES".
# Any other layout is compared whole, as before.
head = _HEAD_RE.match(payload) if timestamp is not None else None
digest = zlib.adler32(memoryview(payload)[head.end(1):] if head else payload)
try: try:
# Atomic write to avoid partial/corrupt files # Atomic write to avoid partial/corrupt files
@@ -411,16 +533,10 @@ class DiskCache:
# Skip the disk entirely when this exact payload was already # Skip the disk entirely when this exact payload was already
# written for this key (plugins re-save unchanged API data # written for this key (plugins re-save unchanged API data
# every update cycle — each write is real SD-card wear). # every update cycle — each write is real SD-card wear).
# Refresh the file mtime so records that rely on it for TTL # A metadata touch is journal-cheap compared to rewriting
# (no embedded 'timestamp') don't expire early; a metadata # the data.
# touch is journal-cheap compared to rewriting the data. if self._skip_unchanged(key, cache_path, digest, timestamp):
if self._write_digests.get(key) == digest: return
try:
os.utime(cache_path, None)
return
except OSError:
# File vanished or perms changed — fall through and write
self._write_digests.pop(key, None)
tmp_dir = os.path.dirname(cache_path) tmp_dir = os.path.dirname(cache_path)
# Try to create temp file in cache directory first # Try to create temp file in cache directory first
@@ -458,7 +574,7 @@ class DiskCache:
# opened it in between was refused. # opened it in between was refused.
_share_open_file(tmp_file.fileno(), _shared_group(tmp_dir)) _share_open_file(tmp_file.fileno(), _shared_group(tmp_dir))
os.replace(tmp_path, cache_path) os.replace(tmp_path, cache_path)
self._write_digests[key] = digest self._remember_write(key, cache_path, digest, timestamp)
finally: finally:
if os.path.exists(tmp_path): if os.path.exists(tmp_path):
try: try:
@@ -471,13 +587,13 @@ class DiskCache:
with open(cache_path, 'wb') as cache_file: with open(cache_path, 'wb') as cache_file:
cache_file.write(payload) cache_file.write(payload)
_share_open_file(cache_file.fileno(), _shared_group(tmp_dir)) _share_open_file(cache_file.fileno(), _shared_group(tmp_dir))
self._write_digests[key] = digest self._remember_write(key, cache_path, digest, timestamp)
self.logger.debug("Wrote cache for %s directly (non-atomic)", key) self.logger.debug("Wrote cache for %s directly (non-atomic)", key)
except (IOError, OSError, PermissionError) as write_error: except (IOError, OSError, PermissionError) as write_error:
# If direct write also fails, try fallback location # If direct write also fails, try fallback location
self.logger.warning("Direct write failed for key '%s' to %s: %s", key, cache_path, write_error) self.logger.warning("Direct write failed for key '%s' to %s: %s", key, cache_path, write_error)
raise # Re-raise to trigger fallback logic raise # Re-raise to trigger fallback logic
except (IOError, OSError, PermissionError): except (IOError, OSError, PermissionError) as primary_error:
# Attempt one-time fallback write to user's home cache directory # Attempt one-time fallback write to user's home cache directory
try: try:
# Try user's home cache directory as fallback # Try user's home cache directory as fallback
@@ -503,11 +619,14 @@ class DiskCache:
self.logger.debug("Fallback cache write also failed for key '%s': %s", key, e2) self.logger.debug("Fallback cache write also failed for key '%s': %s", key, e2)
# If all write attempts failed, log warning but don't raise exception # If all write attempts failed, log warning but don't raise exception
# Cache is a performance optimization, not critical for operation # Cache is a performance optimization, not critical for operation.
# Name the real error: this used to say "permission denied"
# whatever happened, which sent a too-long filename off to
# be debugged as a directory-ownership problem.
self.logger.warning( self.logger.warning(
"Could not write cache for key '%s' to %s (permission denied). " "Could not write cache for key '%s' to %s (%s). "
"Cache will be unavailable for this key, but application will continue.", "Cache will be unavailable for this key, but application will continue.",
key, cache_path key, cache_path, primary_error.strerror or primary_error
) )
return # Exit gracefully without raising exception return # Exit gracefully without raising exception
@@ -520,6 +639,66 @@ class DiskCache:
) )
return # Exit gracefully without raising exception return # Exit gracefully without raising exception
def _skip_unchanged(self, key: str, cache_path: str, digest: int,
timestamp: Optional[float]) -> bool:
"""Stand in for a write of an unchanged record by touching the file.
True when the file already holds this record bar its timestamp and the
touch landed; False means write it. The touch sets mtime to the
record's timestamp -- what a rewrite would have stored -- or to now
for a record without one, whose age readers already take from mtime.
Caller holds _lock.
"""
last = self._write_digests.get(key)
if last is None or last[0] != digest:
return False
_, written_ts, ino, size = last
if (timestamp is None) != (written_ts is None):
return False
if timestamp is not None and written_ts is not None:
# Never backwards (a rewrite would make the record older), and
# never further than readers will trust the mtime: past that the
# record is rewritten, so its own timestamp catches up.
if not written_ts <= timestamp <= written_ts + _MAX_TIMESTAMP_LIFT:
return False
try:
st = os.stat(cache_path)
if (st.st_ino, st.st_size) != (ino, size):
# Replaced since our write (another process, a restore):
# its contents are not the ones the digest describes.
self._write_digests.pop(key, None)
return False
# Setting an explicit time needs the file's owner; a file someone
# else wrote fails here and is rewritten (as our own file) instead.
os.utime(cache_path, None if timestamp is None else (timestamp, timestamp))
return True
except OSError:
# File vanished or perms changed — fall through and write
self._write_digests.pop(key, None)
return False
def _remember_write(self, key: str, cache_path: str, digest: int,
timestamp: Optional[float]) -> None:
"""After a real write: pin mtime to the record's timestamp and note
what was written, so the next unchanged save can be skipped.
Pinning keeps a record saved with an older timestamp from looking as
fresh as the moment it was written (see "UNCHANGED RE-SAVES"); for
CacheManager.set's records the two differ only by the write itself.
A timestamp in the future is left alone, mtime already being older.
Never raises: the data is on disk, and anything failing here only
costs the next save its skip. Caller holds _lock.
"""
self._write_digests.pop(key, None)
try:
if timestamp is not None and timestamp <= time.time():
os.utime(cache_path, (timestamp, timestamp))
st = os.stat(cache_path)
except OSError as e:
self.logger.debug("Could not pin mtime of %s: %s", cache_path, e)
return
self._write_digests[key] = (digest, timestamp, st.st_ino, st.st_size)
def clear(self, key: Optional[str] = None) -> None: def clear(self, key: Optional[str] = None) -> None:
""" """
Clear cache entry or all entries. Clear cache entry or all entries.
+83 -12
View File
@@ -43,6 +43,35 @@ from src.logging_config import get_logger
# it from either path. # it from either path.
from src.cache.disk_cache import DateTimeEncoder # noqa: F401 - deliberate re-export from src.cache.disk_cache import DateTimeEncoder # noqa: F401 - deliberate re-export
# CacheManager.config_manager not built yet (None means "not available").
_UNSET: Any = object()
def _outlived(record: Any, max_age: Optional[float], now: float) -> bool:
"""Whether a record's own timestamp puts it past max_age.
The memory tier times an entry from when it was put there, and a record
loaded from disk is put there when it is read, not when it was written: a
record 290 s old, read after a restart, could be served for another
max_age from memory. This is the age check DiskCache.get makes, with the
same rule that a stored ttl wins over the caller's max_age. A record that
carries no timestamp is left to the memory tier's own clock.
"""
if not isinstance(record, dict):
return False
stored_ttl = record.get('ttl')
if isinstance(stored_ttl, (int, float)) and not isinstance(stored_ttl, bool) \
and stored_ttl >= 0:
max_age = stored_ttl
stamp = record.get('timestamp')
if max_age is None or stamp is None or isinstance(stamp, bool):
return False
try:
return now - float(stamp) > max_age
except (TypeError, ValueError):
return False
class CacheManager: class CacheManager:
"""Manages caching of API responses to reduce API calls.""" """Manages caching of API responses to reduce API calls."""
@@ -73,21 +102,19 @@ class CacheManager:
self.logger.error("Could not find or create a writable cache directory. Caching will be disabled.") self.logger.error("Could not find or create a writable cache directory. Caching will be disabled.")
self.cache_dir = None self.cache_dir = None
# Initialize config manager for sport-specific intervals # The config manager is built on first use of self.config_manager; see
try: # the property. Nothing in the cache reads it any more.
from src.config_manager import ConfigManager self._config_manager: Any = _UNSET
self.config_manager: Optional[Any] = ConfigManager() self._config_manager_lock = threading.Lock()
self.config_manager.load_config()
except ImportError:
self.config_manager: Optional[Any] = None
self.logger.warning("ConfigManager not available, using default cache intervals")
# Initialize cache components using composition # Initialize cache components using composition
self._memory_cache_component = MemoryCache( self._memory_cache_component = MemoryCache(
max_size=default_max_size(), cleanup_interval=300.0 max_size=default_max_size(), cleanup_interval=300.0
) )
self._disk_cache_component = DiskCache(cache_dir=self.cache_dir, logger=self.logger) self._disk_cache_component = DiskCache(cache_dir=self.cache_dir, logger=self.logger)
self._strategy_component = CacheStrategy(config_manager=self.config_manager, logger=self.logger) # No config manager: CacheStrategy keeps the parameter for callers but
# reads nothing from it, and passing ours would build it eagerly.
self._strategy_component = CacheStrategy(logger=self.logger)
self._metrics_component = CacheMetrics(logger=self.logger) self._metrics_component = CacheMetrics(logger=self.logger)
# Disk cleanup configuration # Disk cleanup configuration
@@ -115,6 +142,44 @@ class CacheManager:
if self.cache_dir: if self.cache_dir:
self.start_cleanup_thread() self.start_cleanup_thread()
@property
def config_manager(self) -> Optional[Any]:
"""A loaded ConfigManager, built the first time it is asked for.
Every CacheManager used to build one and load the whole config in
__init__, for a cache strategy that stopped reading it -- startup paid
a config load (and the web interface another) per manager for nothing.
It is still public: the sports plugins resolve the global timezone and
display settings through ``cache_manager.config_manager``, and they get
the same object they always did, on first access instead of at
construction. None when ConfigManager cannot be imported, as before.
Assigning replaces it, as assigning the attribute always did.
"""
# getattr: a manager made with __new__ (some tests) has no slot yet.
value = getattr(self, '_config_manager', _UNSET)
if value is not _UNSET:
return value
lock = getattr(self, '_config_manager_lock', None) or threading.Lock()
with lock:
value = getattr(self, '_config_manager', _UNSET)
if value is _UNSET:
try:
from src.config_manager import ConfigManager
except ImportError:
self.logger.warning("ConfigManager not available, using default cache intervals")
value = None
else:
value = ConfigManager()
# Raises as it did from __init__; nothing is kept, so the
# next access tries again.
value.load_config()
self._config_manager = value
return value
@config_manager.setter
def config_manager(self, value: Optional[Any]) -> None:
self._config_manager = value
def _get_writable_cache_dir(self) -> Optional[str]: def _get_writable_cache_dir(self) -> Optional[str]:
"""Tries to find or create a writable cache directory, preferring a system path when available.""" """Tries to find or create a writable cache directory, preferring a system path when available."""
# Attempt 1: System-wide persistent cache directory (preferred for services) # Attempt 1: System-wide persistent cache directory (preferred for services)
@@ -245,7 +310,11 @@ class CacheManager:
# 1) Memory cache # 1) Memory cache
cached = self._memory_cache_component.get(key, max_age=in_memory_ttl) cached = self._memory_cache_component.get(key, max_age=in_memory_ttl)
if cached is not None: if cached is not None:
return cached if not _outlived(cached, max_age, time.time()):
return cached
# Too old for this reader. Disk may hold a newer write (from the
# other process), and if it does not, the miss is the right answer.
self._memory_cache_component.clear(key)
# 2) Disk cache # 2) Disk cache
record = self._disk_cache_component.get(key, max_age=max_age) record = self._disk_cache_component.get(key, max_age=max_age)
@@ -279,7 +348,9 @@ class CacheManager:
# Check memory cache first (1 minute TTL) # Check memory cache first (1 minute TTL)
cached = self._memory_cache_component.get(key, max_age=60) cached = self._memory_cache_component.get(key, max_age=60)
if cached is not None: if cached is not None:
return cached if not _outlived(cached, 3600, time.time()):
return cached
self._memory_cache_component.clear(key)
# Check disk cache # Check disk cache
data = self._disk_cache_component.get(key, max_age=3600) # 1 hour for load_cache data = self._disk_cache_component.get(key, max_age=3600) # 1 hour for load_cache
+8 -3
View File
@@ -17,7 +17,9 @@ Rules for the package:
- `from src.common import ...` re-exports `APIHelper`, `ScrollHelper`, - `from src.common import ...` re-exports `APIHelper`, `ScrollHelper`,
`LogoHelper`, `TextHelper`, `scroll_config` (plus `ScrollSettings`, `LogoHelper`, `TextHelper`, `scroll_config` (plus `ScrollSettings`,
`configure_scroll`, `resolve_scroll_settings`, `refresh_hz_from_config`) and `configure_scroll`, `resolve_scroll_settings`, `refresh_hz_from_config`) and
the adaptive layout names below ([`__init__.py`](__init__.py)). the adaptive layout names below ([`__init__.py`](__init__.py)). Each is
imported on first use, so `import src.common` or a submodule import stays
cheap; add a new re-export to `_LAZY` there as well as `__all__`.
## Summary ## Summary
@@ -107,8 +109,11 @@ and the plugin test harness all use it. Most plugins get BDF text through
[`espn_dates.py`](espn_dates.py). ESPN's site API rejects `dates=` ranges [`espn_dates.py`](espn_dates.py). ESPN's site API rejects `dates=` ranges
and truncates results when `limit` is above 500. `fetch_espn_scoreboard()` and truncates results when `limit` is above 500. `fetch_espn_scoreboard()`
splits a range into month and day requests ESPN accepts and merges the splits a range into month and day requests ESPN accepts and merges the
results; `espn_date_chunks()`, `fetch_espn_date_chunks()`, results; `espn_date_chunks()`, `espn_request_chunks()`,
`clamp_espn_limit()` and `merge_scoreboard_payloads()` are the pieces. `fetch_espn_date_chunks()`, `clamp_espn_limit()` and
`merge_scoreboard_payloads()` are the pieces. A window's partial edge months
are asked whole and trimmed to its days (US Eastern), and chunk requests share
one process-wide cap of `ESPN_CHUNK_WORKERS` in flight.
Every request goes through [`fetch_service`](#fetch_service), the chunks Every request goes through [`fetch_service`](#fetch_service), the chunks
counted against the plugin that asked. Scoreboard plugins also bundle a copy counted against the plugin that asked. Scoreboard plugins also bundle a copy
for older cores. for older cores.
+103 -36
View File
@@ -6,45 +6,90 @@ This package provides reusable functionality for plugins and core modules:
- Logo helpers - Logo helpers
- Text/scroll helpers - Text/scroll helpers
- Adaptive layout and image helpers - Adaptive layout and image helpers
The names below are imported on first use (PEP 562), not when the package is
imported. ``from src.common import ScrollHelper`` and
``src.common.ScrollHelper`` work as before and return the same objects, but
``import src.common`` -- or importing any submodule, such as
``src.common.path_safety`` -- no longer loads numpy, requests and freetype
along with every helper. The web interface imports src.common only for a few
small modules and never needs those.
""" """
# Export commonly used utilities import importlib
from src.common.api_helper import APIHelper from typing import TYPE_CHECKING, Any, Dict, List, Optional, Tuple
from src.common.scroll_helper import ScrollHelper
from src.common import scroll_config
from src.common.scroll_config import (
ScrollSettings,
configure as configure_scroll,
resolve as resolve_scroll_settings,
refresh_hz_from_config,
)
from src.common.logo_helper import LogoHelper
from src.common.text_helper import TextHelper
# Adaptive layout & images (canonical homes: src.adaptive_layout / if TYPE_CHECKING:
# src.adaptive_images — re-exported here so plugin authors find them in the # What mypy and editors see: the real names and their types.
# blessed-helpers package). See docs/ADAPTIVE_LAYOUT.md. from src.common.api_helper import APIHelper
from src.adaptive_layout import ( from src.common.scroll_helper import ScrollHelper
Region, from src.common import scroll_config
LayoutContext, from src.common.scroll_config import (
FontStep, ScrollSettings,
FontLadder, configure as configure_scroll,
LADDER_GRID, resolve as resolve_scroll_settings,
LADDER_ARCADE, refresh_hz_from_config,
FitResult, )
draw_fitted_text, from src.common.logo_helper import LogoHelper
ScoreboardRegions, from src.common.text_helper import TextHelper
scoreboard_regions,
MediaRow, # Adaptive layout & images (canonical homes: src.adaptive_layout /
media_row, # src.adaptive_images — re-exported here so plugin authors find them in the
) # blessed-helpers package). See docs/ADAPTIVE_LAYOUT.md.
from src.adaptive_images import ( from src.adaptive_layout import (
ImageFitResult, Region,
fit_image, LayoutContext,
draw_fitted_image, FontStep,
RESAMPLE_LANCZOS, FontLadder,
RESAMPLE_NEAREST, LADDER_GRID,
) LADDER_ARCADE,
FitResult,
draw_fitted_text,
ScoreboardRegions,
scoreboard_regions,
MediaRow,
media_row,
)
from src.adaptive_images import (
ImageFitResult,
fit_image,
draw_fitted_image,
RESAMPLE_LANCZOS,
RESAMPLE_NEAREST,
)
#: Exported name -> (module it lives in, attribute name there). An attribute
#: of None means the name is the module itself. Keep in step with the
#: TYPE_CHECKING imports above and with __all__.
_LAZY: Dict[str, Tuple[str, Optional[str]]] = {
'APIHelper': ('src.common.api_helper', 'APIHelper'),
'ScrollHelper': ('src.common.scroll_helper', 'ScrollHelper'),
'scroll_config': ('src.common.scroll_config', None),
'ScrollSettings': ('src.common.scroll_config', 'ScrollSettings'),
'configure_scroll': ('src.common.scroll_config', 'configure'),
'resolve_scroll_settings': ('src.common.scroll_config', 'resolve'),
'refresh_hz_from_config': ('src.common.scroll_config', 'refresh_hz_from_config'),
'LogoHelper': ('src.common.logo_helper', 'LogoHelper'),
'TextHelper': ('src.common.text_helper', 'TextHelper'),
# adaptive layout & images
'Region': ('src.adaptive_layout', 'Region'),
'LayoutContext': ('src.adaptive_layout', 'LayoutContext'),
'FontStep': ('src.adaptive_layout', 'FontStep'),
'FontLadder': ('src.adaptive_layout', 'FontLadder'),
'LADDER_GRID': ('src.adaptive_layout', 'LADDER_GRID'),
'LADDER_ARCADE': ('src.adaptive_layout', 'LADDER_ARCADE'),
'FitResult': ('src.adaptive_layout', 'FitResult'),
'draw_fitted_text': ('src.adaptive_layout', 'draw_fitted_text'),
'ScoreboardRegions': ('src.adaptive_layout', 'ScoreboardRegions'),
'scoreboard_regions': ('src.adaptive_layout', 'scoreboard_regions'),
'MediaRow': ('src.adaptive_layout', 'MediaRow'),
'media_row': ('src.adaptive_layout', 'media_row'),
'ImageFitResult': ('src.adaptive_images', 'ImageFitResult'),
'fit_image': ('src.adaptive_images', 'fit_image'),
'draw_fitted_image': ('src.adaptive_images', 'draw_fitted_image'),
'RESAMPLE_LANCZOS': ('src.adaptive_images', 'RESAMPLE_LANCZOS'),
'RESAMPLE_NEAREST': ('src.adaptive_images', 'RESAMPLE_NEAREST'),
}
__all__ = [ __all__ = [
'APIHelper', 'APIHelper',
@@ -75,3 +120,25 @@ __all__ = [
'RESAMPLE_LANCZOS', 'RESAMPLE_LANCZOS',
'RESAMPLE_NEAREST', 'RESAMPLE_NEAREST',
] ]
def __getattr__(name: str) -> Any:
"""Import an exported name on first access (PEP 562).
Only called for names not already in the module namespace, so after the
first access the cached value below is returned directly. Unknown names
raise AttributeError, which ``from src.common import <submodule>`` relies
on to fall through to importing the submodule.
"""
try:
module_name, attr = _LAZY[name]
except KeyError:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}") from None
module = importlib.import_module(module_name) # nosemgrep: python.lang.security.audit.non-literal-import.non-literal-import -- module_name comes from the fixed _LAZY table
value = module if attr is None else getattr(module, attr)
globals()[name] = value
return value
def __dir__() -> List[str]:
return sorted(set(globals()) | set(__all__))
+3 -2
View File
@@ -17,6 +17,7 @@ from src.common.espn_dates import (
store_espn_scoreboard_cache, store_espn_scoreboard_cache,
) )
from src.common.fetch_service import fetch_get, fetch_post, share_connection_pool from src.common.fetch_service import fetch_get, fetch_post, share_connection_pool
from src.common.json_body import response_json
from typing import TYPE_CHECKING, Any, Dict, Mapping, Optional, cast from typing import TYPE_CHECKING, Any, Dict, Mapping, Optional, cast
import requests import requests
@@ -157,7 +158,7 @@ class APIHelper:
response.raise_for_status() response.raise_for_status()
# Parse JSON response # Parse JSON response
data: Dict[Any, Any] = response.json() data: Dict[Any, Any] = response_json(response)
# Cache response if cache key provided # Cache response if cache key provided
if cache_key and self.cache_manager: if cache_key and self.cache_manager:
@@ -304,7 +305,7 @@ class APIHelper:
) )
response.raise_for_status() response.raise_for_status()
return cast(Optional[Dict[Any, Any]], response.json()) return cast(Optional[Dict[Any, Any]], response_json(response))
except requests.exceptions.RequestException as e: except requests.exceptions.RequestException as e:
self.logger.error(f"POST request failed for {url}: {e}") self.logger.error(f"POST request failed for {url}: {e}")
+189 -37
View File
@@ -26,10 +26,30 @@ A month can hold more than 500 events (college baseball's March does), and
ESPN answers that with exactly ``limit`` events and no hint that more exist. A ESPN answers that with exactly ``limit`` events and no hint that more exist. A
month chunk that comes back full is therefore re-asked day by day. month chunk that comes back full is therefore re-asked day by day.
A window's *partial* edge months are asked for whole, too, once the window
covers ``ESPN_MONTH_COVER_MIN_DAYS`` or more of their days, and the answer is
trimmed back to the window's days. A scoreboard's default fortnight either side
of today (29 days, two partial months) was 29 day requests per league; it is
now 2. Trimming needs ESPN's "game day", which is the event's start in US
Eastern time -- checked against the live API on 2026-10-03: 417 of 417 soccer
events across five leagues and three months (one of them spanning the end of
daylight saving) came back from exactly the day query their Eastern date
names. A short window (a live poll's one or two days) stays day by day, so it
never downloads a whole month to read a day of it.
Chunk requests share one process-wide budget of ``ESPN_CHUNK_WORKERS`` in
flight, however many windows are being fetched at once. Each window used to get
its own six, so a scoreboard starting eight leagues -- each with a recent and
an upcoming manager -- had ~40 requests in flight, every one beyond a session's
pool a new connection and a new DNS lookup. On a Pi whose resolver could not
keep up, that was ~90 ``NameResolutionError`` lines within a minute of every
start.
Once a range has been rejected, later ranges skip straight to chunks for Once a range has been rejected, later ranges skip straight to chunks for
``RANGE_RETRY_SECONDS`` instead of spending a doomed request first -- live ``RANGE_RETRY_SECONDS`` instead of spending a doomed request first -- live
scoreboards ask every 30 seconds. After that the range is tried again, so the scoreboards ask every 30 seconds. After that the range is tried again, so the
workaround retires itself if ESPN reverts. workaround retires itself if ESPN reverts. A process starts inside that
period, as if a range had just been rejected.
ONE CACHE KEY PER SCOREBOARD ONE CACHE KEY PER SCOREBOARD
---------------------------- ----------------------------
@@ -55,7 +75,7 @@ import re
import threading import threading
import time import time
from concurrent.futures import ThreadPoolExecutor from concurrent.futures import ThreadPoolExecutor
from datetime import date, datetime, timedelta from datetime import date, datetime, timedelta, tzinfo
from functools import partial from functools import partial
from typing import Any, Callable, Dict, Iterable, List, Optional, Tuple, cast from typing import Any, Callable, Dict, Iterable, List, Optional, Tuple, cast
@@ -100,16 +120,57 @@ RANGE_RETRY_SECONDS = 6 * 60 * 60
# pool_maxsize of 10 so the shared Session never has to discard connections. # pool_maxsize of 10 so the shared Session never has to discard connections.
ESPN_CHUNK_WORKERS = 6 ESPN_CHUNK_WORKERS = 6
#: An edge month the window covers at least this many days of is asked for
#: whole and trimmed, instead of one request per day (see module docstring).
#: Below it the days are cheaper than the month: a whole month is two to
#: three times the bytes of the half of it a fortnight window holds.
ESPN_MONTH_COVER_MIN_DAYS = 7
# Every chunk request in the process holds one of these while it is in flight
# -- the cap is per process, not per window (see module docstring).
_chunk_slots = threading.BoundedSemaphore(ESPN_CHUNK_WORKERS)
def _eastern_zone() -> Optional[tzinfo]:
"""US Eastern, the zone ESPN's ``dates=YYYYMMDD`` means, or None when
this Python has no time zone data (no edge month is trimmed then)."""
zone: Optional[tzinfo] = None
try:
from zoneinfo import ZoneInfo
zone = ZoneInfo("America/New_York")
except Exception: # noqa: BLE001 - no zoneinfo module or no tz database
zone = None
if zone is not None:
return zone
try:
import pytz
return cast(tzinfo, pytz.timezone("America/New_York"))
except Exception: # noqa: BLE001
return None
_EASTERN = _eastern_zone()
# What _fetch_one_chunk returns for a month that came back at the cap.
_CAPPED: Any = object()
_range_lock = threading.Lock() _range_lock = threading.Lock()
_ranges_rejected_until = 0.0 # A process starts out assuming ranges are still rejected, as they have been
# since 2026-09-15, and tries one again RANGE_RETRY_SECONDS in. Starting
# from "unknown" cost one doomed range request per window at every start --
# eleven 400s at once from a soccer board, each fetching before any had
# answered -- to learn what every start learns.
_ranges_rejected_until = time.monotonic() + RANGE_RETRY_SECONDS
__all__ = [ __all__ = [
"ESPN_MAX_LIMIT", "ESPN_MAX_LIMIT",
"ESPN_CHUNK_WORKERS", "ESPN_CHUNK_WORKERS",
"ESPN_MONTH_COVER_MIN_DAYS",
"RANGE_RETRY_SECONDS", "RANGE_RETRY_SECONDS",
"clamp_espn_limit", "clamp_espn_limit",
"parse_espn_date_range", "parse_espn_date_range",
"espn_date_chunks", "espn_date_chunks",
"espn_request_chunks",
"merge_scoreboard_payloads", "merge_scoreboard_payloads",
"fetch_espn_date_chunks", "fetch_espn_date_chunks",
"fetch_espn_scoreboard", "fetch_espn_scoreboard",
@@ -220,6 +281,79 @@ def espn_date_chunks(start: date, end: date) -> List[str]:
return chunks return chunks
def espn_request_chunks(
start: date,
end: date,
month_cover_min_days: Optional[int] = None,
) -> List[Tuple[str, Optional[Tuple[date, date]]]]:
"""The requests that fetch ``[start, end]``, as ``(dates, trim)`` pairs.
:func:`espn_date_chunks`, except that a partial edge month with
``month_cover_min_days`` (default ``ESPN_MONTH_COVER_MIN_DAYS``) or more
of its days in the window becomes one ``YYYYMM`` request whose ``trim``
is the first and last of those days: its events that start outside them
(US Eastern) are dropped. ``trim`` is None for every other request.
Without time zone data nothing can be trimmed, so the edge days stay day
requests.
"""
if month_cover_min_days is None:
month_cover_min_days = ESPN_MONTH_COVER_MIN_DAYS
planned: List[Tuple[str, Optional[Tuple[date, date]]]] = []
run: List[str] = []
def flush() -> None:
if (_EASTERN is not None and month_cover_min_days > 0
and len(run) >= month_cover_min_days):
planned.append((run[0][:6], (_parse_day(run[0]), _parse_day(run[-1]))))
else:
planned.extend((day, None) for day in run)
run.clear()
for chunk in espn_date_chunks(start, end):
if run and (len(chunk) != 8 or chunk[:6] != run[0][:6]):
flush()
if len(chunk) == 8:
run.append(chunk)
else:
planned.append((chunk, None))
flush()
return planned
def _parse_day(text: str) -> date:
return date(int(text[:4]), int(text[4:6]), int(text[6:8]))
def _eastern_day(stamp: Any) -> Optional[date]:
"""The US Eastern date of an ESPN event ``date`` ("2026-10-10T11:30Z"),
or None when it cannot be read."""
if not isinstance(stamp, str) or _EASTERN is None:
return None
try:
moment = datetime.fromisoformat(stamp.strip().replace("Z", "+00:00"))
except ValueError:
return None
if moment.tzinfo is None:
return None
return moment.astimezone(_EASTERN).date()
def _trim_to_days(payload: Any, first: date, last: date) -> Any:
"""Drop the events of a month payload that start outside ``[first, last]``
(US Eastern). An event whose date cannot be read is kept: its day query
might well have returned it, and a game is never dropped on a guess.
"""
if not isinstance(payload, dict) or not isinstance(payload.get("events"), list):
return payload
kept = []
for event in payload["events"]:
day = _eastern_day(event.get("date")) if isinstance(event, dict) else None
if day is None or first <= day <= last:
kept.append(event)
payload["events"] = kept
return payload
def merge_scoreboard_payloads(payloads: List[Any]) -> Dict[str, Any]: def merge_scoreboard_payloads(payloads: List[Any]) -> Dict[str, Any]:
"""Fold chunk responses into one scoreboard payload. """Fold chunk responses into one scoreboard payload.
@@ -250,37 +384,54 @@ def merge_scoreboard_payloads(payloads: List[Any]) -> Dict[str, Any]:
def _fetch_one_chunk( def _fetch_one_chunk(
session, url: str, params: Dict[str, Any], headers, timeout, logger, chunk: str, session, url: str, params: Dict[str, Any], headers, timeout, logger, chunk: str,
cache_max_age: Optional[float] = None, cache_max_age: Optional[float] = None,
) -> Optional[Dict[str, Any]]: trims: Optional[Dict[str, Tuple[date, date]]] = None,
) -> Any:
"""GET a single ``dates=`` chunk, or None when it failed. """GET a single ``dates=`` chunk, or None when it failed.
One bad chunk must not sink the rest of the season, so every error is One bad chunk must not sink the rest of the season, so every error is
logged and swallowed here rather than raised to the gather below. logged and swallowed here rather than raised to the gather below.
A month that comes back at the cap is truncated: it returns ``_CAPPED``,
its payload dropped here before it is ever held beside the others. A
month in ``trims`` loses its events outside the days given there.
The request holds one of the process-wide ``_chunk_slots`` while it runs.
""" """
try: try:
response = fetch_get( with _chunk_slots:
session, response = fetch_get(
url, session,
params=dict(params, dates=chunk, limit=ESPN_MAX_LIMIT), url,
headers=headers, params=dict(params, dates=chunk, limit=ESPN_MAX_LIMIT),
timeout=timeout, headers=headers,
**_memo_kwargs(cache_max_age), timeout=timeout,
) **_memo_kwargs(cache_max_age),
response.raise_for_status() )
return cast(Optional[Dict[str, Any]], response_json(response)) response.raise_for_status()
payload = response_json(response)
except Exception as exc: # noqa: BLE001 - see docstring except Exception as exc: # noqa: BLE001 - see docstring
if logger: if logger:
logger.warning("ESPN chunk %s failed, skipping it: %s", chunk, exc) logger.warning("ESPN chunk %s failed, skipping it: %s", chunk, exc)
return None return None
if len(chunk) == 6 and isinstance(payload, dict):
if len(payload.get("events") or []) >= ESPN_MAX_LIMIT:
return _CAPPED
trim = (trims or {}).get(chunk)
if trim is not None:
payload = _trim_to_days(payload, *trim)
return payload
def _fetch_chunks( def _fetch_chunks(
session, url: str, params: Dict[str, Any], headers, timeout, logger, session, url: str, params: Dict[str, Any], headers, timeout, logger,
chunks: List[str], cache_max_age: Optional[float] = None, chunks: List[str], cache_max_age: Optional[float] = None,
) -> List[Optional[Dict[str, Any]]]: trims: Optional[Dict[str, Tuple[date, date]]] = None,
) -> List[Any]:
"""Fetch every chunk, returning payloads positionally aligned with ``chunks``. """Fetch every chunk, returning payloads positionally aligned with ``chunks``.
Requests go out ``ESPN_CHUNK_WORKERS`` at a time because a cold season is Requests go out ``ESPN_CHUNK_WORKERS`` at a time because a cold season is
over a hundred of them. The order they come back in is not significant -- over a hundred of them -- and no more than that across every window the
process is fetching, which ``_fetch_one_chunk``'s slot enforces. The order they come back in is not significant --
callers keep ``chunks`` order from the returned list -- but it does mean callers keep ``chunks`` order from the returned list -- but it does mean
the session is shared across threads, which is why this only ever issues the session is shared across threads, which is why this only ever issues
GETs and never touches session state. GETs and never touches session state.
@@ -293,7 +444,7 @@ def _fetch_chunks(
return [] return []
fetch = partial( fetch = partial(
_fetch_one_chunk, session, url, params, headers, timeout, logger, _fetch_one_chunk, session, url, params, headers, timeout, logger,
cache_max_age=cache_max_age, cache_max_age=cache_max_age, trims=trims,
) )
if len(chunks) == 1: if len(chunks) == 1:
return [fetch(chunks[0])] return [fetch(chunks[0])]
@@ -340,7 +491,9 @@ def fetch_espn_date_chunks(
if span is None: if span is None:
return None return None
chunks = espn_date_chunks(*span) planned = espn_request_chunks(*span)
chunks = [chunk for chunk, _ in planned]
trims = {chunk: trim for chunk, trim in planned if trim is not None}
if logger: if logger:
logger.debug( logger.debug(
"Fetching ESPN date range %s as %d month/day chunks", "Fetching ESPN date range %s as %d month/day chunks",
@@ -349,32 +502,31 @@ def fetch_espn_date_chunks(
results = _fetch_chunks( results = _fetch_chunks(
session, url, params, headers, timeout, logger, chunks, cache_max_age, session, url, params, headers, timeout, logger, chunks, cache_max_age,
trims,
) )
attempted = len(chunks) attempted = len(chunks)
# A month that came back at the cap is truncated; its days replace it in # A month that came back at the cap is truncated; its days (only the
# place, so merged events stay in chunk order however the requests raced. # window's, for a trimmed edge month) replace it in place, so merged
# events stay in chunk order however the requests raced. Its payload was
# already dropped in the worker: a capped college-baseball month is ~2MB
# of parsed JSON, and holding four of them through ~120 day requests added
# ~25MB to the peak -- more than the concurrency itself. Low-memory boards
# (docs/LOW_MEMORY_BOARDS.md) have under 200MB of headroom.
slots: List[Any] = results slots: List[Any] = results
capped: Dict[int, List[str]] = {} capped: Dict[int, List[str]] = {}
for index, chunk in enumerate(chunks): for index, chunk in enumerate(chunks):
payload = slots[index] if slots[index] is not _CAPPED:
if payload is None or len(chunk) != 6:
continue continue
events = payload.get("events") if isinstance(payload, dict) else None if logger:
if len(events or []) >= ESPN_MAX_LIMIT: logger.info(
if logger: "ESPN month %s hit the %d-event cap; re-asking it day by day",
logger.info( chunk, ESPN_MAX_LIMIT,
"ESPN month %s hit the %d-event cap; re-asking it day by day", )
chunk, ESPN_MAX_LIMIT, trim = trims.get(chunk)
) capped[index] = (_days_of_month(chunk) if trim is None
capped[index] = _days_of_month(chunk) else espn_date_chunks(*trim))
# Drop the truncated month now rather than after its days arrive: slots[index] = None
# a capped college-baseball month is ~2MB of parsed JSON, and
# holding four of them through ~120 day requests added ~25MB to
# the peak -- more than the concurrency itself. Low-memory boards
# (docs/LOW_MEMORY_BOARDS.md) have under 200MB of headroom.
slots[index] = None
payload = events = None
if capped: if capped:
days = [day for index in sorted(capped) for day in capped[index]] days = [day for index in sorted(capped) for day in capped[index]]
+79 -3
View File
@@ -121,6 +121,7 @@ three times per threshold, so keep it to diagnostic runs, not soaks.
from __future__ import annotations from __future__ import annotations
import atexit
import copy import copy
import gc import gc
import json import json
@@ -134,6 +135,8 @@ import time
import traceback import traceback
from typing import Any, Callable, Dict, List, Optional, Tuple, TypedDict from typing import Any, Callable, Dict, List, Optional, Tuple, TypedDict
from src.common import scroll_config
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
#: Bumped when a field changes meaning, so a reader can refuse stale files. #: Bumped when a field changes meaning, so a reader can refuse stale files.
@@ -178,6 +181,12 @@ MAX_REFRESH_DROP = 0.2
#: trusted -- about a second of scrolling. #: trusted -- about a second of scrolling.
MIN_FRAMES_FOR_REFRESH = 90 MIN_FRAMES_FOR_REFRESH = 90
#: Trusted windows, counting the one that adopted the period, before a panel
#: slower than its cap is reported. The estimate can still fall (the refresh
#: rate rise) by up to MAX_REFRESH_DROP per window early on; the warning
#: should not name a rate one more window would have corrected.
REFRESH_CHECK_WINDOWS = 3
FLUSH_INTERVAL = 10.0 FLUSH_INTERVAL = 10.0
#: A scroll's last frame older than this is a stall worth a stack dump. #: A scroll's last frame older than this is a stall worth a stack dump.
@@ -213,10 +222,19 @@ class GcMonitor:
lock: the render thread and the stats writer only read them. lock: the render thread and the stats writer only read them.
Install it once per process with :func:`install_gc_monitor`. Install it once per process with :func:`install_gc_monitor`.
Collections still run while the interpreter shuts down, after module
globals such as ``time`` may already be torn down to ``None``. The clock
and ``sys.is_finalizing`` are bound here so the callback never looks a
global up, it does nothing once finalization has begun, and
:func:`install_gc_monitor` unregisters it at exit anyway.
""" """
def __init__(self, threshold: float = GC_PAUSE_SECONDS): def __init__(self, threshold: float = GC_PAUSE_SECONDS,
clock: Callable[[], float] = time.perf_counter):
self.threshold = threshold self.threshold = threshold
self._clock = clock
self._is_finalizing = sys.is_finalizing
self._started: Optional[float] = None self._started: Optional[float] = None
#: Per generation (0, 1, 2), since the monitor was installed. #: Per generation (0, 1, 2), since the monitor was installed.
self.collections = [0, 0, 0] self.collections = [0, 0, 0]
@@ -232,7 +250,9 @@ class GcMonitor:
self.last_long: Optional[Tuple[float, float]] = None self.last_long: Optional[Tuple[float, float]] = None
def __call__(self, phase: str, info: Dict[str, Any]) -> None: def __call__(self, phase: str, info: Dict[str, Any]) -> None:
now = time.perf_counter() if self._is_finalizing():
return
now = self._clock()
if phase == "start": if phase == "start":
self._started = now self._started = now
return return
@@ -267,15 +287,38 @@ _gc_monitor_lock = threading.Lock()
def install_gc_monitor() -> GcMonitor: def install_gc_monitor() -> GcMonitor:
"""The process's GcMonitor, installed in ``gc.callbacks`` on first call.""" """The process's GcMonitor, installed in ``gc.callbacks`` on first call.
It is unregistered at exit (:func:`uninstall_gc_monitor`), before the
interpreter tears module globals down.
"""
global _gc_monitor global _gc_monitor
with _gc_monitor_lock: with _gc_monitor_lock:
if _gc_monitor is None: if _gc_monitor is None:
_gc_monitor = GcMonitor() _gc_monitor = GcMonitor()
gc.callbacks.append(_gc_monitor) gc.callbacks.append(_gc_monitor)
atexit.register(uninstall_gc_monitor)
return _gc_monitor return _gc_monitor
def uninstall_gc_monitor() -> None:
"""Take the process's GcMonitor out of ``gc.callbacks``; safe to repeat.
A recorder that still holds the monitor keeps its counters; they just
stop moving. The next :func:`install_gc_monitor` installs a fresh one.
"""
global _gc_monitor
with _gc_monitor_lock:
monitor, _gc_monitor = _gc_monitor, None
if monitor is None:
return
atexit.unregister(uninstall_gc_monitor)
try:
gc.callbacks.remove(monitor)
except ValueError:
pass
#: One presented frame's interval: (interval, blit, wait, hold, ops), where #: One presented frame's interval: (interval, blit, wait, hold, ops), where
#: ops is the work noted before it (kind -> bytes) or None. #: ops is the work noted before it (kind -> bytes) or None.
_Frame = Tuple[float, float, float, int, Optional[Dict[str, int]]] _Frame = Tuple[float, float, float, int, Optional[Dict[str, int]]]
@@ -406,6 +449,12 @@ class FrameTimingRecorder:
1.0 / refresh_hz if refresh_hz and refresh_hz > 0 else None) 1.0 / refresh_hz if refresh_hz and refresh_hz > 0 else None)
# The first estimate, until a second window agrees with it. # The first estimate, until a second window agrees with it.
self._refresh_candidate: Optional[float] = None self._refresh_candidate: Optional[float] = None
# The rate scroll speeds are solved against; see plan_refresh().
self.planned_refresh_hz: Optional[float] = None
# Trusted windows seen since the period was adopted, until the
# shortfall check has run.
self._refresh_windows = 0
self._shortfall_checked = True
self.totals: Dict[str, Any] = { self.totals: Dict[str, Any] = {
"static_frames": 0, "static_frames": 0,
"scroll_frames": 0, "scroll_frames": 0,
@@ -604,7 +653,12 @@ class FrameTimingRecorder:
self._refresh_candidate = estimate self._refresh_candidate = estimate
elif current * (1.0 - MAX_REFRESH_DROP) <= estimate < current: elif current * (1.0 - MAX_REFRESH_DROP) <= estimate < current:
self.refresh_period = estimate self.refresh_period = estimate
if self.refresh_period is not None:
self._refresh_windows += 1
period = self.refresh_period period = self.refresh_period
if (period and not self._shortfall_checked
and self._refresh_windows >= REFRESH_CHECK_WINDOWS):
self._check_refresh_shortfall(1.0 / period)
histograms = self.histograms histograms = self.histograms
for frame in batch: for frame in batch:
@@ -653,6 +707,25 @@ class FrameTimingRecorder:
elif missed <= -1: elif missed <= -1:
totals["early_frames"] += 1 totals["early_frames"] += 1
def plan_refresh(self, hz: Optional[float]) -> None:
"""Say what rate scroll speeds are solved against, before frames arrive.
``DisplayManager.refresh_hz``: the configured cap. Once the measured
rate has held for :data:`REFRESH_CHECK_WINDOWS` windows, a panel that
falls short of it is logged once, with a cap it can hold (see
:func:`src.common.scroll_config.refresh_shortfall`). The display
manager calls this only for a real panel.
"""
self.planned_refresh_hz = hz
self._shortfall_checked = not hz
def _check_refresh_shortfall(self, measured_hz: float) -> None:
"""Log, once, a panel that cannot reach the rate speeds assume."""
self._shortfall_checked = True
shortfall = scroll_config.refresh_shortfall(measured_hz, self.planned_refresh_hz)
if shortfall:
logger.warning(scroll_config.describe_refresh_shortfall(shortfall))
def snapshot(self) -> Dict[str, Any]: def snapshot(self) -> Dict[str, Any]:
"""The JSON document: cumulative since this process started.""" """The JSON document: cumulative since this process started."""
if not self._binding_checked: if not self._binding_checked:
@@ -669,6 +742,9 @@ class FrameTimingRecorder:
"bucket_ms": BUCKET_MS, "bucket_ms": BUCKET_MS,
"freeze_seconds": FREEZE_SECONDS, "freeze_seconds": FREEZE_SECONDS,
"measured_refresh_hz": round(1.0 / period, 2) if period else None, "measured_refresh_hz": round(1.0 / period, 2) if period else None,
# Additive: what scroll speeds were solved against, so a reader can
# tell a stale file (written under another cap) from this one.
"planned_refresh_hz": self.planned_refresh_hz,
"binding_releases_gil": self._binding_gil, "binding_releases_gil": self._binding_gil,
"info": info, "info": info,
"totals": copy.deepcopy(self.totals), "totals": copy.deepcopy(self.totals),
+63
View File
@@ -538,3 +538,66 @@ def speed_advice(
"smooth": smooth, "smooth": smooth,
"alternatives": [as_dict(c) for c in alternatives], "alternatives": [as_dict(c) for c in alternatives],
} }
#: A measured refresh this far below the rate speeds are planned for means
#: the panel cannot reach its cap. Smaller gaps are the cap's own slack and
#: the estimate's: one rig measured 99.95 Hz under a 100 Hz cap.
REFRESH_SHORTFALL = 0.03
#: How far under the measured rate a suggested cap sits. The measurement is
#: the fast end of the panel's refreshes (frame_timing takes the 10th
#: percentile of intervals), and an uncapped panel drifts: one read
#: 107.6-113.1 Hz over 15 seconds. A cap inside that band would not hold.
CAP_HEADROOM = 0.05
def holdable_cap(measured_hz: Any) -> Optional[int]:
"""A refresh cap the panel can hold: a multiple of 10, 5% under what it measured.
A multiple of 10 because its whole-pixel speeds are round numbers (a
100 Hz cap gives 50 and 100 px/s). None without a usable measurement, or
when the panel is too slow for any cap of 10 Hz or more.
"""
hz = _coerce(measured_hz)
if hz is None:
return None
cap = int(hz * (1.0 - CAP_HEADROOM) // 10) * 10
return cap if cap >= 10 else None
def refresh_shortfall(measured_hz: Any, planned_hz: Any) -> Optional[Dict[str, Any]]:
"""When the panel refreshes measurably slower than speeds are planned for.
``planned_hz`` is what :func:`configure` solves against -- the
``limit_refresh_rate_hz`` cap, or :data:`DEFAULT_REFRESH_HZ` when it is 0.
A panel that cannot reach it still moves whole pixels per frame, but every
speed runs slow by the shortfall and the ladder of smooth speeds is the
cap's, not the panel's. None when there is no measurement, or the panel
reaches the cap (or beats it, as some do by a few Hz).
"""
measured, planned = _coerce(measured_hz), _coerce(planned_hz)
if measured is None or planned is None:
return None
if measured >= planned * (1.0 - REFRESH_SHORTFALL):
return None
return {
"measured_hz": round(measured, 1),
"planned_hz": round(planned, 1),
"suggested_cap_hz": holdable_cap(measured),
"slow_percent": round((1.0 - measured / planned) * 100),
}
def describe_refresh_shortfall(shortfall: Dict[str, Any]) -> str:
"""One log line for :func:`refresh_shortfall`'s answer."""
text = (
f"The panel refreshes at about {shortfall['measured_hz']:.0f} Hz, below "
f"the {shortfall['planned_hz']:.0f} Hz that scroll speeds are planned "
f"for (display.hardware.limit_refresh_rate_hz), so every scroll runs "
f"about {shortfall['slow_percent']}% slower than configured and the "
f"smooth speeds are worked out for a rate this panel never reaches.")
if shortfall.get("suggested_cap_hz"):
text += (f" Set Limit Refresh Rate to {shortfall['suggested_cap_hz']} Hz "
f"(web UI, Display tab), which this panel can hold, and restart.")
return text
+76 -31
View File
@@ -28,6 +28,30 @@ import numpy as np
# long over one frame, so a sample this large is an idle gap between scrolls. # long over one frame, so a sample this large is an idle gap between scrolls.
FPS_LOG_INTERVAL = 5.0 FPS_LOG_INTERVAL = 5.0
# The stats line goes to INFO only when a window is worth an operator's
# attention, as Vegas's FPS line does (src/vegas_mode/coordinator.py): every
# 5s from every scroller was most of the journal on a healthy rig. A window is
# degraded when its frame rate falls below this fraction of the rate it was
# locked to (1 / its own median frame time; same 0.9 as Vegas) ...
STATS_HEALTHY_FRACTION = 0.9
# ... or when more than this share of its frames stalled (past 1.5x the
# median). A 1% stall rate barely moves the mean, so the fps test alone would
# miss the judder this line exists to show.
STATS_DEGRADED_STALL_RATE = 0.01
# A healthy scroller still logs at INFO this often, so silence in the journal
# means stopped rather than fine. Every window is still logged at DEBUG.
STATS_HEARTBEAT_INTERVAL = 300.0
def frame_stats_degraded(stats: Dict[str, Any]) -> bool:
"""Whether one frame_stats() window is worth logging at INFO."""
n = stats["frames"]
if n == 0 or stats["median"] <= 0:
return False
locked_fps = 1.0 / stats["median"]
return (stats["fps"] < locked_fps * STATS_HEALTHY_FRACTION
or stats["stalls"] > n * STATS_DEGRADED_STALL_RATE)
def _rgb_pixels(item) -> np.ndarray: def _rgb_pixels(item) -> np.ndarray:
"""An appended item's pixels as an RGB array, as pasting it would draw them.""" """An appended item's pixels as an RGB array, as pasting it would draw them."""
@@ -189,6 +213,11 @@ class ScrollHelper:
# Every frame time since the last stats line, so the 5s summary can # Every frame time since the last stats line, so the 5s summary can
# report the tail rather than one arbitrary sample. Cleared on log. # report the tail rather than one arbitrary sample. Cleared on log.
self._window: list = [] self._window: list = []
# INFO-level stats bookkeeping (see STATS_HEARTBEAT_INTERVAL). Kept
# across reset_scroll(): a heartbeat per scroll start would bring the
# chatter back. 0.0 so the first window after start-up is at INFO.
self._stats_last_info_log = 0.0
self._stats_was_degraded = False
# Scrolling state management # Scrolling state management
self.is_scrolling = False self.is_scrolling = False
@@ -532,7 +561,7 @@ class ScrollHelper:
width = self.display_width width = self.display_width
strip_width = self.cached_array.shape[1] strip_width = self.cached_array.shape[1]
if start_x + width + 1 <= strip_width: if 0 <= start_x and start_x + width + 1 <= strip_width:
# Slice the backing array directly. Going via # Slice the backing array directly. Going via
# _get_visible_portion_integer would build two PIL images only for # _get_visible_portion_integer would build two PIL images only for
# them to be converted straight back to arrays, which measured 15x # them to be converted straight back to arrays, which measured 15x
@@ -540,9 +569,10 @@ class ScrollHelper:
near = self.cached_array[:, start_x:start_x + width] near = self.cached_array[:, start_x:start_x + width]
far = self.cached_array[:, start_x + 1:start_x + 1 + width] far = self.cached_array[:, start_x + 1:start_x + 1 + width]
else: else:
# Close enough to the end that one of the slices wraps; let the # One of the slices wraps (close to the end, or a strip narrower
# integer path handle that and pay the conversion. Continuous mode # than the panel); let the integer path handle that and pay the
# extends the strip before reaching here, so this is the rare case. # conversion. Continuous mode extends the strip before reaching
# here, so this is the rare case.
near = np.asarray( near = np.asarray(
self._get_visible_portion_integer(start_x, start_x + width)) self._get_visible_portion_integer(start_x, start_x + width))
far = np.asarray( far = np.asarray(
@@ -572,31 +602,33 @@ class ScrollHelper:
_size = (self.display_width, self.display_height) _size = (self.display_width, self.display_height)
img_w = self.cached_array.shape[1] img_w = self.cached_array.shape[1]
if end_x <= img_w: if 0 <= start_x and end_x <= img_w:
# Normal case: single contiguous slice (fastest path) # Normal case: single contiguous slice (fastest path). tobytes()
frame_array = np.ascontiguousarray(self.cached_array[:, start_x:end_x]) # on the column-slice view already returns C-order bytes, so
return Image.frombytes('RGB', _size, frame_array.tobytes()) # ascontiguousarray() first only added a second full-frame copy.
return Image.frombytes(
'RGB', _size,
self.cached_array[:, start_x:end_x].tobytes())
# Ensure frame buffer is allocated for all non-simple paths
if self._frame_buffer is None or self._frame_buffer.shape != (self.display_height, self.display_width, 3):
self._frame_buffer = np.zeros((self.display_height, self.display_width, 3), dtype=np.uint8)
if img_w == 0:
self._frame_buffer[:] = 0
else: else:
# Ensure frame buffer is allocated for all non-simple paths # The frame runs off the strip, so it carries on from the head:
if self._frame_buffer is None or self._frame_buffer.shape != (self.display_height, self.display_width, 3): # frame column j is strip column (start_x + j) modulo the strip's
self._frame_buffer = np.zeros((self.display_height, self.display_width, 3), dtype=np.uint8) # width -- the tail and then the head, and a strip narrower than
# the panel repeated across it. Copying the tail and then the rest
# of the frame from the head assumed the head was that wide, and
# raised at every position for a strip narrower than the panel
# (Vegas composes one, with no lead-in, when its content is
# narrower than the chain).
np.take(self.cached_array, np.arange(start_x, end_x), axis=1,
mode='wrap', out=self._frame_buffer)
width1 = img_w - start_x return Image.frombytes('RGB', _size, self._frame_buffer.tobytes())
if width1 > 0:
# Wrap-around: tail of image + head of image
self._frame_buffer[:, :width1] = self.cached_array[:, start_x:]
remaining_width = self.display_width - width1
self._frame_buffer[:, width1:] = self.cached_array[:, :remaining_width]
else:
# Edge case: start_x at or past image end — show from beginning,
# clamped to available width (scroll_position should wrap before
# reaching this state in normal operation).
available = min(self.display_width, img_w)
self._frame_buffer[:, :available] = self.cached_array[:, :available]
if available < self.display_width:
self._frame_buffer[:, available:] = 0
return Image.frombytes('RGB', _size, self._frame_buffer.tobytes())
def calculate_dynamic_duration(self) -> int: def calculate_dynamic_duration(self) -> int:
""" """
@@ -1207,10 +1239,23 @@ class ScrollHelper:
# as an idle gap. There is nothing to report, and reporting the # as an idle gap. There is nothing to report, and reporting the
# gap itself is the bug above. # gap itself is the bug above.
if self._window: if self._window:
self.logger.info( # INFO when degraded, on the window that recovers from it, and
"Scroll frame stats - %s", # as a slow heartbeat; DEBUG otherwise.
format_frame_stats(self._window), degraded = frame_stats_degraded(frame_stats(self._window))
) if (degraded or self._stats_was_degraded
or current_time - self._stats_last_info_log
>= STATS_HEARTBEAT_INTERVAL):
self.logger.info(
"Scroll frame stats - %s",
format_frame_stats(self._window),
)
self._stats_last_info_log = current_time
elif self.logger.isEnabledFor(logging.DEBUG):
self.logger.debug(
"Scroll frame stats - %s",
format_frame_stats(self._window),
)
self._stats_was_degraded = degraded
self.last_fps_log_time = current_time self.last_fps_log_time = current_time
self.frame_count = 0 self.frame_count = 0
self._window = [] self._window = []
+43 -4
View File
@@ -18,7 +18,7 @@ the extra guard only stops a None size raising TypeError.
""" """
import logging import logging
from datetime import datetime, timezone from datetime import datetime, timedelta, timezone
from typing import Any, Dict, Optional, Tuple from typing import Any, Dict, Optional, Tuple
from zoneinfo import ZoneInfo from zoneinfo import ZoneInfo
@@ -338,10 +338,46 @@ def format_game_date(config: Optional[Dict[str, Any]], logger, date_text: str,
if not raw: if not raw:
return "" return ""
fmt = str(scroll_card_option(config, "date_format", "abbrev") or "abbrev") fmt = str(scroll_card_option(config, "date_format", "abbrev") or "abbrev")
return _format_date_as(fmt, raw, lambda: weekday_for(config, logger, game)) return _format_date_as(fmt, raw, lambda: weekday_for(config, logger, game),
game=game)
def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR) -> str: def _printed_weekday(game: Optional[Dict], month: int, day: int) -> str:
"""The weekday of the date a card prints as month/day, or '' if unknown.
The extractor prints "M/D" in the plugin's resolved zone (its own setting,
else the global one, else the system zone). The card cannot see that zone:
it is handed the plugin's config, whose ``timezone`` ships as "", so
card_tzinfo answers UTC and an evening kickoff in the Americas got the
next day's weekday ("Sat Oct 2" for a Friday game). Every zone is within
a day of UTC, so the printed date is the start's UTC date or a neighbour
of it; the one with that month and day is the date on the card.
"""
if not isinstance(game, dict):
return ""
raw = game.get("start_time_utc") or game.get("start_time")
if not raw:
return ""
try:
start = raw if isinstance(raw, datetime) else datetime.fromisoformat(
str(raw).replace("Z", "+00:00"))
if start.utcoffset() is None:
return "" # naive: no instant to place the date against
utc_day = start.astimezone(timezone.utc).date()
except (ValueError, TypeError, OverflowError):
return ""
for offset in (0, -1, 1):
try:
candidate = utc_day + timedelta(days=offset)
except OverflowError:
continue
if (candidate.month, candidate.day) == (month, day):
return WEEKDAY_ABBR[candidate.weekday()]
return ""
def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR,
game: Optional[Dict] = None) -> str:
"""Render a stripped, non-empty "M/D" *raw* in style *fmt*. """Render a stripped, non-empty "M/D" *raw* in style *fmt*.
The body both date formatters share. They differ in which setting names the The body both date formatters share. They differ in which setting names the
@@ -349,6 +385,9 @@ def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR) -> str:
``SportsCoreSharedMixin._format_game_date``), so those arrive as arguments: ``SportsCoreSharedMixin._format_game_date``), so those arrive as arguments:
*weekday* is a zero-argument callable, only called for the "weekday" style. *weekday* is a zero-argument callable, only called for the "weekday" style.
*months* lets the mixin keep reading its (overridable) ``_MONTH_ABBR``. *months* lets the mixin keep reading its (overridable) ``_MONTH_ABBR``.
With *game*, the "weekday" style names the printed date's own weekday
(:func:`_printed_weekday`), and *weekday* is only the fallback for a
date its start time cannot place.
""" """
if fmt == "numeric": if fmt == "numeric":
return raw return raw
@@ -364,7 +403,7 @@ def _format_date_as(fmt: str, raw: str, weekday, months=MONTH_ABBR) -> str:
if fmt == "day_first": if fmt == "day_first":
return f"{day} {name}" return f"{day} {name}"
if fmt == "weekday": if fmt == "weekday":
day_name = weekday() day_name = _printed_weekday(game, month, day) or weekday()
return f"{day_name} {name} {day}" if day_name else f"{name} {day}" return f"{day_name} {name} {day}" if day_name else f"{name} {day}"
return f"{name} {day}" return f"{name} {day}"
+16 -2
View File
@@ -323,7 +323,7 @@ class SportsScrollDisplay:
:returns: True if a frame was drawn; False when there is no content or :returns: True if a frame was drawn; False when there is no content or
the frame could not be rendered. the frame could not be rendered.
""" """
if not self.scroll_helper.cached_image: if not self._has_strip():
return False return False
try: try:
@@ -416,7 +416,21 @@ class SportsScrollDisplay:
def has_cached_content(self) -> bool: def has_cached_content(self) -> bool:
"""Whether content is prepared and ready to scroll.""" """Whether content is prepared and ready to scroll."""
return bool(self.scroll_helper.cached_image) return self._has_strip()
def _has_strip(self) -> bool:
"""Whether the helper holds a strip, without building its PIL image.
Reading ``cached_image`` after the strip was extended or trimmed builds
the image from the array and keeps it, so the strip is held twice;
display_scroll_frame asks this every frame. ``has_strip()`` answers
from the helper's bookkeeping. A helper without it (a plugin's own, a
test double) is asked the old way.
"""
helper = self.scroll_helper
if callable(getattr(type(helper), "has_strip", None)):
return bool(helper.has_strip())
return bool(helper.cached_image)
# ------------------------------------------------------------------ # ------------------------------------------------------------------
# Live Vegas cards # Live Vegas cards
+4 -2
View File
@@ -360,14 +360,16 @@ class SportsCoreSharedMixin:
The formatting is sports_card's. What differs from the card's The formatting is sports_card's. What differs from the card's
``format_game_date`` is passed in: the setting (``switch_date_format``, ``format_game_date`` is passed in: the setting (``switch_date_format``,
see :meth:`_switch_date_format`) and the weekday, which comes from see :meth:`_switch_date_format`) and the weekday, which comes from
:meth:`_weekday_for` and so from this plugin's resolved timezone. :meth:`_weekday_for` and so from this plugin's resolved timezone
when the game's start cannot place the printed date. The game goes
in too, so both formatters name the printed date's own weekday.
""" """
raw = str(date_text or "").strip() raw = str(date_text or "").strip()
if not raw: if not raw:
return raw return raw
return _card._format_date_as(self._switch_date_format(), raw, return _card._format_date_as(self._switch_date_format(), raw,
lambda: self._weekday_for(game), lambda: self._weekday_for(game),
self._MONTH_ABBR) self._MONTH_ABBR, game=game)
def _weekday_for(self, game: Optional[Dict]) -> str: def _weekday_for(self, game: Optional[Dict]) -> str:
"""Weekday abbreviation from the game's start time, or ''.""" """Weekday abbreviation from the game's start time, or ''."""
+6 -1
View File
@@ -29,7 +29,6 @@ import time
import logging import logging
from enum import Enum from enum import Enum
from typing import Callable, Optional from typing import Callable, Optional
import numpy as np
from PIL import Image from PIL import Image
from src.config_manager_atomic import _replace from src.config_manager_atomic import _replace
@@ -434,6 +433,12 @@ class DisplaySyncManager:
return return
if self._leader_state != LeaderState.CONNECTED or not self._peer_ip: if self._leader_state != LeaderState.CONNECTED or not self._peer_ip:
return return
# numpy is imported here, not at module level: the web interface
# imports this module for its constants (STATUS_FILE, SYNC_PORT) and
# would otherwise load numpy for nothing. Only a connected leader
# gets this far, and after the first frame the import is a
# sys.modules lookup.
import numpy as np
try: try:
arr = np.asarray(image.convert("RGB"), dtype=np.uint8) arr = np.asarray(image.convert("RGB"), dtype=np.uint8)
header = _RAW_MAGIC + _RAW_HEADER.pack(image.width, image.height) header = _RAW_MAGIC + _RAW_HEADER.pack(image.width, image.height)
+43 -10
View File
@@ -46,6 +46,35 @@ from src.common.permission_utils import (
get_config_dir_mode get_config_dir_mode
) )
def _private_copy(config: Dict[str, Any]) -> Dict[str, Any]:
"""A deep copy of ``config`` that shares nothing with it.
load_config() hands one out per call, and the saves keep one, so the
cached config is never an object a caller holds. A web handler edits what
it loaded, validates, and may refuse the save; when the cache was that
same object, the refused edit stayed in it, and the next save of any
other setting wrote it to config.json -- a nested secret included, in
plain text, since it had never reached config_secrets.json to be
stripped.
The config is JSON data, so only its dicts and lists need copying; every
other value in it is immutable. On a Pi 4 with a real 60 KiB config this
takes 2.1 ms against copy.deepcopy's 6.8 ms, on a path ~30 handlers call
(a pickle round trip is no faster, 1.9 ms, and brings pickle into the
config path for nothing).
"""
return _copy_containers(config)
def _copy_containers(value: Any) -> Any:
if isinstance(value, dict):
return {key: _copy_containers(item) for key, item in value.items()}
if isinstance(value, list):
return [_copy_containers(item) for item in value]
return value
class ConfigManager: class ConfigManager:
""" """
Reads and writes the main application configuration files. Reads and writes the main application configuration files.
@@ -126,9 +155,10 @@ class ConfigManager:
validate_after_write=validate_after_write validate_after_write=validate_after_write
) )
# Update in-memory config if save was successful # Update in-memory config if save was successful. A copy: the caller
# still holds new_config_data (see _private_copy).
if result.status == SaveResultStatus.SUCCESS: if result.status == SaveResultStatus.SUCCESS:
self.config = new_config_data self.config = _private_copy(new_config_data)
# In-memory config now matches what was just written, so the # In-memory config now matches what was just written, so the
# load_config fast path may return it. It still carries the # load_config fast path may return it. It still carries the
# merged secrets that were stripped on disk; that matches a full # merged secrets that were stripped on disk; that matches a full
@@ -208,14 +238,16 @@ class ConfigManager:
Fast path: when config.json, config_secrets.json and the template Fast path: when config.json, config_secrets.json and the template
are all unchanged since the last successful load (mtime_ns + size), are all unchanged since the last successful load (mtime_ns + size),
the already-parsed self.config is returned without touching the a copy of the already-parsed self.config is returned without
files — same aliasing semantics as the full path, which also touching the files.
returns self.config.
Either way the caller gets its own copy (see _private_copy): editing
it changes nothing here until it is saved.
""" """
try: try:
current_sig = self._files_signature() current_sig = self._files_signature()
if self.config and self._loaded_sig == current_sig: if self.config and self._loaded_sig == current_sig:
return self.config return _private_copy(self.config)
# Check if config file exists, if not create from template # Check if config file exists, if not create from template
if not os.path.exists(self.config_path): if not os.path.exists(self.config_path):
@@ -249,8 +281,8 @@ class ConfigManager:
# Signature taken AFTER load + migration (migration may write the # Signature taken AFTER load + migration (migration may write the
# config back), so it reflects exactly what was read/written. # config back), so it reflects exactly what was read/written.
self._loaded_sig = self._files_signature() self._loaded_sig = self._files_signature()
return self.config return _private_copy(self.config)
except FileNotFoundError as e: except FileNotFoundError as e:
# Only config.json can get here: a missing or unreadable secrets # Only config.json can get here: a missing or unreadable secrets
# file is handled where it is read. # file is handled where it is read.
@@ -355,8 +387,9 @@ class ConfigManager:
try: try:
atomic_write_json(self.config_path, config_to_write) atomic_write_json(self.config_path, config_to_write)
# Update the in-memory config to the new state (which includes secrets for runtime) # Update the in-memory config to the new state (which includes
self.config = new_config_data # secrets for runtime), as a copy -- see _private_copy
self.config = _private_copy(new_config_data)
self._loaded_sig = self._files_signature() self._loaded_sig = self._files_signature()
self.logger.info(f"Configuration successfully saved to {os.path.abspath(self.config_path)}") self.logger.info(f"Configuration successfully saved to {os.path.abspath(self.config_path)}")
if secrets_content: if secrets_content:
+92 -42
View File
@@ -14,7 +14,7 @@ import json
import time import time
import threading import threading
from pathlib import Path from pathlib import Path
from typing import Dict, Any, Optional, List, Callable from typing import Dict, Any, Optional, List, Callable, Tuple
from collections import defaultdict from collections import defaultdict
import logging import logging
import hashlib import hashlib
@@ -52,7 +52,18 @@ class ConfigService:
# Thread safety # Thread safety
self._lock: threading.RLock = threading.RLock() self._lock: threading.RLock = threading.RLock()
# Held across a whole reload -- read, swap, notify -- so one reload's
# notifications finish before the next one's start. Subscribers run
# under this lock and never under _lock: the display's per-plugin
# subscriber can wait seconds for a busy plugin, and get_config(),
# subscribe() and unsubscribe() -- called from the render thread --
# must not wait behind it.
self._notify_lock: threading.RLock = threading.RLock()
# (key, callback, thread id) of the callback a notification is running,
# so unsubscribe() can wait for that one call; signalled on its return.
self._running_callback: Optional[Tuple[str, Callable[..., None], int]] = None
self._callback_done = threading.Condition(self._lock)
# Current configuration # Current configuration
self._current_config: Dict[str, Any] = {} self._current_config: Dict[str, Any] = {}
self._current_checksum: Optional[str] = None self._current_checksum: Optional[str] = None
@@ -87,32 +98,33 @@ class ConfigService:
True if config changed, False otherwise True if config changed, False otherwise
""" """
try: try:
new_config = self.config_manager.load_config() with self._notify_lock:
new_checksum = self._calculate_checksum(new_config) new_config = self.config_manager.load_config()
new_checksum = self._calculate_checksum(new_config)
with self._lock:
# Check if config actually changed with self._lock:
if new_checksum == self._current_checksum: # Check if config actually changed
self.logger.debug("Configuration unchanged, skipping reload") if new_checksum == self._current_checksum:
return False self.logger.debug("Configuration unchanged, skipping reload")
return False
# Store old config for change detection
old_config = self._current_config.copy() # Store old config for change detection
old_config = self._current_config.copy()
# Update current config
self._current_config = new_config # Update current config
self._current_checksum = new_checksum self._current_config = new_config
self._current_checksum = new_checksum
# Notify subscribers
# Notify subscribers, outside _lock (see _notify_lock)
self._notify_subscribers(old_config, new_config) self._notify_subscribers(old_config, new_config)
self.logger.info( self.logger.info(
"Configuration reloaded (checksum: %s)", "Configuration reloaded (checksum: %s)",
new_checksum[:8] new_checksum[:8]
) )
return True return True
except ConfigError as e: except ConfigError as e:
self.logger.error("Error loading configuration: %s", e, exc_info=True) self.logger.error("Error loading configuration: %s", e, exc_info=True)
return False return False
@@ -127,35 +139,64 @@ class ConfigService:
Args: Args:
old_config: Previous configuration old_config: Previous configuration
new_config: New configuration new_config: New configuration
Called without _lock held. The subscriber lists are copied under it,
and each callback is checked against them again just before it runs.
""" """
with self._lock:
subscribers = {key: list(callbacks) for key, callbacks in self._subscribers.items()}
# Notify global subscribers (key: '*') # Notify global subscribers (key: '*')
for callback in self._subscribers.get('*', []): for callback in subscribers.get('*', []):
try: self._call_subscriber('*', callback, old_config, new_config)
callback(old_config, new_config)
except Exception as e:
self.logger.error("Error in global config change callback: %s", e, exc_info=True)
# Notify plugin-specific subscribers # Notify plugin-specific subscribers
for plugin_id in self._subscribers.keys(): for plugin_id, callbacks in subscribers.items():
if plugin_id == '*': if plugin_id == '*':
continue continue
old_plugin_config = old_config.get(plugin_id, {}) old_plugin_config = old_config.get(plugin_id, {})
new_plugin_config = new_config.get(plugin_id, {}) new_plugin_config = new_config.get(plugin_id, {})
# Only notify if plugin config actually changed # Only notify if plugin config actually changed
if old_plugin_config != new_plugin_config: if old_plugin_config != new_plugin_config:
for callback in self._subscribers[plugin_id]: for callback in callbacks:
try: self._call_subscriber(plugin_id, callback,
callback(old_plugin_config, new_plugin_config) old_plugin_config, new_plugin_config)
except Exception as e:
self.logger.error( def _call_subscriber(
"Error in config change callback for %s: %s", self,
plugin_id, key: str,
e, callback: Callable[[Dict[str, Any], Dict[str, Any]], None],
exc_info=True old_config: Dict[str, Any],
) new_config: Dict[str, Any],
) -> None:
"""Run one callback, unless it was unsubscribed since the snapshot.
unsubscribe() promises that once it returns the callback is neither
running nor will run: the display unloads the plugin straight after.
"""
with self._lock:
if callback not in self._subscribers.get(key, ()):
return
self._running_callback = (key, callback, threading.get_ident())
try:
callback(old_config, new_config)
except Exception as e:
if key == '*':
self.logger.error("Error in global config change callback: %s", e, exc_info=True)
else:
self.logger.error(
"Error in config change callback for %s: %s",
key,
e,
exc_info=True
)
finally:
with self._lock:
self._running_callback = None
self._callback_done.notify_all()
def _check_file_changes(self) -> bool: def _check_file_changes(self) -> bool:
""" """
Check if configuration files have been modified. Check if configuration files have been modified.
@@ -276,6 +317,11 @@ class ConfigService:
""" """
Unsubscribe from configuration changes. Unsubscribe from configuration changes.
Once this returns the callback is not running and will not be called
again. A notification that is running this very callback is waited
for (unless the callback is the caller); one running any other
callback is not.
Args: Args:
callback: Callback function to remove callback: Callback function to remove
plugin_id: Optional plugin ID (must match subscription) plugin_id: Optional plugin ID (must match subscription)
@@ -285,6 +331,10 @@ class ConfigService:
if callback in self._subscribers[key]: if callback in self._subscribers[key]:
self._subscribers[key].remove(callback) self._subscribers[key].remove(callback)
self.logger.debug("Unsubscribed from config changes for %s", key) self.logger.debug("Unsubscribed from config changes for %s", key)
while (self._running_callback is not None
and self._running_callback[:2] == (key, callback)
and self._running_callback[2] != threading.get_ident()):
self._callback_done.wait()
def shutdown(self) -> None: def shutdown(self) -> None:
"""Shutdown the configuration service.""" """Shutdown the configuration service."""
+252 -26
View File
@@ -25,11 +25,12 @@ import os
import inspect import inspect
import signal import signal
import json import json
import math
import threading import threading
import types import types
from collections import deque from collections import deque
from contextlib import contextmanager from contextlib import contextmanager
from typing import Dict, Any, List, Optional, Callable, Set, Tuple from typing import Dict, Any, FrozenSet, List, Optional, Callable, Set, Tuple
from datetime import datetime from datetime import datetime
from concurrent.futures import ThreadPoolExecutor, as_completed # pylint: disable=no-name-in-module from concurrent.futures import ThreadPoolExecutor, as_completed # pylint: disable=no-name-in-module
import pytz import pytz
@@ -55,7 +56,7 @@ from src.ipc.contract import (
PluginReloadArgs, PluginReloadArgs,
PluginReloadResult, PluginReloadResult,
) )
from src.ipc.server import ControlServer, QueuedCommand, start_control_server from src.ipc.server import ControlServer, QueuedCommand, StateHub, start_control_server
from src.vegas_mode.render_pipeline import SYNC_SEND_INTERVAL from src.vegas_mode.render_pipeline import SYNC_SEND_INTERVAL
# Get logger with consistent configuration # Get logger with consistent configuration
@@ -65,6 +66,13 @@ logger = get_logger(__name__)
# treats display_current_state older than 120 s as unknown. # treats display_current_state older than 120 s as unknown.
CURRENT_STATE_REFRESH_SECONDS = 30 CURRENT_STATE_REFRESH_SECONDS = 30
# While the control socket serves the web interface's state readers
# (StateHub.readers_active), display_current_state is only their fallback:
# it is then rewritten at this interval and on a change of the flags, not on
# every mode change. Below the readers' 120 s max_age, so the fallback copy
# never reads as unknown.
CURRENT_STATE_RELAXED_REFRESH_SECONDS = 60
# How long startup will wait for plugins to fetch their first data before # How long startup will wait for plugins to fetch their first data before
# showing anything. Each plugin's update blocks for up to the executor's 30s # showing anything. Each plugin's update blocks for up to the executor's 30s
# timeout and they run one after another, so the uncapped total is the sum of # timeout and they run one after another, so the uncapped total is the sum of
@@ -82,6 +90,19 @@ _MIN_INITIAL_UPDATE_TIMEOUT_SECONDS = 2.0
DEFAULT_DYNAMIC_DURATION_CAP = 180.0 DEFAULT_DYNAMIC_DURATION_CAP = 180.0
def _finite_seconds(value: Any) -> Optional[float]:
"""``value`` as seconds when it is a finite number or a numeric string,
else None. A bool is not a number here, though it is an int: True would
read as a one-second screen."""
if isinstance(value, bool):
return None
try:
seconds = float(value)
except (TypeError, ValueError, OverflowError):
return None
return seconds if math.isfinite(seconds) else None
class _PluginReloadJob: class _PluginReloadJob:
"""A ``plugin.reload`` whose slow half runs off the render thread. """A ``plugin.reload`` whose slow half runs off the render thread.
@@ -323,6 +344,9 @@ class DisplayController:
# Monotonic stamp of the last _service_pending_changes pass; same # Monotonic stamp of the last _service_pending_changes pass; same
# "None means never" convention as _last_on_demand_poll. # "None means never" convention as _last_on_demand_poll.
self._last_pending_service: Optional[float] = None self._last_pending_service: Optional[float] = None
# Monotonic stamp of the last scheduled-update pass; see
# _tick_plugin_updates_if_due. Same "None means never" convention.
self._last_plugin_update_tick: Optional[float] = None
# The control socket (src/ipc), started by run(). None when it is not # The control socket (src/ipc), started by run(). None when it is not
# served (Windows, LEDMATRIX_CONTROL_SOCKET=off, a bind failure); # served (Windows, LEDMATRIX_CONTROL_SOCKET=off, a bind failure);
# the file mailbox works either way. # the file mailbox works either way.
@@ -344,6 +368,10 @@ class DisplayController:
self.on_demand_last_error: Optional[str] = None self.on_demand_last_error: Optional[str] = None
self.on_demand_last_event: Optional[str] = None self.on_demand_last_event: Optional[str] = None
self.on_demand_schedule_override = False self.on_demand_schedule_override = False
# The mode the request named, when it named one (not a mode resolved
# from a bare plugin id). Shown even when the plugin's live checks
# would leave it out of the session (_on_demand_modes_for_plugin).
self._on_demand_named_mode: Optional[str] = None
# Plugins that are disabled in config and loaded only because an # Plugins that are disabled in config and loaded only because an
# on-demand request named them. The main loop unloads each one once # on-demand request named them. The main loop unloads each one once
# on-demand has moved off it (_release_on_demand_plugins). # on-demand has moved off it (_release_on_demand_plugins).
@@ -537,6 +565,17 @@ class DisplayController:
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Plugin system initialization failed") logger.exception("Plugin system initialization failed")
self.plugin_manager = None self.plugin_manager = None
# A restored session has no plugin to resume on. It may have been
# read already (on_demand_active) or not yet, if initialization
# failed before the restore ran; either way, end it visibly.
try:
cached_session = self.cache_manager.get('display_on_demand_config',
max_age=3600)
except Exception: # pylint: disable=broad-except
cached_session = None
if self.on_demand_active or cached_session:
self.cache_manager.clear_cache('display_on_demand_config')
self._set_on_demand_error('restore-failed')
# Its state machine no longer describes what runs; let the last # Its state machine no longer describes what runs; let the last
# snapshot go stale (readers then say unknown) rather than keep # snapshot go stale (readers then say unknown) rather than keep
# refreshing it. # refreshing it.
@@ -1089,10 +1128,37 @@ class DisplayController:
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Error marking plugin %s updated for Vegas", plugin_id) logger.exception("Error marking plugin %s updated for Vegas", plugin_id)
#: Shortest gap between scheduled-update passes from the frame loops and
#: the dwell sleep. The pass (PluginManager.run_scheduled_updates) copies
#: the plugin dict and takes several locks per plugin to find, almost
#: always, that nothing is due: about 95 us with 20 plugins on a Pi, or
#: 1.2% of the render thread at 125 frames a second. No interval is
#: shorter than PluginManager.MIN_DYNAMIC_UPDATE_INTERVAL (5 s), and the
#: 1 Hz frame loop already ticks once a second, so a quarter second late
#: is not noticed.
PLUGIN_UPDATE_TICK_INTERVAL = 0.25
#: Class-level default for controllers built without __init__ (tests).
_last_plugin_update_tick: Optional[float] = None
def _tick_plugin_updates_if_due(self) -> None:
"""_tick_plugin_updates, at most once per PLUGIN_UPDATE_TICK_INTERVAL.
For the per-frame callers. The top of each loop pass calls
_tick_plugin_updates itself, unthrottled, because that is where a
plugin just loaded, reloaded or enabled for on-demand gets its first
update, and it must not wait out the floor.
"""
last = self._last_plugin_update_tick
if last is not None and time.monotonic() - last < self.PLUGIN_UPDATE_TICK_INTERVAL:
return
self._tick_plugin_updates()
def _tick_plugin_updates(self): def _tick_plugin_updates(self):
"""Run any plugin updates that are due.""" """Run any plugin updates that are due."""
if not self.plugin_manager: if not self.plugin_manager:
return return
self._last_plugin_update_tick = time.monotonic()
try: try:
self.plugin_manager.run_scheduled_updates() self.plugin_manager.run_scheduled_updates()
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
@@ -1164,6 +1230,13 @@ class DisplayController:
note = getattr(self.plugin_manager, 'note_display_duration', None) note = getattr(self.plugin_manager, 'note_display_duration', None)
if note is not None and plugin_id: if note is not None and plugin_id:
note(plugin_id, time.monotonic() - started) note(plugin_id, time.monotonic() - started)
# A screen that drew once and holds makes no more
# update_display() calls, so a frame the preview throttle
# skipped would otherwise never reach the snapshot.
write_owed = getattr(getattr(self, 'display_manager', None),
'write_owed_snapshot', None)
if write_owed is not None:
write_owed()
def _health_tracker(self): def _health_tracker(self):
"""The plugin circuit breaker, or None when it is not enabled.""" """The plugin circuit breaker, or None when it is not enabled."""
@@ -1260,7 +1333,7 @@ class DisplayController:
# A dwell can be a minute long (sixty seconds while scheduled # A dwell can be a minute long (sixty seconds while scheduled
# off); the watchdog must hear from this thread throughout. # off); the watchdog must hear from this thread throughout.
display_watchdog.watchdog.beat() display_watchdog.watchdog.beat()
self._tick_plugin_updates() self._tick_plugin_updates_if_due()
self._service_pending_changes() self._service_pending_changes()
self._check_live_takeover() self._check_live_takeover()
if (self.current_display_mode != mode if (self.current_display_mode != mode
@@ -1303,6 +1376,12 @@ class DisplayController:
"until one does", self.EMPTY_ROTATION_PAUSE) "until one does", self.EMPTY_ROTATION_PAUSE)
self._sleep_with_plugin_updates(self.EMPTY_ROTATION_PAUSE) self._sleep_with_plugin_updates(self.EMPTY_ROTATION_PAUSE)
#: Plugins already warned about a display duration that is not a number,
#: so a bad setting logs once, not at every one of its screens. A
#: frozenset, replaced rather than mutated; class-level default for
#: controllers built without __init__ (tests).
_duration_warned: FrozenSet[str] = frozenset()
def _get_display_duration(self, mode_key): def _get_display_duration(self, mode_key):
"""Seconds to show a mode: the Rotation & Durations page's value for it """Seconds to show a mode: the Rotation & Durations page's value for it
(display.display_durations), else the plugin's own duration. (display.display_durations), else the plugin's own duration.
@@ -1310,6 +1389,17 @@ class DisplayController:
The saved value has to win. Every plugin inherits The saved value has to win. Every plugin inherits
get_display_duration(), so checking the plugin first meant the page's get_display_duration(), so checking the plugin first meant the page's
values were never read. values were never read.
The plugin's answer is checked here, not trusted. Several plugins
return their display_duration setting straight from config.json, so
one saved as "20" or null (the raw config editor, a hand edit) came
back as a string or None; _resolve_durations compared it with 0, and
the TypeError went past every handler in the loop and stopped the
display service, which systemd restarted into the same screen. A
numeric string counts, as in BasePlugin.get_display_duration; any
other value that is not a finite number, or a raise, gets the 30 s a
mode without a plugin gets. A number at or below zero is passed on:
_resolve_durations has its own rule for that.
""" """
display_durations = self.config.get('display', {}).get('display_durations', {}) or {} display_durations = self.config.get('display', {}).get('display_durations', {}) or {}
override = display_durations.get(mode_key) override = display_durations.get(mode_key)
@@ -1317,8 +1407,22 @@ class DisplayController:
return float(override) return float(override)
plugin_instance = self.plugin_modes.get(mode_key) plugin_instance = self.plugin_modes.get(mode_key)
if plugin_instance is not None and hasattr(plugin_instance, 'get_display_duration'): if plugin_instance is None or not hasattr(plugin_instance, 'get_display_duration'):
return plugin_instance.get_display_duration() return 30
try:
value = plugin_instance.get_display_duration()
except Exception as err: # pylint: disable=broad-except
problem = f"get_display_duration() raised {type(err).__name__}: {err}"
else:
seconds = _finite_seconds(value)
if seconds is not None:
return seconds
problem = f"display duration {value!r} is not a number"
plugin_id = getattr(plugin_instance, 'plugin_id', None) or mode_key
if plugin_id not in self._duration_warned:
self._duration_warned = self._duration_warned | {plugin_id}
logger.warning("Plugin %s: %s; showing its modes for 30s (logged once)",
plugin_id, problem)
return 30 return 30
def _get_global_dynamic_cap(self) -> Optional[float]: def _get_global_dynamic_cap(self) -> Optional[float]:
@@ -1430,18 +1534,60 @@ class DisplayController:
return None return None
return max(0.0, expires_at - time.time()) return max(0.0, expires_at - time.time())
#: The control socket's state stream (src/ipc/server.StateHub), while the
#: socket is served. Class-level default for controllers built without
#: __init__ (tests) and for a display with no socket.
_state_hub: Optional[StateHub] = None
def _current_mode_state(self) -> Dict[str, Any]:
"""What display_current_state and the socket's ``display`` section hold."""
return {
'mode': self.current_display_mode,
'plugin_id': self.mode_to_plugin_id.get(self.current_display_mode),
'mode_index': self.current_mode_index,
'total_modes': len(self.available_modes),
'on_demand_active': self.on_demand_active,
'is_display_active': self.is_display_active,
'last_updated': time.time(),
}
def _push_live_state(self, display_state: Optional[Dict[str, Any]] = None) -> None:
"""Hand the current mode and the brightness to the control socket's
state stream. In memory, no disk: the hub only bumps its version (and
wakes subscribers) when something other than ``last_updated`` changed.
Called on every pass of the publish points below, so
``display.last_updated`` doubles as the render thread's proof of life
for the socket's readers, as the cache key's max_age does today.
"""
hub = self._state_hub
if hub is None:
return
try:
hub.publish('display', display_state or self._current_mode_state(),
volatile=('last_updated',))
hub.publish('brightness', {
'brightness': getattr(self, '_normal_brightness', None),
'panel_brightness': getattr(self, 'current_brightness', None),
'dimmed': bool(getattr(self, 'is_dimmed', False)),
})
except Exception as err: # pylint: disable=broad-except
logger.debug("Could not publish the display state to the control socket: %s",
err, exc_info=True)
def _state_readers_on_socket(self) -> bool:
"""Is the control socket serving the web interface's state readers?"""
hub = self._state_hub
try:
return bool(hub is not None and hub.readers_active())
except Exception: # pylint: disable=broad-except
return False
def _publish_current_mode_state(self) -> None: def _publish_current_mode_state(self) -> None:
"""Publish the currently active display mode/plugin to cache for the web UI.""" """Publish the currently active display mode/plugin to cache for the web UI."""
try: try:
state = { state = self._current_mode_state()
'mode': self.current_display_mode, self._push_live_state(state)
'plugin_id': self.mode_to_plugin_id.get(self.current_display_mode),
'mode_index': self.current_mode_index,
'total_modes': len(self.available_modes),
'on_demand_active': self.on_demand_active,
'is_display_active': self.is_display_active,
'last_updated': time.time(),
}
self.cache_manager.set('display_current_state', state) self.cache_manager.set('display_current_state', state)
self._last_published_mode = self.current_display_mode self._last_published_mode = self.current_display_mode
self._last_published_flags = self._current_state_flags() self._last_published_flags = self._current_state_flags()
@@ -1467,11 +1613,23 @@ class DisplayController:
priority, a single enabled plugin -- has to be republished or the UI priority, a single enabled plugin -- has to be republished or the UI
reports it as unknown. Otherwise this writes only on a change, not on reports it as unknown. Otherwise this writes only on a change, not on
every render tick. every render tick.
While the control socket serves the web interface's state readers,
the socket's in-memory copy is updated on every call and the cache
key is only their fallback: a mode change alone is then written at
the relaxed refresh (CURRENT_STATE_RELAXED_REFRESH_SECONDS), still
inside the readers' max_age. The flags are still written at once.
When the socket stops serving them, the next call writes a changed
mode again.
""" """
if (self.current_display_mode != self._last_published_mode relaxed = self._state_readers_on_socket()
refresh = CURRENT_STATE_RELAXED_REFRESH_SECONDS if relaxed else CURRENT_STATE_REFRESH_SECONDS
if ((not relaxed and self.current_display_mode != self._last_published_mode)
or self._current_state_flags() != getattr(self, '_last_published_flags', None) or self._current_state_flags() != getattr(self, '_last_published_flags', None)
or time.monotonic() - self._last_published_at >= CURRENT_STATE_REFRESH_SECONDS): or time.monotonic() - self._last_published_at >= refresh):
self._publish_current_mode_state() self._publish_current_mode_state() # pushes to the socket as well
else:
self._push_live_state()
def _on_demand_state(self) -> Dict[str, Any]: def _on_demand_state(self) -> Dict[str, Any]:
"""The on-demand state as published to the cache and the control socket.""" """The on-demand state as published to the cache and the control socket."""
@@ -1494,6 +1652,11 @@ class DisplayController:
"""Publish current on-demand state to cache for external consumers.""" """Publish current on-demand state to cache for external consumers."""
try: try:
state = self._on_demand_state() state = self._on_demand_state()
hub = self._state_hub
if hub is not None:
# In memory, first: a subscriber hears the outcome of an
# on-demand command even if the cache write below fails.
hub.publish('on_demand', state, volatile=('last_updated', 'remaining'))
self.cache_manager.set('display_on_demand_state', state) self.cache_manager.set('display_on_demand_state', state)
except (OSError, RuntimeError, ValueError, TypeError) as err: except (OSError, RuntimeError, ValueError, TypeError) as err:
logger.error("Failed to publish on-demand state: %s", err, exc_info=True) logger.error("Failed to publish on-demand state: %s", err, exc_info=True)
@@ -1524,6 +1687,7 @@ class DisplayController:
self.on_demand_expires_at = None self.on_demand_expires_at = None
self.on_demand_pinned = False self.on_demand_pinned = False
self.on_demand_schedule_override = False self.on_demand_schedule_override = False
self._on_demand_named_mode = None
# While the session ran, _evaluate_schedule may have forced # While the session ran, _evaluate_schedule may have forced
# is_display_active on over a scheduled-off answer. Drop the minute # is_display_active on over a scheduled-off answer. Drop the minute
# gate so the next _check_schedule recomputes it; otherwise the panel # gate so the next _check_schedule recomputes it; otherwise the panel
@@ -1690,6 +1854,7 @@ class DisplayController:
self.on_demand_pinned = on_demand_config.get('pinned', False) self.on_demand_pinned = on_demand_config.get('pinned', False)
self.on_demand_requested_at = on_demand_config.get('requested_at') self.on_demand_requested_at = on_demand_config.get('requested_at')
self.on_demand_expires_at = on_demand_config.get('expires_at') self.on_demand_expires_at = on_demand_config.get('expires_at')
self._on_demand_named_mode = on_demand_config.get('named_mode')
self.on_demand_status = 'active' self.on_demand_status = 'active'
self.on_demand_schedule_override = True self.on_demand_schedule_override = True
logger.info("On-demand mode detected during initialization: resuming on plugin '%s'; " logger.info("On-demand mode detected during initialization: resuming on plugin '%s'; "
@@ -1738,12 +1903,31 @@ class DisplayController:
""" """
if self._control_server is not None: if self._control_server is not None:
return return
hub = StateHub(loop_probe=display_watchdog.watchdog.liveness)
try: try:
self._control_server = start_control_server( self._control_server = start_control_server(
status_provider=self._control_status, status_provider=self._control_status,
cache_dir=getattr(self.cache_manager, 'cache_dir', None)) cache_dir=getattr(self.cache_manager, 'cache_dir', None),
state_hub=hub)
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Control socket not started; using the file mailbox only") logger.exception("Control socket not started; using the file mailbox only")
if self._control_server is not None:
self._start_state_stream(hub)
def _start_state_stream(self, hub: StateHub) -> None:
"""Start publishing to the socket's state stream (``state.get`` and
``state.subscribe``): everything a reader would see, now, then on
every publish. Never raises; without it readers use the cache keys."""
try:
self._state_hub = hub
self._push_live_state()
hub.publish('on_demand', self._on_demand_state(),
volatile=('last_updated', 'remaining'))
publisher = getattr(self, '_plugin_runtime_publisher', None)
if publisher is not None:
publisher.attach_hub(hub)
except Exception: # pylint: disable=broad-except
logger.exception("Control socket state stream not started; readers use the cache")
def _control_status(self) -> Dict[str, Any]: def _control_status(self) -> Dict[str, Any]:
"""The socket's on_demand.status answer. Runs on the socket's thread: reads only.""" """The socket's on_demand.status answer. Runs on the socket's thread: reads only."""
@@ -2157,13 +2341,25 @@ class DisplayController:
return modes[0] return modes[0]
return plugin_id return plugin_id
def _on_demand_modes_for_plugin(self, plugin_id: str) -> List[str]: def _on_demand_modes_for_plugin(self, plugin_id: str,
named_mode: Optional[str] = None) -> List[str]:
"""Every loaded display mode belonging to `plugin_id`, in rotation order. """Every loaded display mode belonging to `plugin_id`, in rotation order.
Live modes that actually have content lead, then the rest, then live Live modes that actually have content lead, then the rest, then live
modes with nothing to show -- so an on-demand request for a sports modes with nothing to show -- so an on-demand request for a sports
plugin opens on a game in progress rather than an empty live screen. plugin opens on a game in progress rather than an empty live screen.
Returns an empty list when the plugin has no loaded modes. Returns an empty list when the plugin has no loaded modes.
`named_mode` is a mode the request asked for by name. It is always
in the list, first when the checks below would have dropped it.
Those checks ask has_live_content(), which is the live-priority
question -- "should this plugin take the panel from the rotation?"
-- and the sports plugins answer it for favourite teams only. Asking
for ncaa_fb_live with fifteen games on and no favourite playing got
a 200 and nfl_recent on the panel. The plugin's display() is what
knows whether the mode has anything to draw; when it has not, the
session moves to the plugin's next mode like any empty on-demand
mode.
""" """
plugin_modes = self.plugin_display_modes.get(plugin_id, []) plugin_modes = self.plugin_display_modes.get(plugin_id, [])
if not plugin_modes: if not plugin_modes:
@@ -2204,6 +2400,15 @@ class DisplayController:
# Only live modes available but no content - use them anyway # Only live modes available but no content - use them anyway
ordered_modes = live_modes ordered_modes = live_modes
if named_mode and named_mode in available_plugin_modes:
# The named mode leads whether or not the live check kept it: a
# second live mode with content is already in the list, but behind
# the first, so the session would rotate away before reaching it.
if named_mode not in ordered_modes:
logger.info("On-demand: showing %s as requested; plugin '%s' reports no "
"live-priority content for it", named_mode, plugin_id)
ordered_modes = [named_mode] + [m for m in ordered_modes if m != named_mode]
return ordered_modes return ordered_modes
def _apply_on_demand_pin(self, ordered_modes: List[str], resolved_mode: Optional[str], def _apply_on_demand_pin(self, ordered_modes: List[str], resolved_mode: Optional[str],
@@ -2232,10 +2437,20 @@ class DisplayController:
plugin_id = self.on_demand_plugin_id plugin_id = self.on_demand_plugin_id
ordered_modes = self._on_demand_modes_for_plugin(plugin_id) ordered_modes = self._on_demand_modes_for_plugin(plugin_id, self._on_demand_named_mode)
if not ordered_modes: if not ordered_modes:
logger.warning("No valid display modes found for on-demand plugin '%s' after restoration", plugin_id) # The plugin did not load this time (seen on a rig: its config
self.on_demand_modes = [] # failed validation after the crash that caused the restart), so
# there is nothing to resume. Leaving the session active with no
# modes published it as active for a plugin that was not running
# until the first pass ended it as an ordinary 'idle', and kept
# the cached request for the next restart to trip over. End it
# as a failure the status endpoint reports, and drop the cache.
logger.error("On-demand session for plugin '%s' cannot resume after the "
"restart: the plugin has no loaded display modes (did it "
"fail to load?); ending it", plugin_id)
self.cache_manager.clear_cache('display_on_demand_config')
self._set_on_demand_error('restore-failed')
return return
# A restart must not silently un-pin: the pin is part of the request # A restart must not silently un-pin: the pin is part of the request
@@ -2436,7 +2651,10 @@ class DisplayController:
if resolved_mode in self.available_modes: if resolved_mode in self.available_modes:
self.current_mode_index = self.available_modes.index(resolved_mode) self.current_mode_index = self.available_modes.index(resolved_mode)
ordered_modes = self._on_demand_modes_for_plugin(resolved_plugin_id) # Named: the request gave this mode itself, rather than a plugin id
# (or a mode the plugin doesn't have) that resolved to a default.
named_mode = resolved_mode if mode == resolved_mode else None
ordered_modes = self._on_demand_modes_for_plugin(resolved_plugin_id, named_mode)
if not ordered_modes: if not ordered_modes:
logger.error("No valid display modes found for plugin '%s'", resolved_plugin_id) logger.error("No valid display modes found for plugin '%s'", resolved_plugin_id)
self._set_on_demand_error("no-modes") self._set_on_demand_error("no-modes")
@@ -2453,6 +2671,7 @@ class DisplayController:
self.on_demand_requested_at = now self.on_demand_requested_at = now
self.on_demand_expires_at = (now + duration) if duration else None self.on_demand_expires_at = (now + duration) if duration else None
self.on_demand_pinned = pinned self.on_demand_pinned = pinned
self._on_demand_named_mode = named_mode
self.on_demand_status = 'active' self.on_demand_status = 'active'
self.on_demand_last_error = None self.on_demand_last_error = None
self.on_demand_last_event = 'started' self.on_demand_last_event = 'started'
@@ -2481,6 +2700,7 @@ class DisplayController:
'mode': resolved_mode, 'mode': resolved_mode,
'duration': duration, 'duration': duration,
'pinned': pinned, 'pinned': pinned,
'named_mode': named_mode,
'requested_at': now, 'requested_at': now,
'expires_at': self.on_demand_expires_at 'expires_at': self.on_demand_expires_at
} }
@@ -2820,7 +3040,10 @@ class DisplayController:
self._follower_local_x = local_x self._follower_local_x = local_x
if rp and rp.scroll_helper.cached_image is not None: # has_strip(), not cached_image: every follower frame asks, and
# reading a strip the follower's own rebuild deferred would build
# and keep a second copy of it as a PIL image.
if rp and rp.scroll_helper.has_strip():
# Hold last frame until TCP image arrives after cycle reset # Hold last frame until TCP image arrives after cycle reset
if not self._follower_pending_new_image and local_x >= width: if not self._follower_pending_new_image and local_x >= width:
rp.scroll_helper.scroll_position = ( rp.scroll_helper.scroll_position = (
@@ -3510,6 +3733,8 @@ class DisplayController:
self._release_on_demand_plugins() self._release_on_demand_plugins()
if not self.available_modes: if not self.available_modes:
continue # it was all there was; idle as above continue # it was all there was; idle as above
# Unthrottled, unlike the frame loops: a plugin loaded,
# reloaded or enabled for on-demand above is due at once.
self._tick_plugin_updates() self._tick_plugin_updates()
# Clean up expired WiFi status messages # Clean up expired WiFi status messages
@@ -3772,7 +3997,7 @@ class DisplayController:
# Multi-display sync: send follower frame after each render # Multi-display sync: send follower frame after each render
self._send_follower_frame(manager_to_display) self._send_follower_frame(manager_to_display)
self._tick_plugin_updates() self._tick_plugin_updates_if_due()
# Throttled: one clock compare between passes. # Throttled: one clock compare between passes.
self._service_pending_changes() self._service_pending_changes()
self._check_live_takeover() self._check_live_takeover()
@@ -3831,7 +4056,7 @@ class DisplayController:
"breaking early", active_mode, "breaking early", active_mode,
self.current_display_mode) self.current_display_mode)
break break
self._tick_plugin_updates() self._tick_plugin_updates_if_due()
elapsed = time.time() - start_time elapsed = time.time() - start_time
if elapsed >= target_duration: if elapsed >= target_duration:
@@ -4532,6 +4757,7 @@ class DisplayController:
except Exception as e: except Exception as e:
logger.warning("Error closing the control socket: %s", e) logger.warning("Error closing the control socket: %s", e)
self._control_server = None self._control_server = None
self._state_hub = None
# Stop the async update worker first so no in-flight update() call # Stop the async update worker first so no in-flight update() call
# is still touching display/cache-backed resources while they're # is still touching display/cache-backed resources while they're
# torn down below. # torn down below.
+64 -4
View File
@@ -57,7 +57,8 @@ import freetype
from src.common import snapshot_policy from src.common import snapshot_policy
from src import display_watchdog from src import display_watchdog
from src.common.frame_timing import FrameTimingRecorder, install_gc_monitor from src.common.frame_timing import (
FrameTimingRecorder, install_gc_monitor, uninstall_gc_monitor)
if TYPE_CHECKING: if TYPE_CHECKING:
from src.common.render_gate import RenderGate from src.common.render_gate import RenderGate
@@ -316,6 +317,11 @@ class DisplayManager:
# is handed to the writer; this only once it has been saved, so an # is handed to the writer; this only once it has been saved, so an
# mtime touch never vouches for a frame still waiting to be written. # mtime touch never vouches for a frame still waiting to be written.
self._saved_snapshot_digest: Optional[int] = None self._saved_snapshot_digest: Optional[int] = None
# A changed frame reached _write_snapshot_if_due() inside the write
# interval and was skipped. Nothing writes it unless update_display()
# runs again, and a screen that draws once and holds never calls it
# again -- see write_owed_snapshot().
self._snapshot_owed = False
self._snapshot_dir_prepared = False self._snapshot_dir_prepared = False
# Background writer used mid-scroll; see _write_snapshot_if_due. # Background writer used mid-scroll; see _write_snapshot_if_due.
self._snapshot_cond = threading.Condition() self._snapshot_cond = threading.Condition()
@@ -387,6 +393,11 @@ class DisplayManager:
self._setup_matrix() self._setup_matrix()
logger.info("Matrix setup completed in %.3f seconds", time.time() - start_time) logger.info("Matrix setup completed in %.3f seconds", time.time() - start_time)
# Only a real panel's swaps wait on its refresh: the emulator and the
# fallback canvas pace themselves, so "slower than the cap" would be
# noise there.
if self.matrix is not None and os.environ.get('EMULATOR', 'false') != 'true':
self.frame_timing.plan_refresh(self.refresh_hz)
self._setup_scan_order_compensation() self._setup_scan_order_compensation()
font_time = time.time() font_time = time.time()
@@ -1407,6 +1418,8 @@ class DisplayManager:
# The stall watchdog would otherwise outlive this manager. # The stall watchdog would otherwise outlive this manager.
if getattr(self, 'frame_timing', None) is not None: if getattr(self, 'frame_timing', None) is not None:
self.frame_timing.close() self.frame_timing.close()
# Installed with the recorder; stop timing collections with it.
uninstall_gc_monitor()
# Reset the singleton state when cleaning up # Reset the singleton state when cleaning up
DisplayManager._instance = None DisplayManager._instance = None
@@ -1493,7 +1506,9 @@ class DisplayManager:
fractional-pixel motion. See src/common/scroll_config.py. fractional-pixel motion. See src/common/scroll_config.py.
Note this is the configured *cap*, not necessarily what the panel Note this is the configured *cap*, not necessarily what the panel
achieves -- scripts/scroll_speeds.py --measure reports the real rate. achieves -- scripts/scroll_speeds.py --measure reports the real rate,
and the frame-timing recorder logs a warning, with a cap the panel can
hold, once it has measured a panel that falls short of this.
""" """
hardware = (self.config.get('display') or {}).get('hardware') or {} hardware = (self.config.get('display') or {}).get('hardware') or {}
try: try:
@@ -1785,9 +1800,10 @@ class DisplayManager:
if frame_checksum is not None: if frame_checksum is not None:
digest = frame_checksum digest = frame_checksum
frame_changed = digest != self._last_snapshot_digest
action = snapshot_policy.decide( action = snapshot_policy.decide(
now, self._last_snapshot_ts, self._last_snapshot_touch_ts, now, self._last_snapshot_ts, self._last_snapshot_touch_ts,
viewer_fresh, digest != self._last_snapshot_digest) viewer_fresh, frame_changed)
else: else:
# Ask as if the frame had changed before paying to find out. # Ask as if the frame had changed before paying to find out.
# decide() is monotone in frame_changed -- a SKIP for a # decide() is monotone in frame_changed -- a SKIP for a
@@ -1799,22 +1815,36 @@ class DisplayManager:
now, self._last_snapshot_ts, self._last_snapshot_touch_ts, now, self._last_snapshot_ts, self._last_snapshot_touch_ts,
viewer_fresh, True) viewer_fresh, True)
if action is snapshot_policy.SnapshotAction.SKIP: if action is snapshot_policy.SnapshotAction.SKIP:
# Not hashed, so not known to be unchanged: owed until a
# later look finds it written or unchanged.
self._snapshot_owed = True
return return
digest = zlib.adler32(self.image.tobytes()) digest = zlib.adler32(self.image.tobytes())
if digest == self._last_snapshot_digest: frame_changed = digest != self._last_snapshot_digest
if not frame_changed:
# Unchanged after all: the decision an unchanged frame gets. # Unchanged after all: the decision an unchanged frame gets.
action = snapshot_policy.decide( action = snapshot_policy.decide(
now, self._last_snapshot_ts, now, self._last_snapshot_ts,
self._last_snapshot_touch_ts, viewer_fresh, False) self._last_snapshot_touch_ts, viewer_fresh, False)
if action is snapshot_policy.SnapshotAction.SKIP: if action is snapshot_policy.SnapshotAction.SKIP:
# A changed frame inside the write interval stays owed: the
# next update_display() would write it, but a static screen
# may not make one -- write_owed_snapshot() covers that.
self._snapshot_owed = frame_changed
return return
if (action is snapshot_policy.SnapshotAction.TOUCH if (action is snapshot_policy.SnapshotAction.TOUCH
and self._saved_snapshot_digest == digest): and self._saved_snapshot_digest == digest):
# mtime bump only: keeps the health check (snapshot age) # mtime bump only: keeps the health check (snapshot age)
# green without paying for a PNG encode of an unchanged frame # green without paying for a PNG encode of an unchanged frame
# (this frame is already on disk, so nothing is owed).
self._snapshot_owed = False
os.utime(self._snapshot_path, None) os.utime(self._snapshot_path, None)
self._last_snapshot_touch_ts = now self._last_snapshot_touch_ts = now
return return
# Owed until the write below succeeds: if it raises, the frame
# stays owed and write_owed_snapshot() retries it, rather than a
# held screen leaving the preview stale after one failed write.
self._snapshot_owed = True
# (A TOUCH for a frame that isn't on disk yet -- still queued, or # (A TOUCH for a frame that isn't on disk yet -- still queued, or
# its write failed -- is written instead: touching would make the # its write failed -- is written instead: touching would make the
# older file on disk look current.) # older file on disk look current.)
@@ -1839,9 +1869,39 @@ class DisplayManager:
self._last_snapshot_ts = now self._last_snapshot_ts = now
self._last_snapshot_touch_ts = now self._last_snapshot_touch_ts = now
self._last_snapshot_digest = digest self._last_snapshot_digest = digest
self._snapshot_owed = False
except Exception as e: except Exception as e:
self._log_snapshot_failure(e) self._log_snapshot_failure(e)
def write_owed_snapshot(self) -> None:
"""Write a frame the snapshot throttle skipped, once it is due.
The preview snapshot is only ever written from update_display(), and
at most once per write interval (snapshot_policy). A frame pushed
inside that interval is skipped, and is written by the next
update_display() that comes after it -- but a screen that draws its
card once and then holds it makes no further call. Its frame was on
the panel and never in the preview: soccer's recent/upcoming cards
skip redundant redraws, and the first one after an on-demand start
(pushed a few milliseconds after the controller's clear) left
/api/v3/display/current and the web preview black for the whole
screen while the panel showed the card.
The render loop calls this after each frame. Cheap when nothing is
owed (one attribute read); otherwise the usual policy decides, so
the write still waits out the interval and an unchanged frame is
never re-encoded.
"""
if not self._snapshot_owed:
return
try:
if self._writes_suppressed():
return
with self._update_lock:
self._write_snapshot_if_due()
except Exception as e: # pylint: disable=broad-except
self._log_snapshot_failure(e)
def _log_snapshot_failure(self, error: Exception) -> None: def _log_snapshot_failure(self, error: Exception) -> None:
# Snapshot failures must never break display — but they must not # Snapshot failures must never break display — but they must not
# be silent either: the snapshot's mtime is the web UI's display # be silent either: the snapshot's mtime is the web UI's display
+17
View File
@@ -216,6 +216,23 @@ class RenderWatchdog:
def armed(self) -> bool: def armed(self) -> bool:
return self._armed return self._armed
def liveness(self) -> Dict[str, Any]:
"""The heartbeat, in memory: what the control socket's state stream
reports as ``loop``.
``heartbeat_age_seconds`` is the age of the render thread's last beat,
the beat that writes the heartbeat file, so it ages at the same rate
and is judged by the same ``HEARTBEAT_STALE_SECONDS``. None until the
loop has drawn its first frame. Any thread may call this: it only
reads two attributes.
"""
last = self._last_beat
age = None
if self._armed and last is not None:
age = max(self._clock() - last, 0.0)
return {'heartbeat_age_seconds': age, 'armed': self._armed,
'stale_after': HEARTBEAT_STALE_SECONDS}
def _on_render_thread(self) -> bool: def _on_render_thread(self) -> bool:
return self._render_thread is not None and threading.get_ident() == self._render_thread return self._render_thread is not None and threading.get_ident() == self._render_thread
+2 -1
View File
@@ -23,6 +23,7 @@ import requests
from typing import Any, Dict, List from typing import Any, Dict, List
from src.common.api_helper import DEFAULT_HTTP_HEADERS from src.common.api_helper import DEFAULT_HTTP_HEADERS
from src.common.json_body import response_json
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -157,7 +158,7 @@ class DynamicTeamResolver:
response = requests.get(rankings_url, headers=dict(DEFAULT_HTTP_HEADERS), response = requests.get(rankings_url, headers=dict(DEFAULT_HTTP_HEADERS),
timeout=self.request_timeout) timeout=self.request_timeout)
response.raise_for_status() response.raise_for_status()
data = response.json() data = response_json(response)
rankings = {} rankings = {}
rankings_data = data.get('rankings', []) rankings_data = data.get('rankings', [])
+1 -1
View File
@@ -485,7 +485,7 @@ def record_error(
# and only the display service's ever records anything (plugin_executor runs # and only the display service's ever records anything (plugin_executor runs
# the plugins there). The web interface therefore reads a snapshot the display # the plugins there). The web interface therefore reads a snapshot the display
# service publishes to the shared cache directory -- the same channel, and the # service publishes to the shared cache directory -- the same channel, and the
# same file permissions, as display_current_state and plugin_metrics:*: files # same file permissions, as display_current_state and plugin_metrics_snapshot: files
# are 0660 and carry the cache directory's group, so root writes and the web # are 0660 and carry the cache directory's group, so root writes and the web
# user reads, and the other way round for the clear request. # user reads, and the other way round for the clear request.
# #
+262 -1
View File
@@ -10,20 +10,24 @@ blocks for longer than ``timeout`` in total.
from __future__ import annotations from __future__ import annotations
import socket import socket
import threading
import time import time
import uuid import uuid
from typing import Any, Dict, List, Mapping, Optional, Sequence from typing import Any, Callable, Dict, List, Mapping, Optional, Sequence
from src.ipc.contract import ( from src.ipc.contract import (
AWAIT_SECONDS, AWAIT_SECONDS,
MAX_MESSAGE_BYTES, MAX_MESSAGE_BYTES,
PROTOCOL_VERSION, PROTOCOL_VERSION,
SUBSCRIBE_KEEPALIVE_SECONDS,
SUPPORTED_VERSIONS, SUPPORTED_VERSIONS,
Command, Command,
FrameReader, FrameReader,
ProtocolError, ProtocolError,
Request, Request,
Response, Response,
StateEvent,
StateEventKind,
client_socket_paths, client_socket_paths,
decode_message, decode_message,
encode_message, encode_message,
@@ -233,3 +237,260 @@ def hello(client: str = 'web', *, timeout: float = DEFAULT_TIMEOUT_SECONDS,
"""Version negotiation: the result's ``version`` is the one both sides speak.""" """Version negotiation: the result's ``version`` is the one both sides speak."""
return request(Command.HELLO, {'versions': list(SUPPORTED_VERSIONS), 'client': client}, return request(Command.HELLO, {'versions': list(SUPPORTED_VERSIONS), 'client': client},
timeout=timeout, paths=paths) timeout=timeout, paths=paths)
# -- the state stream (stage 3) ---------------------------------------------------------
def state_get(since: Optional[int] = None, epoch: Optional[str] = None, *,
timeout: float = DEFAULT_TIMEOUT_SECONDS,
paths: Optional[Sequence[str]] = None) -> Dict[str, Any]:
"""The display's state now, as a :class:`~src.ipc.contract.StateSnapshot`.
With ``since``/``epoch`` from an earlier answer, an unchanged state comes
back in the short ``changed: false`` form. Raises :class:`ControlError`
(``unknown_command`` from a display older than stage 3).
"""
args: Dict[str, Any] = {}
if since is not None:
args['since'] = since
if epoch is not None:
args['epoch'] = epoch
return request(Command.STATE_GET, args, timeout=timeout, paths=paths)
def snapshot_age(snapshot: Mapping[str, Any], now_mono: Optional[float] = None) -> float:
"""Seconds since ``snapshot`` arrived: ``received_mono`` (set by
:meth:`StateSubscription.latest`) to now; 0 for a one-shot answer."""
received = snapshot.get('received_mono')
if isinstance(received, (int, float)) and not isinstance(received, bool):
now_mono = time.monotonic() if now_mono is None else now_mono
return max(now_mono - float(received), 0.0)
return 0.0
def snapshot_loop_age(snapshot: Mapping[str, Any],
now_mono: Optional[float] = None) -> Optional[float]:
"""The render loop's heartbeat age now, from a state snapshot: the age the
display measured when it answered, plus the time since the answer
arrived. None when the display has no beat to report yet."""
loop = snapshot.get('loop')
if not isinstance(loop, dict):
state = snapshot.get('state')
loop = state.get('loop') if isinstance(state, dict) else None
age = loop.get('heartbeat_age_seconds') if isinstance(loop, dict) else None
if not isinstance(age, (int, float)) or isinstance(age, bool):
return None
return max(float(age), 0.0) + snapshot_age(snapshot, now_mono)
def _merge_volatile(state: Dict[str, Any], volatile: Any) -> None:
"""Fold a tick's ``volatile`` values (``{section: {key: value}}``) into
``state``, copying each section it touches.
These are the timestamps the hub leaves out of its version --
``display.last_updated``, ``on_demand.last_updated``/``remaining``,
``plugins.published_at`` -- and the readers judge freshness by them, so
a copy that only full ``state`` events updated would go stale while the
same mode stayed on screen. Only keys the section already has are taken:
a tick never adds a section or a key the last snapshot did not carry
(a section left out of a truncated snapshot stays out).
"""
if not isinstance(volatile, dict):
return # a display from before ticks carried them
for name, values in volatile.items():
section = state.get(name)
if not isinstance(section, dict) or not isinstance(values, dict):
continue
fresh = {k: v for k, v in values.items() if k in section}
if fresh:
state[name] = dict(section, **fresh)
#: A subscription that has heard nothing for this long is not trusted: the
#: display sends a tick at least every SUBSCRIBE_KEEPALIVE_SECONDS.
SUBSCRIPTION_SILENCE_SECONDS = 3 * SUBSCRIBE_KEEPALIVE_SECONDS
#: Reconnect backoff: the first retry, and the cap. A display that does not
#: know state.subscribe (stage 2 or older) is retried at the cap.
_RECONNECT_MIN_SECONDS = 1.0
_RECONNECT_MAX_SECONDS = 30.0
#: Failures that another try soon will not fix.
_SLOW_RETRY_REASONS = frozenset({'unknown_command', 'unsupported_version', 'disabled',
'unsupported'})
class StateSubscription:
"""One ``state.subscribe`` connection, held on a daemon thread.
Keeps the latest snapshot the display pushed, so a reader answers from
memory (:meth:`latest`). Reconnects with a backoff when the display goes
away. Never raises into the caller: :meth:`latest` is None whenever the
copy cannot be vouched for (not connected, or silent for longer than
``silence``), and the caller falls back.
"""
def __init__(self, paths: Optional[Sequence[str]] = None, *,
silence: float = SUBSCRIPTION_SILENCE_SECONDS,
connect_timeout: float = DEFAULT_TIMEOUT_SECONDS,
clock: Callable[[], float] = time.monotonic):
self._paths = list(paths) if paths is not None else None
self._silence = silence
self._connect_timeout = connect_timeout
self._clock = clock
self._lock = threading.Lock()
self._snapshot: Optional[Dict[str, Any]] = None
self._received: Optional[float] = None
self._connected = False
self._stop = threading.Event()
self._sock: Optional[socket.socket] = None
self._thread: Optional[threading.Thread] = None
#: The reason the last connection ended (a ControlError reason).
self.last_error: Optional[str] = None
#: Full snapshots received: the subscribe answer and each state event.
self.snapshots = 0
# -- the reader's side ---------------------------------------------------
@property
def connected(self) -> bool:
return self._connected
def latest(self) -> Optional[Dict[str, Any]]:
"""A copy of the latest snapshot, with ``received_mono`` (this
process's monotonic clock when it arrived); None when not trusted."""
with self._lock:
if not self._connected or self._snapshot is None or self._received is None:
return None
if self._clock() - self._received > self._silence:
return None
snap = dict(self._snapshot)
snap['received_mono'] = self._received
return snap
# -- lifecycle -----------------------------------------------------------
def start(self) -> 'StateSubscription':
if self._thread is None or not self._thread.is_alive():
self._stop.clear()
self._thread = threading.Thread(target=self._run, name='ledmatrix-state-feed',
daemon=True)
self._thread.start()
return self
def stop(self, timeout: float = 2.0) -> None:
self._stop.set()
sock = self._sock
if sock is not None:
try:
sock.shutdown(socket.SHUT_RDWR)
except OSError:
pass
thread = self._thread
if thread is not None and thread is not threading.current_thread():
thread.join(timeout)
self._thread = None
# -- the feed thread -----------------------------------------------------
def _run(self) -> None:
backoff = _RECONNECT_MIN_SECONDS
while not self._stop.is_set():
snapshots = self.snapshots
try:
self._follow()
except ControlError as e:
self.last_error = e.reason
if e.reason in _SLOW_RETRY_REASONS:
backoff = _RECONNECT_MAX_SECONDS
except Exception as e: # pylint: disable=broad-except
self.last_error = type(e).__name__
finally:
with self._lock:
self._connected = False
sock, self._sock = self._sock, None
if sock is not None:
try:
sock.close()
except OSError:
pass
if self.snapshots != snapshots:
# This connection got as far as the display's state: whatever
# ended it (a restart, most often), it was working, so the
# next try starts from the shortest wait again.
backoff = _RECONNECT_MIN_SECONDS
if self._stop.wait(backoff):
return
backoff = min(backoff * 2, _RECONNECT_MAX_SECONDS)
def _follow(self) -> None:
"""Subscribe, then read events until the connection ends. Raises ControlError."""
if not socket_supported():
raise ControlError('unsupported', 'no Unix sockets on this platform')
candidates = list(self._paths) if self._paths is not None else client_socket_paths()
if not candidates:
raise ControlError('disabled', 'the control socket is turned off')
request_id = str(uuid.uuid4())
payload = encode_message(Request(id=request_id, cmd=Command.STATE_SUBSCRIBE,
args={}).to_dict())
sock = _connect(candidates, time.monotonic() + self._connect_timeout)
self._sock = sock
try:
sock.settimeout(self._connect_timeout)
sock.sendall(payload)
# A read waits for the next event; the display sends one at least
# every keepalive, so this much silence means it is gone.
sock.settimeout(self._silence)
reader = FrameReader(MAX_MESSAGE_BYTES)
first = True
while not self._stop.is_set():
data = sock.recv(65536)
if not data:
raise ControlError('closed', 'the display closed the connection')
for line in reader.feed(data):
obj = decode_message(line)
if first:
response = Response.from_dict(obj)
if not response.ok:
error = response.error
raise ControlError(error.code if error else 'bad_response',
error.message if error else '')
self._store(dict(response.result or {}), full=True)
first = False
continue
event = StateEvent.from_dict(obj)
self._store(event.result, full=event.event == StateEventKind.STATE)
except socket.timeout:
raise ControlError('timeout', 'the display went quiet') from None
except ProtocolError as e:
raise ControlError('bad_response', e.message) from None
except OSError as e:
if self._stop.is_set():
return
raise ControlError('closed', str(e)) from None
def _store(self, result: Dict[str, Any], full: bool) -> None:
now = self._clock()
with self._lock:
if full and isinstance(result.get('state'), dict):
self._snapshot = result
self.snapshots += 1
elif (self._snapshot is not None
and result.get('epoch') == self._snapshot.get('epoch')):
# A tick: nothing changed but the render loop's liveness and
# the volatile keys (timestamps) the writers keep refreshing.
snap = dict(self._snapshot)
state = dict(snap.get('state') or {})
if result.get('version') == snap.get('version'):
_merge_volatile(state, result.get('volatile'))
loop = result.get('loop')
if isinstance(loop, dict):
state['loop'] = loop
snap['loop'] = loop
snap['state'] = state
snap['served_at'] = result.get('served_at', snap.get('served_at'))
self._snapshot = snap
else:
return # a tick before any state, or from another epoch
self._received = now
self._connected = True
+192 -3
View File
@@ -30,6 +30,11 @@ later the state stream). A few commands (:data:`AWAITED_COMMANDS`) are
answered only once the render thread has applied them, or with ``pending`` answered only once the render thread has applied them, or with ``pending``
when it has not within :data:`AWAIT_SECONDS`. when it has not within :data:`AWAIT_SECONDS`.
``state.subscribe`` is the one exception to "one response per request": its
response is followed, on the same connection, by :class:`StateEvent` lines
the display pushes until either side hangs up. Events carry ``event``
instead of ``ok``.
New commands are added within a protocol version: a display that does not New commands are added within a protocol version: a display that does not
know one answers ``unknown_command``, the client falls back, and ``hello`` know one answers ``unknown_command``, the client falls back, and ``hello``
lists the commands a display knows. The version changes only when the lists the commands a display knows. The version changes only when the
@@ -142,11 +147,14 @@ class Command:
ON_DEMAND_STATUS = 'on_demand.status' ON_DEMAND_STATUS = 'on_demand.status'
BRIGHTNESS_SET = 'brightness.set' BRIGHTNESS_SET = 'brightness.set'
PLUGIN_RELOAD = 'plugin.reload' PLUGIN_RELOAD = 'plugin.reload'
STATE_GET = 'state.get'
STATE_SUBSCRIBE = 'state.subscribe'
#: Every command version 1 defines, in the order ``hello`` reports them. #: Every command version 1 defines, in the order ``hello`` reports them.
#: ``brightness.set`` and ``plugin.reload`` came in stage 2, within version 1 #: ``brightness.set`` and ``plugin.reload`` came in stage 2, and ``state.get``
#: (see the module docstring on adding commands). #: and ``state.subscribe`` in stage 3, all within version 1 (see the module
#: docstring on adding commands).
COMMANDS: Tuple[str, ...] = ( COMMANDS: Tuple[str, ...] = (
Command.HELLO, Command.HELLO,
Command.PING, Command.PING,
@@ -155,6 +163,8 @@ COMMANDS: Tuple[str, ...] = (
Command.ON_DEMAND_STATUS, Command.ON_DEMAND_STATUS,
Command.BRIGHTNESS_SET, Command.BRIGHTNESS_SET,
Command.PLUGIN_RELOAD, Command.PLUGIN_RELOAD,
Command.STATE_GET,
Command.STATE_SUBSCRIBE,
) )
#: Commands that are queued for the render thread. #: Commands that are queued for the render thread.
@@ -173,6 +183,25 @@ AWAIT_SECONDS: Dict[str, float] = {
} }
AWAITED_COMMANDS = frozenset(AWAIT_SECONDS) AWAITED_COMMANDS = frozenset(AWAIT_SECONDS)
#: The state stream (stage 3). ``state.subscribe`` turns its connection into
#: a one-way stream of :class:`StateEvent` lines. Subscribers have their own
#: bound, separate from the short request connections, so they can never
#: take the slots a command needs.
MAX_SUBSCRIBERS = 4
#: A subscriber hears from the display at least this often: a ``state``
#: event when something changed, else a ``tick`` carrying the render loop's
#: liveness and the latest volatile timestamps. A client that has heard nothing for a few of these treats its
#: copy as unknown.
SUBSCRIBE_KEEPALIVE_SECONDS = 5.0
#: The shape of the ``state`` object in a state snapshot. Bumped only when a
#: field changes meaning; new fields are added within a schema.
STATE_SCHEMA = 1
#: The sections of a state snapshot, in the order they are documented.
STATE_SECTIONS: Tuple[str, ...] = ('display', 'on_demand', 'brightness', 'plugins', 'loop')
#: Brightness, in percent, as the display's hardware setting takes it. #: Brightness, in percent, as the display's hardware setting takes it.
MIN_BRIGHTNESS = 0 MIN_BRIGHTNESS = 0
MAX_BRIGHTNESS = 100 MAX_BRIGHTNESS = 100
@@ -477,8 +506,67 @@ class PluginReloadArgs:
return cls(plugin_id=plugin_id) return cls(plugin_id=plugin_id)
def _optional_version(args: Mapping[str, Any], key: str) -> Optional[int]:
value = args.get(key)
if value is None:
return None
if not _is_int(value) or value < 0:
raise ProtocolError(ErrorCode.INVALID_ARGS, f'{key} must be a non-negative integer')
return value
def _optional_epoch(args: Mapping[str, Any]) -> Optional[str]:
value = args.get('epoch')
if value is None or value == '':
return None
if not _valid_id(value):
raise ProtocolError(ErrorCode.INVALID_ARGS,
f'epoch must be a printable string of 1-{MAX_ID_LENGTH} characters')
return str(value)
@dataclass(frozen=True)
class StateGetArgs:
"""``state.get``: the display's state, as a versioned snapshot.
With ``since`` and the ``epoch`` it came from, the answer is only
``{changed: false, version, epoch, served_at, loop, volatile}`` while the
state is still at that version, so a poller that already has it is sent
no state -- only the latest values of the keys that do not count as a
change (``volatile``, see :class:`StateSnapshot`).
"""
since: Optional[int] = None
epoch: Optional[str] = None
def to_dict(self) -> Dict[str, Any]:
return {'since': self.since, 'epoch': self.epoch}
@classmethod
def from_dict(cls, args: Mapping[str, Any]) -> 'StateGetArgs':
return cls(since=_optional_version(args, 'since'), epoch=_optional_epoch(args))
@dataclass(frozen=True)
class StateSubscribeArgs:
"""``state.subscribe``: the snapshot now, then a push stream of changes.
The response is the snapshot ``state.get`` returns. After it the
connection carries only :class:`StateEvent` lines from the display: a
``state`` event whenever the state changes (always the latest version,
so a reader that falls behind skips versions instead of queueing them),
and a ``tick`` at least every :data:`SUBSCRIBE_KEEPALIVE_SECONDS`.
"""
def to_dict(self) -> Dict[str, Any]:
return {}
@classmethod
def from_dict(cls, args: Mapping[str, Any]) -> 'StateSubscribeArgs':
return cls()
CommandArgs = Union[HelloArgs, OnDemandStartArgs, OnDemandStopArgs, NoArgs, CommandArgs = Union[HelloArgs, OnDemandStartArgs, OnDemandStopArgs, NoArgs,
BrightnessSetArgs, PluginReloadArgs] BrightnessSetArgs, PluginReloadArgs, StateGetArgs, StateSubscribeArgs]
#: The arguments of a command that goes on the render thread's queue. #: The arguments of a command that goes on the render thread's queue.
QueuedArgs = Union[OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs] QueuedArgs = Union[OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs]
@@ -491,6 +579,8 @@ _ARG_TYPES: Dict[str, Any] = {
Command.ON_DEMAND_STATUS: NoArgs, Command.ON_DEMAND_STATUS: NoArgs,
Command.BRIGHTNESS_SET: BrightnessSetArgs, Command.BRIGHTNESS_SET: BrightnessSetArgs,
Command.PLUGIN_RELOAD: PluginReloadArgs, Command.PLUGIN_RELOAD: PluginReloadArgs,
Command.STATE_GET: StateGetArgs,
Command.STATE_SUBSCRIBE: StateSubscribeArgs,
} }
@@ -563,6 +653,105 @@ class PluginReloadResult(TypedDict):
modes: List[str] modes: List[str]
class LoopState(TypedDict):
"""``loop``: is the render loop still going round?
``heartbeat_age_seconds`` is the age of the render thread's last beat,
measured in memory by the display when it answered -- the same beat that
writes ``display-heartbeat.json``. None until the loop has drawn its first
frame. At ``stale_after`` or more the loop is stalled: the threshold
``/api/v3/health`` uses.
"""
heartbeat_age_seconds: Optional[float]
armed: bool
stale_after: float
class StateSnapshot(TypedDict, total=False):
"""The answer to ``state.get`` and ``state.subscribe``, and the
``result`` of a ``state`` event.
``version`` counts changes to the state within one ``epoch`` (one run of
the display process): a reader that sees a new epoch starts over.
``changed`` is False only for a ``state.get`` whose ``since`` is still
current, and then ``state`` is absent and ``volatile`` is there instead:
``{section: {key: value}}``, the current values of the keys the version
ignores (``display.last_updated``, ``on_demand.last_updated`` and
``remaining``, ``plugins.published_at``). A reader merges them into the
copy it has; they are how it can tell the writers are still publishing.
``served_at`` is the display's wall clock when it answered. ``loop`` is
measured at that moment, so it is also inside ``state``.
``state`` holds the sections in :data:`STATE_SECTIONS`:
* ``display``: what ``display_current_state`` holds (mode, plugin_id,
mode_index, total_modes, on_demand_active, is_display_active,
last_updated);
* ``on_demand``: what ``display_on_demand_state`` holds;
* ``brightness``: ``{brightness, panel_brightness, dimmed}``;
* ``plugins``: the plugin runtime snapshot (``plugin_runtime_snapshot``),
or None when there is none (or it was too large to send);
* ``loop``: :class:`LoopState`.
A section the display has not published yet is None.
"""
schema: int
version: int
epoch: str
pid: int
served_at: float
changed: bool
state: Dict[str, Any]
volatile: Dict[str, Dict[str, Any]]
loop: LoopState
class StateEventKind:
STATE = 'state' # result: a full StateSnapshot, the latest version
TICK = 'tick' # result: {version, epoch, pid, served_at, loop, volatile}; nothing changed
@dataclass(frozen=True)
class StateEvent:
"""One message the display pushes to a subscriber.
``{"v": 1, "id": "<the subscribe request's id>", "event": "state" | "tick",
"result": {...}}``. It has no ``ok``, which is how a reader tells it from
a response.
"""
id: str
event: str
result: Dict[str, Any]
v: int = PROTOCOL_VERSION
def to_dict(self) -> Dict[str, Any]:
return {'v': self.v, 'id': self.id, 'event': self.event, 'result': dict(self.result)}
@classmethod
def from_dict(cls, obj: Any) -> 'StateEvent':
"""Validate an event. Raises :class:`ProtocolError` (BAD_REQUEST)."""
if not isinstance(obj, dict):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'an event must be a JSON object')
version = obj.get('v')
if not _is_int(version):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'v must be an integer')
raw_id = obj.get('id')
if not isinstance(raw_id, str):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'id must be a string')
event = obj.get('event')
if event not in (StateEventKind.STATE, StateEventKind.TICK):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'event must be "state" or "tick"')
result = obj.get('result')
if not isinstance(result, dict):
raise ProtocolError(ErrorCode.BAD_REQUEST, 'result must be a JSON object')
return cls(id=raw_id, event=event, result=result, v=version)
def is_event(obj: Any) -> bool:
"""Whether a decoded message is a pushed event rather than a response."""
return isinstance(obj, dict) and 'event' in obj and 'ok' not in obj
def negotiate_version(client_versions: Tuple[int, ...]) -> Optional[int]: def negotiate_version(client_versions: Tuple[int, ...]) -> Optional[int]:
"""The highest version both sides speak, or None.""" """The highest version both sides speak, or None."""
common = set(client_versions) & set(SUPPORTED_VERSIONS) common = set(client_versions) & set(SUPPORTED_VERSIONS)
+369 -16
View File
@@ -14,6 +14,13 @@ every kind of screen. An awaited command (``brightness.set``,
``plugin.reload``) carries a :class:`CommandOutcome` that the render thread ``plugin.reload``) carries a :class:`CommandOutcome` that the render thread
fills in; its connection thread waits for that, bounded, before answering. fills in; its connection thread waits for that, bounded, before answering.
The state stream (stage 3): the display publishes what it is doing into a
:class:`StateHub`, in memory, and ``state.get`` / ``state.subscribe`` read
it. A subscriber's connection gives back its request slot, takes one of
:data:`~src.ipc.contract.MAX_SUBSCRIBERS`, and is pushed the latest version
on every change plus a keepalive tick, from its own thread: publishing never
waits for a reader, and a reader that stops reading is dropped.
Robustness rules, because this runs inside the display process: Robustness rules, because this runs inside the display process:
* every connection has its own daemon thread, at most :data:`MAX_CLIENTS` at * every connection has its own daemon thread, at most :data:`MAX_CLIENTS` at
@@ -37,6 +44,7 @@ server checks them again: root, its own user, or a member of that group.
from __future__ import annotations from __future__ import annotations
import json
import logging import logging
import os import os
import queue import queue
@@ -45,8 +53,9 @@ import stat
import struct import struct
import threading import threading
import time import time
import uuid
from dataclasses import dataclass, field from dataclasses import dataclass, field
from typing import Any, Callable, Dict, FrozenSet, List, Mapping, Optional from typing import Any, Callable, Dict, FrozenSet, Iterable, List, Mapping, Optional, Tuple
from src.ipc.contract import ( from src.ipc.contract import (
AWAIT_SECONDS, AWAIT_SECONDS,
@@ -55,8 +64,12 @@ from src.ipc.contract import (
DEFAULT_SOCKET_DIR, DEFAULT_SOCKET_DIR,
DEFAULT_SOCKET_PATH, DEFAULT_SOCKET_PATH,
MAX_MESSAGE_BYTES, MAX_MESSAGE_BYTES,
MAX_SUBSCRIBERS,
PROTOCOL_VERSION, PROTOCOL_VERSION,
QUEUED_COMMANDS, QUEUED_COMMANDS,
STATE_SCHEMA,
STATE_SECTIONS,
SUBSCRIBE_KEEPALIVE_SECONDS,
SUPPORTED_VERSIONS, SUPPORTED_VERSIONS,
AckResult, AckResult,
BrightnessSetArgs, BrightnessSetArgs,
@@ -72,6 +85,9 @@ from src.ipc.contract import (
QueuedArgs, QueuedArgs,
Request, Request,
Response, Response,
StateEvent,
StateEventKind,
StateGetArgs,
configured_socket_path, configured_socket_path,
decode_message, decode_message,
dev_socket_path, dev_socket_path,
@@ -181,6 +197,241 @@ class QueuedCommand:
self.outcome.fail(code, message) self.outcome.fail(code, message)
# -- the state stream (stage 3) ----------------------------------------------------------
#: How long a ``state.get`` keeps the display counting its readers as served
#: over the socket (:meth:`StateHub.readers_active`). A subscriber counts for
#: as long as it is connected.
READER_WINDOW_SECONDS = 60.0
#: Room kept for the envelope (``v``, ``id``, ``event``) around a snapshot,
#: within MAX_MESSAGE_BYTES.
_ENVELOPE_ROOM = 512
_MISSING = object()
LoopProbe = Callable[[], Mapping[str, Any]]
def _fingerprint(value: Optional[Mapping[str, Any]], volatile: Iterable[str]) -> Any:
"""What a section's version is judged on: the value minus its volatile keys
(timestamps that move on every publish without anything changing)."""
if value is None:
return None
skip = frozenset(volatile)
return {k: v for k, v in value.items() if k not in skip} if skip else dict(value)
def _volatile_values(sections: Mapping[str, Optional[Dict[str, Any]]],
volatile: Mapping[str, FrozenSet[str]]) -> Dict[str, Dict[str, Any]]:
"""``{section: {key: value}}``: the volatile keys each published section
has now. The sections are never mutated after publish (a publish swaps
in a new dict), so reading them outside the lock is safe."""
values: Dict[str, Dict[str, Any]] = {}
for name, keys in volatile.items():
value = sections.get(name)
if not keys or not isinstance(value, dict):
continue
present = {k: value[k] for k in keys if k in value}
if present:
values[name] = present
return values
def _unknown_loop() -> Dict[str, Any]:
return {'heartbeat_age_seconds': None, 'armed': False, 'stale_after': None}
def fit_snapshot(snapshot: Dict[str, Any]) -> Dict[str, Any]:
"""``snapshot``, or a copy without the plugin runtime section when the
message would be over MAX_MESSAGE_BYTES (hundreds of plugins). The
reader then falls back to the cache for that section only; ``truncated``
says which was left out."""
state = snapshot.get('state')
if not isinstance(state, dict) or state.get('plugins') is None:
return snapshot
try:
size = len(json.dumps(snapshot, separators=(',', ':'), ensure_ascii=True,
allow_nan=False))
except (TypeError, ValueError):
size = MAX_MESSAGE_BYTES
if size <= MAX_MESSAGE_BYTES - _ENVELOPE_ROOM:
return snapshot
logger.warning("State snapshot is %d bytes; sending it without the plugin runtime "
"section", size)
trimmed = dict(snapshot)
trimmed['state'] = dict(state, plugins=None)
trimmed['truncated'] = ['plugins']
return trimmed
class StateHub:
"""The display's live state, in memory, for ``state.get`` and ``state.subscribe``.
Writers publish whole sections (:meth:`publish`): the render thread
publishes ``display``, ``on_demand`` and ``brightness``, and the plugin
runtime publisher's thread publishes ``plugins``. Each section has one
writer. ``loop`` is not published: it is measured when a reader asks
(``loop_probe``), so it keeps ageing while the render thread is stuck.
The version goes up when a section's value changes, ignoring the keys
the publisher names as volatile (timestamps). Those keys still carry
news -- ``display.last_updated`` is the render thread's proof of life --
so the short ``changed: false`` answer, which is what a subscriber's
tick carries, has their current values in ``volatile``; a reader merges
them into its copy. Publishing never blocks on
a reader: the lock is held only to swap a dict reference and compare it,
and every socket write happens on the reader's own thread, outside it.
A reader that is slow gets the latest version when it next asks, not
every version in between.
"""
def __init__(self, loop_probe: Optional[LoopProbe] = None, *,
clock: Callable[[], float] = time.monotonic,
wall_clock: Callable[[], float] = time.time,
epoch: Optional[str] = None, pid: Optional[int] = None,
reader_window: float = READER_WINDOW_SECONDS):
self._cond = threading.Condition(threading.Lock())
self._sections: Dict[str, Optional[Dict[str, Any]]] = {}
self._fingerprints: Dict[str, Any] = {}
self._volatile: Dict[str, FrozenSet[str]] = {}
self._version = 0
self.epoch = epoch or uuid.uuid4().hex[:16]
self.pid = os.getpid() if pid is None else pid
self._loop_probe = loop_probe
self._clock = clock
self._wall_clock = wall_clock
self._reader_window = reader_window
self._last_read: Optional[float] = None
self._subscribers = 0
@property
def version(self) -> int:
return self._version
@property
def subscribers(self) -> int:
return self._subscribers
# -- writers -------------------------------------------------------------
def publish(self, section: str, value: Optional[Mapping[str, Any]],
volatile: Iterable[str] = ()) -> bool:
"""Store a section's latest value; True when that is a new version.
The value is copied (one level), so the caller may reuse its dict.
"""
stored = None if value is None else dict(value)
skip = frozenset(volatile)
fingerprint = _fingerprint(stored, skip)
with self._cond:
self._sections[section] = stored
self._volatile[section] = skip
if self._fingerprints.get(section, _MISSING) == fingerprint:
return False
self._fingerprints[section] = fingerprint
self._version += 1
self._cond.notify_all()
return True
def wake(self) -> None:
"""Wake every waiting reader (the server is closing)."""
with self._cond:
self._cond.notify_all()
# -- readers -------------------------------------------------------------
def loop(self) -> Dict[str, Any]:
"""The render loop's liveness now. Never raises."""
if self._loop_probe is None:
return _unknown_loop()
try:
return dict(self._loop_probe())
except Exception: # pylint: disable=broad-except
logger.debug("Render loop liveness probe failed", exc_info=True)
return _unknown_loop()
def snapshot(self, since: Optional[int] = None,
epoch: Optional[str] = None) -> Dict[str, Any]:
"""The :class:`~src.ipc.contract.StateSnapshot` now.
``since`` with this hub's ``epoch``, still the current version, gives
the short ``changed: false`` form, with ``volatile``: each section's
volatile keys at their latest values (the rest of the section is
what the reader already has).
"""
with self._cond:
version = self._version
sections = dict(self._sections)
volatile = dict(self._volatile)
loop = self.loop()
result: Dict[str, Any] = {
'schema': STATE_SCHEMA,
'version': version,
'epoch': self.epoch,
'pid': self.pid,
'served_at': self._wall_clock(),
'loop': loop,
}
if since is not None and epoch == self.epoch and since == version:
result['changed'] = False
result['volatile'] = _volatile_values(sections, volatile)
return result
state: Dict[str, Any] = {name: sections.get(name) for name in STATE_SECTIONS
if name != 'loop'}
state['loop'] = loop
result['changed'] = True
result['state'] = state
return result
def wait_for_change(self, version: int, timeout: float,
stop: Optional[threading.Event] = None) -> bool:
"""Block up to ``timeout`` for a version other than ``version``."""
with self._cond:
self._cond.wait_for(
lambda: self._version != version or (stop is not None and stop.is_set()),
timeout)
return self._version != version
# -- who is reading ------------------------------------------------------
def note_read(self) -> None:
self._last_read = self._clock()
def subscriber_joined(self) -> None:
with self._cond:
self._subscribers += 1
def subscriber_left(self) -> None:
with self._cond:
self._subscribers = max(0, self._subscribers - 1)
self._last_read = self._clock()
def readers_active(self) -> bool:
"""Is the socket serving state readers? A subscriber is connected, or a
``state.get`` came within the reader window. The display uses this to
write the cache copies of the same state less often."""
if self._subscribers > 0:
return True
last = self._last_read
return last is not None and self._clock() - last < self._reader_window
class _Slot:
"""A connection slot, released once (a subscriber gives its back early)."""
def __init__(self, semaphore: threading.BoundedSemaphore):
self._semaphore = semaphore
self._held = True
self._lock = threading.Lock()
def release(self) -> None:
with self._lock:
if self._held:
self._held = False
self._semaphore.release()
# -- peer credentials ------------------------------------------------------------------ # -- peer credentials ------------------------------------------------------------------
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -315,8 +566,14 @@ class ControlServer:
message_timeout: float = MESSAGE_TIMEOUT_SECONDS, message_timeout: float = MESSAGE_TIMEOUT_SECONDS,
idle_timeout: float = IDLE_TIMEOUT_SECONDS, idle_timeout: float = IDLE_TIMEOUT_SECONDS,
check_peer: bool = True, check_peer: bool = True,
await_seconds: Optional[Mapping[str, float]] = None): await_seconds: Optional[Mapping[str, float]] = None,
state_hub: Optional[StateHub] = None,
max_subscribers: int = MAX_SUBSCRIBERS,
keepalive: float = SUBSCRIBE_KEEPALIVE_SECONDS):
self.path = path self.path = path
self.state_hub = state_hub
self._subscriber_slots = threading.BoundedSemaphore(max_subscribers)
self._keepalive = keepalive
self._await_seconds: Dict[str, float] = dict(AWAIT_SECONDS) self._await_seconds: Dict[str, float] = dict(AWAIT_SECONDS)
if await_seconds: if await_seconds:
self._await_seconds.update(await_seconds) self._await_seconds.update(await_seconds)
@@ -378,6 +635,8 @@ class ControlServer:
"""Stop accepting and remove the socket file (only if it is still ours).""" """Stop accepting and remove the socket file (only if it is still ours)."""
self._stopping.set() self._stopping.set()
self._close_socket() self._close_socket()
if self.state_hub is not None:
self.state_hub.wake() # subscribers see _stopping and hang up
thread = self._thread thread = self._thread
if thread is not None and thread is not threading.current_thread(): if thread is not None and thread is not threading.current_thread():
thread.join(timeout=2.0) thread.join(timeout=2.0)
@@ -549,6 +808,7 @@ class ControlServer:
def _serve(self, conn: socket.socket) -> None: def _serve(self, conn: socket.socket) -> None:
"""One connection: authenticate, then answer requests until it ends.""" """One connection: authenticate, then answer requests until it ends."""
slot = _Slot(self._slots)
try: try:
conn.settimeout(self._io_timeout) conn.settimeout(self._io_timeout)
peer = peer_credentials(conn) peer = peer_credentials(conn)
@@ -557,7 +817,7 @@ class ControlServer:
"this user or group %s", peer.pid, peer.uid, peer.gid, self._group) "this user or group %s", peer.pid, peer.uid, peer.gid, self._group)
self._send(conn, Response.failure(None, ErrorCode.FORBIDDEN, 'not permitted')) self._send(conn, Response.failure(None, ErrorCode.FORBIDDEN, 'not permitted'))
return return
self._read_requests(conn, peer) self._read_requests(conn, peer, slot)
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Control socket connection failed") logger.exception("Control socket connection failed")
finally: finally:
@@ -565,7 +825,7 @@ class ControlServer:
conn.close() conn.close()
except OSError: except OSError:
pass pass
self._slots.release() slot.release()
def _peer_ok(self, peer: PeerCredentials) -> bool: def _peer_ok(self, peer: PeerCredentials) -> bool:
groups = None groups = None
@@ -573,7 +833,8 @@ class ControlServer:
groups = process_groups(peer.pid) groups = process_groups(peer.pid)
return peer_allowed(peer, self._own_uid, self._group, groups) return peer_allowed(peer, self._own_uid, self._group, groups)
def _read_requests(self, conn: socket.socket, peer: Optional[PeerCredentials]) -> None: def _read_requests(self, conn: socket.socket, peer: Optional[PeerCredentials],
slot: Optional[_Slot] = None) -> None:
reader = FrameReader(MAX_MESSAGE_BYTES) reader = FrameReader(MAX_MESSAGE_BYTES)
idle_since = time.monotonic() idle_since = time.monotonic()
message_started: Optional[float] = None message_started: Optional[float] = None
@@ -598,7 +859,13 @@ class ControlServer:
self._send(conn, Response.failure(None, e.code, e.message)) self._send(conn, Response.failure(None, e.code, e.message))
return # can't find the next message boundary: hang up return # can't find the next message boundary: hang up
for line in lines: for line in lines:
if not self._send(conn, self.handle_line(line, peer)): response, cmd = self._handle(line, peer)
if cmd == Command.STATE_SUBSCRIBE and response.ok:
# The connection becomes a one-way stream; anything the
# client sent after the subscribe is ignored.
self._subscribe(conn, response, slot)
return
if not self._send(conn, response):
return return
if reader.pending: if reader.pending:
if message_started is None or lines: if message_started is None or lines:
@@ -624,20 +891,91 @@ class ControlServer:
# -- requests -------------------------------------------------------------------- # -- requests --------------------------------------------------------------------
def handle_line(self, line: bytes, peer: Optional[PeerCredentials] = None) -> Response: def handle_line(self, line: bytes, peer: Optional[PeerCredentials] = None) -> Response:
"""Answer one request line. Never raises.""" """Answer one request line. Never raises.
A ``state.subscribe`` answered here gets its snapshot only; the
stream that follows needs a connection (``_read_requests``).
"""
return self._handle(line, peer)[0]
def _handle(self, line: bytes,
peer: Optional[PeerCredentials]) -> Tuple[Response, Optional[str]]:
"""The response to one line, and the command it answered (when known)."""
request_id: Optional[str] = None request_id: Optional[str] = None
cmd: Optional[str] = None
try: try:
obj = decode_message(line) obj = decode_message(line)
raw_id = obj.get('id') raw_id = obj.get('id')
request_id = raw_id if isinstance(raw_id, str) and len(raw_id) <= 128 else None request_id = raw_id if isinstance(raw_id, str) and len(raw_id) <= 128 else None
request = Request.from_dict(obj) request = Request.from_dict(obj)
request_id = request.id request_id = request.id
return self._dispatch(request, peer) cmd = request.cmd
return self._dispatch(request, peer), cmd
except ProtocolError as e: except ProtocolError as e:
return Response.failure(e.request_id or request_id, e.code, e.message) return Response.failure(e.request_id or request_id, e.code, e.message), cmd
except Exception: # pylint: disable=broad-except except Exception: # pylint: disable=broad-except
logger.exception("Control socket handler failed") logger.exception("Control socket handler failed")
return Response.failure(request_id, ErrorCode.INTERNAL, 'internal error') return Response.failure(request_id, ErrorCode.INTERNAL, 'internal error'), cmd
# -- the state stream ----------------------------------------------------------
def _subscribe(self, conn: socket.socket, response: Response,
slot: Optional[_Slot]) -> None:
"""Answer a ``state.subscribe`` and push state events until it ends.
Subscribers have their own bound (MAX_SUBSCRIBERS) and give their
request slot back, so a few browsers watching never use up the slots
commands need. Everything here runs on this connection's thread: a
reader that does not keep up only stalls its own sends, and one that
stops reading for a whole IO timeout is dropped. The render thread
only ever publishes into the hub.
"""
hub = self.state_hub
if hub is None or not self._subscriber_slots.acquire(blocking=False):
self._send(conn, Response.failure(response.id, ErrorCode.BUSY,
'too many state subscribers', v=response.v))
return
if slot is not None:
slot.release()
hub.subscriber_joined()
try:
if not self._send(conn, response):
return
result = response.result or {}
version = result.get('version', -1)
sub_id = response.id or ''
logger.debug("Control socket: state subscriber joined at version %s", version)
while not self._stopping.is_set():
hub.wait_for_change(version, self._keepalive, self._stopping)
if self._stopping.is_set():
return
snap = hub.snapshot(since=version, epoch=hub.epoch)
if snap.get('changed'):
version = snap['version']
event = StateEvent(sub_id, StateEventKind.STATE, fit_snapshot(snap),
v=response.v)
else:
event = StateEvent(sub_id, StateEventKind.TICK, snap, v=response.v)
if not self._send_event(conn, event):
return
finally:
hub.subscriber_left()
self._subscriber_slots.release()
def _send_event(self, conn: socket.socket, event: StateEvent) -> bool:
try:
data = encode_message(event.to_dict())
except ProtocolError as e:
logger.error("Control socket state event not sent: %s", e.message)
return False
try:
conn.sendall(data)
return True
except socket.timeout:
logger.info("Control socket: dropping a state subscriber that stopped reading")
return False
except OSError:
return False
def _dispatch(self, request: Request, peer: Optional[PeerCredentials]) -> Response: def _dispatch(self, request: Request, peer: Optional[PeerCredentials]) -> Response:
if request.cmd == Command.HELLO: if request.cmd == Command.HELLO:
@@ -678,6 +1016,18 @@ class ControlServer:
v=request.v) v=request.v)
return Response.success(request.id, self._status_provider(), v=request.v) return Response.success(request.id, self._status_provider(), v=request.v)
if request.cmd in (Command.STATE_GET, Command.STATE_SUBSCRIBE):
hub = self.state_hub
if hub is None:
return Response.failure(request.id, ErrorCode.INTERNAL, 'no state available',
v=request.v)
if isinstance(args, StateGetArgs):
hub.note_read()
snap = hub.snapshot(since=args.since, epoch=args.epoch)
else:
snap = hub.snapshot()
return Response.success(request.id, fit_snapshot(snap), v=request.v)
if request.cmd in QUEUED_COMMANDS and isinstance(args, ( if request.cmd in QUEUED_COMMANDS and isinstance(args, (
OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs)): OnDemandStartArgs, OnDemandStopArgs, BrightnessSetArgs, PluginReloadArgs)):
awaited = request.cmd in AWAITED_COMMANDS awaited = request.cmd in AWAITED_COMMANDS
@@ -728,22 +1078,25 @@ class ControlServer:
def start_control_server(status_provider: Optional[StatusProvider] = None, def start_control_server(status_provider: Optional[StatusProvider] = None,
cache_dir: Optional[str] = None, cache_dir: Optional[str] = None,
environ: Optional[Mapping[str, str]] = None) -> Optional[ControlServer]: environ: Optional[Mapping[str, str]] = None,
state_hub: Optional[StateHub] = None) -> Optional[ControlServer]:
"""Start the display's control socket, or return None when it can't run. """Start the display's control socket, or return None when it can't run.
None covers Windows, ``LEDMATRIX_CONTROL_SOCKET=off`` and any failure to None covers Windows, ``LEDMATRIX_CONTROL_SOCKET=off`` and any failure to
bind; in every case the web interface falls back to the file mailbox. bind; in every case the web interface falls back to the file mailbox
and to the cache keys the display still writes.
""" """
path = server_socket_path(environ) path = server_socket_path(environ)
if path is None: if path is None:
logger.debug("Control socket disabled or unsupported here; using the file mailbox only") logger.debug("Control socket disabled or unsupported here; using the file mailbox only")
return None return None
server = ControlServer(path, status_provider, resolve_socket_group(cache_dir)) server = ControlServer(path, status_provider, resolve_socket_group(cache_dir),
state_hub=state_hub)
return server if server.start() else None return server if server.start() else None
__all__ = [ __all__ = [
'CommandOutcome', 'ControlServer', 'PeerCredentials', 'QueuedCommand', 'StatusProvider', 'CommandOutcome', 'ControlServer', 'PeerCredentials', 'QueuedCommand', 'StateHub',
'peer_allowed', 'peer_credentials', 'process_groups', 'resolve_socket_group', 'StatusProvider', 'fit_snapshot', 'peer_allowed', 'peer_credentials', 'process_groups',
'server_socket_path', 'start_control_server', 'PROTOCOL_VERSION', 'resolve_socket_group', 'server_socket_path', 'start_control_server', 'PROTOCOL_VERSION',
] ]
+3 -2
View File
@@ -21,6 +21,7 @@ from PIL.PngImagePlugin import PngInfo
from requests.adapters import HTTPAdapter from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry from urllib3.util.retry import Retry
from src.common.api_helper import DEFAULT_HTTP_HEADERS from src.common.api_helper import DEFAULT_HTTP_HEADERS
from src.common.json_body import response_json
from src.common.logo_helper import MAX_LOGO_BYTES from src.common.logo_helper import MAX_LOGO_BYTES
from src.common.permission_utils import ( from src.common.permission_utils import (
ensure_directory_permissions, ensure_directory_permissions,
@@ -481,7 +482,7 @@ class LogoDownloader:
logger.info(f"Fetching team data for {league} from ESPN API...") logger.info(f"Fetching team data for {league} from ESPN API...")
response = self.session.get(api_url, params={'limit':1000},headers=self.headers, timeout=self.request_timeout) response = self.session.get(api_url, params={'limit':1000},headers=self.headers, timeout=self.request_timeout)
response.raise_for_status() response.raise_for_status()
data: Dict = response.json() data: Dict = response_json(response)
logger.info(f"Successfully fetched team data for {league}") logger.info(f"Successfully fetched team data for {league}")
return data return data
@@ -505,7 +506,7 @@ class LogoDownloader:
logger.info(f"Fetching team data for team {team_id} in {league} from ESPN API...") logger.info(f"Fetching team data for team {team_id} in {league} from ESPN API...")
response = self.session.get(f"{api_url}/{team_id}", headers=self.headers, timeout=self.request_timeout) response = self.session.get(f"{api_url}/{team_id}", headers=self.headers, timeout=self.request_timeout)
response.raise_for_status() response.raise_for_status()
data: Dict = response.json() data: Dict = response_json(response)
logger.info(f"Successfully fetched team data for {team_id} in {league}") logger.info(f"Successfully fetched team data for {team_id} in {league}")
return data return data
+33 -2
View File
@@ -3,15 +3,46 @@ LEDMatrix Plugin System
This module provides the core plugin infrastructure for the LEDMatrix project. This module provides the core plugin infrastructure for the LEDMatrix project.
It enables dynamic loading, management, and discovery of display plugins. It enables dynamic loading, management, and discovery of display plugins.
BasePlugin and PluginManager are imported on first use (PEP 562), not when
the package is imported: the web interface imports several submodules
(store_manager, schema_manager, ...) and never needs PluginManager, which
pulls in the loader, executor and the shared helpers behind them.
``from src.plugin_system import BasePlugin`` works as before and returns the
same class.
""" """
import importlib
from typing import TYPE_CHECKING, Any, Dict, List, Tuple
__version__ = "1.0.0" __version__ = "1.0.0"
from .base_plugin import BasePlugin if TYPE_CHECKING:
from .plugin_manager import PluginManager from .base_plugin import BasePlugin
from .plugin_manager import PluginManager
#: Exported name -> (module it lives in, attribute name there).
_LAZY: Dict[str, Tuple[str, str]] = {
'BasePlugin': ('src.plugin_system.base_plugin', 'BasePlugin'),
'PluginManager': ('src.plugin_system.plugin_manager', 'PluginManager'),
}
__all__ = [ __all__ = [
'BasePlugin', 'BasePlugin',
'PluginManager', 'PluginManager',
] ]
def __getattr__(name: str) -> Any:
"""Import an exported name on first access (PEP 562); see src.common."""
try:
module_name, attr = _LAZY[name]
except KeyError:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}") from None
value = getattr(importlib.import_module(module_name), attr) # nosemgrep: python.lang.security.audit.non-literal-import.non-literal-import -- module_name comes from the fixed _LAZY table
globals()[name] = value
return value
def __dir__() -> List[str]:
return sorted(set(globals()) | set(__all__))
+9 -2
View File
@@ -48,12 +48,19 @@ class PluginOperation:
completed_at: Optional[datetime] = None completed_at: Optional[datetime] = None
def to_dict(self) -> Dict[str, Any]: def to_dict(self) -> Dict[str, Any]:
"""Convert operation to dictionary for serialization.""" """Convert operation to dictionary for serialization.
Parameters whose name starts with ``_`` are internal and left out:
PluginOperationQueue keeps the operation's callback there as
``_callback`` until its worker runs it, and a pending operation's
status answered 500 because that function cannot be serialized.
"""
return { return {
'operation_id': self.operation_id, 'operation_id': self.operation_id,
'operation_type': self.operation_type.value, 'operation_type': self.operation_type.value,
'plugin_id': self.plugin_id, 'plugin_id': self.plugin_id,
'parameters': self.parameters, 'parameters': {key: value for key, value in self.parameters.items()
if not str(key).startswith('_')},
'status': self.status.value, 'status': self.status.value,
'progress': self.progress, 'progress': self.progress,
'message': self.message, 'message': self.message,
+4 -1
View File
@@ -92,7 +92,10 @@ class PluginExecutor:
with plugin_scope(plugin_id): with plugin_scope(plugin_id):
result_container['value'] = operation() result_container['value'] = operation()
result_container['completed'] = True result_container['completed'] = True
except Exception as e: except BaseException as e: # pylint: disable=broad-except
# asyncio.CancelledError and SystemExit too: uncaught, one
# ended this thread with 'completed' unset, and an operation
# that failed at once was reported as timing out.
result_container['exception'] = e result_container['exception'] = e
result_container['completed'] = True result_container['completed'] = True
+75 -4
View File
@@ -199,9 +199,22 @@ def contained_plugin_dir(plugin_dir: Path, plugins_dir: Path) -> Optional[str]:
name that came out of ``os.scandir()`` on the trusted root carries no name that came out of ``os.scandir()`` on the trusted root carries no
taint, which is a real containment guarantee (and one CodeQL's taint, which is a real containment guarantee (and one CodeQL's
path-injection query can follow), not a string sanitiser. path-injection query can follow), not a string sanitiser.
The entry looked for is the one ``plugin_dir`` itself names when it sits
directly in ``plugins_dir``: for a dev plugin symlinked in under its id,
the link's name. Resolving the link first and looking for the target's
folder name refused ``plugins/foo -> ~/.ledmatrix-dev-plugins/ledmatrix-foo``
(what ``dev_plugin_setup.sh link-github foo <url>`` makes), so the plugin
never loaded. Any other path is resolved and matched by its final name,
as before.
""" """
plugin_dir_real = os.path.realpath(str(plugin_dir))
plugins_dir_real = os.path.realpath(str(plugins_dir)) plugins_dir_real = os.path.realpath(str(plugins_dir))
plugin_dir_abs = os.path.abspath(str(plugin_dir))
if os.path.realpath(os.path.dirname(plugin_dir_abs)) == plugins_dir_real:
matched_name = find_trusted_subdir(plugins_dir_real, os.path.basename(plugin_dir_abs))
if matched_name is not None:
return os.path.join(plugins_dir_real, matched_name)
plugin_dir_real = os.path.realpath(str(plugin_dir))
matched_name = find_trusted_subdir(plugins_dir_real, os.path.basename(plugin_dir_real)) matched_name = find_trusted_subdir(plugins_dir_real, os.path.basename(plugin_dir_real))
if matched_name is None: if matched_name is None:
return None return None
@@ -243,6 +256,10 @@ class PluginLoader:
self.logger = logger or get_logger(__name__) self.logger = logger or get_logger(__name__)
self._loaded_modules: Dict[str, Any] = {} self._loaded_modules: Dict[str, Any] = {}
self._plugin_module_registry: Dict[str, set] = {} # Maps plugin_id to set of module names self._plugin_module_registry: Dict[str, set] = {} # Maps plugin_id to set of module names
# plugin_id -> {dotted name: module} for the modules of the plugin's
# own packages (``providers.feed``). They keep their names while the
# plugin runs and are dropped with it; see _iter_plugin_submodules.
self._plugin_submodules: Dict[str, Dict[str, Any]] = {}
# Lock to serialize module loading when plugins share module names # Lock to serialize module loading when plugins share module names
# (e.g., scroll_display.py, game_renderer.py across sport plugins). # (e.g., scroll_display.py, game_renderer.py across sport plugins).
# During exec_module, bare-name sub-modules temporarily appear in # During exec_module, bare-name sub-modules temporarily appear in
@@ -449,6 +466,45 @@ class PluginLoader:
continue continue
return result return result
@staticmethod
def _iter_plugin_submodules(
plugin_dir: Path, before_keys: set
) -> list:
"""Return dotted-name modules from plugin_dir added after before_keys.
The modules of a package the plugin ships (``providers.feed`` from
``providers/feed.py``). _iter_plugin_bare_modules skips them, so the
bare ``providers`` was namespaced and dropped on unload while
``providers.feed`` stayed in sys.modules: a reload after a store update
imported a fresh ``providers`` and then got the old ``feed`` back from
the cache, running the new manager.py against the old helpers until the
display restarted.
A module counts when its ``__file__`` -- or, for a namespace package,
which has none, every ``__path__`` entry -- is inside plugin_dir, so a
library the plugin imports (``requests.adapters``) never does.
Returns a list of (mod_name, module) tuples.
"""
resolved_dir = plugin_dir.resolve()
result = []
for key in set(sys.modules.keys()) - before_keys:
if "." not in key:
continue
mod = sys.modules.get(key)
if mod is None:
continue
mod_file = getattr(mod, "__file__", None)
locations = [mod_file] if mod_file else list(getattr(mod, "__path__", None) or [])
if not locations:
continue
try:
if all(Path(loc).resolve().is_relative_to(resolved_dir) for loc in locations):
result.append((key, mod))
except (ValueError, TypeError, OSError):
continue
return result
def _evict_stale_bare_modules(self, plugin_dir: Path) -> dict: def _evict_stale_bare_modules(self, plugin_dir: Path) -> dict:
"""Temporarily remove bare-name sys.modules entries from other plugins. """Temporarily remove bare-name sys.modules entries from other plugins.
@@ -527,6 +583,13 @@ class PluginLoader:
# Track for cleanup during unload # Track for cleanup during unload
self._plugin_module_registry[plugin_id] = namespaced_names self._plugin_module_registry[plugin_id] = namespaced_names
# The modules of the plugin's own packages keep their dotted names
# while it runs -- as they always have, so the package and its
# children stay a matching set in sys.modules -- and are dropped
# with the plugin by unregister_plugin_modules().
self._plugin_submodules[plugin_id] = dict(
self._iter_plugin_submodules(plugin_dir, before_keys))
if namespaced_names: if namespaced_names:
self.logger.info( self.logger.info(
"Namespace-isolated %d module(s) for plugin %s", "Namespace-isolated %d module(s) for plugin %s",
@@ -537,10 +600,16 @@ class PluginLoader:
"""Remove namespaced sub-modules and cached module for a plugin from sys.modules. """Remove namespaced sub-modules and cached module for a plugin from sys.modules.
Called by PluginManager during unload to clean up all module entries Called by PluginManager during unload to clean up all module entries
that were created when the plugin was loaded. that were created when the plugin was loaded, including the dotted
modules of its packages. A dotted name is dropped only while it still
holds this plugin's module: the name is not namespaced, so another
plugin may have put its own there since.
""" """
for ns_name in self._plugin_module_registry.pop(plugin_id, set()): for ns_name in self._plugin_module_registry.pop(plugin_id, set()):
sys.modules.pop(ns_name, None) sys.modules.pop(ns_name, None)
for name, mod in self._plugin_submodules.pop(plugin_id, {}).items():
if sys.modules.get(name) is mod:
sys.modules.pop(name, None)
self._loaded_modules.pop(plugin_id, None) self._loaded_modules.pop(plugin_id, None)
def load_module( def load_module(
@@ -646,11 +715,13 @@ class PluginLoader:
if evicted_name not in sys.modules: if evicted_name not in sys.modules:
sys.modules[evicted_name] = evicted_mod sys.modules[evicted_name] = evicted_mod
# Clean up the partially-initialized main module and any # Clean up the partially-initialized main module and any
# bare-name sub-modules that were added during exec_module # bare-name or package sub-modules that were added during
# so they don't leak into subsequent plugin loads. # exec_module so they don't leak into subsequent plugin loads.
sys.modules.pop(module_name, None) sys.modules.pop(module_name, None)
for key, _ in self._iter_plugin_bare_modules(plugin_dir, before_keys): for key, _ in self._iter_plugin_bare_modules(plugin_dir, before_keys):
sys.modules.pop(key, None) sys.modules.pop(key, None)
for key, _ in self._iter_plugin_submodules(plugin_dir, before_keys):
sys.modules.pop(key, None)
raise raise
self._loaded_modules[plugin_id] = module self._loaded_modules[plugin_id] = module
+10 -4
View File
@@ -1163,7 +1163,7 @@ class PluginManager:
def _record_update_failure( def _record_update_failure(
self, self,
plugin_id: str, plugin_id: str,
exc: Optional[Exception] = None, exc: Optional[BaseException] = None,
log: bool = True, log: bool = True,
count_failure: bool = True, count_failure: bool = True,
) -> None: ) -> None:
@@ -1187,7 +1187,7 @@ class PluginManager:
""" """
failure_time = time.time() failure_time = time.time()
if exc is not None: if exc is not None:
err: Exception = exc err: BaseException = exc
error_type = type(exc).__name__ error_type = type(exc).__name__
else: else:
err = Exception(f"Plugin {plugin_id} execution failed (timeout or executor error)") err = Exception(f"Plugin {plugin_id} execution failed (timeout or executor error)")
@@ -1653,7 +1653,7 @@ class PluginManager:
finish_guard = threading.Lock() finish_guard = threading.Lock()
finished = {'done': False} finished = {'done': False}
def _finish(success: bool, exc: Optional[Exception] = None) -> None: def _finish(success: bool, exc: Optional[BaseException] = None) -> None:
with finish_guard: with finish_guard:
if finished['done']: if finished['done']:
return return
@@ -1727,7 +1727,13 @@ class PluginManager:
self.resource_monitor.monitor_call(plugin_id, plugin_instance.update) self.resource_monitor.monitor_call(plugin_id, plugin_instance.update)
else: else:
plugin_instance.update() plugin_instance.update()
except Exception as exc: except BaseException as exc: # pylint: disable=broad-except
# BaseException, not just Exception: asyncio.CancelledError
# and SystemExit derive from it. Either one skipped _finish,
# so the plugin kept its lock and stayed RUNNING for good --
# never rescheduled, and every display() skipped as busy.
# Re-raised for the executor, which reports it as this
# update's failure.
_finish(False, exc=exc) _finish(False, exc=exc)
raise raise
else: else:
+118 -10
View File
@@ -36,6 +36,13 @@ tmpfs. A missing heartbeat (dev server, emulator, Windows, a display still
starting up) or one from another process (a display restarted after a starting up) or one from another process (a display restarted after a
watchdog kill) says nothing, and the snapshot is judged on its own. watchdog kill) says nothing, and the snapshot is judged on its own.
The control socket. Where the display serves its state stream (stage 3,
docs/IPC_CONTROL_SOCKET.md), every tick also hands the snapshot to it, in
memory, and the web interface reads it there first
(``view_from_socket_state``, judged by the same rules). While the socket
serves those readers, the cache copy is their fallback and an unchanged
snapshot is rewritten every ``RELAXED_REFRESH_INTERVAL`` instead.
A dead publisher. systemd removes the heartbeat's directory when the A dead publisher. systemd removes the heartbeat's directory when the
service stops, so after a watchdog kill there is no heartbeat to go stale. service stops, so after a watchdog kill there is no heartbeat to go stale.
The reader then asks whether the snapshot's ``pid`` still exists (POSIX The reader then asks whether the snapshot's ``pid`` still exists (POSIX
@@ -48,7 +55,7 @@ import math
import os import os
import threading import threading
import time import time
from dataclasses import dataclass, field from dataclasses import dataclass, field, replace
from typing import Any, Callable, Dict, Optional from typing import Any, Callable, Dict, Optional
from src import display_watchdog from src import display_watchdog
@@ -75,6 +82,15 @@ TICK_INTERVAL = 5.0
#: A snapshot older than this is stale: three missed refreshes. #: A snapshot older than this is stale: three missed refreshes.
STALE_AFTER = 3 * REFRESH_INTERVAL STALE_AFTER = 3 * REFRESH_INTERVAL
#: The refresh while the control socket serves the web interface's readers
#: (``StateHub.readers_active``). The cache copy is then only their fallback,
#: so an unchanged snapshot is rewritten half as often; the snapshot says so
#: in its own ``refresh_interval`` and ``stale_after``.
RELAXED_REFRESH_INTERVAL = 2 * REFRESH_INTERVAL
#: The control socket's section for this snapshot (``state.plugins``).
STATE_SECTION = "plugins"
#: Bounds on a published ``stale_after``, so a corrupt value can make a #: Bounds on a published ``stale_after``, so a corrupt value can make a
#: reader neither trust a dead display for hours nor distrust a live one. #: reader neither trust a dead display for hours nor distrust a live one.
_STALE_AFTER_MIN = 30.0 _STALE_AFTER_MIN = 30.0
@@ -140,7 +156,8 @@ def summarize_error(error_info: Optional[Dict[str, Any]]) -> Optional[Dict[str,
def build_runtime_snapshot(state_manager: Any, *, started_at: float, def build_runtime_snapshot(state_manager: Any, *, started_at: float,
now: Optional[float] = None, now: Optional[float] = None,
running: bool = True) -> Dict[str, Any]: running: bool = True,
refresh_interval: float = REFRESH_INTERVAL) -> Dict[str, Any]:
"""The snapshot for ``state_manager`` (a plugin_state.PluginStateManager). """The snapshot for ``state_manager`` (a plugin_state.PluginStateManager).
A stopped snapshot (``running=False``) lists no plugins: nothing is A stopped snapshot (``running=False``) lists no plugins: nothing is
@@ -162,8 +179,8 @@ def build_runtime_snapshot(state_manager: Any, *, started_at: float,
"running": running, "running": running,
"published_at": time.time() if now is None else now, "published_at": time.time() if now is None else now,
"started_at": started_at, "started_at": started_at,
"refresh_interval": REFRESH_INTERVAL, "refresh_interval": refresh_interval,
"stale_after": STALE_AFTER, "stale_after": 3 * refresh_interval,
"pid": os.getpid(), "pid": os.getpid(),
"plugins": plugins, "plugins": plugins,
} }
@@ -197,30 +214,87 @@ class PluginRuntimePublisher:
self._tick_lock = threading.Lock() self._tick_lock = threading.Lock()
self._stop = threading.Event() self._stop = threading.Event()
self._thread: Optional[threading.Thread] = None self._thread: Optional[threading.Thread] = None
# The control socket's state stream (src/ipc/server.StateHub), when
# the display serves one: every tick also hands it the snapshot, in
# memory, and the cache refresh relaxes while it has readers.
self._hub: Any = None
self._hub_change: Optional[int] = None
self._hub_snapshot: Optional[Dict[str, Any]] = None
self.relaxed_refresh_interval = RELAXED_REFRESH_INTERVAL
def _write(self, running: bool) -> None: def attach_hub(self, hub: Any) -> None:
snapshot = build_runtime_snapshot(self.state_manager, started_at=self.started_at, """Also publish to the control socket's state hub, starting now."""
now=self._wall_clock(), running=running) with self._tick_lock:
self._hub = hub
self._hub_change = None
self._hub_snapshot = None
try:
self._push_to_hub(self.state_manager.change_count)
except Exception as err: # never let reporting break the display
logger.debug("Could not publish the plugin runtime state: %s", err,
exc_info=True)
def _push_to_hub(self, change: int) -> None:
"""The snapshot to the state hub: rebuilt when the state machine
changed, otherwise the last one with a new ``published_at``, which
the hub does not count as a new version. In memory, every tick, so
the socket's copy is never more than a tick old."""
hub = self._hub
if hub is None:
return
now = self._wall_clock()
if self._hub_snapshot is None or change != self._hub_change:
snapshot = build_runtime_snapshot(self.state_manager, started_at=self.started_at,
now=now)
else:
snapshot = dict(self._hub_snapshot, published_at=now)
hub.publish(STATE_SECTION, snapshot, volatile=("published_at",))
self._hub_snapshot = snapshot
self._hub_change = change
def _cache_refresh_interval(self) -> float:
"""The cache refresh: relaxed while the socket serves the readers."""
hub = self._hub
try:
if hub is not None and hub.readers_active():
return self.relaxed_refresh_interval
except Exception: # pylint: disable=broad-except
# The normal interval is the safe answer: it only writes more.
logger.debug("State hub readers_active() failed; using the normal refresh", exc_info=True)
return self.refresh_interval
def _write(self, running: bool, refresh_interval: Optional[float] = None) -> None:
snapshot = build_runtime_snapshot(
self.state_manager, started_at=self.started_at, now=self._wall_clock(),
running=running,
refresh_interval=self.refresh_interval if refresh_interval is None
else refresh_interval)
self.cache_manager.set(PLUGIN_RUNTIME_KEY, snapshot) self.cache_manager.set(PLUGIN_RUNTIME_KEY, snapshot)
def tick(self) -> bool: def tick(self) -> bool:
"""Publish if something changed (throttled) or the refresh is due. """Publish if something changed (throttled) or the refresh is due.
True if a snapshot was written.""" True if a snapshot was written to the cache."""
with self._tick_lock: with self._tick_lock:
try: try:
change = self.state_manager.change_count change = self.state_manager.change_count
try:
self._push_to_hub(change)
except Exception as err: # the cache copy still goes out below
logger.debug("Could not publish the plugin runtime state: %s", err,
exc_info=True)
now = self._clock() now = self._clock()
refresh = self._cache_refresh_interval()
since = None if self._last_attempt is None else now - self._last_attempt since = None if self._last_attempt is None else now - self._last_attempt
if since is not None: if since is not None:
if change == self._published_change: if change == self._published_change:
if since < self.refresh_interval: if since < refresh:
return False return False
elif since < self.min_interval: elif since < self.min_interval:
return False return False
# Stamp the attempt before writing: a cache that keeps failing # Stamp the attempt before writing: a cache that keeps failing
# is retried at the throttled rate, not on every tick. # is retried at the throttled rate, not on every tick.
self._last_attempt = now self._last_attempt = now
self._write(running=True) self._write(running=True, refresh_interval=refresh)
self._published_change = change self._published_change = change
return True return True
except Exception as err: # never let reporting break the display except Exception as err: # never let reporting break the display
@@ -317,6 +391,9 @@ class PluginRuntimeView:
plugins: Dict[str, Dict[str, Any]] = field(default_factory=dict) plugins: Dict[str, Dict[str, Any]] = field(default_factory=dict)
#: Age of the render loop's heartbeat, when it was taken into account. #: Age of the render loop's heartbeat, when it was taken into account.
heartbeat_age_seconds: Optional[float] = None heartbeat_age_seconds: Optional[float] = None
#: Where the snapshot came from: ``cache`` (the shared cache file and the
#: heartbeat file) or ``socket`` (the control socket's state stream).
source: str = "cache"
@property @property
def live(self) -> bool: def live(self) -> bool:
@@ -348,6 +425,7 @@ class PluginRuntimeView:
"stale_after": self.stale_after, "stale_after": self.stale_after,
"heartbeat_age_seconds": (None if self.heartbeat_age_seconds is None "heartbeat_age_seconds": (None if self.heartbeat_age_seconds is None
else round(self.heartbeat_age_seconds, 1)), else round(self.heartbeat_age_seconds, 1)),
"source": self.source,
} }
@@ -445,6 +523,36 @@ def view_from_snapshot(snapshot: Any, now: Optional[float] = None,
) )
def view_from_socket_state(snapshot: Any, now: Optional[float] = None,
now_mono: Optional[float] = None) -> Optional[PluginRuntimeView]:
"""Judge the ``plugins`` section of a control-socket state snapshot by
the same rules as the cache copy; None when it has none (an older
display, or a snapshot too large to carry it), so the caller reads the
cache instead.
The display measured its render loop's heartbeat age when it answered
(``state.loop``); that is the heartbeat here, aged by the time since the
answer arrived. A live snapshot with a stalled loop is ``stalled``, and a
snapshot older than its ``stale_after`` (the publisher thread stopped)
is ``stale``, exactly as for the cache. The display answered, so its
process is alive: there is no pid check.
"""
from src.ipc.client import snapshot_loop_age # stdlib-only module
if not isinstance(snapshot, dict):
return None
state = snapshot.get("state")
plugins = state.get(STATE_SECTION) if isinstance(state, dict) else None
if not isinstance(plugins, dict):
return None
now_mono = time.monotonic() if now_mono is None else now_mono
beat_age = snapshot_loop_age(snapshot, now_mono=now_mono)
heartbeat = None
if beat_age is not None:
heartbeat = {"pid": plugins.get("pid"), "mono": now_mono - beat_age}
view = view_from_snapshot(plugins, now=now, heartbeat=heartbeat, now_mono=now_mono)
return replace(view, source="socket")
def read_plugin_runtime(cache_manager: Any, now: Optional[float] = None, def read_plugin_runtime(cache_manager: Any, now: Optional[float] = None,
heartbeat_path: Optional[str] = None) -> PluginRuntimeView: heartbeat_path: Optional[str] = None) -> PluginRuntimeView:
"""The display's latest snapshot, judged for staleness and against the """The display's latest snapshot, judged for staleness and against the
+109 -36
View File
@@ -8,7 +8,7 @@ Provides resource limits and performance monitoring.
import math import math
import time import time
import threading import threading
from typing import Dict, Optional, Any, Callable, cast from typing import Dict, Optional, Any, Callable, Set, cast
from dataclasses import dataclass, field, fields from dataclasses import dataclass, field, fields
from src.logging_config import get_logger from src.logging_config import get_logger
@@ -99,18 +99,33 @@ class ResourceMetrics:
last_update_time: float = field(default_factory=time.time) last_update_time: float = field(default_factory=time.time)
#: How often a plugin's metrics are written to the cache, in seconds. #: How often the metrics snapshot is written to the cache, in seconds.
#: #:
#: Persisting on every call meant a small file rewritten roughly nine times a #: Persisting on every call meant a small file rewritten roughly nine times a
#: minute per plugin. On a rig with fourteen active plugins that was ~126 #: minute per plugin. On a rig with fourteen active plugins that was ~126
#: writes a minute for metrics alone, and since each ~350-byte file costs a #: writes a minute for metrics alone, and since each ~350-byte file costs a
#: 4KB block plus an ext4 journal entry, it dominated the device's write #: 4KB block plus an ext4 journal entry, it dominated the device's write
#: volume -- on an SD card, which wears out. #: volume -- on an SD card, which wears out. Throttling each plugin's own
#: record to once per 30 s still left two writes a minute per plugin, so all
#: plugins now share one record (METRICS_SNAPSHOT_KEY), written at most once
#: a minute: one write a minute however many plugins there are.
#: #:
#: The in-memory copy stays authoritative and exact; only the cross-process #: The in-memory copy stays authoritative and exact; only the cross-process
#: snapshot the web UI reads is delayed, and telemetry up to half a minute old #: snapshot the web UI reads is delayed, and telemetry up to a minute old is
#: is still a fair description of a long-running plugin. #: still a fair description of a long-running plugin.
_METRICS_PERSIST_INTERVAL = 30.0 _METRICS_PERSIST_INTERVAL = 60.0
#: The one cache record holding every plugin's metrics:
#: ``{"schema": 1, "plugins": {plugin_id: <metrics record>}}``, each metrics
#: record shaped as the per-plugin ``plugin_metrics:<id>`` records were. Those
#: older records are still read for a plugin the snapshot does not have yet
#: (an upgrade, or a plugin that has not run since), never written.
METRICS_SNAPSHOT_KEY = "plugin_metrics_snapshot"
_METRICS_SNAPSHOT_SCHEMA = 1
#: A plugin with no call for this long is dropped from the snapshot -- what
#: the cache's 30-day default retention did to its own record before.
_METRICS_SNAPSHOT_ENTRY_MAX_AGE = 30 * 86400
class PluginResourceMonitor: class PluginResourceMonitor:
@@ -140,10 +155,15 @@ class PluginResourceMonitor:
self._metrics: Dict[str, ResourceMetrics] = {} self._metrics: Dict[str, ResourceMetrics] = {}
self._limits: Dict[str, ResourceLimits] = {} self._limits: Dict[str, ResourceLimits] = {}
self._bad_limits_warned: set = set() self._bad_limits_warned: set = set()
# When each plugin's metrics last reached the cache. Metrics change on # When the metrics snapshot last reached the cache (monotonic), None
# every call, so they cannot be de-duplicated the way health state can; # until it has. Metrics change on every call, so they cannot be
# they are rate-limited instead. See _METRICS_PERSIST_INTERVAL. # de-duplicated the way health state can; they are rate-limited
self._metrics_persisted_at: Dict[str, float] = {} # instead. See _METRICS_PERSIST_INTERVAL.
self._snapshot_persisted_at: Optional[float] = None
# Plugins whose metrics this process recorded since the last snapshot
# write: only their entries are overwritten, the rest are kept as
# found on disk.
self._metrics_dirty: Set[str] = set()
# Lock for thread-safe access # Lock for thread-safe access
self._lock = threading.Lock() self._lock = threading.Lock()
@@ -247,10 +267,14 @@ class PluginResourceMonitor:
with self._lock: with self._lock:
if force_reload or plugin_id not in self._metrics: if force_reload or plugin_id not in self._metrics:
# Try to load from cache # Try to load from cache
cache_key = self._get_metrics_key(plugin_id) memory_ttl = 0 if force_reload else None
cached = self.cache_manager.get( cached = self._read_snapshot(memory_ttl).get(plugin_id)
cache_key, max_age=None, memory_ttl=0 if force_reload else None if cached is None:
) # Not in the snapshot: the per-plugin record an older
# version wrote, if there is one.
cached = self.cache_manager.get(
self._get_metrics_key(plugin_id), max_age=None,
memory_ttl=memory_ttl)
if cached: if cached:
metrics = self._metrics_from_cache(plugin_id, cached) metrics = self._metrics_from_cache(plugin_id, cached)
else: else:
@@ -498,12 +522,70 @@ class PluginResourceMonitor:
summaries[plugin_id] = self.get_metrics_summary(plugin_id) summaries[plugin_id] = self.get_metrics_summary(plugin_id)
return summaries return summaries
def _persist_metrics(self, plugin_id: str, metrics: ResourceMetrics, def _read_snapshot(self, memory_ttl: Optional[int] = None) -> Dict[str, Any]:
force: bool = False) -> None: """The snapshot's per-plugin records, or {} if there is none usable.
"""Write a plugin's metrics to the cache, at most once per interval.
Caller must hold ``self._lock``. Caller must hold ``self._lock``.
""" """
cached = self.cache_manager.get(
METRICS_SNAPSHOT_KEY, max_age=None, memory_ttl=memory_ttl)
if not isinstance(cached, dict) or cached.get('schema') != _METRICS_SNAPSHOT_SCHEMA:
return {}
plugins = cached.get('plugins')
return plugins if isinstance(plugins, dict) else {}
@staticmethod
def _metrics_record(metrics: ResourceMetrics) -> Dict[str, Any]:
"""One plugin's entry in the snapshot."""
return {
'memory_mb': metrics.memory_mb,
'cpu_percent': metrics.cpu_percent,
'execution_time': metrics.execution_time,
'call_count': metrics.call_count,
'total_execution_time': metrics.total_execution_time,
'max_execution_time': metrics.max_execution_time,
'min_execution_time': (metrics.min_execution_time
if metrics.min_execution_time != float('inf')
else 0.0),
'last_update_time': metrics.last_update_time,
}
def _write_snapshot(self, drop: Optional[str] = None) -> None:
"""Write the snapshot: what is on disk, with this process's recorded
plugins updated and ``drop`` removed.
Starting from the disk copy rather than from memory keeps the entries
of plugins this process has not run -- disabled ones, which the web UI
still shows -- and a reset made from the other process.
Caller must hold ``self._lock``.
"""
plugins = dict(self._read_snapshot(memory_ttl=0))
if drop is not None:
plugins.pop(drop, None)
for plugin_id in self._metrics_dirty:
if plugin_id in self._metrics:
plugins[plugin_id] = self._metrics_record(self._metrics[plugin_id])
cutoff = time.time() - _METRICS_SNAPSHOT_ENTRY_MAX_AGE
for plugin_id, record in list(plugins.items()):
last = record.get('last_update_time') if isinstance(record, dict) else None
if isinstance(last, (int, float)) and last < cutoff:
del plugins[plugin_id]
self.cache_manager.set(METRICS_SNAPSHOT_KEY, {
'schema': _METRICS_SNAPSHOT_SCHEMA,
'plugins': plugins,
})
# Only once the write has landed, so a failed one is retried in full.
self._metrics_dirty.clear()
def _persist_metrics(self, plugin_id: str, metrics: ResourceMetrics,
force: bool = False) -> None:
"""Record that a plugin's metrics changed, and write the snapshot if
the last write is at least an interval old.
Caller must hold ``self._lock``.
"""
self._metrics_dirty.add(plugin_id)
# Monotonic, not wall clock: these devices have no RTC, so the clock # Monotonic, not wall clock: these devices have no RTC, so the clock
# jumps by however far off boot-time was the moment NTP first syncs. # jumps by however far off boot-time was the moment NTP first syncs.
# A forward jump would allow an early write, a backward one would # A forward jump would allow an early write, a backward one would
@@ -515,35 +597,26 @@ class PluginResourceMonitor:
# single run -- the throttle swallowed the very first snapshot, which # single run -- the throttle swallowed the very first snapshot, which
# is the one that matters most after a restart. # is the one that matters most after a restart.
now = time.monotonic() now = time.monotonic()
last_written = self._metrics_persisted_at.get(plugin_id) last_written = self._snapshot_persisted_at
if (not force and last_written is not None if (not force and last_written is not None
and now - last_written < _METRICS_PERSIST_INTERVAL): and now - last_written < _METRICS_PERSIST_INTERVAL):
return return
cache_key = self._get_metrics_key(plugin_id) self._write_snapshot()
self.cache_manager.set(cache_key, {
'memory_mb': metrics.memory_mb,
'cpu_percent': metrics.cpu_percent,
'execution_time': metrics.execution_time,
'call_count': metrics.call_count,
'total_execution_time': metrics.total_execution_time,
'max_execution_time': metrics.max_execution_time,
'min_execution_time': (metrics.min_execution_time
if metrics.min_execution_time != float('inf')
else 0.0),
'last_update_time': metrics.last_update_time,
})
# Only after the write lands. Marking it first would mean a failed # Only after the write lands. Marking it first would mean a failed
# set() bought the next interval's silence without leaving a snapshot. # set() bought the next interval's silence without leaving a snapshot.
self._metrics_persisted_at[plugin_id] = now self._snapshot_persisted_at = now
def reset_metrics(self, plugin_id: str) -> None: def reset_metrics(self, plugin_id: str) -> None:
"""Reset metrics for a plugin.""" """Reset metrics for a plugin."""
with self._lock: with self._lock:
if plugin_id in self._metrics: if plugin_id in self._metrics:
self._metrics[plugin_id] = ResourceMetrics() self._metrics[plugin_id] = ResourceMetrics()
cache_key = self._get_metrics_key(plugin_id) self._metrics_dirty.discard(plugin_id)
self.cache_manager.delete(cache_key) self._write_snapshot(drop=plugin_id)
# The record an older version wrote, so the reader's fallback
# cannot bring the old numbers back.
self.cache_manager.delete(self._get_metrics_key(plugin_id))
# Let the next call persist immediately rather than leaving the # Let the next call persist immediately rather than leaving the
# deleted key absent for the rest of the interval. # plugin absent from the snapshot for the rest of the interval.
self._metrics_persisted_at.pop(plugin_id, None) self._snapshot_persisted_at = None
+14
View File
@@ -344,12 +344,26 @@ class PluginStoreManager(_RegistryMixin, _InstallMixin, _UpdateMixin):
2. Fix permissions via os.chmod() then retry (works for same-owner files) 2. Fix permissions via os.chmod() then retry (works for same-owner files)
3. Use sudo rm -rf as last resort (works for root-owned __pycache__, etc.) 3. Use sudo rm -rf as last resort (works for root-owned __pycache__, etc.)
A symlink -- a dev plugin linked in by scripts/dev/dev_plugin_setup.sh
-- is removed as a link, before any of that: rmtree refuses one, and
stage 2 would walk through it and chmod the developer's checkout.
Args: Args:
path: Path to directory to remove path: Path to directory to remove
Returns: Returns:
True if directory was removed successfully, False otherwise True if directory was removed successfully, False otherwise
""" """
if path.is_symlink():
# Checked before exists(), which follows the link: a dangling one
# would read as already removed and be left behind.
try:
path.unlink()
return True
except OSError as e:
self.logger.error(f"Could not remove the symlink {path}: {e}")
return False
if not path.exists(): if not path.exists():
return True # Already removed return True # Already removed
+11 -8
View File
@@ -95,11 +95,13 @@ class StartupValidator:
def _validate_systemd_units(self) -> None: def _validate_systemd_units(self) -> None:
"""Warn when an installed unit has drifted from the repo's template. """Warn when an installed unit has drifted from the repo's template.
Nothing re-applies these after the first install. `git pull` -- which is Before updates refreshed units, nothing re-applied these after the
what the web UI's update button runs -- brings a new template into the first install: `git pull` brought a new template into the checkout,
checkout, but nothing copies it to /etc/systemd/system and nothing runs but nothing copied it to /etc/systemd/system, so the unit that
`systemctl daemon-reload`, so the unit that actually runs is whatever actually ran was whatever first_time_install.sh wrote on day one.
first_time_install.sh wrote on day one. Updates now install changed units through the root helper
ledmatrix-refresh-units (web_interface/unit_refresh.py) -- but only on
a device whose installer granted it, so this still catches the rest.
That makes every hardening added to a unit inert on existing installs. That makes every hardening added to a unit inert on existing installs.
Measured on one rig: the installed unit was thirteen days older than the Measured on one rig: the installed unit was thirteen days older than the
@@ -141,10 +143,11 @@ class StartupValidator:
if self._unit_body(expected) != self._unit_body(actual): if self._unit_body(expected) != self._unit_body(actual):
self.warnings.append( self.warnings.append(
f"{installed.name} differs from {template_rel}; the " f"{installed.name} differs from {template_rel}, so "
"installed unit is not refreshed by an update, so "
"settings added to the template are not in effect. " "settings added to the template are not in effect. "
"Re-run scripts/install/install_service.sh to apply them." "Updates apply them only once the installer has granted "
"ledmatrix-refresh-units: re-run "
"scripts/install/install_service.sh (or first_time_install.sh) to apply them."
) )
except OSError as e: except OSError as e:
self.logger.debug("Could not compare systemd units: %s", e) self.logger.debug("Could not compare systemd units: %s", e)
+5
View File
@@ -19,6 +19,11 @@ Type=simple
User=__USER__ User=__USER__
WorkingDirectory=__PROJECT_ROOT_DIR__ WorkingDirectory=__PROJECT_ROOT_DIR__
Environment=USE_THREADING=1 Environment=USE_THREADING=1
# Cap glibc's malloc arenas, as ledmatrix.service does: each allocating thread
# can get its own arena, up to 8 x CPU count (24 on a 3-core Pi), and a grown
# arena is never handed back to the OS. This threaded Flask process would hold
# that memory the same way. See ledmatrix.service for the measurement.
Environment=MALLOC_ARENA_MAX=2
ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/scripts/utils/start_web_conditionally.py ExecStart=/usr/bin/python3 __PROJECT_ROOT_DIR__/scripts/utils/start_web_conditionally.py
Restart=on-failure Restart=on-failure
RestartSec=10 RestartSec=10
+7 -2
View File
@@ -724,11 +724,14 @@ class RunLoopHarness:
self.clock.at(t, post) self.clock.at(t, post)
def restore_on_demand(self, plugin_id: str, mode: Optional[str] = None, def restore_on_demand(self, plugin_id: str, mode: Optional[str] = None,
duration: Optional[float] = None, pinned: bool = False): duration: Optional[float] = None, pinned: bool = False,
named_mode: Optional[str] = None):
"""Start with an on-demand session resumed from the cache, as after """Start with an on-demand session resumed from the cache, as after
a restart: the state _select_startup_plugins restores, then a restart: the state _select_startup_plugins restores, then
_populate_on_demand_modes_from_plugin, as __init__ calls it.""" _populate_on_demand_modes_from_plugin, as __init__ calls it. A
session that cannot resume is logged as ``on-demand-error``."""
dc = self.controller dc = self.controller
dc._on_demand_named_mode = named_mode
dc.on_demand_active = True dc.on_demand_active = True
dc.on_demand_plugin_id = plugin_id dc.on_demand_plugin_id = plugin_id
dc.on_demand_mode = mode dc.on_demand_mode = mode
@@ -739,6 +742,8 @@ class RunLoopHarness:
dc.on_demand_status = 'active' dc.on_demand_status = 'active'
dc.on_demand_schedule_override = True dc.on_demand_schedule_override = True
dc._populate_on_demand_modes_from_plugin() dc._populate_on_demand_modes_from_plugin()
if dc.on_demand_status == 'error':
self.log("on-demand-error", dc.on_demand_last_error)
def wifi_message(self, t: float, message: str, duration: float = 5): def wifi_message(self, t: float, message: str, duration: float = 5):
def write(): def write():
+15
View File
@@ -328,6 +328,21 @@ def _hermetic_control_socket(monkeypatch):
monkeypatch.setenv(SOCKET_PATH_ENV, 'off') monkeypatch.setenv(SOCKET_PATH_ENV, 'off')
@pytest.fixture(autouse=True)
def _hermetic_unit_refresh(monkeypatch, tmp_path_factory):
"""Keep updates' systemd unit refresh off the host.
perform_core_update runs web_interface/unit_refresh.py after any update
that moves HEAD, and several tests run the real one against a test clone.
On a device -- or a machine where install_service.sh was tried out -- it
would compare the clone's templates with the real /etc/systemd/system and
run the real sudo helper. Point it at a folder that does not exist: no units
installed, nothing to do. The unit refresh tests pass their own.
"""
from web_interface import unit_refresh
monkeypatch.setattr(unit_refresh, 'SYSTEMD_DIR', str(tmp_path_factory.getbasetemp() / 'no-systemd'))
@pytest.fixture(autouse=True) @pytest.fixture(autouse=True)
def reset_logging(): def reset_logging():
"""Reset logging configuration before each test.""" """Reset logging configuration before each test."""
+9
View File
@@ -175,6 +175,15 @@
"POST" "POST"
] ]
], ],
[
"/api/v3/config/refresh-rate",
"api_v3.get_refresh_rate",
[
"GET",
"HEAD",
"OPTIONS"
]
],
[ [
"/api/v3/config/schedule", "/api/v3/config/schedule",
"api_v3.get_schedule_config", "api_v3.get_schedule_config",
+29
View File
@@ -0,0 +1,29 @@
{
"screens": [
[0.0, "clock", 5.0, "on-demand-start", 6, false],
[5.0, "sports_live", 15.0, "duration", 15, true],
[20.0, "sports_recent", 15.0, "duration", 15, true],
[35.0, "sports_upcoming", 5.0, "on-demand-requested-stop", 6, true],
[40.0, "clock", 20.0, "duration", 20, true],
[60.0, "sports_live", 15.0, "display-false", 11, true],
[75.0, "sports_recent", 15.0, "duration", 15, true],
[90.0, "sports_upcoming", 10.0, "on-demand-start", 11, true],
[100.0, "sports_live", 0.0, "empty", 1, true],
[100.0, "sports_recent", 15.0, "duration", 15, true],
[115.0, "sports_upcoming", 15.0, "duration", 15, true],
[130.0, "sports_live", 0.0, "empty", 1, true],
[130.0, "sports_recent", 10.0, "on-demand-requested-stop", 11, true],
[140.0, "sports_upcoming", 15.0, "duration", 15, true],
[155.0, "clock", 5.0, "horizon", 5, true]
],
"events": [
[5.0, "request", "start:n1"],
[5.0, "on-demand-start", "sports"],
[40.0, "request", "stop:n2"],
[40.0, "on-demand-requested-stop"],
[100.0, "request", "start:n3"],
[100.0, "on-demand-start", "sports"],
[140.0, "request", "stop:n4"],
[140.0, "on-demand-requested-stop"]
]
}
@@ -0,0 +1,10 @@
{
"screens": [
[0.0, "clock", 20.0, "duration", 20, false],
[20.0, "weather", 20.0, "duration", 20, true],
[40.0, "clock", 20.0, "horizon", 20, true]
],
"events": [
[0.0, "on-demand-error", "restore-failed"]
]
}
+8
View File
@@ -45,7 +45,12 @@ server has none.
|---|---|---| |---|---|---|
| `unit/test_list_filter.js` | no | `ListFilter` search/filter/sort/count/sticky, and the installed-plugins config **extracted verbatim** from `plugins_manager.js` so the test can't drift from it | | `unit/test_list_filter.js` | no | `ListFilter` search/filter/sort/count/sticky, and the installed-plugins config **extracted verbatim** from `plugins_manager.js` so the test can't drift from it |
| `unit/test_update_all.js` | no | `PluginInstallManager.updateAll` from `plugins/install_manager.js`: Check & Update All sends only plugin ids (never `starlark:` app entries), re-sends a request that got no HTTP answer (web service restarting) instead of skipping that plugin, never re-sends one that got any HTTP answer (the real `api_client.js` classifies a proxy 502 or a JSON error without `error_code` as `API_ERROR`), and counts a no-op update as already up to date in the summary. Also run by `test/web_interface/test_update_all_plugins.py` so CI covers it | | `unit/test_update_all.js` | no | `PluginInstallManager.updateAll` from `plugins/install_manager.js`: Check & Update All sends only plugin ids (never `starlark:` app entries), re-sends a request that got no HTTP answer (web service restarting) instead of skipping that plugin, never re-sends one that got any HTTP answer (the real `api_client.js` classifies a proxy 502 or a JSON error without `error_code` as `API_ERROR`), and counts a no-op update as already up to date in the summary. Also run by `test/web_interface/test_update_all_plugins.py` so CI covers it |
| `unit/test_store_install.js` | no | The store's Install button, with the whole of `plugins_manager.js` run by `plugins_manager_sandbox.js` (a vm context, fake DOM and API): a fresh install reloads the list, then enables the id the plugin was installed as -- the answer's `plugin_id`, else the installed entry the store entry matches (Weather installs as `ledmatrix-weather`); a Reinstall leaves the enabled state alone |
| `unit/test_install_polling.js` | no | How long Install waits for a queued install (sandbox): at least the server's 300 s dependency-install timeout; when it stops waiting it reloads the installed list and warns, rather than reporting a failure or enabling anything |
| `unit/test_store_categories.js` | no | The store's category filter (sandbox): the template ships only All Categories, the rest come from the store's plugins (one per category whatever its case), choosing one filters to it, and a swapped-in select is refilled from the cache keeping the choice |
| `unit/test_github_url_install.js` | no | Install Single Plugin (sandbox, the button as `plugins.html` ships it): no inline `onclick`, so a click or Enter sends exactly one `install-from-url` request and raises no error |
| `unit/test_render_cards.js` | no | `renderInstalledCards` markup, both empty states, and HTML-escaping of hostile plugin metadata | | `unit/test_render_cards.js` | no | `renderInstalledCards` markup, both empty states, and HTML-escaping of hostile plugin metadata |
| `unit/test_plugin_order_list.js` | no | `widgets/plugin-order-list.js` (the Vegas and rotation order lists): a disabled plugin, which gets no row, keeps its slot in the saved order and its Vegas exclusion when the list rewrites its hidden inputs, around reordering and include/exclude; an uninstalled plugin's id is dropped, a failed plugin list leaves the inputs as saved, and only string ids are carried over, once each |
| `unit/test_style_editor_element_keys.js` | no | `elementKeys()`/`styleRows()`/`positionRows()` from `widgets/style-editor.js`: every `customization.layout` entry gets exactly one row -- paired with its style element through core's `x-layout-key` (so `score` belongs to `score_text`, not a second row), or a position row of its own, leaves included -- since the widget claims the whole `layout` block from the generic fallback renderer | | `unit/test_style_editor_element_keys.js` | no | `elementKeys()`/`styleRows()`/`positionRows()` from `widgets/style-editor.js`: every `customization.layout` entry gets exactly one row -- paired with its style element through core's `x-layout-key` (so `score` belongs to `score_text`, not a second row), or a position row of its own, leaves included -- since the widget claims the whole `layout` block from the generic fallback renderer |
| `unit/test_style_editor_layout_leaf_columns.js` | no | `columnsFor()` from `widgets/style-editor.js`: a layout-only key whose own value is a leaf (no x/y sub-object, e.g. a `show_logo` toggle) gets a self-keyed column instead of a blank, uneditable row | | `unit/test_style_editor_layout_leaf_columns.js` | no | `columnsFor()` from `widgets/style-editor.js`: a layout-only key whose own value is a leaf (no x/y sub-object, e.g. a `show_logo` toggle) gets a self-keyed column instead of a blank, uneditable row |
| `unit/test_style_editor_layout_leaf_collision.js` | no | `columnsFor()` from `widgets/style-editor.js`: a layout-only leaf key still gets its own column even when its name collides with an unrelated element's style sub-field or another layout axis's sub-field | | `unit/test_style_editor_layout_leaf_collision.js` | no | `columnsFor()` from `widgets/style-editor.js`: a layout-only leaf key still gets its own column even when its name collides with an unrelated element's style sub-field or another layout axis's sub-field |
@@ -53,6 +58,9 @@ server has none.
| `unit/test_store_registry_fields.js` | no | The store card's registry fields from `plugins_manager.js`: the commit that introduced the listed version (a hex SHA only, linked to that tree), the "Needs LEDMatrix X+" warning, a card from an older registry without either, and `isStorePluginInstalled` answering to `aliases` | | `unit/test_store_registry_fields.js` | no | The store card's registry fields from `plugins_manager.js`: the commit that introduced the listed version (a hex SHA only, linked to that tree), the "Needs LEDMatrix X+" warning, a card from an older registry without either, and `isStorePluginInstalled` answering to `aliases` |
| `unit/test_page_registry.js` | no | The page lifecycle in `js/core/registry.js` (a minimal DOM shim): one `init` per `data-page` root, `destroy` and an aborted `ctx.signal` when htmx swaps it away, a vetoed swap keeps it, lazy page modules, a root removed without htmx swept on the next swap | | `unit/test_page_registry.js` | no | The page lifecycle in `js/core/registry.js` (a minimal DOM shim): one `init` per `data-page` root, `destroy` and an aborted `ctx.signal` when htmx swaps it away, a vetoed swap keeps it, lazy page modules, a root removed without htmx swept on the next swap |
| `unit/test_core_modules.js` | no | `js/core/api.js` (JSON envelope, HTTP/`status: error`/network errors, abort passthrough, the #683 login redirect, same-server paths only) and `js/core/facade.js` (`window.LEDMatrix`, deprecated aliases) | | `unit/test_core_modules.js` | no | `js/core/api.js` (JSON envelope, HTTP/`status: error`/network errors, abort passthrough, the #683 login redirect, same-server paths only) and `js/core/facade.js` (`window.LEDMatrix`, deprecated aliases) |
| `unit/test_overview_reconciliation_poll.js` | no | The Overview's reconciliation-banner poll from `partials/overview.html`, run in a vm: it gives up after a bounded number of requests when the status never says done, runs only while the Overview is on screen (`LEDVisibility`, its own key), and stops once the banner is shown |
| `unit/test_display_partial_ids.js` | no | `partials/display.html`: every literal `getElementById()` in its inline scripts names an id the partial renders, and moving the brightness slider (the shipped script, in a vm with a fake DOM) updates its label without throwing |
| `unit/test_general_web_login_token.js` | no | `window.webLogin.createToken` from `partials/general.html`, run in a vm: a created API token clears the form's `data-dirty` mark (so a reload does not ask "Leave site?"), a refused one keeps it |
| `unit/test_plugin_action_delegation.js` | no | The document-level card-action delegation and `handlePluginAction` from `plugins_manager.js`, run with the handler inside an IIFE as in the real file: each action is handled once, a Starlark app uninstall goes to `DELETE /starlark/apps/<id>`, and an uninstall is confirmed once | | `unit/test_plugin_action_delegation.js` | no | The document-level card-action delegation and `handlePluginAction` from `plugins_manager.js`, run with the handler inside an IIFE as in the real file: each action is handled once, a Starlark app uninstall goes to `DELETE /starlark/apps/<id>`, and an uninstall is confirmed once |
| `dom/test_installed_dom.js` | yes | The toolbar in a real DOM: pill/search/sort interaction, the HTMX partial re-swap, and a `getComputedStyle` check that `.filter-pill[data-active]` really matches the emitted markup | | `dom/test_installed_dom.js` | yes | The toolbar in a real DOM: pill/search/sort interaction, the HTMX partial re-swap, and a `getComputedStyle` check that `.filter-pill[data-active]` really matches the emitted markup |
| `dom/test_store_dom.js` | yes | Store pagination, per-page, category, tri-state Installed button, and persistence across a re-boot, against the live registry | | `dom/test_store_dom.js` | yes | Store pagination, per-page, category, tri-state Installed button, and persistence across a re-boot, against the live registry |
+5 -1
View File
@@ -98,7 +98,11 @@ const ok = (l, c, x) => c ? (pass++, console.log(' ok ' + l))
const lists = () => requests.filter(r => r.url === '/api/v3/plugins/installed').length; const lists = () => requests.filter(r => r.url === '/api/v3/plugins/installed').length;
const $ = id => doc.getElementById(id); const $ = id => doc.getElementById(id);
const order = () => JSON.parse($('rotation_plugin_order_value').value || '[]'); // The rows' ids, in order. The input also keeps saved ids that have no row
// (a disabled plugin's place, see test/js/unit/test_plugin_order_list.js),
// and the saved order comes from whatever config the server has.
const SHOWN = plugins.filter(p => p.enabled).map(p => p.id);
const order = () => JSON.parse($('rotation_plugin_order_value').value || '[]').filter(id => SHOWN.includes(id));
async function swap() { async function swap() {
panel.dispatchEvent(new window.CustomEvent('htmx:beforeSwap', { bubbles: true, detail: { target: panel, shouldSwap: true } })); panel.dispatchEvent(new window.CustomEvent('htmx:beforeSwap', { bubbles: true, detail: { target: panel, shouldSwap: true } }));
panel.innerHTML = partial; panel.innerHTML = partial;
+42
View File
@@ -98,8 +98,50 @@ const get = p => new Promise((res, rej) =>
window.saveMqttBridge(); window.saveMqttBridge();
await tick(150); await tick(150);
ok('save includes password once typed', sent && sent.mqtt_password === 'typed-secret'); ok('save includes password once typed', sent && sent.mqtt_password === 'typed-secret');
// A password with TLS off is refused unless allow_insecure_mqtt is set
// (CWE-319, api_v3/misc.py). The form has to be able to send it, or a
// plain-LAN broker with a password can never be saved from here.
const allowRow = () => $('mqtt-allow-insecure-row');
const shown = el => !!el && !el.classList.contains('hidden');
ok('allow-without-TLS control rendered', !!$('mqtt-allow-insecure'));
ok('allow-without-TLS starts as saved',
!!$('mqtt-allow-insecure') && $('mqtt-allow-insecure').checked === !!bridge.data.config.allow_insecure_mqtt);
ok('allow-without-TLS shown only while TLS is off',
shown(allowRow()) === !$('mqtt-tls').checked);
$('mqtt-tls').checked = true;
$('mqtt-tls').dispatchEvent(new window.Event('change', { bubbles: true }));
ok('ticking TLS hides it', !shown(allowRow()));
$('mqtt-tls').checked = false;
$('mqtt-tls').dispatchEvent(new window.Event('change', { bubbles: true }));
ok('unticking TLS shows it again', shown(allowRow()));
const setAllow = v => { if ($('mqtt-allow-insecure')) $('mqtt-allow-insecure').checked = v; };
setAllow(false);
window.saveMqttBridge();
await tick(150);
ok('save sends allow_insecure_mqtt false when unticked', !!sent && sent.allow_insecure_mqtt === false, sent);
setAllow(true);
window.saveMqttBridge();
await tick(150);
ok('save sends allow_insecure_mqtt true when ticked', !!sent && sent.allow_insecure_mqtt === true, sent);
onPut = null; onPut = null;
// Prefilled from the saved settings, and hidden while TLS is saved on.
bridgePayload = JSON.parse(JSON.stringify(bridge));
bridgePayload.data.config.allow_insecure_mqtt = true;
bridgePayload.data.config.mqtt_tls = false;
window.loadMqttBridge();
await tick(150);
ok('a saved opt-in is prefilled', !!$('mqtt-allow-insecure') && $('mqtt-allow-insecure').checked === true);
bridgePayload.data.config.mqtt_tls = true;
window.loadMqttBridge();
await tick(150);
ok('hidden on load when TLS is saved on', !shown(allowRow()));
bridgePayload = bridge;
window.loadMqttBridge();
await tick(150);
// ── Pixlet editor, idle ──────────────────────────────────────────────── // ── Pixlet editor, idle ────────────────────────────────────────────────
const appIds = (apps.data.apps || []).map(a => a.id); const appIds = (apps.data.apps || []).map(a => a.id);
ok('editor lists the apps on disk', ok('editor lists the apps on disk',
+212
View File
@@ -0,0 +1,212 @@
// The whole of plugins_manager.js (and list_filter.js before it, as the page
// loads them), evaluated in a node vm context against a small fake DOM.
//
// For suites that drive the plugin manager's real flows -- install, polling,
// store filters, the GitHub-URL button -- rather than one function sliced
// out of the file. Nothing is mocked inside the script: only what the page
// gives it (document, fetch, timers, showNotification, LEDEscape).
//
// const sb = create({ route: (method, url, body) => ({ status, json }) });
// sb.el('plugin-store-grid'); // make an element exist by id
// sb.window.installPlugin('weather');
// await sb.until(() => sb.requests.some(r => r.url.includes('/toggle')));
//
// Timers ignore their delays and run on the next turn, so a poll loop that
// would take minutes in a browser finishes in milliseconds. The page is in
// readyState "loading" with no #installed-plugins-grid, so the script's own
// start-up does nothing until a suite asks for it (window.initPluginsPage()).
const fs = require('fs');
const path = require('path');
const vm = require('vm');
const ledEscape = require('./led_escape');
const V3 = path.resolve(__dirname, '../../web_interface/static/v3');
const PLUGINS_HTML = path.resolve(__dirname, '../../web_interface/templates/v3/partials/plugins.html');
class FakeClassList {
constructor() { this.set = new Set(); }
add(...c) { c.forEach(x => this.set.add(x)); }
remove(...c) { c.forEach(x => this.set.delete(x)); }
contains(c) { return this.set.has(c); }
toggle(c, force) {
const on = force === undefined ? !this.set.has(c) : !!force;
if (on) this.set.add(c); else this.set.delete(c);
return on;
}
}
function create({ route } = {}) {
const elements = new Map();
const requests = [];
const toasts = [];
const errors = [];
const restartNotes = [];
class FakeElement {
constructor(id, tag = 'div', attributes = {}) {
this.id = id;
this.tagName = tag.toUpperCase();
this.attributes = { ...attributes };
this.listeners = {};
this.children = [];
this.classList = new FakeClassList();
this.style = { removeProperty() {} };
this.dataset = {};
this.value = '';
this.textContent = '';
this.disabled = false;
this.parentNode = null;
this._html = '';
}
get innerHTML() { return this._html; }
set innerHTML(v) { this._html = String(v); this.children = []; }
getAttribute(n) { return n in this.attributes ? this.attributes[n] : null; }
setAttribute(n, v) { this.attributes[n] = String(v); }
hasAttribute(n) { return n in this.attributes; }
removeAttribute(n) { delete this.attributes[n]; }
addEventListener(type, fn) { (this.listeners[type] = this.listeners[type] || []).push(fn); }
removeEventListener(type, fn) {
this.listeners[type] = (this.listeners[type] || []).filter(f => f !== fn);
}
appendChild(child) { this.children.push(child); child.parentNode = this; return child; }
querySelector() { return null; }
querySelectorAll() { return []; }
closest() { return null; }
cloneNode() {
const copy = new FakeElement(this.id, this.tagName, this.attributes);
copy._html = this._html;
copy.value = this.value;
return copy;
}
replaceChild(next, prev) {
next.parentNode = this;
prev.parentNode = null;
if (next.id) elements.set(next.id, next);
return prev;
}
replaceWith(next) { if (this.parentNode) this.parentNode.replaceChild(next, this); }
// A browser runs an inline on<type> attribute first (it was set before
// any listener was added), then the listeners, and an exception in one
// does not stop the next: it is reported, which is what `errors` holds.
dispatch(type, init = {}) {
const event = {
type, target: this, currentTarget: this, key: init.key,
defaultPrevented: false,
preventDefault() { this.defaultPrevented = true; },
stopPropagation() {}, stopImmediatePropagation() {},
};
const inline = this.getAttribute('on' + type);
const handlers = [];
if (inline !== null) {
handlers.push(vm.runInContext(`(function(event) {\n${inline}\n})`, ctx));
}
handlers.push(...(this.listeners[type] || []));
for (const h of handlers) {
try { h.call(this, event); } catch (e) { errors.push(e); }
}
return event;
}
click() { return this.dispatch('click'); }
}
function el(id, tag, attributes) {
if (!elements.has(id)) {
const parent = new FakeElement(null);
parent.appendChild(new FakeElement(id, tag, attributes));
elements.set(id, parent.children[0]);
}
return elements.get(id);
}
const timers = [];
const ctx = {
// Warnings are the script noting elements this fake page doesn't have.
console: { log: console.log.bind(console), error: console.error.bind(console),
warn: () => {}, info: () => {}, debug: () => {} },
debugLog: () => {},
addEventListener() {},
URL,
document: {
readyState: 'loading',
body: { addEventListener() {} },
getElementById: id => elements.get(id) || null,
querySelector: () => null,
querySelectorAll: () => [],
addEventListener() {},
dispatchEvent() { return true; },
createElement: tag => new FakeElement(null, tag),
},
CustomEvent: class { constructor(type, init) { this.type = type; this.detail = init && init.detail; } },
setTimeout: (fn, _ms, ...args) => { timers.push(setImmediate(() => fn(...args))); return timers.length; },
clearTimeout: () => {},
setInterval: () => 0,
clearInterval: () => {},
requestAnimationFrame: fn => setImmediate(fn),
getComputedStyle: () => ({ display: 'block' }),
scrollTo() {},
sessionStorage: { getItem: () => null, setItem() {}, removeItem() {} },
localStorage: { getItem: () => null, setItem() {}, removeItem() {} },
confirm: () => true,
alert: () => {},
showNotification: (message, type) => {
toasts.push({ message: String(message),
type: type && typeof type === 'object' ? type.type : type });
},
noteRestartRequired: (body) => { restartNotes.push(body); },
fetch: async (url, opts = {}) => {
const method = (opts.method || 'GET').toUpperCase();
let body = null;
try { body = opts.body ? JSON.parse(opts.body) : null; } catch (e) { body = opts.body; }
requests.push({ method, url: String(url), body });
const answer = (route && route(method, String(url), body)) || { status: 200, json: { status: 'success' } };
const status = answer.status || 200;
return { ok: status < 400, status, json: async () => answer.json };
},
};
ctx.window = ctx;
vm.createContext(ctx);
ledEscape.install(ctx);
for (const file of ['js/plugins/list_filter.js', 'plugins_manager.js']) {
vm.runInContext(fs.readFileSync(path.join(V3, file), 'utf8'), ctx, { filename: file });
}
// Resolves once cond() is true, letting timers and promises run between
// checks; rejects if it never is.
async function until(cond, label = 'condition', turns = 20000) {
for (let i = 0; i < turns; i++) {
if (cond()) return;
await new Promise(r => setImmediate(r));
}
throw new Error('timed out waiting for ' + label);
}
// Lets every pending timer and promise run.
async function settle(turns = 50) {
for (let i = 0; i < turns; i++) await new Promise(r => setImmediate(r));
}
return { window: ctx, el, FakeElement, requests, toasts, errors, restartNotes, until, settle };
}
// The attributes of the element with this id in partials/plugins.html, as
// the template ships them (no Jinja on the tags these suites read).
function templateAttributes(id) {
const html = fs.readFileSync(PLUGINS_HTML, 'utf8');
const at = html.indexOf(`id="${id}"`);
if (at < 0) throw new Error(`no element with id ${id} in plugins.html`);
const start = html.lastIndexOf('<', at);
let end = start, quote = null;
for (; end < html.length; end++) {
const ch = html[end];
if (quote) { if (ch === quote) quote = null; } else if (ch === '"' || ch === "'") quote = ch;
else if (ch === '>') break;
}
const tag = html.slice(start, end + 1);
const attrs = {};
const re = /([\w:-]+)\s*=\s*("([^"]*)"|'([^']*)')/g;
let m;
while ((m = re.exec(tag))) attrs[m[1]] = m[3] !== undefined ? m[3] : m[4];
return { tag: tag.match(/^<(\w+)/)[1], attrs, source: tag };
}
module.exports = { create, templateAttributes };
+11 -2
View File
@@ -17,13 +17,22 @@ const fs = require('fs');
const BASE = process.env.BASE || 'http://localhost:5000'; const BASE = process.env.BASE || 'http://localhost:5000';
const UNIT = ['unit/test_list_filter.js', 'unit/test_render_cards.js', const UNIT = ['unit/test_list_filter.js', 'unit/test_render_cards.js',
'unit/test_plugin_order_list.js',
'unit/test_html_escaping.js', 'unit/test_style_editor_element_keys.js', 'unit/test_html_escaping.js', 'unit/test_style_editor_element_keys.js',
'unit/test_style_editor_layout_leaf_columns.js', 'unit/test_style_editor_layout_leaf_columns.js',
'unit/test_style_editor_layout_leaf_collision.js', 'unit/test_style_editor_layout_leaf_collision.js',
'unit/test_update_all.js', 'unit/test_inline_handler_escaping.js', 'unit/test_update_all.js',
'unit/test_store_install.js',
'unit/test_install_polling.js',
'unit/test_store_categories.js',
'unit/test_github_url_install.js',
'unit/test_inline_handler_escaping.js',
'unit/test_plugin_action_delegation.js', 'unit/test_file_upload_widget.js', 'unit/test_plugin_action_delegation.js', 'unit/test_file_upload_widget.js',
'unit/test_store_registry_fields.js', 'unit/test_restart_banner.js', 'unit/test_store_registry_fields.js', 'unit/test_restart_banner.js',
'unit/test_page_registry.js', 'unit/test_core_modules.js']; 'unit/test_page_registry.js', 'unit/test_core_modules.js',
'unit/test_overview_reconciliation_poll.js',
'unit/test_display_partial_ids.js',
'unit/test_general_web_login_token.js'];
const DOM = ['dom/test_installed_dom.js', 'dom/test_store_dom.js', 'dom/test_no_double_fetch.js', const DOM = ['dom/test_installed_dom.js', 'dom/test_store_dom.js', 'dom/test_no_double_fetch.js',
'dom/test_tools_sections.js', 'dom/test_cache_page.js', 'dom/test_tools_sections.js', 'dom/test_cache_page.js',
'dom/test_durations_page.js', 'dom/test_operation_history_page.js', 'dom/test_durations_page.js', 'dom/test_operation_history_page.js',
+108
View File
@@ -0,0 +1,108 @@
// The Display tab's inline script must only look up elements the partial
// renders.
//
// Its brightness slider handler also wrote to #brightness-display, a "LED
// brightness: N%" line that #387 removed from partials/display.html. The
// lookup returned null, so every movement of the slider threw a TypeError.
// This checks every literal getElementById() in the partial's inline scripts
// against the ids its markup renders, and runs the shipped script in a vm
// with a fake DOM (null for an id the markup lacks, as in a browser) to move
// the slider.
//
// No jsdom and no server needed.
const fs = require('fs');
const path = require('path');
const vm = require('vm');
const PARTIAL = path.resolve(__dirname, '../../../web_interface/templates/v3/partials/display.html');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' ' + JSON.stringify(extra) : '')));
const html = fs.readFileSync(PARTIAL, 'utf8');
const blocks = [...html.matchAll(/<script\b[^>]*>([\s\S]*?)<\/script[^>]*>/gi)];
const scripts = blocks.map(m => m[1]);
// The markup is what lies between the script blocks (sliced around them, not
// a replace(), which CodeQL reads as an incomplete HTML sanitizer).
let markup = '';
let from = 0;
for (const m of blocks) {
markup += html.slice(from, m.index);
from = m.index + m[0].length;
}
markup += html.slice(from);
const rendered = new Set([...markup.matchAll(/\bid="([^"{}]+)"/g)].map(m => m[1]));
console.log('\n── Display partial: element lookups ──');
// 1. Static: every literal lookup names an id the partial renders.
const lookups = scripts.flatMap(s => [...s.matchAll(/getElementById\('([^']+)'\)/g)].map(m => m[1]));
const missing = [...new Set(lookups.filter(id => !rendered.has(id)))];
ok('the inline scripts look elements up', lookups.length > 0, lookups.length);
ok('every looked-up id is rendered by the partial', missing.length === 0, missing);
// 2. Behaviour: moving the brightness slider updates its label and throws nothing.
function fakeElement(id) {
const listeners = {};
const classes = new Set();
return {
id, value: '', textContent: '', min: '', max: '', checked: false,
style: {}, dataset: {}, className: '',
classList: {
add: c => classes.add(c), remove: c => classes.delete(c),
toggle: (c, on) => (on === undefined ? (classes.has(c) ? classes.delete(c) : classes.add(c)) : (on ? classes.add(c) : classes.delete(c))),
contains: c => classes.has(c),
},
addEventListener: (type, fn) => { (listeners[type] ||= []).push(fn); },
dispatchEvent() { return true; },
appendChild() {},
listeners,
};
}
const main = scripts.find(s => s.includes("getElementById('brightness')"));
ok('found the script that wires the brightness slider', !!main);
if (main) {
const elements = new Map();
const document = {
readyState: 'complete',
hidden: false,
getElementById: id => {
if (!rendered.has(id)) return null;
if (!elements.has(id)) elements.set(id, fakeElement(id));
return elements.get(id);
},
createElement: () => fakeElement(''),
createTextNode: () => ({}),
addEventListener() {},
};
const window = {
LEDEscape: { html: v => String(v), attr: v => String(v) },
LEDVisibility: { onActive() {} },
};
const context = {
window, document, console, URLSearchParams,
fetch: () => new Promise(() => {}),
setTimeout: () => 0, clearTimeout() {}, setInterval: () => 0, clearInterval() {},
};
vm.createContext(context);
let loadError = null;
try { vm.runInContext(main, context); } catch (e) { loadError = e; }
ok('the script loads', !loadError, loadError && String(loadError));
const slider = elements.get('brightness');
const handlers = (slider && slider.listeners.input) || [];
ok('the slider has an input handler', handlers.length > 0);
let thrown = null;
slider.value = '42';
try { handlers.forEach(fn => fn.call(slider, { target: slider })); } catch (e) { thrown = e; }
ok('moving the slider throws nothing', !thrown, thrown && String(thrown));
ok('...and shows the new value', elements.get('brightness-value').textContent === '42',
elements.get('brightness-value').textContent);
}
console.log(`\n${pass} passed, ${fail} failed\n`);
process.exit(fail ? 1 : 0);
@@ -0,0 +1,107 @@
// Creating an API token on the General tab must leave its form clean.
//
// app.js marks a form data-dirty on any input in it and clears the mark only
// after a successful htmx request; its beforeunload handler then asks "Leave
// site?" while any visible form is still dirty. The token form posts with
// fetch (window.webLogin.createToken in partials/general.html), so after a
// token was created the form stayed dirty and reloading the page while the
// General tab was open prompted about changes that had been saved.
//
// Runs the shipped inline script in a vm with a fake fetch and DOM -- no jsdom
// and no server needed.
const fs = require('fs');
const path = require('path');
const vm = require('vm');
const PARTIAL = path.resolve(__dirname, '../../../web_interface/templates/v3/partials/general.html');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' ' + JSON.stringify(extra) : '')));
function webLoginScript() {
const html = fs.readFileSync(PARTIAL, 'utf8');
const scripts = [...html.matchAll(/<script\b[^>]*>([\s\S]*?)<\/script[^>]*>/gi)].map(m => m[1]);
const found = scripts.find(s => s.includes('window.webLogin = {'));
if (!found) throw new Error('webLogin script not found in general.html');
return found;
}
function el() {
const classes = new Set(['hidden']);
return {
textContent: '', dataset: {}, style: {}, className: '',
classList: { add: c => classes.add(c), remove: c => classes.delete(c), contains: c => classes.has(c) },
appendChild() {}, addEventListener() {}, querySelector: () => null,
};
}
function load(answer) {
const elements = {
'web-login-tokens': el(),
'web-login-new-token-value': el(),
'web-login-new-token': el(),
};
const notes = [];
const window = { showNotification: (m, t) => notes.push([m, t]), alert() {}, confirm: () => true };
const context = {
window, console,
document: {
getElementById: id => elements[id] || null,
createElement: () => el(),
querySelectorAll: () => [],
},
fetch: () => Promise.resolve({
ok: answer.ok, status: answer.ok ? 200 : 400,
json: () => Promise.resolve(answer.body),
}),
};
vm.createContext(context);
vm.runInContext(webLoginScript(), context);
return { webLogin: context.window.webLogin, elements, notes };
}
function dirtyForm() {
const attrs = new Map([['data-dirty', '']]);
return {
querySelector: sel => (sel === '[name="name"]' ? { value: 'Home Assistant' } : null),
reset() {},
hasAttribute: name => attrs.has(name),
setAttribute: (name, value) => attrs.set(name, String(value)),
removeAttribute: name => attrs.delete(name),
};
}
const flush = async () => { for (let i = 0; i < 10; i++) await new Promise(r => setImmediate(r)); };
(async () => {
console.log('\n── General tab: API token form ──');
{
const t = load({ ok: true, body: {
status: 'success', message: 'Token created',
data: { token: 'lmx_secret', record: { id: 't1', name: 'Home Assistant', prefix: 'lmx_sec' } },
} });
const form = dirtyForm();
t.webLogin.createToken(form);
await flush();
ok('the new token is shown', t.elements['web-login-new-token-value'].textContent === 'lmx_secret');
ok('a created token leaves the form clean (no "Leave site?" on reload)',
!form.hasAttribute('data-dirty'));
}
{
const t = load({ ok: false, body: { status: 'error', message: 'Name is required' } });
const form = dirtyForm();
t.webLogin.createToken(form);
await flush();
ok('a refused request reports the error', t.notes.some(([m, type]) => type === 'error' && /Name is required/.test(m)),
t.notes);
ok('...and keeps the form dirty: nothing was saved', form.hasAttribute('data-dirty'));
}
console.log(`\n${pass} passed, ${fail} failed\n`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.log('HARNESS ERROR: ' + e.stack); process.exit(1); });
+93
View File
@@ -0,0 +1,93 @@
// Plugin Manager > Install from GitHub > Install Single Plugin: one click,
// one request, no errors.
//
// The Install button carried an inline onclick calling
// window.handleGitHubPluginInstall, and attachInstallButtonHandler also gave
// it a click listener that installs. Both ran on every click. The inline one
// threw a ReferenceError (it called isGithubUrl, which lives inside the
// plugin-manager IIFE, from outside it), so only the listener's request went
// out -- and fixing that scope alone would have sent every install twice.
// The button now has the listener only.
//
// Runs the whole of plugins_manager.js in the sandbox, with the button as
// partials/plugins.html ships it.
const { create, templateAttributes } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
const URL = 'https://github.com/someone/ledmatrix-demo';
function route(method, url) {
if (method === 'POST' && url === '/api/v3/plugins/install-from-url') {
return { json: { status: 'success', message: 'Plugin demo installed successfully', plugin_id: 'demo' } };
}
if (url.startsWith('/api/v3/plugins/installed')) return { json: { status: 'success', data: { plugins: [] } } };
return { json: { status: 'success' } };
}
function page() {
const sb = create({ route });
const button = templateAttributes('install-plugin-from-url');
sb.el('install-plugin-from-url', button.tag, button.attrs);
sb.el('github-plugin-url', 'input');
sb.el('github-plugin-status');
sb.el('plugin-branch-input', 'input');
return sb;
}
const installs = sb => sb.requests.filter(r => r.url === '/api/v3/plugins/install-from-url');
(async () => {
console.log('\nthe template');
{
const { attrs } = templateAttributes('install-plugin-from-url');
ok('the Install button has no inline onclick', !('onclick' in attrs), attrs.onclick);
}
console.log('\na click');
{
const sb = page();
sb.window.attachInstallButtonHandler();
// htmx:afterSettle runs it again on every swap; that must not add a handler.
sb.window.attachInstallButtonHandler();
sb.el('github-plugin-url').value = URL;
sb.window.document.getElementById('install-plugin-from-url').click();
await sb.settle();
ok('raises no error', sb.errors.length === 0, sb.errors.map(String));
ok('sends exactly one install request', installs(sb).length === 1, installs(sb));
ok('for the URL typed', installs(sb)[0] && installs(sb)[0].body.repo_url === URL, installs(sb));
ok('and reports the result', /Successfully installed: demo/.test(sb.el('github-plugin-status').innerHTML),
sb.el('github-plugin-status').innerHTML);
}
console.log('\nEnter in the URL field');
{
const sb = page();
sb.window.attachInstallButtonHandler();
const input = sb.el('github-plugin-url');
input.value = URL;
input.dispatch('keypress', { key: 'Enter' });
await sb.settle();
ok('raises no error', sb.errors.length === 0, sb.errors.map(String));
ok('sends exactly one install request', installs(sb).length === 1, installs(sb));
}
console.log('\na URL that is not GitHub');
{
const sb = page();
sb.window.attachInstallButtonHandler();
sb.el('github-plugin-url').value = 'https://example.com/x';
sb.window.document.getElementById('install-plugin-from-url').click();
await sb.settle();
ok('is refused without a request or an error',
installs(sb).length === 0 && sb.errors.length === 0 && /valid GitHub URL/.test(sb.el('github-plugin-status').innerHTML),
{ errors: sb.errors.map(String), status: sb.el('github-plugin-status').innerHTML });
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+89
View File
@@ -0,0 +1,89 @@
// How long the store's Install waits for a queued install, and what it says
// when it stops waiting.
//
// It polled the operation 60 times, a second apart, then reported "Install
// operation timed out" as an error and did nothing else. The server is
// allowed far longer: the plugin's dependency install alone may take 300 s
// (install_requirements_file in src/plugin_system/store_install.py), after
// a download that fetches the plugin one file at a time. So an install that
// went on to succeed was reported as failed, never enabled, and missing from
// the installed list until the page was reloaded.
//
// Runs the whole of plugins_manager.js in the sandbox; its timers ignore
// their delays, so each poll here stands for one second on a real page.
const { create } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
// The server's dependency-install timeout, in polls (one a second).
const DEPENDENCY_INSTALL_TIMEOUT_POLLS = 300;
function server(completesAfterPolls) {
const state = { polls: 0, installed: [] };
state.route = (method, url, body) => {
if (url.startsWith('/api/v3/plugins/installed')) {
return { json: { status: 'success', data: { plugins: state.installed.map(p => ({ ...p })) } } };
}
if (method === 'POST' && url === '/api/v3/plugins/install') {
return { json: { status: 'success', message: 'queued', data: { operation_id: 'op-1' } } };
}
if (url === '/api/v3/plugins/operation/op-1') {
state.polls++;
if (completesAfterPolls === null || state.polls < completesAfterPolls) {
return { json: { status: 'success', data: { status: 'running' } } };
}
state.installed = [{ id: 'clock-simple', name: 'Clock', enabled: false }];
return { json: { status: 'success', data: { status: 'completed',
result: { success: true, message: 'installed', plugin_id: 'clock-simple' } } } };
}
if (method === 'POST' && url === '/api/v3/plugins/toggle') {
return { json: { status: 'success', message: 'enabled' } };
}
return { json: { status: 'success' } };
};
return state;
}
(async () => {
console.log('\nan install that takes longer than a minute');
{
// 200 s: well inside what the server allows.
const srv = server(200);
const sb = create({ route: srv.route });
sb.window.installPlugin('clock-simple');
await sb.until(() => sb.toasts.some(t => /installed and enabled|enabling it failed|timed out|still/i.test(t.message)),
'the install to finish');
await sb.settle();
ok('is waited for until it completes', srv.polls === 200, srv.polls);
ok('is not reported as an error', !sb.toasts.some(t => t.type === 'error'), sb.toasts);
ok('and is enabled', sb.requests.some(r => r.url === '/api/v3/plugins/toggle' && r.body.plugin_id === 'clock-simple'),
sb.requests.filter(r => r.method === 'POST'));
}
console.log('\nan install that never reports back');
{
const srv = server(null);
const sb = create({ route: srv.route });
sb.window.installPlugin('clock-simple');
await sb.until(() => sb.toasts.length >= 3, 'the poller to give up');
await sb.settle();
ok(`is polled for at least the ${DEPENDENCY_INSTALL_TIMEOUT_POLLS} s dependency-install timeout`,
srv.polls >= DEPENDENCY_INSTALL_TIMEOUT_POLLS, srv.polls);
ok('...but not forever', srv.polls <= 1200, srv.polls);
const lastPoll = sb.requests.map(r => r.url).lastIndexOf('/api/v3/plugins/operation/op-1');
ok('then the installed list is reloaded, to show what actually happened',
sb.requests.slice(lastPoll + 1).some(r => r.url === '/api/v3/plugins/installed'),
sb.requests.slice(lastPoll + 1).map(r => r.url));
const last = sb.toasts[sb.toasts.length - 1];
ok('it says the install may still be running, as a warning, not a failure',
last && last.type === 'warning' && !/fail|timed out/i.test(last.message), sb.toasts);
ok('nothing is enabled on a guess', !sb.requests.some(r => r.url === '/api/v3/plugins/toggle'));
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
@@ -0,0 +1,137 @@
// The Overview's "Plugin Config Warning" poll must end.
//
// The banner script in partials/overview.html asks
// /api/v3/plugins/reconciliation-status every 2 s until startup reconciliation
// says it is done. The route answers done: false whenever its status file is
// missing -- reconciliation raised before writing it, or /tmp was cleaned
// under a long-running web service -- so the poll used to run every 2 s for
// as long as the page stayed open, on every tab. It now gives up after a
// bounded number of tries and runs only while the Overview is on screen
// (LEDVisibility, like the other partials' pollers).
//
// Runs the shipped inline script in a vm with fake timers, fetch and DOM --
// no jsdom and no server needed.
const fs = require('fs');
const path = require('path');
const vm = require('vm');
const PARTIAL = path.resolve(__dirname, '../../../web_interface/templates/v3/partials/overview.html');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' ' + JSON.stringify(extra) : '')));
function bannerScript() {
const html = fs.readFileSync(PARTIAL, 'utf8');
const scripts = [...html.matchAll(/<script\b[^>]*>([\s\S]*?)<\/script[^>]*>/gi)].map(m => m[1]);
const found = scripts.find(s => s.includes('ledmatrix-recon-dismissed'));
if (!found) throw new Error('reconciliation banner script not found in overview.html');
return found;
}
const flush = async () => { for (let i = 0; i < 10; i++) await new Promise(r => setImmediate(r)); };
function load({ payload, visibility = true }) {
const timers = new Map();
let nextId = 1;
const calls = [];
const banner = { style: { setProperty() {} }, dataset: {} };
const text = { textContent: '' };
const registrations = [];
const window = {};
if (visibility) {
window.LEDVisibility = {
onActive(tab, start, stop, key) { registrations.push({ tab, start, stop, key }); start(); },
};
}
const context = {
window,
document: {
getElementById: id => ({ 'reconciliation-banner': banner, 'reconciliation-banner-text': text })[id] || null,
},
sessionStorage: { getItem: () => null, setItem() {} },
fetch: (url) => {
calls.push(url);
return Promise.resolve({ json: () => Promise.resolve(payload()) });
},
setTimeout: (fn) => { const id = nextId++; timers.set(id, fn); return id; },
clearTimeout: (id) => { timers.delete(id); },
};
vm.createContext(context);
vm.runInContext(bannerScript(), context);
const fireTimers = async () => {
const due = [...timers.entries()];
timers.clear();
due.forEach(([, fn]) => fn());
await flush();
};
return { calls, timers, registrations, banner, text, window, fireTimers };
}
(async () => {
console.log('\n── Overview reconciliation poll ──');
// 1. A status file that never says done: the poll stops on its own.
{
const t = load({ payload: () => ({ status: 'success', data: { done: false, unresolved: [] } }) });
await flush();
for (let i = 0; i < 200; i++) await t.fireTimers();
ok('a status that never turns done stops being polled', t.timers.size === 0,
{ pending: t.timers.size, requests: t.calls.length });
ok('...after a bounded number of requests (at most 30, a minute at 2 s)',
t.calls.length > 1 && t.calls.length <= 30, t.calls.length);
}
// 2. Runs only while the Overview is on screen.
{
const t = load({ payload: () => ({ status: 'success', data: { done: false, unresolved: [] } }) });
await flush();
const reg = t.registrations[0];
ok('registers with LEDVisibility for the overview tab', !!reg && reg.tab === 'overview', reg && reg.tab);
ok('under its own key, so it does not replace another overview poller',
!!reg && !!reg.key && reg.key !== 'overview', reg && reg.key);
ok('first request goes out at once', t.calls.length === 1, t.calls.length);
if (reg) {
reg.stop();
ok('leaving the tab cancels the pending retry', t.timers.size === 0, t.timers.size);
for (let i = 0; i < 5; i++) await t.fireTimers();
ok('no requests while another tab is active', t.calls.length === 1, t.calls.length);
reg.start();
await flush();
ok('coming back asks again at once', t.calls.length === 2, t.calls.length);
ok('...and keeps polling', t.timers.size === 1, t.timers.size);
}
}
// 3. A finished reconciliation with findings shows the banner and stops.
{
let done = false;
const t = load({ payload: () => (done
? { status: 'success', data: { done: true, unresolved: [{ plugin_id: 'clock', type: 'plugin_missing_on_disk' }] } }
: { status: 'success', data: { done: false, unresolved: [] } }) });
await flush();
await t.fireTimers();
done = true;
await t.fireTimers();
ok('the banner names the finding once reconciliation is done',
t.text.textContent.includes('clock'), t.text.textContent);
const before = t.calls.length;
for (let i = 0; i < 5; i++) await t.fireTimers();
ok('no more requests once it is done', t.calls.length === before && t.timers.size === 0,
{ before, after: t.calls.length, pending: t.timers.size });
}
// 4. Without LEDVisibility (base.html always has it) it still runs, bounded.
{
const t = load({ visibility: false, payload: () => ({ status: 'success', data: { done: false } }) });
await flush();
ok('runs without LEDVisibility', t.calls.length === 1, t.calls.length);
for (let i = 0; i < 200; i++) await t.fireTimers();
ok('...and is still bounded', t.timers.size === 0 && t.calls.length <= 30, t.calls.length);
}
console.log(`\n${pass} passed, ${fail} failed\n`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.log('HARNESS ERROR: ' + e.stack); process.exit(1); });
+189
View File
@@ -0,0 +1,189 @@
// The shared plugin order list (widgets/plugin-order-list.js) keeps what it
// does not show.
//
// It lists enabled plugins only, and rewrites its hidden inputs from those
// rows as soon as it has drawn them. A disabled plugin's place in the order
// and its Vegas exclusion used to vanish from the inputs on that rewrite, so
// any later save of the Display or Rotation & Durations tab stored them
// without it: re-enabled, the plugin came back at the end of the rotation and
// scrolling in Vegas again. An uninstalled plugin's id is still dropped, as
// before, so the lists don't collect ids nothing can show. Runs the shipped
// widget in a vm with a minimal fake DOM -- no jsdom and no server needed, so
// it runs under test/test_js_unit_suites.py too.
const fs = require('fs');
const path = require('path');
const vm = require('vm');
const WIDGET = path.resolve(__dirname, '../../../web_interface/static/v3/js/widgets/plugin-order-list.js');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' ' + JSON.stringify(extra) : '')));
const same = (a, b) => JSON.stringify(a) === JSON.stringify(b);
class FakeElement {
constructor(tag) {
this.tagName = tag.toUpperCase();
this.children = [];
this.parent = null;
this.dataset = {};
this.style = {};
this.className = '';
this.value = '';
this.checked = false;
this.listeners = {};
this._text = '';
}
appendChild(child) {
if (child.parent) child.parent.children = child.parent.children.filter(c => c !== child);
child.parent = this;
this.children.push(child);
return child;
}
insertBefore(child, ref) {
if (!ref) return this.appendChild(child);
if (child.parent) child.parent.children = child.parent.children.filter(c => c !== child);
child.parent = this;
this.children.splice(this.children.indexOf(ref), 0, child);
return child;
}
get previousElementSibling() {
const siblings = this.parent ? this.parent.children : [];
return siblings[siblings.indexOf(this) - 1] || null;
}
get nextElementSibling() {
const siblings = this.parent ? this.parent.children : [];
const i = siblings.indexOf(this);
return i < 0 ? null : siblings[i + 1] || null;
}
set textContent(value) { this._text = value; this.children = []; }
get textContent() { return this._text; }
setAttribute() {}
focus() {}
addEventListener(type, fn) { (this.listeners[type] ||= []).push(fn); }
fire(type, event) { (this.listeners[type] || []).forEach(fn => fn.call(this, event || {})); }
descendants() { return this.children.flatMap(c => [c, ...c.descendants()]); }
querySelectorAll(selector) {
const cls = selector.replace(/^\./, '');
return this.descendants().filter(e => e.className.split(/\s+/).includes(cls));
}
querySelector(selector) { return this.querySelectorAll(selector)[0] || null; }
}
/** Run the widget over `plugins` with the given saved inputs; resolves once it has drawn. */
async function mount({ plugins, order, excluded, fetchFails }) {
const els = {
list: new FakeElement('div'),
order: Object.assign(new FakeElement('input'), { value: JSON.stringify(order) }),
};
if (excluded !== undefined) {
els.excluded = Object.assign(new FakeElement('input'), { value: JSON.stringify(excluded) });
}
const context = {
// The widget logs a failed list; expected there, so kept off the output.
console: fetchFails ? Object.assign({}, console, { error: () => {} }) : console,
window: {},
document: {
getElementById: (id) => els[id] || null,
createElement: (tag) => new FakeElement(tag),
createTextNode: (text) => new FakeElement('#text'),
},
fetch: () => (fetchFails ? Promise.reject(new Error('service restarting')) : Promise.resolve({
json: () => Promise.resolve({ status: 'success', data: { plugins } }),
})),
};
vm.createContext(context);
vm.runInContext(fs.readFileSync(WIDGET, 'utf8'), context);
context.window.PluginOrderList.init({
containerId: 'list', orderInputId: 'order',
excludedInputId: excluded !== undefined ? 'excluded' : undefined,
});
await new Promise(resolve => setTimeout(resolve, 0));
const rows = () => els.list.querySelectorAll('.plugin-order-item');
return {
rows,
rowIds: () => rows().map(r => r.dataset.pluginId),
order: () => JSON.parse(els.order.value),
excluded: () => JSON.parse(els.excluded.value),
row: (id) => rows().find(r => r.dataset.pluginId === id),
};
}
const PLUGINS = [
{ id: 'weather', name: 'Weather', enabled: true },
{ id: 'clock', name: 'Clock', enabled: false },
{ id: 'stocks', name: 'Stocks', enabled: true },
];
(async () => {
console.log('\nVegas: a disabled plugin keeps its place and its exclusion');
{
const t = await mount({ plugins: PLUGINS, order: ['weather', 'clock', 'stocks'], excluded: ['clock'] });
ok('only enabled plugins get a row', same(t.rowIds(), ['weather', 'stocks']), t.rowIds());
ok('drawing the list keeps the disabled plugin in the order, in its place',
same(t.order(), ['weather', 'clock', 'stocks']), t.order());
ok('drawing the list keeps its exclusion', same(t.excluded(), ['clock']), t.excluded());
// Move Stocks up: the rows swap, and Clock stays in its saved slot.
const up = t.row('stocks').querySelectorAll('.plugin-order-move')[0];
up.fire('click');
ok('reordering the rows fills the other slots in the new order',
same(t.order(), ['stocks', 'clock', 'weather']), t.order());
const include = t.row('weather').querySelector('.plugin-order-include');
include.checked = false;
include.fire('change');
ok('unchecking a row adds it, and the disabled exclusion stays',
same([...t.excluded()].sort(), ['clock', 'weather']), t.excluded());
include.checked = true;
include.fire('change');
ok('checking it again removes only that one', same(t.excluded(), ['clock']), t.excluded());
}
console.log('\nRotation order: the same, without exclusions');
{
const plugins = [
{ id: 'clock', enabled: true },
{ id: 'off', enabled: false },
{ id: 'weather', enabled: true },
{ id: 'new', enabled: true },
];
const t = await mount({ plugins, order: ['clock', 'off', 'weather'] });
ok('the disabled plugin keeps its slot; a plugin not in the saved order goes last',
same(t.order(), ['clock', 'off', 'weather', 'new']), t.order());
}
console.log('\nAn uninstalled plugin is dropped; a failed list keeps everything');
{
const t = await mount({ plugins: PLUGINS, order: ['weather', 'gone', 'clock', 'stocks'],
excluded: ['gone', 'clock'] });
ok('the disabled plugin is kept and the uninstalled one dropped from the order',
same(t.order(), ['weather', 'clock', 'stocks']), t.order());
ok('and from the exclusions', same(t.excluded(), ['clock']), t.excluded());
}
{
const t = await mount({ plugins: PLUGINS, order: ['weather', 'gone', 'clock', 'stocks'],
excluded: ['gone', 'clock'], fetchFails: true });
// No installed list, so nothing can be told apart: no rows, and the
// inputs keep what was saved, uninstalled ids included.
ok('a failed plugin list draws no rows', t.rowIds().length === 0, t.rowIds());
ok('and leaves the saved order as it was',
same(t.order(), ['weather', 'gone', 'clock', 'stocks']), t.order());
ok('and the saved exclusions', same(t.excluded(), ['gone', 'clock']), t.excluded());
}
console.log('\nOnly what the server would accept is carried over');
{
const t = await mount({ plugins: PLUGINS, order: ['weather', 7, 'clock', null, 'clock', 'stocks'],
excluded: ['clock', 3, 'clock'] });
// /config/main refuses a list holding anything but strings, which would
// block every later Display save; a repeated id is kept once.
ok('non-string and repeated saved ids are dropped from the order',
same(t.order(), ['weather', 'clock', 'stocks']), t.order());
ok('and from the exclusions', same(t.excluded(), ['clock']), t.excluded());
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+105
View File
@@ -0,0 +1,105 @@
// The Plugin Store's category filter offers the categories its plugins have.
//
// The template listed seven fixed categories. The registry uses about
// twenty (productivity, utility, transit, finance, ...), so roughly a third
// of the store could not be filtered to at all, and "Financial" missed the
// plugin filed under "finance". The options are now built from the store's
// plugins, as the Starlark section builds its own; the template ships only
// "All Categories".
//
// Runs the whole of plugins_manager.js in the sandbox.
const fs = require('fs');
const path = require('path');
const { create, templateAttributes } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
const STORE = [
{ id: 'nfl', name: 'NFL', category: 'sports' },
{ id: 'nba', name: 'NBA', category: 'Sports' },
{ id: 'todo', name: 'Todo', category: 'productivity' },
{ id: 'stocks', name: 'Stocks', category: 'finance' },
{ id: 'crypto', name: 'Crypto', category: 'financial' },
{ id: 'bus', name: 'Bus', category: 'transit' },
{ id: 'mystery', name: 'Mystery' },
];
function route(method, url) {
if (url.startsWith('/api/v3/plugins/store/list')) return { json: { status: 'success', data: { plugins: STORE } } };
if (url.startsWith('/api/v3/plugins/installed')) return { json: { status: 'success', data: { plugins: [] } } };
if (url.startsWith('/api/v3/plugins/store/github-status')) {
return { json: { status: 'success', data: { token_status: 'valid', authenticated: true, rate_limit: 5000 } } };
}
if (url.startsWith('/api/v3/plugins/saved-repositories')) {
return { json: { status: 'success', data: { repositories: [] } } };
}
if (url.startsWith('/api/v3/display/on-demand/status')) {
return { json: { status: 'success', data: { state: {}, service: {} } } };
}
return { json: { status: 'success' } };
}
const options = sel => sel.children.map(o => o.value);
const cardIds = sb => [...sb.el('plugin-store-grid').innerHTML.matchAll(/<h4[^>]*>([^<]*)<\/h4>/g)].map(m => m[1]);
(async () => {
console.log('\nthe template');
{
const html = fs.readFileSync(path.resolve(__dirname,
'../../../web_interface/templates/v3/partials/plugins.html'), 'utf8');
const start = html.indexOf('<select id="plugin-category"');
const block = html.slice(start, html.indexOf('</select>', start));
const shipped = [...block.matchAll(/<option value="([^"]*)"/g)].map(m => m[1]);
ok('ships only "All Categories"', JSON.stringify(shipped) === JSON.stringify(['']), shipped);
}
const sb = create({ route });
const attrs = templateAttributes('plugin-category').attrs;
const select = sb.el('plugin-category', 'select', attrs);
sb.el('plugin-store-grid');
sb.el('installed-plugins-grid');
sb.window.initPluginsPage();
await sb.until(() => sb.requests.some(r => r.url.startsWith('/api/v3/plugins/store/list')), 'the store list');
await sb.settle();
console.log('\noptions come from the store\'s plugins');
ok('"All Categories" is still the first choice', /<option value="">All Categories<\/option>/.test(select.innerHTML),
select.innerHTML);
ok('every category a plugin has is offered once, whatever its case',
JSON.stringify(options(select)) === JSON.stringify(['finance', 'financial', 'productivity', 'sports', 'transit']),
options(select));
ok('labels are capitalised',
(select.children.find(o => o.value === 'productivity') || {}).textContent === 'Productivity');
console.log('\nchoosing one filters to it');
select.value = 'productivity';
select.dispatch('change');
ok('productivity shows its plugin', JSON.stringify(cardIds(sb)) === JSON.stringify(['Todo']), cardIds(sb));
select.value = 'finance';
select.dispatch('change');
ok('finance is not lost to "financial"', JSON.stringify(cardIds(sb)) === JSON.stringify(['Stocks']), cardIds(sb));
select.value = 'sports';
select.dispatch('change');
ok('one option covers both spellings of sports',
JSON.stringify(cardIds(sb).sort()) === JSON.stringify(['NBA', 'NFL']), cardIds(sb));
console.log('\nthe partial is swapped back in (tab switch)');
{
// A fresh <select> from the template, the store list still cached.
const fresh = new sb.FakeElement('plugin-category', 'select', attrs);
select.parentNode.replaceChild(fresh, select);
sb.window.searchPluginStore(false);
await sb.settle();
ok('the new select is filled from the cache',
JSON.stringify(options(fresh)) === JSON.stringify(['finance', 'financial', 'productivity', 'sports', 'transit']),
options(fresh));
ok('keeping the chosen category', fresh.value === 'sports', fresh.value);
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+138
View File
@@ -0,0 +1,138 @@
// The store's Install button: which plugin it enables afterwards, and when.
//
// 1. Weather, Music, Stocks and Leaderboard are registry entries (`weather`)
// whose manifests declare another id (`ledmatrix-weather`). The plugin
// list, its config section and /plugins/toggle know them by that id, but
// the button enabled the registry id: /plugins/toggle answered 404
// "Plugin not found" and the plugin stayed disabled behind "installed,
// but enabling it failed". It now enables the id the install answer
// names (`plugin_id`), or, from an answer without one, the installed
// entry the store entry matches (its id, plugin_path name or aliases).
//
// 2. Reinstall (the same button on an installed plugin) enabled it too, so
// reinstalling a plugin the user had switched off switched it back on.
// Only a fresh install enables.
//
// Runs the whole of plugins_manager.js in the sandbox against a fake API.
const { create } = require('../plugins_manager_sandbox');
let pass = 0, fail = 0;
const ok = (label, cond, extra) => cond
? (pass++, console.log(' ok ' + label))
: (fail++, console.log(' FAIL ' + label + (extra !== undefined ? ' -> ' + JSON.stringify(extra).slice(0, 400) : '')));
const STORE = [
{ id: 'weather', name: 'Weather', category: 'weather', plugin_path: 'plugins/ledmatrix-weather',
aliases: ['ledmatrix-weather'] },
{ id: 'clock-simple', name: 'Clock', category: 'time', plugin_path: 'plugins/clock-simple', aliases: [] },
];
// A server with one install in flight. `queue` false answers the install
// directly; `names` false leaves plugin_id out of the answer (an older
// server); `installsAs` is the id the installed manifest declares.
function server({ installed = [], queue = true, names = true, installsAs }) {
const state = { installed: installed.map(p => ({ ...p })), polls: 0 };
const done = (id) => {
if (!state.installed.some(p => p.id === installsAs)) {
state.installed.push({ id: installsAs, name: id, enabled: false });
}
const result = { success: true, message: `Plugin ${id} installed successfully`, restart_required: false };
if (names) result.plugin_id = installsAs;
return result;
};
state.route = (method, url, body) => {
if (url.startsWith('/api/v3/plugins/store/list')) {
return { json: { status: 'success', data: { plugins: STORE } } };
}
if (url.startsWith('/api/v3/plugins/installed')) {
return { json: { status: 'success', data: { plugins: state.installed.map(p => ({ ...p })) } } };
}
if (method === 'POST' && url === '/api/v3/plugins/install') {
if (queue) return { json: { status: 'success', message: 'queued', data: { operation_id: 'op-1' } } };
return { json: { status: 'success', message: 'Plugin installed successfully', ...done(body.plugin_id) } };
}
if (url === '/api/v3/plugins/operation/op-1') {
state.polls++;
if (state.polls < 3) return { json: { status: 'success', data: { status: 'running' } } };
return { json: { status: 'success', data: { status: 'completed', result: done('weather') } } };
}
if (method === 'POST' && url === '/api/v3/plugins/toggle') {
const plugin = state.installed.find(p => p.id === body.plugin_id);
if (!plugin) return { status: 404, json: { status: 'error', message: 'Plugin not found' } };
plugin.enabled = body.enabled;
return { json: { status: 'success', message: `Plugin ${body.plugin_id} enabled successfully` } };
}
return { json: { status: 'success' } };
};
return state;
}
async function install(pluginId, opts) {
const srv = server(opts);
const sb = create({ route: srv.route });
sb.window.searchPluginStore();
await sb.until(() => sb.requests.some(r => r.url.startsWith('/api/v3/plugins/store/list')), 'store list');
await sb.window.pluginManager.loadInstalledPlugins(true);
await sb.settle();
sb.requests.length = 0;
sb.toasts.length = 0;
sb.window.installPlugin(pluginId);
await sb.until(() => sb.toasts.some(t => /installed and enabled|enabling it failed|reinstalled/.test(t.message)),
'the install to finish');
await sb.settle();
const toggles = sb.requests.filter(r => r.url === '/api/v3/plugins/toggle').map(r => r.body);
return { sb, srv, toggles };
}
(async () => {
console.log('\n1. a fresh install enables the id the plugin was installed as');
{
const { srv, toggles, sb } = await install('weather', { installsAs: 'ledmatrix-weather' });
ok('enables ledmatrix-weather, not the registry id',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'ledmatrix-weather', enabled: true }]), toggles);
ok('...which the server enabled', srv.installed.find(p => p.id === 'ledmatrix-weather').enabled === true, srv.installed);
ok('says so', sb.toasts.some(t => t.type === 'success' && /installed and enabled/.test(t.message)), sb.toasts);
ok('no "Plugin not found"', !sb.toasts.some(t => /not found|failed/.test(t.message)), sb.toasts);
const lastList = sb.requests.map(r => r.url).lastIndexOf('/api/v3/plugins/installed');
const toggleAt = sb.requests.findIndex(r => r.url === '/api/v3/plugins/toggle');
ok('the installed list is reloaded before enabling, so the new card is there to update',
lastList >= 0 && lastList < toggleAt, sb.requests.map(r => r.method + ' ' + r.url));
}
{
const { toggles } = await install('weather', { installsAs: 'ledmatrix-weather', names: false });
ok('an answer without plugin_id: the installed entry the store entry matches (its alias)',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'ledmatrix-weather', enabled: true }]), toggles);
}
{
const { toggles } = await install('weather', { installsAs: 'ledmatrix-weather', queue: false });
ok('without the operation queue, from the direct answer',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'ledmatrix-weather', enabled: true }]), toggles);
}
{
const { toggles } = await install('clock-simple', { installsAs: 'clock-simple', names: false });
ok('a plugin installed under its registry id is enabled by that id',
JSON.stringify(toggles) === JSON.stringify([{ plugin_id: 'clock-simple', enabled: true }]), toggles);
}
console.log('\n2. a reinstall leaves the plugin as the user had it');
{
const { srv, toggles, sb } = await install('weather', {
installsAs: 'ledmatrix-weather', installed: [{ id: 'ledmatrix-weather', name: 'Weather', enabled: false }],
});
ok('sends no toggle', toggles.length === 0, toggles);
ok('the plugin stays disabled', srv.installed.find(p => p.id === 'ledmatrix-weather').enabled === false);
ok('says it was reinstalled', sb.toasts.some(t => t.type === 'success' && /reinstalled/.test(t.message)), sb.toasts);
ok('and reloads the list',
sb.requests.some(r => r.url === '/api/v3/plugins/installed'), sb.requests.map(r => r.url));
}
{
const { toggles } = await install('weather', {
installsAs: 'ledmatrix-weather', installed: [{ id: 'ledmatrix-weather', name: 'Weather', enabled: true }],
});
ok('an enabled plugin is not toggled either', toggles.length === 0, toggles);
}
console.log(`\n${pass} passed, ${fail} failed`);
process.exit(fail ? 1 : 0);
})().catch(e => { console.error(e); process.exit(1); });
+2 -1
View File
@@ -52,7 +52,8 @@ global.installedPlugins = [];
// eslint-disable-next-line no-eval // eslint-disable-next-line no-eval
eval([ eval([
'function escapeHtml(text) {', 'function escapeAttribute(text) {', 'function jsStringAttr(value) {', 'function escapeHtml(text) {', 'function escapeAttribute(text) {', 'function jsStringAttr(value) {',
'function isStorePluginInstalled(pluginIdOrPlugin) {', 'function renderPluginStore(plugins) {', 'function isStorePluginInstalled(pluginIdOrPlugin) {',
'function findInstalledStorePlugin(pluginIdOrPlugin) {', 'function renderPluginStore(plugins) {',
].map(extract).join('\n') + '\nglobal.renderPluginStore = renderPluginStore;' ].map(extract).join('\n') + '\nglobal.renderPluginStore = renderPluginStore;'
+ '\nglobal.isStorePluginInstalled = isStorePluginInstalled;'); + '\nglobal.isStorePluginInstalled = isStorePluginInstalled;');
+60 -3
View File
@@ -52,7 +52,7 @@ function fakeApi(behaviour = {}) {
}; };
} }
function setup(api, { stateList, windowList } = {}) { function setup(api, { stateList, windowList, pluginManager } = {}) {
global.window = { global.window = {
PluginAPI: api, PluginAPI: api,
installedPlugins: windowList, installedPlugins: windowList,
@@ -60,6 +60,7 @@ function setup(api, { stateList, windowList } = {}) {
installedPlugins: stateList, installedPlugins: stateList,
loadInstalledPlugins: async () => stateList, loadInstalledPlugins: async () => stateList,
}, },
pluginManager,
}; };
} }
@@ -92,12 +93,68 @@ const noSleep = { sleep: async () => {} };
ok('progress total counts only what is sent', ok('progress total counts only what is sent',
progress.length === EXPECTED.length && progress.every(([, n]) => n === EXPECTED.length), progress); progress.length === EXPECTED.length && progress.every(([, n]) => n === EXPECTED.length), progress);
} }
{
// A page without the plugin manager has no window.installedPlugins.
const api = fakeApi();
setup(api, { stateList: INSTALLED });
await Manager.updateAll(null, noSleep);
ok('the PluginStateManager list (no live list) is filtered the same way',
JSON.stringify(api.calls) === JSON.stringify(EXPECTED), api.calls);
}
console.log('\na second run sends the live list, not the first run\'s snapshot');
{
// Run 1 leaves PluginStateManager holding a, b, c. Then c is uninstalled
// and d installed: plugins_manager.js publishes that only as
// window.installedPlugins. Run 2 used to send a, b, c -- c failed as
// "plugin not found" and d, which had an update waiting, was skipped.
const api = fakeApi({
c: () => { throw { error_code: 'PLUGIN_UPDATE_FAILED', message: 'Plugin update failed: plugin not found' }; },
});
const stale = [{ id: 'a' }, { id: 'b' }, { id: 'c' }];
setup(api, { stateList: stale, windowList: [{ id: 'a' }, { id: 'b' }, { id: 'd' }] });
const results = await Manager.updateAll(null, noSleep);
ok('sends exactly what is installed now',
JSON.stringify(api.calls) === JSON.stringify(['a', 'b', 'd']), api.calls);
ok('...so nothing fails over an uninstalled plugin', results.every(r => r.success), results);
}
{ {
const api = fakeApi(); const api = fakeApi();
setup(api, { stateList: INSTALLED, windowList: [] }); setup(api, { stateList: INSTALLED, windowList: [] });
const results = await Manager.updateAll(null, noSleep);
ok('an empty live list means nothing is installed: nothing is sent',
api.calls.length === 0 && results.length === 0, api.calls);
}
console.log('\nthe end-of-run refresh redraws the installed grid');
{
// PluginStateManager's refresh replaced window.installedPlugins and
// nothing else: the cards kept "Update to vX" and the Updates badge
// kept its count. The plugin manager's load renders the grid.
const loads = [];
let stateLoads = 0;
const pluginManager = { loadInstalledPlugins: async (force) => { loads.push(force); } };
setup(fakeApi(), { stateList: INSTALLED, windowList: INSTALLED, pluginManager });
window.PluginStateManager.loadInstalledPlugins = async () => { stateLoads++; };
await Manager.updateAll(null, noSleep); await Manager.updateAll(null, noSleep);
ok('the PluginStateManager list is filtered the same way', ok('reloads through the plugin manager once, forced past its caches',
JSON.stringify(api.calls) === JSON.stringify(EXPECTED), api.calls); JSON.stringify(loads) === JSON.stringify([true]), loads);
ok('...instead of PluginStateManager', stateLoads === 0, stateLoads);
}
{
const pluginManager = { loadInstalledPlugins: async () => { throw new Error('offline'); } };
const answer = { status: 'success', data: { update_status: 'updated' }, restart_required: true };
setup(fakeApi({ 'ledmatrix-flights': () => answer }), { windowList: INSTALLED, pluginManager });
const warn = console.warn;
console.warn = () => {};
let results;
try {
results = await Manager.updateAll(null, noSleep);
} finally {
console.warn = warn;
}
ok('a failed plugin-manager reload still returns the results with their restart flag',
Array.isArray(results) && Manager.restartRequest(results) === answer);
} }
{ {
const api = fakeApi(); const api = fakeApi();
+41 -5
View File
@@ -145,13 +145,49 @@ class TestContextImageCache:
a = ctx.fit_image(img, (20, 20)) a = ctx.fit_image(img, (20, 20))
assert ctx.fit_image(img, (20, 20)) is a assert ctx.fit_image(img, (20, 20)) is a
def test_id_safety_pins_source(self, ctx): def test_id_keyed_entry_does_not_pin_source(self, ctx):
# id()-keyed entries must pin the source image so a recycled id import gc
# can't alias a dead image's cache entry. import weakref
img = _solid(10, 10) img = _solid(10, 10)
ctx.fit_image(img, (20, 20)) ctx.fit_image(img, (20, 20))
pinned = [entry[1] for entry in ctx._image_cache.values()] assert len(ctx._image_cache) == 1
assert img in pinned watch = weakref.ref(img)
del img
gc.collect()
assert watch() is None # the cache did not keep it alive
assert len(ctx._image_cache) == 0 # and its entry went with it
def test_fresh_image_each_frame_holds_nothing(self, ctx):
# draw_image(Image.open(path), box) every frame: the old pinning
# filled all 64 slots with dead-weight sources.
for _ in range(ctx._IMAGE_CACHE_MAX * 2):
ctx.fit_image(_solid(50, 50), (20, 20))
assert len(ctx._image_cache) == 0
def test_recycled_id_does_not_alias(self, ctx):
# A same-size image at a recycled address must not get the dead
# image's fit, even if the entry somehow outlived its source.
red = _solid(10, 10, (255, 0, 0, 255))
first = ctx.fit_image(red, (20, 20))
key = next(iter(ctx._image_cache))
entry = ctx._image_cache[key]
ctx._image_cache[key] = (entry[0], lambda: None) # source "gone"
again = ctx.fit_image(red, (20, 20))
assert again is not first # refit, not a stale hit
assert ctx.fit_image(red, (20, 20)) is again
def test_unweakrefable_source_is_pinned(self, ctx, monkeypatch):
import src.adaptive_layout as layout_mod
def no_weakref(*_args, **_kwargs):
raise TypeError("cannot create weak reference")
monkeypatch.setattr(layout_mod.weakref, "ref", no_weakref)
img = _solid(10, 10)
a = ctx.fit_image(img, (20, 20))
assert ctx.fit_image(img, (20, 20)) is a
assert next(iter(ctx._image_cache.values()))[1]() is img
def test_cache_key_entries_do_not_pin(self, ctx): def test_cache_key_entries_do_not_pin(self, ctx):
img = _solid(10, 10) img = _solid(10, 10)
@@ -0,0 +1,104 @@
"""POST /plugins/install says which id the plugin was installed as.
A registry entry can install under another id: `weather` (aliases
`ledmatrix-weather`) installs a directory whose manifest declares
`ledmatrix-weather`, and that is the id the plugin list, the plugin's config
section and /plugins/toggle know it by. The store's Install button enabled
the new plugin by the registry id, which /plugins/toggle answered with 404
"Plugin not found", so Weather, Music, Stocks and Leaderboard installed
disabled behind an "enabling it failed" warning.
The answer -- the queued operation's result, or the direct response --
carries `plugin_id`: the id the installed manifest declares, found the way
the store's update and uninstall find an install.
"""
import json
from unittest.mock import MagicMock
import pytest
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401
INSTALL = "/api/v3/plugins/install"
@pytest.fixture
def store(api_v3_module, tmp_path):
manager = api_v3_module.api_v3.plugin_store_manager
manager.install_plugin.return_value = True
manager.get_registry_info.return_value = None
manager._find_plugin_path.return_value = None
def installed_as(directory, manifest):
path = tmp_path / directory
path.mkdir()
(path / "manifest.json").write_text(json.dumps(manifest), encoding="utf-8")
manager._find_plugin_path.side_effect = (
lambda pid: path if pid == "weather" else None)
return path
manager.installed_as = installed_as
return manager
@pytest.fixture
def queued(api_v3_module):
queue = MagicMock()
def enqueue(operation_type, plugin_id, operation_callback=None):
queue.callback_result = operation_callback(MagicMock())
return "op-1"
queue.enqueue_operation.side_effect = enqueue
api_v3_module.api_v3.operation_queue = queue
return queue
class TestDirectInstall:
def test_an_aliased_entry_reports_the_manifest_id(self, api_v3_client, store):
store.installed_as("ledmatrix-weather", {"id": "ledmatrix-weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["status"] == "success"
assert body["plugin_id"] == "ledmatrix-weather"
store._find_plugin_path.assert_called_with("weather")
def test_an_entry_installed_under_its_own_id_reports_that(self, api_v3_client, store):
store.installed_as("weather", {"id": "weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["plugin_id"] == "weather"
def test_an_install_that_cannot_be_found_reports_the_requested_id(self, api_v3_client, store):
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["status"] == "success"
assert body["plugin_id"] == "weather"
def test_a_manifest_id_that_is_not_a_plain_name_is_not_passed_on(self, api_v3_client, store):
store.installed_as("ledmatrix-weather", {"id": "../elsewhere"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["plugin_id"] == "weather"
def test_an_unreadable_manifest_reports_the_requested_id(self, api_v3_client, store):
path = store.installed_as("ledmatrix-weather", {})
(path / "manifest.json").write_text("[not json", encoding="utf-8")
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["plugin_id"] == "weather"
def test_the_restart_fields_are_still_sent(self, api_v3_client, store):
store.installed_as("ledmatrix-weather", {"id": "ledmatrix-weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert "restart_required" in body
class TestQueuedInstall:
def test_the_operation_result_names_the_manifest_id(self, api_v3_client, store, queued):
store.installed_as("ledmatrix-weather", {"id": "ledmatrix-weather"})
body = api_v3_client.post(INSTALL, json={"plugin_id": "weather"}).get_json()
assert body["data"]["operation_id"] == "op-1"
assert queued.callback_result["success"] is True
assert queued.callback_result["plugin_id"] == "ledmatrix-weather"
def test_an_install_that_cannot_be_found_names_the_requested_id(
self, api_v3_client, store, queued):
api_v3_client.post(INSTALL, json={"plugin_id": "weather"})
assert queued.callback_result["plugin_id"] == "weather"
@@ -0,0 +1,59 @@
"""GET /api/v3/plugins/installed carries each plugin's ``display_modes``.
The on-demand modal (plugins_manager.js) fills its Display Mode list from
``plugin.display_modes``, but the route never included the field, so every
plugin offered one option -- its own id -- under "This plugin exposes a
single display mode". The display turns that id into the plugin's first
mode, so a multi-mode plugin could only be started, and pinned, on that one.
The modes come from the plugin catalog (the manifests the web process
discovered), the same source /display/modes and on-demand/start use.
"""
from unittest.mock import MagicMock
import pytest
from test._api_v3_test_helpers import ( # noqa: F401 - fixtures
api_v3_client, api_v3_module,
)
@pytest.fixture
def installed(api_v3_module, api_v3_client, tmp_path):
def _get(declared_modes):
api = api_v3_module.api_v3
# The listing's own metadata says nothing about modes: what the
# route reports must come from the catalog.
info = {'id': 'football-scoreboard', 'name': 'Football', 'version': '1.0.0'}
api.plugin_catalog.plugins_dir = str(tmp_path) # no manifest on disk
api.plugin_catalog.get_all_plugin_info = MagicMock(return_value=[info])
api.plugin_catalog.get_plugin_display_modes = MagicMock(return_value=declared_modes)
api.plugin_store_manager.get_registry_info = MagicMock(return_value=None)
api.config_manager.load_config = MagicMock(return_value={})
response = api_v3_client.get('/api/v3/plugins/installed')
assert response.status_code == 200
plugins = [p for p in response.get_json()['data']['plugins']
if p['id'] == 'football-scoreboard']
assert len(plugins) == 1
api.plugin_catalog.get_plugin_display_modes.assert_any_call('football-scoreboard')
return plugins[0]
return _get
def test_every_declared_mode_is_listed_in_order(installed):
modes = ['nfl_live', 'nfl_recent', 'nfl_upcoming']
assert installed(modes)['display_modes'] == modes
def test_a_single_mode_plugin_lists_its_one_mode(installed):
assert installed(['clock-simple'])['display_modes'] == ['clock-simple']
def test_no_declared_modes_is_an_empty_list(installed):
# The modal falls back to the plugin id for an empty list.
assert installed([])['display_modes'] == []
def test_a_hand_edited_manifest_cannot_put_non_strings_in_the_list(installed):
assert installed(['nfl_live', 7, None, {'x': 1}])['display_modes'] == ['nfl_live']
+97
View File
@@ -37,6 +37,7 @@ from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401
START_URL = "/api/v3/display/on-demand/start" START_URL = "/api/v3/display/on-demand/start"
STOP_URL = "/api/v3/display/on-demand/stop" STOP_URL = "/api/v3/display/on-demand/stop"
MAILBOX = "display_on_demand_request" MAILBOX = "display_on_demand_request"
DISPLAY = "web_interface.blueprints.api_v3.display"
@pytest.fixture @pytest.fixture
@@ -150,6 +151,102 @@ class TestStartWhileTheServiceIsStopped:
assert response.get_json()["status"] == "error" assert response.get_json()["status"] == "error"
class _Mailbox:
"""The CacheManager calls the routes make, over a dict."""
def __init__(self):
self.entries = {}
def set(self, key, value, ttl=None):
self.entries[key] = value
def get(self, key, max_age=300, memory_ttl=None):
return self.entries.get(key)
def delete(self, key):
self.entries.pop(key, None)
class TestARefusedStartLeavesNoRequestBehind:
"""A start the route answers with an error must not run later.
The request was posted (to the mailbox, with the display stopped) before
the route refused it, and the display reads the mailbox for an hour
without looking at a request's age. So "Display service is not running"
(start_service off) or "Failed to start display service" left the
request waiting, and the next time the display started -- minutes later,
by hand -- it ran that plugin, pinned if the request said so.
A socket acknowledgement is the other side of it: the display answered,
so it is running and has the request, whatever systemd says (a display
run by hand or in the emulator has no active unit). That is a success,
not "not running", and no unit is started beside it.
"""
@pytest.fixture
def mailbox(self, api_v3_module, service):
box = _Mailbox()
api_v3_module.api_v3.cache_manager = box
service["state"]["active"] = False
return box
@pytest.mark.parametrize("body", [
{"plugin_id": "weather", "start_service": False},
{"plugin_id": "weather"}, # start_service defaults on
])
def test_a_socket_ack_is_a_success_whatever_systemd_says(
self, api_v3_client, service, mailbox, body):
with patch(f"{DISPLAY}.control_client.on_demand_start",
side_effect=lambda request_id, *a: {"accepted": True}):
response = api_v3_client.post(START_URL, json=body)
assert response.status_code == 200, response.get_json()
assert response.get_json()["data"]["transport"] == "socket"
assert MAILBOX not in mailbox.entries
assert _systemctl_verbs(service["systemctl"]) == [], (
"a unit was started beside a display that answered the socket")
def test_without_start_service_the_request_is_taken_back(
self, api_v3_client, service, mailbox):
response = api_v3_client.post(START_URL, json={
"plugin_id": "weather", "pinned": True, "start_service": False})
assert response.status_code == 400
assert response.get_json()["status"] == "error"
assert MAILBOX not in mailbox.entries
def test_a_start_that_fails_takes_its_request_back(self, api_v3_client, service, mailbox):
service["systemctl"].side_effect = lambda args: {
"returncode": 1, "stdout": "", "stderr": "denied"}
response = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert response.status_code == 500
assert MAILBOX not in mailbox.entries
def test_a_newer_request_is_left_alone_on_the_400(self, api_v3_client, service, mailbox):
newer = {"request_id": "someone-else", "action": "start", "plugin_id": "clock"}
def stopped_and_another_post_lands(*args):
mailbox.entries[MAILBOX] = newer
return {"active": False}
with patch(f"{DISPLAY}._get_display_service_status",
side_effect=stopped_and_another_post_lands):
response = api_v3_client.post(START_URL, json={
"plugin_id": "weather", "start_service": False})
assert response.status_code == 400
assert mailbox.entries[MAILBOX] is newer
def test_a_newer_request_is_left_alone_on_the_500(self, api_v3_client, service, mailbox):
newer = {"request_id": "someone-else", "action": "start", "plugin_id": "clock"}
def start_fails_after_another_post(args):
mailbox.entries[MAILBOX] = newer
return {"returncode": 1, "stdout": "", "stderr": "denied"}
service["systemctl"].side_effect = start_fails_after_another_post
response = api_v3_client.post(START_URL, json={"plugin_id": "weather"})
assert response.status_code == 500
assert mailbox.entries[MAILBOX] is newer
class TestStop: class TestStop:
def test_stop_posts_a_stop_request_and_leaves_the_service_running( def test_stop_posts_a_stop_request_and_leaves_the_service_running(
self, api_v3_client, service): self, api_v3_client, service):
@@ -0,0 +1,83 @@
"""GET /api/v3/plugins/operation/<id> answers for an operation still waiting.
PluginOperationQueue keeps an operation's callback in its parameters, under
``_callback``, until the worker takes it to run. PluginOperation.to_dict()
returned the parameters as they were, so for a pending operation the route
handed jsonify a function and answered 500 "A system error occurred". That
is every poll of an install queued behind another plugin's: the second of
two installs read as broken until the first one finished.
"""
import json
import sys
import threading
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent))
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
from src.plugin_system.operation_queue import PluginOperationQueue # noqa: E402
from src.plugin_system.operation_types import ( # noqa: E402
OperationType, PluginOperation,
)
def _callback(op):
return {"success": True, "message": "done"}
class TestToDict:
def test_private_parameters_are_left_out(self):
op = PluginOperation(OperationType.INSTALL, "demo",
parameters={"_callback": _callback, "branch": "main"})
assert op.to_dict()["parameters"] == {"branch": "main"}
json.dumps(op.to_dict()) # serializable
def test_the_operation_keeps_its_callback_for_the_worker(self):
op = PluginOperation(OperationType.INSTALL, "demo",
parameters={"_callback": _callback})
op.to_dict()
assert op.parameters["_callback"] is _callback
def test_the_other_fields_are_unchanged(self):
op = PluginOperation(OperationType.UNINSTALL, "demo", operation_id="op-1")
assert op.to_dict() == {
"operation_id": "op-1", "operation_type": "uninstall", "plugin_id": "demo",
"parameters": {}, "status": "pending", "progress": 0.0, "message": "",
"error": None, "result": None,
"created_at": op.created_at.isoformat(), "started_at": None,
"completed_at": None,
}
class TestTheRoute:
@pytest.fixture
def busy_queue(self, api_v3_module):
"""A real queue whose worker is held by another plugin's operation."""
queue = PluginOperationQueue(max_history=10)
api_v3_module.api_v3.operation_queue = queue
started, release = threading.Event(), threading.Event()
def blocker(op):
started.set()
release.wait(10)
return {"success": True, "message": "done"}
queue.enqueue_operation(OperationType.INSTALL, "busy", operation_callback=blocker)
assert started.wait(5)
yield queue
release.set()
queue.shutdown()
def test_a_pending_operation_reports_pending(self, api_v3_client, busy_queue):
op_id = busy_queue.enqueue_operation(
OperationType.INSTALL, "demo", operation_callback=_callback)
response = api_v3_client.get(f"/api/v3/plugins/operation/{op_id}")
assert response.status_code == 200, response.get_json()
data = response.get_json()["data"]
assert data["status"] == "pending"
assert data["plugin_id"] == "demo"
assert "_callback" not in data["parameters"]
+26
View File
@@ -253,6 +253,32 @@ class TestVegasCycleDurations:
assert saved['config']['display']['display_durations'] == {'clock': 45} assert saved['config']['display']['display_durations'] == {'clock': 45}
class TestMalformedBody:
"""A JSON body that does not parse is the caller's mistake: a 400.
get_json() raised Werkzeug's BadRequest inside the handler's try, whose
catch-all answered 500 CONFIG_SAVE_FAILED with "check file permissions"
advice and logged a traceback at ERROR.
"""
def test_is_a_400_in_the_raw_routes_shape(self, api_v3_client, saved, api_v3_module):
api_v3_module.api_v3.config_manager.get_raw_file_content.return_value = {}
resp = api_v3_client.post('/api/v3/config/main', data='{not json',
content_type='application/json')
assert resp.status_code == 400
assert resp.get_json() == {'status': 'error', 'message': 'Invalid JSON in request body'}
assert 'config' not in saved
raw = api_v3_client.post('/api/v3/config/raw/main', data='{not json',
content_type='application/json')
assert (raw.status_code, raw.get_json()) == (400, resp.get_json())
def test_an_empty_json_post_is_still_no_data(self, api_v3_client, saved):
resp = api_v3_client.post('/api/v3/config/main', data='',
content_type='application/json')
assert resp.status_code == 400
assert resp.get_json()['message'] == 'No data provided'
class TestRawSaveStartsAutoUpdateSetup: class TestRawSaveStartsAutoUpdateSetup:
@pytest.fixture @pytest.fixture
def raw_env(self, api_v3_module, monkeypatch): def raw_env(self, api_v3_module, monkeypatch):
+93
View File
@@ -0,0 +1,93 @@
"""POST /api/v3/plugins/action hands ``params`` to the plugin's script intact.
The route runs the script through a generated wrapper, and the params went
into that wrapper as Python source: ``params = {json.dumps(params)}``. JSON is
not Python. ``true``, ``false`` and ``null`` are undefined names there, so any
params holding a boolean or a null died with a NameError before the script
ran. The plugin file manager's category toggle sends ``{"category_name": ...,
"enabled": true}``, so of-the-day's category toggle failed every time with
"Action failed".
The script's side of the contract is unchanged and pinned here too: the
params arrive on stdin as one JSON document, LEDMATRIX_ROOT is set, and what
the script prints to stdout is what the route parses.
"""
import json
import subprocess
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent))
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
ACTION_URL = "/api/v3/plugins/action"
# The action script: report what it was handed, as JSON on stdout.
ECHO_SCRIPT = (
"import json, os, sys\n"
"raw = sys.stdin.read()\n"
"print(json.dumps({'status': 'success', 'got': json.loads(raw),\n"
" 'root': os.environ.get('LEDMATRIX_ROOT')}))\n"
)
@pytest.fixture
def echo_plugin(tmp_path, api_v3_module, monkeypatch):
plugin_dir = tmp_path / "demo"
plugin_dir.mkdir()
(plugin_dir / "manifest.json").write_text(json.dumps({
"id": "demo",
"web_ui_actions": [{"id": "toggle", "type": "script", "script": "echo.py"}],
}), encoding="utf-8")
(plugin_dir / "echo.py").write_text(ECHO_SCRIPT, encoding="utf-8")
api_v3_module.api_v3.plugin_catalog.get_plugin_directory.return_value = str(plugin_dir)
# The route runs `python3`; use this interpreter, so the test does not
# depend on what that name resolves to here.
real_run = subprocess.run
def run(cmd, *args, **kwargs):
if isinstance(cmd, list) and cmd and cmd[0] == "python3":
cmd = [sys.executable] + cmd[1:]
return real_run(cmd, *args, **kwargs)
monkeypatch.setattr(subprocess, "run", run)
return plugin_dir
@pytest.mark.parametrize("params", [
{"category_name": "jokes", "enabled": True}, # the file manager's toggle
{"category_name": "jokes", "enabled": False},
{"filename": None},
{"nested": {"list": [1, None, True, 2.5], "empty": {}}},
{"text": "café ✓ \U0001F600"},
{"text": "he said \"hi\" and 'bye' \\ ''' \"\"\" \n\t end"},
], ids=["true", "false", "null", "nested", "unicode", "quotes"])
def test_the_script_receives_the_params_it_was_sent(api_v3_client, echo_plugin, params):
response = api_v3_client.post(ACTION_URL, json={
"plugin_id": "demo", "action_id": "toggle", "params": params})
body = response.get_json()
assert response.status_code == 200, body
assert body["got"] == params
def test_a_param_cannot_run_code_in_the_wrapper(api_v3_client, echo_plugin, tmp_path):
marker = tmp_path / "PWNED"
hostile = "\"}\nopen(%r, 'w').write('ran')\n#" % str(marker)
params = {"name": hostile, "flag": True}
response = api_v3_client.post(ACTION_URL, json={
"plugin_id": "demo", "action_id": "toggle", "params": params})
assert response.status_code == 200, response.get_json()
assert response.get_json()["got"] == params
assert not marker.exists(), "a param value ran as code"
def test_the_script_still_gets_ledmatrix_root(api_v3_client, echo_plugin, api_v3_module):
response = api_v3_client.post(ACTION_URL, json={
"plugin_id": "demo", "action_id": "toggle", "params": {"enabled": True}})
assert response.status_code == 200, response.get_json()
assert response.get_json()["root"] == str(api_v3_module.PROJECT_ROOT)
@@ -0,0 +1,89 @@
"""A second install or uninstall while one is in progress is a 409, not a 500.
PluginOperationQueue refuses a second operation for a plugin that already
has one waiting or running (test_operation_queue_pending_and_trim.py), and
says so by raising ValueError. /plugins/install let that escape to the
blueprint's catch-all, so a double-clicked Install answered 500 "An error
occurred; see logs for details" while the first install carried on.
/plugins/uninstall caught it in its own catch-all: a 500 "Failed to
uninstall plugin", and an "uninstall failed" entry in the operation
history for an uninstall that never started.
"""
import sys
import threading
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent))
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
from src.plugin_system.operation_queue import PluginOperationQueue # noqa: E402
INSTALL = "/api/v3/plugins/install"
UNINSTALL = "/api/v3/plugins/uninstall"
@pytest.fixture
def installing(api_v3_module, tmp_path):
"""A real queue with an install of "clock" running and held there."""
queue = PluginOperationQueue(max_history=10)
api_v3_module.api_v3.operation_queue = queue
started, release = threading.Event(), threading.Event()
def slow_install(plugin_id, branch=None):
started.set()
release.wait(10)
return True
store = api_v3_module.api_v3.plugin_store_manager
store.install_plugin.side_effect = slow_install
store.get_registry_info.return_value = None
store.plugins_dir = str(tmp_path)
api_v3_module.api_v3.plugin_catalog.get_plugin_directory.return_value = None
yield {"queue": queue, "started": started, "store": store}
release.set()
queue.shutdown()
def _start_first_install(client, installing):
response = client.post(INSTALL, json={"plugin_id": "clock"})
assert response.status_code == 200, response.get_json()
assert installing["started"].wait(5)
def _failed_history(api_v3_module):
return [c for c in api_v3_module.api_v3.operation_history.record_operation.call_args_list
if c.kwargs.get("status") == "failed"]
def test_a_second_install_click_is_a_conflict(api_v3_client, api_v3_module, installing):
_start_first_install(api_v3_client, installing)
response = api_v3_client.post(INSTALL, json={"plugin_id": "clock"})
assert response.status_code == 409, response.get_json()
body = response.get_json()
assert body["status"] == "error"
assert body["error_code"] == "PLUGIN_OPERATION_CONFLICT"
assert "clock" in body["message"]
assert installing["store"].install_plugin.call_count == 1
assert _failed_history(api_v3_module) == []
def test_an_uninstall_during_the_install_is_a_conflict(api_v3_client, api_v3_module,
installing):
_start_first_install(api_v3_client, installing)
response = api_v3_client.post(UNINSTALL, json={"plugin_id": "clock"})
assert response.status_code == 409, response.get_json()
assert response.get_json()["error_code"] == "PLUGIN_OPERATION_CONFLICT"
assert _failed_history(api_v3_module) == [], (
"an uninstall that never started was recorded as failed")
api_v3_module.api_v3.plugin_store_manager.uninstall_plugin.assert_not_called()
def test_another_plugin_is_still_queued(api_v3_client, installing):
_start_first_install(api_v3_client, installing)
response = api_v3_client.post(INSTALL, json={"plugin_id": "weather"})
assert response.status_code == 200, response.get_json()
assert response.get_json()["data"]["operation_id"]
+73
View File
@@ -0,0 +1,73 @@
"""GET /api/v3/plugins/<plugin_id>/static/<path> serves binary files too.
The route opened every file as UTF-8 text, so an image -- what the API
reference says it is for, plugin previews and icons -- failed to decode and
answered 500 "UnicodeDecodeError". Files are now sent as bytes. The text
types the route always set are unchanged, and the path checks are pinned in
test_path_traversal_guards.py::TestServePluginStatic.
"""
import json
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent))
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
PNG = (b"\x89PNG\r\n\x1a\n\x00\x00\x00\rIHDR\x00\x00\x00\x01\x00\x00\x00\x01"
b"\x08\x06\x00\x00\x00\x1f\x15\xc4\x89")
@pytest.fixture
def plugin_dir(tmp_path, api_v3_module):
d = tmp_path / "demo"
(d / "web_ui").mkdir(parents=True)
(d / "manifest.json").write_text(json.dumps({"id": "demo"}), encoding="utf-8")
api_v3_module.api_v3.plugin_catalog.get_plugin_directory.side_effect = (
lambda pid: str(d) if pid == "demo" else None)
return d
def _get(client, path):
return client.get(f"/api/v3/plugins/demo/static/{path}")
def test_an_image_is_served_as_its_bytes(api_v3_client, plugin_dir):
(plugin_dir / "web_ui" / "icon.png").write_bytes(PNG)
response = _get(api_v3_client, "web_ui/icon.png")
assert response.status_code == 200, response.get_json(silent=True)
assert response.mimetype == "image/png"
assert response.data == PNG
def test_an_unknown_binary_file_is_served_too(api_v3_client, plugin_dir):
blob = bytes(range(256))
(plugin_dir / "data.bin").write_bytes(blob)
response = _get(api_v3_client, "data.bin")
assert response.status_code == 200, response.get_json(silent=True)
assert response.data == blob
@pytest.mark.parametrize("name,mimetype", [
("page.html", "text/html"),
("app.js", "application/javascript"),
("style.css", "text/css"),
("data.json", "application/json"),
("notes.txt", "text/plain"),
("README.md", "text/plain"),
("helper.py", "text/plain"),
])
def test_text_files_keep_their_types(api_v3_client, plugin_dir, name, mimetype):
content = "caf\u00e9 \u2713 <p>hi</p>\n"
(plugin_dir / name).write_bytes(content.encode("utf-8"))
response = _get(api_v3_client, name)
assert response.status_code == 200
assert response.mimetype == mimetype
assert response.data == content.encode("utf-8")
def test_a_missing_file_is_still_a_404(api_v3_client, plugin_dir):
assert _get(api_v3_client, "nope.png").status_code == 404
+326
View File
@@ -0,0 +1,326 @@
"""Four web answers that disagreed with the rig they describe (found on ledpi).
1. POST /config/schedule refused the schedule GET returns on a fresh install
(config.template.json: per-day, every day off, schedule disabled) with
"At least one day must be enabled", as did /config/dim-schedule. A
disabled schedule needs no enabled day.
2. A brightness-only POST /config/main answered ``restart_required: true``,
though the display applies brightness live (brightness.set over the
socket, and the config watcher). The flag now says whether anything
changed that the running display does not pick up by itself.
3. /health stayed "healthy" with the display service stopped: only the
sub-checks changed. Service inactive, no socket and no live heartbeat
is now ``display_loop: stopped`` and "degraded".
4. /display/current-status kept answering ``is_display_active: true`` from
the cache for up to 120 s after the display stopped. With no socket and
no live heartbeat it is now unknown.
"""
import copy
import json
import os
import sys
import time
from pathlib import Path
from unittest.mock import patch
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent))
from test._api_v3_test_helpers import api_v3_client, api_v3_module # noqa: F401,E402
from src import display_watchdog # noqa: E402
from src.ipc import client as control_client # noqa: E402
from web_interface import display_state # noqa: E402
REPO = Path(__file__).resolve().parent.parent
TEMPLATE = json.loads((REPO / 'config' / 'config.template.json').read_text(encoding='utf-8'))
@pytest.fixture
def store(api_v3_module, monkeypatch):
state = {'config': {}, 'saves': 0}
api_v3_module.api_v3.config_manager.load_config.side_effect = \
lambda *a, **k: copy.deepcopy(state['config'])
def fake_save(_manager, config, **_kwargs):
state['config'] = copy.deepcopy(config)
state['saves'] += 1
return True, ''
monkeypatch.setattr(api_v3_module, '_save_config_atomic', fake_save)
return state
# --- 1. schedules ---------------------------------------------------------------
SCHEDULE_ROUTES = [('/api/v3/config/schedule', 'schedule'),
('/api/v3/config/dim-schedule', 'dim_schedule')]
@pytest.mark.parametrize('route,section', SCHEDULE_ROUTES)
def test_the_templates_disabled_per_day_schedule_saves_back(api_v3_client, store,
route, section):
stored = copy.deepcopy(TEMPLATE[section])
stored['mode'] = 'per-day'
assert stored['enabled'] is False
assert not any(day['enabled'] for day in stored['days'].values())
store['config'] = {section: copy.deepcopy(stored)}
read = api_v3_client.get(route).get_json()['data']
resp = api_v3_client.post(route, json=read)
assert resp.status_code == 200, resp.get_json()
saved = store['config'][section]
assert saved['enabled'] is False and saved['mode'] == 'per-day'
# The disabled days keep their times: switching one on finds them.
assert saved['days'] == stored['days']
@pytest.mark.parametrize('route,section', SCHEDULE_ROUTES)
def test_an_enabled_per_day_schedule_still_needs_a_day(api_v3_client, store, route, section):
body = copy.deepcopy(TEMPLATE[section])
body.update(enabled=True, mode='per-day')
resp = api_v3_client.post(route, json=body)
assert resp.status_code == 400
assert 'At least one day must be enabled' in resp.get_json()['message']
assert store['saves'] == 0
@pytest.mark.parametrize('route', [r for r, _ in SCHEDULE_ROUTES])
def test_the_pickers_form_post_with_every_day_off_saves(api_v3_client, store, route):
"""What schedule-picker.js posts: flat hidden inputs, booleans as strings,
times for every day."""
body = {'enabled': 'false', 'mode': 'per_day', 'start_time': '07:00', 'end_time': '23:00'}
for day in ('monday', 'tuesday', 'wednesday', 'thursday', 'friday', 'saturday', 'sunday'):
body.update({f'{day}_enabled': 'false', f'{day}_start': '06:30', f'{day}_end': '22:15'})
resp = api_v3_client.post(route, json=body)
assert resp.status_code == 200, resp.get_json()
def test_an_invalid_time_on_a_disabled_day_is_dropped_not_refused(api_v3_client, store):
body = {'enabled': False, 'mode': 'per-day',
'days': {'monday': {'enabled': False, 'start_time': 'soon', 'end_time': '22:00'}}}
resp = api_v3_client.post('/api/v3/config/schedule', json=body)
assert resp.status_code == 200, resp.get_json()
assert store['config']['schedule']['days']['monday'] == {'enabled': False,
'end_time': '22:00'}
# --- 2. restart_required on /config/main ------------------------------------------
STORED_MAIN = {
'timezone': 'America/Chicago',
'display': {
'hardware': {'rows': 32, 'cols': 64, 'chain_length': 2, 'brightness': 90,
'disable_hardware_pulsing': False, 'inverse_colors': False,
'show_refresh_rate': False},
'runtime': {'gpio_slowdown': 4},
'display_durations': {'clock': 15},
'use_short_date_format': False,
},
}
@pytest.fixture
def main_store(store):
store['config'] = copy.deepcopy(STORED_MAIN)
return store
def _save_main(client, body):
with patch('web_interface.blueprints.api_v3.control_client.brightness_set',
side_effect=control_client.ControlError('no_socket', 'x')):
resp = client.post('/api/v3/config/main', data=json.dumps(body),
content_type='application/json')
assert resp.status_code == 200, resp.get_json()
return resp.get_json()
def test_a_brightness_only_save_needs_no_restart(api_v3_client, main_store):
body = _save_main(api_v3_client, {'brightness': 40})
assert main_store['config']['display']['hardware']['brightness'] == 40
assert body['restart_required'] is False
def test_a_brightness_save_on_a_config_without_a_display_section(api_v3_client, store):
"""The route creates display.hardware and display.runtime on the way;
empty sections are not a change."""
store['config'] = {}
assert _save_main(api_v3_client, {'brightness': 40})['restart_required'] is False
def test_the_display_form_with_only_brightness_changed_needs_no_restart(api_v3_client,
main_store):
hw = STORED_MAIN['display']['hardware']
body = {'__form_section': 'display', 'rows': 32, 'cols': 64, 'chain_length': 2,
'brightness': 55, 'gpio_slowdown': 4}
body.update({k: 'on' for k in ('disable_hardware_pulsing', 'inverse_colors',
'show_refresh_rate') if hw[k]})
assert _save_main(api_v3_client, body)['restart_required'] is False
def test_a_mode_duration_needs_no_restart(api_v3_client, main_store):
body = _save_main(api_v3_client, {'duration__clock': 40})
assert main_store['config']['display']['display_durations']['clock'] == 40
assert body['restart_required'] is False
@pytest.mark.parametrize('change', [{'rows': 64}, {'brightness': 40, 'chain_length': 3},
{'gpio_slowdown': 2}, {'timezone': 'UTC'}])
def test_a_setting_the_display_reads_at_startup_still_needs_one(api_v3_client, main_store,
change):
assert _save_main(api_v3_client, change)['restart_required'] is True
def test_restart_needed_compares_leaves():
from web_interface.blueprints.api_v3.config import restart_needed
before = {'display': {'hardware': {'brightness': 90, 'rows': 32}}}
assert not restart_needed(before, copy.deepcopy(before))
assert not restart_needed(before, {'display': {'hardware': {'brightness': 10, 'rows': 32},
'runtime': {}}})
assert restart_needed(before, {'display': {'hardware': {'brightness': 90}}}) # removed
assert not restart_needed({}, {'clock': {'enabled': True}}, live_paths=[('clock',)])
assert restart_needed({}, {'clockwork': {'enabled': True}}, live_paths=[('clock',)])
# --- 3 and 4. a stopped display -----------------------------------------------------
@pytest.fixture
def no_display(monkeypatch, tmp_path):
"""A Pi whose display service has stopped: the socket is expected here
but does not answer, and systemd took the heartbeat's directory away."""
monkeypatch.setattr(display_state, 'socket_supported', lambda: True)
monkeypatch.setattr(display_state, 'client_socket_paths', lambda: [str(tmp_path / 'gone')])
monkeypatch.setattr(display_state, 'read_state', lambda: None)
path = tmp_path / 'display-heartbeat.json'
monkeypatch.setattr(display_watchdog, 'HEARTBEAT_PATH', str(path))
def beat(age, pid=None):
path.write_text(json.dumps({'pid': os.getpid() if pid is None else pid,
'mono': time.monotonic() - age,
'wall': time.time() - age}))
return beat
@pytest.fixture
def service(monkeypatch):
status = {'active': False, 'returncode': 3, 'stdout': 'inactive', 'stderr': ''}
monkeypatch.setattr('web_interface.blueprints.api_v3.misc._get_display_service_status',
lambda: dict(status))
return status
@pytest.fixture
def fresh_preview(tmp_path, monkeypatch):
"""The preview frame the display left behind, under 60 s old: on its own
it kept the hardware check "connected"."""
from web_interface import display_preview
snapshot = tmp_path / 'preview.png'
snapshot.write_bytes(b'png')
monkeypatch.setattr(display_preview, 'SNAPSHOT_PATH', str(snapshot))
def _health(client):
resp = client.get('/api/v3/health')
assert resp.status_code == 200, resp.get_json()
return resp.get_json()['data']
class TestHealth:
def test_a_stopped_display_service_is_degraded(self, api_v3_client, no_display, service,
fresh_preview):
data = _health(api_v3_client)
assert data['services']['display_service']['status'] == 'inactive'
assert data['checks']['display_loop']['status'] == 'stopped'
assert data['status'] == 'degraded'
def test_a_service_still_starting_is_not(self, api_v3_client, no_display, service,
fresh_preview):
"""Active, before its socket and first heartbeat: not stopped."""
service.update(active=True, stdout='active', returncode=0)
data = _health(api_v3_client)
assert data['checks']['display_loop']['status'] == 'not_reported'
assert data['status'] == 'healthy'
def test_a_display_run_by_hand_is_not_stopped(self, api_v3_client, no_display, service,
fresh_preview):
"""The service is off but a display process beats (sudo python3 run.py)."""
no_display(age=2)
data = _health(api_v3_client)
assert data['checks']['display_loop']['status'] == 'running'
assert data['status'] == 'healthy'
@pytest.mark.parametrize('platform', ['no_unix_sockets', 'socket_off'])
def test_without_a_socket_to_expect_nothing_changes(self, api_v3_client, no_display,
service, fresh_preview, monkeypatch,
platform):
"""Windows and the dev server (no systemd unit), or the socket
deliberately off: no heartbeat is no signal, as before."""
if platform == 'no_unix_sockets':
monkeypatch.setattr(display_state, 'socket_supported', lambda: False)
else:
monkeypatch.setattr(display_state, 'client_socket_paths', lambda: [])
service.update(returncode=-1, stdout='', stderr='systemctl not found')
data = _health(api_v3_client)
assert data['checks']['display_loop']['status'] == 'not_reported'
assert data['status'] == 'healthy'
def test_the_status_only_answer_says_degraded(self, api_v3_client, no_display, service,
fresh_preview, monkeypatch):
monkeypatch.setattr('web_interface.blueprints.api_v3.misc.request_is_authenticated',
lambda: False)
resp = api_v3_client.get('/api/v3/health')
assert resp.get_json()['data'] == {'status': 'degraded'}
class TestCurrentStatus:
CACHED = {'mode': 'clock', 'plugin_id': 'clock', 'is_display_active': True,
'on_demand_active': False, 'last_updated': None}
@pytest.fixture
def cached(self, api_v3_module):
entry = dict(self.CACHED, last_updated=time.time() - 30)
cache = api_v3_module.api_v3.cache_manager
cache.get.side_effect = lambda key, *a, **kw: (
dict(entry) if key == 'display_current_state' else None)
return entry
def _status(self, client):
resp = client.get('/api/v3/display/current-status')
assert resp.status_code == 200
return resp.get_json()['data']
def test_a_stopped_display_is_not_reported_active(self, api_v3_client, no_display, cached):
data = self._status(api_v3_client)
assert not data.get('is_display_active')
assert data['mode'] is None and data['last_updated'] is None
assert data['source'] == 'cache'
def test_a_stale_heartbeat_is_not_active_either(self, api_v3_client, no_display, cached):
no_display(age=display_watchdog.HEARTBEAT_STALE_SECONDS + 5)
assert self._status(api_v3_client)['mode'] is None
@pytest.mark.skipif(os.name != 'posix', reason='process_exists answers only on POSIX')
def test_a_heartbeat_from_a_dead_process_is_not_active(self, api_v3_client, no_display,
cached):
no_display(age=1, pid=2 ** 22 + 12345)
assert self._status(api_v3_client)['mode'] is None
def test_a_live_heartbeat_without_a_socket_reads_the_cache(self, api_v3_client,
no_display, cached):
"""An older display with no socket, still running."""
no_display(age=2)
data = self._status(api_v3_client)
assert data['mode'] == 'clock' and data['is_display_active'] is True
@pytest.mark.parametrize('platform', ['no_unix_sockets', 'socket_off'])
def test_without_a_socket_to_expect_the_cache_answers(self, api_v3_client, no_display,
cached, monkeypatch, platform):
if platform == 'no_unix_sockets':
monkeypatch.setattr(display_state, 'socket_supported', lambda: False)
else:
monkeypatch.setattr(display_state, 'client_socket_paths', lambda: [])
data = self._status(api_v3_client)
assert data['mode'] == 'clock' and data['is_display_active'] is True
+27
View File
@@ -13,6 +13,7 @@ The invariants that keep this change safe:
inline path exactly. inline path exactly.
""" """
import asyncio
import os import os
import sys import sys
import threading import threading
@@ -209,6 +210,32 @@ class TestFailurePaths:
assert pm.get_plugin_lock(plugin_id).acquire(blocking=False) is True assert pm.get_plugin_lock(plugin_id).acquire(blocking=False) is True
pm.get_plugin_lock(plugin_id).release() pm.get_plugin_lock(plugin_id).release()
@pytest.mark.parametrize("raised", [asyncio.CancelledError, SystemExit])
def test_update_raising_a_base_exception_still_releases_the_plugin(self, pm, raised):
"""asyncio.CancelledError and SystemExit derive from BaseException,
not Exception. Raised from update() on the worker, one skipped the
bookkeeping entirely: the plugin kept its lock and stayed RUNNING for
the life of the process -- never updated again, and every display()
skipped as busy."""
class CancellingPlugin(SlowPlugin):
def update(self):
self.update_calls += 1
raise raised()
plugin_id = _install(pm, CancellingPlugin())
pm.run_scheduled_updates()
deadline = time.monotonic() + 3
while pm.plugins[plugin_id].update_calls == 0 and time.monotonic() < deadline:
time.sleep(0.05)
time.sleep(0.2)
assert pm.get_plugin_lock(plugin_id).acquire(blocking=False) is True
pm.get_plugin_lock(plugin_id).release()
assert pm.state_manager.can_execute(plugin_id) is True
assert pm.plugin_last_update.get(plugin_id, 0) > 0
error = pm.state_manager.get_error_info(plugin_id)
assert error is not None and error["error_type"] == raised.__name__
def test_unloaded_while_queued_is_harmless(self, pm): def test_unloaded_while_queued_is_harmless(self, pm):
"""Exercise the public unload_plugin() lifecycle rather than """Exercise the public unload_plugin() lifecycle rather than
deleting pm.plugins directly: queue the target's update behind a deleting pm.plugins directly: queue the target's update behind a
+26
View File
@@ -517,6 +517,32 @@ class TestUpdateIsVerified:
h.updater.run() h.updater.run()
assert h.pending['dependency_failures'] == ['requirements.txt'] assert h.pending['dependency_failures'] == ['requirements.txt']
@pytest.mark.parametrize('unit_refresh, expected', [
({'status': 'refreshed', 'message': '', 'units': ['ledmatrix.service']}, True),
({'status': 'needs_reinstall', 'message': '', 'units': ['ledmatrix.service']}, False),
(None, False),
])
def test_the_health_check_learns_whether_the_update_installed_units(self, tmp_path, unit_refresh,
expected):
"""Its rollback restores the previous units only when this update replaced them."""
repo = Repo(tmp_path)
repo.publish()
h = Harness(tmp_path, repo, core_update=real_pull(repo.device, unit_refresh=unit_refresh))
h.updater.run()
assert h.pending['units_refreshed'] is expected
def test_a_health_check_that_never_starts_also_restores_the_units(self, tmp_path):
repo = Repo(tmp_path)
old = repo.head()
repo.publish()
refreshed = {'status': 'refreshed', 'message': '', 'units': ['ledmatrix.service']}
h = Harness(tmp_path, repo, pickup=False,
core_update=real_pull(repo.device, unit_refresh=refreshed))
h.updater.run()
assert repo.head() == old
assert [a for a in h.sudo if a[-1] == '--restore'] == [
['sudo', '-n', '/usr/local/sbin/ledmatrix-refresh-units', '--restore']]
def test_a_health_check_that_never_starts_means_the_update_is_undone(self, tmp_path): def test_a_health_check_that_never_starts_means_the_update_is_undone(self, tmp_path):
repo = Repo(tmp_path) repo = Repo(tmp_path)
old = repo.head() old = repo.head()
+48
View File
@@ -80,6 +80,8 @@ class FakeHost:
self.nrestarts = 0 self.nrestarts = 0
self.heartbeat = heartbeat self.heartbeat = heartbeat
self.display_started_at = -1000.0 # the pre-update display, long running self.display_started_at = -1000.0 # the pre-update display, long running
self.unit_restores = [] # (argv, commit checked out, restarts so far)
self.restore_ok = True
def broken(self, kind): def broken(self, kind):
if self.running_head is None: if self.running_head is None:
@@ -103,6 +105,10 @@ class FakeHost:
if self.pip: if self.pip:
return self.pip(args, self) return self.pip(args, self)
return done(args, rc=0 if self.pip_ok else 1) return done(args, rc=0 if self.pip_ok else 1)
if args[:3] == ['sudo', '-n', av.REFRESH_UNITS_PATH]:
self.unit_restores.append((list(args), git(self.repo, 'rev-parse', 'HEAD'),
len(self.restarts)))
return done(args, rc=0 if self.restore_ok else 1)
if args[:4] == ['sudo', '-n', 'systemctl', 'restart']: if args[:4] == ['sudo', '-n', 'systemctl', 'restart']:
if self.restart_failures: if self.restart_failures:
self.restart_failures -= 1 self.restart_failures -= 1
@@ -405,3 +411,45 @@ def test_the_heartbeat_location_and_freshness_match_the_display():
# window it has to stay healthy for. # window it has to stay healthy for.
assert av.HEARTBEAT_FRESH_SECONDS + av.POLL_SECONDS < av.STABLE_SECONDS assert av.HEARTBEAT_FRESH_SECONDS + av.POLL_SECONDS < av.STABLE_SECONDS
assert av.HEARTBEAT_FRESH_SECONDS > display_watchdog.BEAT_INTERVAL_SECONDS * 2 assert av.HEARTBEAT_FRESH_SECONDS > display_watchdog.BEAT_INTERVAL_SECONDS * 2
# -- systemd units the update installed ------------------------------------------
def test_a_rollback_restores_the_units_the_update_installed(tmp_path):
"""The update installed new units (web_interface/unit_refresh.py); the
rollback puts the old ones back before restarting onto the old code."""
code, result, host, head, old, new = check(tmp_path, 'display_down', units_refreshed=True)
assert result['status'] == 'rolled_back' and head == old
assert len(host.unit_restores) == 1
argv, commit, restarts_before = host.unit_restores[0]
assert argv == ['sudo', '-n', av.REFRESH_UNITS_PATH, '--restore']
assert commit == old, 'restored after the code was rolled back'
assert restarts_before == 2, 'restored before the services restart onto the old code'
assert host.restarts[-2:] == [('ledmatrix.service', old), ('ledmatrix-web.service', old)]
assert result['detail'] is None
@pytest.mark.parametrize('pending', [{}, {'units_refreshed': False}])
def test_a_rollback_leaves_units_alone_when_the_update_did_not_change_them(tmp_path, pending):
# {} is what an updater from before this change writes.
code, result, host, head, old, new = check(tmp_path, 'display_down', **pending)
assert result['status'] == 'rolled_back' and host.unit_restores == []
def test_a_healthy_update_keeps_its_new_units(tmp_path):
code, result, host, head, old, new = check(tmp_path, units_refreshed=True)
assert result['status'] == 'success' and host.unit_restores == []
def test_a_failed_unit_restore_is_reported_but_the_rollback_stands(tmp_path):
repo, old, new = updated_repo(tmp_path)
av.write_pending(av.pending_path(repo), {'status': 'pending', 'old_head': old, 'new_head': new,
'display_was_active': True, 'dependency_failures': [],
'units_refreshed': True})
host = FakeHost(repo, new, 'display_down')
host.restore_ok = False
host.verifier().verify()
result = av.read_pending(av.pending_path(repo))
assert result['status'] == 'rolled_back'
assert git(repo, 'rev-parse', 'HEAD') == old
assert 'install_service.sh' in result['detail']
+18
View File
@@ -200,6 +200,24 @@ class TestGetOdds:
manager.get_odds('football', 'nfl', '401') # hit manager.get_odds('football', 'nfl', '401') # hit
assert [r for r in caplog.records if r.levelno == logging.INFO] == [] assert [r for r in caplog.records if r.levelno == logging.INFO] == []
def test_debug_off_does_not_serialize_the_response(
self, manager, mock_get, caplog):
# json.dumps(indent=2) of every odds body ran even with DEBUG off.
with caplog.at_level(logging.INFO, logger=manager.logger.name), \
patch('src.base_odds_manager.json.dumps') as dumps:
assert manager.get_odds('football', 'nfl', '401') == FULL_EXTRACTED
dumps.assert_not_called()
def test_debug_on_still_logs_the_raw_response(
self, manager, mock_get, caplog):
with caplog.at_level(logging.DEBUG, logger=manager.logger.name):
manager.get_odds('football', 'nfl', '401')
messages = [r.getMessage() for r in caplog.records]
assert any(m.startswith('Received raw odds data from ESPN: {')
for m in messages), messages
assert any(m.startswith('Returning extracted odds data: {')
for m in messages), messages
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# _extract_espn_data # _extract_espn_data
+99
View File
@@ -0,0 +1,99 @@
"""A cache key too long to be a filename still gets a cache file.
The calendar plugin's key joins every calendar id the user picked; on a real
install it passed 300 bytes, and since ext4 caps a filename at 255 every write
failed with ENAMETOOLONG -- logged as "permission denied", every update.
"""
import logging
import os
from unittest.mock import patch
from src.cache.disk_cache import DiskCache, _MAX_KEY_FILENAME_BYTES, _filename_stem
from src.cache_manager import CacheManager
# The shape of the key that failed on hdpi, ids anonymised.
CALENDAR_KEY = (
"calendar_events_someone@example.com_en.usa#holiday@group.v.calendar.google.com_"
"family13997378751670666433@group.calendar.google.com_ncaaf_-m-07kbp5_"
"%47eorgia+%42ulldogs+football#sports@group.v.calendar.google.com_nfl_-m-07l24_"
"%54ampa+%42ay+%42uccaneers#sports@group.v.calendar.google.com_primary"
)
# ext4/xfs/btrfs NAME_MAX; set()'s temp file adds 15 bytes to the stem.
NAME_MAX = 255
TEMP_OVERHEAD = len(".") + len(".json") + len(".") + 8
def test_the_real_key_was_too_long_to_write():
assert len((CALENDAR_KEY + ".json").encode()) > NAME_MAX - 10
def test_a_long_key_round_trips(tmp_path):
cache = DiskCache(str(tmp_path))
cache.set(CALENDAR_KEY, {"events": [1, 2, 3]})
assert cache.get(CALENDAR_KEY, max_age=None) == {"events": [1, 2, 3]}
path = cache.get_cache_path(CALENDAR_KEY)
assert os.path.isfile(path)
stem = os.path.basename(path)[:-len(".json")]
assert len(stem.encode()) + TEMP_OVERHEAD <= NAME_MAX
def test_short_keys_keep_their_filename(tmp_path):
cache = DiskCache(str(tmp_path))
exactly = "k" * _MAX_KEY_FILENAME_BYTES
assert cache.get_cache_path("weather_current") == str(tmp_path / "weather_current.json")
assert cache.get_cache_path(exactly) == str(tmp_path / f"{exactly}.json")
assert cache.get_cache_path(exactly + "k") != str(tmp_path / f"{exactly}k.json")
def test_long_keys_sharing_a_prefix_stay_apart(tmp_path):
cache = DiskCache(str(tmp_path))
first, second = CALENDAR_KEY + "_a", CALENDAR_KEY + "_b"
cache.set(first, {"which": "a"})
cache.set(second, {"which": "b"})
assert cache.get_cache_path(first) != cache.get_cache_path(second)
assert cache.get(first, max_age=None) == {"which": "a"}
assert cache.get(second, max_age=None) == {"which": "b"}
def test_the_prefix_never_splits_a_character():
key = "news_" + "é" * 300 # two bytes each, so the cut lands mid-character
stem = _filename_stem(key)
assert stem.startswith("news_é")
assert len(stem.encode("utf-8")) <= _MAX_KEY_FILENAME_BYTES
stem.encode("utf-8").decode("utf-8") # well-formed
def test_a_stem_listed_by_the_web_ui_deletes_the_same_file(tmp_path):
with patch('src.cache_manager.CacheManager._get_writable_cache_dir', return_value=str(tmp_path)):
manager = CacheManager()
try:
manager.save_cache(CALENDAR_KEY, {"events": []})
listed = [entry["key"] for entry in manager.list_cache_files()]
assert len(listed) == 1
manager.clear_cache(listed[0])
assert [n for n in os.listdir(tmp_path) if n.endswith(".json")] == []
finally:
manager.stop_cleanup_thread()
def test_a_failed_write_names_the_real_error(tmp_path, monkeypatch, caplog):
blocker = tmp_path / "a-file"
blocker.write_text("")
# No writable fallback either, so set() gives up and says why.
monkeypatch.setattr(os.path, "expanduser", lambda _p: str(blocker / "home"))
cache = DiskCache(str(tmp_path / "missing"))
with caplog.at_level(logging.WARNING):
cache.set("weather_current", {"t": 1})
gave_up = [r.getMessage() for r in caplog.records if "Could not write cache" in r.getMessage()]
assert len(gave_up) == 1
assert "permission denied" not in gave_up[0]
assert os.strerror(2) in gave_up[0] # ENOENT: the directory does not exist
+98
View File
@@ -0,0 +1,98 @@
"""CacheManager builds its ConfigManager on first use, not in __init__.
Every CacheManager built a ConfigManager and loaded the whole config for a
cache strategy that no longer reads it. The attribute stays public -- the
sports plugins resolve the global timezone through
``cache_manager.config_manager`` -- so it is now built on first access.
"""
from unittest.mock import MagicMock, patch
import pytest
import src.config_manager as config_manager_module
from src.cache_manager import CacheManager
@pytest.fixture
def built(monkeypatch):
"""Count ConfigManager constructions and load_config calls."""
made = []
class CountingConfigManager:
def __init__(self):
made.append(self)
self.loads = 0
def load_config(self):
self.loads += 1
return {}
monkeypatch.setattr(config_manager_module, "ConfigManager", CountingConfigManager)
return made
@pytest.fixture
def manager(tmp_path):
with patch('src.cache_manager.CacheManager._get_writable_cache_dir',
return_value=str(tmp_path)):
cm = CacheManager()
cm.stop_cleanup_thread()
return cm
def test_construction_does_not_load_the_config(built, manager):
manager.set("k", {"v": 1})
assert manager.get("k") == {"v": 1}
assert manager.get_cache_strategy("sports_live")["max_age"] > 0
assert built == []
def test_first_access_builds_and_loads_it_once(built, manager):
first = manager.config_manager
assert manager.config_manager is first
assert getattr(manager, "config_manager", None) is first
assert len(built) == 1 and first.loads == 1
def test_assignment_still_wins(built, manager):
replacement = MagicMock()
manager.config_manager = replacement
assert manager.config_manager is replacement
assert built == []
def test_an_unimportable_config_manager_is_none(manager, monkeypatch):
import builtins
real_import = builtins.__import__
def refuse(name, *args, **kwargs):
if name == "src.config_manager":
raise ImportError("no config manager here")
return real_import(name, *args, **kwargs)
monkeypatch.setattr(builtins, "__import__", refuse)
assert manager.config_manager is None
def test_a_failed_load_is_retried_on_the_next_access(manager, monkeypatch):
attempts = []
class Flaky:
def load_config(self):
attempts.append(1)
if len(attempts) == 1:
raise RuntimeError("config.json unreadable")
return {}
monkeypatch.setattr(config_manager_module, "ConfigManager", Flaky)
with pytest.raises(RuntimeError):
manager.config_manager
assert isinstance(manager.config_manager, Flaky)
assert len(attempts) == 2
def test_a_manager_made_without_init_still_answers(built):
bare = CacheManager.__new__(CacheManager)
bare.logger = MagicMock()
assert bare.config_manager is built[0]

Some files were not shown because too many files have changed in this diff Show More